Friday, May 19, 2006

New ideas about human-chimp speciation - the power of comparative genomics

There is a fascinating paper up on the journal Nature's website. (View the abstract here; you need a subscription for the whole paper. The NY Times has a nice piece about it here.) This paper is making the rounds on the blogosphere as well. (Most readers here probably know where to look, but for more info, the blogs I follow are on Science Blogs; the Panda's Thumb [the link's on the sidebar] is another one I read.)

I'll try to cover this paper in three parts:

1. I'll talk about the basic ideas and conclusions in the paper,
2. I'll go over some of the technical details in more depth, so we can see how this group did their analysis,
3. Finally, I'll talk about what I think this does and doesn't imply about evolution - Intelligent Design groups seem ready to jump all over every high-profile paper touching on evolution with one convoluted misinterpretation after another, so we need to deal with potential ways to misread this paper.

So, part one - what this paper is all about:

A group at the Broad Institute (which is part of both MIT and Harvard) has generated about 87 million bases of new gorilla genome sequence; this new sequence data has enabled them to line up large sections of great ape genomes and make an extremely detailed inventory of the DNA base differences among them. They conclude that these results suggest that human and chimp speciation was more complex than previously appreciated, and that there might have been some cross-breeding between the two lineages for some time after they diverged. (And no, this doesn't suggest that humans were having sex with chimps 5 million years ago - there were no humans and chimps then; there were sets of closely related species that might have interbred.)

What I find most exciting about this paper is its application of the power of comparative genomics to primates. Researchers have been comparing the genomes of different yeast species (we have dozens of yeast genomes now), different fly species, etc. and in the process we have learned a lot about both evolution and the basic biology of the cell. Over the last ten years, scientists have developed some powerful computational and statistical tools for doing this kind of analysis; now, we finally have enough primate genome sequence to really start applying these tools to the species we are most interested in - humans!

How do you compare genomes? Even before any genome sequences were available, people were saying that we are "98% chimpanzee." But this figure came from comparing the sequence of a limited number of genes; the catch is that different genes give different answers because genes can evolve at different rates. (In fact, different regions of a gene can evolve at different rate - sections of the gene crucial for a specific function tend to exhibit very few changes, while other regions can evolve fairly rapidly.) Now we can attempt to line up extensive regions of the genome, side by side, and count the number of differences in each region.

Because chimpanzees have their genetic material arranged somewhat differently (for example, some genes which are on chromosome 21 in humans are on chromosome 22 in chimps; you can also have extensive rearrangements within chromosomes), you can't just line up the entire genome from each species and compare them base by base. You have to choose regions that you can line up - regions that haven't undergone extensive rearrangements. In addition, it is helpful to have several species lined up together, to help resolve uncertainties about what kinds of changes took place. Thus, in order to look in detail at human-chimp differences, it is helpful to include genome sequence from gorillas, orangutans, and the much more genetically distant macaques.

In the current paper, the authors were able to line up thousands of regions from chimp, human, gorilla, macaque, and sometimes orangutan genomes, adding up to over thirty million DNA bases that they could directly compare among these species. Previous studies of great apes, according to the paper, covered only about 25 thousand bases.

Once these sequences were lined up, the authors basically counted up the number of bases where the sequences differ (after applying certain filters to eliminate sequence that could confound the analysis - for example, you have to pull out 'hypermutable' regions where the mutation rate is too high to make a valid comparison). The authors could divide the differences into categories - for example, you can have places where:

- the human genome differs from the other four genomes
- the chimp genome differs from the other four genomes
- humans and chimps are the same, but different from the other three
- humans and gorillas are the same, but different from the other three
- chimps and gorillas are the same, but different from the other three
etc...

These researchers found that the divergence (basically how many differences are in a given region) between the human and chimp genomes varies greatly across the genome, but if you take the average for any given chromosome, the divergence for that chromosome is fairly close to the average for the whole genome... but there is a big exception - the X chromosome, which showed a much lower divergence (circled in red in the figure below, from Patterson, et al. - the y-axis is relative divergence, with 1 being the genome average). In other words, human and chimp X-chromosomes are much more similar to each other than expected if the lineages leading to humans and chimps split off from each other 6-7 million years ago.




This is the most surprising finding of the study, and it is what leads the authors to suggest that interbreeding occurred between the chimp and human lineages for some time after initially diverging from each other. They suggest that there was an initial split 6-7 million years ago, roughly in line with the fossil record, but that later (less that 6.3 million years ago) there was hybridization between the two lineages resulting in some gene exchange. This paper is still new, so reactions to this scenario are just starting to trickle in, but this hypothesis is the most controversial part of the paper.

It's important to note what is not controversial though: the genomic analysis is fairly standard. The authors have used well-established statistical and computational tools to compare these genomes; the high similarity between human and chimp X chromosomes is real. These genetic analysis techniques are solid, even though some old-school anthropologists and paleontologists still resist them. The challenge now is to decide what this low X divergence is saying about how chimp-human speciation occurred, and that is where the most controversy about this paper will be.

And this is basically what scientists were saying in the NY Times article:

"David Page, a human geneticist at the Whitehead Institute in Cambridge, said the design of the new analysis was "really beautiful, with all the pieces of the puzzle laid out." Whether the hybridization will turn out to be the right solution to the puzzle remains to be seen, "but for the moment I can't think of a better explanation," he said."

In the next post, I'll go into some more technical detail about how these authors did their analysis, as an example of how useful it is to be able to compare multiple genomes.

Monday, May 01, 2006

NY Times distorts peer review

The NY Times has an odd article called "For Science's Gatekeepers, a Credibility Gap". For starters, this article (by a physician named Lawrence Altman) is not that well written. There is very little serious reporting in the article; instead, it's filled with general sentences like this one:

"But many authors have still withheld information for fear that journals would pull their papers for an infraction. Increasingly, journals and authors' institutions also send out news releases ahead of time about a peer-reviewed discovery so that reports from news organizations coincide with a journal's date of issue."

Normally, at this point, a genuine reporter would then cite sources who were interviewed for the article, or go into more detail. Instead, this piece just continues on with more of the same:

"A barrage of news reports can follow. But often the news release is sent without the full paper, so reports may be based only on the spin created by a journal or an institution."

Again, no sources or specific examples are cited.

Besides the fact that this article has no serious reporting, it completely mischaracterizes the peer-review system that is the standard for publishing serious research. Here are some gems from the article:

"The publication process is complex. Many factors can allow error, even fraud, to slip through. They include economic pressures for journals to avoid investigating suspected errors; the desire to avoid displeasing the authors and the experts who review manuscripts; and the fear that angry scientists will withhold the manuscripts that are the lifeline of the journals, putting them out of business. By promoting the sanctity of peer review and using it to justify a number of their actions in recent years, journals have added to their enormous power."


Huh??? "They fear that angry scientists will withold the manuscripts..." I personally know prominent scientists who frequently have papers rejected by top journals; I also personally know editors of journals. The idea that editors accept fraudulent science because they are afraid of angering scientists who will never send them a manuscript again is just total crap. Rejection is a normal part of life for all scientists. Nobody gets all of their papers into Science or JAMA, and fear of authors' anger is not really a major factor in editors' decisions.


Is peer review just part of a cynical game on the part of journals to add 'to their enormous power?" This article tries to make you think so:

"The release of news about scientific and medical findings is among the most tightly managed in country. Journals control when the public learns about findings from taxpayer-supported research by setting dates when the research can be published."

This is absurd - journals are trying to exert tight control over scientific information "by setting dates when the research can be published"???? What happens in reality is this: authors submit a paper, and when it's accepted, it enters the publication pipeline and gets published on a routine timescale. Sometimes a journal will receive two closely related papers in a short period of time; the editors will try to publish both articles back to back in the same issue. In other cases, the findings of a particular study are considered unusually significant, and the manuscript will go through an accelerated publication process (after the paper makes it through peer-review). While a manuscript is under review or awaiting publication, authors frequently present their findings in seminars and conferences, which are generally open to any taxpayer or journalist who cares to attend and listen to a highly technical talk. There is no conspiracy to control what the public learns.

And then there's the paragraph that I quoted before: [I'm not sure what the logical connection is between these two sentences in this paragraph]

"But many authors have still withheld information for fear that journals would pull their papers for an infraction. Increasingly, journals and authors' institutions also send out news releases ahead of time about a peer-reviewed discovery so that reports from news organizations coincide with a journal's date of issue."

It's hard to resist the temptation to just quote bad paragraph after bad paragraph (just go read the article yourself.) It's just so odd that this was considered fit to print by the NY Times.

Here is where I think Altman really goes wrong (other than his attempt to ascribe peer-review's power to a conspiracy of power-hungry journal editors):

Peer-review isn't really set up to catch fraud. And this is not because editors are trying to avoid offending authors, or increase their power over the scientific and medical information available to the public. The real reason is this: peer-reviewers check a manuscript for flaws in reasoning and methodology (did the authors leave out any crucial experiments, was the experimental design subject to unacknowledged error, or did they misinterpret their results?), and they evaluate a paper's scientific merit (they ask, is this a significant achievement in the field?). That's basically it. If authors have falsified their data, reviewers may very well not catch it. Peer-review is still based on trust.

Is scientific fraud a growing problem? Only because the population of scientists is growing. If you have more practicing scientists, you're going to have more fraud. And fraud is caught (always by other scientists) and punished severely. Unlike in politics, business, and even law and medicine, fraud is a complete career-ender in science. Scientists guilty of serious fraud cannot get grants, they lose their academic positions, and can never publish in a reputable journal again. They can no longer be scientists. Fraud is punished more harshly in science than in almost any other profession.


Maybe the NY Times can use a little peer-review for some of their pieces.

Wednesday, April 12, 2006

Slowly getting back in business

I've finally managed to blog about science again. I'm not sure if I'll get around to the dog genome paper, but there are plenty of other fascinating things to write about - stay tuned.

This is the reason it's been hard to make entries:

Adaptive evolution to a new ecological niche: The GAL pathway in yeast

Yeast is a fantastic model organism for all sorts of biological studies, ranging from the cell division cycle, to protein synthesis, to gene regulation. With the advent of fast, relatively cheap genome sequencing methods, yeast has become an excellent model system for studying evolution. Different species of budding yeast alone represent more than 300 million years of evolutionary innovation since their divergence from non-budding yeasts. To put this in perspective, the evolution of budding yeasts has continued uninterrupted since before the earliest dinosaurs appeared. Although during this period all these yeast lineages remained unicellular fungal organisms, a huge amount of genome evolution took place as various lines adapted to different environments.

We currently have genome sequences for over two dozen yeast species, which occupy a wide range of ecological niches. By comparing these genomes we can learn about the basic evolutionary processes at the level of genes. (This paper is a nice but technical introduction to the subject; the Genolevures website is one of the best online sources of information about yeast evolution.)

I came across an interesting paper about budding yeast that discusses a specific instance of adaptive evolution in stunning detail - the alteration of galactose utilization pathway in several yeast species that have adapted to different environments. As the authors of the paper state,

"We have found that at least three independent lineages of yeast have inactivated or lost most or all of the genes of the GAL pathway while leaving interacting genes intact. The parallel losses of this entire pathway are extreme examples of an emerging general principle that gene-inactivation events reveal specific functions affected by recent changes in the ecological pressures acting on a species."

Let's back up and explain what the galactose utilization pathway is: Yeast metabolize the sugar galactose (see the picture below) through several steps mediated by specific enyzmes that convert galactose into a modified form of glucose, which is then further metabolized to generate high energy molecules and byproducts like alcohol or lactic acid. This is a well characterized pathway, meaning that the specific proteins involved and their functions are known, as are the modifcations made to the galactose molecule. In rough outline, here is the galactose utilization pathway: The protein Gal2p allows galactose into the cell, where it is transformed in sequence by the proteins Gal1p, Gal7p, and Gal5p. By this point the galactose molecule has been changed around to become a modified form of glucose, which can then be further metabolized. This pathway is tightly regulated by a set of control proteins (such as Gal4p), which act to turn the synthesis of the pathway proteins on or off.



A yeast species that lives in an environment where galactose is a major food source needs a functioning galacose utilization pathway. Species that adapt to a difference niche, where galactose utilization is not essential, would be expected to lose the genes involved in this pathway, because those genes would no longer be preserved by natural selection.

Now, back to the paper:

The authors looked for the GAL genes in the genomes of four yeast species that cannot metabolize galactose, and found that in three of these species, all of the dedicated GAL genes were completely missing, with on exception:

The yeast species E. gossypii harbors a form of the GAL4 gene, but the DNA sequence of this gene suggests that it has undergone rapid evolutionary change, possibly an indication that this GAL4 gene has been recruited for a new function in E. gossypii. This is an excellent hypothesis begging to be tested - one could delete this gene in E. gossypi and look at the consequences.

[Note that I'm skipping over a lot of technical detail - determining that this particular gene has undergone rapid DNA sequence change involves some statistical testing - you can find more details here. The point is, when we say that a gene has undergone rapid evolutionary change, it involves more than just looking at a genes sequence and saying "yep, looks like evolution to me."]

These three yeast species which were missing any trace of most of the GAL genes were species that broke off early from the evolutionary branch leading to baker's yeast. They lost their ability to use galactose a long time ago, and the relentless erosion by mutation in the absence of selection has erased the GAL genes from their genomes. (See the figure below - the species that can't use galactose are listed in red, and the red vertical bars represent points in time where each lineage lost the ability to use galactose)



The species S. kudriavzevii however, is much more closely related to baker's yeast. It cannot use galactose, even though all of its close species cousins still can. This indicates that it lost the ability to use galactose more recently. As a result, the non-functional remnants of the GAL genes have not yet been erased from its genome - they are all still present as pseudogenes. (Pseudoegenes are genes that look like protein-coding genes, except they have been rendered non-functional by some inactivating mutation.)

S. kudriavzevii appears to have recently adapted to a new ecological niche - unlike other close relatives of baker's yeast, S. kudriavzevii is not found in sugar-rich environments; it was isolated from soil and decaying leaves. As a result of this change in lifestyle, there was no longer any selective pressure to maintain the GAL genes, thus they are slowly being erased by mutation.

To sum up:

1) 4 yeast species among the set of 11 studied cannot use galactose for metabolic energy. Each of these 4 species has lost its functional GAL genes, while species that can use galactose have retained their GAL genes.

2) Species that lost the ability to use galactose a long time ago have had nearly all traces of their GAL genes erased by the relentless force of mutation over time (we're on a scale of 100 million years or more). In one species, the GAL4 gene appears to have been recruited for a new (as yet unknown) function.

3) One species lost its ability to use galactose more recently (the last few million years); in this species non-functional GAL genes are still present as pseudogenes.

4) This is a fascinating example of gene inactivation resulting from the relaxation of the force of natural selection that occurs as organisms adapt to a new environment.

The real bottom line: What we find when we look at genomes is absolutely unexplainable unless we accept the bedrock tenets of evolution: 1) today's diverse life forms diverged from common ancestors, and 2) mutation and natural selection are major driving forces of evolutionary change.

Tuesday, March 07, 2006

Back in business soon...

This blog break has been a little longer than I anticipated, but I plan on getting back to some more science posts this week. Our twins were born January 8, I turned in the final, corrected copy of my doctoral thesis on January 16th (my defense was in December), we packed up our stuff and left NY with a Penske truck on January 26th, and finally arrived at our new home in St. Louis on January 28th. I started my postdoc position at the Center for Genome Sciences at the Wash. U. School of Medicine on Feb. 1.

Now that our family's recovered a bit from that last insane month, it's time to get back to some of the interesting things that are going on in biology. The next installment on the dog genome should be up by Thursday.