Showing posts with label species. Show all posts
Showing posts with label species. Show all posts

Friday, April 13, 2018

Monophyletic species, kind of

A paper by bryologist Brent Mishler and philosopher of biology John Wilkins has just come out, with the title The Hunting of the SnaRC: A Snarky Solution to the Species Problem. It is open access in the journal Philosophy Theory and Practice in Biology, so anybody with internet access can check it out.

Many bloggers have issues that they return to again and again even if they are not necessarily the nominal topics of their blogs - for example, Jerry Coyne frequently posts about Free Will and about students trying to shut down talks by speakers they don't like, and Larry Moran regularly takes apart papers claiming that junk DNA has been disproved. This much less widely known blogger can reliably be coaxed out from behind the oven by at least two such recurring issues: bad arguments for the acceptance of paraphyletic taxa, and the in my eyes incoherent concept of "monophyletic species".

As the title indicates, Mishler & Wilkins present a solution for the species problem, i.e. the perennial question in biology of what 'a species' even is. Especially as the paper is freely accessible it would serve no purpose to summarise its introduction, so I will move immediately to what I find most interesting: their views on how to view species and some pointers on how to do classification at the lowest levels in practice.

Note that I say "their views", plural, deliberately, because this is one aspect of the paper that I have not quite understood yet:

Wilkins has argued in the past that the popular approach of developing a theoretical species concept and then applying it to a potentially recalcitrant reality is a dead end. What biologists should do is the opposite, i.e. consider species as empirical phenomena in need of individual explanations. And here in this paper, Wilkins' argument is reiterated concisely in section 3, A Way Forward: Species Are at Least Initially Phenomena.

What I like about this flip in perspective is that it allows much more flexibility; obviously the empirical phenomena that we generally identify as species, be it popularly or as biologists - generally gaps in morphological or genetic variation - need a different scientific explanation for example in asexual than in sexual species, making one-size-fits-all species concepts difficult to apply.

Mishler, in turn, has argued in the past that species are not a special biological category different from e.g. monophyletic genera and families. The species category is arbitrary, and we should just classify all organisms into nested monophyletic groups, AKA clades, all the way down to the individual specimens. And here in this paper, Mishler's argument is reiterated in sections 4, Rankless Taxonomy, 5, Capturing the SNaRC, and 6, Using SNaRCs in Systematic, Evolutionary, and Ecological Studies.

The thing is, while there is perhaps technically no direct contradiction between those two arguments to the degree that there is a contradiction between "all taxa should be monophyletic" and "taxa should be allowed to be paraphyletic", they appear to be two rather different prescriptions. If I understand correctly, the first says,
  • We should treat species as empirical phenomena in need of explanation instead of indiscriminately applying a given theoretical concept to them.
The second says,
  • It makes no sense to even talk of species, we should stop doing so, and here is a single theoretical concept (everything is clades) that we should indiscriminately apply to all specimens.
In fact I am currently unable to see how sections 4-6 and the conclusions of this paper would have to change if section 3 were to be deleted in its entirety. What am I missing?

What I found most useful about this paper was that it has some thoughts on how to do classification into nested clades all the way down to the individual specimens in practice, because that was completely unclear to me in all past instances when this approach was suggested. There are some apparent problems with it, particularly that we need items forming a tree structure to even have clades. It is sometimes difficult to illustrate the issue, but it can perhaps be presented as follows:
  1. The prescription is, as mentioned above, that a classification should be clades (= monophyletic groups) all the way down to individual specimens.
  2. A clade is a complete branch in a tree structure, and usually understood to be specifically a complete branch of a species phylogeny.
  3. In other words, the way the term clade is defined, it applies only in a tree-structure but is inapplicable in a net-like structure.
  4. Sexually reproducing species are systems consisting of individual specimens that have net-like relationships with each other, because they share numerous ancestors instead of one ancestor in each sufficiently earlier generation.
  5. It follows necessarily from the previous two points that the term clade cannot be applied to describe the relationship between specimens if what we are looking at includes multiple specimens from the same sexually reproducing species.
  6. If follows then that it is logically impossible to classify into clades all the way down to these specimens, unless the meaning of the word clade is changed to a degree that the whole purpose of having that word is defeated.
To my understanding this is why Hennig spent so much time discussing the different ways that specimens (or snapshots of them, which he called semaphoronts) can be related to each other. The relationship between four (non-hybridogenic) species is tree-like, so they can, and should, be classified into clades. But relationships between individuals within a sexually reproducing species are net-like, so they cannot possibly be classified into clades, as the word does not even have a meaning in that structure.

The point at which approaches to classification change is approximately at the species level. Phylogenetic systematics applies only above it, and it uses species as the units that it groups into clades, because if it used any smaller units there would not be clades. This is also why in my opinion one cannot coherently reject the reality of species and be a phylogenetic systematist and, conversely, coherently accept the reality of species and promote paraphyletic taxa, because clades are species that have diversified. Many others, of course, disagree.

Now, what is the practical approach suggested by the present paper? It argues that the terminal units of classification should be "the finest-scale clades that can be convincingly demonstrated with current data", here called Smallest Named and Registered Clades (SNaRCs). Obviously such a 'clade' cannot be based on information from a single gene, as it may show a different history than other genes, for example because of introgression or incomplete lineage sorting. The solution is to use as evidence for monophyly "the preponderance of gene lineages making up a clade", or in other words "congruence among the majority of gene trees and other types of phylogenetic characters available".

On the plus side, this is a very empirical and testable prescription. But consider two thought experiments. First, take three samples A, B and C, look at, say, 100 gene trees, and if 51 of them show ((A,B),C) then A and B form a 'clade', even if all three of them are members of the same sexually reproducing species. Again, that is doable, empirical and testable, and we get a clear answer.

Nonetheless this approach does not convince me at the moment, nor will it even if we assume a scenario of 100 gene trees supporting (A,B), simply because no matter what the gene trees say, in reality there is no tree-structure inside the species. Yes, we can easily sequence for example the DNA of three siblings and run an analysis that will produce a phylogenetic tree for each gene, but in reality these three people just don't have a tree-relationship with each other, so it does not make sense to me to use terminology or a classification that implies there is one.

For the second thought experiment, take three samples D, E, and F, and if 33 gene trees say ((D,E),F), 33 say (D,(E,F)), and 34 say (E,(D,F)), we are inside a SNaRC and should not delimit any more narrowly, even if D is a specimen from an arid zone ephemeral, E from an alpine perennial, and F from a narrow endemic of the northwestern Blue Mountains that only occurs on ironstone-sandstone outcrops, and all three of them are geographically isolated from each other.

This hypothetical case has three very distinct entities that show a lot of gene tree discordance for the genes we used for our analysis. This is a much weaker problem than the previous one because Mishler & Wilkins argue that SNaRCs are, as all scientific hypotheses, tentative and await revision after the examination of more data. Maybe the next 100 gene trees will clinch it for (A,(B,C)), and then at least we could separate out A; more realistically, sampling more individuals of all three species will presumably resolve the three species as three SNaRCs, even if we cannot figure out the relationship of those three SNaRCs with each other (they may even form a true polytomy, and that's fine).

Still it bothers me that in a situation where we unfortunately have only one sample per species available for analysis the approach promoted in the present paper might lead to the tentative lumping of clearly distinct entities. And unless something is added to the approach, or unless I am missing something, it would have to, because it does not seem to include a way of recognising single-specimen SNaRCs except in the case of one being left alone as sister to another SNaRCs, that, in turn, would still consist of two potentially vastly different specimens. But maybe I am taking this too literally.

On top of that there is perhaps another methodological issue, or again maybe just something I don't understand. It seems to me as if "majority vote of the gene trees" is not actually how multi-locus phylogenetic analyses generally work. To the best of my understanding they reconcile gene trees in rather more complex ways, even in the case of such a simple approach as Gene Tree Parsimony, let alone the multi-gene coalescent model. Many of these approaches actually presuppose the existence of species or populations, and for the same reason as I argued above: what happens within a sexually reproducing lineage is rather different from what happens between such lineages.

More than anything what I find uncomfortable about the approach presented here is that it seems to care not so much about the actual patterns of common descent of what it classifies as about character or gene tree distribution. The difference may come across as subtle, admittedly. What I am trying to say is that I believe phylogenetic systematics should be about classifying organisms by relatedness, by exclusivity of common descent.

I do not, for example, care very much about the fact that most of the ancestral chloroplast genome has been moved over into the nucleus of the host cell, because the chloroplasts are directly descended in an unbroken line from the first cyanobacterium that colonised a plant cell, and the plant species we have today are descended in an unbroken line from that plant cell. To me chloroplasts are a subclade of cyanobacteria and plants are a subclade of eucaryotes, all regardless of what happened to the individual genes.

To use an example from within a species, I have mentioned in the past that it is possible, although statistically unlikely, that I have inherited no genetic material whatsoever from my maternal grandfather, if it just so happened that all the chromosomes my mother gave me were those she got from her mother (the Y chromosome is of course always from the paternal grandfather, by necessity). But even if that were the case we would nonetheless consider it to be an important piece of information that I descended from my maternal grandfather, and I would nonetheless not exist without his involvement. So yes, we use the genes to infer common descent, but the point is really the common descent itself, and the genes are just a data source that can potentially mislead us. Sometimes the right answer may be (A,(B,C)) even if most genes say ((A,B),C).

The "majority vote of the gene trees" approach, however, feels as if its practical concern starts and ends at the pattern shown by the genes, regardless of what the patterns of descent are. To me that feels the wrong way around.

Another way of looking at the issue may be this: If we truly accept the argument made in section 3, that we should look at natural phenomena, consider them to be explananda, and find the most appropriate scientific explanation for each of them, would the logical result not be Hennig's original approach? The phenomenon that a beetle specimen shares more traits with a bee specimen than either share with a slug specimen has an explanation, and that is that the former two share a much more recent common ancestor from which they inherited the shared traits. We express that reality by grouping the former two into a taxon called 'insects' while leaving the slug out.

The fact that I may easily in some cases share more genetic similarity with somebody born in Italy than with another northern German, however, would most likely be due to the stochastic nature of allele inheritance inside our sexually reproducing species. There is no clade wherein two specimens of humanity - the hypothetical Italian and I - share one and only one most recent common ancestor. Instead, beyond some point in the past we share thousands of ancestral 'specimens' in each generation. Because this is a different biological phenomenon than ((beetle,bee),slug), we need a different approach to classification at that level.

Thursday, March 23, 2017

Species delimitation using the coalescent model

For two weeks or so now a new paper has been making the rounds, and we discussed it in our journal club today:
Sukumaran J, Knowles LL, 2017. Multispecies coalescent delimits structure, not species. PNAS 7: 1607-1612.
The context is species delimitation: given a bunch of individuals, how many species are there, and which individuals belong to which species? There are a number of ways to address these questions, and they partly depend on the available data and technology and partly on the species concept the researcher is using.

Very traditionally, of course, a taxonomist would look at the morphology of the specimens and more or less intuitively try to form clusters of similar specimens separated from each other by gaps in morphological variation. In other words, a qualitative application of the Genotypic Cluster Species Concept. More dubious approaches would involve ideal "types" (in a Platonic sense), "central identities", or rules of thumb on the lines of "one difference means subspecies, two differences means species", none of which seem to have much basis in what we know about genetics or evolutionary biology.

More formally, one can take the same theoretical approach but conduct an explicit, quantitative analysis. Score the morphological data and produce a pair-wise distance matrix for example with the Gower metric, then do a Principal Coordinates Analysis to visualise potential clusters and gaps between them, or do hierarchical or non-hierarchical clustering. The same can be done with non-morphological data, such as environmental data from the collecting localities, in that case to show that putative species have different ecological niches.

A clustering approach can also, of course, be used for genetic data. In that case one would use some kind of genotyping approach, for example microsatellites, AFLP or genome-wide SNPs, and do hierarchical clustering or use a software such as STRUCTURE. Although using a population genetics model, the results produced by the latter are at a practical level comparable to the non-hierarchical clustering in that we get an optimal number of clusters and information on what sample belongs to what cluster; we then need to make the additional interpretative step of assuming that the clusters are the species. (Meaning we have solved the grouping problem but need additional arguments to solve the ranking problem.)

But today more and more people have multi-locus sequence data at their disposal. They are used for phylogenetics under the coalescent model and using species tree approaches, so it was probably unavoidable that the coalescent model would also be applied to species delimitation. The idea behind the relevant software tools such as the currently very popular BPP (disclosure: I have never used it) is that the information from multiple loci can be used to figure out how many species there are among the samples, under the assumption that samples belonging to the same species should have a history of reticulation but samples belonging to different species should have a history of (permanent) lineage divergence.

That sounds logical, but the aforementioned paper seems to hit this idea under the waterline: as the title suggests, the authors conclude that species delimitation under the coalescent resolves population structure, not species limits.

Frankly, although the method has been extremely popular lately, there has also been a lot of scepticism in the community. After all, its application has produced rather one-sided results, nearly always splitting species into several smaller species. I have heard a talk that amounted to a scathing criticism of the approach, arguing that genetic isolation of a small population for less than 200 years would be enough to make it show up as a separate "species" under the coalescent, surely a ridiculous outcome.

Consequently, the present paper fits my thinking on the issue; I, personally, would rather use clustering approaches to search for gaps in variation. But that being said, the way the authors addressed the issue still seems a bit odd to me and leaves me wondering how far their particular argument will carry.

The thing is, the study does not involve any empirical data, it is entirely based on simulations. The authors used a model under which at first only populations split and then some of them may turn into separate species after varying lag times; although there does not appear to be an explicit process in the model I guess the assumption is that it needs a bit of time to accumulate enough differences that a population cannot reunite with its sister population even if they get back into contact with each other. They then simulated species lineages under that model, and then gene trees in those species lineages, and then sequence matrices for those gene trees. And then they analysed the sequence matrices with the coalescent-based species delimitation approach trying to get the original species back.

Surprise, surprise, the coalescent species delimitation approach recovered the population splits, not the species splits. But what has this really shown? As far as I can tell, it has shown that an approach using a model counting all population splits immediately as species splits will not produce the results expected under a model not counting all population splits immediately as species splits.

Maybe I am missing something, but that is exactly what I would have expected before complex simulations on supercomputers had been conducted. If I simulate bicycle rides under a model that assumes I cycle to work at 20 km/h and then try to fit the results back to a model that assumes I cycle to work at 100 km/h I will also likely find that there is poor fit, right? But that does not tell me anything about how fast I really cycle to work, or in other words, anything about which of the two models is a better fit to reality.

Consequently I have to admit that arguments on the lines of "this real-life population that is clearly not a separate species but has merely been isolated for 200 years comes out as a new species under the coalescent approach" seem to be more impressive.

Friday, January 6, 2017

Ochlospecies

Happy new year, everybody! Although many people are variously fed up with 2016 for the supposedly high number of celebrity deaths, Brexit and the US election, for me personally the past year was very enjoyable and successful, so I cannot really complain. Let's hope that 2017 will at least be better than so many of us expect.

Anyway, quite some time ago I wrote a series of posts on species concepts. Because a reviewer mentioned it, I had now reason to look into a concept that I was not familiar with, that of ochlospecies.

It was apparently developed by a tropical tree taxonomist called White in the 1960s, but the most useful source that is available online appears to be a 1998 review by Q.C.B. Cronk.

The idea is that there are different kinds of species that can be distinguished based on their patterns of variation. There are well-behaved, very distinctive species. There are species that are variable but in a way that is easily understood, for example because variation is nicely hierarchical, because several characters show correlation, or because there is a clear geographic pattern. And then there are the bad apples, species that show variation but without any clear structure. Several crucial characters appear to vary independently, no geographic pattern, just a mess. Those are then the ochlospecies.

What do I take from this?

Well, first to me this is a rather unexpected way of going about species concepts. The whole point, as far as I am concerned, is that before you embark on a taxonomic study you should set up clear criteria of how you will evaluate the evidence that you are going to collect.

That makes it science. So if you are dealing with the circumscription of genera, instead of making it up as you go you say in advance: I will know that a genus is acceptable if it turns out to be monophyletic. And if you are dealing with the circumscription of species, again, instead of making it up as you go you say in advance: In my study I will follow the Genotypic Cluster Species Concept, meaning I will circumscribe species based on gaps in morphological and/or genetic variation. Or whatever other species concept works for the group and the data that you can actually collect.

When I look at the ochlospecies, however, it looks to me as if it is not a criterion but a conclusion. Reading the Cronk review, it appears as if the taxonomist circumscribes species somehow (magical asterisk) and then labels some of the resulting species as ochlospecies after the fact. At a minimum that means that the concept has a different utility, if any, than the species concepts I have considered previously, than that of serving as a potential guideline.

So at the moment I am a bit at a loss as to what to do with the concept regarding my paper. As the concept is not directly useful methodologically, my options seems to reduce to name-checking it either in introduction or in discussion.

Sunday, August 21, 2016

Monophyletic species yet again: a recent example

Recently I received a publication alert for Ja soon to be published manuscript. It gave me reason once more to write about the issue of "monophyletic" species.

I do not want to give the impression I am deliberately picking on this particular paper. On the one hand, for all I know its data are completely awesome and its conclusions are valid; the issue I am writing about here is somewhat tangential to the paper anyway, its main focus being on genus level phylogeny. On the other hand, this same issue can be seen in many, many other papers in the field, as all too many systematics lectures at universities seem to leave it at "stuff must be monophyletic", without explaining the relevant background like the various relationships that OTUs can have to each other and, crucially, that different classification approaches apply to different relationships. So the occasion here is really only that the present paper showcases the issue in an extremely compact format, all condensed down to a mere three sentences in the discussion.

In full they run as follows:
However, while cladists debate whether higher level taxonomic groups should be monophyletic (e.g. Horandl and Stuessy, 2010; Schmidt-Lebuhn, 2012), it is conceivable that species need not be monophyletic, as different modes of speciation may have different phylogenetic outcomes (Rieseberg and Brouillet, 1994); non-monophyly is an expected intermediate state as taxa diverge (Avise and Ball, 1990). Indeed, in a morphology - based survey of 206 Australian plant species and subspecies (Proteaceae and Fabaceae), Crisp and Chandler (1996) estimated that 21% were paraphyletic. In addition, eucalypt taxonomists generally follow the ecological species concept that allows for hybridisation between taxa (Johnson, 1976), and such reticulation can cause non-monophyly and incongruence between morphological and genetic markers (e.g. Rutherford et al., 2016).
So what do I find problematic about these three sentences? Going through again in order...
However, while cladists debate whether higher level taxonomic groups should be monophyletic (e.g. Horandl and Stuessy, 2010; Schmidt-Lebuhn, 2012),
First, Hoerandl and Stuessy are not cladists. Second, of course cladists do not debate if supraspecific taxa should be monophyletic, because a cladist is defined as somebody who has already decided that supraspecific taxa should be monophyletic. If you still debate it you are by definition Not A Cladist. Compare "vegetarians debate whether they should stop eating meat".  If they are still pondering that question they are Not Vegetarians.
it is conceivable that species need not be monophyletic,
This is where we get to the real issue: monophyly of species. Admittedly we first have to ask, are we talking about sexually reproducing species here? But I assume we are, because the article is about eucalypts, and they are mostly sexual. And the thing is, if that is the case then the concept of monophyly just simply does not apply. It is a word that describes a group of terminals on a rooted tree graph, like so:



A good analogy to describe what is going on is this. Imagine you have a real tree in front of you, and you are taking a group of twigs off using secateurs. To get a monophyletic group, you need to cut exactly once. To get a non-monophyletic group, you need to cut more than once. (If after cutting off a non-monophyletic group you have kept only one piece in your hand, it is paraphyletic; if you have kept several pieces but thrown out what used to connect them, it is polyphyletic. But that just as an aside.)

Within a sexually reproducing species, AKA a breeding group, there is no tree structure but a network structure, as individuals have numerous ancestors in each generation as opposed to one. That looks like this:


How do you get a monophyletic group? Well, you cannot, it is impossible. You could argue that my secateurs analogy would work for paraphyly even in a network, but only by jumping over the "imagine you have a real tree in front of you" part. You don't have a tree, you have a fishnet.

So to me this fragment - and everything that follows - makes as much sense as "it is conceivable that songs need not be yellow with purple stripes". Of course they don't - Amy Winehouse's Rehab, for example, is not yellow with purple stripes and yet it is a perfectly acceptable song. But then again I have no idea how it could be yellow with purple stripes, even if one were to try and make it so.
as different modes of speciation may have different phylogenetic outcomes (Rieseberg and Brouillet, 1994); non-monophyly is an expected intermediate state as taxa diverge (Avise and Ball, 1990).
Opponents of Phylogenetic Systematics regularly make the argument that "non-monophyly is an expected intermediate state as taxa diverge" at all taxonomic levels. At the supraspecific level, this constitutes a rather clear example of circular reasoning. A cladist would argue that a subclade cannot diverge from the larger clade it is part of, ever, because it is by definition part of that clade. That is what the sub- part of subclade means.

At the species level, on the other hand, the above fragment makes sense if we think of incomplete lineage sorting. Barring recombination, the copies or alleles of an individual gene do indeed evolve in a tree-like fashion, and the alleles found in one species will at first generally be paraphyletic to the copies found in its sister species. Only over time will selection or even loss through purely stochastic processes (genetic drift) make the alleles from each species monophyletic on the gene tree, a process known as lineage sorting.

It is possible that this is what the authors are referring to. But to me it still does not mean that it makes sense to call a species paraphyletic, because the components of a species are not gene copies but individuals, and individuals of the same sexually reproducing species stand in a network-relationship to each other, so that the word paraphyletic does not apply.
Indeed, in a morphology - based survey of 206 Australian plant species and subspecies (Proteaceae and Fabaceae), Crisp and Chandler (1996) estimated that 21% were paraphyletic.
Although I would not use the terms as they did (see above), the conclusions of the Crisp & Chandler paper are completely in accord with what I am saying here: species are special, because they are the level at which and from which on downwards it does not make sense any more to try and make stuff monophyletic. That being said, however, the paper does show species as paraphyletic on several morphology-based trees. How did it arrive at that result? Or in other words, given what I wrote earlier, what is the difference in perspective?

First, the terminals on the trees in Crisp & Chandler, the OTUs, are not actually individuals but groups of individuals, such as populations or subspecies; second, the authors conducted phylogenetic analyses on these OTUs. What that means is that the OTUs are forced into a tree-relationship even if the true relationship is net-like, because that is what phylogenetic analyses do. But if we are really talking about structures within a breeding group, within a sexually reproducing species, then in my eyes that analysis was just not appropriate because yes, the true relationship is net-like instead of tree-like. (And if the OTUs are not in a net-like, reticulating relationship, but instead genetically isolated, separate evolutionary linages, then why aren't they recognised as species?)

For example, I can jot down some morphological traits of a bunch of fellow humans, make a data matrix, and do a phylogenetic analysis. Because I do a phylogenetic analysis, the analysis will invariably return a tree. But does that mean that each of my OTUs - individual humans - had only a single parent, and only a single grandparent? Of course not, because we humans do not have a branching, tree-like relationship to each other either. The analysis simply made assumptions that do not hold up against reality.
In addition, eucalypt taxonomists generally follow the ecological species concept that allows for hybridisation between taxa (Johnson, 1976), and such reticulation can cause non-monophyly
Unfortunately it is left unclear what items are forming a non-monophyletic group in those situations. If it is alleles, see a few paragraphs further up; if it is individuals, see the immediately preceding section.
and incongruence between morphological and genetic markers (e.g. Rutherford et al., 2016).
That is true, but we could also mention several other processes, like aforementioned incomplete lineage sorting, meaning that we can have such incongruence even in the complete absence of hybridisation.

To close this post I would like to present a little paragraph that shows how the three sentences discussed above read to me:
However, while opponents of Scottish independence debate whether Scotland should be independent, it is conceivable that citizens need not be independent nations as different ways of acquiring citizenship may have different political outcomes; not being a geographic entity is an expected intermediate state as nations become independent countries. Indeed, in a survey of 206 individual citizens, Doe & Average (2010) estimated that 21% of them were not independent nations. In addition, political scientists usually follow a concept of citizenship that allows dual citizenship, and such reticulation can cause nations not being independent from other nations and incongruence between native language and nationality.
Again, this could be part of a great article on the Scottish independence movement, just like the present paper presents interesting genomic data on its study genus. But does this read as if the author was just a tiny bit confused about the difference between nations and the citizens that nations consist of? Quite so.

A lot of unproductive controversy and confusion among systematists and evolutionary biologists could be avoided if it became a bit more widely known what even ur-cladist Willi Hennig himself had in mind when he came up with the idea of "making stuff monophyletic". He was only arguing that supraspecific taxa should be monophyletic groups of species; the concept of species being monophyletic groups of individuals would not have made any sense to him, as he was very clear on the difference between usually tree-like (phylogenetic) relationships between species and net-like relationships within them.

Reference

Crisp MD, Chandler GT, 1996. Paraphyletic species. Telopea 6(4): 813–844.

Wednesday, May 25, 2016

Species delimitation once more

Today I attended a kind of mini-symposium featuring five short talks from an ecology focused university department. The main topic was the use of molecular data in ecology, and I learned a lot of fascinating things especially about recovery of animal populations after bush fires and about Antarctic ice age refugia.

The discussion, however, turned very strongly towards species delimitation, and this is where things got a bit weird. It was remarkable how self-confidently some, say, colleagues specialising on the behaviour of a few selected fluffy mammals pronounce some black and white opinion on a problem that systematists have been struggling with for centuries.

Let's just say that if the situation were reversed - perhaps me being asked about fire ecology after giving a talk on plant systematics - I hope I would be able to stick to something on the lines of "well, this is what I think, but really you should ask an ecologist". I sure hope I would not simply say "based on what I observed in one plant species, fire has no effects whatsoever on the genetic diversity of animals". Some of the ecologists here had no such reservations.

At one end of the spectrum several audience members appeared to expect that now that we have genomic data we should finally figure out an objective cut-off for how much genetic difference between two individuals puts them into different species. This is one of those ideas that look superficially attractive but reveal new layers of wrongness the longer one thinks about them.

The easiest answer was provided by one of the speakers: Difference in what molecular marker(s)? Different genes or regions evolve at very different speeds. Taking the whole genome also seems a bit suspect as most of it is junk DNA, the exact amount varying by species. But again, there are layers. The speed of molecular evolution would also differ vastly between different lineages depending on their generation times, mode of reproduction, perhaps other life history traits, and the environment they find themselves in. The genetic diversity within populations is also vastly different from species to species.

But what mostly blows the idea out of the water is quite simply that there is necessarily a grey zone between being one species and having split into two species. ANY cut-off would have to be entirely arbitrary.

That brings me to the other end of the spectrum. After the talks, one of the speakers declared over snacks all of the following: (a) we should use the Biological Species Concept, (b) species are arbitrary human constructs and have no empirical reality, and (c) genetic data will always give you two clear groups, only we must be careful how we interpret that.

My first observation is that it should not really be possible to believe these things at the same time, because (a) and (c) are in direct contradiction to (b).

The second is that I consider all three to be wrong. It is clear that the BSC simply does not apply in many cases, especially asexual species and fossils.

As for the arbitrariness of species, well, it depends. (This is also the only answer to the species problem that I find useful.) I wrote above that there is a grey area, that a cut-off for "moment of speciation" is arbitrary. That being said, the same is true for many other things that we happily classify. We can probably all agree that the cut-off between child and adult is arbitrarily placed at the eighteenth birthday; it could just as well be a month earlier or three years later.

Now here is the question: Would you say that toddlers should be drafted into the army? No? Seems as if the difference is not so arbitrary after all, even if there is a gradient between the two categories. Likewise, there is nothing arbitrary about the question whether humans and horses are the same species or not.

Finally, genetic data always giving two discrete clusters? Ha, I wish.

In summary, species are complicated. I would be very skeptical of any claim that somebody has sorted it all out.

Wednesday, May 20, 2015

'Monophyletic species' once more

Originally I had decided not to write anything on the "open letter to scientific community" [sic] from about two weeks ago.

Because the author not so much refutes but rather rejects the logic of the arguments of those he criticises, a response could at best be a rephrasing in different words of what I wrote in the paper that incensed him so. It consequently seemed pointless to write a reply, and indeed I would just refer again to that paper anybody who is interested in arguing about the information content of 'evolutionary' classifications, the feasibility of delimiting taxa based on long branches, and the relevance of a distinction between paraphyletic and polyphyletic taxa.

However, a few days later I picked the manuscript attached to the letter up another time and gave more attention than before to the second half, the one where the author's ire is directed towards the publications of Frank Zachos. There I found a section that motivated me to write something after all; not much, but something, simply because the section in question is so depressingly typical of much of the opposition to phylogenetic systematics:
Another subterfuge is to exempt species from the ban on paraphyly ("the concept of paraphyly does not apply to the species category"), so the "actual common ancestor is (or was) a species, but it does (did) not belong to any supraspecific subdivision of the descendant group" - again a desctructive [sic] (making classification cripple [sic], with millions - one for each "accepted" non-monotypic taxon! - of species "not belonging anywhere") and illogical "convention" designed only to defend the indefensibly harmful dogma
It contains two claims of interest.

Thursday, April 30, 2015

CBA Conference 2015, last day

The final day of the 2015 CBA conference on species delimitation featured another set of very interesting talks.

Sally Potter presented her and colleagues' exome capture pipeline for the study of skinks. Admittedly this was already so specific and applied that one would not expect to walk away with many concrete ideas for one's own work unless one wants to use pretty much the same method, but she also mentioned a species tree method that had so far escaped my attention.

ASTRAL is supposedly very fast for large numbers of loci; it is said to be "improving on MP-EST and the population tree from BUCKy, two statistically consistent leading coalescent-based methods". Statistical consistency sounds nice, but when I played around with them for a bit back when those precursor methods didn't really convince me. Still, ASTRAL may be worth a try at some point in the future.

Wednesday, April 29, 2015

CBA Conference, day 2

Continuing with the 2015 CBA conference on species delimitation.

Today's first speaker was Sasha Mikheyev, who presented genome sequencing data on hybridisation zones in social insects. Most of his talk focused on Africanised bees in America, where he was lucky enough to find collaborators who had a time series of samples in the freezer capturing exactly the moment when the African genes swamped populations in the southern United States. Important research, but the results were still preliminary, so he could not go into what he ultimately wants to find out about the dynamics of hybridisation.

He ended his talk with a really weird story about an ant species in which queens are clones of their mothers, males are clones of their fathers (how does that work - is there a zygote but the female chromosomes are discarded?), and only the workers are half-half. This means that the males and females of the same "species" are genetically completely isolated, and indeed the male lineage is more closely related to a different species than to the females they have sex with.

You just can't make anything up that is more bizarre than what reality holds in stock for us.

Tuesday, April 28, 2015

CBA Conference 2015, day 1

Most of this week I am at the 2015 CBA conference Species delimitation in the age of genomics.

Although the title suggests that you'd die of alcohol poisoning in short order if you took a drink every time you heard somebody claim that genomic data are going to solve all our problems, the actual talks happily do not fit that stereotype. Maybe people have seen enough genomic data now to dispense with the hyperbole.

Today started off with a talk by the Australian philosopher of science John Wilkins. He argued that the traditional approach taken by most people who are explicitly grappling with the species problem is to declare a theory-based species concept, try to force it onto reality, and then consider the species-ness to be an explanation for what we see in nature: This individual is the way it is because it is a member of that species. In Wilkins' opinion, that is precisely the wrong way around.

Thursday, March 19, 2015

What is the age of a species?

We had journal club today, and I realised that something that I had always considered quite obvious and logical is not as obvious and logical to everybody else. The question here is: how old is a species?

If you are like me, you will visualise the phylogenetic tree of the species and its closest relatives, point to the moment where it diverged from its sister group, and say that that is its age. See the following diagram:


We know today of the existence of species A, C and D, and if we reconstructed the phylogeny I would say that the divergence time of A on the one side and the ancestor of C and D on the other side, marked with the black line and the number 1, is the age of species A.

However, there is a little snag: What if there were side branches on the tree of life that we do not know about because they are now extinct? In the hypothetical case above, there existed a species B for some time but then it died out. So really the age of species A should be the divergence time from B, indicated with the purple line and the number 2. But if we only know about A, C and D, we can at least say that A is at the most as old as indicated by the black line.

So far so good. What I learned today is that there are apparently people who believe that the age of a species is where the gene copies it is carrying or the family lineages inside it coalesce back in time. This is indicated by the red line and the number 3: the yellow gene tree shows that all extant gene copies (of that gene) in species A are derived from an ancestral copy that existed at time 3.

This time would often be much closer to the present than times 2 and 3. In the case of us humans, we are usually inferred to have diverged from the chimpanzee lineage a few million years ago, but the last female and male common ancestors of all of humanity - the individuals from whom we have all inherited our mitochondria and Y chromosomes, respectively - have been inferred to have existed a few ten to hundred thousand years ago. And indeed some people seem to assume that that is then the age of our species. (See also my earlier dissection of the Doomsday Argument.)

When I was reminded of this concept today I couldn't believe that this would make sense to any biologist. Mike Crisp pointed out the first major problem: this coalescence point constantly moves through time as gene lineages and family lineages proliferate and die out again. As one can see even in my little diagram above, there were several earlier times at which the same would have been true. So how can the age of the same species constantly be corrected downwards as it gets older?

Worse, if speciation events are spaced closely in time and effective population sizes are large, it is well possible that coalescence time might actually be in an ancestor of species A instead of itself; this is known as incomplete lineage sorting or ancestral polymorphism. In other words, species A would then be older than its existence as an independent lineage. Surely that is immediately recognisable as nonsensical.

However, Mike also pointed out that (as I would add: in cases where it is after the lineage divergence time that we can infer), this coalescence point provides a lower limit on the age of a species. So it must then at least be as old as that, and it is at most as old as the divergence time from its living relatives.

Thursday, October 3, 2013

Speciation in plants

On his Sandwalk blog, biochemist Larry Moran recently posted about the many different definitions of the term evolution. Very interesting topic but I see no reason to duplicate his efforts on this much less frequently visited blog. What got me thinking instead was a somewhat incoherent creationist who immediately jumped into the comment thread to declare
Evolution is not a fact or a theory in explaining changes in populations beyond reproducing types at any one point in history. Show where a population changed so that it could no longer reproduce with its parent population??
So he denies the possibility of speciation. Would be laughable in face of all the evidence we have, but the problem is of course that somebody like that will never be satisfied with the kind of evidence a biologist can actually produce. Speciation rarely happens over a human lifespan (although sometimes it can, see below), and so the retort will always be, "were you there?", or to let the creationist speak,
Of coarse extrapolating from these is not scientific evidence for a history of biology. Even if true history.
That is a fairly myopic stance, entirely comparable to doubting that the Roman Empire existed because nobody who is currently alive has seen it and historical documents and archeological artifacts do not constitute direct evidence. I think Richard Dawkins used the same analogy in a recent book, actually.

No, the standard of evidence for processes that take hundreds of thousands of years to happen cannot possibly be "have you seen it with your own eyes?" But just as with the Roman empire or with a murder trial lacking direct witnesses, we can infer events beyond reasonable doubt. The nested structure of biological diversity is our friend here - we can produce phylogenetic trees to find out which species are most closely related, and sometimes we find two species that are genetically so extremely close together that we can infer what may have happened to make them different species, perhaps as little time ago as before the last ice age. In other cases, we find groups in the incipient stages of speciation, somewhere uncomfortably between being able to interbreed and not being able to interbreed.

Anyway, this got me thinking. The evolution/creationism controversy is usually fought over animals, whereas plants are usually cruelly neglected. So what is the state of evidence in plants? I am not talking about a formal review here, but merely a short list of spectacular cases, well understood mechanisms of speciation, and some well researched examples. Here goes.

Tuesday, July 2, 2013

My own take of the species problem

This is the final post in my short series on species concepts. After discussing a selection of synchronous species concepts, with particular focus on the biological one and the Genotypic Cluster Species Concept, and a rather short summary of asynchronous concepts, I will now sketch out my own current thinking on species. Note again that I am not a specialist on speciation or suchlike, and that my conclusions are obviously tentative. But well, as a systematist and taxonomist I cannot avoid dealing with the issue every day, and that includes sometimes problematic controversies even with collaborators, so I have to have an opinion.

Monday, June 24, 2013

Asynchronous species concepts: internodal and composite

As indicated before in this series about species concepts, most of them apply only to contemporary organisms or more generally to those existing in the same time slice. That becomes clear quite quickly if we try to apply them in an asynchronous fashion, i.e. through time.

The Biological Species Concept (BSC), for example, sees species as breeding groups. Because we humans do not interbreed with desert oaks, and indeed would find it hard to do so, we are clearly separate species. If we let our gaze drift over all of evolutionary history, however, it becomes clear that there must have been an unbroken chain of individuals connecting me, the desert oak I photographed on the Great Central Road in 2010, and some common ancestor the two of us had sometime a few hundred million years ago. In other words, seen through deep time all of life on earth is a breeding community, and thus the BSC would fail to cleave the diversity of life into species.

As another example, the Genotypic Cluster Species Concept (GCSC) identifies species as clusters of individuals in some morphological or genetic analysis that have no or few intermediates with other such clusters in the same analysis. Again, because all organisms on the planet appear to be parts of the same great tree of life, and because evolutionary change happens gradually through the change of allele frequencies in populations, there just is no place along the branches of the tree of life where there are "no or few intermediates". Just as the BSC, the GCSC would not work.

Of course, if you see species as breeding groups, you might immediately conclude that applying species concepts through time must be absurd. Still, there are those who have tried to formulate ideas on how to make the word "species" work in this context. There aren't many, and one of them I have already discussed before is Willi Hennig's Internodal Species Concept, so I will keep it short on that one. A few more words are then needed for the newer alternative.

Monday, June 3, 2013

Fundamental and realized niche

A short follow-up on the previous post mentioning the Ecological Species Concept: As discussed, ecological niches are not entirely unproblematic. It sounds like a useful mental model to view each species as "fulfilling a role" in its community but reality is more complex.

To mention just one issue, there is a difference between fundamental niches and realized niches. What does that mean? Well, it could be that a plant, for example, could happily grow under wet, mesic and dry conditions if it were the only plant around. However, there is another plant that can also grow under mesic conditions but it is much more competitive than the first plant. That means that if they occur in different areas, the first plant will grow in all three habitats and the second one only in the mesic habitat, but if they happen to occur in the same area, you will find the first species in wet and dry places and the second one in the mesic places where it excludes the first.

For illustration, this is how well both species do along a moisture gradient, which we could call their fundamental niches:


And this is how many individuals of each species you would really be able to find along the moisture gradient if they both occur together, which would show their realized niches:


What this shows is first that no organism should be expected to come with a clearly defined role but that it has to settle into one depending on what other organisms are around. Also, which of these two is the niche relevant for the Ecological Species Concept?

If we go with the realized niche, for example because we have so far only observed the species together in nature and don't realize that the orange species would feel even happier under mesic conditions than where it is actually forced to live, we would probably intuit that the orange populations in dry and the wet habitats are two Ecospecies. Perhaps there is a bit of circular reasoning involved on my part (i.e. I may be smuggling the criteria of the Biological Species Concept into my argument), but this seems plainly absurd to me. But if we take the fundamental niche as a guide, what use is that if it may never be realized in nature?

Wednesday, May 29, 2013

Mopping up the remaining synchronous species concepts

It is time to pick up writing about species concepts again (other posts in this series can be found under the eponymous tag). I have dealt at some length with Biological Species, Genotypic Cluster Species, Morphospecies / "typological" / autapomorphic species, and the idea that species must be monophyletic. There are numerous synchronous species concepts left before I can tackle the asynchronous ones, but unfortunately I do not feel that I can treat all of them in the same depth. Some of them are simply variants of others, some apply only to very special cases, and some I simply don't really understand. I will therefore deal with all of them together in this post.

Monday, April 29, 2013

"Monophylogenetic" species

This continues a series on species. The previous episodes introduced the topic, provided an intuitive classification of species concepts, and dealt with biological, genotypic cluster and "typological" species.

 The term "phylogenetic" became so popular after phylogenetic systematics gained ascendency in the systematic and taxonomic community that several quite unrelated species concepts were published under that label. In the previous post you may have noticed that I call something the autoapomorphic species concept following this list compiled by a philosopher of science although it was really published as “The phylogenetic species concept (sensu Wheeler and Platnick)”. Not only does that clarification in the brackets nicely demonstrate the problem of homonymy here, but I am also unsure what exactly is so phylogenetic about a concept considering species to be groups of samples with a unique character combination.

I will therefore limit this post to discussing the so-called phylogenetic species concepts that demand species be monophyletic (i.e. the Phylogenetic Taxon Species in the list mentioned above), although that necessarily means that I will in part reiterate what I already wrote before.

Thursday, April 18, 2013

Typological species (?) and autapomorphic species

This post continuing my series on species will perhaps not do its topic full justice, but unfortunately I currently cannot seem to find the time to write as much.

The term "Typological Species Concept" (TSC) is one that you can hear rather often from taxonomists, generally accompanied by a sneer, but it appears surprisingly ill-defined. Googleing around a bit one can, for example, encounter these teaching materials from Southern Illinois University which describe it as follows: "typological species are defined by similarity to a type specimen or ideal type." This is then dismissed as biologically unjustified and followed by several alternative species concepts.

Defining the TSC like this is in my opinion rather strange because it confuses systematics and nomenclature. No matter how a systematist circumscribes species and no matter what species concept they use to do that, afterwards a name has to be assigned to each of these species, and that is always done with reference to a type specimen. A species that contains the type specimen of, say, Bellis perennis will be called Bellis perennis regardless of whether it is a biological species, a phenotypic cluster species or a phylogenetic species*. Those are simply the rules of nomenclature and not another distinct species concept.

So, be the TSC formally defined as such somewhere or not, I will now describe how we in the business appear to use the term. As mentioned above, it is generally used as an accusation, as in "yeah, that guy used a typological species concept, his treatment is completely useless". What the detractors mean in that case is that the taxonomist in question has used a very unbiological and schematic approach to circumscribing species. Typical criticism would include:
  • Excessive splitting, especially "single character taxonomy" in which a specimen showing any morphological deviation whatsoever is immediately given taxonomic recognition as a variety, subspecies or even species
  • Failure to take the plasticity of the organisms into account, basing taxonomic decisions on characters that are variable under varying environmental conditions; this is often due to the taxonomist not having seen the study group in the field
  • Failure to reflect the hypothesized relatedness of the samples in the classification, in extreme cases assigning varieties that are admitted to be most closely related to each other to separate species because they differ in a character that was arbitrarily considered "important enough" to define the species level
As an example, we might consider a breeding group of daisies in which some populations have a pappus and others don't. Now although there are such cases this happens rarely - usually an entire species of daisies either has a pappus or it hasn't. A competent taxonomist using the Phenotypic Cluster Concept would look at overall variation and conclude that the character "pappus presence" does not correlate with anything else and that the overlap in morphological variability between the populations with and those without a pappus is otherwise so huge that they should be one species. A colleague using the Biological Species Concept would find that the two forms interbreed and make them one species. In contrast, a taxonomist with a typological approach (as commonly understood) might well consider them two different species. Their gut feeling would be that presence or absence of a pappus has always been a "taxonomically important character" in daisies, and that is that.

Monday, April 8, 2013

Genotypic cluster species (and similar)

This post continues a small series on species that started here and continued with this post.

When I wrote about the Biological Species Concept and its relatives, I wrote that it is what most non-scientists would spontaneously suggest when asked for a definition of species, and that it is also very popular especially with zoologists. At least in theory - it is not very practical to conduct crossing experiments every time you write a monograph or describe a novelty you found as a single specimen on a field trip. It is, however, clearly the concept that evolutionary biologists and everybody studying speciation have in their minds. In contrast, I would argue the concept to be discussed today is the one that most competent taxonomists use in actual practice, often intuitively but increasingly explicitly and supported by quantitative analysis.

The Genotypic Cluster Concept of species (subsequently GCC) was formalized by Mallet (1995). It sees species as groups of individuals that form genetic or morphological clusters with few or no intermediates to other such clusters.

Saturday, March 30, 2013

Biological species and related concepts

This post is part of a series on species. Read the previous parts here and here.

The Biological Species Concept (BSC) is perhaps the most famous of them all. In science, it is one of the things evolutionary biologists Theodosius Dobzhansky and (in particular) Ernst Mayr are best known for. It is basically what most non-scientists would also come up with when asked what a species is, unless they have never given the topic a thought before: species are breeding groups, i.e. groups of individuals that form a reproductive community.

Many colleagues will tell you that the BSC is very popular with zoologists but not so much with botanists, supposedly because it works well with animals but not so much with plants. And indeed I know many well accepted plant species that can interbreed, at least potentially. But that is of course circular reasoning: should they then really be separate species? Maybe botanists are simply splitting species too much.

On the other hand, the idea that some interbreeding or even potential interbreeding already makes two populations conspecific may be a caricature of the BSC as promoted, for example, by Coyne & Orr's book Speciation (which, by the way, is a fantastic resource for those interested in species); really the BSC is often understood to allow for some rare gene flow between biological species.

Saturday, March 23, 2013

Classifying the species concepts

This post continues a series on species that began with this introduction.

So looking over the literature and the aforementioned list of 26 species concepts, let's try to classify the concepts into different groups:

Synchronous concepts
    Reproductive community concepts
            *Compilospecies
        Biological species
        Cohesion species
        Genetic species
        Genic species (?)
        Recognition species
        Reproductive competition species
    Phenetic and cluster concepts
        Autapomorphic species
        Ecospecies
        Genotypic cluster
            *Agamospecies
        Morphospecies or typological species
        Phenospecies
    Phylogenetic concepts
        Phylogenetic species
        Genealogical concordance species
Asynchronous concepts
    Internodal species
    Composite species
        *Successional species
Incertae sedis
    Evolutionary species

The first division, indicated by bold font, is whether the concepts deal only with species in the here and now or, more generally speaking, with those existing together in one time-slice (synchronous) or whether they have a historical dimension (asynchronous). Note that the latter are rarely of interest for anything beyond theoretical discussion unless we are talking about a group with an excellent continuous fossil record. I am not sure where to place the Evolutionary Species Concept because I may not quite understand it. It sounds a bit like a reproductive community seen from an asynchronous perspective, so I have placed it apart for the moment. (Similarly, I previously understood the genic species to mean something different from how it is described in the list I linked to, so there is a question mark behind it.)

Within the synchronous concepts that deal with delimiting and circumscribing species in the present we have a second division, indicated in italics, where the concepts are grouped by their main grouping and, if defined, ranking criteria. The first group is what we might call the family of the biological species concept, different concepts that attempt to delimit species by the ability or propensity of individual organisms to interbreed. The second group defines species via some kind of character combination or overall similarity. In all these cases, species are groups of individuals that are ecologically, morphologically or genetically more similar to each other than they are to individuals belonging to any other species. In contrast to this essentially phenetic approach, the small third group employs some kind of phylogenetic approach to species circumscription.

I have left out several concepts that are silly or too similar to others. The ones that are further indented and marked with asterisks are also a bit dubious as they appear to represent merely special cases of the concepts directly above them. Even having reduced the number already, it will still be useful to lump them further and discuss, for example, all reproductive community concepts together. On the other hand, it appears useful to take a more granular approach to the phenetic concepts - but that may just reflect my personal biases.