GENESPACE tracks regions of interest and gene copy number variation across multiple genomes

Abstract
Editor's evaluation
eLife digest
Introduction
Results and discussion
Materials and methods
Data availability
References
Article and author information
Metrics

Abstract

The development of multiple chromosome-scale reference genome sequences in many taxonomic groups has yielded a high-resolution view of the patterns and processes of molecular evolution. Nonetheless, leveraging information across multiple genomes remains a significant challenge in nearly all eukaryotic systems. These challenges range from studying the evolution of chromosome structure, to finding candidate genes for quantitative trait loci, to testing hypotheses about speciation and adaptation. Here, we present GENESPACE, which addresses these challenges by integrating conserved gene order and orthology to define the expected physical position of all genes across multiple genomes. We demonstrate this utility by dissecting presence–absence, copy-number, and structural variation at three levels of biological organization: spanning 300 million years of vertebrate sex chromosome evolution, across the diversity of the Poaceae (grass) plant family, and among 26 maize cultivars. The methods to build and visualize syntenic orthology in the GENESPACE R package offer a significant addition to existing gene family and synteny programs, especially in polyploid, outbred, and other complex genomes.

Editor's evaluation

GENESPACE is a new and straightforward computational tool to include synteny information in the calculation of genome-wide sets of orthologs. The development of this tool is very timely as more and more complete chromosome-scale assembled genomes are becoming available. While the assembly problem has been solved, this is not the case for multiple genome comparisons, and GENESPACE is an important step to help remedy this gap in our comparative genomics toolbox.

https://doi.org/10.7554/eLife.78526.sa0

eLife digest

The genome is the complete DNA sequence of an individual. It is a crucial foundation for many studies in medicine, agriculture, and conservation biology. Advances in genetics have made it possible to rapidly sequence, or read out, the genome of many organisms. For closely related species, scientists can then do detailed comparisons, revealing similar genes with a shared past or a common role, but comparing more distantly related organisms remains difficult.

One major challenge is that genes are often lost or duplicated over evolutionary time. One way to be more confident is to look at ‘synteny’, or how genes are organized or ordered within the genome. In some groups of species, synteny persists across millions of years of evolution. Combining sequence similarity with gene order could make comparisons between distantly related species more robust.

To do this, Lovell et al. developed GENESPACE, a software that links similarities between DNA sequences to the order of genes in a genome. This allows researchers to visualize and explore related DNA sequences and determine whether genes have been lost or duplicated. To demonstrate the value of GENESPACE, Lovell et al. explored evolution in vertebrates and flowering plants. The software was able to highlight the shared sequences between unique sex chromosomes in birds and mammals, and it was able to track the positions of genes important in the evolution of grass crops including maize, wheat, and rice.

Exploring the genetic code in this way could lead to a better understanding of the evolution of important sections of the genome. It might also allow scientists to find target genes for applications like crop improvement. Lovell et al. have designed the GENESPACE software to be easy for other scientists to use, allowing them to make graphics and perform analyses with few programming skills.

Introduction

De novo genome assemblies and gene model annotations represent increasingly common resources that describe the sequence and positions of protein coding and intergenic regions within a single genotype. Evolutionary relationships among these DNA sequences form the foundation of many molecular tools in modern medical, breeding, and evolutionary biology research.

Perhaps the most crucial inference to make when comparing genomes revolves around homologous genes, which share an evolutionary common ancestor and ensuing sequence or protein structure similarity. Analyses of homologs, including comparative gene expression, epigenetics, and sequence evolution, require the distinction between orthologs, which arise from speciation events, and paralogs, which arise from sequence duplications (see Box 1 for definitions). In some systems, this is a simple task where most genes are single copy, and orthologs are synonymous with reciprocal best-scoring protein BLAST hits. Other sequence similarity approaches such as OrthoFinder (Emms and Kelly, 2019; Emms and Kelly, 2015) leverage graphs and gene trees to test for orthology, permitting more robust analyses in systems with gene copy number (CNV) or presence–absence variation (PAV). However, whole-genome duplications (WGDs), chromosomal deletions, and variable rates of sequence evolution, such as subgenome dominance in polyploids, can confound the evidence of orthology from sequence similarity alone.

Box 1

Definitions

Orthogroup — a set of genes across multiple genomes derived from a single ancestral gene.

Ortholog — a pair of orthogroup members in two species derived from a single gene in their most recent common ancestor.

Paralog — orthogroup members derived from a duplication event since speciation.

Homeolog — paralogs derived from a whole-genome duplication.

Tandem array — paralogs in proximity on a chromosome within a genome.

Gene collinearity — retained order of genes across species due to common ancestry.

Synteny — like collinearity but at larger scales, like chromosome arms.

Pan-genome annotation — A set of orthogroups across multiple genomes, placed along the coordinate system of a specified reference genome.

The physical position of homologs offers a second line of evidence that can help overcome challenges posed by WGDs, tandem arrays, heterozygous-duplicated regions, and other genomic complexities (Drillon et al., 2020; Haug-Baltzell et al., 2017; Wang et al., 2012). Synteny, or the conserved order of DNA sequences among chromosomes that share a common ancestor, is a typical feature of eukaryotic genomes. In some taxa, synteny is preserved across hundreds of millions of years of evolution and is retained over multiple WGDs (Jiao et al., 2014; Simakov et al., 2020; Zhao and Schranz, 2019). Like chromosome-scale synteny, conserved gene order collinearity along local regions of chromosomes can provide evidence of homology, and in some cases enable determination of whether two regions diverged as a result of speciation or a large-scale duplication event (Drillon et al., 2020). Combined, evidence of gene collinearity and sequence similarity should improve the ability to classify paralogous and orthologous relationships beyond either approach in isolation.

Integrating synteny and collinearity into comparative genomics pipelines also physically anchors the positions of related gene sequences onto the assemblies of each genome. For example, by exploring only syntenic orthologs it is possible to examine all putatively functional variants within a genomic region of interest, even those in genes that are absent in the focal reference genome (Lovell et al., 2018). Such a ‘pan-genome annotation’ framework (Lovell et al., 2021a) permits access to multi-genome networks of high-confidence orthologs and paralogs, regardless of ploidy or other complicating aspects of genome biology. Here, we present GENESPACE, an analytical pipeline that explicitly links synteny and sequence similarity to provide high-confidence inference about networks of genes that share a common ancestor and represents these networks as a ‘pan-genome annotation’. We then leverage this framework to explore gene family evolution in flowering plants, mammals, and reptiles.

Results and discussion

GENESPACE syntenic orthology methods to compare multiple complex genomes

Comparative genomics across the complex evolutionary histories of eukaryotes typically requires equally varied input and analytical pipelines depending on researchers’ goals and study systems. For example, synteny between closely related haploid assemblies is often inferred by exploring only 1:1 reciprocal best-scoring hits with MCScanX (Wang et al., 2012). Alternatively, polyploid genomes are typically split into subgenomes so that homeologs are viewed by clustering algorithms like OrthoFinder (Emms and Kelly, 2019; Emms and Kelly, 2015) as orthologous and not paralogous. While expert knowledge that informs these analytical decisions can dramatically improve precision, this knowledge is not available in many systems. These issues boil down to a simple circular problem: a priori knowledge of gene copy number is needed to effectively infer orthology and synteny, yet measures of synteny and orthology are needed to infer copy number between a pair of sequences.

GENESPACE resolves this circular problem by operating on a foundational assumption: homologs should be exactly single copy within any syntenic region between a pair of genomes. There are two major violations of this assumption that cause copy number variation within a syntenic block: (1) tandem arrays and (2) gene PAV. GENESPACE addresses these complexities directly in two ways. First, physically proximate multigene families (hereon, ‘tandem arrays’, Box 1) are condensed to the physically most central gene of the array. Gene rank order along the genome is recalculated on these ‘array representative’ genes, effectively masking copy number variation due to tandem arrays. Second, synteny is inferred in a pairwise manner only using ‘potential anchor’ protein BLAST hits (hereon ‘hits’) where both the query and target genes are in the same orthogroup. The rank-order positions of these potential anchor genes are also recalculated prior to synteny inference, which effectively masks orthogroups missing a gene in one genome (i.e., PAV). Thus, GENESPACE operates on OrthoFinder-derived orthogroups with one-to-one relationships for any accurately defined syntenic region, regardless of ploidy or level of sequence conservation.

Given GENESPACE’s reliance on syntenic regions between genomes, errors in syntenic block coordinates can have major effects on downstream estimates. Therefore, we crafted a sensitive pipeline to infer syntenic regions (see pipeline overview in Figure 1). To demonstrate its functionality, we ran GENESPACE on seven pairs of genomes spanning closely related maize cultivars to ancient mammal–bird divergence (~300 M ya; Table 1, see Materials and methods). The three comparisons between vertebrate genomes have no recent history of WGDs, while WGDs predated evolutionary divergence of the four plant contrasts: maize is a 12 M ya paleo-tetraploid, cotton is a 1.6 M ya meso-tetraploid, and the ~70 M ya Rho WGD predated grass diversification. Paralogs derived from these WGDs notoriously obscure contrasts between orthologs in these systems. As such, we treated all genomes as haploid and did not include an outgroup; therefore, orthogroups should not include homeologs since the genomes share a common ancestor that arose after the WGD (see ‘additional considerations’ methods section: Outgroups and the phylogenetic context of orthology inference). In each run, we contrasted GENESPACE-derived orthogroups and syntenic blocks to the defaults from OrthoFinder and MCScanX, respectively (Table 1). Since it is a common practice to refine hits prior to running MCScanX, we also included syntenic block coverage from MCScanX run on orthogroup-constrained hits where both the target and query genes must be in the same orthogroup. To contrast each approach, we calculated the percent of genes in (1) orthogroups and (2) syntenic blocks that were placed on exactly one chromosome per genome (Table 1).

Figure 1

Download asset Open asset

GENESPACE synteny and pan-genome annotation methods.

(A, grey panel) GENESPACE runs and parses OrthoFinder results into a synteny-constrained pan-genome annotation. (B, purple panel) Chromosome, gene rank order, and orthogroup membership are added to BLAST hits, which allows direct integration between estimates of orthology and synteny. The three dotplots present the efficacy of GENESPACE syntenic blocks by exploring a particularly challenging region on human (x-axis) and chimpanzee (y-axis) chr. 6. Each point is a BLAST hit rank-order position, colored by syntenic block; colors are recycled if there are more than eight blocks. (C, green panel) Synteny-constrained orthogroups and optionally non-syntenic orthologs are decomposed into a pan-genome annotation where each orthogroup is placed at its inferred syntenic position.

Table 1

Comparison of synteny and orthogroup methods.

To test the precision of GENESPACE syntenic orthogroups estimates, we contrasted seven pairs of haploid genome assemblies. We present the percent of genes that were found in an orthogroup that hit a single chromosome per genome from the default OrthoFinder and GENESPACE runs. The precision of syntenic block breakpoint estimates was calculated similarly, where the percentage of genes that are placed in a single syntenic block per genome are presented for MCScanX run on all hits, those where both the query and target genes are in the same orthogroup (‘OG’) or via the GENESPACE pipeline.

		(a) % genes in single-copy OGs		(b) % genes in single-copy syntenic blocks
	Age (~M ya)	OrthoFinder	GENESPACE	MCScanX	MCScanX OG	GENESPACE
B73 vs. B97 maize*	<0.01	51.5	73.6	50.8	79.0	93.4
Human Hg38 vs. T2T	0–0.1	87.7	95.9	81.1	95.0	97.7
Cotton*^,+	0.5	35.6	85.7	2.7	14.1	96.2
HAL2 vs. FIL2 panicgrass*	1.1	74.8	83.2	62.3	89.3	92.0
Human-chimpanzee	7	81.1	90.2	78.6	91.2	93.3
Sorghum-Brachypodium*	50	46.7	50.2	49.3	67.4	76.3
Human-chicken	310	66.7	68.5	66.4	71.2	73.0

*

The plant genomes all have one or more WGDs that predate divergence of the genomes,.⁺Cotton species Gossypium barbadense and G. darwinii have the most recent WGD of ~1.6 M ya, which causes a large number of blocks to be included as two copies; to avoid confusion between subgenomes, blkSize, and nGaps parameters were increased from 5 (default) to 10 genes.

For every pair of genomes, GENESPACE produced a greater percentage of single-copy syntenic blocks and genes in single chromosome orthogroups than either OrthoFinder or MCScanX in isolation. GENESPACE also outperformed simple integration between the two methods through MCScanX on orthogroup-constrained hits. The improved performance of GENESPACE synteny-constrained orthogroups was most subtle between highly diverged haploid animal assemblies. For example, in the comparison between human and chicken, GENESPACE resolved 2% and 1.8% more single-copy orthogroups and syntenic blocks than default OrthoFinder and orthogroup-constrained MCScanX, respectively. In contrast, the benefits of GENESPACE were most pronounced in recently diverged genomes with a history of WGDs. Single-copy orthogroups between two meso-tetraploid cotton species that share a WGD that predated speciation by ~1 M ya, were uncommon in the default OrthoFinder run (35.6%) but far more prevalent in GENESPACE syntenic orthogroups (85.7%). Similarly, homeologs derived from the cotton WGD impacted estimates of syntenic blocks: only 14% of the genomes were single-copy syntenic in orthogroup-constrained MCScanX but 96.2% were single copy in GENESPACE syntenic blocks. Combined, these results demonstrate significant flexibility and utility of GENESPACE across a range of evolutionary histories and divergence.

It is important to note that some evolutionary processes, including small-scale translocations, can cause true orthologs to exist outside of syntenic regions. Closely related genomes without a history of WGDs tend to have few non-syntenic orthologs. For example, there are 1096 non-syntenic orthologs (6.3% of all orthologs) between human and chimpanzee. In contrast, the 50 M ya diverged Sorghum and Brachypodium genomes have 9002 non-syntenic orthologs, many of which are the result of over-retained Rho WGD-derived paralogs (see below). Since the non-syntenic orthologs can be important in some situations, GENESPACE embraces this complexity by including and flagging non-syntenic orthologs within the pan-genome annotation (Figure 1).

Synteny-anchored exploration of vertebrate sex chromosomes

GENESPACE facilitates the exploration and analysis of sequence evolution across multiple genomes within regions of interest (e.g., quantitative trait loci [QTL] intervals, see the next section). One particularly instructive example comes from the origin and evolution of the mammalian XY and avian ZW sex chromosome systems. To explore these chromosomes, we ran GENESPACE on 15 haploid avian and mammalian genome assemblies, spanning most major clades of birds, placental mammals, monotremes, and marsupials with available chromosome-scale annotated reference genomes (See Materials and methods). We also included two reptile genomes as outgroups to the avian genomes. The heteromorphic chromosomes (Y and W) are often unassembled, or, where assemblies exist, lack sufficient synteny to provide a useful metric for comparative genomics. As such, we chose to focus on the homomorphic X and Z chromosomes, which have remained surprisingly intact over the >100 M years of independent mammalian (Murphy et al., 1999) and avian evolution (Zhou et al., 2014; Figure 2, Figure 2—figure supplement 1).

Figure 2 with 1 supplement see all

Download asset Open asset

Sex chromosome syntenic network across 17 representative vertebrate genomes.

The plot was generated by the plot_riparian GENESPACE function. Genomes are ordered vertically to minimize the number of translocations between each pairwise combination. Chromosomes are ordered horizontally to maximize synteny with the human chromosomes [X, 1–22]. Regions containing syntenic orthogroup members to the mammalian X (gold) or avian Z (blue) chromosomes are highlighted. All sex chromosomes are represented by red segments while autosomes are white. Chromosome segment sizes are scaled by the total number of genes in syntenic networks and positions of the braids are the gene order along the chromosome sequence. See Figure 2—figure supplement 1 for the full synteny graph including autosomes and chromosome labels.

While the same or similar genomic regions often recurrently evolve into sex chromosomes, perhaps due to ancestral gene functions involved in gonadogenesis, evidence about the nonrandomness of sex chromosome evolution is still contentious (Kratochvíl et al., 2021). Given our analysis, the avian Z chromosome clearly did not evolve from either of the two reptile Z chromosomes sampled here, but instead likely arose from autosomal regions or unsampled ancestral sex chromosomes. The situation in mammals is less clear, in part because both reptile genomes are more closely related to avian than mammalian genomes, which makes ancestral state reconstructions between the two groups less accurate. Nonetheless, the mammalian X and sand lizard Z chromosomes partially share syntenic orthology, an outcome that would be consistent with common descent from a shared ancestral sex chromosome or autosome containing sex-related genes. The shared 91.7 M bp region between the human X and sand lizard Z represents 59.0% of the human X chromosome genic sequence. The remaining 64.0 M bp of human X linked sequences are syntenic with autosomes in the sand lizard and garter snake genomes (Figure 2).

The eutherian mammalian X chromosome is largely composed of two regions, an X-conserved ancestral sex chromosome region that arose in the common ancestor of therian mammals, and an X-added region that arose in the common ancestor of eutherians (Ross et al., 2005). Consistent with this evolutionary history, the X chromosome is syntenic across all five eutherian mammals studied here. Further, a 107.2 M bp (68.8%) segment of the human X, which corresponds with the X-conserved region, is syntenic with 77.8 M bp (93.9%) of the Tasmanian devil X chromosome and represents the entire syntenic region between the human and all three marsupial X chromosomes (Figure 2).

Similarly, the chicken Z chromosome is retained in its entirety across all five avian genomes. The only notable exception being the budgie Z chromosome, which features a partial fusion between the Z and an otherwise autosomal 19.5 M bp segment of chicken chromosome 11 (Figure 2), potentially representing a neo-sex chromosome fusion that has not yet been described.

In contrast to conserved eutherian and avian sex chromosomes, the complex monotreme XnYn sex chromosomes are only partially syntenic between the two sampled genomes. Only the first X chromosomes are ancestral to both echidna and platypus (Rens et al., 2007), and all are unrelated to the mammalian X chromosomes (Figure 2—figure supplement 1), consistent with their independent evolution (Rens et al., 2007). Interestingly, the entirety of the echidna X4 and 47.6 M bp (67.9%) of the genic region of the platypus X5 chromosomes are syntenic with the avian Z chromosome (Figure 2). The phylogenetic scale of the genomes presented here precludes evolutionary inference about the origin of these shared sex chromosome sequences; however, the possibility of parallel evolution of sex chromosomes between such diverged lineages may prove an interesting future line of inquiry.

Exploiting synteny to track candidate genes in grasses

The Poaceae grass plant family is one of the best studied lineages of all multicellular eukaryotes and includes experimental model species (Brachypodium distachyon; Panicum hallii; Setaria viridis) and many of the most productive (Zea mays – maize/corn; Triticum aestivum – wheat, Oryza sativa – rice) and emerging (Sorghum bicolor – sorghum; Panicum virgatum – switchgrass) agricultural crops. Despite the tremendous genetic resources of these and other grasses, genomic comparisons among grasses are difficult, in part because of an ancient polyploid origin (see the next section), and because subsequent WGDs are a feature of most clades of grasses. For example, maize is an 11.4 M ya paleo-polyploid (Gaut and Doebley, 1997), allo-tetraploid switchgrass formed 4–6 M ya (Lovell et al., 2021b), and allo-hexaploid bread wheat arose about 8 k ya (Haas et al., 2019). In some cases, homeologous gene duplications from polyploidy have generated genetic diversity that can be targeted for crop improvement; however, in other cases the genetic basis of trait variation may be restricted to sequences that arose in a single subgenome. Thus, it is crucial to contextualize comparative–quantitative genomics and explicitly explore only the orthologous or homeologous regions of interest when searching for markers or candidate genes underlying heritable trait variation — a significant challenge in the complex and polyploid grass genomes.

To help overcome this challenge and provide tools for grass comparative genomics, we conducted a GENESPACE run and built an interactive viewer hosted on Phytozome (Goodstein et al., 2012). Owing to its use of within-block orthology and synteny constraints (Figure 1), GENESPACE is ideally suited to conduct comparisons across species with diverse polyploidy events. Default parameters produced a largely contiguous map of synteny even across notoriously difficult comparisons like the paleo-homeologs between the maize subgenomes (Figure 3A). Furthermore, the sensitive synteny construction pipeline implemented by GENESPACE effectively masks additional paralogous regions like those from the Rho duplication that gave rise to all extant grasses.

Figure 3 with 1 supplement see all

Download asset Open asset

Comparative–quantitative genomics in the grasses.

(A) The GENESPACE syntenic map (‘riparian plot’) of orthologous regions among eight grass genomes. Chromosomes are ordered horizontally to maximize synteny with rice and ribbons are color coded by synteny to rice chromosomes. Genomes are ordered vertically by general phylogenetic positions. (B) The upper bars display the proportion of maize gene models without syntenic orthologs (‘absent’) in each genome, split by the full background (dark colors) and 86 C₃/C₄ genes (light colors). (C) The proportion of absent genes is higher in the C₃ genomes (green bars), even when controlling for more global gene absences (lower odds ratios). (D) Syntenic orthologs, excluding homeologs among the 26 maize nested association mapping (NAM) founder genomes, with two quantitative trait loci (QTL) intervals highlighted on chromosome 3 (‘Chr3’) and chromosome 6 (’Chr6’). (E) Focal QTL regions that affect productivity in drought where only the genome that drives the QTL effect (middle), the top (B73) and bottom (Tzi8) genomes are presented and the region plotted is restricted to the physical B73 QTL interval and a 25 M bp buffer on either side. Note that the Chr3 QTL disarticulates into two intervals. Due to a larger number of potential candidate genes, the larger Chr3 region, flagged with **, is explored separately in Figure 3—figure supplement 1. (F) Presence–absence and copy number variation are presented for two of the three intervals as heatmaps where each row is a genome (order following panel D), each column is a pan-genome entry (see Figure 1), and the color of each tile indicates absence (gray), single copy (light blue), and multicopy (dark blue). PAV/CNV of the focal genome is outlined. For each interval, the estimated QTL allelic effect relative to B73 of each genome is plotted as bars to the right of the heatmap.

Breeders and molecular biologists can take two general approaches to understand the genetic basis of complex traits: studying variation caused by a priori-defined genes of interest or determining candidate genes from genomic regions of interest. As an example of the exploration of lists of a priori-defined candidate genes, we analyzed PAV of 86 genes shown to be involved in the transition between C₃ and C₄ photosynthesis (Ding et al., 2015), the latter permitting ecological dominance in arid climates and agricultural productivity under forecasted increased heat load of the next century. To conduct this analysis, we built pan-genome annotations across the seven grasses anchored to C₄ maize, which was the genome in which these genes were discovered. This resulted in 159 pan-genome entries: nearly always two placements for each gene in the paleo-tetraploid maize genome. Given that many of these genes were discovered in part because of sequence similarity to genes in Arabidopsis and other diverged plant species, it is not surprising that PAV among C₃/C₄ genes was lower than the background (9.7% vs. 38.2%, odds ratio = 5.7, p < 1 × 10⁻¹⁶; Figure 3B). However, these ratios were highly variable among genomes, particularly among the C₃ species (wheat, rice, B. distachyon), which had far more absences than the C₄ species (15.3% vs. 5.5%, odds ratio = 3.1, p = 6.25 × 10⁻⁸, Figure 3B). This effect is undoubtedly due in part to the increased evolutionary distance between maize and the C₃ species compared to the other C₄ species. However, when controlling for the elevated level of absent genes globally in C₃ species, the effect was still very strong: the odds of C₃ species having more of these C₃/C₄ genes at syntenic pan-genome positions than the background was always lower than the C₄ species (Figure 3C). Despite these interesting patterns, given only a single C₃/C₄ phylogenetic split in this dataset, it is not possible to test evolutionary hypotheses regarding the causes of such PAV. Nonetheless, this result suggests a possible role of gene loss or gain as an evolutionary mechanism for drought- and heat-adapted photosynthesis.

Like the exploration of a priori-defined sets of genes, finding candidate genes within QTL intervals usually involves querying a single reference genome and extracting genes with promising annotations or putatively functional polymorphism. In the case of a biparental mapping population genotyped against a single reference, this is a trivial process where genes within physical bounds of a QTL are the candidates. However, many genetic mapping populations now have reference genome sequences for all parents; this offers an opportunity to explore variation among functional alleles and PAV, which would be impossible with a single reference genome. GENESPACE is ideally suited for this type of exploration, and indeed was originally designed to solve this problem between the two P. hallii reference genomes and their F₂ progeny (Lovell et al., 2018) using synteny to project the positions of genes across multiple genomes onto the physical positions of a reference. To illustrate this approach, we reanalyzed QTL generated from the 26-parent USA maize nested association mapping (NAM) population (Li et al., 2016). Originally, candidates for these QTL were defined by the proximate gene models only in the B73 reference genome (Li et al., 2016); however, with GENESPACE and the recently released NAM parent genomes (Hufford et al., 2021), it is now possible to evaluate candidate genes present in the genomes of other NAM founder lines but either absent or unannotated in the B73 reference genome.

To explore this possibility, we built a single-copy synteny map across all 26 NAM founders anchored to the B73 genome (Figure 3D). We opted to focus on QTL where the allelic effect of a single parental genome was an outlier relative to all other alleles. Such ‘private’ allelic contributions in multi-parent populations offer a powerful opportunity to define high-confidence candidates as genes with parent-specific sequences that match parent-specific allelic contributions to phenotypic trait variation (Abdulkina et al., 2019). Among the 190 QTLs, three displayed outlier effects of one parent: cultivar ‘Mo18W’ contributed an allele for delayed anthesis-silking interval at two adjacent chromosome 3 (‘Chr3’) QTLs and cultivar ‘Ki11’ contributed an allele that reduced plant height under drought at a QTL on chromosome 6 (‘Chr6’) (Li et al., 2016). Given that these QTL were chosen only due to their parental allelic effects, we were surprised to find that the two Mo18W QTL regions exist within a 11.7 M bp derived inversion that is only found in the Mo18W genome (Figure 3D, E). Since inversions reduce recombination, it is possible that multiple Mo18W causal variants have been fixed in linkage disequilibrium in this NAM population.

In addition to this chromosomal mutation and sequence variation between the parents and B73 (Li et al., 2016), we sought to define candidate genes from the patterns of presence–absence and copy number variation, explicitly looking for genes that were private to the focal genome. Two genes in the smaller Chr3 and one gene in the larger Chr3 interval were private to Mo18W, and four genes in three pan-genome entries (one two-member array) were private to Ki11 in the Chr6 interval (Figure 3F, Figure 3—figure supplement 1). While these genes do not have functional annotations relating to drought, this method provides additional candidates that would not have been discovered by B73-only candidate gene exploration.

Studying the WGD that led to the diversification of the grasses

Like most plant families (Barker et al., 2016; Stebbins, 1950; One Thousand Plant Transcriptomes Initiative, 2019), but unlike nearly all animal lineages (Muller, 1925), the grasses radiated following a whole-genome duplication: the ~70 M ya Rho WGD. The resulting gene family redundancy and gene-function subfunctionalization is hypothesized to underlie the tremendous ecological and morphological diversity of grasses (Preston et al., 2009; Preston and Kellogg, 2006; Wu et al., 2008).

To explore sequence variation among Rho-derived paralogs, we used GENESPACE to build a ploidy-aware syntenic pan-genome annotation among eight species, using the built-in functionality that allows the user to mask primary (likely orthologous) syntenic regions and search for secondary hits (likely paralogous, Figure 4A). This method acts similarly to using an outgroup that predated the WGD (see ‘additional considerations’ methods), but using the same OrthoFinder run as in Figure 3A. Overall, the peptide identity between Rho-derived paralogous regions was much lower than orthologs between species (e.g., S. viridis vs. P. hallii: Wilcoxon W = 88,094,632, p < 1 × 10⁻¹⁶), consistent with the previous discovery that the Rho duplication predated the split among most extant grasses (Ma et al., 2021). However, as has been previously observed (Wang et al., 2011), there is significant variation in the relative similarity of Rho-duplicated sequences. As an example, the peptide sequences of single-copy gene hits in primary syntenic regions (median identity = 90.6%) between chromosome 8 of P. hallii and S. viridis, were 26.9% more similar than the secondary Rho-derived regions (Figure 4B; median identity = 71.4%, Wilcoxon W = 87,842, p < 1 × 10⁻¹⁶). However, S. viridis chromosome 8 contained a single over-retained paralogous region. Unlike all other Rho-derived blocks, the P. hallii paralogs to this 2.7 M bp chromosome 8 region were not significantly less conserved than the primary orthologous region (91.6% vs. 91.9%, W = 14,830, p = 0.13). Outside of this region, the peptide identity of paralogs returned to the genome-wide average (Figure 4B). In line with this observation, the GENESPACE run treating the eight genomes as haploid representations could not distinguish between the Rho-derived paralogs in the over-retained region across all grasses (Figures 3A and 4C), except for of all chromosome pairs between B. distachyon, wheat, and blocks connecting Maize chromosome 10 to sorghum chromosome 5.

Figure 4

Download asset Open asset

Analysis of the grass *Rho* WGD.

(A) BLAST hits between *P. hallii* and *S. viridis* where the target and query genes were in the same orthogroup are plotted and color coded by sequence similarity. Two over-retained regions are highlighted in the red and yellow boxes. (B) The protein identity of *S. viridis* chromosome 8 primary orthologous (blue line) hits against *P. hallii* chromosome 8 and the secondary hits (orange line) against *P. hallii* chromosome 3 demonstrate sequence conservation heterogeneity. The region between the two red vertical lines corresponds to the red-boxed over-retained primary block in panel A. (C) The two boxed regions in panel A were tracked from *P. hallii* chromosomes 3 (red) and 8 (yellow); 50% transparency of the braids means that overlapping regions appear orange.

It is interesting to note that all syntenic over-retained regions were at the extreme termini of the chromosomes outside of maize, B. distachyon and wheat; further, the only genome with complete segregation of the two paralogs, wheat, also retains these regions in the center of all six chromosomes (Figure 4C). These results are consistent with the proposed evolutionary mechanism (Wang et al., 2011) where concerted evolution and ‘illegitimate’ homeologous recombination may have homogenized these paralogous regions. This process would be less effective in pericentromeric regions than the chromosome tails, where a single crossover event would be sufficient to homogenize two paralogous regions.

Conclusions

Combined, the historical abundance of genetic mapping studies and ongoing proliferation of genome resources provides a strong foundation for the integration of comparative and quantitative genomics to accelerate discoveries in evolutionary biology, medicine, and agriculture. The incorporation of synteny and orthology into comparative genomics and quantitative genetics pipelines offers a mechanism to bridge these disparate disciplines. Here, we presented the GENESPACE R package as a framework to help bridge the current gaps between comparative and quantitative genomics, especially in complex evolutionary systems. We hope that the examples presented here will inspire further work to leverage the powerful genome-wide annotations that are coming online, both within and among species.

Materials and methods

GENESPACE pipeline and analysis overview

Request a detailed protocol

GENESPACE operates on gff3-formatted annotation files and accompanying peptide fasta files for primary gene models. There are convenience functions for reformatting the gff and peptide files to simplify the naming scheme and reduce redundant gene models to the primary longest transcript. With these data in hand, GENESPACE calculates BLAST-like hits from DIAMOND2 and runs OrthoFinder (Emms and Kelly, 2019) to infer orthogroups and orthologues. GENESPACE then extracts syntenic regions from the hits using a combination of graph- and cluster-based approaches, producing syntenic orthogroups for each unique (not reciprocal) pair of genomes. Syntenic orthogroups are then collapsed into a pan-genome annotation, which is a matrix of positions against a reference genome (rows) and unique gene models in each syntenic orthogroup for each genome (columns). Detailed step-by-step pipeline methods can be found below.

All analyses were performed in R 4.1.2 on macOS Big Sur 10.16. The following R packages were used for visualization or within GENESPACE v0.9.3 (February 11, 2022 release): data.table v1.14.0 (Dowle and Srinivasan, 2021), dbscan v1.1-8 (Hahsler et al., 2019), igraph v1.2.6 (Csardi and Nepusz, 2006), Biostrings v2.58.0 (Pagès et al., 2020), and rtracklayer v1.50.0 (Lawrence et al., 2009). GENESPACE also calls the following third party software: DIAMOND v2.0.8.146 (Buchfink et al., 2021), OrthoFinder v2.5.4 (Emms and Kelly, 2019), and MCScanX no version installed on October 23, 2021 (Wang et al., 2012). All results were generated programmatically; the accompanying scripts and key output are available on github: jtlovell/GENESPACE_data. Minor adjustments to figures to improve clarity were accomplished in Adobe Illustrator v26.01. A full description of each step in GENESPACE is provided in the documentation that accompanies the package source code on github (jtlovell/GENESPACE).

Description of the vignettes

Request a detailed protocol

Publicly available genome annotations were downloaded on or before October 8, 2021. See Table 2 for data sources, citations, and metadata. All GENESPACE runs used default parameterization, with the following exceptions: (1) the Rho grass run allowed a single secondary hit (default is 0, this is how the paralogs are explicitly searched for) and maximum number of gaps in secondary regions of 10 (default is 5, relaxed to reduce ancient paralogous block splitting), and (2) the maize run used the ‘fast’ OrthoFinder method since all genomes are closely related and haploid. Some maize genomes contained small alternative haplotype scaffolds, which were dropped for all analyses.

Table 2

Raw data sources.

A list of the genomes used in analyses here. Genome version IDs are taken from those posted on the respective data sources and may not reflect the name of the genome in the publication. Where multiple haplotypes are available, only the primary was used for these analyses. All polyploids presented here have only a primary haplotype assembled into chromosomes.

ID	Species	Genome version	Data source	Ploidy*	Reference
garter snake	Thamnophis elegans	rThaEle1.pri	NCBI	1	Rhie et al., 2021
sand lizard	Lacerta_agilis	rLacAgi1.pri	NCBI	1	Rhie et al., 2021
chicken	Gallus gallus	mat.broiler.GRCg7b	NCBI	1	https://www.ncbi.nlm.nih.gov/grc
hummingbird	Calypte anna	bCalAnn1_v1.p	NCBI	1	Rhie et al., 2021
budgie	Melopsittacus undulatus	bMelUnd1.mat.Z	NCBI	1	Unpublished VGP
swan	Cygnus olor	bCygOlo1.pri.v2	NCBI	1	Rhie et al., 2021
zebra finch	Taeniopygia guttata	bTaeGut1.4.pri	NCBI	1	Rhie et al., 2021
echidna	Tachyglossus aculeatus	mTacAcu1.pri	NCBI	1	Zhou et al., 2021
platypus	Ornithorhynchus anatinus	mOrnAna1.pri.v4	NCBI	1	Zhou et al., 2021
brushtail possum	Trichosurus vulpecula	mmTriVul1.pri	NCBI	1	Rhie et al., 2021
opossum	Monodelphis domestica	MonDom5	NCBI	1	Mikkelsen et al., 2007
Tasmanian devil	Sarcophilus harrisii	mSarHar1.11	NCBI	1	Rhie et al., 2021
human (Hg38)	Homo sapiens	GRCh38.p13	NCBI	1	https://www.ncbi.nlm.nih.gov/grc
human (t2t)	Homo sapiens	CHM13-T2T v2.1	NCBI	1	Nurk et al., 2022
chimpanzee	Pan troglodytes	Clint_PTRv2	NCBI	1	Chimpanzee Sequencing and Analysis Consortium, 2005
mouse	Mus musculus	GRCm39	NCBI	1	https://www.ncbi.nlm.nih.gov/grc
dog	Canis lupus familiaris	Dog10K_Boxer_Tasha	NCBI	1	Jagannathan et al., 2021
sloth	Choloepus didactylus	mChoDid1.pri	NCBI	1	Rhie et al., 2021
horseshoe bat	Rhinolophus ferrumequinum	mRhiFer1_v1.p	NCBI	1	Rhie et al., 2021
dolphin	Tursiops truncatus	mTurTru1.mat.Y	NCBI	1	Unpublished VGP
P. hallii	Panicum hallii var. hallii	HAL2_v2.1	Phytozome	1	Lovell et al., 2018
P. hallii (FIL)	Panicum hallii var. filipes	FIL2_v3.1	Phytozome	1	Lovell et al., 2018
switchgrass	Panicum virgatum	AP13_v5.1	Phytozome	2	Lovell et al., 2021b
S viridis	Setaria viridis	v2.1	Phytozome	1	Mamidi et al., 2020
Sorghum	Sorghum bicolor	BTx623_v3.1	Phytozome	1	Paterson et al., 2009
maize	Zea mays	B73_refgen_v5	NCBI	*2	Hufford et al., 2021
rice	Oryza sativa cv ‘kitaake’	kitaake_v2.1	Phytozome	1	Jain et al., 2019
Brachypodium	Brachypodium distachyon	Bd21_v3.1	Phytozome	1	International Brachypodium Initiative, 2010
wheat	Triticum aestivum	V4 (Chinese Spring)	NCBI	3	Zhu et al., 2021
G barbadense	Gossypium barbadense	v1.1	Phytozome	2	Chen et al., 2020
G. darwinii	Gossypium darwinii	v1.1	Phytozome	2	Chen et al., 2020
26 NAM parents	Zea mays	see data on NCBI	NCBI	*1	Hufford et al., 2021

*

Ploidy indicates how the genome was treated in the analyses. All values match the ploidy of the primary assembly haplotype except maize, where the refgen_v5 was treated as diploid (to match both homeologs) in the multispecies run, but as haploid in the nested association mapping (NAM) founder population to track only meiotic homologs across the population. This parameterization is to match the phylogenetic position of the whole-genome duplication (WGD) in the terminal branch of the grass-wide analysis, but ancestral in the 26-NAM analysis.

The publicly available C₃/C₄ gene lists and QTL intervals were generated against the v2 maize assembly. To make this comparable to the across grass and NAM parent GENESPACE runs, we also accomplished a fast GENESPACE run between v2 and the two v5 versions used here. The orthologs and syntenic mapping between these versions are included as text files in the data repository.

All statistics presented here were calculated within R. To compare nonnormal distributions (e.g., sequence identity), we used the nonparametric signed Wilcoxon ranked sum test. To measure sequence divergence, we calculated ungapped percent peptide sequence identity from pairwise Needleman–Wunsch global alignments, implemented in Biostrings (Pagès et al., 2020). To determine single outliers from a unimodal distribution, we applied the Grubbs test implemented in the outliers R package (Komsta, 2011). Some figures were constructed outside of GENESPACE using base R plotting routines and ggplot2 v3.3.3 (Wickham, 2016). Some color palettes were chosen with RColorBrewer (Neuwirth, 2014) and viridis (Garnier et al., 2021).

GENESPACE pipeline: estimating syntenic orthogroups

Request a detailed protocol

GENESPACE relies on high-confidence homologs, which are typically defined as members of the same orthogroup via OrthoFinder. In addition to the original method, which clusters genes and builds a graph from closely related genes based on BLAST scores (Emms and Kelly, 2015), OrthoFinder can also use gene trees to split orthogroups into pairwise orthologs (Emms and Kelly, 2019), which represent a stricter definition of orthology. GENESPACE attempts to merge the benefits of each of these methods by first only considering orthogroups for synteny, which allows users to optionally include paralogs in the scan, then including non-syntenic gene tree-inferred orthologs into the pan-genome annotation during its final steps (Figure 1, see pan-genome pipeline description below). The initial orthogroups are inferred ‘globally’ using pairwise reciprocal hits, either using the default OrthoFinder specification or ‘fast’ by only looking at unique pairs of genomes. Depending on the ploidy and user specifications, orthogroups can be re-estimated within syntenic regions. These three steps are detailed below.

Method 1: default global orthogroups. The default behavior of GENESPACE is to run OrthoFinder using its default parameters. GENESPACE then builds a vector of OrthoFinder geneIDs and their corresponding orthogroup membership. When synteny is called, the GENESPACE-formatted gff text files are read in and merged with OrthoFinder sequenceIDs.

Method 2: fast orthogroup estimation. GENESPACE offers a ‘fast’ orthofinderMethod (Table 3), which performs only one-way DIAMOND2 searches, where the genome annotation with more gene models serves as the query and the smaller annotation is the target. The hits then are mirrored, and each stored as OrthoFinder-formatted BLAST text files. OrthoFinder is then run to the orthogroup step on the precomputed BLAST files. This method results in significant speed improvements with little loss of fidelity among closely related haploid genomes. The user can also specify the sensitivity of DIAMOND2 via the diamondMode parameter during GENESPACE initialization.

Table 3

Comparison of GENESPACE setting performance.

The mirrored ‘fast’ method significantly speeds up OrthoFinder runs by calling DIAMOND2 on each nonredundant pairwise combination of genomes. However, this approach is less sensitive than the default performance and is suggested for only closely related haploid genomes, as the recall of 2:2:2 OGs is less sensitive than the default specification.

	Default OrthoFinder	GENESPACE ‘fast’
n.1:1:1 OGs	22,050	22,444
n.2:2:2 OGs	13,793	13,511
n.tandem arrays	10,597 (4433)	10,599 (4426)
*Run time (min)	59.95	12.45

*

Run time is for ortholog/orthogroup inference (not the GENESPACE pipeline as a whole) using the three cotton genomes, running on 6 2 Gb cores.

Method 3: orthogroups within syntenic blocks. In addition to the above global OrthoFinder runs, GENESPACE can rerun OrthoFinder (only through the -og step) in syntenic regions between pairs of genomes. This is accomplished through four additional steps in the synteny pipeline: (1) syntenic hits are split by large syntenic regions; (2) hits from each syntenic region are passed to OrthoFinder; (3) the within-block orthogroups assignments are aggregated into a single vector across all genomes; and finally, (4) the synteny pipeline (see below) is rerun using the newly defined combined syntenic and inBlkOG vector. While using significant computational resources, this process can improve the sensitivity of orthogroup discovery in paralogous or homeologous regions; however, it is not clear that this offers much improvement over global orthogroups between purely haploid genomes. As such, the default behavior of GENESPACE is to only use orthofinderInBlk when any genome has a ploidy >1.

GENESPACE pipeline: extracting syntenic blast hits

Request a detailed protocol

The GENESPACE function ‘synteny’ accepts several user-defined parameters, which allow for flexibility; however, the defaults are sufficient for most high-quality genomes and evolutionary scenarios. For example, we used the same default parameters for 300 M years of vertebrate evolution, 50 M years and multiple WGDs of grasses, and 10 k years of Maize divergence. For a full list of parameters, see documentation of the GENESPACE function ‘set_syntenyParams’, but here, we will discuss (1) the minimum number of unique hits within a syntenic block (‘blkSize’, default = 5), (2) the maximum number of gaps within a block alignment (‘nGaps’, default = 5), and (3) the radius around a syntenic anchor for a hit to be considered syntenic (‘synBuff’, default = 100). Synteny begins by processing all gene annotations (steps 1–3), then proceeds to process blast hits for each unique pair of genomes (steps 4–7).

Step 1: flag the gff3-formatted annotation. GENESPACE adds the following information to a data matrix that initially contains the name and physical coordinates of each gene for each genome: (1) OrthoFinder IDs (from the sequenceIDs.txt file), (2) gene rank order, and (3) orthogroup IDs.

Step 2: quality control the annotation. Chromosomes with fewer than blkSize genes are dropped so that they will not be used for synteny inference.

Step 3: find and parse tandem arrays. Tandem arrays are defined through six steps: (1) define potential tandem arrays as orthogroups containing >1 gene for each genome-by-chromosome combination; (2) calculate the maximum gene rank-order gap between each adjacent gene in a potential tandem array; (3) split those with a gap between genes > synBuff into separate arrays using dbscan; (4) flag clusters of >1 genes within synBuff as tandem arrays; (5) choose representative genes for each tandem array as the gene closest to the median position of the array and secondarily as the longest peptide sequence; and (6) recalculate ‘arrayOrd’ as the gene rank order of tandem array representative genes.

Step 4: read and process raw blast files. To reduce unnecessary computation, GENESPACE concatenates both reciprocal BLAST hit files so that all unique combinations of query and target genes are represented in a single matrix. For simplicity, the genome with more gene models is treated as the query and the genome with fewer, the target. For each line (‘hit’) in the BLAST file, the following data and statistics are added for both the query (genome1) and target (genome2) gene: (1) positional information, including gene rank order; (2) tandem array representatives; (3) hits where both the query and target are in the same orthogroup; and (4) the relative strength of a hit, where scrRank = 1 represents the single highest bit score blast hit for each query and target gene.

Step 5: define initial syntenic anchors. The following additional parameters are used to define initial syntenic anchors: nHits1 (the top n hits for each gene in the query genome, default ploidy of target genome), nHits2 (the top n hits for each gene in the target genome, default = ploidy of query genome), onlyOgAnchors (logical specifying whether anchors hits must be isOg, default = TRUE), and maskTheseHits (what hits should be masked in the search, default = none; see below on modifications for secondary hits and polyploid self hits). These parameters are applied to find regions that are large and collinear enough to be classified as syntenic through three steps: (1) potential anchor hits are defined as those where both the query and target are array representatives, the hit is not in maskTheseHits, the query scrRank ≤ nHits1, the target scrRank ≤ nHits2, and, optionally if onlyOgAnchors, both the query and target are in the same orthogroup; (2) gene rank order is recalculated for the potential anchor hits; (3) MCScanX_h (-s = blkSize, -m = nGaps) is called from R using condensed gene rank-order positions as the physical location, and those hits within collinear blocks are flagged as initial anchors.

Step 6: clean up initial syntenic anchors. In some cases, syntenic anchors from step 5 can be broken up or do not extend to the ends of syntenic regions. To resolve this, we run five additional block finalization steps: (1) potential anchor hits within synBuff of the initial anchors are extracted and collinear hits are recalled into ‘cleaned’ anchors; (2) the cleaned anchor hits are clustered into blocks via dbscan using synBuff as the radius and blkSize as minPts arguments, dropping blocks with fewer anchors than blkSize and merging nearby syntenic regions; (3) final syntenic anchors are flagged by rerunning MCScanX within each syntenic region; (4) initial block breakpoints are generated by dbscan first with a large radius of synBuff, then on reranked gene order within regions with a radius of blkSize; and (5) overlapping blocks that are not duplicated are split using run-length equivalent decoding. Since the final coding of blocks is conducted within broad syntenic regions, the parameters passed should be robust to variation in ploidy and sequence divergence between genomes.

Step 7: flag anchors, syntenic region hits, and block coordinates. With the final syntenic anchors from step 6 in hand we finalize and annotate the tab-delimited blast hits. We define two sets of syntenic qualifications: blocks, which are fine grained runs of completely collinear hits, and ‘regions’, which are clustered blocks that average across minor inversions. Blocks are defined from the step 6 anchors. Anchor hits are reclustered via dbscan with a radius of synBuff and each cluster between each pair of genomes and chromosomes is assigned a ‘regID’. All hits within synBuff of an anchor hit are flagged as ‘inBuffer’. The coordinates of blocks and regions are calculated from the bounding anchors for each syntenic block.

Modifications: steps 5–7 are modified depending on whether the BLAST hit file is intra- or inter-genomic, the ploidy of the query and target genome, and the user-defined number of ‘secondary hits’. These modifications are as follows. (1) step 5–7(1): [if necessary] syntenic regions are called for self blast hits within haploid genomes. This is the simplest case where the steps are ignored and syntenic anchors are defined as self hits between tandem array representatives. Block and region IDs are chromosome IDs. (2) Modification step 5–7(2): [if necessary] syntenic regions are called for self blast hits within genomes with ploidy >1. Here, the modified syntenic hits from 5-7(1) are specified in maskTheseHits with a synBuff radius of 500 genes. This excludes potentially problematic large tandem arrays in some genomes. The number of hits after masking is set to ploidy – 1. Then the standard step 5–7 processes are run. (3) Modification step 5–7(3): [if necessary] syntenic regions are called for blast hits where nSecondaryHits >0. Here, the methods for syntenic hit calculation from 5-7(2) (if intragenomic) or 5-7 (as a mask of the rerun if intergenomic) are conducted except that the parameters specified are given as those ending in ‘Second’.

GENESPACE pipeline: constructing pan-genome annotations

Request a detailed protocol

The GENESPACE function ‘pangenome’ decodes pairwise syntenic orthologs into a multigenome pan-annotation. The output is a long-formatted text file, where each gene is given a reference genome syntenic position and chromosome, and flags that are described below. In addition, pangenome also returns a wide-formatted data.table (R object) where each row is a pan-genome entry with positional information and each column is a list of genes by genome (e.g., Figure 1).

Step 1: build a reference-anchored map of syntenic hits. The chromosome and positions of the tandem array representatives from the reference genome annotation form the foundation of mapping across all genomes. Upon building this positional backbone, syntenic hits are pulled for each genome and placed against the anchor genes that are in the same orthogroup.

Step 2: interpolate syntenic reference positions. In most pairwise comparisons between genomes, some syntenic orthogroups will be missing in the reference. Since we want to extract PAV by physical coordinates, all syntenic orthogroups, even those that do not include a reference genome gene model, need to have reference positional information. To fill this gap, GENESPACE interpolates the syntenic reference position of all array representative genes in all genomes through a three step pipeline: (1) subset the syntenic anchor hits to ungapped collinear hits following a 1:1 synteny mapping ratio (i.e., perfect diagonal of hits by gene rank order); (2) cluster the 1:1 syntenic anchors so that any jump of >1 gene is split into its own cluster; (3) fill in missing positions through linear interpolation.

Step 3: determine the reference mapping positions of each syntenic orthogroup. In the case where the reference genome is purely haploid with no segmental duplicates, this is a straightforward step: the orthogroups that contain reference genome genes are placed at that backbone position, and those without a reference gene are clustered into the most likely interpolated placement. However, we do not want to rely on single-copy references, so GENESPACE allows multiple placements not only for each nonreference gene, but also for genes in the reference genome itself. For example, in a polyploid, we would want to know if there is PAV between two homeologous chromosomes. To do this, we take a three-step approach: (1) the interpolated (or actual reference-anchored) positions of genes in each syntenic orthogroup are extracted and split by interpolated chromosome and maximum gap between any two inferred positions >synBuff; (2) unsplit orthogroups are set aside; (3) split orthogroups are checked and retained if they contain ≥ propAssignThresh (default = 0.5) proportion of genes in the syntenic orthogroup with an interpolated position in that cluster; (4) clustered positions are dropped if there are more than maxPlacementsPerRefChr (default = 2) positional placements for that syntenic orthogroup, retaining the top clusters ranked by propAssignThresh; (5) the culled positional clusters are merged with those from step (2) to create the initial pan-genome annotation.

Step 4: add and flag other forms of orthogroups. The reference pan-genome built in step 3 only acts on direct syntenic blast hits (edges), which allows for strict construction of interpolated positions without the potential polluting positional effects of orthogroups in minor or miss-assembled syntenic blocks. GENESPACE fixes these and other complexities of syntenic orthogroups through a four step pipeline to finalize the pan-genome annotation: (1) ‘indirect syntenic’ orthogroup members are added back into the initial pan-genome annotation by parsing the gff-like text file that contains a vector of syntenic orthogroups; (2) syntenic orthogroups that are missing all interpolated positions are added and given NA positions; (3) tandem array members that are not representative genes are added into the pan-genome annotation; (4) if available, orthologues from the initial OrthoFinder run that are not present in the pan-genome annotation are added and flagged. Note that if OrthoFinder was run using the ‘fast’ orthofinderMethod, orthologs will not be produced nor added into the pan-genome annotation.

Additional considerations for comparative genomics parameterization

Request a detailed protocol

There are several factors to consider when constructing your GENESPACE run, generating syntenic block breakpoints, or looking at comparative genomics through protein similarity estimates like OrthoFinder. We detail two of these below:

(1) Outgroups and the phylogenetic context of orthology inference. OrthoFinder defines orthogroups as the set of genes that are descended from a single gene in the last common ancestor of all the species being considered. As such, the scale of the run matters, often significantly. For example, an orthogroup would not be likely to contain homeologs across the two ancient subgenomes for a run that included only two maize genomes. Since the coalescence of any two maize genotypes occurred more recently than the ~12M ya WGD, few homeologs would both be descended from the same common ancestor when considering only maize genotypes. Hence, the within-maize NAM parent run (Figure 3D) excludes homeologs. However, if an outgroup to maize was included in the run, both maize homeologs would be likely to show common ancestry to a single gene in the outgroup, thus connecting the maize homeologs into a single orthogroup. Hence, both maize homeologous regions are present in the across grasses synteny graph (Figure 3A) despite using identical synteny parameters to the maize NAM parent run. Given the potentially significant role of outgroups on the results of the global run, GENESPACE offers an ‘outgroup’ parameter, which specifies the genomes that should be included in the orthofinder run but excluded from all downstream analyses.

(2) Studying homoelogs in polyploids and other paralogs. Non-homeologous paralogs are typically excluded in the default GENESPACE parameterization. This is in part because GENESPACE was originally designed for work with plants, and most plant lineages have undergone one or more WGDs. However, there are two ways to study paralogs by altering GENESPACE parameters: (1) specify an ‘outgroup’ genome (see above), which is only used for the global OrthoFinder run and set the genomes’ ploidy as that expected by the number of duplications relative to the outgroup; (2) if an outgroup is not available (or too distantly related to be of use), specify the synteny parameter ‘nSecondaryHits’ as the number of paralogous copies per genome. In the second case, secondaryHits are inferred after masking out the normal syntenic hits, then synteny is rerun without requiring anchors to be global orthogroups. In both cases, it will be far better to set orthofinderInBlk to TRUE, so that pairs of genomes are considered and genes without a solid hit in the outgroup are not excluded.

Data availability

Raw data were sourced entirely from NCBI and Phytozome. Processed data, intermediate files, scripts, plots, and source data are all available in the data repository: https://github.com/jtlovell/GENESPACE_data (copy archived at swh:1:rev:77612f8c59fbfd43ef3f4c1719933bf0cbca3261). All source code and documentation for the GENESPACE R package can be found at https://github.com/jtlovell/GENESPACE (copy archieved at swh:1:rev:390341499ee1d2ccd5e1a894c4bd7c1bd20a3dda). An interactive viewer for the plant genomes can be found on phytozome at https://phytozome-next.jgi.doe.gov/tools/dotplot/synteny.html.

References

(2019) Components of the ribosome biogenesis pathway underlie establishment of telomere length set point in Arabidopsis
Nature Communications 10:5479.

https://doi.org/10.1038/s41467-019-13448-z
- PubMed
- Google Scholar
1. Barker MS
2. Arrigo N
3. Baniaga AE
4. Li Z
5. Levin DA
(2016) On the relative abundance of autopolyploids and allopolyploids
The New Phytologist 210:391–398.

https://doi.org/10.1111/nph.13698
- PubMed
- Google Scholar
(2021) Sensitive protein alignments at tree-of-life scale using DIAMOND
Nature Methods 18:366–368.

https://doi.org/10.1038/s41592-021-01101-x
- PubMed
- Google Scholar
1. Chen ZJ
2. Sreedasyam A
3. Ando A
4. Song Q
5. De Santiago LM
6. Hulse-Kemp AM
7. Ding M
8. Ye W
9. Kirkbride RC
10. Jenkins J
11. Plott C
12. Lovell J
13. Lin YM
14. Vaughn R
15. Liu B
16. Simpson S
17. Scheffler BE
18. Wen L
19. Saski CA
20. Grover CE
21. Hu G
22. Conover JL
23. Carlson JW
24. Shu S
25. Boston LB
26. Williams M
27. Peterson DG
28. McGee K
29. Jones DC
30. Wendel JF
31. Stelly DM
32. Grimwood J
33. Schmutz J
(2020) Genomic diversifications of five Gossypium allopolyploid species and their impact on cotton improvement
Nature Genetics 52:525–533.

https://doi.org/10.1038/s41588-020-0614-5
- PubMed
- Google Scholar
1. Chimpanzee Sequencing and Analysis Consortium
(2005) Initial sequence of the chimpanzee genome and comparison with the human genome
Nature 437:69–87.

https://doi.org/10.1038/nature04072
- PubMed
- Google Scholar
1. Csardi G
2. Nepusz T
(2006)
The igraph software package for complex network research

InterJournal, Complex Systems 1695:1–9.
- Google Scholar
1. Ding Z
2. Weissmann S
3. Wang M
4. Du B
5. Huang L
6. Wang L
7. Tu X
8. Zhong S
9. Myers C
10. Brutnell TP
11. Sun Q
12. Li P
(2015) Identification of photosynthesis-associated C4 candidate genes through comparative leaf gradient transcriptome in multiple lineages of C3 and C4 species
PLOS ONE 10:e0140629.

https://doi.org/10.1371/journal.pone.0140629
- PubMed
- Google Scholar
Software
1. Dowle M
2. Srinivasan A
(2021) Data.table: extension of “data.frame.”, version 1.14.0
CRAN.

https://cran.r-project.org/web/packages/data.table/index.html
(2020) Phylogenetic reconstruction based on synteny block and gene adjacencies
Molecular Biology and Evolution 37:2747–2762.

https://doi.org/10.1093/molbev/msaa114
- PubMed
- Google Scholar
1. Emms DM
2. Kelly S
(2015) OrthoFinder: solving fundamental biases in whole genome comparisons dramatically improves orthogroup inference accuracy
Genome Biology 16:157.

https://doi.org/10.1186/s13059-015-0721-2
- PubMed
- Google Scholar
1. Emms DM
2. Kelly S
(2019) OrthoFinder: phylogenetic orthology inference for comparative genomics
Genome Biology 20:238.

https://doi.org/10.1186/s13059-019-1832-y
- PubMed
- Google Scholar
Software
1. Garnier S
2. Ross N
3. Rudis R
4. Camargo PA
5. Sciaini M
6. Scherer C
(2021) Viridis - Colorblind-Friendly Color Maps for R, version 0.6.2
Zenodo.

https://doi.org/10.5281/zenodo.4679424
1. Gaut BS
2. Doebley JF
(1997) DNA sequence evidence for the segmental allotetraploid origin of maize
PNAS 94:6809–6814.

https://doi.org/10.1073/pnas.94.13.6809
- PubMed
- Google Scholar
1. Goodstein DM
2. Shu S
3. Howson R
4. Neupane R
5. Hayes RD
6. Fazo J
7. Mitros T
8. Dirks W
9. Hellsten U
10. Putnam N
11. Rokhsar DS
(2012) Phytozome: a comparative platform for green plant genomics
Nucleic Acids Research 40:D1178–D1186.

https://doi.org/10.1093/nar/gkr944
- PubMed
- Google Scholar
(2019) Domestication and crop evolution of wheat and barley: Genes, genomics, and future directions
Journal of Integrative Plant Biology 61:204–225.

https://doi.org/10.1111/jipb.12737
- PubMed
- Google Scholar
(2019) dbscan: Fast density-based clustering with R
Journal of Statistical Software 91:i01.

https://doi.org/10.18637/jss.v091.i01
- Google Scholar
(2017) SynMap2 and SynMap3D: web-based whole-genome synteny browsers
Bioinformatics 33:2197–2198.

https://doi.org/10.1093/bioinformatics/btx144
- PubMed
- Google Scholar
1. Hufford MB
2. Seetharam AS
3. Woodhouse MR
4. Chougule KM
5. Ou S
6. Liu J
7. Ricci WA
8. Guo T
9. Olson A
10. Qiu Y
11. Della Coletta R
12. Tittes S
13. Hudson AI
14. Marand AP
15. Wei S
16. Lu Z
17. Wang B
18. Tello-Ruiz MK
19. Piri RD
20. Wang N
21. Kim DW
22. Zeng Y
23. O’Connor CH
24. Li X
25. Gilbert AM
26. Baggs E
27. Krasileva KV
28. Portwood JL
29. Cannon EKS
30. Andorf CM
31. Manchanda N
32. Snodgrass SJ
33. Hufnagel DE
34. Jiang Q
35. Pedersen S
36. Syring ML
37. Kudrna DA
38. Llaca V
39. Fengler K
40. Schmitz RJ
41. Ross-Ibarra J
42. Yu J
43. Gent JI
44. Hirsch CN
45. Ware D
46. Dawe RK
(2021) De novo assembly, annotation, and comparative analysis of 26 diverse maize genomes
Science 373:655–662.

https://doi.org/10.1126/science.abg5289
- PubMed
- Google Scholar
1. International Brachypodium Initiative
(2010) Genome sequencing and analysis of the model grass Brachypodium distachyon
Nature 463:763–768.

https://doi.org/10.1038/nature08747
- PubMed
- Google Scholar
1. Jagannathan V
2. Hitte C
3. Kidd JM
4. Masterson P
5. Murphy TD
6. Emery S
7. Davis B
8. Buckley RM
9. Liu YH
10. Zhang XQ
11. Leeb T
12. Zhang YP
13. Ostrander EA
14. Wang GD
(2021) Dog10K_Boxer_Tasha_1.0: A long-read assembly of the dog reference genome
Genes 12:847.

https://doi.org/10.3390/genes12060847
- PubMed
- Google Scholar
1. Jain R
2. Jenkins J
3. Shu S
4. Chern M
5. Martin JA
6. Copetti D
7. Duong PQ
8. Pham NT
9. Kudrna DA
10. Talag J
11. Schackwitz WS
12. Lipzen AM
13. Dilworth D
14. Bauer D
15. Grimwood J
16. Nelson CR
17. Xing F
18. Xie W
19. Barry KW
20. Wing RA
21. Schmutz J
22. Li G
23. Ronald PC
(2019) Genome sequence of the model rice variety KitaakeX
BMC Genomics 20:905.

https://doi.org/10.1186/s12864-019-6262-4
- PubMed
- Google Scholar
1. Jiao Y
2. Li J
3. Tang H
4. Paterson AH
(2014) Integrated syntenic and phylogenomic analyses reveal an ancient genome duplication in monocots
The Plant Cell 26:2792–2802.

https://doi.org/10.1105/tpc.114.127597
- PubMed
- Google Scholar
Software
1. Komsta L
(2011) Outliers: tests for outliers
CRAN.

https://cran.r-project.org/web/packages/outliers/index.html
(2021) Sex chromosome evolution among amniotes: is the origin of sex chromosomes non-random?
Philosophical Transactions of the Royal Society of London. Series B, Biological Sciences 376:20200108.

https://doi.org/10.1098/rstb.2020.0108
- PubMed
- Google Scholar
(2009) rtracklayer: an R package for interfacing with genome browsers
Bioinformatics 25:1841–1842.

https://doi.org/10.1093/bioinformatics/btp328
- PubMed
- Google Scholar
1. Li C
2. Sun B
3. Li Y
4. Liu C
5. Wu X
6. Zhang D
7. Shi Y
8. Song Y
9. Buckler ES
10. Zhang Z
11. Wang T
12. Li Y
(2016) Numerous genetic loci identified for drought tolerance in the maize nested association mapping populations
BMC Genomics 17:894.

https://doi.org/10.1186/s12864-016-3170-8
- PubMed
- Google Scholar
1. Lovell JT
2. Jenkins J
3. Lowry DB
4. Mamidi S
5. Sreedasyam A
6. Weng X
7. Barry K
8. Bonnette J
9. Campitelli B
10. Daum C
11. Gordon SP
12. Gould BA
13. Khasanova A
14. Lipzen A
15. MacQueen A
16. Palacio-Mejía JD
17. Plott C
18. Shakirov EV
19. Shu S
20. Yoshinaga Y
21. Zane M
22. Kudrna D
23. Talag JD
24. Rokhsar D
25. Grimwood J
26. Schmutz J
27. Juenger TE
(2018) The genomic landscape of molecular responses to natural drought stress in Panicum hallii
Nature Communications 9:5213.

https://doi.org/10.1038/s41467-018-07669-x
- PubMed
- Google Scholar
1. Lovell JT
2. Bentley NB
3. Bhattarai G
4. Jenkins JW
5. Sreedasyam A
6. Alarcon Y
7. Bock C
8. Boston LB
9. Carlson J
10. Cervantes K
11. Clermont K
12. Duke S
13. Krom N
14. Kubenka K
15. Mamidi S
16. Mattison CP
17. Monteros MJ
18. Pisani C
19. Plott C
20. Rajasekar S
21. Rhein HS
22. Rohla C
23. Song M
24. Hilaire RS
25. Shu S
26. Wells L
27. Webber J
28. Heerema RJ
29. Klein PE
30. Conner P
31. Wang X
32. Grauke LJ
33. Grimwood J
34. Schmutz J
35. Randall JJ
(2021a) Four chromosome scale genomes and a pan-genome annotation to accelerate pecan tree breeding
Nature Communications 12:4125.

https://doi.org/10.1038/s41467-021-24328-w
- PubMed
- Google Scholar
1. Lovell JT
2. MacQueen AH
3. Mamidi S
4. Bonnette J
5. Jenkins J
6. Napier JD
7. Sreedasyam A
8. Healey A
9. Session A
10. Shu S
11. Barry K
12. Bonos S
13. Boston L
14. Daum C
15. Deshpande S
16. Ewing A
17. Grabowski PP
18. Haque T
19. Harrison M
20. Jiang J
21. Kudrna D
22. Lipzen A
23. Pendergast TH
24. Plott C
25. Qi P
26. Saski CA
27. Shakirov EV
28. Sims D
29. Sharma M
30. Sharma R
31. Stewart A
32. Singan VR
33. Tang Y
34. Thibivillier S
35. Webber J
36. Weng X
37. Williams M
38. Wu GA
39. Yoshinaga Y
40. Zane M
41. Zhang L
42. Zhang J
43. Behrman KD
44. Boe AR
45. Fay PA
46. Fritschi FB
47. Jastrow JD
48. Lloyd-Reilley J
49. Martínez-Reyna JM
50. Matamala R
51. Mitchell RB
52. Rouquette FM
53. Ronald P
54. Saha M
55. Tobias CM
56. Udvardi M
57. Wing RA
58. Wu Y
59. Bartley LE
60. Casler M
61. Devos KM
62. Lowry DB
63. Rokhsar DS
64. Grimwood J
65. Juenger TE
66. Schmutz J
(2021b) Genomic mechanisms of climate adaptation in polyploid bioenergy switchgrass
Nature 590:438–444.

https://doi.org/10.1038/s41586-020-03127-1
- PubMed
- Google Scholar
1. Ma PF
2. Liu YL
3. Jin GH
4. Liu JX
5. Wu H
6. He J
7. Guo ZH
8. Li DZ
(2021) The Pharus latifolius genome bridges the gap of early grass evolution
The Plant Cell 33:846–864.

https://doi.org/10.1093/plcell/koab015
- PubMed
- Google Scholar
1. Mamidi S
2. Healey A
3. Huang P
4. Grimwood J
5. Jenkins J
6. Barry K
7. Sreedasyam A
8. Shu S
9. Lovell JT
10. Feldman M
11. Wu J
12. Yu Y
13. Chen C
14. Johnson J
15. Sakakibara H
16. Kiba T
17. Sakurai T
18. Tavares R
19. Nusinow DA
20. Baxter I
21. Schmutz J
22. Brutnell TP
23. Kellogg EA
(2020) A genome resource for green millet Setaria viridis enables discovery of agronomically valuable loci
Nature Biotechnology 38:1203–1210.

https://doi.org/10.1038/s41587-020-0681-2
- PubMed
- Google Scholar
1. Mikkelsen TS
2. Wakefield MJ
3. Aken B
4. Amemiya CT
5. Chang JL
6. Duke S
7. Garber M
8. Gentles AJ
9. Goodstadt L
10. Heger A
11. Jurka J
12. Kamal M
13. Mauceli E
14. Searle SMJ
15. Sharpe T
16. Baker ML
17. Batzer MA
18. Benos PV
19. Belov K
20. Clamp M
21. Cook A
22. Cuff J
23. Das R
24. Davidow L
25. Deakin JE
26. Fazzari MJ
27. Glass JL
28. Grabherr M
29. Greally JM
30. Gu W
31. Hore TA
32. Huttley GA
33. Kleber M
34. Jirtle RL
35. Koina E
36. Lee JT
37. Mahony S
38. Marra MA
39. Miller RD
40. Nicholls RD
41. Oda M
42. Papenfuss AT
43. Parra ZE
44. Pollock DD
45. Ray DA
46. Schein JE
47. Speed TP
48. Thompson K
49. VandeBerg JL
50. Wade CM
51. Walker JA
52. Waters PD
53. Webber C
54. Weidman JR
55. Xie X
56. Zody MC
57. Broad Institute Genome Sequencing Platform
58. Broad Institute Whole Genome Assembly Team
59. Graves JAM
60. Ponting CP
61. Breen M
62. Samollow PB
63. Lander ES
64. Lindblad-Toh K
(2007) Genome of the marsupial Monodelphis domestica reveals innovation in non-coding sequences
Nature 447:167–177.

https://doi.org/10.1038/nature05805
- PubMed
- Google Scholar
1. Muller HJ
(1925) Why polyploidy is rarer in animals than in plawhy polyploidy is rarer in animals than in plants
The American Naturalist 59:346–353.

https://doi.org/10.1086/280047
- Google Scholar
(1999) Extensive conservation of sex chromosome organization between cat and human revealed by parallel radiation hybrid mapping
Genome Research 9:1223–1230.

https://doi.org/10.1101/gr.9.12.1223
- PubMed
- Google Scholar
Software
1. Neuwirth E
(2014) RColorBrewer: colorbrewer palettes
CRAN.

https://cran.r-project.org/web/packages/RColorBrewer/index.html
1. Nurk S
2. Koren S
3. Rhie A
4. Rautiainen M
5. Bzikadze AV
6. Mikheenko A
7. Vollger MR
8. Altemose N
9. Uralsky L
10. Gershman A
11. Aganezov S
12. Hoyt SJ
13. Diekhans M
14. Logsdon GA
15. Alonge M
16. Antonarakis SE
17. Borchers M
18. Bouffard GG
19. Brooks SY
20. Caldas GV
21. Chen N-C
22. Cheng H
23. Chin C-S
24. Chow W
25. de Lima LG
26. Dishuck PC
27. Durbin R
28. Dvorkina T
29. Fiddes IT
30. Formenti G
31. Fulton RS
32. Fungtammasan A
33. Garrison E
34. Grady PGS
35. Graves-Lindsay TA
36. Hall IM
37. Hansen NF
38. Hartley GA
39. Haukness M
40. Howe K
41. Hunkapiller MW
42. Jain C
43. Jain M
44. Jarvis ED
45. Kerpedjiev P
46. Kirsche M
47. Kolmogorov M
48. Korlach J
49. Kremitzki M
50. Li H
51. Maduro VV
52. Marschall T
53. McCartney AM
54. McDaniel J
55. Miller DE
56. Mullikin JC
57. Myers EW
58. Olson ND
59. Paten B
60. Peluso P
61. Pevzner PA
62. Porubsky D
63. Potapova T
64. Rogaev EI
65. Rosenfeld JA
66. Salzberg SL
67. Schneider VA
68. Sedlazeck FJ
69. Shafin K
70. Shew CJ
71. Shumate A
72. Sims Y
73. Smit AFA
74. Soto DC
75. Sović I
76. Storer JM
77. Streets A
78. Sullivan BA
79. Thibaud-Nissen F
80. Torrance J
81. Wagner J
82. Walenz BP
83. Wenger A
84. Wood JMD
85. Xiao C
86. Yan SM
87. Young AC
88. Zarate S
89. Surti U
90. McCoy RC
91. Dennis MY
92. Alexandrov IA
93. Gerton JL
94. O’Neill RJ
95. Timp W
96. Zook JM
97. Schatz MC
98. Eichler EE
99. Miga KH
100. Phillippy AM
(2022) The complete sequence of a human genome
Science 376:44–53.

https://doi.org/10.1126/science.abj6987
- PubMed
- Google Scholar
1. One Thousand Plant Transcriptomes Initiative
(2019) One thousand plant transcriptomes and the phylogenomics of green plants
Nature 574:679–685.

https://doi.org/10.1038/s41586-019-1693-2
- PubMed
- Google Scholar
Software
(2020) Biostrings: efficient manipulation of biological strings, version 2.58.0
Bioconductor.

https://bioconductor.org/packages/Biostrings/
1. Paterson AH
2. Bowers JE
3. Bruggmann R
4. Dubchak I
5. Grimwood J
6. Gundlach H
7. Haberer G
8. Hellsten U
9. Mitros T
10. Poliakov A
11. Schmutz J
12. Spannagl M
13. Tang H
14. Wang X
15. Wicker T
16. Bharti AK
17. Chapman J
18. Feltus FA
19. Gowik U
20. Grigoriev IV
21. Lyons E
22. Maher CA
23. Martis M
24. Narechania A
25. Otillar RP
26. Penning BW
27. Salamov AA
28. Wang Y
29. Zhang L
30. Carpita NC
31. Freeling M
32. Gingle AR
33. Hash CT
34. Keller B
35. Klein P
36. Kresovich S
37. McCann MC
38. Ming R
39. Peterson DG
40. Mehboob-ur-Rahman R
41. Ware D
42. Westhoff P
43. Mayer KFX
44. Messing J
45. Rokhsar DS
(2009) The Sorghum bicolor genome and the diversification of grasses
Nature 457:551–556.

https://doi.org/10.1038/nature07723
- PubMed
- Google Scholar
1. Preston JC
2. Kellogg EA
(2006) Reconstructing the evolutionary history of paralogous APETALA1/FRUITFULL-like genes in grasses (Poaceae)
Genetics 174:421–437.

https://doi.org/10.1534/genetics.106.057125
- PubMed
- Google Scholar
(2009) MADS-box gene expression and implications for developmental origins of the grass spikelet
American Journal of Botany 96:1419–1429.

https://doi.org/10.3732/ajb.0900062
- PubMed
- Google Scholar
(2007) The multiple sex chromosomes of platypus and echidna are not completely identical and several share homology with the avian Z
Genome Biology 8:R243.

https://doi.org/10.1186/gb-2007-8-11-r243
- Google Scholar
1. Rhie A
2. McCarthy SA
3. Fedrigo O
4. Damas J
5. Formenti G
6. Koren S
7. Uliano-Silva M
8. Chow W
9. Fungtammasan A
10. Kim J
11. Lee C
12. Ko BJ
13. Chaisson M
14. Gedman GL
15. Cantin LJ
16. Thibaud-Nissen F
17. Haggerty L
18. Bista I
19. Smith M
20. Haase B
21. Mountcastle J
22. Winkler S
23. Paez S
24. Howard J
25. Vernes SC
26. Lama TM
27. Grutzner F
28. Warren WC
29. Balakrishnan CN
30. Burt D
31. George JM
32. Biegler MT
33. Iorns D
34. Digby A
35. Eason D
36. Robertson B
37. Edwards T
38. Wilkinson M
39. Turner G
40. Meyer A
41. Kautt AF
42. Franchini P
43. Detrich HW
44. Svardal H
45. Wagner M
46. Naylor GJP
47. Pippel M
48. Malinsky M
49. Mooney M
50. Simbirsky M
51. Hannigan BT
52. Pesout T
53. Houck M
54. Misuraca A
55. Kingan SB
56. Hall R
57. Kronenberg Z
58. Sović I
59. Dunn C
60. Ning Z
61. Hastie A
62. Lee J
63. Selvaraj S
64. Green RE
65. Putnam NH
66. Gut I
67. Ghurye J
68. Garrison E
69. Sims Y
70. Collins J
71. Pelan S
72. Torrance J
73. Tracey A
74. Wood J
75. Dagnew RE
76. Guan D
77. London SE
78. Clayton DF
79. Mello CV
80. Friedrich SR
81. Lovell PV
82. Osipova E
83. Al-Ajli FO
84. Secomandi S
85. Kim H
86. Theofanopoulou C
87. Hiller M
88. Zhou Y
89. Harris RS
90. Makova KD
91. Medvedev P
92. Hoffman J
93. Masterson P
94. Clark K
95. Martin F
96. Howe K
97. Flicek P
98. Walenz BP
99. Kwak W
100. Clawson H
101. Diekhans M
102. Nassar L
103. Paten B
104. Kraus RHS
105. Crawford AJ
106. Gilbert MTP
107. Zhang G
108. Venkatesh B
109. Murphy RW
110. Koepfli K-P
111. Shapiro B
112. Johnson WE
113. Di Palma F
114. Marques-Bonet T
115. Teeling EC
116. Warnow T
117. Graves JM
118. Ryder OA
119. Haussler D
120. O’Brien SJ
121. Korlach J
122. Lewin HA
123. Howe K
124. Myers EW
125. Durbin R
126. Phillippy AM
127. Jarvis ED
(2021) Towards complete and error-free genome assemblies of all vertebrate species
Nature 592:737–746.

https://doi.org/10.1038/s41586-021-03451-0
- PubMed
- Google Scholar
1. Ross MT
2. Grafham DV
3. Coffey AJ
4. Scherer S
5. McLay K
6. Muzny D
7. Platzer M
8. Howell GR
9. Burrows C
10. Bird CP
11. Frankish A
12. Lovell FL
13. Howe KL
14. Ashurst JL
15. Fulton RS
16. Sudbrak R
17. Wen G
18. Jones MC
19. Hurles ME
20. Andrews TD
21. Scott CE
22. Searle S
23. Ramser J
24. Whittaker A
25. Deadman R
26. Carter NP
27. Hunt SE
28. Chen R
29. Cree A
30. Gunaratne P
31. Havlak P
32. Hodgson A
33. Metzker ML
34. Richards S
35. Scott G
36. Steffen D
37. Sodergren E
38. Wheeler DA
39. Worley KC
40. Ainscough R
41. Ambrose KD
42. Ansari-Lari MA
43. Aradhya S
44. Ashwell RIS
45. Babbage AK
46. Bagguley CL
47. Ballabio A
48. Banerjee R
49. Barker GE
50. Barlow KF
51. Barrett IP
52. Bates KN
53. Beare DM
54. Beasley H
55. Beasley O
56. Beck A
57. Bethel G
58. Blechschmidt K
59. Brady N
60. Bray-Allen S
61. Bridgeman AM
62. Brown AJ
63. Brown MJ
64. Bonnin D
65. Bruford EA
66. Buhay C
67. Burch P
68. Burford D
69. Burgess J
70. Burrill W
71. Burton J
72. Bye JM
73. Carder C
74. Carrel L
75. Chako J
76. Chapman JC
77. Chavez D
78. Chen E
79. Chen G
80. Chen Y
81. Chen Z
82. Chinault C
83. Ciccodicola A
84. Clark SY
85. Clarke G
86. Clee CM
87. Clegg S
88. Clerc-Blankenburg K
89. Clifford K
90. Cobley V
91. Cole CG
92. Conquer JS
93. Corby N
94. Connor RE
95. David R
96. Davies J
97. Davis C
98. Davis J
99. Delgado O
100. Deshazo D
101. Dhami P
102. Ding Y
103. Dinh H
104. Dodsworth S
105. Draper H
106. Dugan-Rocha S
107. Dunham A
108. Dunn M
109. Durbin KJ
110. Dutta I
111. Eades T
112. Ellwood M
113. Emery-Cohen A
114. Errington H
115. Evans KL
116. Faulkner L
117. Francis F
118. Frankland J
119. Fraser AE
120. Galgoczy P
121. Gilbert J
122. Gill R
123. Glöckner G
124. Gregory SG
125. Gribble S
126. Griffiths C
127. Grocock R
128. Gu Y
129. Gwilliam R
130. Hamilton C
131. Hart EA
132. Hawes A
133. Heath PD
134. Heitmann K
135. Hennig S
136. Hernandez J
137. Hinzmann B
138. Ho S
139. Hoffs M
140. Howden PJ
141. Huckle EJ
142. Hume J
143. Hunt PJ
144. Hunt AR
145. Isherwood J
146. Jacob L
147. Johnson D
148. Jones S
149. de Jong PJ
150. Joseph SS
151. Keenan S
152. Kelly S
153. Kershaw JK
154. Khan Z
155. Kioschis P
156. Klages S
157. Knights AJ
158. Kosiura A
159. Kovar-Smith C
160. Laird GK
161. Langford C
162. Lawlor S
163. Leversha M
164. Lewis L
165. Liu W
166. Lloyd C
167. Lloyd DM
168. Loulseged H
169. Loveland JE
170. Lovell JD
171. Lozado R
172. Lu J
173. Lyne R
174. Ma J
175. Maheshwari M
176. Matthews LH
177. McDowall J
178. McLaren S
179. McMurray A
180. Meidl P
181. Meitinger T
182. Milne S
183. Miner G
184. Mistry SL
185. Morgan M
186. Morris S
187. Müller I
188. Mullikin JC
189. Nguyen N
190. Nordsiek G
191. Nyakatura G
192. O’Dell CN
193. Okwuonu G
194. Palmer S
195. Pandian R
196. Parker D
197. Parrish J
198. Pasternak S
199. Patel D
200. Pearce AV
201. Pearson DM
202. Pelan SE
203. Perez L
204. Porter KM
205. Ramsey Y
206. Reichwald K
207. Rhodes S
208. Ridler KA
209. Schlessinger D
210. Schueler MG
211. Sehra HK
212. Shaw-Smith C
213. Shen H
214. Sheridan EM
215. Shownkeen R
216. Skuce CD
217. Smith ML
218. Sotheran EC
219. Steingruber HE
220. Steward CA
221. Storey R
222. Swann RM
223. Swarbreck D
224. Tabor PE
225. Taudien S
226. Taylor T
227. Teague B
228. Thomas K
229. Thorpe A
230. Timms K
231. Tracey A
232. Trevanion S
233. Tromans AC
234. d’Urso M
235. Verduzco D
236. Villasana D
237. Waldron L
238. Wall M
239. Wang Q
240. Warren J
241. Warry GL
242. Wei X
243. West A
244. Whitehead SL
245. Whiteley MN
246. Wilkinson JE
247. Willey DL
248. Williams G
249. Williams L
250. Williamson A
251. Williamson H
252. Wilming L
253. Woodmansey RL
254. Wray PW
255. Yen J
256. Zhang J
257. Zhou J
258. Zoghbi H
259. Zorilla S
260. Buck D
261. Reinhardt R
262. Poustka A
263. Rosenthal A
264. Lehrach H
265. Meindl A
266. Minx PJ
267. Hillier LW
268. Willard HF
269. Wilson RK
270. Waterston RH
271. Rice CM
272. Vaudin M
273. Coulson A
274. Nelson DL
275. Weinstock G
276. Sulston JE
277. Durbin R
278. Hubbard T
279. Gibbs RA
280. Beck S
281. Rogers J
282. Bentley DR
(2005) The DNA sequence of the human X chromosome
Nature 434:325–337.

https://doi.org/10.1038/nature03440
- PubMed
- Google Scholar
1. Simakov O
2. Marlétaz F
3. Yue J-X
4. O’Connell B
5. Jenkins J
6. Brandt A
7. Calef R
8. Tung C-H
9. Huang T-K
10. Schmutz J
11. Satoh N
12. Yu J-K
13. Putnam NH
14. Green RE
15. Rokhsar DS
(2020) Deeply conserved synteny resolves early events in vertebrate evolution
Nature Ecology & Evolution 4:820–830.

https://doi.org/10.1038/s41559-020-1156-z
- PubMed
- Google Scholar
Book
1. Stebbins GL
(1950) Variation and Evolution in Plants
Columbia University Press.

https://doi.org/10.7312/steb94536
- Google Scholar
(2011) Seventy million years of concerted evolution of a homoeologous chromosome pair, in parallel, in major Poaceae lineages
The Plant Cell 23:27–37.

https://doi.org/10.1105/tpc.110.080622
- PubMed
- Google Scholar
1. Wang Y
2. Tang H
3. Debarry JD
4. Tan X
5. Li J
6. Wang X
7. Lee T
8. Jin H
9. Marler B
10. Guo H
11. Kissinger JC
12. Paterson AH
(2012) MCScanX: a toolkit for detection and evolutionary analysis of gene synteny and collinearity
Nucleic Acids Research 40:e49.

https://doi.org/10.1093/nar/gkr1293
- PubMed
- Google Scholar
Software
1. Wickham H
(2016) ggplot2: create elegant data visualisations using the grammar of graphics, version 3.3.3
CRAN.

https://cran.r-project.org/web/packages/ggplot2/index.html
1. Wu Y
2. Zhu Z
3. Ma L
4. Chen M
(2008) The preferential retention of starch synthesis genes reveals the impact of whole-genome duplication on grass evolution
Molecular Biology and Evolution 25:1003–1006.

https://doi.org/10.1093/molbev/msn052
- PubMed
- Google Scholar
1. Zhao T
2. Schranz ME
(2019) Network-based microsynteny analysis identifies major differences and genomic outliers in mammalian and angiosperm genomes
PNAS 116:2165–2174.

https://doi.org/10.1073/pnas.1801757116
- PubMed
- Google Scholar
1. Zhou Q
2. Zhang J
3. Bachtrog D
4. An N
5. Huang Q
6. Jarvis ED
7. Gilbert MTP
8. Zhang G
(2014) Complex evolutionary trajectories of sex chromosomes across bird taxa
Science 346:1246338.

https://doi.org/10.1126/science.1246338
- PubMed
- Google Scholar
1. Zhou Y
2. Shearwin-Whyatt L
3. Li J
4. Song Z
5. Hayakawa T
6. Stevens D
7. Fenelon JC
8. Peel E
9. Cheng Y
10. Pajpach F
11. Bradley N
12. Suzuki H
13. Nikaido M
14. Damas J
15. Daish T
16. Perry T
17. Zhu Z
18. Geng Y
19. Rhie A
20. Sims Y
21. Wood J
22. Haase B
23. Mountcastle J
24. Fedrigo O
25. Li Q
26. Yang H
27. Wang J
28. Johnston SD
29. Phillippy AM
30. Howe K
31. Jarvis ED
32. Ryder OA
33. Kaessmann H
34. Donnelly P
35. Korlach J
36. Lewin HA
37. Graves J
38. Belov K
39. Renfree MB
40. Grutzner F
41. Zhou Q
42. Zhang G
(2021) Platypus and echidna genomes reveal mammalian biology and evolution
Nature 592:756–762.

https://doi.org/10.1038/s41586-020-03039-0
- PubMed
- Google Scholar
1. Zhu T
2. Wang L
3. Rimbert H
4. Rodriguez JC
5. Deal KR
6. De Oliveira R
7. Choulet F
8. Keeble-Gagnère G
9. Tibbits J
10. Rogers J
11. Eversole K
12. Appels R
13. Gu YQ
14. Mascher M
15. Dvorak J
16. Luo MC
(2021) Optical maps refine the bread wheat Triticum aestivum cv. Chinese Spring genome assembly
The Plant Journal 107:303–314.

https://doi.org/10.1111/tpj.15289
- PubMed
- Google Scholar

Article and author information

Author details

John T Lovell
1. Genome Sequencing Center, HudsonAlpha Institute for Biotechnology, Huntsville, United States
2. Joint Genome Institute, Lawrence Berkeley National Laboratory, Berkeley, United States
Contribution
Conceptualization, Data curation, Software, Formal analysis, Visualization, Methodology, Writing – original draft, Project administration, Writing – review and editing

For correspondence
jlovell@hudsonalpha.org

Competing interests
No competing interests declared

"This ORCID iD identifies the author of this article:" 0000-0002-8938-1166
Avinash Sreedasyam

Genome Sequencing Center, HudsonAlpha Institute for Biotechnology, Huntsville, United States

Contribution
Conceptualization, Software, Formal analysis, Visualization, Methodology, Writing – original draft, Writing – review and editing

Competing interests
No competing interests declared

"This ORCID iD identifies the author of this article:" 0000-0001-7336-7012
M Eric Schranz

Biosystematics Group, Wageningen University and Research, Wageningen, Netherlands

Contribution
Conceptualization, Formal analysis, Methodology, Writing – review and editing

Competing interests
No competing interests declared

"This ORCID iD identifies the author of this article:" 0000-0001-6777-6565
Melissa Wilson

Center for Evolution and Medicine, School of Life Sciences, Arizona State University, Tempe, United States

Contribution
Formal analysis, Visualization, Methodology, Writing – original draft, Writing – review and editing

Competing interests
No competing interests declared

"This ORCID iD identifies the author of this article:" 0000-0002-2614-0285
Joseph W Carlson

Joint Genome Institute, Lawrence Berkeley National Laboratory, Berkeley, United States

Contribution
Data curation, Formal analysis, Methodology, Writing – original draft, Writing – review and editing

Competing interests
No competing interests declared
Alex Harkess
1. Genome Sequencing Center, HudsonAlpha Institute for Biotechnology, Huntsville, United States
2. Department of Crop, Soil, and Environmental Sciences, Auburn University, Auburn, United States
Contribution
Conceptualization, Investigation, Writing – original draft, Writing – review and editing

Competing interests
No competing interests declared

"This ORCID iD identifies the author of this article:" 0000-0002-2035-0871
David Emms

Oxford University, Oxford, United Kingdom

Contribution
Software, Formal analysis, Methodology

Competing interests
No competing interests declared

"This ORCID iD identifies the author of this article:" 0000-0002-9065-8978
David M Goodstein

Joint Genome Institute, Lawrence Berkeley National Laboratory, Berkeley, United States

Contribution
Conceptualization, Software, Supervision, Funding acquisition, Visualization

Competing interests
No competing interests declared

"This ORCID iD identifies the author of this article:" 0000-0001-6287-2697
Jeremy Schmutz
1. Genome Sequencing Center, HudsonAlpha Institute for Biotechnology, Huntsville, United States
2. Joint Genome Institute, Lawrence Berkeley National Laboratory, Berkeley, United States
Contribution
Conceptualization, Software, Funding acquisition, Methodology, Writing – original draft, Project administration, Writing – review and editing

Competing interests
No competing interests declared

"This ORCID iD identifies the author of this article:" 0000-0001-8062-9172

Funding

U.S. Department of Energy (DE-AC02-05CH1123)

John T Lovell
Avinash Sreedasyam
Joseph W Carlson
David M Goodstein
Jeremy Schmutz

National Institute of General Medical Sciences (R35GM124827)

Melissa Wilson

The funders had no role in study design, data collection, and interpretation, or the decision to submit the work for publication.

Acknowledgements

The GENESPACE pipeline has been improved by advice and testing by A Healey, N Walden, V Scarlett, R Walstead, S Carey, L Smith, J Vogel, J Willis, J Jenkins, T Juenger, and many others. Thanks to J Schnable, J Leebens-Mack, JG Monroe, CH Li, R Tarvin, and M Hufford for help refining the datasets and analyses presented in this manuscript. Thank you to Erich D Jarvis and the Vertebrate Genome Project members for advice and prepublication access to several genomes (budgerigar and dolphin). The work conducted by the US Department of Energy Joint Genome Institute is supported by the Office of Science of the US Department of Energy under Contract No. DE-AC02-05CH1123. Visualization was inspired in part by MCScanX and pairwise ‘river’ plots generated by other software. The use of syntenic orthogroups was originally inspired by work developed by CoGe; similar syntenic homology approaches have been implemented by other software, including pSONIC. MAW’s work on this was supported by the National Institute of General Medical Sciences (NIGMS) of the National Institutes of Health (NIH) grant R35GM124827. JTL would like to thank Ashley Lovell, our friends and family for their support, which allowed him to work on this project during the difficult past 2 years.

Copyright

This article is distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use and redistribution provided that the original author and source are credited.