Ultra-sensitive detection of transposon insertions across multiple families by transposable element display sequencing

TE display sequencing

Classic transposon display assays rely on genomic DNA digestion using restriction enzymes followed by ligation of an end-compatible custom adapter, and PCR amplification using a primer specific to the target TE sequence and another complementary to the adapter [22, 26]. To adapt this assay for high-throughput detection of TE insertions using next generation sequencing (NGS), we replaced enzymatic digestion with random fragmentation by sonication, producing fragments of approximately 200–500 bp (Fig. 1a). Similar to standard NGS library preparations, the fragmented DNA is end-repaired and A-tailed to enable the ligation of custom adapters, which are compatible with Illumina sequencing. We designed the adapters to avoid amplification by adapter primers only, which would otherwise lead to the generation of random whole genomic libraries. Specifically, the priming sites on the adapter are not present ab initio, but generated through a first round of extension from the TE primer, analogous to the strategy used in vectorette/splinkerette PCR strategies [27, 28]. Unlike these strategies, which utilize hairpins or partially mismatched “bubbled” adapters, the present method relies on asymmetric Illumina forked-adapters. These adapters are much shorter and readily compatible with NGS libraries, providing advantages compared to vectorette/splinkerette PCR adapters. To block extension from the asymmetric forked-adapter, its short arm includes a 3′ dideoxy nucleotide (Fig. 1b). Even in the unlikely situation where an adapter is extended first, producing templates with full adapter sequences at both ends, these fragment ends will preferentially anneal to form an intramolecular “panhandle” structure, which cannot be amplified further. The “suppression PCR” effect occurs because the complementarity of the adapter’s primer is shorter than that of the adapter itself, favoring intramolecular pairing. In combination, the universal adapter’s design and primer sequences ensure that only bona fide PCR templates containing both the target TE site and adapter sequence are generated during the first rounds of extension.

Fig. 1figure 1

Transposable display sequencing method. a Schematic representation of the initial steps to prepare adapter-containing whole genome libraries. b Schematic representation of specific amplification of TE-containing adapted fragments and incorporation of P5 and P7 adapters compatible with Illumina sequencing. c Examples of TE insertions detected at different dilution factors. d Sensitivity of TEd-seq to detect TE insertions at different dilution factors and in relation to sequencing variable depth. e Number of reads supporting each TE insertion present at different dilution factors in relation to variable sequencing depth

To increase further the sensitivity and specificity of the assay, we included a nested PCR step using a TE-specific primer closer to the TE edge together with an adapter’s primer (Fig. 1b), respectively containing the P5 and P7 sequences required for cluster recognition on Illumina sequencing machines. The adapter’s primers also include sequence barcodes, enabling the multiplexing of up to hundreds of samples in the same sequencing run. Critically, to improve sequencing quality due to the low diversity of the amplicon library, nested TE-specific primers contain different numbers of spacer nucleotides before the P5 adapter. Such straightforward modifications significantly improve library diversity and overall sequencing quality (Additional file 1: Fig. S1).

To analyze the data generated by this assay, we developed a ready-to-use bioinformatic tool that takes as input the sequenced reads, targeted TE extremity, and reference genome sequence to detect TE insertions. Specifically, following demultiplexing, reads are processed bioinformatically to eliminate PCR duplicates and select paired-reads containing the target TE sequence end, which is trimmed to retrieved informative insertion sites sequences. Alignment of reads on the reference genome and their clustering allows the detection of reference and non-reference insertion sites (Fig. 1c).

Highly specific and ultra-sensitive detection of new TE insertions

To assess the sensitivity and specificity of the transposable element display sequencing (TEd-seq) approach to detect non-reference TE insertions, we took advantage of a population of TE accumulation (TEA) lines in A. thaliana that have underwent extensive transposition following the epigenetic activation of TEs [12]. The identity and genomic location of heritable TE insertions in each of these lines were previously characterized using Illumina mate-pair libraries [18, 29] and TE sequence capture [18], providing a gold standard data set to test the sensitivity and specificity of TEd-seq. We selected eight lines carrying distinct repertoires of non-reference TE insertions of the LTR retrotransposon ATCOPIA93 (also known as Évadé). We extracted genomic DNA (gDNA) from each of these TEA lines and serially diluted them in DNA extracted from the reference wild-type Col-0 plants, leading to different concentrations of TEA line’s DNAs (from 1/2 to 1/6400). Pooling together these differentially diluted gDNA thus leads to a mixed gDNA with 212 non-reference TE insertions present at different dilutions (Additional file 1: Table S1).

The resulting pooled DNA was subjected to TEd-seq using primers specific to the mobile ATCOPIA93 copy (AT5TE20395), sequenced to a depth of 400 M pair-end reads, and analyzed bioinformatically to detect non-reference TE insertion sites. Comparison between the set of non-reference ATCOPIA93 insertions detected by TEd-seq to the gold standard set of previously identified insertions [29] showed that 87–100% of TE insertions were detected for all the dilutions (Fig. 1c, d). Importantly, the number of reads supporting each ATCOPIA93 insertion correlated well with the dilution factor (Fig. 1c, e), indicating that the number of supporting reads is a reliable estimator of the TE insertion representation.

To evaluate the influence of sequencing depth on the sensitivity of TEd-seq, we subsampled decreasing numbers of row paired-end reads obtained in the TEd-seq experiment and detected non-reference TE insertions. As expected, sensitivity decreased in relation to sequencing depth (Fig. 1d). Specifically, 80% of the TE insertions present in the 1/6400 dilution were detected with as little as 6 M pair-end reads. On the basis of the empirical sensitivity obtained from the serially diluted DNA, we applied a subsampling approach to estimate the theoretical limit of detection for TEd-seq (Additional file 1: Fig. S2). This analysis indicated that 400 M reads are sufficient to detect TE insertions present at dilutions as low as 1/250,000.

Sensitive detection of TE insertions using ONT sequencing

Recent developments in sequencing technologies based on Oxford Nanopore Technologies (ONT) are revolutionizing the output and speed of genomic analysis. We reasoned that in addition to allowing the sequencing of long DNA fragments, which can facilitate the identification of insertion sites within highly repetitive or low complexity regions, its portability and real-time generation of sequencing data can significantly reduce the time-to-answer. We thus tested the possibility of sequencing a TEd-seq library using ONT (Fig. 2b). We optimized the fragmentation and size-selection steps to obtain longer DNA fragments, which are more suitable for ONT sequencing.

Fig. 2figure 2

Transposable elements display long-read sequencing. a Schematic representation of TEd-seq sequencing using ONT. b Fragment size distribution of short-read and ONT-based TEd-seq. c Example of a TE insertion detected by short-read and ONT-based TEd-seq. d Sensitivity of ONT-based TEd-seq to detect TE insertions at different dilution factors. e Number of reads supporting each TE insertion present at different dilution factors for both short-read and ONT-based TEd-seq

We performed a TEd-seq library of the retrotransposon ATCOPIA93 using gDNA extracted from a single TEA line (TEA500) diluted 1/200 in Col-0 gDNA. As expected, the average size of ONT reads is longer than the fragments obtained by short-read sequencing (814.27 vs 350 nt, respectively), with some reads reaching more than 3000 nt (Fig. 2b and c). We next detected non-reference TE insertions using a modified bioinformatic pipeline compatible with ONT reads (see Methods). Comparison between the set of non-reference ATCOPIA93 insertions detected by TEd-seq to the gold standard set of previously identified insertions [29] in this TEA line showed 100% correspondence, demonstrating that combination of TEd-seq with long-read ONT sequencing is highly specific. To test the sensitivity of ONT sequencing to detect non-reference TE insertions, we generated and sequenced the TEd-seq library containing the mixed pool of serially diluted gDNA from the different TEA lines. With as little as 1 M ONT reads, we identified more than 70% of insertions present in 1/400 dilution (Fig. 2d). This sensitivity is close to the one obtained by ~ 1 M short-reads, indicating that the major factor influencing the sensitivity of TEd-seq is the depth of coverage. In combination, these results establish that TEd-seq libraries are compatible with ONT sequencing, and that the results are comparable to the ones obtained with short-read sequencing. Notably, the portability and short hand-off time of ONT sequencing offer a turnaround time of less than 24 h from DNA extraction to insertion identification, significantly expediting the investigation of TE mobilization in routine experiments.

Simultaneous detection of insertions from multiple TE families using multiplex TEd-seq

Eukaryotic genomes typically contain several active TE families. To effectively investigate transposition activity, it is desirable to simultaneously detect non-reference insertions for multiple TE families. To explore this possibility, we designed TEd-seq primers specific for the CACTA and MuDR DNA transposons ATENSPM3 and VANDAL21, which are highly mobile in the population of TEA lines [29]. Using gDNA from TEA line 454 (dilution 1/10) and 500 (dilution 1/200), we generated, for each TEA line, independent TEd-seq libraries for ATCOPIA93, VANDAL21, and ATENSPM3, together with multiplexed TEd-seq libraries using an equimolar mix of the three TE-specific primers during the amplification steps (Fig. 3a). Based on 1 M ONT reads, we detected insertions for the three TE families (Fig. 3b). Remarkably, between 97.77 and 100% of the homozygous non-reference insertions previously identified in the TEA lines [29] were detected in our multiplexed TEd-seq experiments (Fig. 3c, d), providing a global sensitivity of ~ 98% and enabling the possibility to study several TE families simultaneously. In addition, out of the 96 non-reference TE insertions detected in the multiplexed libraries, only three appear to be false positives, leading to an overall precision rate of 96.8% (Fig. 3c, d). Altogether, these results established multiplex TEd-seq as a remarkably cost-effective, highly sensitive, and specific method to simultaneously detect non-reference insertions of different TE families, enabling the investigation of rich mobilomes.

Fig. 3figure 3

Simultaneous detection of insertions for different families by multiplexed transposable element display sequencing. a Schematic representation of specific multiplexed amplification of TE-containing adapted fragments and incorporation of P5 and P7 adapters compatible with Illumina sequencing. b Examples of TE insertions from different TE families detected by multiplex and simplex TEd-seq. c and d Precision and sensitivity of multiplex TEd-seq to detect TE insertions present in different TEA lines

Multiplex TEd-seq empowers evolve and resequencing experiments

Although most TE-induced mutations are likely deleterious, empirical and theoretical evidence suggest that a continuum of selective effects must exist [30]. However, direct estimations of the evolutionary dynamics of new TE insertions remain largely underexplored. Whole genome sequencing (WGS) of pools of individuals from experimentally evolved populations, such as in common garden competition experiments or experimental evolution setups, has enabled the identification of genomic responses to selection. The so-called evolve and resequencing (E&R) approaches use short-read WGS to identify segregating polymorphisms and characterize their changes in allele frequency across generations [31, 32]. Detection of complex DNA sequence variations, such as TE insertion polymorphisms, as well as of rare variants remains very challenging using WGS, limiting the wide implementation of E&R approaches to investigate the evolutionary dynamics of new TE insertions. The ultra-high sensitivity of TEd-seq combined with its ability to simultaneously detect non-reference TE insertions for multiple TE families in a cost-effective manner may therefore empower E&R experiments. As a proof-of-concept for the utility of TEd-seq in quantifying the presence and frequency of non-reference TE insertions in these experimental experiments, we generated a largely isogenic wild-type population of Arabidopsis plants undergoing a burst of transposition for the endogenous TE families ATCOPIA93, VANDAL21, and ATENSPM3. Specifically, we crossed wild-type Col-0 plants with distinct TEA lines (TEA 523, 54, and 393) containing active copies for these three TEs families. F1 plants were selfed to obtain thousands of F2 seeds. An equal number of F2 seeds (3500) from each cross, together with naïve Col-0 seeds (500), were mixed to obtain a so-called TE invaded population (G0) (Fig. 4a). We performed two parallel E&R experiments by growing 2500 plants from the TE invaded population in environmental chambers reproducing realistic climates, including temperature, humidity, and light intensity and quality (Additional file 1: Fig. S3) recorded at an experimental station near Paris (see Methods). All produced seeds at the end of the life cycle were collected from each population. This was repeated for another generation to obtain G2 offspring.

Fig. 4figure 4

Investigation of the evolutionary dynamics of new TE insertions using TEd-seq. a Schematic representation of the evolve and resequencing (E&R) experiment developed to study the evolutionary dynamics of new TE insertions. Two independent populations (Pop. 1 and Pop. 2) were used to run two E&R replicated experiments. b Number of TE insertions identified by multiplex TEd-seq in two independent samples of 1000 offspring of G1 Pop. 1 for the three TE families (ATCOPIA93, VANDAL21, and ATENSPM3) transposing actively in the experimental population. c Number of private and shared insertions detected in the two samples. d Comparison of allele frequencies of all TE insertions estimated in the two populations. e Allele frequencies of shared and private TE insertions. Statistically significant differences are indicated (MWU test). f Meta-analysis of levels of H2A.Z levels around shared and private ATCOPIA93 insertion sites. g Estimated transposition rate across generations in both replicated populations. Three independent samples were analyzed in each case. h Allele frequencies change across generations in both populations. Insertions showing increase, decrease, or no change allele frequencies were selected by comparison with allele frequencies trajectories calculated under random genetic drift. i Comparison of TE insertions detected as increase, decrease, or no change among populations. j Fraction of TE insertions identified as recurrent increase, decrease, or no change within genic features and intergenic regions. k Circos representation of TE insertions identified as recurrent increase, decrease, or no change in chromosome 5 of Arabidopsis. Density of annotated genes and TEs are indicated as heatmaps

To assess the ability of TEd-seq to accurately estimate the allele frequency of non-reference TE insertions, we performed multiplex TEd-seq (ATCOPIA93, VANDAL21, and ATENSPM3) on two random samples of 1000 G1 offsprings from population 1. In total, we identified 340 and 321 non-reference insertions in each sample, most of which belonged to the LTR retrotransposon ATCOPIA93 and were shared between samples (Fig. 4b and c). We next estimated the allele frequency of each TE insertion by comparing the numbers of reads supporting new (i.e., non-reference) TE insertions in relation to that for fixed (i.e., reference) insertions. This analysis revealed that allele frequency estimations were highly consistent between samples (Fig. 4d), demonstrating that TEd-seq offers reliable quantification of the allele frequency of TE insertions.

Shared insertions among offspring samples should correspond to segregating alleles in the G1, and as such they are expected to be present at higher allele frequencies than private insertions. To test this expectation, we compared the allele frequency of shared and private non-reference TE insertions. As predicted, shared insertions were covered by a much higher number of reads than private ones (Additional file 1: Fig. S4), and we estimated that segregate at frequencies higher than 0.025, compared to 0.004 for the latter (Fig. 4e). Thus, private insertions represent de novo heritable transposition events.

It was previously shown that ATCOPIA93 integrates preferentially within nucleosomal DNA containing the histone variant H2A.Z [29]. Analysis of the integration sites of ATCOPIA93 in relation to the presence of H2A.Z confirmed this preference and revealed a higher occupancy of H2A.Z at de novo compared to shared TE insertion sites (Fig. 4f). This result are consistent with counterselection of TE insertions within H2A.Z containing regions.

Next, we set out to investigate the transgenerational dynamics of active transposition on the E&R experiment. We performed multiplex TEd-seq on DNA extracted from independent samples of 1000 offspring of G0, G1, and G2 generations for both replicated populations (Fig. 4a). In total, we identified 5799 insertions, 48% of which were private to individual samples. Using the number of private insertions as a proxy for de novo TE mobilization events, we estimated the transposition rate across replicates and generations. Estimated transposition rates were very similar between populations and generations (Fig. 4g), consistent with lineal, rather than exponential [29], accumulation of new TE copies across generations.

To characterize the evolutionary dynamics of segregating TE insertions, we estimated the allele frequency across generations of insertions shared among samples. To detect insertions potentially subjected to selection, we compared their allele frequencies with neutral expectations based on random genetic drift and considering less than 5% of outcrossing rate [33]. Segregation dynamics of most (~ 63%) TE insertions were indistinguishable from the expected random genetic drift. Conversely, 19% and 14% of the TE insertions show a statistically significant increase or decrease in allele frequencies in at least one population, respectively. In addition, 26 (3%) and 33 (3.9%) of these insertions show paralleled allele frequency changes in both populations (Fig. 4i). Such parallel changes are unlikely to be produced by chance (hypergeometric test p value < 6.3E − 6 and 6.9E − 8 for increases and decreases, respectively), suggesting recurrent positive or negative selection of these TE insertions across replicated populations. In addition, 493 TE insertions show no changes in allele frequency across populations, suggesting that they correspond to neutral, or nearly neutral, mutations (Fig. 4i). The distribution of TE insertions with reproducible increase, decrease, or no changes in allele frequencies shows no chromosome-wide bias (Fig. 4k), suggesting that the observed changes in allele frequencies are not due to selection on underlying haplotypes (i.e., background selection), but rather on individual TE insertions. Moreover, TE insertions with recurrent decrease in allele frequency show overrepresentation over exonic features compared to neutral TE insertion mutations (Fig. 4j), consistent with exonic insertions being more likely associated with deleterious effects.

Among the 18 genes with insertions in their gene body or promoter that are increasing in population frequency, we identified 14 GO terms. These terms are primarily associated with transmembrane transporters, suggesting a role for these TE insertions in cell-environment interaction and transport processes. We identified 56 GO terms associated with the 28 genes carrying counterselected TE insertions. These terms are involved in various binding activities (e.g., nucleotides and nucleosides), cell wall structure and organization, which are essential for maintaining cell integrity and ensuring proper cellular interactions, as well as responses to environmental changes. For instance, we found an ATENSPM3 insertion with consistent decrease in allele frequency across populations within the gene encoding KNUCKLES (KNU), which is involved in an essential regulation of flower determinacy [34] (Additional file 1: Fig. S5). In combination, these results indicate that the vast majority of new TE insertions are neutral or nearly neutral and that as much as 4% of new TE insertions are affected by positive or negative selection in our E&R experiment.

Comments (0)

No login
gif