Uploaded by common.user19301

Novel UCE Probe Set for RTA Spider Phylogenomics

Cladistics (2023) 1–13
doi:10.1111/cla.12523
A novel probe set for the phylogenomics and evolution of RTA
spiders
Junxia Zhanga*, Zhaoyi Lia, Jiaxing Laia, Zhisheng Zhangb and Feng Zhanga*
a
Key Laboratory of Zoological Systematics and Application of Hebei Province, Institute of Life Science and Green Development, College of Life
Sciences, Hebei University, Baoding, Hebei, 071002, China; bSchool of Life Sciences, Southwest University, Chongqing, 400700, China
Accepted 21 December 2022
Abstract
Spiders are important models for evolutionary studies of web building, sexual selection and adaptive radiation. The recent
development of probes for UCE (ultra-conserved element)-based phylogenomic studies has shed light on the phylogeny and evolution of spiders. However, the two available UCE probe sets for spider phylogenomics (Spider and Arachnida probe sets) have
relatively low capture efficiency within spiders, and are not optimized for the retrolateral tibial apophysis (RTA) clade, a hyperdiverse lineage that is key to understanding the evolution and diversification of spiders. In this study, we sequenced 15 genomes
of species in the RTA clade, and using eight reference genomes, we developed a new UCE probe set (41 845 probes targeting
3802 loci, labelled as the RTA probe set). The performance of the RTA probes in resolving the phylogeny of the RTA clade
was compared with the Spider and Arachnida probes through an in-silico test on 19 genomes. We also tested the new probe set
empirically on 28 spider species of major spider lineages. The results showed that the RTA probes recovered twice and four
times as many loci as the other two probe sets, and the phylogeny from the RTA UCEs provided higher support for certain
relationships. This newly developed UCE probe set shows higher capture efficiency empirically and is particularly advantageous
for phylogenomic and evolutionary studies of RTA clade and jumping spiders.
© 2023 Willi Hennig Society.
Introduction
Spiders (Order Araneae) are among the most diverse
terrestrial predators with over 50 000 species already
described (World Spider Catalog, 2022) and many
more waiting to be discovered (Agnarsson
et al., 2013). As an ancient group, spiders can be dated
back to the Devonian (>380 Ma), and have adapted to
diverse ecosystems with remarkable behaviour and
morphology during their evolutionary history. Over
the years, arachnologists have worked long and hard
to understand the diversification and evolutionary history of spiders (Garrison et al., 2016; Fernandez
et al., 2018; Shao and Li, 2018; Dimitrov and Hormiga, 2020).
*Corresponding author:
E-mail address: [email protected]; [email protected]
© 2023 Willi Hennig Society.
Spiders are well known for their production of silk
and the utility of foraging webs, which have been
hypothesized as the key innovation for the diversification of spiders (Bond and Opell, 1998). However, a
major radiation within spiders, the retrolateral tibial
apophysis (RTA) clade, are mainly wandering hunters
without foraging webs. Spiders in this clade are characterized by the presence of an RTA on the male palp
for mating stabilization, and trichobothria on the tarsi
and metatarsi for vibration sensitivity (Wheeler
et al., 2017). This spider lineage is extremely diverse
with over 25 000 described species (Dimitrov and Hormiga, 2020; World Spider Catalog, 2022), including
the most species-rich spider family Salticidae (>6000
described species; World Spider Catalog, 2022), which
are well known for their acute vision and spectacular
courtship dances as well as the recently discovered
milk provision (Richman and Jackson, 1992; Foelix, 1996; Chen et al., 2018). Recent divergence dating
10960031, 0, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/cla.12523 by University Of California, Riverside, Wiley Online Library on [31/01/2023]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
Cladistics
Zhang J. et al. / Cladistics 0 (2023) 1–13
analyses suggested that the RTA clade is relatively
young (139–161 Ma) compared with the Araneoidea
clade of largely orb-weaving spiders, but the main drivers for its diversification remain contentious (Garrison et al., 2016; Fern
andez et al., 2018; Shao and
Li, 2018; Dimitrov and Hormiga, 2020; Magalhaes
et al., 2020). In order to better understand the evolution of this major clade of spiders, we need to build
on recent work (Miller et al., 2010; Agnarsson
et al., 2013; Maddison, 2015; Wheeler et al., 2017;
Azevedo et al., 2022) to resolve its phylogeny more
fully and with better support.
The rapid development of sequencing technology
and analytical pipelines have strongly promoted progress in building the tree of life. Among various phylogenetic approaches applying genomic-scale data, the
UCE (ultra-conserved element) method targets thousands of homologous loci by designing probes from
highly conserved regions of representative taxa (Faircloth et al., 2012) and has been widely utilized in phylogenomic studies of vertebrates (e.g. McCormack
et al., 2012, 2013; Guillory et al., 2019; Pie
et al., 2019; Ochoa et al., 2020) and a variety of invertebrate groups, such as Arachnida (e.g. Starrett
et al., 2017; Van Dam et al., 2018; Hedin et al., 2019;
Kulkarni et al., 2020a), Pycnogonida (Ballesteros
et al., 2020), Collembola (Sun et al., 2020), Hymenoptera (Faircloth et al., 2015; Branstetter et al., 2017;
Cruaud et al., 2018) and Hemiptera (Forthman
et al., 2019). This approach has proved to be successful for resolving both deep and shallow relationships
(e.g. Blaimer et al., 2015; Branstetter et al., 2017;
Jesovnik et al., 2017; Van Dam et al., 2017; Blair
et al., 2019; Branstetter and Longino, 2019).
Probes are critical for UCE phylogenomic approach.
Currently two probe sets have been widely applied in
spider phylogenomics, the Arachnida probes (Faircloth, 2017; Starrett et al., 2017) and the Spider probes
(Kulkarni et al., 2020a). The Arachnida probe set was
designed based on 10 exemplar taxa across Arachnida,
including five spider species of Theridiidae, Eresidae,
Sicariidae and Theraphosidae, and contains 14 799
probes targeting 1120 loci (Faircloth, 2017). This set
of probes has been applied in phylogenetic analyses of
Arachnida (Starrett et al., 2017), harvestmen (Derkarabetian et al., 2018, 2019) and spiders (e.g. Wood
et al., 2018; Hedin et al., 2019; Ramırez et al., 2020;
Maddison et al., 2020b; Azevedo et al., 2022). The Spider probe set was designed using four exemplar taxa
of the spider families Theridiidae, Araneidae, Sicariidae and Eresidae, contains 15 051 probes harvesting
2021 UCEs (Kulkarni et al., 2020a), and has been used
in phylogenetic studies of spiders (Kulkarni
et al., 2020a, b) and Salticidae (Maddison
et al., 2020a). Comparing the performance of these
two probe sets, Kulkarni et al. (2020a) found that the
Spider probe set captured more loci than the Arachnida probe set, and the phylogenetic tree inferred by
the Spider UCEs gained higher bootstrap values for
certain nodes, and therefore showed higher potential
for solving some challenging relationships than the
Arachnida probe set. However, taxa from the RTA
clade were not included in the design of these two
probe sets. In the study by Kulkarni et al. (2020a),
UCEs were enriched for three species of the RTA
clade using the Spider probe set, and about 800–1000
loci were obtained (<50% of the targeted 2021 loci).
Applying these probe sets in jumping spiders, the
Arachnida probes usually harvested 300–700 UCEs
and the Spider probes 890–1200 UCEs (Maddison
et al., 2020a, b). By designing probes specifically targeting a clade, we should be able to obtain more loci
for resolving difficult phylogenetic relationships. For
example, Xu et al. (2021) designed the liphistiidspecific probe set (19 740 probes targeting 3111 ultraconserved loci) that was streamlined for the UCE phylogenomics of the segmented trapdoor spiders
(Suborder Mesothelae: Liphistiidae).
In this study, we aim to develop a new UCE probe
set to recover more loci and improve capture efficiency
for resolving the recalcitrant relationships within the
RTA clade of spiders and in particular within the family of jumping spiders (Salticidae). With the 15 genomes of the RTA clade sequenced and assembled in
this study, in combination with the publicly available
genomes, we (i) apply a modified probe design and
optimization procedure to develop a new UCE probe
set for spider phylogenomic and evolutionary studies
(referred as RTA probe set); (ii) compare the performance of RTA probes in resolving the phylogeny of
the RTA clade and Salticidae with Spider and Arachnida probes through an in-silico test on 19 genomes;
and (iii) test empirically the applicability of the newly
developed probe set over a broad set of spider lineages.
Materials and methods
A general flowchart for the analytical procedure of probe design,
optimization and phylogenetic analyses applied in this study is provided in Fig. S1.
Taxon sampling, DNA extraction and sequencing
In total, 57 spider species of 27 families were included in this
study, among which 40 species of 17 families belong to the RTA
clade. See Table S1 for details about the species and specimen information. We sequenced genomes of 15 spider species of the RTA
clade including the families Cheiracanthiidae (one species), Clubionidae (one species) and Salticidae (13 species). The jumping spiders
were biased during genome sequencing and probe design because we
are launching a large-scale project exploring the phylogeny and
10960031, 0, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/cla.12523 by University Of California, Riverside, Wiley Online Library on [31/01/2023]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
2
3
evolution of this fascinating spider group. In addition, the sequencing data obtained at Southwest University (Professor Zhisheng
Zhang) for genome annotation studies of Argyroneta aquatica (Clerck, 1757) (Dictynidae) and Pardosa laura Karsch, 1879 (Lycosidae),
as well as the genome data of 12 spider species (three Araneidae,
two Eresidae, three Theridiidae, one Lycosidae, one Dysderidae, one
Sicariidae and one Theraphosidae) were also downloaded from
NCBI to assess the applicability of the newly developed probe set in
a wide range of spider lineages.
Genomic DNA was extracted using QIAGEN DNeasy Blood &
Tissue Kit, and 1 ll of RNase A (Solarbio) was added to the DNA
extraction and then left at room temperature for 2 min to remove
RNA. The quantity of DNA was checked using a QubitTM fluorometer. The genomic DNA was sent to Novogene Co. Ltd for library
preparation using a Truseq Nano DNA HT sample preparation kit
(Illumina USA), and then sequenced on an Illumina NovaSeq platform with 150 bp paired-end reads and insert size around 350 bp.
sequences (160 bp per locus) for a temporary probe set (39 tiling
density; repetitive regions <15%; 30% < GC content <70%; two
probes per locus). The putative duplicated probes were identified and
removed from the temporary probe set with the default setting (identity and coverage both as 50).
The temporary probes were aligned back to all the reference genomes at a 50% sequence identity and the conserved loci were
extracted (buffered to 180 bp) from all genomes. If the probe
matched to different regions of the genome, the locus was deleted.
Again, a database was built and the conserved loci shared by different number of taxa were calculated. The extracted conserved loci of
all genomes shared by all of the eight reference taxa were used for a
final probe design (39 tiling density; repetitive regions <15%;
30% < GC content <70%; two probes per locus), and the putative
duplicated probes were identified and removed with identity and coverage both as 50. These probes were titled the “RTA_v1” probe set
for clarity.
Genome assembly
In-silico test and comparison of the three probe sets
Genome assembly followed the Phylogenomics from Lowcoverage Whole-genome Sequencing (PLWS) pipeline as in Zhang
et al. (2019). In brief, the sequenced reads were first compressed into
clumps with duplicates removed using clumpify.sh (BBTools) (Bushnell, 2014). Quality trimming of reads was completed using bbduk.sh
(BBTools) with the reads shorter than 15 bp or with more than 5 Ns
as well as the poly-A or poly-T tails of at least 10 bp being trimmed.
The bbnorm.sh (BBTools) was then used to normalize the reads in
order to accelerate the assembly. Genome contigs were assembled
with multiple k-mer strategies in Minia v3.2.1 (Chikhi and
Rizk, 2013). The contigs representing high heterozygosity were identified and deleted by Redundans v0.13c (Pryszcz and
Gabald
on, 2016). Contig scaffolding and gap filling were performed
with BESST v2.2.8 (Sahlin et al., 2014) and GapCloser v1.12 in the
SOAPdenovo2 suite (Luo et al., 2012) respectively. Sequences shorter
than 500 bp were deleted by reformat.sh (Bushnell, 2014).
We compared the performance of the “RTA_v1”, “Spider” and
“Arachnida” probe sets through the in-silico test following the PHYLUCE workflow (Faircloth, 2016). The UCEs were harvested from
19 genomes (see Table S1) using the three probe sets, respectively.
First, the genome data were converted from fasta to 2bit format
using FaToTwoBit (http://hgdownload.soe.ucsc.edu/admin/exe/),
and the corresponding sizes.tab was built with TwoBitInfo (https://
genome.ucsc.edu/goldenPath/help/twoBit.html). We then aligned
each of the three sets of probes to the 19 genomes with the coverage
and identity both as 75 and extracted 500 bp on either side. The
extracted UCE loci were aligned to the corresponding probe set
(min-coverage and min-identity both as 65) to remove duplicated
UCEs and determine the final orthologues for phylogenetic analyses.
The remaining UCE loci were imported into Geneious Prime
v2019.1.3 (Kearse et al., 2012) and the loci with <15 taxa were
excluded from downstream analyses. Sequence alignments were carried out using Mafft v7.313 (Katoh and Standley, 2013) with the LINS-I strategy.
Probe design
Eight genomes were used for identifying UCEs and designing
probes: one each of Cheiracanthiidae, Clubionidae, Dictynidae,
Lycosidae and Theridiidae, and three Salticidae (see Table S1). The
probe design followed the PHYLUCE (Faircloth, 2016) workflow.
We selected the jumping spider species Attulus fasciger (Simon, 1880)
as the base genome because it was sequenced with a higher depth.
Six genomes of the RTA clade were used as exemplars to represent
the genetic diversity of this group, and the genome of the theridiid
Parasteatoda tepidariorum was also included in probe design as outgroup. Short reads (100 bp paired-end and 29 coverage) of these
assembled genomes were simulated with ART (Huang et al., 2012).
The exemplar genomes were then aligned to the base genome using
Stampy v1.0.32 (Lunter and Goodson, 2011) with the substitution
rate as 0.05 and the insert size as 200. The resulting alignments were
saved in BAM format and reduced using Samtools v1.10 (Li
et al., 2009).
BEDtools v2.28.0 (Quinlan and Hall, 2010) was used to convert
the BAM files to BED format, which allowed us to sort and merge
overlapping or nearly overlapping alignment positions. Alignments
that were shorter than 80 bp or contained a high proportion of
repetitive regions (>15% of length) or ambiguous (N) bases were
removed. The retained alignments were put into an SQLite database
for identification of the conserved loci shared between the base genome and a different number of exemplar genomes. We selected the
conserved loci shared between the base genome and all seven exemplar genomes and extracted the corresponding base genome
Overlap check and identification of coding regions
A previous study found that some targeted UCEs were actually
different regions of the same gene (Hedin et al., 2019), which could
possibly cause redundancy or overlap of UCE loci in the phylogenetic dataset. We checked for overlap of UCEs in each of the three
datasets obtained from different probe sets. First, the sequences of
A. aquatica were extracted from the alignments of each dataset using
Seqkit v0.13.2 (Shen et al., 2016). If the A. aquatica sequence in the
alignment was missing, the sequence of its close relative species (Pardosa laura or Pardosa pseudoannulata) was extracted instead. These
extracted sequences (with gaps at both ends converted to N and
internal gaps removed) were then mapped to the annotated genome
of A. aquatica (publication pending) in Geneious Prime v2019.1.3.
We then inspected the mapping results to identify the UCEs that are
overlapping or nearly overlapping with each other, of which we only
retained the one with a longer sequence and more taxa. A few UCEs
that were not able to be mapped to the annotated genome of
A. aquatica (usually owing to the large number of Ns at the ends)
were also excluded from downstream analyses.
While conducting the overlap check, we also recorded if a UCE is
in a coding or non-coding region according to the annotated genome. The proportion of coding vs. non-coding UCEs was then calculated for each dataset (after excluding the overlapping and
unmapped UCEs). In addition, we checked the congruence between
the RTA and Spider UCEs by cross-checking the mapping results
10960031, 0, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/cla.12523 by University Of California, Riverside, Wiley Online Library on [31/01/2023]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
Zhang J. et al. / Cladistics 0 (2023) 1–13
Zhang J. et al. / Cladistics 0 (2023) 1–13
and identified the UCE loci that are targeted by both the Spider and
RTA probe sets.
Optimization of the RTA probe set
We conducted an additional optimization on the “RTA_v1” probe
set to improve the capture efficiency and compatibility with data
captured using other probe sets. First, the probes associated with the
loci of low taxon-occupancy (<75%) as well as the overlapping and
unmapped UCEs identified during the overlap-check were removed.
The UCEs harvested from the 19 genomes (≥4 genomes) that were
unique to the “Spider” probe set (UCEs also targeted by the
“RTA_v1” probe set were excluded) and their associated probes for
the two species P. tepidariorum (Theridiidae) and Stegodyphus mimosarum Pavesi, 1883 (Eresidae) were added to the optimized RTA
probe set. Only the Spider probes were considered during the optimization since many of the Arachnida probes were already incorporated into the Spider probe set (Kulkarni et al., 2020a). This will
help to efficiently integrate the data generated with different probe
sets. After combining the remaining RTA probes with the Spider
probes, we identified and removed the putative duplicated probes
with identity and coverage both as 70. We added “20 000 000” to
the UCE number associated with the Spider UCEs (e.g. “uce-7” converted to “uce-20000007”, “uce-10061” converted to “uce-20010061”)
to differentiate them from the “RTA-specific” UCEs. The optimized
probe set was labelled as “RTA_v2”.
Test of RTA_v2 probe set
The performance of the RTA_v2 probe set was tested on the 19
genomes focusing on the RTA clade and jumping spiders (see above)
and a broader sampling of spiders that includes 57 species of Mygalomorphae and Aranemorphae (Synspermiata, Eresidae, Araneoidea
and RTA clade). For the available genomic data, the UCEs were
directly extracted from the genomes following the above protocol.
To test the optimized probes empirically, the UCE loci were also
enriched and sequenced for 28 species using the RTA_v2 probe set
manufactured by Daicel Arbor Biosciences.
For empirically capturing UCEs, the genomic DNA extracted
from each specimen was first fragmentated by sonication (target size
of approximately 300–600 bp). The fragmented DNA (21 lL) was
then used as input for DNA library preparation with the NEXTFLEXÒ Rapid DNA-Seq Kit 2.0 (Bioo Scientific) following the
manufacturer’s protocol with minor modifications. After the adapter
ligation using the NEXTFLEXÒ Unique Dual Index Barcodes (Set
C), a 0.89 beads clean was conducted (NEXTFLEX Cleanup Beads
2.0, Bioo Scientific) followed by a PCR reaction using the HiFi
HotStart ReadyMix (Kapa Biosystems). The PCR system was as follows: 25 lL postligation library, 26 lL HiFi HotStart ReadyMix
and 2 lL primer mix. The following thermal protocol was applied:
98°C for 45 s; 18 cycles of 98°C for 15 s, 60°C for 30 s and 72°C for
1 min; and final extension at 72°C for 5 min. PCR clean-up was
done with a 0.89 beads clean. The 28 libraries were divided into
three pools with nine or 10 libraries being combined into one pool
(final volume of 7 lL and final concentration of 171.6–208 ng/lL) at
equimolar ratios for UCE enrichment following the myBaits protocol 5.01 (Daicel Arbor Biosciences). For one pool (nine libraries), we
conducted the UCE enrichment using the RTA_v2 probes and the
Spider probes independently in order to directly compare the capture
efficiency of the two probe sets; the other two pools were only captured with the RTA_v2 probes. The enriched UCE libraries were
then sent to Novogene Co. Ltd for sequencing using the Illumina
NovaSeq platform with 150 bp paired-end reads.
The adapters and low-quality bases were removed from the
sequenced raw reads for each species using the bbduk.sh (BBTools).
The trimmed reads were then assembled using SPAdes v3.14.1 (Prjibelski et al., 2020) with “--cov-cutoff auto”. The assembled contigs
shorter than 200 bp were removed for subsequent analyses using
reformat.sh. Finding and extracting UCEs from the assembled contigs followed the PHYLUCE workflow with min-coverage and minidentity both set to 65. The UCEs extracted from genomes and target enrichment data were combined and organized by locus, and
then aligned using Mafft v7.313 with the L-INS-I strategy.
Phylogenomic analyses
In total, 11 datasets were generated for phylogenetic reconstruction: nine were from the in-silico test and comparison of three probe
sets (RTA_v1_Full_UCE, Spider_Full_UCE, Arachnida_Full _UCE,
RTA_v1_Coding_UCE, Spider_Coding_UCE, Arachnida_Coding
_UCE, RTA_v1_Non-coding_UCE, Spider_Non-coding_UCE and
Arachnida_Non-Coding _UCE) and two from the final optimized
RTA_v2 probe set (RTA_v2_19genomes_UCE and RTA_v2_57spp_UCE). The combination of Spruceup v2020.2.19 (Borowiec, 2019)
and Seqtools (PASTA; Mirarab et al., 2014) was used for alignment
trimming. First, we applied Spruceup v2020.2.19 to convert the obviously misaligned fragments in each alignment to gaps (cutoffs as
0.85). The gappy regions in each alignment were then masked using
Seqtools in PASTA package with “masksites = 30” for the
“RTA_v2_57spp_UCE” dataset and “masksites = 10” for all other
datasets.
The “gene tree and alignment” method as in Zhang et al. (2020)
was also applied to check and remove the putative paralogous or
contamination sequences in each dataset. A gene tree was first constructed for each MSA (Multiple Sequence Alignment) using
RAxML v8.2.12 (Stamatakis, 2014) with the GTRGAMMA model.
Gene trees were inspected using the customized Python script (Zhang
et al., 2020) to flag taxa with odd sequences that may be subject to
contamination errors (identical sequences for unrelated taxa) or
resulted in abnormally long branches on the gene tree. These flagged
sequences were then removed from the corresponding MSAs. For
the “RTA_v2_57spp_UCE” dataset, alignments with <30 taxa or
150 bp were removed from subsequent phylogenetic reconstruction
analyses.
All UCE loci for each dataset were concatenated by FASconCAT
v1.0 (K€
uck and Meusemann, 2010). Maximum likelihood (ML) analyses were performed on each concatenated matrix using IQ-TREE
v2.0.6 (Minh et al., 2020). The best-fitting model and optimized partition scheme were inferred for each supermatrix in IQ-TREE v2.0.6
using the option “-m MF+MERGE”. For each dataset, we ran 40
independent ML tree searches (20 with random starting trees and 20
with parsimonious starting trees) in IQ-TREE v2.0.6. Nonparametric bootstrap analyses (100 replicates) were conducted to
assess node support. Bayesian inferences were also conducted for the
concatenated full loci matrices in ExaBayes v1.5.1 (Aberer
et al., 2014). The Markov Chain Monte Carlo (MCMC) chain of
each Bayesian analysis was set for 200 000 generations
(“RTA_v2_57spp_UCE” dataset) or 150 000 generations (the other
datasets) and two independent runs were employed. Trees were sampled every 500 generations. The Effective Sampling Size (ESS) and
Potential Scale Reduction Factor (PSRF) were inspected to ensure
the convergence following the ExaBayes manual. The first 25% of
sampled trees were discarded as burn-in. The coding and non-coding
regions were also concatenated and then analysed separately in IQTREE v2.0.6 (following above procedure) to inspect their performance on phylogenetic reconstruction. Maximum parsimony (MP)
analyses were conducted on each concatenated dataset using TNT
(Goloboff et al., 2008), and the recently developed scripts for running TNT analyses on phylogenomic dataset were applied (Torres
et al., 2021). The MP tree searches were carried out with “New
Technology” search scheme, searching level at 5 and number of hits
10960031, 0, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/cla.12523 by University Of California, Riverside, Wiley Online Library on [31/01/2023]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
4
to best length at 3, and 500 replicates of bootstrap analyses with a
fast approximate resampling algorithm were conducted to assess
node support (searching level at 3).
The coalescent-based species-tree method to account for potential
gene tree heterogeneity and discordance was also applied to the three
concatenated full loci datasets for the comparison of three different
probe sets. First, the ML tree and 100 non-parametric bootstrap
replicates were inferred for each alignment in IQ-TREE v.2.0.6 using
the best-fitting model selected by ModelFinder (Kalyaanamoorthy
et al., 2017). For each gene tree, the branches with bootstrap ≤30%
were collapsed by Newick Utils v1.6 (Junier and Zdobnov, 2010).
The Accurate Species Tree Algorithm (ASTRAL-III v5.7.1; Zhang
et al., 2018) was then applied to estimate the species tree with 100
replicates of bootstrapping to assess the node support.
Results
Genome assembly and probe design
Genome sequencing and assembly results were provided in Table S1. About 20–90 million reads were
generated for the genome assembly of selected taxa.
The assembled genome size varies across taxa (0.6–3.5
Gb), which may be affected by the actual genome size
and sequencing depth, and the GC contents of the
assembled genomes are all around 30%.
The UCE number shared among different numbers
of taxa during probe design is presented in Table S2.
We selected the 3900 UCEs shared by Attulus fasciger
and all seven exemplar taxa for the probe design.
The “RTA_v1” probe set contains 60 018 probes targeting for 3856 UCEs, which was applied in the subsequent in-silico test and comparison of three probe
sets.
In-silico test and comparison of three probe sets
The number of UCE loci harvested from different
probe sets for each taxon is shown in Table S1. On
average, ~700 UCEs were obtained from the Arachnida probe set, ~1340 UCEs from the Spider probe set
and ~2600 UCEs from the “RTA_v1” probe set. The
“RTA_v1” probes harvest about twice as many UCEs
as the Spider probes and four times as many as the
Arachnida probes (Fig. 1a; Table S1).
The statistics for the UCE datasets from different
probes are provided in Table S3. After the overlap
check, the RTA_v1, Spider and Arachnida datasets
contain 2347, 961 and 424 UCEs respectively (matrix
with ~75% of taxon-completeness). The total sites and
parsimony-informative sites of the supermatrix of the
RTA_v1 dataset are both about six times the Arachnida supermatrix and three times the Spider supermatrix. The missing data of the three datasets are similar
(Fig. 1b) with the alignments in the Spider dataset
having the lowest average missing data (Spider dataset,
10.28%, ranging from 0.42 to 21.70%; RTA_v1
5
dataset, 11.51%, ranging from 0.03 to 23.24%; Arachnida dataset, 11.91%, ranging from 1.10 to 26.72%).
Mapping the representative UCE sequences from
each dataset to the annotated A. aquatica genome, the
results show that most of the UCE loci (over twothirds) in each dataset are in coding regions. The
RTA_v1 dataset contains a lower proportion (69%) of
coding UCE loci than the Spider (85%) and Arachnida (90%) dataset (Fig. 1c).
Optimization of the RTA probe set
After deleting the probes associated with lowoccupancy, overlapping and unmapped UCEs and
adding the probes for the “Spider-specific” UCEs, the
final optimized RTA probe set (“RTA_v2”) contains
41 845 probes targeting 3802 loci, of which 36 385
probes are for the RTA UCEs (2334 loci) and 5460
probes are associated with the “Spider-specific” UCEs
(1468 loci).
On average, 3058 UCEs were obtained from the 29
genomes and 2485 UCEs from the 28 species enriched
empirically using the RTA_v2 probe set (Table S1).
For the nine samples that were enriched with both
probe sets, on average the RTA_v2 probes (~2597
UCEs) captured about 2.5 times as many loci as the
Spider probes (~1006 UCEs), and the average capture
success rate for the RTA_v2 probes was higher than
that of the Spider probes (68.3 vs. 49.8%) (Table S1).
Phylogenetic inference
The results of phylogenetic inference from different
datasets and methods are provided in Figs 2 and 3
and Figs S2–S8. For comparison of the three probe
sets, the ML analyses on the concatenated full and
coding datasets recovered the same topology (Fig. 2)
and the RTA_v1 dataset shows higher average bootstrap support than the corresponding Spider or Arachnida dataset (Table S3). In addition, the RTA_v1 data
show higher support for some nodes, e.g. the clade
with sampled Dionycha taxa (RTA_v1, 100% in all
sets vs. Arachnida, 90% in the full set and coding set)
and the node with Myrmarachne formicaria (JXZ414),
Marpissa milleri (JXZ425) and Mendoza nobilis
(JXZ419) (RTA_v1, 100% in all sets vs. Spider, 98%
in the full set and 86% in the coding set; Arachnida,
98% in the full set). The concatenated non-coding loci
from the three datasets all fail to recover certain
nodes. For example, the clade with Myrmarachne
formicaria (JXZ414), Marpissa milleri (JXZ425) and
Mendoza nobilis (JXZ419) was not recovered by the
non-coding loci from the Arachnida set, and the node
with Maratus sp. (JXZ159a) and Parabathippus shelfordi (JXZ417) was not recovered by the non-coding
loci from both RTA_v1 and Spider sets. For the
10960031, 0, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/cla.12523 by University Of California, Riverside, Wiley Online Library on [31/01/2023]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
Zhang J. et al. / Cladistics 0 (2023) 1–13
Zhang J. et al. / Cladistics 0 (2023) 1–13
RTA_v1 probes
(a)
Spider probes
Arachnida probes
(b)
RTA_v1 Probes
Number of loci
Spider Probes
4000
0.09
Arachnida Probes
2000
Pa
Density
rt
r
ge ia h egiu
lid ete
s
es
r
ag oide
or
a
ifo
rm
is
na
h
Pa
n
ra ne f obil
ba
th orm is
Pl ippu icar
ex s
i
ip she a
p o lf
o
i
Sy Po des rdi
ra
ste
at
od
a
Ar tepi
gy da
0
ro rio
ne
ta rum
Pa
r
P aq
Cl dos ard uat
ub a p
os ica
io
a
s
na eud la
u
Ch pse oan ra
ud
n
ei
ra
og ula
ca
t
e
nt rm a
an
hi
u
ic
At m in a
tu
s
ig
lu
n
At s fa e
tu
sc
l
i
Co us s ger
in
ry
en
th
a
sis
Ev lia
H arc opim
ab
ha
a
ro
a
na lb
ttu ari
so a
ph
r
M Mar ys
ar
atu
pi
ss
ss
p.
a
M
M
m
yrm en il
ar doz leri
ac a
0.06
0.03
(c)
15%
31%
69%
10%
85%
90%
0.00
RTA_v1 Probes
Coding Region
Spider Probes
Arachnida Probes
0
10
20
Missing percentage
Non-coding Region
Fig. 1. Comparison of UCE (ultra-conserved element) data from RTA, Spider and Arachnid probe sets. (a) Number of UCE loci extracted from
19 genomes; (b) distribution of missing percentage in the alignments of three datasets; and (c) proportion of coding vs. non-coding loci in three
datasets.
deeper nodes the non-coding loci tend to show lower
support than the coding loci. The Bayesian analyses
on the three concatenated full loci datasets recovered
the same topology as in Fig. 2 with posterior probability for all the nodes being 1.0. However, the MP analyses often recovered different topologies as in Fig. 2,
for instance the placement of Clubiona pseudogermanica (YCH205) and Myrmarachne formicaria (JXZ414),
sometimes with relatively high bootstrap support
(Fig. S6).
The species trees from ASTRAL analyses sometimes
show different relationships for the deeper nodes
(Fig. S5). For instance, Clubiona pseudogermanica
(YCH205) was placed as the sister to the clade with
Argyroneta and Pardosa in the species tree based on
the RTA_v1 dataset rather than within the Dionycha
clade as in the concatenated analyses. Within Salticidae, the species tree from the Spider dataset shows the
sister relationship of Maratus sp. (JXZ159a) with
Corythalia opima (JXZ418) but with low bootstrap
support (31%).
The concatenated dataset of 19 genomes from the
final optimized probe set (“RTA_v2_19genomes_UCE”) with about 50% of taxon-completeness contains 3645 UCE loci (Table S3), and the ML analysis
resulted in the same topology as that from the
RTA_v1 concatenated dataset but with all nodes having bootstrap supports of 100% and posterior probabilities of 1.0 (Fig. 2). The concatenated dataset of 57
species (“RTA_v2_57spp_UCE”) with about 50% of
taxon-completeness
contains
3271
UCE
loci
(Table S3), and the results from ML and Bayesian
analyses are shown in Fig. 3. The relationships within
the RTA clade are congruent with the tree from the
“RTA_v2_19genomes_UCE” dataset. In addition, the
MP analysis on the “RTA_v2_19genomes_UCE” dataset recovered different placements for Clubiona pseudogermanica (YCH205) and Myrmarachne formicaria
(JXZ414) with strong bootstrap supports (100%); the
MP analysis on the “RTA_v2_57spp_UCE” dataset
recovered slightly different placements for Orthobula
crucifera (JXZ601) and Parabathippus shelfordi
(JXZ417) (Figs S7 and S8).
Discussion
This study aims to present a new UCE probe set
that provides more loci for future comprehensive phylogenomic and evolutionary studies of spiders. The
10960031, 0, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/cla.12523 by University Of California, Riverside, Wiley Online Library on [31/01/2023]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
6
7
Fig. 2. Summary of phylogenetic inferences from different datasets and analytical methods. The tree shown is the ML tree (log-likelihood = 24
526 285.98) on the concatenated dataset (3645 loci) from the RTA_v2 probe set with the black circles indicating that the bootstrap supports and
posterior probabilities are 100% and 1.0 respectively for all the nodes. The scale bar is in substitutions per position.
final optimized probe set (RTA_v2) can target up to
3802 UCEs, about two to three times as many as the
previously available Spider (up to 2021 loci) and
Arachnida (up to 1120 loci) probe sets (Faircloth, 2017;
Starrett et al., 2017; Kulkarni et al., 2020a). We also
accommodated the loci targeted by the other probe
sets in this newly designed probe set, so that the data
obtained in previous studies (e.g. Starrett et al., 2017;
Wood et al., 2018; Kulkarni et al., 2020a; Maddison
et al., 2020a; Girard et al., 2021) can be efficiently
integrated with sequences captured using the new
probe set.
We specifically tailored the new probe set for the
RTA clade (spiders with retrolateral tibial apophyses
on the male palpi; Spagna and Gillespie, 2008) and
jumping spiders (a major radiation within the RTA
clade). The RTA clade is hyperdiverse, containing
more than half of the recorded spider species. However, neither the Arachnida nor the Spider probe set
included species of the RTA clade during probe design
largely owing to a lack of genomes of RTA spiders
(Faircloth, 2017; Kulkarni et al., 2020a). The target
efficiency of these probe sets in RTA species is low.
For instance, in jumping spiders (Salticidae) the
Arachnida probes can harvest 300–700 UCEs (Maddison et al., 2020a, b), and the Spider probes can harvest 890–1200 UCEs (Maddison et al., 2020a), at most
around 60% of the total targeted loci. Using the optimized probe set (RTA_v2), up to over 82% of the targeted loci (3802 UCEs) were successfully enriched.
Direct comparison of the capture efficiency using nine
libraries shows that the new probe set (RTA_v2) could
obtain from RTA spiders on average ~1600 more
UCEs with ~20% higher capture efficiency than the
Spider probe set. Low capture efficiency usually results
in phylogenetic datasets with fewer loci and more
missing data. For instance, the dataset generated from
the spider probes in Kulkarni et al. (2020a) had 1010
loci with 25% of taxon-occupancy, but when the
taxon-occupancy requirement was increased to 50%,
only 276 loci were retained. Here, the dataset generated from the RTA_v2 probes and 57 species with
10960031, 0, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/cla.12523 by University Of California, Riverside, Wiley Online Library on [31/01/2023]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
Zhang J. et al. / Cladistics 0 (2023) 1–13
Zhang J. et al. / Cladistics 0 (2023) 1–13
Fig. 3. The ML tree (log-likelihood = 23 158 038.986) on the concatenated dataset from the RTA_v2 probe set and 57 spider species with the
circles at the nodes indicating the bootstrap supports (posterior probabilities for all the nodes are 1.0). The * along the taxon name indicates that
the UCEs were empirically captured using the manufactured probes for the species. The schematic tree shown at the left bottom corner is modified from Kulkarni et al. (2020b), indicating the different placement of Eresidae on the phylogeny. The scale bar is in substitutions per position.
about 50% taxon occupancy contained over 3000 loci
(Table S3). With relatively poorly assembled genomes
(e.g. Habronattus ophrys and Plexippoides regius), the
RTA_v2 probes can harvest many more loci than the
Spider (~800 more loci) and Arachnida (~1200 more
loci) probes (Table S1).
The analyses in this study with relatively few exemplar taxa are not intended to tackle the problematic
10960031, 0, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/cla.12523 by University Of California, Riverside, Wiley Online Library on [31/01/2023]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
8
nodes in the spider phylogeny, but to show that the
new probe set from this study has great potential to
provide more sequence data for resolving the recalcitrant relationships and exploring evolutionary questions within spiders, especially the RTA clade and
Salticidae. The relationships of major spider lineages
resulting from the 57-species dataset is consistent with
previous studies using transcriptome data (Garrison
et al., 2016; Fern
andez et al., 2018). In the previous
UCE-based phylogenomic studies using the Spider
probe set (Kulkarni et al., 2020a, b), Eresidae was
recovered as closely related to the RTA clade rather
than Araneoidea (Fig. 3). However, our analyses with
the newly developed probe set suggested that Eresidae
is more closely related to Araneoidea than the RTA
clade (Fig. 3), which was also supported from
transcriptome-based phylogenomic analyses of spiders
(Garrison et al., 2016; Fern
andez et al., 2018). The
monophyly of the RTA clade has been supported by
both Sanger-based sequence data and the genomescale datasets, but the placement of RTA clade on the
spider phylogeny and the relationships within the
RTA clade differ dramatically among studies and
analyses (Miller et al., 2010; Agnarsson et al., 2013;
Moradmand et al., 2014; Wheeler et al., 2017;
Fern
andez et al., 2018; Kulkarni et al., 2020b).
Although the concatenated analyses of this study converged to the same topology regarding the relationships of the five RTA families included (Dictynidae,
Lycosidae, Cheiracanthiidae, Clubionidae and Salticidae) in the 19-genome datasets, some ASTRAL analyses revealed different relationships (Fig. 2 and
Fig. S5). In the 57-species dataset with more RTA
families included, the phylogenetic relationships among
major lineages, such as Zodariidae, Sparassidae, the
Marronoid clade, the Oval Calamistrum clade and the
Dionycha, are consistent with the recent phylogenomic
studies (Kulkarni et al., 2020b; Azevedo et al., 2022).
More jumping spider species were sampled in this
study to ensure that this new probe kit is appropriate
for a recently launched project to resolve the phylogeny and evolution of Salticidae. The recovered relationships among the sampled salticids are largely
congruent with previous studies (e.g. Maddison, 2015;
Maddison et al., 2017, 2020a), but with increased support. A major difference among results from different
analyses involves the relationships of the three taxa in
the tribe Euophryini (Corythalia opima, Maratus sp.
and Parabathippus shelfordi) (Fig. 2, Figs S4 and S5).
Euophryini is the most diverse tribe in jumping spiders
with about 120 genera and over 1000 species reported
worldwide (Zhang and Maddison, 2013, 2015). A
much denser taxon sampling will be needed to clarify
the relationships within this lineage.
In addition to UCE, other phylogenomic approaches
such as Anchored Hybrid Enrichment (AHE) and
9
transcriptomes have been applied in building the spider tree of life (Garrison et al., 2016; Hamilton
et al., 2016; Maddison et al., 2017; Fern
andez
et al., 2018; Leduc-Robert and Maddison, 2018; Opatova et al., 2019). The AHE probes for spiders can target up to 585 loci (Hamilton et al., 2016), which is
only about one-sixth of the number of loci that the
RTA_v2 probes can target. The limited number of
AHE loci may not be sufficient to resolve relationships
for lineages with rapid radiations such as jumping spiders. Transcriptomic data have provided incredible
insights on the phylogeny and evolution of spiders
(Fernandez et al., 2018; Leduc-Robert and Maddison, 2018). However, transcriptome-based phylogenomics using RNA-seq procedures needs RNA as a
template for library preparation and sequencing, and
therefore requires high-quality tissues or specimens
being flash-frozen in liquid nitrogen or directly preserved in RNAlater, which often prohibits the direct
utilization of museum collections and limits taxon
sampling in a phylogenomic study (Zhang and
Lai, 2020), whereas the target enrichment approaches,
such as UCE and AHE, use DNA for library preparation, and studies have shown that museum materials
with degraded DNA are still applicable (Wood
et al., 2018; Derkarabetian et al., 2019). Next generation sequencing of transcriptomes often needs high
sequencing depth to recover a more complete set of
single-copy genes in an organism, and therefore is
more expensive than the target-enrichment-based
approaches such as UCE and AHE (Zhang and
Lai, 2020). These drawbacks often prevent the utility
of transcriptomes on a large-scale phylogenomic project. Recently, Zhang et al. (2019) proposed a novel
phylogenomic pipeline (PLWS) extracting markers
from low-coverage whole-genome sequencing data,
which has proved to be of great value for phylogenomic studies of organisms with small genomes
(<1 Gbp) (Sun et al., 2020). Genome size varies across
spiders (Gregory and Shorthouse, 2003). For instance,
the genome of Latrodectus hesperus (Theridiidae) is
about 1.1 Gbp, whereas the assembled genome of
Acanthoscurria geniculata (Theraphosidae) is over 7
Gbp (see Table S1). Our study suggests that the species in the RTA clade tend to have rather large genomes:
the
published
genome
for
Pardosa
pseudoannulata is about 4 Gbp (Yu et al., 2019) and
the assembled genome for Evarcha albaria in this study
is about 3.5 Gbp (the genome survey indicated that its
actual genome size may reach 5 Gbp). This hinders the
direct application of PLWS in the phylogenomic studies of these spiders (Zhang and Lai, 2020).
A previous study by Hedin et al. (2019) showed that
most of the loci targeted by the Arachnida probe set
are in coding regions and some loci are actually different regions of the same coding gene. By mapping the
10960031, 0, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/cla.12523 by University Of California, Riverside, Wiley Online Library on [31/01/2023]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
Zhang J. et al. / Cladistics 0 (2023) 1–13
Zhang J. et al. / Cladistics 0 (2023) 1–13
representative UCE sequences to the annotated genome of A. aquatica, we find that the UCEs from the
RTA and Spider datasets are also largely coding genes
(Fig. 1c). We inspected the potential overlapping of
RTA probe set loci and removed 69 redundant UCEs
from the RTA probe set. Analysing the coding and
non-coding regions separately, the results show that
the coding regions usually result in a topology congruent with the full dataset and obtain higher support values for certain nodes than the non-coding regions
(Fig. 2, Figs S2–S4). The exonic nature of the RTA
and Spider UCEs enables the combination of vast data
sets of UCEs and transcriptomes, which facilitates an
expanded taxon sampling as well as the reconciliation
of UCE and transcriptome data (see the procedure
applied in Bossert et al., 2019; Kulkarni et al., 2020b).
An efficient probe set that can successfully capture
as many loci as possible in most of the sampled taxa is
key for a target-enrichment based approach including
UCE phylogenomics. Previous UCE probe design
often applied the procedure outlined in the PHYLUCE workflow (Faircloth, 2016). Here we employed
an additional optimization during the development of
the new probe set (Fig. S1). To potentially increase the
capture success of the probes in a wider range of taxa
and result in more complete phylogenetic dataset, we
incorporated the results from the in-silico test on the
19 genomes and excluded the probes for overlap and
unmapped loci and the loci with taxon occupancy
lower than 75%. Probes from the Spider probe set
were also added to this new probe set for the compatibility of data obtained from different probe sets. In
addition, the newly developed probe set was tested in
a wide range of spider taxa other than the RTA clade,
including Araneidae, Atypidae, Dysderidae, Eresidae,
Linyphiidae, Pholcidae, Tetragnathidae, Theridiidae,
Sicariidae and Theraphosidae. About 1500–3000 UCE
loci were recovered for these taxa and the resulting
phylogeny is consistent with previous studies. Therefore, this novel probe set can further advance study of
the spider tree of life as a whole.
Acknowledgements
This work was funded by the National Natural
Science Foundation of China to Junxia Zhang (grant
no. 32070422), the Advanced Talents Incubation Program of the Hebei University to Junxia Zhang (grant
no. 521000981324) and the Institute of Life Sciences
and Green Development of Hebei University. We
would like to express our gratitude to Dr Wayne P.
Maddison, Dr Feng Zhang, Dr Shahan Derkarabetian,
Dr Marshal C. Hedin and Dr Zheng Fan for their
help and suggestions for this study; to Jigang Li and
Wenwen Song for their assistance in computational
issues; to Dr Martın J. Ramırez for his valuable comments on the manuscript; to Dr Wayne P. Maddison
and Dr John M. Heraty for help with polishing the
English of the manuscript; to Dr Xiangbo Guo for
assistance in submitting data to NCBI; to the anonymous reviewers for their valuable comments and suggestions
to
improve
the
manuscript.
The
Supercomputer Centre at Hebei University provided
computational resources for part of the analyses.
Conflict of interest
None declared.
Data availability statement
The probe sequences of the “RTA_v2” probe set,
assembled genomes and the data matrices and trees
from this study are deposited in Dryad Digital Repository (https://doi.org/10.5061/dryad.xksn02vkj). The
sequenced reads are deposited in the NCBI Sequence
Read Archive as project PRJNA907200.
References
Aberer, A.J., Kobert, K. and Stamatakis, A., 2014. Exa Bayes:
Massively parallel bayesian tree inference for the whole-genome
era. Mol. Biol. Evol. 31, 2553–2556.
Agnarsson, I., Coddington, J.A. and Kuntner, M., 2013.
Systematics, progress in the study of spider diversity and
evolution. In: Penney, D. (Ed.), Spider Research in the 21st
Century. Siri Scientific Press, Manchester, pp. 58–111.
Azevedo, G.H.F., Bougie, T., Carboni, M., Hedin, M. and Ramırez,
M.J., 2022. Combining genomic, phenotypic and Sanger
sequencing data to elucidate the phylogeny of the two-clawed
spiders (Dionycha). Mol. Phylogenet. Evol. 166, 107327.
Ballesteros, J.A., Setton, E.V.W., L
opez, C.E.S., Arango, C.P.,
Brenneis, G., Brix, S., Corbett, K.F., Cano-S
anchez, E.,
Dandouch, M., Dilly, G.F., Eleaume, M.P., Gainett, G., Gallut,
C., McAtee, S., McIntyre, L., Moran, A.L., Moran, R., L
opezGonzalez, P.J., Scholtz, G., Williamson, C., Woods, H.A.,
Zehms, J.T., Wheeler, W.C. and Sharma, P.P., 2020.
Phylogenomic resolution of sea spider diversification through
integration of multiple data classes. Mol. Biol. Evol. 38,
686–701.
Blaimer, B.B., Brady, S.G., Schultz, T.R., Lloyd, M.W., Fisher, B.L. and
Ward, P.S., 2015. Phylogenomic methods outperform traditional
multi-locus approaches in resolving deep evolutionary history: A case
study of formicine ants. BMC Evol. Biol. 15, 271–285.
Blair, C., Bryson, R.W., Jr., Linkem, C.W., Lazcano, D., Klicka, J.
and McCormack, J.E., 2019. Cryptic diversity in the Mexican
highlands: Thousands of UCE loci help illuminate phylogenetic
relationships, species limits and divergence times of montane
rattlesnakes (Viperidae: Crotalus). Mol. Ecol. Resour. 19, 349–
365.
Bond, J.E. and Opell, D., 1998. Testing adaptive radiation and key
innovation hypotheses in spiders. Evolution 52, 403–144.
Borowiec, M.L., 2019. Spruceup: Fast and flexible identification,
visualization, and removal of outliers from large multiple
sequence alignments. J. Open Source Softw. 4, 1635.
Bossert, S., Murray, E.A., Almeida, E.A.B., Brady, S.G., Blaimer,
B.B. and Danforth, B.N., 2019. Combining transcriptomes and
10960031, 0, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/cla.12523 by University Of California, Riverside, Wiley Online Library on [31/01/2023]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
10
ultraconserved elements to illuminate the phylogeny of Apidae.
Mol. Phylogenet. Evol. 130, 121–131.
Branstetter, M.G. and Longino, J.T., 2019. Ultra-conserved element
phylogenomics of new world Ponera (Hymenoptera: Formicidae)
illuminates the origin and phylogeographic history of the
endemic exotic ant Ponera exotica. Insect Syst. Divers. 3, 1–13.
Branstetter, M.G., Longino, J.T., Ward, P.S. and Faircloth, B.C.,
2017. Enriching the ant tree of life: Enhanced UCE bait set for
genome-scale phylogenetics of ants and other Hymenoptera.
Methods Ecol. Evol. 8, 768–776.
Bushnell, B., 2014. BBtools. Retrieved from https://sourceforge.net/
projects/bbmap/ (accessed on September 20, 2021).
Chen, D., Brau, E.L., Forthman, M., Kimball, R.T. and Zhang, Z.,
2018. A simple strategy for recovering ultraconserved elements,
exons, and introns from low coverage shotgun sequencing of
museum specimens: Placement of the partridge genus
Tropicoperdix within the Galliformes. Mol. Phylogenet. Evol.
129, 304–314.
Chikhi, R. and Rizk, G., 2013. Space-efficient and exact de Bruijn
graph representation based on a Bloom filter. Algorithms Mol.
Biol. 8, 1–9.
Cruaud, A., Nidelet, S., Arnal, P., Weber, A., Fusu, L., Gumovsky,
A. and Rasplus, J.-Y., 2018. Optimised DNA extraction and
library preparation for minute arthropods: Application to target
enrichment in chalcid wasps used for biocontrol. Mol. Ecol.
Resour. 19, 702–710.
Derkarabetian, S., Starrett, J., Tsurusaki, N., Ubick, D., Castillo, S.
and Hedin, M., 2018. A stable phylogenomic classification of
Travunioidea (Arachnida, Opiliones, Laniatores) based on
sequence capture of ultraconserved elements. Zookeys 2018, 1–
36.
Derkarabetian, S., Benavides, L.R. and Giribet, G., 2019. Sequence
capture phylogenomics of historical ethanol-preserved museum
specimens: Unlocking the rest of the vault. Mol. Ecol. Resour.
19, 1531–1544.
Dimitrov, D. and Hormiga, H., 2020. Spider diversification through
space and time. Annu. Rev. Entomol. 7, 225–241.
Faircloth, B.C., 2016. PHYLUCE is a software package for the
analysis of conserved genomic loci. Bioinformatics 32, 786–788.
Faircloth, B.C., 2017. Identifying conserved genomic elements and
designing universal bait sets to enrich them. Methods Ecol. Evol.
8, 1103–1112.
Faircloth, B.C., McCormack, J.E., Crawford, N.G., Harvey, M.G.,
Brumfield, R.T. and Glenn, T.C., 2012. Ultraconserved elements
anchor thousands of genetic markers spanning multiple
evolutionary timescales. Syst. Biol. 61, 717–726.
Faircloth, B.C., Branstetter, M.G., White, N.D. and Brady, S.G.,
2015. Target enrichment of ultraconserved elements from
arthropods provides a genomic perspective on relationships
among Hymenoptera. Mol. Ecol. Resour. 15, 489–501.
Fern
andez, R., Kallal, R.J., Dimitrov, D., Ballesteros, J.A., Arnedo,
M.A., Giribet, G. and Hormiga, G., 2018. Phylogenomics,
diversification dynamics, and comparative transcriptomics across
the spider tree of life. Curr. Biol. 28, 1489–1497.e5.
Foelix, R.F., 1996. Biology of Spiders, 2nd edition. Oxford
University Press, New York, p. 330.
Forthman, M., Miller, C.W. and Kimball, R.T., 2019. Phylogenomic
analysis suggests Coreidae and Alydidae (Hemiptera:
Heteroptera) are not monophyletic. Zool. Scr. 48, 520–534.
Garrison, N.L., Rodriguez, J., Agnarsson, I., Coddington, J.A.,
Griswold, C.E., Hamilton, C.A., Hedin, M., Kocot, K.M.,
Ledford, J.M. and Bond, J.E., 2016. Spider phylogenomics:
Untangling the spider tree of life. PeerJ 4, e1719.
Girard, M.B., Elias, D.O., Azevedo, G., Bi, K., Kasumovic, M.M.,
Waldock, J.M. and Hedin, M., 2021. Phylogenomics of peacock
spiders and their kin (Salticidae: Maratus), with implications for
the evolution of male courtship displays. Biol. J. Linn. Soc. 132,
471–494.
Goloboff, P.A., Farris, J.S. and Nixon, K.C., 2008. TNT, a free
program for phylogenetic analysis. Cladistics 24, 774–786.
11
Gregory, T.R. and Shorthouse, D.P., 2003. Genome sizes of spiders.
J. Hered. 94, 285–290.
Guillory, W.X., Muell, M.R., Summer, K. and Brown, J.L., 2019.
Phylogenomic reconstruction of the Neotropical poison frogs
(Dendrobatidae) and their conservation. Diversity 11, 126–140.
Hamilton, C.A., Lemmon, A.R., Lemmon, E.M. and Bond, J.E.,
2016. Expanding anchored enrichment to resolve both deep and
shallow relationships within the spider tree of life. BMC Evol.
Biol. 16, 212.
Hedin, M., Derkarabetian, S., Alfaro, A., Ramırez, M.J. and Bond,
J.E., 2019. Phylogenomic analysis and revised classification of
atypoid mygalomorph spiders (Araneae, Mygalomorphae), with
notes on arachnid ultraconserved element loci. PeerJ 7, e6864.
Huang, W., Li, L., Myers, J.R. and Marth, J.T., 2012. ART: A
next-generation sequencing read simulator. Bioinformatics 28,
593–594.
Jesovnik, A., Sosa-Calvo, J., Lloyd, M.W., Branstetter, M.G., Fern
Andez,
F. and Schultz, T.R., 2017. Phylogenomic species
delimitation and host-symbiont coevolution in the fungusfarming ant genus Sericomyrmex Mayr (Hymenoptera:
Formicidae): Ultraconserved elements (UCEs) resolve a recent
radiation. Syst. Entomol. 42, 523–542.
Junier, T. and Zdobnov, E.M., 2010. The Newick utilities: Highthroughput phylogenetic tree processing in the UNIX shell.
Bioinformatics 26, 1669–1670.
Kalyaanamoorthy, S., Minh, B.Q., Wong, T.K.F., von Haeseler, A.
and Jermiin, L.S., 2017. ModelFinder: Fast model selection for
accurate phylogenetic estimates. Nat. Methods 14, 587–589.
Katoh, K. and Standley, D.M., 2013. MAFFT multiple sequence
alignment software version 7: Improvements in performance and
usability. Mol. Biol. Evol. 30, 772–780.
Kearse, M., Moir, R., Wilson, A., Stones-Havas, S., Cheung, M.,
Sturrock, S., Buxton, S., Cooper, A., Markowitz, S., Duran, C.,
Thierer, T., Ashton, B., Meintjes, P. and Drummond, A., 2012.
Geneious basic: An integrated and extendable desktop software
platform for the organization and analysis of sequence data.
Bioinformatics 28, 1647–1649.
K€
uck, P. and Meusemann, K., 2010. FASconCAT: Convenient
handling of datamatrices. Mol. Phylogenet. Evol. 56, 1115–1118.
Kulkarni, S., Wood, H., Lloyd, M. and Hormiga, G., 2020a. Spiderspecific probe set for ultraconserved elements offers new
perspectives on the evolutionary history of spiders (Arachnida,
Araneae). Mol. Ecol. Resour. 20, 185–203.
Kulkarni, S., Kallal, R.J., Wood, H., Dimitrov, D., Giribet, G. and
Hormiga, G., 2020b. Interrogating genomic-scale data to resolve
recalcitrant nodes in the spider tree of life. Mol. Biol. Evol. 38,
891–903.
Leduc-Robert, G. and Maddison, W.P., 2018. Phylogeny with
introgression in Habronattus jumping spiders (Araneae:
Salticidae). BMC Evol. Biol. 18, 24–47.
Li, H., Handsaker, B., Wysoker, A., Fennell, T., Ruan, J., Homer, N.,
Marth, G., Abecasis, G., Durbin, R. and 1000 Genome Project
Data Processing Subgroup, 2009. The sequence alignment/map
(SAM) format and SAMtools. Bioinformatics 25, 2078–2079.
Lunter, G. and Goodson, M., 2011. Stampy: A statistical algorithm
for sensitive and fast mapping of Illumina sequence reads.
Genome Res. 21, 936–939.
Luo, R., Liu, B., Xie, Y., Li, Z., Huang, W., Yuan, J. and Wang, J.,
2012. SOAPdenovo2: An empirically improved memory-efficient
short-read de novo assembler. Gigascience 1, 1–6.
Maddison, W.P., 2015. A phylogenetic classification of jumping
spiders (Araneae: Salticidae). J. Arachnol. 43, 231–292.
Maddison, W.P., Evans, S.C., Hamilton, C.A., Bond, J.E., Lemmon,
A.R. and Lemmon, E.M., 2017. A genome-wide phylogeny of
jumping spiders (Araneae, Salticidae) using anchored hybrid
enrichment. Zookeys 695, 89–101.
Maddison, W.P., Beattie, I., Marathe, K., Ng, P.Y.C.,
Kanesharatnam, N., Benjamin, S.P. and Kunte, K., 2020a. A
phylogenetic and taxonomic review of baviine jumping spiders
(Araneae, Salticidae, Baviini). Zookeys 1004, 27–97.
10960031, 0, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/cla.12523 by University Of California, Riverside, Wiley Online Library on [31/01/2023]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
Zhang J. et al. / Cladistics 0 (2023) 1–13
Zhang J. et al. / Cladistics 0 (2023) 1–13
Maddison, W.P., Maddison, D.R., Derkarabetia, S. and Hedin, M.,
2020b. Sitticine jumping spiders: Phylogeny, classification, and
chromosomes (Araneae, Salticidae, Sitticini). Zookeys 925, 1–54.
Magalhaes, I.L.F., Azevedo, G.H.F., Michalik, P. and Ramırez,
M.J., 2020. The fossil record of spiders revisited: Implications for
calibrating trees and evidence for a major faunal turnover since
the Mesozoic. Biol. Rev. 95, 184–217.
McCormack, J.E., Faircloth, B.C., Crawford, N.G., Gowaty, P.A.,
Brumfield, R.T. and Glenn, T.C., 2012. Ultraconserved elements
are novel phylogenomic markers that resolve placental mammal
phylogeny when combined with species-tree analysis. Genome
Res. 22, 746–754.
McCormack, J.E., Harvey, M.J. and Faircloth, B.C., 2013. A
phylogeny of birds based on over 1,500 loci collected by target
enrichment and high-throughput sequencing. PLoS One 8, 51–67.
Miller, J.A., Carmichael, A., Ramırez, M.J., Spagna, J.C., Haddad,
C.R., Rezac, M., Johannesen, J., Kral, J., Wang, X.P. and
Griswold, C.E., 2010. Phylogeny of entelegyne spiders: Affinities
of the family Penestomidae (NEW RANK), generic phylogeny of
Eresidae, and asymmetric rates of change in spinning organ
evolution (Araneae, Araneoidea, Entelegynae). Mol. Phylogenet.
Evol. 55, 786–804. https://doi.org/10.1016/j.ympev.2010.02.021.
Minh, B.Q., Schmidt, H.A., Chernomor, O., Schrempf, D.,
Woodhams, M.D., von Haeseler, A. and Lanfear, R., 2020. IQTREE 2: New models and efficient methods for phylogenetic
inference in the genomic era. Mol. Biol. Evol. 37, 1530–1534.
Mirarab, S., Nguyen, N. and Warnow, T., 2014. PASTA: Ultralarge multiple sequence alignment. Res. Comput. Mol. Biol. 22,
177–191.
Moradmand, M., Sch€
onhofer, A.L. and J€ager, P., 2014. Molecular
phylogeny of the spider family Sparassidae with focus on the
genus Eusparassus and notes on the RTA-clade and
‘Laterigradae’. Mol. Phylogenet. Evol. 74, 48–65.
Ochoa, L.E., Datovo, A., DoNascimiento, C., Roxo, F.F., Sabaj,
M.H., Chang, J., Melo, B.F., Silva, G.S.C., Foresti, F., Alfaro,
M. and Oliveira, C., 2020. Phylogenomic analysis of
trichomycterid catfishes (Teleostei: Siluriformes) inferred from
ultraconserved elements. Sci. Rep. 10, 2697.
Opatova, V., Hamilton, C.A., Hedin, M., De Oca, L.M., Kral, J.
and Bond, J.E., 2019. Phylogenetic systematics and evolution of
the spider infraorder Mygalomorphae using genomic scale data.
Syst. Biol. 69, 671–707.
Pie, M.R., Bornschein, M.R., Ribeiro, L.F., Faircloth, B.C. and
McCormac, J.E., 2019. Phylogenomic species delimitation in
microendemic frogs of the Brazilian Atlantic Forest. Mol.
Phylogenet. Evol. 141, 106627.
Prjibelski, A., Antipov, D., Meleshko, D., Lapidus, A. and
Korobeynikov, A., 2020. Using SPAdes de novo assembler. Curr.
Protoc. Bioinformatics 70, e102.
Pryszcz, L.P. and Gabald
on, T., 2016. Redundans: An assembly
pipeline for highly heterozygous genomes. Nucleic Acids Res. 44,
e113–e123.
Quinlan, A.R. and Hall, I.M., 2010. BEDTools: A flexible suite of
utilities for comparing genomic features. Bioinformatics 26, 841–
842.
Ramırez, M.J., Magalhaes, I.L.F., Derkarabetian, S., Ledford, J.,
Griswold, C.E., Wood, H.M. and Hedin, M., 2020. Sequence
capture phylogenomics of true spiders reveals convergent
evolution of respiratory systems. Syst. Biol. 70, 14–20.
Richman, D.B. and Jackson, R.R., 1992. A review of the ethology
of jumping spiders (Araneae, Salticidae). Bull. Br. Arachnol. Soc.
9, 33–37.
Sahlin, K., Vezzi, F., Nystedt, B., Lundeberg, J. and Arvestad, L.,
2014. BESST-efficient scaffolding of large fragmented assemblies.
BMC Bioinformatics 15, 281–292.
Shao, L. and Li, S., 2018. Early Cretaceous greenhouse pumped
higher taxa diversification in spiders. Mol. Phylogenet. Evol. 127,
146–155.
Shen, W., Le, S. and Li, Y., 2016. SeqKit: A cross-platform and
ultrafast toolkit for fasta/q file manipulation. PLoS One 11,
e0163962.
Spagna, J.C. and Gillespie, R.G., 2008. More data, fewer shifts:
Molecular insights into the evolution of the spinning apparatus
in non-orb-weaving spiders. Mol. Phylogenet. Evol. 46, 347–368.
Stamatakis, A., 2014. RAxML version 8: A tool for phylogenetic
analysis and post-analysis of large phylogenies. Bioinformatics
30, 1312–1313.
Starrett, J., Derkarabetian, S., Hedin, M., Bryson, R.W.,
McCormack, J.E. and Faircloth, B.C., 2017. High phylogenetic
utility of an ultraconserved element probe set designed for
Arachnida. Mol. Ecol. Resour. 17, 812–823.
Sun, X., Ding, Y., Orr, M.C. and Zhang, F., 2020. Streamlining
universal single-copy orthologue and ultraconserved element
design: A case study in Collembola. Mol. Ecol. Resour. 20, 706–
717.
Torres, A., Goloboff, P.A. and Catalano, S.A., 2021. Parsimony
analysis of phylogenomic datasets (I): Scripts and guidelines for
using TNT (Tree Analysis using New Technology). Cladistics 38,
103–125.
Van Dam, M.H., Lam, A.W., Sagata, K., Gewa, B., Laufa, R.,
Balke, M., Faircloth, B.C. and Riedel, A., 2017. Ultraconserved
elements (UCEs) resolve the phylogeny of australasian smurfweevils. PLoS One 12, 1–21.
Van Dam, M.H., Trautwein, M., Spicer, G.S. and Esposito, L.,
2018. Advancing mite phylogenomics: Designing ultraconserved
elements for Acari phylogeny. Mol. Ecol. Resour. 19, 465–475.
Wheeler, W.C., Coddington, J.A., Crowley, L.M., Dimitrov, D.,
Goloboff, P.A., Griswold, C.E., Hormiga, G., Prendini, L.,
Ramırez, M.J., Sierwald, P., Almeida-Silva, L., Alvarez-Padilla,
F., Arnedo, M.A., Silva, L.R.B., Benjamin, S.P., Bond, J.E.,
Grismado, C.J., Hasan, E., Hedin, M., Izquierdo, M.A.,
Labarque, F.M., Ledford, J., Lopardo, L., Maddison, W.P.,
Miller, J.A., Piacentini, L.N., Platnick, N.I., Polotow, D., SilvaDavila, D., Scharff, N., Sz}
uts, T., Ubick, D., Vink, C.J., Wood,
H.M. and Zhang, J., 2017. The spider tree of life: Phylogeny of
Araneae based on target-gene analyses from an extensive taxon
sampling. Cladistics 33, 574–616.
Wood, H.M., Gonzalez, V.L., Lloyd, M., Coddington, J. and
Scharff, N., 2018. Next-generation museum genomics:
Phylogenetic relationships among palpimanoid spiders using
sequence capture techniques (Araneae: Palpimanoidea). Mol.
Phylogenet. Evol. 127, 907–918.
World Spider Catalog, 2022. World Spider Catalog. Version 23.5.
Natural History Museum Bern. Retrieved from http://wsc.nmbe.
ch (accessed on October 20, 2022).
Xu, X., Su, Y.-C., Ho, S.Y.W., Kuntner, M., Ono, H., Liu, F.,
Chang, C.-C., Warrit, N., Sivayyapram, V., Aung, K.P.P., Pham,
D.S., Norma-Rashid, Y. and Li, D., 2021. Phylogenomic analysis
of ultraconserved elements resolves the evolutionary and
biogeographic history of segmented trapdoor spiders. Syst. Biol.
70, 1110–1122.
Yu, N., Li, J., Liu, M., Huang, L.X., Bao, H.B., Yang, Z.M.,
Zhang, Y., Gao, H., Wang, Z., Yang, Y., Van Leeuwen, T.,
Millar, N.S. and Liu, Z.W., 2019. Genome sequencing and
neurotoxin diversity of a wandering spider Pardosa
pseudoannulata (pond wolf spider). bioRxiv, 747147. https://doi.
org/10.1101/747147.
Zhang, J. and Lai, J., 2020. Phylogenomic approaches in systematic
studies. Zool. Syst. 45, 151–162.
Zhang, J. and Maddison, W.P., 2013. Molecular phylogeny,
divergence times and biogeography of spiders of the subfamily
Euophryinae (Araneae: Salticidae). Mol. Phylogenet. Evol. 68,
81–92.
Zhang, J. and Maddison, W.P., 2015. Genera of euophryine jumping
spiders (Araneae: Salticidae), with a combined molecularmorphological phylogeny. Zootaxa 3938, 1–147.
Zhang, C., Rabiee, M., Sayyari, E. and Mirarab, S., 2018.
ASTRAL- III: Polynomial time species tree reconstruction from
partially resolved gene trees. BMC Bioinformatics 19, 15–30.
Zhang, F., Ding, Y., Zhu, C.-D., Zhou, X., Orr, M.C., Scheu, S.
and Luan, Y.-X., 2019. Phylogenomics from low-coverage wholegenome sequencing. Methods Ecol. Evol. 10, 507–517.
10960031, 0, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/cla.12523 by University Of California, Riverside, Wiley Online Library on [31/01/2023]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
12
Zhang, J., Lindsey, A.R.I., Peters, R.S., Heraty, J.M., Hopper,
K.R., Werren, J.H., Martinson, E.O., Woolley, J.B., Yoder, M.J.
and Krogmann, L., 2020. Conflicting signal in transcriptomic
markers leads to a poorly resolved backbone phylogeny of
chalcidoid wasps. Syst. Entomol. 45, 783–802.
Supporting Information
Additional supporting information may be found
online in the Supporting Information section at the
end of the article.
Fig. S1. Flowchart of probe design, in-silico test and
optimization procedure.
Fig. S2. The best trees from ML analyses on the full
loci concatenated datasets from RTA_v1, Spider and
Arachnida probes. The numbers along the branches
are shown as: ML bootstrap/posterior probability. The
scale bar is in substitutions per position.
Fig. S3. The best trees from ML analyses on the
concatenated coding loci datasets from RTA_v1, Spider and Arachnida probes. The numbers along the
branches are ML bootstrap. The scale bar is in substitutions per position.
Fig. S4. The best trees from ML analyses on the
concatenated non-coding loci datasets from RTA_v1,
Spider and Arachnida probes (the red branches indicate the different topologies from the concatenated full
loci ML analyses). The numbers along the branches
are ML bootstrap. The scale bar is in substitutions per
position.
Fig. S5. The ASTRAL species trees from the full
loci datasets from RTA_v1, Spider and Arachnida
probes (the red branches indicate the different
13
topologies from the concatenated analyses). The numbers along the branches are bootstrap values.
Fig. S6. Strict consensus of all equally parsimonious
trees from TNT analyses on datasets from RTA_v1,
Spider and Arachnida probes (the red branches indicate different topologies from the ML analyses). The
numbers along the branches are bootstrap support values (only show >75%).
Fig. S7. Maximum parsimonious tree from TNT
analysis on the dataset from RTA_v2 probes and 19
genomes (the red branches indicate different topologies
from the ML analyses). The numbers along the
branches are bootstrap support values (only show
>75%).
Fig. S8. Maximum parsimonious tree from TNT
analysis on the dataset from RTA_v2 probes and 57
species (the red branches indicate different topologies
from the ML analyses). The numbers along the
branches are bootstrap support values (only show
>75%).
Table S1. Specimen information and summary of
genome assembly and harvested UCE loci number
from different probe sets. *Denotes the species
enriched empirically in the lab.
Table S2. Number of UCE loci shared between taxa
during probe design. *Indicates that the set was chosen
for this probe design, 3900 loci shared by eight taxa.
Table S3. Summary statistics for the UCE datasets
in the concatenation analyses.
10960031, 0, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/cla.12523 by University Of California, Riverside, Wiley Online Library on [31/01/2023]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
Zhang J. et al. / Cladistics 0 (2023) 1–13