Cladistics (2023) 1–13 doi:10.1111/cla.12523 A novel probe set for the phylogenomics and evolution of RTA spiders Junxia Zhanga*, Zhaoyi Lia, Jiaxing Laia, Zhisheng Zhangb and Feng Zhanga* a Key Laboratory of Zoological Systematics and Application of Hebei Province, Institute of Life Science and Green Development, College of Life Sciences, Hebei University, Baoding, Hebei, 071002, China; bSchool of Life Sciences, Southwest University, Chongqing, 400700, China Accepted 21 December 2022 Abstract Spiders are important models for evolutionary studies of web building, sexual selection and adaptive radiation. The recent development of probes for UCE (ultra-conserved element)-based phylogenomic studies has shed light on the phylogeny and evolution of spiders. However, the two available UCE probe sets for spider phylogenomics (Spider and Arachnida probe sets) have relatively low capture efficiency within spiders, and are not optimized for the retrolateral tibial apophysis (RTA) clade, a hyperdiverse lineage that is key to understanding the evolution and diversification of spiders. In this study, we sequenced 15 genomes of species in the RTA clade, and using eight reference genomes, we developed a new UCE probe set (41 845 probes targeting 3802 loci, labelled as the RTA probe set). The performance of the RTA probes in resolving the phylogeny of the RTA clade was compared with the Spider and Arachnida probes through an in-silico test on 19 genomes. We also tested the new probe set empirically on 28 spider species of major spider lineages. The results showed that the RTA probes recovered twice and four times as many loci as the other two probe sets, and the phylogeny from the RTA UCEs provided higher support for certain relationships. This newly developed UCE probe set shows higher capture efficiency empirically and is particularly advantageous for phylogenomic and evolutionary studies of RTA clade and jumping spiders. © 2023 Willi Hennig Society. Introduction Spiders (Order Araneae) are among the most diverse terrestrial predators with over 50 000 species already described (World Spider Catalog, 2022) and many more waiting to be discovered (Agnarsson et al., 2013). As an ancient group, spiders can be dated back to the Devonian (>380 Ma), and have adapted to diverse ecosystems with remarkable behaviour and morphology during their evolutionary history. Over the years, arachnologists have worked long and hard to understand the diversification and evolutionary history of spiders (Garrison et al., 2016; Fernandez et al., 2018; Shao and Li, 2018; Dimitrov and Hormiga, 2020). *Corresponding author: E-mail address: [email protected]; [email protected] © 2023 Willi Hennig Society. Spiders are well known for their production of silk and the utility of foraging webs, which have been hypothesized as the key innovation for the diversification of spiders (Bond and Opell, 1998). However, a major radiation within spiders, the retrolateral tibial apophysis (RTA) clade, are mainly wandering hunters without foraging webs. Spiders in this clade are characterized by the presence of an RTA on the male palp for mating stabilization, and trichobothria on the tarsi and metatarsi for vibration sensitivity (Wheeler et al., 2017). This spider lineage is extremely diverse with over 25 000 described species (Dimitrov and Hormiga, 2020; World Spider Catalog, 2022), including the most species-rich spider family Salticidae (>6000 described species; World Spider Catalog, 2022), which are well known for their acute vision and spectacular courtship dances as well as the recently discovered milk provision (Richman and Jackson, 1992; Foelix, 1996; Chen et al., 2018). Recent divergence dating 10960031, 0, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/cla.12523 by University Of California, Riverside, Wiley Online Library on [31/01/2023]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License Cladistics Zhang J. et al. / Cladistics 0 (2023) 1–13 analyses suggested that the RTA clade is relatively young (139–161 Ma) compared with the Araneoidea clade of largely orb-weaving spiders, but the main drivers for its diversification remain contentious (Garrison et al., 2016; Fern andez et al., 2018; Shao and Li, 2018; Dimitrov and Hormiga, 2020; Magalhaes et al., 2020). In order to better understand the evolution of this major clade of spiders, we need to build on recent work (Miller et al., 2010; Agnarsson et al., 2013; Maddison, 2015; Wheeler et al., 2017; Azevedo et al., 2022) to resolve its phylogeny more fully and with better support. The rapid development of sequencing technology and analytical pipelines have strongly promoted progress in building the tree of life. Among various phylogenetic approaches applying genomic-scale data, the UCE (ultra-conserved element) method targets thousands of homologous loci by designing probes from highly conserved regions of representative taxa (Faircloth et al., 2012) and has been widely utilized in phylogenomic studies of vertebrates (e.g. McCormack et al., 2012, 2013; Guillory et al., 2019; Pie et al., 2019; Ochoa et al., 2020) and a variety of invertebrate groups, such as Arachnida (e.g. Starrett et al., 2017; Van Dam et al., 2018; Hedin et al., 2019; Kulkarni et al., 2020a), Pycnogonida (Ballesteros et al., 2020), Collembola (Sun et al., 2020), Hymenoptera (Faircloth et al., 2015; Branstetter et al., 2017; Cruaud et al., 2018) and Hemiptera (Forthman et al., 2019). This approach has proved to be successful for resolving both deep and shallow relationships (e.g. Blaimer et al., 2015; Branstetter et al., 2017; Jesovnik et al., 2017; Van Dam et al., 2017; Blair et al., 2019; Branstetter and Longino, 2019). Probes are critical for UCE phylogenomic approach. Currently two probe sets have been widely applied in spider phylogenomics, the Arachnida probes (Faircloth, 2017; Starrett et al., 2017) and the Spider probes (Kulkarni et al., 2020a). The Arachnida probe set was designed based on 10 exemplar taxa across Arachnida, including five spider species of Theridiidae, Eresidae, Sicariidae and Theraphosidae, and contains 14 799 probes targeting 1120 loci (Faircloth, 2017). This set of probes has been applied in phylogenetic analyses of Arachnida (Starrett et al., 2017), harvestmen (Derkarabetian et al., 2018, 2019) and spiders (e.g. Wood et al., 2018; Hedin et al., 2019; Ramırez et al., 2020; Maddison et al., 2020b; Azevedo et al., 2022). The Spider probe set was designed using four exemplar taxa of the spider families Theridiidae, Araneidae, Sicariidae and Eresidae, contains 15 051 probes harvesting 2021 UCEs (Kulkarni et al., 2020a), and has been used in phylogenetic studies of spiders (Kulkarni et al., 2020a, b) and Salticidae (Maddison et al., 2020a). Comparing the performance of these two probe sets, Kulkarni et al. (2020a) found that the Spider probe set captured more loci than the Arachnida probe set, and the phylogenetic tree inferred by the Spider UCEs gained higher bootstrap values for certain nodes, and therefore showed higher potential for solving some challenging relationships than the Arachnida probe set. However, taxa from the RTA clade were not included in the design of these two probe sets. In the study by Kulkarni et al. (2020a), UCEs were enriched for three species of the RTA clade using the Spider probe set, and about 800–1000 loci were obtained (<50% of the targeted 2021 loci). Applying these probe sets in jumping spiders, the Arachnida probes usually harvested 300–700 UCEs and the Spider probes 890–1200 UCEs (Maddison et al., 2020a, b). By designing probes specifically targeting a clade, we should be able to obtain more loci for resolving difficult phylogenetic relationships. For example, Xu et al. (2021) designed the liphistiidspecific probe set (19 740 probes targeting 3111 ultraconserved loci) that was streamlined for the UCE phylogenomics of the segmented trapdoor spiders (Suborder Mesothelae: Liphistiidae). In this study, we aim to develop a new UCE probe set to recover more loci and improve capture efficiency for resolving the recalcitrant relationships within the RTA clade of spiders and in particular within the family of jumping spiders (Salticidae). With the 15 genomes of the RTA clade sequenced and assembled in this study, in combination with the publicly available genomes, we (i) apply a modified probe design and optimization procedure to develop a new UCE probe set for spider phylogenomic and evolutionary studies (referred as RTA probe set); (ii) compare the performance of RTA probes in resolving the phylogeny of the RTA clade and Salticidae with Spider and Arachnida probes through an in-silico test on 19 genomes; and (iii) test empirically the applicability of the newly developed probe set over a broad set of spider lineages. Materials and methods A general flowchart for the analytical procedure of probe design, optimization and phylogenetic analyses applied in this study is provided in Fig. S1. Taxon sampling, DNA extraction and sequencing In total, 57 spider species of 27 families were included in this study, among which 40 species of 17 families belong to the RTA clade. See Table S1 for details about the species and specimen information. We sequenced genomes of 15 spider species of the RTA clade including the families Cheiracanthiidae (one species), Clubionidae (one species) and Salticidae (13 species). The jumping spiders were biased during genome sequencing and probe design because we are launching a large-scale project exploring the phylogeny and 10960031, 0, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/cla.12523 by University Of California, Riverside, Wiley Online Library on [31/01/2023]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License 2 3 evolution of this fascinating spider group. In addition, the sequencing data obtained at Southwest University (Professor Zhisheng Zhang) for genome annotation studies of Argyroneta aquatica (Clerck, 1757) (Dictynidae) and Pardosa laura Karsch, 1879 (Lycosidae), as well as the genome data of 12 spider species (three Araneidae, two Eresidae, three Theridiidae, one Lycosidae, one Dysderidae, one Sicariidae and one Theraphosidae) were also downloaded from NCBI to assess the applicability of the newly developed probe set in a wide range of spider lineages. Genomic DNA was extracted using QIAGEN DNeasy Blood & Tissue Kit, and 1 ll of RNase A (Solarbio) was added to the DNA extraction and then left at room temperature for 2 min to remove RNA. The quantity of DNA was checked using a QubitTM fluorometer. The genomic DNA was sent to Novogene Co. Ltd for library preparation using a Truseq Nano DNA HT sample preparation kit (Illumina USA), and then sequenced on an Illumina NovaSeq platform with 150 bp paired-end reads and insert size around 350 bp. sequences (160 bp per locus) for a temporary probe set (39 tiling density; repetitive regions <15%; 30% < GC content <70%; two probes per locus). The putative duplicated probes were identified and removed from the temporary probe set with the default setting (identity and coverage both as 50). The temporary probes were aligned back to all the reference genomes at a 50% sequence identity and the conserved loci were extracted (buffered to 180 bp) from all genomes. If the probe matched to different regions of the genome, the locus was deleted. Again, a database was built and the conserved loci shared by different number of taxa were calculated. The extracted conserved loci of all genomes shared by all of the eight reference taxa were used for a final probe design (39 tiling density; repetitive regions <15%; 30% < GC content <70%; two probes per locus), and the putative duplicated probes were identified and removed with identity and coverage both as 50. These probes were titled the “RTA_v1” probe set for clarity. Genome assembly In-silico test and comparison of the three probe sets Genome assembly followed the Phylogenomics from Lowcoverage Whole-genome Sequencing (PLWS) pipeline as in Zhang et al. (2019). In brief, the sequenced reads were first compressed into clumps with duplicates removed using clumpify.sh (BBTools) (Bushnell, 2014). Quality trimming of reads was completed using bbduk.sh (BBTools) with the reads shorter than 15 bp or with more than 5 Ns as well as the poly-A or poly-T tails of at least 10 bp being trimmed. The bbnorm.sh (BBTools) was then used to normalize the reads in order to accelerate the assembly. Genome contigs were assembled with multiple k-mer strategies in Minia v3.2.1 (Chikhi and Rizk, 2013). The contigs representing high heterozygosity were identified and deleted by Redundans v0.13c (Pryszcz and Gabald on, 2016). Contig scaffolding and gap filling were performed with BESST v2.2.8 (Sahlin et al., 2014) and GapCloser v1.12 in the SOAPdenovo2 suite (Luo et al., 2012) respectively. Sequences shorter than 500 bp were deleted by reformat.sh (Bushnell, 2014). We compared the performance of the “RTA_v1”, “Spider” and “Arachnida” probe sets through the in-silico test following the PHYLUCE workflow (Faircloth, 2016). The UCEs were harvested from 19 genomes (see Table S1) using the three probe sets, respectively. First, the genome data were converted from fasta to 2bit format using FaToTwoBit (http://hgdownload.soe.ucsc.edu/admin/exe/), and the corresponding sizes.tab was built with TwoBitInfo (https:// genome.ucsc.edu/goldenPath/help/twoBit.html). We then aligned each of the three sets of probes to the 19 genomes with the coverage and identity both as 75 and extracted 500 bp on either side. The extracted UCE loci were aligned to the corresponding probe set (min-coverage and min-identity both as 65) to remove duplicated UCEs and determine the final orthologues for phylogenetic analyses. The remaining UCE loci were imported into Geneious Prime v2019.1.3 (Kearse et al., 2012) and the loci with <15 taxa were excluded from downstream analyses. Sequence alignments were carried out using Mafft v7.313 (Katoh and Standley, 2013) with the LINS-I strategy. Probe design Eight genomes were used for identifying UCEs and designing probes: one each of Cheiracanthiidae, Clubionidae, Dictynidae, Lycosidae and Theridiidae, and three Salticidae (see Table S1). The probe design followed the PHYLUCE (Faircloth, 2016) workflow. We selected the jumping spider species Attulus fasciger (Simon, 1880) as the base genome because it was sequenced with a higher depth. Six genomes of the RTA clade were used as exemplars to represent the genetic diversity of this group, and the genome of the theridiid Parasteatoda tepidariorum was also included in probe design as outgroup. Short reads (100 bp paired-end and 29 coverage) of these assembled genomes were simulated with ART (Huang et al., 2012). The exemplar genomes were then aligned to the base genome using Stampy v1.0.32 (Lunter and Goodson, 2011) with the substitution rate as 0.05 and the insert size as 200. The resulting alignments were saved in BAM format and reduced using Samtools v1.10 (Li et al., 2009). BEDtools v2.28.0 (Quinlan and Hall, 2010) was used to convert the BAM files to BED format, which allowed us to sort and merge overlapping or nearly overlapping alignment positions. Alignments that were shorter than 80 bp or contained a high proportion of repetitive regions (>15% of length) or ambiguous (N) bases were removed. The retained alignments were put into an SQLite database for identification of the conserved loci shared between the base genome and a different number of exemplar genomes. We selected the conserved loci shared between the base genome and all seven exemplar genomes and extracted the corresponding base genome Overlap check and identification of coding regions A previous study found that some targeted UCEs were actually different regions of the same gene (Hedin et al., 2019), which could possibly cause redundancy or overlap of UCE loci in the phylogenetic dataset. We checked for overlap of UCEs in each of the three datasets obtained from different probe sets. First, the sequences of A. aquatica were extracted from the alignments of each dataset using Seqkit v0.13.2 (Shen et al., 2016). If the A. aquatica sequence in the alignment was missing, the sequence of its close relative species (Pardosa laura or Pardosa pseudoannulata) was extracted instead. These extracted sequences (with gaps at both ends converted to N and internal gaps removed) were then mapped to the annotated genome of A. aquatica (publication pending) in Geneious Prime v2019.1.3. We then inspected the mapping results to identify the UCEs that are overlapping or nearly overlapping with each other, of which we only retained the one with a longer sequence and more taxa. A few UCEs that were not able to be mapped to the annotated genome of A. aquatica (usually owing to the large number of Ns at the ends) were also excluded from downstream analyses. While conducting the overlap check, we also recorded if a UCE is in a coding or non-coding region according to the annotated genome. The proportion of coding vs. non-coding UCEs was then calculated for each dataset (after excluding the overlapping and unmapped UCEs). In addition, we checked the congruence between the RTA and Spider UCEs by cross-checking the mapping results 10960031, 0, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/cla.12523 by University Of California, Riverside, Wiley Online Library on [31/01/2023]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License Zhang J. et al. / Cladistics 0 (2023) 1–13 Zhang J. et al. / Cladistics 0 (2023) 1–13 and identified the UCE loci that are targeted by both the Spider and RTA probe sets. Optimization of the RTA probe set We conducted an additional optimization on the “RTA_v1” probe set to improve the capture efficiency and compatibility with data captured using other probe sets. First, the probes associated with the loci of low taxon-occupancy (<75%) as well as the overlapping and unmapped UCEs identified during the overlap-check were removed. The UCEs harvested from the 19 genomes (≥4 genomes) that were unique to the “Spider” probe set (UCEs also targeted by the “RTA_v1” probe set were excluded) and their associated probes for the two species P. tepidariorum (Theridiidae) and Stegodyphus mimosarum Pavesi, 1883 (Eresidae) were added to the optimized RTA probe set. Only the Spider probes were considered during the optimization since many of the Arachnida probes were already incorporated into the Spider probe set (Kulkarni et al., 2020a). This will help to efficiently integrate the data generated with different probe sets. After combining the remaining RTA probes with the Spider probes, we identified and removed the putative duplicated probes with identity and coverage both as 70. We added “20 000 000” to the UCE number associated with the Spider UCEs (e.g. “uce-7” converted to “uce-20000007”, “uce-10061” converted to “uce-20010061”) to differentiate them from the “RTA-specific” UCEs. The optimized probe set was labelled as “RTA_v2”. Test of RTA_v2 probe set The performance of the RTA_v2 probe set was tested on the 19 genomes focusing on the RTA clade and jumping spiders (see above) and a broader sampling of spiders that includes 57 species of Mygalomorphae and Aranemorphae (Synspermiata, Eresidae, Araneoidea and RTA clade). For the available genomic data, the UCEs were directly extracted from the genomes following the above protocol. To test the optimized probes empirically, the UCE loci were also enriched and sequenced for 28 species using the RTA_v2 probe set manufactured by Daicel Arbor Biosciences. For empirically capturing UCEs, the genomic DNA extracted from each specimen was first fragmentated by sonication (target size of approximately 300–600 bp). The fragmented DNA (21 lL) was then used as input for DNA library preparation with the NEXTFLEXÒ Rapid DNA-Seq Kit 2.0 (Bioo Scientific) following the manufacturer’s protocol with minor modifications. After the adapter ligation using the NEXTFLEXÒ Unique Dual Index Barcodes (Set C), a 0.89 beads clean was conducted (NEXTFLEX Cleanup Beads 2.0, Bioo Scientific) followed by a PCR reaction using the HiFi HotStart ReadyMix (Kapa Biosystems). The PCR system was as follows: 25 lL postligation library, 26 lL HiFi HotStart ReadyMix and 2 lL primer mix. The following thermal protocol was applied: 98°C for 45 s; 18 cycles of 98°C for 15 s, 60°C for 30 s and 72°C for 1 min; and final extension at 72°C for 5 min. PCR clean-up was done with a 0.89 beads clean. The 28 libraries were divided into three pools with nine or 10 libraries being combined into one pool (final volume of 7 lL and final concentration of 171.6–208 ng/lL) at equimolar ratios for UCE enrichment following the myBaits protocol 5.01 (Daicel Arbor Biosciences). For one pool (nine libraries), we conducted the UCE enrichment using the RTA_v2 probes and the Spider probes independently in order to directly compare the capture efficiency of the two probe sets; the other two pools were only captured with the RTA_v2 probes. The enriched UCE libraries were then sent to Novogene Co. Ltd for sequencing using the Illumina NovaSeq platform with 150 bp paired-end reads. The adapters and low-quality bases were removed from the sequenced raw reads for each species using the bbduk.sh (BBTools). The trimmed reads were then assembled using SPAdes v3.14.1 (Prjibelski et al., 2020) with “--cov-cutoff auto”. The assembled contigs shorter than 200 bp were removed for subsequent analyses using reformat.sh. Finding and extracting UCEs from the assembled contigs followed the PHYLUCE workflow with min-coverage and minidentity both set to 65. The UCEs extracted from genomes and target enrichment data were combined and organized by locus, and then aligned using Mafft v7.313 with the L-INS-I strategy. Phylogenomic analyses In total, 11 datasets were generated for phylogenetic reconstruction: nine were from the in-silico test and comparison of three probe sets (RTA_v1_Full_UCE, Spider_Full_UCE, Arachnida_Full _UCE, RTA_v1_Coding_UCE, Spider_Coding_UCE, Arachnida_Coding _UCE, RTA_v1_Non-coding_UCE, Spider_Non-coding_UCE and Arachnida_Non-Coding _UCE) and two from the final optimized RTA_v2 probe set (RTA_v2_19genomes_UCE and RTA_v2_57spp_UCE). The combination of Spruceup v2020.2.19 (Borowiec, 2019) and Seqtools (PASTA; Mirarab et al., 2014) was used for alignment trimming. First, we applied Spruceup v2020.2.19 to convert the obviously misaligned fragments in each alignment to gaps (cutoffs as 0.85). The gappy regions in each alignment were then masked using Seqtools in PASTA package with “masksites = 30” for the “RTA_v2_57spp_UCE” dataset and “masksites = 10” for all other datasets. The “gene tree and alignment” method as in Zhang et al. (2020) was also applied to check and remove the putative paralogous or contamination sequences in each dataset. A gene tree was first constructed for each MSA (Multiple Sequence Alignment) using RAxML v8.2.12 (Stamatakis, 2014) with the GTRGAMMA model. Gene trees were inspected using the customized Python script (Zhang et al., 2020) to flag taxa with odd sequences that may be subject to contamination errors (identical sequences for unrelated taxa) or resulted in abnormally long branches on the gene tree. These flagged sequences were then removed from the corresponding MSAs. For the “RTA_v2_57spp_UCE” dataset, alignments with <30 taxa or 150 bp were removed from subsequent phylogenetic reconstruction analyses. All UCE loci for each dataset were concatenated by FASconCAT v1.0 (K€ uck and Meusemann, 2010). Maximum likelihood (ML) analyses were performed on each concatenated matrix using IQ-TREE v2.0.6 (Minh et al., 2020). The best-fitting model and optimized partition scheme were inferred for each supermatrix in IQ-TREE v2.0.6 using the option “-m MF+MERGE”. For each dataset, we ran 40 independent ML tree searches (20 with random starting trees and 20 with parsimonious starting trees) in IQ-TREE v2.0.6. Nonparametric bootstrap analyses (100 replicates) were conducted to assess node support. Bayesian inferences were also conducted for the concatenated full loci matrices in ExaBayes v1.5.1 (Aberer et al., 2014). The Markov Chain Monte Carlo (MCMC) chain of each Bayesian analysis was set for 200 000 generations (“RTA_v2_57spp_UCE” dataset) or 150 000 generations (the other datasets) and two independent runs were employed. Trees were sampled every 500 generations. The Effective Sampling Size (ESS) and Potential Scale Reduction Factor (PSRF) were inspected to ensure the convergence following the ExaBayes manual. The first 25% of sampled trees were discarded as burn-in. The coding and non-coding regions were also concatenated and then analysed separately in IQTREE v2.0.6 (following above procedure) to inspect their performance on phylogenetic reconstruction. Maximum parsimony (MP) analyses were conducted on each concatenated dataset using TNT (Goloboff et al., 2008), and the recently developed scripts for running TNT analyses on phylogenomic dataset were applied (Torres et al., 2021). The MP tree searches were carried out with “New Technology” search scheme, searching level at 5 and number of hits 10960031, 0, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/cla.12523 by University Of California, Riverside, Wiley Online Library on [31/01/2023]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License 4 to best length at 3, and 500 replicates of bootstrap analyses with a fast approximate resampling algorithm were conducted to assess node support (searching level at 3). The coalescent-based species-tree method to account for potential gene tree heterogeneity and discordance was also applied to the three concatenated full loci datasets for the comparison of three different probe sets. First, the ML tree and 100 non-parametric bootstrap replicates were inferred for each alignment in IQ-TREE v.2.0.6 using the best-fitting model selected by ModelFinder (Kalyaanamoorthy et al., 2017). For each gene tree, the branches with bootstrap ≤30% were collapsed by Newick Utils v1.6 (Junier and Zdobnov, 2010). The Accurate Species Tree Algorithm (ASTRAL-III v5.7.1; Zhang et al., 2018) was then applied to estimate the species tree with 100 replicates of bootstrapping to assess the node support. Results Genome assembly and probe design Genome sequencing and assembly results were provided in Table S1. About 20–90 million reads were generated for the genome assembly of selected taxa. The assembled genome size varies across taxa (0.6–3.5 Gb), which may be affected by the actual genome size and sequencing depth, and the GC contents of the assembled genomes are all around 30%. The UCE number shared among different numbers of taxa during probe design is presented in Table S2. We selected the 3900 UCEs shared by Attulus fasciger and all seven exemplar taxa for the probe design. The “RTA_v1” probe set contains 60 018 probes targeting for 3856 UCEs, which was applied in the subsequent in-silico test and comparison of three probe sets. In-silico test and comparison of three probe sets The number of UCE loci harvested from different probe sets for each taxon is shown in Table S1. On average, ~700 UCEs were obtained from the Arachnida probe set, ~1340 UCEs from the Spider probe set and ~2600 UCEs from the “RTA_v1” probe set. The “RTA_v1” probes harvest about twice as many UCEs as the Spider probes and four times as many as the Arachnida probes (Fig. 1a; Table S1). The statistics for the UCE datasets from different probes are provided in Table S3. After the overlap check, the RTA_v1, Spider and Arachnida datasets contain 2347, 961 and 424 UCEs respectively (matrix with ~75% of taxon-completeness). The total sites and parsimony-informative sites of the supermatrix of the RTA_v1 dataset are both about six times the Arachnida supermatrix and three times the Spider supermatrix. The missing data of the three datasets are similar (Fig. 1b) with the alignments in the Spider dataset having the lowest average missing data (Spider dataset, 10.28%, ranging from 0.42 to 21.70%; RTA_v1 5 dataset, 11.51%, ranging from 0.03 to 23.24%; Arachnida dataset, 11.91%, ranging from 1.10 to 26.72%). Mapping the representative UCE sequences from each dataset to the annotated A. aquatica genome, the results show that most of the UCE loci (over twothirds) in each dataset are in coding regions. The RTA_v1 dataset contains a lower proportion (69%) of coding UCE loci than the Spider (85%) and Arachnida (90%) dataset (Fig. 1c). Optimization of the RTA probe set After deleting the probes associated with lowoccupancy, overlapping and unmapped UCEs and adding the probes for the “Spider-specific” UCEs, the final optimized RTA probe set (“RTA_v2”) contains 41 845 probes targeting 3802 loci, of which 36 385 probes are for the RTA UCEs (2334 loci) and 5460 probes are associated with the “Spider-specific” UCEs (1468 loci). On average, 3058 UCEs were obtained from the 29 genomes and 2485 UCEs from the 28 species enriched empirically using the RTA_v2 probe set (Table S1). For the nine samples that were enriched with both probe sets, on average the RTA_v2 probes (~2597 UCEs) captured about 2.5 times as many loci as the Spider probes (~1006 UCEs), and the average capture success rate for the RTA_v2 probes was higher than that of the Spider probes (68.3 vs. 49.8%) (Table S1). Phylogenetic inference The results of phylogenetic inference from different datasets and methods are provided in Figs 2 and 3 and Figs S2–S8. For comparison of the three probe sets, the ML analyses on the concatenated full and coding datasets recovered the same topology (Fig. 2) and the RTA_v1 dataset shows higher average bootstrap support than the corresponding Spider or Arachnida dataset (Table S3). In addition, the RTA_v1 data show higher support for some nodes, e.g. the clade with sampled Dionycha taxa (RTA_v1, 100% in all sets vs. Arachnida, 90% in the full set and coding set) and the node with Myrmarachne formicaria (JXZ414), Marpissa milleri (JXZ425) and Mendoza nobilis (JXZ419) (RTA_v1, 100% in all sets vs. Spider, 98% in the full set and 86% in the coding set; Arachnida, 98% in the full set). The concatenated non-coding loci from the three datasets all fail to recover certain nodes. For example, the clade with Myrmarachne formicaria (JXZ414), Marpissa milleri (JXZ425) and Mendoza nobilis (JXZ419) was not recovered by the non-coding loci from the Arachnida set, and the node with Maratus sp. (JXZ159a) and Parabathippus shelfordi (JXZ417) was not recovered by the non-coding loci from both RTA_v1 and Spider sets. For the 10960031, 0, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/cla.12523 by University Of California, Riverside, Wiley Online Library on [31/01/2023]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License Zhang J. et al. / Cladistics 0 (2023) 1–13 Zhang J. et al. / Cladistics 0 (2023) 1–13 RTA_v1 probes (a) Spider probes Arachnida probes (b) RTA_v1 Probes Number of loci Spider Probes 4000 0.09 Arachnida Probes 2000 Pa Density rt r ge ia h egiu lid ete s es r ag oide or a ifo rm is na h Pa n ra ne f obil ba th orm is Pl ippu icar ex s i ip she a p o lf o i Sy Po des rdi ra ste at od a Ar tepi gy da 0 ro rio ne ta rum Pa r P aq Cl dos ard uat ub a p os ica io a s na eud la u Ch pse oan ra ud n ei ra og ula ca t e nt rm a an hi u ic At m in a tu s ig lu n At s fa e tu sc l i Co us s ger in ry en th a sis Ev lia H arc opim ab ha a ro a na lb ttu ari so a ph r M Mar ys ar atu pi ss ss p. a M M m yrm en il ar doz leri ac a 0.06 0.03 (c) 15% 31% 69% 10% 85% 90% 0.00 RTA_v1 Probes Coding Region Spider Probes Arachnida Probes 0 10 20 Missing percentage Non-coding Region Fig. 1. Comparison of UCE (ultra-conserved element) data from RTA, Spider and Arachnid probe sets. (a) Number of UCE loci extracted from 19 genomes; (b) distribution of missing percentage in the alignments of three datasets; and (c) proportion of coding vs. non-coding loci in three datasets. deeper nodes the non-coding loci tend to show lower support than the coding loci. The Bayesian analyses on the three concatenated full loci datasets recovered the same topology as in Fig. 2 with posterior probability for all the nodes being 1.0. However, the MP analyses often recovered different topologies as in Fig. 2, for instance the placement of Clubiona pseudogermanica (YCH205) and Myrmarachne formicaria (JXZ414), sometimes with relatively high bootstrap support (Fig. S6). The species trees from ASTRAL analyses sometimes show different relationships for the deeper nodes (Fig. S5). For instance, Clubiona pseudogermanica (YCH205) was placed as the sister to the clade with Argyroneta and Pardosa in the species tree based on the RTA_v1 dataset rather than within the Dionycha clade as in the concatenated analyses. Within Salticidae, the species tree from the Spider dataset shows the sister relationship of Maratus sp. (JXZ159a) with Corythalia opima (JXZ418) but with low bootstrap support (31%). The concatenated dataset of 19 genomes from the final optimized probe set (“RTA_v2_19genomes_UCE”) with about 50% of taxon-completeness contains 3645 UCE loci (Table S3), and the ML analysis resulted in the same topology as that from the RTA_v1 concatenated dataset but with all nodes having bootstrap supports of 100% and posterior probabilities of 1.0 (Fig. 2). The concatenated dataset of 57 species (“RTA_v2_57spp_UCE”) with about 50% of taxon-completeness contains 3271 UCE loci (Table S3), and the results from ML and Bayesian analyses are shown in Fig. 3. The relationships within the RTA clade are congruent with the tree from the “RTA_v2_19genomes_UCE” dataset. In addition, the MP analysis on the “RTA_v2_19genomes_UCE” dataset recovered different placements for Clubiona pseudogermanica (YCH205) and Myrmarachne formicaria (JXZ414) with strong bootstrap supports (100%); the MP analysis on the “RTA_v2_57spp_UCE” dataset recovered slightly different placements for Orthobula crucifera (JXZ601) and Parabathippus shelfordi (JXZ417) (Figs S7 and S8). Discussion This study aims to present a new UCE probe set that provides more loci for future comprehensive phylogenomic and evolutionary studies of spiders. The 10960031, 0, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/cla.12523 by University Of California, Riverside, Wiley Online Library on [31/01/2023]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License 6 7 Fig. 2. Summary of phylogenetic inferences from different datasets and analytical methods. The tree shown is the ML tree (log-likelihood = 24 526 285.98) on the concatenated dataset (3645 loci) from the RTA_v2 probe set with the black circles indicating that the bootstrap supports and posterior probabilities are 100% and 1.0 respectively for all the nodes. The scale bar is in substitutions per position. final optimized probe set (RTA_v2) can target up to 3802 UCEs, about two to three times as many as the previously available Spider (up to 2021 loci) and Arachnida (up to 1120 loci) probe sets (Faircloth, 2017; Starrett et al., 2017; Kulkarni et al., 2020a). We also accommodated the loci targeted by the other probe sets in this newly designed probe set, so that the data obtained in previous studies (e.g. Starrett et al., 2017; Wood et al., 2018; Kulkarni et al., 2020a; Maddison et al., 2020a; Girard et al., 2021) can be efficiently integrated with sequences captured using the new probe set. We specifically tailored the new probe set for the RTA clade (spiders with retrolateral tibial apophyses on the male palpi; Spagna and Gillespie, 2008) and jumping spiders (a major radiation within the RTA clade). The RTA clade is hyperdiverse, containing more than half of the recorded spider species. However, neither the Arachnida nor the Spider probe set included species of the RTA clade during probe design largely owing to a lack of genomes of RTA spiders (Faircloth, 2017; Kulkarni et al., 2020a). The target efficiency of these probe sets in RTA species is low. For instance, in jumping spiders (Salticidae) the Arachnida probes can harvest 300–700 UCEs (Maddison et al., 2020a, b), and the Spider probes can harvest 890–1200 UCEs (Maddison et al., 2020a), at most around 60% of the total targeted loci. Using the optimized probe set (RTA_v2), up to over 82% of the targeted loci (3802 UCEs) were successfully enriched. Direct comparison of the capture efficiency using nine libraries shows that the new probe set (RTA_v2) could obtain from RTA spiders on average ~1600 more UCEs with ~20% higher capture efficiency than the Spider probe set. Low capture efficiency usually results in phylogenetic datasets with fewer loci and more missing data. For instance, the dataset generated from the spider probes in Kulkarni et al. (2020a) had 1010 loci with 25% of taxon-occupancy, but when the taxon-occupancy requirement was increased to 50%, only 276 loci were retained. Here, the dataset generated from the RTA_v2 probes and 57 species with 10960031, 0, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/cla.12523 by University Of California, Riverside, Wiley Online Library on [31/01/2023]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License Zhang J. et al. / Cladistics 0 (2023) 1–13 Zhang J. et al. / Cladistics 0 (2023) 1–13 Fig. 3. The ML tree (log-likelihood = 23 158 038.986) on the concatenated dataset from the RTA_v2 probe set and 57 spider species with the circles at the nodes indicating the bootstrap supports (posterior probabilities for all the nodes are 1.0). The * along the taxon name indicates that the UCEs were empirically captured using the manufactured probes for the species. The schematic tree shown at the left bottom corner is modified from Kulkarni et al. (2020b), indicating the different placement of Eresidae on the phylogeny. The scale bar is in substitutions per position. about 50% taxon occupancy contained over 3000 loci (Table S3). With relatively poorly assembled genomes (e.g. Habronattus ophrys and Plexippoides regius), the RTA_v2 probes can harvest many more loci than the Spider (~800 more loci) and Arachnida (~1200 more loci) probes (Table S1). The analyses in this study with relatively few exemplar taxa are not intended to tackle the problematic 10960031, 0, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/cla.12523 by University Of California, Riverside, Wiley Online Library on [31/01/2023]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License 8 nodes in the spider phylogeny, but to show that the new probe set from this study has great potential to provide more sequence data for resolving the recalcitrant relationships and exploring evolutionary questions within spiders, especially the RTA clade and Salticidae. The relationships of major spider lineages resulting from the 57-species dataset is consistent with previous studies using transcriptome data (Garrison et al., 2016; Fern andez et al., 2018). In the previous UCE-based phylogenomic studies using the Spider probe set (Kulkarni et al., 2020a, b), Eresidae was recovered as closely related to the RTA clade rather than Araneoidea (Fig. 3). However, our analyses with the newly developed probe set suggested that Eresidae is more closely related to Araneoidea than the RTA clade (Fig. 3), which was also supported from transcriptome-based phylogenomic analyses of spiders (Garrison et al., 2016; Fern andez et al., 2018). The monophyly of the RTA clade has been supported by both Sanger-based sequence data and the genomescale datasets, but the placement of RTA clade on the spider phylogeny and the relationships within the RTA clade differ dramatically among studies and analyses (Miller et al., 2010; Agnarsson et al., 2013; Moradmand et al., 2014; Wheeler et al., 2017; Fern andez et al., 2018; Kulkarni et al., 2020b). Although the concatenated analyses of this study converged to the same topology regarding the relationships of the five RTA families included (Dictynidae, Lycosidae, Cheiracanthiidae, Clubionidae and Salticidae) in the 19-genome datasets, some ASTRAL analyses revealed different relationships (Fig. 2 and Fig. S5). In the 57-species dataset with more RTA families included, the phylogenetic relationships among major lineages, such as Zodariidae, Sparassidae, the Marronoid clade, the Oval Calamistrum clade and the Dionycha, are consistent with the recent phylogenomic studies (Kulkarni et al., 2020b; Azevedo et al., 2022). More jumping spider species were sampled in this study to ensure that this new probe kit is appropriate for a recently launched project to resolve the phylogeny and evolution of Salticidae. The recovered relationships among the sampled salticids are largely congruent with previous studies (e.g. Maddison, 2015; Maddison et al., 2017, 2020a), but with increased support. A major difference among results from different analyses involves the relationships of the three taxa in the tribe Euophryini (Corythalia opima, Maratus sp. and Parabathippus shelfordi) (Fig. 2, Figs S4 and S5). Euophryini is the most diverse tribe in jumping spiders with about 120 genera and over 1000 species reported worldwide (Zhang and Maddison, 2013, 2015). A much denser taxon sampling will be needed to clarify the relationships within this lineage. In addition to UCE, other phylogenomic approaches such as Anchored Hybrid Enrichment (AHE) and 9 transcriptomes have been applied in building the spider tree of life (Garrison et al., 2016; Hamilton et al., 2016; Maddison et al., 2017; Fern andez et al., 2018; Leduc-Robert and Maddison, 2018; Opatova et al., 2019). The AHE probes for spiders can target up to 585 loci (Hamilton et al., 2016), which is only about one-sixth of the number of loci that the RTA_v2 probes can target. The limited number of AHE loci may not be sufficient to resolve relationships for lineages with rapid radiations such as jumping spiders. Transcriptomic data have provided incredible insights on the phylogeny and evolution of spiders (Fernandez et al., 2018; Leduc-Robert and Maddison, 2018). However, transcriptome-based phylogenomics using RNA-seq procedures needs RNA as a template for library preparation and sequencing, and therefore requires high-quality tissues or specimens being flash-frozen in liquid nitrogen or directly preserved in RNAlater, which often prohibits the direct utilization of museum collections and limits taxon sampling in a phylogenomic study (Zhang and Lai, 2020), whereas the target enrichment approaches, such as UCE and AHE, use DNA for library preparation, and studies have shown that museum materials with degraded DNA are still applicable (Wood et al., 2018; Derkarabetian et al., 2019). Next generation sequencing of transcriptomes often needs high sequencing depth to recover a more complete set of single-copy genes in an organism, and therefore is more expensive than the target-enrichment-based approaches such as UCE and AHE (Zhang and Lai, 2020). These drawbacks often prevent the utility of transcriptomes on a large-scale phylogenomic project. Recently, Zhang et al. (2019) proposed a novel phylogenomic pipeline (PLWS) extracting markers from low-coverage whole-genome sequencing data, which has proved to be of great value for phylogenomic studies of organisms with small genomes (<1 Gbp) (Sun et al., 2020). Genome size varies across spiders (Gregory and Shorthouse, 2003). For instance, the genome of Latrodectus hesperus (Theridiidae) is about 1.1 Gbp, whereas the assembled genome of Acanthoscurria geniculata (Theraphosidae) is over 7 Gbp (see Table S1). Our study suggests that the species in the RTA clade tend to have rather large genomes: the published genome for Pardosa pseudoannulata is about 4 Gbp (Yu et al., 2019) and the assembled genome for Evarcha albaria in this study is about 3.5 Gbp (the genome survey indicated that its actual genome size may reach 5 Gbp). This hinders the direct application of PLWS in the phylogenomic studies of these spiders (Zhang and Lai, 2020). A previous study by Hedin et al. (2019) showed that most of the loci targeted by the Arachnida probe set are in coding regions and some loci are actually different regions of the same coding gene. By mapping the 10960031, 0, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/cla.12523 by University Of California, Riverside, Wiley Online Library on [31/01/2023]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License Zhang J. et al. / Cladistics 0 (2023) 1–13 Zhang J. et al. / Cladistics 0 (2023) 1–13 representative UCE sequences to the annotated genome of A. aquatica, we find that the UCEs from the RTA and Spider datasets are also largely coding genes (Fig. 1c). We inspected the potential overlapping of RTA probe set loci and removed 69 redundant UCEs from the RTA probe set. Analysing the coding and non-coding regions separately, the results show that the coding regions usually result in a topology congruent with the full dataset and obtain higher support values for certain nodes than the non-coding regions (Fig. 2, Figs S2–S4). The exonic nature of the RTA and Spider UCEs enables the combination of vast data sets of UCEs and transcriptomes, which facilitates an expanded taxon sampling as well as the reconciliation of UCE and transcriptome data (see the procedure applied in Bossert et al., 2019; Kulkarni et al., 2020b). An efficient probe set that can successfully capture as many loci as possible in most of the sampled taxa is key for a target-enrichment based approach including UCE phylogenomics. Previous UCE probe design often applied the procedure outlined in the PHYLUCE workflow (Faircloth, 2016). Here we employed an additional optimization during the development of the new probe set (Fig. S1). To potentially increase the capture success of the probes in a wider range of taxa and result in more complete phylogenetic dataset, we incorporated the results from the in-silico test on the 19 genomes and excluded the probes for overlap and unmapped loci and the loci with taxon occupancy lower than 75%. Probes from the Spider probe set were also added to this new probe set for the compatibility of data obtained from different probe sets. In addition, the newly developed probe set was tested in a wide range of spider taxa other than the RTA clade, including Araneidae, Atypidae, Dysderidae, Eresidae, Linyphiidae, Pholcidae, Tetragnathidae, Theridiidae, Sicariidae and Theraphosidae. About 1500–3000 UCE loci were recovered for these taxa and the resulting phylogeny is consistent with previous studies. Therefore, this novel probe set can further advance study of the spider tree of life as a whole. Acknowledgements This work was funded by the National Natural Science Foundation of China to Junxia Zhang (grant no. 32070422), the Advanced Talents Incubation Program of the Hebei University to Junxia Zhang (grant no. 521000981324) and the Institute of Life Sciences and Green Development of Hebei University. We would like to express our gratitude to Dr Wayne P. Maddison, Dr Feng Zhang, Dr Shahan Derkarabetian, Dr Marshal C. Hedin and Dr Zheng Fan for their help and suggestions for this study; to Jigang Li and Wenwen Song for their assistance in computational issues; to Dr Martın J. Ramırez for his valuable comments on the manuscript; to Dr Wayne P. Maddison and Dr John M. Heraty for help with polishing the English of the manuscript; to Dr Xiangbo Guo for assistance in submitting data to NCBI; to the anonymous reviewers for their valuable comments and suggestions to improve the manuscript. The Supercomputer Centre at Hebei University provided computational resources for part of the analyses. Conflict of interest None declared. Data availability statement The probe sequences of the “RTA_v2” probe set, assembled genomes and the data matrices and trees from this study are deposited in Dryad Digital Repository (https://doi.org/10.5061/dryad.xksn02vkj). The sequenced reads are deposited in the NCBI Sequence Read Archive as project PRJNA907200. References Aberer, A.J., Kobert, K. and Stamatakis, A., 2014. Exa Bayes: Massively parallel bayesian tree inference for the whole-genome era. Mol. Biol. Evol. 31, 2553–2556. Agnarsson, I., Coddington, J.A. and Kuntner, M., 2013. Systematics, progress in the study of spider diversity and evolution. In: Penney, D. (Ed.), Spider Research in the 21st Century. Siri Scientific Press, Manchester, pp. 58–111. Azevedo, G.H.F., Bougie, T., Carboni, M., Hedin, M. and Ramırez, M.J., 2022. Combining genomic, phenotypic and Sanger sequencing data to elucidate the phylogeny of the two-clawed spiders (Dionycha). Mol. Phylogenet. Evol. 166, 107327. Ballesteros, J.A., Setton, E.V.W., L opez, C.E.S., Arango, C.P., Brenneis, G., Brix, S., Corbett, K.F., Cano-S anchez, E., Dandouch, M., Dilly, G.F., Eleaume, M.P., Gainett, G., Gallut, C., McAtee, S., McIntyre, L., Moran, A.L., Moran, R., L opezGonzalez, P.J., Scholtz, G., Williamson, C., Woods, H.A., Zehms, J.T., Wheeler, W.C. and Sharma, P.P., 2020. Phylogenomic resolution of sea spider diversification through integration of multiple data classes. Mol. Biol. Evol. 38, 686–701. Blaimer, B.B., Brady, S.G., Schultz, T.R., Lloyd, M.W., Fisher, B.L. and Ward, P.S., 2015. Phylogenomic methods outperform traditional multi-locus approaches in resolving deep evolutionary history: A case study of formicine ants. BMC Evol. Biol. 15, 271–285. Blair, C., Bryson, R.W., Jr., Linkem, C.W., Lazcano, D., Klicka, J. and McCormack, J.E., 2019. Cryptic diversity in the Mexican highlands: Thousands of UCE loci help illuminate phylogenetic relationships, species limits and divergence times of montane rattlesnakes (Viperidae: Crotalus). Mol. Ecol. Resour. 19, 349– 365. Bond, J.E. and Opell, D., 1998. Testing adaptive radiation and key innovation hypotheses in spiders. Evolution 52, 403–144. Borowiec, M.L., 2019. Spruceup: Fast and flexible identification, visualization, and removal of outliers from large multiple sequence alignments. J. Open Source Softw. 4, 1635. Bossert, S., Murray, E.A., Almeida, E.A.B., Brady, S.G., Blaimer, B.B. and Danforth, B.N., 2019. Combining transcriptomes and 10960031, 0, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/cla.12523 by University Of California, Riverside, Wiley Online Library on [31/01/2023]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License 10 ultraconserved elements to illuminate the phylogeny of Apidae. Mol. Phylogenet. Evol. 130, 121–131. Branstetter, M.G. and Longino, J.T., 2019. Ultra-conserved element phylogenomics of new world Ponera (Hymenoptera: Formicidae) illuminates the origin and phylogeographic history of the endemic exotic ant Ponera exotica. Insect Syst. Divers. 3, 1–13. Branstetter, M.G., Longino, J.T., Ward, P.S. and Faircloth, B.C., 2017. Enriching the ant tree of life: Enhanced UCE bait set for genome-scale phylogenetics of ants and other Hymenoptera. Methods Ecol. Evol. 8, 768–776. Bushnell, B., 2014. BBtools. Retrieved from https://sourceforge.net/ projects/bbmap/ (accessed on September 20, 2021). Chen, D., Brau, E.L., Forthman, M., Kimball, R.T. and Zhang, Z., 2018. A simple strategy for recovering ultraconserved elements, exons, and introns from low coverage shotgun sequencing of museum specimens: Placement of the partridge genus Tropicoperdix within the Galliformes. Mol. Phylogenet. Evol. 129, 304–314. Chikhi, R. and Rizk, G., 2013. Space-efficient and exact de Bruijn graph representation based on a Bloom filter. Algorithms Mol. Biol. 8, 1–9. Cruaud, A., Nidelet, S., Arnal, P., Weber, A., Fusu, L., Gumovsky, A. and Rasplus, J.-Y., 2018. Optimised DNA extraction and library preparation for minute arthropods: Application to target enrichment in chalcid wasps used for biocontrol. Mol. Ecol. Resour. 19, 702–710. Derkarabetian, S., Starrett, J., Tsurusaki, N., Ubick, D., Castillo, S. and Hedin, M., 2018. A stable phylogenomic classification of Travunioidea (Arachnida, Opiliones, Laniatores) based on sequence capture of ultraconserved elements. Zookeys 2018, 1– 36. Derkarabetian, S., Benavides, L.R. and Giribet, G., 2019. Sequence capture phylogenomics of historical ethanol-preserved museum specimens: Unlocking the rest of the vault. Mol. Ecol. Resour. 19, 1531–1544. Dimitrov, D. and Hormiga, H., 2020. Spider diversification through space and time. Annu. Rev. Entomol. 7, 225–241. Faircloth, B.C., 2016. PHYLUCE is a software package for the analysis of conserved genomic loci. Bioinformatics 32, 786–788. Faircloth, B.C., 2017. Identifying conserved genomic elements and designing universal bait sets to enrich them. Methods Ecol. Evol. 8, 1103–1112. Faircloth, B.C., McCormack, J.E., Crawford, N.G., Harvey, M.G., Brumfield, R.T. and Glenn, T.C., 2012. Ultraconserved elements anchor thousands of genetic markers spanning multiple evolutionary timescales. Syst. Biol. 61, 717–726. Faircloth, B.C., Branstetter, M.G., White, N.D. and Brady, S.G., 2015. Target enrichment of ultraconserved elements from arthropods provides a genomic perspective on relationships among Hymenoptera. Mol. Ecol. Resour. 15, 489–501. Fern andez, R., Kallal, R.J., Dimitrov, D., Ballesteros, J.A., Arnedo, M.A., Giribet, G. and Hormiga, G., 2018. Phylogenomics, diversification dynamics, and comparative transcriptomics across the spider tree of life. Curr. Biol. 28, 1489–1497.e5. Foelix, R.F., 1996. Biology of Spiders, 2nd edition. Oxford University Press, New York, p. 330. Forthman, M., Miller, C.W. and Kimball, R.T., 2019. Phylogenomic analysis suggests Coreidae and Alydidae (Hemiptera: Heteroptera) are not monophyletic. Zool. Scr. 48, 520–534. Garrison, N.L., Rodriguez, J., Agnarsson, I., Coddington, J.A., Griswold, C.E., Hamilton, C.A., Hedin, M., Kocot, K.M., Ledford, J.M. and Bond, J.E., 2016. Spider phylogenomics: Untangling the spider tree of life. PeerJ 4, e1719. Girard, M.B., Elias, D.O., Azevedo, G., Bi, K., Kasumovic, M.M., Waldock, J.M. and Hedin, M., 2021. Phylogenomics of peacock spiders and their kin (Salticidae: Maratus), with implications for the evolution of male courtship displays. Biol. J. Linn. Soc. 132, 471–494. Goloboff, P.A., Farris, J.S. and Nixon, K.C., 2008. TNT, a free program for phylogenetic analysis. Cladistics 24, 774–786. 11 Gregory, T.R. and Shorthouse, D.P., 2003. Genome sizes of spiders. J. Hered. 94, 285–290. Guillory, W.X., Muell, M.R., Summer, K. and Brown, J.L., 2019. Phylogenomic reconstruction of the Neotropical poison frogs (Dendrobatidae) and their conservation. Diversity 11, 126–140. Hamilton, C.A., Lemmon, A.R., Lemmon, E.M. and Bond, J.E., 2016. Expanding anchored enrichment to resolve both deep and shallow relationships within the spider tree of life. BMC Evol. Biol. 16, 212. Hedin, M., Derkarabetian, S., Alfaro, A., Ramırez, M.J. and Bond, J.E., 2019. Phylogenomic analysis and revised classification of atypoid mygalomorph spiders (Araneae, Mygalomorphae), with notes on arachnid ultraconserved element loci. PeerJ 7, e6864. Huang, W., Li, L., Myers, J.R. and Marth, J.T., 2012. ART: A next-generation sequencing read simulator. Bioinformatics 28, 593–594. Jesovnik, A., Sosa-Calvo, J., Lloyd, M.W., Branstetter, M.G., Fern Andez, F. and Schultz, T.R., 2017. Phylogenomic species delimitation and host-symbiont coevolution in the fungusfarming ant genus Sericomyrmex Mayr (Hymenoptera: Formicidae): Ultraconserved elements (UCEs) resolve a recent radiation. Syst. Entomol. 42, 523–542. Junier, T. and Zdobnov, E.M., 2010. The Newick utilities: Highthroughput phylogenetic tree processing in the UNIX shell. Bioinformatics 26, 1669–1670. Kalyaanamoorthy, S., Minh, B.Q., Wong, T.K.F., von Haeseler, A. and Jermiin, L.S., 2017. ModelFinder: Fast model selection for accurate phylogenetic estimates. Nat. Methods 14, 587–589. Katoh, K. and Standley, D.M., 2013. MAFFT multiple sequence alignment software version 7: Improvements in performance and usability. Mol. Biol. Evol. 30, 772–780. Kearse, M., Moir, R., Wilson, A., Stones-Havas, S., Cheung, M., Sturrock, S., Buxton, S., Cooper, A., Markowitz, S., Duran, C., Thierer, T., Ashton, B., Meintjes, P. and Drummond, A., 2012. Geneious basic: An integrated and extendable desktop software platform for the organization and analysis of sequence data. Bioinformatics 28, 1647–1649. K€ uck, P. and Meusemann, K., 2010. FASconCAT: Convenient handling of datamatrices. Mol. Phylogenet. Evol. 56, 1115–1118. Kulkarni, S., Wood, H., Lloyd, M. and Hormiga, G., 2020a. Spiderspecific probe set for ultraconserved elements offers new perspectives on the evolutionary history of spiders (Arachnida, Araneae). Mol. Ecol. Resour. 20, 185–203. Kulkarni, S., Kallal, R.J., Wood, H., Dimitrov, D., Giribet, G. and Hormiga, G., 2020b. Interrogating genomic-scale data to resolve recalcitrant nodes in the spider tree of life. Mol. Biol. Evol. 38, 891–903. Leduc-Robert, G. and Maddison, W.P., 2018. Phylogeny with introgression in Habronattus jumping spiders (Araneae: Salticidae). BMC Evol. Biol. 18, 24–47. Li, H., Handsaker, B., Wysoker, A., Fennell, T., Ruan, J., Homer, N., Marth, G., Abecasis, G., Durbin, R. and 1000 Genome Project Data Processing Subgroup, 2009. The sequence alignment/map (SAM) format and SAMtools. Bioinformatics 25, 2078–2079. Lunter, G. and Goodson, M., 2011. Stampy: A statistical algorithm for sensitive and fast mapping of Illumina sequence reads. Genome Res. 21, 936–939. Luo, R., Liu, B., Xie, Y., Li, Z., Huang, W., Yuan, J. and Wang, J., 2012. SOAPdenovo2: An empirically improved memory-efficient short-read de novo assembler. Gigascience 1, 1–6. Maddison, W.P., 2015. A phylogenetic classification of jumping spiders (Araneae: Salticidae). J. Arachnol. 43, 231–292. Maddison, W.P., Evans, S.C., Hamilton, C.A., Bond, J.E., Lemmon, A.R. and Lemmon, E.M., 2017. A genome-wide phylogeny of jumping spiders (Araneae, Salticidae) using anchored hybrid enrichment. Zookeys 695, 89–101. Maddison, W.P., Beattie, I., Marathe, K., Ng, P.Y.C., Kanesharatnam, N., Benjamin, S.P. and Kunte, K., 2020a. A phylogenetic and taxonomic review of baviine jumping spiders (Araneae, Salticidae, Baviini). Zookeys 1004, 27–97. 10960031, 0, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/cla.12523 by University Of California, Riverside, Wiley Online Library on [31/01/2023]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License Zhang J. et al. / Cladistics 0 (2023) 1–13 Zhang J. et al. / Cladistics 0 (2023) 1–13 Maddison, W.P., Maddison, D.R., Derkarabetia, S. and Hedin, M., 2020b. Sitticine jumping spiders: Phylogeny, classification, and chromosomes (Araneae, Salticidae, Sitticini). Zookeys 925, 1–54. Magalhaes, I.L.F., Azevedo, G.H.F., Michalik, P. and Ramırez, M.J., 2020. The fossil record of spiders revisited: Implications for calibrating trees and evidence for a major faunal turnover since the Mesozoic. Biol. Rev. 95, 184–217. McCormack, J.E., Faircloth, B.C., Crawford, N.G., Gowaty, P.A., Brumfield, R.T. and Glenn, T.C., 2012. Ultraconserved elements are novel phylogenomic markers that resolve placental mammal phylogeny when combined with species-tree analysis. Genome Res. 22, 746–754. McCormack, J.E., Harvey, M.J. and Faircloth, B.C., 2013. A phylogeny of birds based on over 1,500 loci collected by target enrichment and high-throughput sequencing. PLoS One 8, 51–67. Miller, J.A., Carmichael, A., Ramırez, M.J., Spagna, J.C., Haddad, C.R., Rezac, M., Johannesen, J., Kral, J., Wang, X.P. and Griswold, C.E., 2010. Phylogeny of entelegyne spiders: Affinities of the family Penestomidae (NEW RANK), generic phylogeny of Eresidae, and asymmetric rates of change in spinning organ evolution (Araneae, Araneoidea, Entelegynae). Mol. Phylogenet. Evol. 55, 786–804. https://doi.org/10.1016/j.ympev.2010.02.021. Minh, B.Q., Schmidt, H.A., Chernomor, O., Schrempf, D., Woodhams, M.D., von Haeseler, A. and Lanfear, R., 2020. IQTREE 2: New models and efficient methods for phylogenetic inference in the genomic era. Mol. Biol. Evol. 37, 1530–1534. Mirarab, S., Nguyen, N. and Warnow, T., 2014. PASTA: Ultralarge multiple sequence alignment. Res. Comput. Mol. Biol. 22, 177–191. Moradmand, M., Sch€ onhofer, A.L. and J€ager, P., 2014. Molecular phylogeny of the spider family Sparassidae with focus on the genus Eusparassus and notes on the RTA-clade and ‘Laterigradae’. Mol. Phylogenet. Evol. 74, 48–65. Ochoa, L.E., Datovo, A., DoNascimiento, C., Roxo, F.F., Sabaj, M.H., Chang, J., Melo, B.F., Silva, G.S.C., Foresti, F., Alfaro, M. and Oliveira, C., 2020. Phylogenomic analysis of trichomycterid catfishes (Teleostei: Siluriformes) inferred from ultraconserved elements. Sci. Rep. 10, 2697. Opatova, V., Hamilton, C.A., Hedin, M., De Oca, L.M., Kral, J. and Bond, J.E., 2019. Phylogenetic systematics and evolution of the spider infraorder Mygalomorphae using genomic scale data. Syst. Biol. 69, 671–707. Pie, M.R., Bornschein, M.R., Ribeiro, L.F., Faircloth, B.C. and McCormac, J.E., 2019. Phylogenomic species delimitation in microendemic frogs of the Brazilian Atlantic Forest. Mol. Phylogenet. Evol. 141, 106627. Prjibelski, A., Antipov, D., Meleshko, D., Lapidus, A. and Korobeynikov, A., 2020. Using SPAdes de novo assembler. Curr. Protoc. Bioinformatics 70, e102. Pryszcz, L.P. and Gabald on, T., 2016. Redundans: An assembly pipeline for highly heterozygous genomes. Nucleic Acids Res. 44, e113–e123. Quinlan, A.R. and Hall, I.M., 2010. BEDTools: A flexible suite of utilities for comparing genomic features. Bioinformatics 26, 841– 842. Ramırez, M.J., Magalhaes, I.L.F., Derkarabetian, S., Ledford, J., Griswold, C.E., Wood, H.M. and Hedin, M., 2020. Sequence capture phylogenomics of true spiders reveals convergent evolution of respiratory systems. Syst. Biol. 70, 14–20. Richman, D.B. and Jackson, R.R., 1992. A review of the ethology of jumping spiders (Araneae, Salticidae). Bull. Br. Arachnol. Soc. 9, 33–37. Sahlin, K., Vezzi, F., Nystedt, B., Lundeberg, J. and Arvestad, L., 2014. BESST-efficient scaffolding of large fragmented assemblies. BMC Bioinformatics 15, 281–292. Shao, L. and Li, S., 2018. Early Cretaceous greenhouse pumped higher taxa diversification in spiders. Mol. Phylogenet. Evol. 127, 146–155. Shen, W., Le, S. and Li, Y., 2016. SeqKit: A cross-platform and ultrafast toolkit for fasta/q file manipulation. PLoS One 11, e0163962. Spagna, J.C. and Gillespie, R.G., 2008. More data, fewer shifts: Molecular insights into the evolution of the spinning apparatus in non-orb-weaving spiders. Mol. Phylogenet. Evol. 46, 347–368. Stamatakis, A., 2014. RAxML version 8: A tool for phylogenetic analysis and post-analysis of large phylogenies. Bioinformatics 30, 1312–1313. Starrett, J., Derkarabetian, S., Hedin, M., Bryson, R.W., McCormack, J.E. and Faircloth, B.C., 2017. High phylogenetic utility of an ultraconserved element probe set designed for Arachnida. Mol. Ecol. Resour. 17, 812–823. Sun, X., Ding, Y., Orr, M.C. and Zhang, F., 2020. Streamlining universal single-copy orthologue and ultraconserved element design: A case study in Collembola. Mol. Ecol. Resour. 20, 706– 717. Torres, A., Goloboff, P.A. and Catalano, S.A., 2021. Parsimony analysis of phylogenomic datasets (I): Scripts and guidelines for using TNT (Tree Analysis using New Technology). Cladistics 38, 103–125. Van Dam, M.H., Lam, A.W., Sagata, K., Gewa, B., Laufa, R., Balke, M., Faircloth, B.C. and Riedel, A., 2017. Ultraconserved elements (UCEs) resolve the phylogeny of australasian smurfweevils. PLoS One 12, 1–21. Van Dam, M.H., Trautwein, M., Spicer, G.S. and Esposito, L., 2018. Advancing mite phylogenomics: Designing ultraconserved elements for Acari phylogeny. Mol. Ecol. Resour. 19, 465–475. Wheeler, W.C., Coddington, J.A., Crowley, L.M., Dimitrov, D., Goloboff, P.A., Griswold, C.E., Hormiga, G., Prendini, L., Ramırez, M.J., Sierwald, P., Almeida-Silva, L., Alvarez-Padilla, F., Arnedo, M.A., Silva, L.R.B., Benjamin, S.P., Bond, J.E., Grismado, C.J., Hasan, E., Hedin, M., Izquierdo, M.A., Labarque, F.M., Ledford, J., Lopardo, L., Maddison, W.P., Miller, J.A., Piacentini, L.N., Platnick, N.I., Polotow, D., SilvaDavila, D., Scharff, N., Sz} uts, T., Ubick, D., Vink, C.J., Wood, H.M. and Zhang, J., 2017. The spider tree of life: Phylogeny of Araneae based on target-gene analyses from an extensive taxon sampling. Cladistics 33, 574–616. Wood, H.M., Gonzalez, V.L., Lloyd, M., Coddington, J. and Scharff, N., 2018. Next-generation museum genomics: Phylogenetic relationships among palpimanoid spiders using sequence capture techniques (Araneae: Palpimanoidea). Mol. Phylogenet. Evol. 127, 907–918. World Spider Catalog, 2022. World Spider Catalog. Version 23.5. Natural History Museum Bern. Retrieved from http://wsc.nmbe. ch (accessed on October 20, 2022). Xu, X., Su, Y.-C., Ho, S.Y.W., Kuntner, M., Ono, H., Liu, F., Chang, C.-C., Warrit, N., Sivayyapram, V., Aung, K.P.P., Pham, D.S., Norma-Rashid, Y. and Li, D., 2021. Phylogenomic analysis of ultraconserved elements resolves the evolutionary and biogeographic history of segmented trapdoor spiders. Syst. Biol. 70, 1110–1122. Yu, N., Li, J., Liu, M., Huang, L.X., Bao, H.B., Yang, Z.M., Zhang, Y., Gao, H., Wang, Z., Yang, Y., Van Leeuwen, T., Millar, N.S. and Liu, Z.W., 2019. Genome sequencing and neurotoxin diversity of a wandering spider Pardosa pseudoannulata (pond wolf spider). bioRxiv, 747147. https://doi. org/10.1101/747147. Zhang, J. and Lai, J., 2020. Phylogenomic approaches in systematic studies. Zool. Syst. 45, 151–162. Zhang, J. and Maddison, W.P., 2013. Molecular phylogeny, divergence times and biogeography of spiders of the subfamily Euophryinae (Araneae: Salticidae). Mol. Phylogenet. Evol. 68, 81–92. Zhang, J. and Maddison, W.P., 2015. Genera of euophryine jumping spiders (Araneae: Salticidae), with a combined molecularmorphological phylogeny. Zootaxa 3938, 1–147. Zhang, C., Rabiee, M., Sayyari, E. and Mirarab, S., 2018. ASTRAL- III: Polynomial time species tree reconstruction from partially resolved gene trees. BMC Bioinformatics 19, 15–30. Zhang, F., Ding, Y., Zhu, C.-D., Zhou, X., Orr, M.C., Scheu, S. and Luan, Y.-X., 2019. Phylogenomics from low-coverage wholegenome sequencing. Methods Ecol. Evol. 10, 507–517. 10960031, 0, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/cla.12523 by University Of California, Riverside, Wiley Online Library on [31/01/2023]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License 12 Zhang, J., Lindsey, A.R.I., Peters, R.S., Heraty, J.M., Hopper, K.R., Werren, J.H., Martinson, E.O., Woolley, J.B., Yoder, M.J. and Krogmann, L., 2020. Conflicting signal in transcriptomic markers leads to a poorly resolved backbone phylogeny of chalcidoid wasps. Syst. Entomol. 45, 783–802. Supporting Information Additional supporting information may be found online in the Supporting Information section at the end of the article. Fig. S1. Flowchart of probe design, in-silico test and optimization procedure. Fig. S2. The best trees from ML analyses on the full loci concatenated datasets from RTA_v1, Spider and Arachnida probes. The numbers along the branches are shown as: ML bootstrap/posterior probability. The scale bar is in substitutions per position. Fig. S3. The best trees from ML analyses on the concatenated coding loci datasets from RTA_v1, Spider and Arachnida probes. The numbers along the branches are ML bootstrap. The scale bar is in substitutions per position. Fig. S4. The best trees from ML analyses on the concatenated non-coding loci datasets from RTA_v1, Spider and Arachnida probes (the red branches indicate the different topologies from the concatenated full loci ML analyses). The numbers along the branches are ML bootstrap. The scale bar is in substitutions per position. Fig. S5. The ASTRAL species trees from the full loci datasets from RTA_v1, Spider and Arachnida probes (the red branches indicate the different 13 topologies from the concatenated analyses). The numbers along the branches are bootstrap values. Fig. S6. Strict consensus of all equally parsimonious trees from TNT analyses on datasets from RTA_v1, Spider and Arachnida probes (the red branches indicate different topologies from the ML analyses). The numbers along the branches are bootstrap support values (only show >75%). Fig. S7. Maximum parsimonious tree from TNT analysis on the dataset from RTA_v2 probes and 19 genomes (the red branches indicate different topologies from the ML analyses). The numbers along the branches are bootstrap support values (only show >75%). Fig. S8. Maximum parsimonious tree from TNT analysis on the dataset from RTA_v2 probes and 57 species (the red branches indicate different topologies from the ML analyses). The numbers along the branches are bootstrap support values (only show >75%). Table S1. Specimen information and summary of genome assembly and harvested UCE loci number from different probe sets. *Denotes the species enriched empirically in the lab. Table S2. Number of UCE loci shared between taxa during probe design. *Indicates that the set was chosen for this probe design, 3900 loci shared by eight taxa. Table S3. Summary statistics for the UCE datasets in the concatenation analyses. 10960031, 0, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/cla.12523 by University Of California, Riverside, Wiley Online Library on [31/01/2023]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License Zhang J. et al. / Cladistics 0 (2023) 1–13