Pig SINE-RIP probe combination, gene chip, kit and application
By developing a combination of pig SINE-RIP probes and gene chips, the problem of the lack of SINE-RIP detection tools in existing technologies has been solved, enabling more efficient and accurate pig breeding and genome analysis, and improving the effectiveness of molecular marker-assisted selection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-13
- Publication Date
- 2026-03-13
AI Technical Summary
Current technologies lack specific detection probes and gene chips for SINE-RIP structural variations, which limits the accuracy and efficiency of marker-assisted selection in pig breeding.
A porcine SINE-RIP probe combo containing 51,000 SINE-RIP molecular marker sites was developed. 51,000 probes were designed, and liquid-phase chips and kits were prepared based on sequences within 150 base pairs in the upstream and downstream flanking regions of the SINE-RIP sites for porcine whole-genome analysis.
It improves the accuracy and efficiency of pig breeding, provides higher accuracy of typing results and detection throughput, and is suitable for various applications such as pig genome selection breeding, breed tracing, kinship analysis, and genetic diversity analysis. It is low in cost and has a broad market prospect.
Smart Images

Figure CN121653259A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to animal molecular breeding technology, and more particularly to a porcine SINE-RIP probe assembly, gene chip, kit, and application. Background Technology
[0002] In modern pig breeding systems, molecular biology techniques have provided revolutionary tools. Among them, marker-assisted breeding (MABB) significantly improves the accuracy and efficiency of selection by detecting genes or markers associated with important economic traits, enabling early and indirect selection. Breakthroughs internationally in key traits such as litter size have largely benefited from the widespread application of these technologies. Currently, related research is using an increasing number and types of molecular markers, driving the evolution of molecular breeding from an auxiliary method to a routine technique.
[0003] Among numerous molecular markers, single nucleotide polymorphisms (SNPs) have been widely used, but they are mostly neutral variations, and their causal relationship with traits is often unclear, limiting their application. In contrast, short, scattered nuclear element insertion polymorphisms (SINE-RIPs), as a structural variation marker, exhibit more significant advantages. SINE transposons account for approximately 11% of the pig genome, with their length mainly concentrated between 200-300 bases. They often carry functional elements such as promoters and enhancers, directly affecting gene expression and thus potentially representing "causal mutations" leading to phenotypic differences. Furthermore, SINE-RIP markers possess advantages such as high polymorphism, co-dominant inheritance, and stable detection, making them more valuable for genomic selection breeding and quantitative trait locus mapping. However, currently, there is a lack of structural variation detection probes and gene chips specifically targeting SINE-RIPs. Summary of the Invention
[0004] Objectives of the invention: The first objective is to provide a SINE-RIP probe combo that can be used for whole-genome analysis of pigs; the second objective is to provide gene chips and kits containing this pig SINE-RIP probe combo; and the third objective is to provide applications of the above products.
[0005] Technical solution: The porcine SINE-RIP probe combination of the present invention detects porcine short scattered nuclear element insertion polymorphism (SINE-RIP) molecular marker combination. The molecular marker combination includes 51,000 SINE-RIP molecular marker sites. The physical location of the deoxynucleotide preceding the site in the reference genome and the length of the marker in the reference genome are shown in Table 1. The reference genome is Sscrofa11.1.
[0006] Table 1 SINE-RIP molecular marker site information
[0007] In Table 1, the site information format is as follows: chromosome or sequence number: genomic location of the base preceding the SINE-RIP molecular marker site: length of the SINE-RIP molecular marker site; a length of 1 indicates that the site is a SINE deletion site in the reference genome, and a length of more than 100 indicates that the site is a SINE insertion site in the reference genome.
[0008] Preferably, the probes in the probe assemblies use sequences within 150 bases upstream and downstream of the SINE-RIP molecular marker site as design reference sequences, with a length of 100-120 bases; more preferably, the probes in the probe assemblies have a GC content between 20% and 80% and a Tm value between 60℃ and 80℃.
[0009] The application of the porcine SINE-RIP probe combination described in this invention in the preparation of gene chips.
[0010] The gene chip of this invention is loaded with the porcine SINE-RIP probe combination as described above.
[0011] Preferably, the chip is a liquid-phase chip or a solid-phase chip; more preferably, the chip is a liquid-phase chip.
[0012] The kit described in this invention comprises the aforementioned porcine SINE-RIP probe combination or gene chip.
[0013] The application of the porcine SINE-RIP probe combination, gene chip, or kit described in this invention in pig breeding.
[0014] The application of the porcine SINE-RIP probe combination, gene chip, or kit described in this invention in porcine breed tracing or kinship analysis.
[0015] The application of the porcine SINE-RIP probe combination, gene chip, or kit described in this invention in the identification and evaluation of porcine germplasm resources.
[0016] The application of the porcine SINE-RIP probe combination, gene chip, or kit described in this invention in the localization of porcine functional genes or loci.
[0017] The application of the porcine SINE-RIP probe combination, gene chip, or kit described in this invention in porcine genome-wide association analysis.
[0018] The application of the porcine SINE-RIP probe combination, gene chip, or kit described in this invention in porcine genetic diversity analysis.
[0019] Beneficial Effects: Compared with existing technologies, this invention has the following significant advantages: 1. The porcine SINE-RIP probe combination detects a class of structural variation molecular markers that are large in size, easily cause changes in gene activity, and have a large genetic effect. It is suitable for various applications and research, such as porcine genome selection breeding, porcine germplasm resource identification and evaluation, kinship analysis, genetic diversity analysis and evaluation, functional gene or locus localization, and genome-wide association analysis; 2. The porcine SINE-RIP probe combination is based on the reference genome Sscrofa11.1, has wide applicability to breeds and populations, and its genotyping results are accurate, the detection process is simple, high-throughput, and low-cost, with broad market prospects; 3. The porcine SINE-RIP probe combination and gene chip are an effective supplement or even replacement for the current SNP molecular marker system, which can further improve the accuracy of genome selection and provide a new, highly reliable tool for porcine genome-wide related applications and research, with extremely high industrial application value. Attached Figure Description
[0020] Figure 1 Flowchart for the development of liquid-phase microarrays for SINE-RIP site screening and porcine 51K SINE-RIP structural variation;
[0021] Figure 2 Flowchart for data analysis of 51K SINE-RIP structural variation detection liquid phase chip for pigs;
[0022] Figure 3 Density map of sites on a liquid-phase chip for structural variations in porcine 51K SINE-RIP;
[0023] Figure 4 The image shows the gel electrophoresis results of representative genotypes in the PCR validation results. The large bands are the amplification products of SINE insertion alleles, and the small bands are the amplification products of SINE deletion alleles.
[0024] Figure 5 The results of applying the 51K SINE-RIP structural variation liquid-phase chip for pig population phylogenetic tree analysis are shown in the figure.
[0025] Figure 6 The image shows the results of applying the 51K SINE-RIP structural variation liquid phase chip for pig population structure analysis. Detailed Implementation
[0026] The technical solution of the present invention will be further described below.
[0027] Example 1: Establishment of porcine SINE-RIP probe assemblies and preparation of gene chips
[0028] 1. Establishment of probe array
[0029] (1) SINE-RIP site identification
[0030] RepeatMasker (v 4.1.5, -e rmblast -pa 200 -s -cutoff 250 -no_is -nolow) software was used based on the pig transposon database (C. Chen, W. Wang, X. Wang, D. Shen, S. Wang, Y. Wang, B. Gao, K. Wimmers, J. Mao, K. Li, Retrotransposons evolution and impact on lncRNA and protein coding genes in pigs). Mobile DNA The SINE insertion sites in the pig reference genome Sscrofa11.1 (https: / / hgdownload.soe.ucsc.edu / goldenPath / susScr11 / bigZips / susScr11.fa.gz) were marked to obtain all SINE insertion sites and their location information in the reference genome.
[0031] First, SINE-RIP site identification was performed based on the assembled genome data: The pig assembled genome was downloaded from the NCBI public database (https: / / www.ncbi.nlm.nih.gov / datasets / genome / ?taxon=9823), mainly in three steps: ① Comprehensive transposon insertion site labeling was performed on the assembled genome using the RepeatMasker program. SINE insertion sites with a label length exceeding 100 bases were screened; ② For each site in ①, a 200-base flanking region sequence was extracted from its upstream or downstream regions and aligned to the reference genome using the alignment software Blat (v 35) to obtain specific location information; ③ By comparing the location information of each site with each other, sites with differential distribution across genomes were obtained, i.e., SINE-RIP sites.
[0032] Next, second-generation whole-genome sequencing data of domestic pigs were collected from the NCBI Sequencing Data Database (SRA, http: / / www.ncbi.nlm.nih.gov / sra / ). Whole-genome SINE-RIP site identification was performed using the TypeSINE program (https: / / github.com / NaisuYang / TypeSINE).
[0033] It mainly includes two identification modules:
[0034] Module 1: First, obtain the SINE itself (30 bases upstream) and its upstream flanking sequences from all SINE retrotransposon insertion sites in the reference genome. Remove sites where the upstream 30-base flanking sequence is repeated in the reference genome. The remaining 60-base sequences form a feature sequence set. Align the feature sequence set with the raw sequencing data. If a read matches the 60-base feature sequence, then SINE is identified. + Then, the upstream 30-base flanking sequences and the downstream 30-base flanking sequences of SINE transposon insertion sites in all SINE transposon insertion sites in the reference genome are obtained to obtain a combined sequence set. Further checks are performed on whether any reads in the sequencing data can simultaneously align with the upstream 30-base flanking sequences and the downstream 30-base flanking sequences of the combined sequence. If so, it is determined that a SINE transposon insertion site exists. - Finally, SINE was selected from the data sample set used in the analysis. + and SINE - The site is the SINE-RIP site.
[0035] Module 2: Extract the non-redundant first 30 bases of the SINE retrotransposon from the porcine transposon database as a seed sequence. Identify fragments containing the seed sequence in the raw sequencing data. Further extract the upstream 30 bases of the seed sequence from the sequencing data and align it with the genome. Then, screen for sites that can uniquely align to the reference genome. Check if the downstream site in the reference genome is a SINE retrotransposon. If not, mark it as a new site and determine if a SINE retrotransposon exists. + Further, the upstream 30 bases of the seed sequence from the aforementioned new site were extracted to find the corresponding sequence in the reference genome, along with the downstream 30 bases. The sequencing data was then analyzed to identify whether these 60-base reads were present. If they were present, the presence of SINE was determined. - Similarly, the final selection process involved identifying SINE samples within the analyzed data sample set. + and SINE - The site is the SINE-RIP site.
[0036] (2) SINE-RIP verification and screening
[0037] By integrating the SINE-RIP sites obtained from the genome assembly and second-generation resequencing genome data analysis in (1) above, the upstream and downstream sequences were analyzed, and 51,000 SINE-RIP sites with unique flanking sequences and allele frequencies between 5% and 95% and even chromosome distribution were selected from the sample set, as shown in Table 1.
[0038] 2. Preparation of molecular probe combinations and gene chips
[0039] For the 51,000 target SINE-RIP sites obtained above, probes for capture were designed. In order to ensure sufficient specificity and stable hybridization kinetics, we selected sequences within 150 bases of the upstream and downstream flanking regions of the SINE-RIP sites as reference sequences for probe design. Probes with a length of 100-120 bases, a GC content between 20% and 80%, and a Tm value between 60℃ and 80℃ were designed, and finally the "Pig Core No. 1"—51K SINE-RIP structural variation liquid phase chip was obtained.
[0040] Furthermore, corresponding kits can be prepared, which integrate capture probe combinations or 51K SINE-RIP structural variation liquid phase chips, buffers, primers, and other necessary reagents.
[0041] The development process of SINE-RIP site screening and 51K SINE-RIP structural variation liquid-phase chip is as follows: Figure 1 As shown.
[0042] Example 2: Performance Verification of Porcine 51K SINE-RIP Structure-Modified Liquid Chip
[0043] 1. Preparation of hybridization product libraries
[0044] (1) The basic information of the 24 selected individual samples is shown in Table 2.
[0045] Table 2. Basic Individual Information.
[0046]
[0047] (2) Based on the flanking region sequence information of the 51,000 SINE-RIP sites shown in Table 1, corresponding probe sequences were designed. The probe length is 120 bases. The specificity of each designed probe sequence was evaluated to avoid non-specific hybridization.
[0048] (3) Based on the base sequence of the probe, chemical synthesis was carried out and biotin labeling was added to obtain the probe library.
[0049] (4) Ear tissue samples were extracted from the above 24 individuals using a genome extraction kit to obtain porcine genomic DNA. The genomic DNA was then broken down to about 250 bases by enzyme digestion using the CWseq Universal DirectFast DNA Library Prep Kit. After end repair, A-tailing, adapter ligation and PCR enrichment, a DNA library with adapter sequences was obtained.
[0050] (5) The probe with biotin label is specifically hybridized and captured with the DNA library with adapter sequence. Finally, the non-specific hybridization is removed by elution buffer. The elution product is amplified, quantified and quality checked by post-PCR to obtain a qualified targeted capture library.
[0051] (6) Then, deep sequencing of PE150 was performed using a high-throughput sequencer to obtain raw data. Among them, samples BM2, BM3, and DB1 were used for experimental replication.
[0052] 2. Data Processing and Analysis
[0053] Data quality control: Offline data undergoes standard quality control using Fastp (v0.24.0) software, including: ① removing reads with adapters; ② removing read pairs where the nitrogen content exceeds 10% of the read's base count; ③ removing read pairs where the low-quality (Q<=5) base count exceeds 50% of the read's base count. The remaining reads are considered clean data.
[0054] Data preprocessing after sequencing: ① Use the paired-end sequencing data splicing software Flash (v1.2.11) to splice the clean data of each sample; ② Use the Seqkit (v2.1.0) software to convert the spliced data of each sample into FASTA format and generate read length information files; ③ Use makeblastdb in BLAST (v2.16.0+) to build the index files required for alignment.
[0055] Basic material preparation: ① Sequence set of the upstream 50 bases flanking regions of each site in the 51K SINE-RIP and its BLAST alignment index file; ② Sequence set of the SINE front 100 bases of all insertion sites in the 51K SINE-RIP and its BLAST alignment index file; ③ Sequence set of the downstream 110 bases flanking regions of the 51K SINE-RIP and its BLAST alignment index file.
[0056] Data Analysis: Chip testing and off-machine data analysis includes:
[0057] (1) Targeting insertional alleles (SINE) +Identification: ① The upstream flanking region sequence set of each locus in 51K SINE-RIP was BLAST-aligned with the results of the preprocessed offline data. Reads with unique alignment results were retained and converted to BED format; ② In the retained reads, BED tools flank was used to extend the sequence downstream by 30-70 bases and extract the extended sequence; ③ The obtained extended sequence was BLAST-aligned with the SINE front 100 base sequence set of all insertion sites in 51K SINE-RIP; ④ Based on the alignment results obtained in steps ① and ③, the sequences containing insertion alleles (SINE... + The SINE-RIP site.
[0058] (2) Targeting deletion alleles (SINE) - Identification: ① Exclude reads that have been used in SINE from those with a unique alignment result in step ① of the identification of inserted alleles. + ① Determine all reads; ② Extend the remaining reads downstream by 50-100 bases and extract the extended sequence; ③ Compare the obtained sequence with the flanking region of the 110 bases downstream of 51K SINE-RIP; ④ Retain the results that have a unique alignment and have an upstream flanking region (for the results obtained in step ① of the insertion allele identification step), and then screen for the presence of deletion alleles (SINE-RIP). - The SINE-RIP site.
[0059] (3) Based on the identification results of each site in (1) and (2) above, if it only appears in (1), then locate SINE. + / + If it only appears in (2), then locate SINE. - / - If it appears in both (1) and (2), then locate SINE. + / - If it does not appear in (1) and (2), then it is located as NA, that is, it failed to be accurately classified.
[0060] The above-mentioned data analysis process for detecting structural variations in pigs using a 51K SINE-RIP liquid-phase chip is as follows: Figure 2 As shown.
[0061] The analysis results showed that the data output of the tested samples was relatively consistent, averaging 16.57 M reads. The Q20 (Phred) percentage (bases with a Phred value greater than 20) accounted for 99.31% of the total bases, indicating high data quality. Further genotyping analysis showed that 49,016 sites could be successfully genotyped in samples with 12 or more genotypes, with an overall detection rate of 96.11% (49,016 / 51,000). The average number of detected sites across all samples was 48,571, with an average detection rate of 95.24% (48,571 / 51,000). Detailed results are shown in Table 3.
[0062] Table 3 Performance verification results of 51K SINE-RIP structure-variant liquid-phase chip
[0063]
[0064] Raw Reads is the number of reads contained in the offline data; Clean Reads is the number of reads after filtering.
[0065] Further statistical analysis was conducted on the repetition rate of the results between the experimental and duplicate groups of the three samples BM2, BM3, and DB1. The detailed results are shown in Table 4. The results show that the repetition rate of the detected sites in the two tests between the experimental and duplicate groups was approximately 99.43%, and the consistency rate of the typing results between the two groups was approximately 98.29%.
[0066] Table 4. Statistical table of the repetition rate of detected sites in the two tests between the replicate group and the experimental group.
[0067]
[0068] 3. PCR detection verification
[0069] 162 loci were randomly selected, and the specific locus names are shown in Table 5. Primers were designed using Primer-BLAST (https: / / www.ncbi.nlm.nih.gov / tools / primer-blast / index.cgi). PCR detection was performed using the genomes of 12 individuals (BMEI2, MZ3, DB5, PTL2, LW1, BM1, DB3, CB2, DLK1, MS2, BN3, and BM3) as templates, and the products were analyzed by gel electrophoresis.
[0070] Table 5. List of 162 randomly selected site names
[0071]
[0072] Results with clear bands of the expected size and no extraneous bands in the gel electrophoresis results were retained as experimental results that could determine the genotype. A total of 1062 valid experimental results were obtained, of which 980 detection results were consistent with the genotyping results of the chip data, with a consistency rate of 92.28%. The statistical information of the results is shown in Table 6.
[0073] Table 6. Statistical information comparing PCR detection results with genotyping analysis results from chip-based data.
[0074]
[0075] Representative examples of liquid phase chip data analysis results and PCR validation results are shown in Table 7. Figure 3 As shown.
[0076] Table 7. Examples of representative genotypes from liquid phase chip data analysis and PCR validation.
[0077]
[0078] Example 3: Practical Application of Porcine 51K SINE-RIP Structure-Variated Liquid Chip
[0079] A total of 120 samples were collected from 42 breeds or populations, including 39 local pig breeds in China, wild boars from Anhui and Northeast China, and Large White pigs. Sample information is shown in Table 8. Following the method in Example 2, genomic DNA was extracted and analyzed using a 51K SINE-RIP structural variation liquid chromatography-array detection procedure. Genotype data were obtained after analysis of the data. The obtained loci were further screened for quality control according to the following criteria: ① loci used on autosomes; ② SINE-RIP detection rate ≥ 60%, individual detection rate ≥ 90%; ③ MAF ≥ 0.05. The retained loci were used for subsequent phylogenetic analysis.
[0080] Table 6. List of Samples Used in the Evaluation of the Application of Porcine 51K SINE-RIP Structural Variation Liquid Chromatography Chip
[0081]
[0082] The phylogenetic tree was constructed using the Neighbor-Joining (NJ) method, and the results are as follows: Figure 4 As shown, the 120 samples can be roughly divided into 6 groups:
[0083] The Erhualian pig, Meishan pig, Mi pig, and Jiangquhai pig from Jiangsu Province clustered in one branch, but a Hanjiang black pig sample was also found in this branch, suggesting possible bloodline exchange. The Bama fragrant pig, Wuzhishan pig, and the Lantang pig, Dahua white pig, Minbei spotted pig, Guanzhuang spotted pig, and Huai pig from Guangdong and Fujian provinces clustered in another branch, consistent with their geographical location. Tibetan pigs and the Hanjiang black pig, Qingyu pig, Liangshan pig, and Xiangxi black pig from Southwest China clustered in another branch. The Qingping pig, Tongcheng pig, Daweizi pig, Shaziling pig, Ningxiang pig, and Leping spotted pig from the middle and lower reaches of the Yangtze River clustered in another branch, which also included wild boars from Anhui Province, suggesting that this region... Numerous breeds may share the same or similar evolutionary origins. The northern breeds of Ma Shen pig, Shenxian pig, Northeast Min pig, Laiwu black pig, Dingyuan pig, and Dapulian pig cluster in one branch, suggesting a distant evolutionary relationship between northern and southern pigs, but closer kinship than southern local pigs. This branch also includes Wei pigs, suggesting possible hybridization between the collected Wei pigs and northern local pigs. The Baixian pig, Kele pig, Youlong black pig, Jinhua pig, Putian black pig, Yantai black pig, Dabai pig, and Northeast wild boar cluster in one branch. This branch is more complex, suggesting possible bloodline exchange among some breeds, preventing them from clustering with surrounding local pigs.
[0084] In summary, individuals of each breed generally clustered together in groups of 2-3, and local pigs from different regions also clustered together, but were well separated from each other. This provides a relatively objective reflection of the kinship among local pigs and any potential bloodline exchange. Further analysis of the population genetic structure can be achieved through diagrams such as... Figure 5 As shown, the lowest cross-validation error occurs when K=4, and the resulting structural map basically corresponds to the constructed phylogenetic tree, which can clearly show the relationships and composition between individuals and varieties.
Claims
1. A porcine SINE-RIP probe assembly, characterized in that, The probe combination was used to detect a combination of short, scattered nuclear element insertion polymorphism (SNP) molecular markers, which included 51,000 SNP molecular marker sites. The physical location of the preceding deoxynucleotide in the reference genome and the length of the marker in the reference genome are shown in Table 1. The reference genome is Sscrofa11.
1.
2. The porcine SINE-RIP probe assembly according to claim 1, characterized in that, The probes in the probe assemblies are designed with a reference sequence of 100-120 bases in length, which is a short, scattered sequence within 150 bases in the flanking regions upstream and downstream of the nuclear element insertion polymorphic molecular marker site.
3. The porcine SINE-RIP probe assembly according to claim 2, characterized in that, The probe GC content in the probe combination is between 20% and 80%, and the Tm value is between 60℃ and 80℃.
4. A gene chip, characterized in that, The chip load is the porcine SINE-RIP probe assembly as described in any one of claims 1-3.
5. The gene chip according to claim 4, characterized in that, The chip is a liquid-phase chip or a solid-phase chip.
6. A reagent kit, characterized in that, The kit comprises the porcine SINE-RIP probe combination according to any one of claims 1-3, or the gene chip according to any one of claims 4-5.
7. The application of the porcine SINE-RIP probe combination according to any one of claims 1-3, or the gene chip according to any one of claims 4-5, or the kit according to claim 6 in pig breeding.
8. The application of the porcine SINE-RIP probe combination according to any one of claims 1-3, or the gene chip according to any one of claims 4-5, or the kit according to claim 6 in the identification and evaluation of porcine germplasm resources.
9. The application of the porcine SINE-RIP probe combination according to any one of claims 1-3, or the gene chip according to any one of claims 4-5, or the kit according to claim 6 in the localization of porcine functional genes or loci.
10. The use of the porcine SINE-RIP probe combination according to any one of claims 1-3, or the gene chip according to any one of claims 4-5, or the kit according to claim 6 in the analysis of porcine genetic diversity or kinship.