Preparation method, device, equipment and product of multi-target targeted capture probe

By using the annotation database and base replacement technology to generate multi-target probes during the preparation of targeted capture probes, the problem of low capture efficiency of single-target probes is solved, and more efficient target molecule capture is achieved.

CN120412730BActive Publication Date: 2025-10-03SHANGHAI JINFUKANG PHARMACEUTICAL ENGINEERING TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510918749.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-10-03
Estimated Expiration
2045-07-03

AI Technical Summary

Technical Problem

In the prior art, mixed single-target targeted capture probes have low capture efficiency of target molecules due to spatial arrangement conflicts and kinetic competition.

Method used

By obtaining the full-length molecular sequence and molecular identifier of the target molecule, identifying and screening it using a preset annotation database, multiple target structure exposure regions are generated, the initial probe sequence is generated by combining the linker sequence, and the target probe sequence is generated by base substitution, and finally a multi-target targeted capture probe is prepared.

Benefits of technology

The efficiency of capturing target molecules is improved, and the problem of low capture efficiency in the prior art is solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120412730B_ABST
    Figure CN120412730B_ABST
Patent Text Reader

Abstract

The embodiment of the present application provides a method, device, equipment and product for preparing a multi-target targeted capture probe, which relates to the field of medical biotechnology. On the basis of obtaining the full-length molecular sequence and molecular identifier of the target molecule, M original probe sequences are generated; N connecting arm sequences are obtained; an initial probe sequence is generated based on the M original probe sequences and the N connecting arm sequences; a feature analysis is performed on the initial probe sequence to generate a probe sequence feature; based on the probe sequence feature, at least one base in the initial probe sequence is replaced to generate a target probe sequence; based on the probe application information and the target probe sequence, a multi-target targeted capture probe is prepared; the target molecule is captured by the multi-target targeted capture probe, which solves the problem of spatial arrangement conflict and kinetic competition between the mixed single-target targeted capture probes, that is, solves the problem of low capture efficiency of target molecules caused by the solution of the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of medical biotechnology, and in particular to a method, device, equipment and product for preparing a multi-targeted capture probe. Background Art

[0002] With the development of life sciences, biological samples are becoming increasingly complex, containing a large number of biomolecules of different types and states. In order to efficiently and accurately capture target molecules from complex biological samples, the technology of targeted capture probes has emerged.

[0003] In the prior art, a method of identifying and capturing a target molecule is typically used, using a mixture of single-target targeted capture probes. Specifically, a single-target targeted capture probe binds to a single specific sequence of the target molecule, thereby achieving specific recognition and capture of the target molecule.

[0004] However, due to the spatial arrangement conflicts and kinetic competition between the mixed single-target targeted capture probes, the existing technology solutions lead to low capture efficiency of target molecules. Summary of the Invention

[0005] The embodiments of the present application provide a method, device, equipment and product for preparing a multi-target targeted capture probe to solve the problem of low capture efficiency of target molecules.

[0006] In a first aspect, an embodiment of the present application provides a method for preparing a multi-target targeted capture probe, comprising: obtaining the full-length molecular sequence and molecular identifier of the target molecule; identifying the full-length molecular sequence according to a preset annotation database to obtain functional domain annotation information, and each functional region corresponding to the functional domain annotation information; inputting the molecular identifier into a preset region identification database, and outputting the sequence variation region identifier and the structural exposure region identifier corresponding to the full-length molecular sequence; screening the structural exposure region identifier according to the sequence variation region identifier to obtain the screened structural exposure region identifier; performing structural prediction on the full-length molecular sequence according to the functional domain annotation information to generate intron structure information and exon structure information; and The structure exposure area corresponding to the screened structure exposure area identifier is screened to determine multiple target structure exposure areas; according to the target molecule sequences corresponding to the M target structure exposure areas in the multiple target structure exposure areas, M original probe sequences are generated; wherein M is an integer greater than or equal to 3; N connecting arm sequences are obtained; the N is the M minus 1; according to the M original probe sequences and the N connecting arm sequences, an initial probe sequence is generated; the initial probe sequence is feature analyzed to generate a probe sequence feature; according to the probe sequence feature, at least one base in the initial probe sequence is replaced to generate a target probe sequence; according to the probe application information and the target probe sequence, a multi-target targeted capture probe is prepared.

[0007] In one possible embodiment, replacing at least one base in the initial probe sequence according to the probe sequence characteristics to generate a target probe sequence includes: dividing the M original probe sequences into a first number of original probe sequences and a second number of original probe sequences; replacing at least one base in the first number of original probe sequences according to the probe sequence characteristics to generate a corresponding first number of replacement probe sequences; and generating the target probe sequence according to the first number of replacement probe sequences, the second number of original probe sequences, and N connecting arm sequences.

[0008] In one possible embodiment, replacing at least one base in the initial probe sequence according to the probe sequence characteristics to generate a target probe sequence includes: dividing the N tether sequences into a third number of tether sequences and a fourth number of tether sequences; replacing at least one base in the third number of tether sequences according to the probe sequence characteristics to generate a corresponding third number of replacement tether sequences; and generating the target probe sequence based on the third number of replacement tether sequences, the fourth number of tether sequences, and M original probe sequences.

[0009] In one possible embodiment, the replacing at least one base in the initial probe sequence according to the probe sequence characteristics to generate a target probe sequence includes: dividing the M original probe sequences into a fifth number of original probe sequences and a sixth number of original probe sequences; replacing at least one base of the fifth number of original probe sequences according to the probe sequence characteristics to generate a corresponding fifth number of replacement probe sequences; dividing the N tether sequences into a seventh number of tether sequences and an eighth number of tether sequences; replacing at least one base of the seventh number of tether sequences according to the probe sequence characteristics to generate a corresponding seventh number of replacement tether sequences; generating the target probe sequence according to the fifth number of replacement probe sequences, the sixth number of original probe sequences, the seventh number of replacement tether sequences, and the eighth number of tether sequences.

[0010] In one possible embodiment, the preparation of multi-target targeted capture probes based on the probe application information and the target probe sequence includes: obtaining molecular configuration information and application scenario information of the target molecule based on the probe application information; obtaining probe configuration information based on the molecular configuration information and the application scenario information; performing configuration processing on the target probe sequence based on the probe configuration information to obtain a target probe sequence after configuration design; and modifying the sequence head and / or sequence end of the target probe sequence after configuration design based on the application scenario information to prepare the multi-target targeted capture probe.

[0011] In one possible embodiment, the sequence beginning and / or sequence end of the target probe sequence after the configuration design is modified according to the application scenario information to prepare the multi-target targeted capture probe, including: determining the probe carrier type according to the application scenario information; determining the modifier according to the probe carrier type; and adding the modifier to the sequence beginning and / or sequence end of the target probe sequence after the configuration design to prepare the multi-target targeted capture probe.

[0012] In a second aspect, an embodiment of the present application provides a device for preparing a multi-target targeted capture probe, comprising:

[0013] The first processing module is used to obtain the full-length molecular sequence and molecular identifier of the target molecule; identify the full-length molecular sequence according to a preset annotation database to obtain functional domain annotation information and each functional region corresponding to the functional domain annotation information; input the molecular identifier into a preset region identification database, and output the sequence variation region identifier and the structural exposure region identifier corresponding to the full-length molecular sequence; screen the structural exposure region identifier according to the sequence variation region identifier to obtain the screened structural exposure region identifier; perform structural prediction on the full-length molecular sequence according to the functional domain annotation information to generate intron structure information and exon structure information; screen the structural exposure regions corresponding to the screened structural exposure region identifier according to the intron structure information and the exon structure information to determine multiple target structural exposure regions;

[0014] A second processing module is configured to generate M original probe sequences based on target molecule sequences corresponding to M target structure exposure regions among the multiple target structure exposure regions, wherein M is an integer greater than or equal to 3; obtain N tether arm sequences, wherein N is M minus 1; generate an initial probe sequence based on the M original probe sequences and the N tether arm sequences; perform feature analysis on the initial probe sequence to generate a probe sequence feature; and replace at least one base in the initial probe sequence based on the probe sequence feature to generate a target probe sequence;

[0015] The preparation module is used to prepare multi-target targeted capture probes based on the probe application information and the target probe sequence.

[0016] In one possible embodiment, when the second processing module replaces at least one base in the initial probe sequence according to the probe sequence characteristics to generate a target probe sequence, it is specifically used to: divide the M original probe sequences into a first number of original probe sequences and a second number of original probe sequences; replace at least one base in the first number of original probe sequences according to the probe sequence characteristics to generate a corresponding first number of replacement probe sequences; and generate the target probe sequence according to the first number of replacement probe sequences, the second number of original probe sequences, and N connecting arm sequences.

[0017] In one possible embodiment, when the second processing module replaces at least one base in the initial probe sequence according to the probe sequence characteristics to generate a target probe sequence, it is specifically used to: divide the N tether sequences into a third number of tether sequences and a fourth number of tether sequences; replace at least one base in the third number of tether sequences according to the probe sequence characteristics to generate a corresponding third number of replacement tether sequences; and generate the target probe sequence based on the third number of replacement tether sequences, the fourth number of tether sequences, and M original probe sequences.

[0018] In a possible embodiment, when the second processing module replaces at least one base in the initial probe sequence according to the probe sequence characteristics to generate a target probe sequence, it is specifically used to: divide the M original probe sequences into a fifth number of original probe sequences and a sixth number of original probe sequences; replace at least one base of the fifth number of original probe sequences according to the probe sequence characteristics to generate a corresponding fifth number of replacement probe sequences; divide the N tether sequences into a seventh number of tether sequences and an eighth number of tether sequences; replace at least one base of the seventh number of tether sequences according to the probe sequence characteristics to generate a corresponding seventh number of replacement tether sequences; and generate the target probe sequence according to the fifth number of replacement probe sequences, the sixth number of original probe sequences, the seventh number of replacement tether sequences, and the eighth number of tether sequences.

[0019] In one possible embodiment, when the preparation module prepares a multi-target targeted capture probe based on the probe application information and the target probe sequence, it is specifically used to: obtain the molecular configuration information and application scenario information of the target molecule based on the probe application information; obtain the probe configuration information based on the molecular configuration information and the application scenario information; perform configuration processing on the target probe sequence based on the probe configuration information to obtain a target probe sequence after configuration design; modify the sequence head and / or sequence end of the target probe sequence after configuration design based on the application scenario information to prepare the multi-target targeted capture probe.

[0020] In one possible embodiment, when the preparation module modifies the sequence beginning and / or sequence end of the target probe sequence after the configuration design according to the application scenario information to prepare the multi-target targeted capture probe, it is specifically used to: determine the probe carrier type according to the application scenario information; determine the modifier according to the probe carrier type; and add the modifier to the sequence beginning and / or sequence end of the target probe sequence after the configuration design to prepare the multi-target targeted capture probe.

[0021] In a third aspect, an embodiment of the present application provides an electronic device, comprising: a memory, a processor;

[0022] The memory stores computer-executable instructions;

[0023] The processor executes the computer-executable instructions stored in the memory, so that the processor executes the above first aspect and / or various possible implementations of the first aspect.

[0024] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the first aspect above and / or various possible implementation methods of the first aspect.

[0025] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the above first aspect and / or various possible implementation methods of the first aspect.

[0026] In a sixth aspect, an embodiment of the present application provides a multi-target targeted capture probe, which is prepared by the first aspect and / or various possible implementation methods of the first aspect.

[0027] In a seventh aspect, an embodiment of the present application provides a kit comprising the multi-targeted capture probe described in the sixth aspect.

[0028] The preparation method, device, equipment and product of the multi-target targeted capture probe provided in the embodiments of the present application, by obtaining the full-length molecular sequence and molecular identifier of the target molecule, calling a preset annotation database, identifying the full-length molecular sequence, and then obtaining functional domain annotation information, and each functional region corresponding to the functional domain annotation information, wherein each functional region corresponds to each sequence fragment in the target molecular sequence; the molecular identifier of the target molecule is input into a preset region identification database, and the sequence variation region identifier and the structure exposure region identifier corresponding to the full-length molecular sequence are output; then, the structure exposure region identifier is screened according to the sequence variation region identifier to obtain the screened structure exposure region identifier; further, the computer device performs structural prediction on the full-length molecular sequence according to the functional domain annotation information, and generates intron structure information and exon structure information; then, according to the intron structure information and the exon structure information, the screened The structural exposure area corresponding to the structural exposure area identification is screened to determine multiple target structural exposure areas; M original probe sequences are generated according to the target molecule sequences corresponding to M target structural exposure areas in the multiple target structural exposure areas; wherein M is an integer greater than or equal to 3; N connecting arm sequences are obtained; N is M minus 1; an initial probe sequence is generated according to the M original probe sequences and the N connecting arm sequences; a feature analysis is performed on the initial probe sequence to generate a probe sequence feature; according to the probe sequence feature, at least one base in the initial probe sequence is replaced to generate a target probe sequence; according to the probe application information and the target probe sequence, a multi-target targeted capture probe is prepared; the multi-target targeted capture probe solves the problem of spatial arrangement conflict and kinetic competition between the mixed single-target targeted capture probes, that is, it solves the problem of low target molecule capture efficiency caused by the solution of the prior art. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0030] Figure 1 A flowchart of a method for preparing a multi-target targeted capture probe provided in one embodiment of the present application;

[0031] Figure 2 A schematic structural diagram of a device for preparing a multi-target targeted capture probe according to one embodiment of the present application;

[0032] Figure 3 This is a schematic diagram of the structure of the electronic device provided in this application.

[0033] The above drawings illustrate specific embodiments of the present application, which will be described in more detail below. These drawings and the textual description are not intended to limit the scope of the present application in any way, but rather to illustrate the concepts of the present application to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION

[0034] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.

[0035] In the technical solution of this application, the user personal information involved and the collection, storage, use, processing, transmission, provision and disclosure of data are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0036] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0037] With the development of life sciences, the complexity of biological samples is increasing, which contains a large number of biological molecules of different types and states. In order to efficiently and accurately capture target molecules from complex biological samples, the technology of targeted capture probes has emerged. In the prior art, a mixed single-target targeted capture probe is usually used to identify and capture the target molecules. Specifically, the single-target targeted capture probe binds based on the single specific sequence of the target molecule, thereby achieving specific recognition and capture of the target molecule. However, due to the spatial arrangement conflicts and kinetic competition between the mixed single-target targeted capture probes, the solutions of the prior art lead to the problem of low capture efficiency of the target molecules.

[0038] The technical concept of the embodiments of the present application is explained below:

[0039] The execution subject of the method provided in the embodiment of the present application can be any form of electronic device, and a computer device is used as the execution subject for explanation. On the basis of obtaining the full-length molecular sequence and molecular identifier of the target molecule, the computer device calls a preset annotation database to identify the full-length molecular sequence, and then obtains functional domain annotation information, and each functional region corresponding to the functional domain annotation information, wherein each functional region corresponds to each sequence fragment in the target molecular sequence; the molecular identifier of the target molecule is input into a preset region identification database, and the sequence variation region identifier and the structural exposure region identifier corresponding to the full-length molecular sequence are output; and then, the structural exposure region identifier is screened according to the sequence variation region identifier to obtain the screened structural exposure region identifier; further, the computer device performs structural analysis on the full-length molecular sequence according to the functional domain annotation information. Predict and generate intron structure information and exon structure information; then, based on the intron structure information and exon structure information, screen the structure exposure regions corresponding to the screened structure exposure region identifiers to determine multiple target structure exposure regions; generate M original probe sequences based on the target molecular sequences corresponding to M target structure exposure regions in the multiple target structure exposure regions; wherein M is an integer greater than or equal to 3; obtain N linker arm sequences; N is M minus 1; generate an initial probe sequence based on the M original probe sequences and the N linker arm sequences; perform feature analysis on the initial probe sequence to generate a probe sequence feature; based on the probe sequence feature, replace at least one base in the initial probe sequence to generate a target probe sequence; prepare a multi-target targeted capture probe based on the probe application information and the target probe sequence.

[0040] The following specific embodiments describe in detail the technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0041] Figure 1 A flowchart of a method for preparing a multi-target targeted capture probe provided in one embodiment of the present application is shown in FIG. Figure 1 As shown, the execution subject of the method for preparing a multi-target targeted capture probe provided in this embodiment can be any form of electronic device. For example, this embodiment uses a computer device as the execution subject of the method of this embodiment. The method for preparing a multi-target targeted capture probe provided in this embodiment includes the following steps:

[0042] Step S101: Obtain the full-length molecular sequence and molecular identifier of the target molecule.

[0043] Exemplarily, based on the target detection object, the full-length molecular sequence of the target detection object and the molecular identifier of the target molecule are obtained. Specifically, for example, the molecular identifier of the target molecule is the molecular name (and / or gene symbol) of the target molecule and the molecular number (and / or gene number) in the relevant database. For example, the target molecule is the human TP53 gene, whose gene symbol is "TP53" and whose gene number in the National Center for Biotechnology Information database is 7157.

[0044] Step S102: identifying the full-length molecular sequence according to a preset annotation database to obtain functional domain annotation information and functional regions corresponding to the functional domain annotation information.

[0045] Exemplarily, the preset annotation database includes a public database and an annotation database established based on sequencing technology, wherein the public database includes the National Center for Biotechnology Information (NCBI) database and the Ensembl Genome Database.

[0046] Furthermore, by calling a preset annotation database, the computer device can identify the full-length molecular sequence and perform functional annotations, thereby obtaining functional domain annotation information and the functional regions corresponding to the functional domain annotation information, where each functional region corresponds to each sequence fragment in the target molecular sequence.

[0047] Step S103: inputting the molecular identifier into a preset region identification database, and outputting the sequence variation region identifier and the structural exposure region identifier corresponding to the full-length molecular sequence.

[0048] Exemplarily, the preset region identification database includes the disease-related human genome variation database (referred to as ClinVar database) and the protein structure database (Prorein Data Bank, referred to as PDB database) of the NCBI database.

[0049] The ClinVar database matches the input molecular identifiers and outputs the sequence variant region identifiers corresponding to the full-length sequence of the target molecule. Sequence variant region identifiers include mutation site region identifiers, splice variant region identifiers, and methylation region identifiers.

[0050] After receiving the input molecular identifier, the PDB database can be used to obtain the three-dimensional structural data of the target molecule. Combined with subsequent structural analysis, the regions in the target molecule that are in a structurally exposed state can be further identified, thereby determining the identification of its structurally exposed region. Specifically, the PDB structure (the three-dimensional structure of the target molecule) is processed by the structural annotation tool DSSP, and the solvent-accessible surface area (SASA) value of each amino acid residue is calculated to determine whether the region is a structurally exposed region, thereby obtaining the structurally exposed region identification. The judgment logic is: when the SASA value of an amino acid residue is greater than the preset threshold (25 ), the residue is considered to be in a structurally exposed state; otherwise, it is considered to be buried. Furthermore, the scope of the structurally exposed region can be determined based on the positional relationship of multiple residues determined to be exposed. Specifically, when multiple exposed residues are continuous in the primary structure, the continuous region can be defined as a structurally exposed region; or when multiple exposed residues are spatially adjacent in the three-dimensional structure (the farthest atomic distance between the residues is less than a preset threshold of 8 ), they can also be clustered into a structurally exposed region. Based on the above method, the scattered exposed residues can be organized into structurally continuous or spatially clustered exposed segments, and the structurally exposed regions can be output accordingly, that is, the structurally exposed region identifiers can be obtained.

[0051] Step S104 , screening the structural exposure region identifiers according to the sequence variation region identifiers to obtain screened structural exposure region identifiers.

[0052] Illustratively, the structural exposure region includes a sequence variation region, that is, the structural exposure region includes a mutation site region, a splicing variation region, and a methylation region; then, the structural exposure region identifier is screened according to the sequence variation region identifier to obtain a screened structural exposure region identifier, wherein the structural exposure region corresponding to the screened structural exposure region identifier does not include a sequence variation region.

[0053] Specifically, the sequence variation region is the region where abnormal variation exists in the full-length molecular sequence of the target molecule. Therefore, if the selected structural exposure region includes the sequence variation region, the multi-target targeted capture probe prepared based on the selected structural exposure region may have recognition deviations during the molecular capture process, thereby leading to poor detection accuracy and low detection efficiency; wherein, the molecular capture process is the detection process.

[0054] More specifically, for example, if there is a mutation in the bases in the mutation site region, and if the selected structural exposure region includes the mutation site region, then when the multi-target targeted capture probe finally prepared is performing molecular capture, if no corresponding base mutation occurs in the molecular sequence, it may cause the molecule to be missed, while the molecule should actually be captured, resulting in poor detection accuracy and low detection efficiency; therefore, it is necessary to eliminate the structural exposure region including the mutation site region.

[0055] Targeting splice variant regions, which are common during gene expression and produce multiple distinct transcripts, is crucial. If multi-target capture probes are designed based on molecular sequences within splice variant regions, molecular capture may only detect a subset of specific transcripts, failing to capture all expression products. This results in poor detection accuracy and efficiency. Therefore, structurally exposed regions, including those at the mutation site, need to be excluded.

[0056] Targeting methylated regions, methylation is an epigenetic modification that can alter the local structure of a molecular sequence. However, this change is variable and uncertain, and cannot be guaranteed to be stable across all targets. Therefore, if multi-targeted capture probes are designed based on the molecular sequence in methylated regions, they may only detect a subset of molecules, leading to poor detection accuracy and low efficiency. Therefore, structurally exposed regions, including methylated regions, need to be eliminated.

[0057] In one possible implementation, since the sequence variation region is described based on a two-dimensional linear coordinate system, such as the position of the sequence variation region is chr1:100-260, and the structural exposure region is described based on a three-dimensional spatial coordinate system, such as the position coordinates of the structural exposure region are (x, y, z), after obtaining the sequence variation region identifier and multiple structural exposure region identifiers based on step S103, the position of the sequence variation region, position_1, is determined based on the sequence variation region identifier, for example, position_1 is chr1:100-260. Then, based on the base sequence and position_1 of the sequence variation region, the three-dimensional spatial position coordinate position_2 of the sequence variation region is determined. Based on multiple structural exposure region identifiers, label_1, label_2, and label_3, the position coordinate position_3 of the structural exposure region corresponding to label_1, the position coordinate position_4 of the structural exposure region corresponding to label_2, and the position coordinate position_5 of the structural exposure region corresponding to label_3 are determined, respectively. Then, the spatial distance dist_1 between the position coordinate position_2 and the position coordinate position_3 is calculated, and the position coordinates are calculated. The spatial distance dist_2 between position coordinate position_2 and position coordinate position_4 is calculated, and the spatial distance dist_3 between position coordinate position_2 and position coordinate position_5 is calculated; if the spatial distances dist_1, dist_2 and dist_3 are all greater than the distance threshold, it indicates that the structural exposure region corresponding to the structural exposure region identifier determined in step S103 does not include the sequence variation region; if any one of the spatial distances dist_1, dist_2 and dist_3 is less than or equal to the distance threshold, for example, the spatial distance dist_2 is less than the distance threshold, it indicates that the structural exposure region corresponding to the structural exposure region identifier label_2 corresponds to the sequence variation region corresponding to the sequence variation region identifier, and then the structural exposure region identifier label_2 is eliminated to obtain the screened structural exposure region identifiers, i.e., the structural exposure region identifiers label_1 and label_3.

[0058] Step S105 , performing structure prediction on the full-length molecular sequence according to the functional domain annotation information, and generating intron structure information and exon structure information.

[0059] Illustratively, a computer device calls a bioinformatics tool to perform structure prediction on the full-length molecular sequence based on the functional domain annotation information, where the structure prediction includes RNA secondary structure prediction and protein tertiary structure prediction; further, based on the prediction results of the structure prediction, the intron structure and exon structure are predicted, that is, the intron structure information and the exon structure information are obtained.

[0060] Step S106 , screening the structural exposure regions corresponding to the screened structural exposure region identifiers according to the intron structural information and the exon structural information, and determining a plurality of target structural exposure regions.

[0061] For example, based on the intron structure and exon structure, it can be determined whether the structural exposure region after screening can serve as the binding functional region of the probe; for example, if a structural exposure region after screening is located in the intron structure and will be cleaved after normal transcription, then this region cannot serve as the binding functional region of the probe, that is, it cannot be determined as the target structural exposure region; if a structural exposure region is located in the exon structure, or in a region that will not be cleaved in a specific transcript variant, then this region can serve as the binding functional region of the probe, that is, it can be determined as the target structural exposure region. Then, based on the intron structure information and the exon structure information, the structural exposure regions corresponding to the structural exposure region identifier after screening are screened to determine multiple target structural exposure regions.

[0062] Step S107 , generating M original probe sequences according to the target molecule sequences corresponding to the M target structure exposure regions among the plurality of target structure exposure regions; wherein M is an integer greater than or equal to 3.

[0063] For example, reverse complementary pairing is performed on the target molecule sequence corresponding to the target structure exposure region to generate the corresponding original probe sequence; for example, the target molecule sequence SEQ ID NO: 1 corresponding to a certain target structure exposure region is 5'-ATGGGCTGCTGGTTTGTAAAC-3', and reverse complementary pairing is performed on the target molecule sequence, and the corresponding original probe sequence SEQ ID NO: 2 is 3'-TACCCGACGACCAAACATTTG-5'; wherein, the 3' end (3 prime end) is the "third carbon atom" end of the molecular chain, which is usually used as the end of the chain; the 5' end (5 prime end) is the "fifth carbon atom" end of the molecular chain, which is usually used as the starting end of the chain. Furthermore, based on the above-mentioned reverse complementary pairing processing method, M original probe sequences can be generated according to the target molecule sequences corresponding to M target structure exposure regions in multiple target structure exposure regions.

[0064] Step S108, obtaining N connecting arm sequences.

[0065] Exemplarily, the linker sequence is used to connect the original probe sequences so that the original probe sequences are flexible and do not form secondary structure interference, thereby improving the binding efficiency; wherein N is M minus 1.

[0066] Illustratively, the tether sequences include TTTTT, TCTCT, TTTCTTT, ATATAT, TATATA.

[0067] In one possible implementation, N linker sequences are determined based on M original probe sequences; specifically, N linker sequences are determined based on the base composition and secondary structure tendency in the M original probe sequences to avoid the selected linker sequences increasing the probability of nonspecific binding and / or structural interference in the finally generated target probe sequence.

[0068] Step S109: generating an initial probe sequence based on the M original probe sequences and the N linker arm sequences.

[0069] Exemplarily, the initial probe sequence can be generated by connecting M original probe sequences in series based on N linker sequences.

[0070] Exemplarily, M is 3, and the three original probe sequences are SEQ ID NO:2: 3'-TACCCGACGACCAAACATTTG-5', SEQ ID NO:3: 3'-GCCGCGTACTTCTTGGTCT-5', and SEQ ID NO:4: 3'-CTCCAACGAGATCCGTGGA-5'; N is 2, and the two linker sequences are TCTCT and TTTCTTT; furthermore, based on the N linker sequences, the M original probe sequences are connected in series, and the generated initial probe sequence SEQ ID NO:5 is 5'-TACCCGACGACCAAACATTTGTCTCTGCCGCGTACTTCTTGGTCTTTTCTTTCTCCAACGAGATCCGTGGA-3'.

[0071] Step S110 , performing feature analysis on the initial probe sequence to generate probe sequence features.

[0072] For example, the computer device calls the corresponding pre-trained sequence analysis model to perform feature analysis on the initial probe sequence to generate probe sequence features; wherein the probe sequence features include self-complementarity, GC content, melting temperature, free energy G value and poor structural sequence region, poor structural sequence region will lead to the generation of hairpin structure, self-dimer or G-quadruplex.

[0073] Step S111 : replacing at least one base in the initial probe sequence according to the probe sequence characteristics to generate a target probe sequence.

[0074] For example, according to the characteristics of the probe sequence, at least one base in the initial probe sequence is replaced to adjust the characteristics of the probe sequence and generate a target probe sequence. Specifically, for example, by replacing at least one base in the initial probe sequence, the number of self-complementary regions corresponding to self-complementarity is reduced, the GC content is within the target content range, the melting temperature is within the target melting temperature range, and the free energy is increased. G value is at the target free energy The number of regions with poorly structured sequence regions was reduced within the G value range.

[0075] The embodiment steps of the present application ensure the thermodynamic stability, binding affinity and spatial flexibility of the generated target probe sequence by replacing at least one base in the initial probe sequence, thereby improving the capture efficiency of the target molecule by the multi-target targeted capture probe prepared based on the target probe sequence in the subsequent steps.

[0076] Specifically, in a possible implementation, the specific implementation steps of step S111 include:

[0077] Step S201 : dividing M original probe sequences into a first number of original probe sequences and a second number of original probe sequences.

[0078] Step S202: replacing at least one base of a first number of original probe sequences according to the probe sequence characteristics to generate a corresponding first number of replacement probe sequences.

[0079] Step S203 , generating a target probe sequence according to the first number of replacement probe sequences, the second number of original probe sequences, and N tether sequences.

[0080] Exemplarily, M is 3, and the three original probe sequences are SEQ ID NO:2: 3'-TACCCGACGACCAAACATTTG-5', SEQ ID NO:3: 3'-GCCGCGTACTTCTTGGTCT-5', and SEQ ID NO:4: 3'-CTCCAACGAGATCCGTGGA-5'. The M original probe sequences are divided into a first number of original probe sequences and a second number of original probe sequences, where the second number can be 0. When the second number is 0, the first number is M, so that the sum of the first number and the second number is M. For example, the first number is 1 and the second number is 2. Furthermore, at least one base of the original probe sequence SEQ ID NO:3: 3'-GCCGCGTACTTCTTGGTCT-5' is replaced to generate the corresponding replacement probe sequence SEQ ID NO:6: 3'-GCCGCGAACTTCTTGGTCT-5'. The other two original probe sequences are not base replaced. N is 2, and the two linker sequences are TCTCT and TTTCTTT. Furthermore, based on the first number of replacement probe sequences (SEQ ID NO: 6), the second number of original probe sequences (SEQ ID NO: 2, SEQ ID NO: 4), and N linker sequences (TCTCT, TTTCTTT), the generated target probe sequence SEQ ID NO: 7 is 5'-TACCCGACGACCAAACATTTGTCTCTGCCGCGAACTTCTTGGTCTTTTCTTTCTCCAACGAGATCCGTGGA-3'.

[0081] In another possible implementation, the specific implementation steps of step S111 include:

[0082] Step S301: Divide N tether arm sequences into a third number of tether arm sequences and a fourth number of tether arm sequences.

[0083] Step S302 : replacing at least one base of a third number of tether sequences according to the probe sequence characteristics to generate a corresponding third number of replaced tether sequences.

[0084] Step S303 , generating a target probe sequence according to the third number of replacement tether sequences, the fourth number of tether sequences, and the M original probe sequences.

[0085] For example, M is 3, and the three original probe sequences are SEQ ID NO:2: 3'-TACCCGACGACCAAACATTTG-5', SEQ ID NO:3: 3'-GCCGCGTACTTCTTGGTCT-5', and SEQ ID NO:4: 3'-CTCCAACGAGATCCGTGGA-5'; N is 2, and both tether sequences are TTTT; furthermore, the initial probe sequence SEQ ID NO:8 is 5'-TACCCGACGACCAAACATTTGTTTTTGCCGCGTACTTCTTGGTCTTTTTTCTCCAACGAGATCCGTGGA-3'. The N tether sequences are divided into a third number of tether sequences and a fourth number of tether sequences, wherein the fourth number can be 0. When the fourth number is 0, the third number is N, so that the sum of the third number and the fourth number is N. For example, N is 2, the third number is 1, and the fourth number is 1. Furthermore, at least one base in one of the tether sequences TTTTT is replaced to generate a corresponding replacement tether sequence TTAAT, while the other tether sequence is not base-substituted. Furthermore, based on the third number of replacement tether sequences (TTAAT), the fourth number of tether sequences (TTTTT), and M original probe sequences (SEQ ID NO:2, SEQ ID NO:3, and SEQ ID NO:4), the generated target probe sequence SEQ ID NO:9 is 5'-TACCCGACGACCAAACATTTGTTAATGCCGCGTACTTCTTGGTCTTTTTTCTCCAACGAGATCCGTGGA-3'.

[0086] In another possible implementation, the specific implementation steps of step S111 include:

[0087] Step S401 : dividing M original probe sequences into a fifth number of original probe sequences and a sixth number of original probe sequences.

[0088] Step S402 : replacing at least one base of a fifth number of original probe sequences according to the probe sequence characteristics to generate a fifth number of corresponding replacement probe sequences.

[0089] Exemplarily, M is 3, and the three original probe sequences are SEQ ID NO:2: 3'-TACCCGACGACCAAACATTTG-5', SEQ ID NO:3: 3'-GCCGCGTACTTCTTGGTCT-5', and SEQ ID NO:4: 3'-CTCCAACGAGATCCGTGGA-5'; the M original probe sequences are divided into a fifth number of original probe sequences and a sixth number of original probe sequences, where the sixth number can be 0. When the sixth number is 0, the fifth number is M, so that the sum of the fifth number and the sixth number is M. For example, the fifth number is 1 and the sixth number is 2, and then, at least one base of the original probe sequence SEQ ID NO:3: 3'-GCCGCGTACTTCTTGGTCT-5' is replaced to generate a corresponding replacement probe sequence SEQ ID NO:6: 3'-GCCGCGAACTTCTTGGTCT-5'; the other two original probe sequences are not base replaced.

[0090] Step S403: Divide the N tether arm sequences into a seventh number of tether arm sequences and an eighth number of tether arm sequences.

[0091] Step S404 : replacing at least one base of the seventh number of tether sequences according to the probe sequence characteristics to generate a corresponding seventh number of replaced tether sequences.

[0092] For example, N is 2, and both tether sequences are TTTTT. Furthermore, the N tether sequences are divided into a seventh number of tether sequences and an eighth number of tether sequences, where the eighth number can be 0. When the eighth number is 0, the seventh number is N, so that the sum of the seventh number and the eighth number is N. For example, the seventh number is 1, and the eighth number is 1. Furthermore, at least one base in one of the tether sequences TTTTT is replaced to generate a corresponding replaced tether sequence TTAAT, and the other tether sequence is not base-substituted.

[0093] Step S405 , generating a target probe sequence according to the fifth number of replacement probe sequences, the sixth number of original probe sequences, the seventh number of replacement tether sequences, and the eighth number of tether sequences.

[0094] Illustratively, the initial probe sequence SEQ ID NO: 8 is 5'-TACCCGACGACCAAACATTTGTTTTTGCCGCGTACTTCTTGGTCTTTTTTCTCCAACGAGATCCGTGGA-3'; based on the fifth number of replacement probe sequences (SEQ ID NO: 6), the sixth number of original probe sequences (SEQ ID NO: 2, SEQ ID NO: 4), the seventh number of replacement tether sequences (TTAAT), and the eighth number of tether sequences (TTTTT), the generated target probe sequence SEQ ID NO: 10 is 5'-TACCCGACGACCAAACATTTGTTAATGCCGCGAACTTCTTGGTCTTTTTTCTCCAACGAGATCCGTGGA-3'.

[0095] It can be understood that the process of processing the sequence based on the characteristics of the probe sequence provided in the embodiment of the present application only refers to replacing at least one base in the initial probe sequence to generate a target probe sequence. The relative position relationship of the replacement probe sequence, the original probe sequence, the linker sequence, and the replacement linker sequence in the target probe sequence is consistent with the relative position relationship of the original probe sequence and the linker sequence in the initial probe sequence. For example, the original probe sequences in the initial probe sequence are a, b, and c, and the linker sequences are d and e; the initial probe sequence is adbec; at least one base in the original probe sequence b is replaced to generate a replacement probe sequence b', and at least one base in the linker sequence d is replaced to generate a replacement linker sequence d', and the generated target probe sequence is a-d'-b'-ec.

[0096] Furthermore, the method provided in the embodiment of the present application further includes: replacing at least one base in the initial probe sequence according to the probe application information and the probe sequence characteristics to generate a target probe sequence. Specifically, for example, according to the probe application information, determining the characteristic constraint threshold; according to the characteristic constraint threshold, including the characteristic constraint number of the self-complementary region, the characteristic constraint content range, the characteristic constraint melting temperature range, the characteristic constraint free energy G value range, the number of characteristic constraint regions; further, at least one base in the initial probe sequence is replaced according to the characteristic constraint threshold value, so that the number of self-complementary regions corresponding to self-complementarity is less than or equal to the characteristic constraint number of the self-complementary region, the GC content is within the characteristic constraint content range, the melting temperature is within the characteristic constraint melting temperature range, and the free energy is within the characteristic constraint melting temperature range. G value is the characteristic constraint free energy The G value range is within the range, and the number of regions with bad structural sequence regions is less than or equal to the number of feature constraint regions.

[0097] Step S112: preparing multi-target targeted capture probes according to the probe application information and the target probe sequence.

[0098] Specifically, the specific implementation steps of step S112 include:

[0099] Step S1121: Obtain molecular configuration information and application scenario information of the target molecule according to the probe application information.

[0100] Step S1122: Obtain probe configuration information based on the molecular configuration information and application scenario information.

[0101] For example, if the molecular configuration of the target molecule is cyclic, then the probe configuration is cyclic, and the combination of the two is more stable; if the molecular configuration of the target molecule is linear, then the probe configuration is linear, and the combination of the two is more stable. If the application scenario is a rapid on-site detection scenario, the kit with a linear probe configuration is easy to operate and does not require complex instruments and professional operators; if the application scenario is a high-sensitivity detection scenario, the probe with a circular configuration has high stability and high specificity. The circular configuration can reduce nonspecific binding and can tightly bind to the target molecule with high sensitivity. Furthermore, based on the molecular configuration information and the application scenario information, the probe configuration information is obtained.

[0102] Step S1123 , performing configuration processing on the target probe sequence according to the probe configuration information to obtain a target probe sequence after configuration design.

[0103] For example, if the probe configuration information is linear configuration information, the target probe sequence is configured according to the linear configuration information, and the configuration of the target probe sequence after the configuration design is a linear configuration. If the probe configuration information is sandwich configuration information, the target probe sequence is configured according to the sandwich configuration information, and the configuration of the target probe sequence after the configuration design is a sandwich configuration.

[0104] Step S1124 , modifying the sequence head and / or sequence end of the target probe sequence after configuration design according to the application scenario information to prepare a multi-target targeted capture probe.

[0105] Specifically, the specific implementation steps of step S1124 include:

[0106] Step S11241: Determine the probe carrier type according to the application scenario information.

[0107] Step S11242: Determine the modifier according to the probe carrier type.

[0108] Step S11243, adding a modifier to the sequence head and / or sequence end of the target probe sequence after configuration design to prepare a multi-target targeted capture probe.

[0109] Exemplarily, for example, according to the application scenario information, the probe carrier type is a magnetic bead carrier, wherein the magnetic bead carrier is modified with streptavidin, and the corresponding modifier is biotin, and biotin and streptavidin have a strong binding force; then, the sequence head and / or sequence end of the target probe sequence after the configuration design is modified with biotin to prepare a multi-target targeted capture probe.

[0110] For another example, according to the application scenario information, the probe carrier type is a nanocarrier, wherein the nanocarrier is modified with gold nanoparticles, and the corresponding modifier is a thiol modification group, in which the sulfhydryl in the group can form a stable gold-sulfur covalent bond with the surface of the gold nanoparticles; furthermore, the sequence head and / or sequence end of the target probe sequence after the configuration design is modified with the thiol modification group to prepare a multi-target targeted capture probe.

[0111] In this embodiment, on the basis of obtaining the full-length molecular sequence and molecular identifier of the target molecule, a preset annotation database is called to identify the full-length molecular sequence, thereby obtaining functional domain annotation information and each functional region corresponding to the functional domain annotation information, wherein each functional region corresponds to each sequence fragment in the target molecular sequence; the molecular identifier of the target molecule is input into a preset region identification database, and the sequence variation region identifier and the structure exposure region identifier corresponding to the full-length molecular sequence are output; then, the structure exposure region identifier is screened according to the sequence variation region identifier to obtain the screened structure exposure region identifier; further, the computer device performs structure prediction on the full-length molecular sequence according to the functional domain annotation information, and generates intron structure information and exon structure information; then, according to the intron structure information and the exon structure information, the structure exposure region identifier corresponding to the screened structure exposure region identifier is screened. The exposed areas are screened to determine multiple target structure exposed areas; M original probe sequences are generated according to the target molecule sequences corresponding to M target structure exposed areas in the multiple target structure exposed areas; wherein M is an integer greater than or equal to 3; N linker arm sequences are obtained; N is M minus 1; an initial probe sequence is generated according to the M original probe sequences and the N linker arm sequences; a feature analysis is performed on the initial probe sequence to generate a probe sequence feature; according to the probe sequence feature, at least one base in the initial probe sequence is replaced to generate a target probe sequence; according to the probe application information and the target probe sequence, a multi-target targeted capture probe is prepared; the multi-target targeted capture probe solves the problems of spatial arrangement conflict and kinetic competition between the mixed single-target targeted capture probes, that is, solves the problem of low target molecule capture efficiency caused by the solutions of the prior art.

[0112] One embodiment of the present application provides a multi-target targeted capture probe, which is Figure 1 The method is prepared according to the technical solution of the embodiment shown.

[0113] One embodiment of the present application provides a kit, the kit comprising: Figure 1 The multi-target targeted capture probe is prepared according to the technical solution of the method embodiment shown.

[0114] Furthermore, the kit provided in the embodiments of the present application also includes primers for amplifying base sequences having regions that hybridize with the multi-target targeted capture probes.

[0115] Figure 2 A schematic diagram of a device for preparing a multi-targeted capture probe according to an embodiment of the present invention is shown in FIG. Figure 2 As shown, the preparation device 3 of the multi-target targeted capture probe provided in this embodiment includes:

[0116] The first processing module 31 is used to obtain the full-length molecular sequence and molecular identifier of the target molecule; identify the full-length molecular sequence according to a preset annotation database to obtain functional domain annotation information and each functional region corresponding to the functional domain annotation information; input the molecular identifier into a preset region identification database, and output the sequence variation region identifier and the structural exposure region identifier corresponding to the full-length molecular sequence; filter the structural exposure region identifier according to the sequence variation region identifier to obtain filtered structural exposure region identifiers; perform structural prediction on the full-length molecular sequence according to the functional domain annotation information to generate intron structure information and exon structure information; filter the structural exposure regions corresponding to the filtered structural exposure region identifiers according to the intron structure information and the exon structure information to determine multiple target structural exposure regions;

[0117] The second processing module 32 is configured to generate M original probe sequences based on target molecule sequences corresponding to M target structure exposure regions among the plurality of target structure exposure regions, wherein M is an integer greater than or equal to 3; obtain N tether arm sequences, wherein N is M minus 1; generate an initial probe sequence based on the M original probe sequences and the N tether arm sequences; perform feature analysis on the initial probe sequences to generate probe sequence features; and replace at least one base in the initial probe sequence based on the probe sequence features to generate a target probe sequence;

[0118] The preparation module 33 is used to prepare multi-target targeted capture probes based on the probe application information and the target probe sequence.

[0119] In one possible embodiment, when the second processing module 32 replaces at least one base in the initial probe sequence according to the probe sequence characteristics to generate a target probe sequence, it is specifically used to: divide M original probe sequences into a first number of original probe sequences and a second number of original probe sequences; replace at least one base in the first number of original probe sequences according to the probe sequence characteristics to generate a corresponding first number of replacement probe sequences; and generate a target probe sequence based on the first number of replacement probe sequences, the second number of original probe sequences, and N connecting arm sequences.

[0120] In one possible embodiment, when the second processing module 32 replaces at least one base in the initial probe sequence according to the probe sequence characteristics to generate a target probe sequence, it is specifically used to: divide the N tether sequences into a third number of tether sequences and a fourth number of tether sequences; replace at least one base in the third number of tether sequences according to the probe sequence characteristics to generate a corresponding third number of replacement tether sequences; and generate a target probe sequence based on the third number of replacement tether sequences, the fourth number of tether sequences, and the M original probe sequences.

[0121] In one possible embodiment, when the second processing module 32 replaces at least one base in the initial probe sequence according to the probe sequence characteristics to generate a target probe sequence, it is specifically used to: divide the M original probe sequences into a fifth number of original probe sequences and a sixth number of original probe sequences; replace at least one base of the fifth number of original probe sequences according to the probe sequence characteristics to generate a corresponding fifth number of replacement probe sequences; divide the N tether sequences into a seventh number of tether sequences and an eighth number of tether sequences; replace at least one base of the seventh number of tether sequences according to the probe sequence characteristics to generate a corresponding seventh number of replacement tether sequences; and generate a target probe sequence according to the fifth number of replacement probe sequences, the sixth number of original probe sequences, the seventh number of replacement tether sequences, and the eighth number of tether sequences.

[0122] In one possible embodiment, when preparing a multi-target targeted capture probe based on the probe application information and the target probe sequence, the preparation module 33 is specifically used to: obtain the molecular configuration information and application scenario information of the target molecule based on the probe application information; obtain the probe configuration information based on the molecular configuration information and the application scenario information; perform configuration processing on the target probe sequence based on the probe configuration information to obtain the target probe sequence after configuration design; modify the sequence head and / or sequence end of the target probe sequence after configuration design based on the application scenario information to prepare a multi-target targeted capture probe.

[0123] In one possible embodiment, when the preparation module 33 modifies the sequence beginning and / or sequence end of the target probe sequence after configuration design according to the application scenario information to prepare a multi-target targeted capture probe, it is specifically used to: determine the probe carrier type according to the application scenario information; determine the modifier according to the probe carrier type; add the modifier to the sequence beginning and / or sequence end of the target probe sequence after configuration design to prepare a multi-target targeted capture probe.

[0124] The first processing module 31, the second processing module 32 and the preparation module 33 are connected in sequence. The preparation device 3 of the multi-target targeted capture probe provided in this embodiment can perform the following steps: Figure 1 The technical solution of the method embodiment shown has similar implementation principles and technical effects, which will not be repeated here.

[0125] Figure 3 This is a schematic diagram of the structure of the electronic device provided in this application. Figure 3 As shown, the electronic device 50 provided in this embodiment includes: at least one processor 501 and a memory 502. Optionally, the device 50 further includes a communication component 503. The processor 501, the memory 502 and the communication component 503 are connected via a bus 504.

[0126] In a specific implementation process, at least one processor 501 executes the computer-executable instructions stored in the memory 502, so that the at least one processor 501 performs the above method.

[0127] The specific implementation process of the processor 501 can be found in the above method embodiment. Its implementation principle and technical effects are similar and will not be repeated here in this embodiment.

[0128] In the above embodiments, it should be understood that the processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASICs), etc. A general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the present invention may be directly executed by a hardware processor or by a combination of hardware and software modules within the processor.

[0129] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage.

[0130] A bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. Buses can be categorized as address buses, data buses, and control buses. For ease of illustration, the buses in the drawings of this application are not limited to just one bus or just one type of bus.

[0131] The present application also provides a computer program product, including a computer program, which implements the above method when executed by a processor.

[0132] The present application also provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the above method is implemented.

[0133] The readable storage medium may be implemented by any type of volatile or non-volatile memory device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium may be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0134] An exemplary readable storage medium is coupled to a processor so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be an integral part of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor and the readable storage medium can also exist in the device as discrete components.

[0135] The division of units is merely a logical functional division; actual implementations may employ alternative divisions, such as combining or integrating multiple units or components into another system, or omitting or disabling certain features. Furthermore, any direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between devices or units, either through an interface, electrical, mechanical, or other means.

[0136] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0137] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0138] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the method of the present invention. The aforementioned storage medium includes various media that can store program code, such as USB flash drives, mobile hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.

[0139] Those skilled in the art will appreciate that all or part of the steps in the above-described method embodiments can be implemented using hardware associated with program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0140] Finally, it should be noted that those skilled in the art will readily identify other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. The present invention is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of the present invention and include common knowledge or customary techniques in the art not disclosed herein. The present invention is not limited to the precise structure described above and illustrated in the accompanying drawings, and various modifications and variations may be made without departing from the scope thereof. The scope of the present invention is limited solely by the appended claims.

Claims

1. A method for preparing a multi-target capture probe, characterized in that: The method comprises: Obtain the full-length molecular sequence and molecular identity of the target molecule; Identify the full-length molecular sequence according to a preset annotation database to obtain functional domain annotation information and functional regions corresponding to the functional domain annotation information; Inputting the molecular identifier into a preset region identification database, and outputting the sequence variation region identifier and the structural exposure region identifier corresponding to the full-length molecular sequence; screening the structural exposure region identifiers according to the sequence variation region identifiers to obtain screened structural exposure region identifiers; Performing structural prediction on the full-length molecular sequence according to the functional domain annotation information to generate intron structure information and exon structure information; Screening the structural exposure regions corresponding to the screened structural exposure region identifiers according to the intron structural information and the exon structural information to determine a plurality of target structural exposure regions; Generate M original probe sequences according to target molecule sequences corresponding to M target structure exposure regions among the multiple target structure exposure regions; wherein M is an integer greater than or equal to 3; Obtaining N tether arm sequences, wherein N is M minus 1; Generating an initial probe sequence according to the M original probe sequences and the N connecting arm sequences; performing feature analysis on the initial probe sequence to generate probe sequence features; replacing at least one base in the initial probe sequence according to the probe sequence characteristics to generate a target probe sequence; Preparing multi-target targeted capture probes according to the probe application information and the target probe sequence; The method of preparing a multi-target targeted capture probe according to the probe application information and the target probe sequence comprises: Obtaining molecular configuration information and application scenario information of the target molecule according to the probe application information; Obtaining probe configuration information according to the molecular configuration information and the application scenario information; Performing configuration processing on the target probe sequence according to the probe configuration information to obtain a target probe sequence after configuration design; According to the application scenario information, the sequence head and / or sequence end of the target probe sequence after the configuration design is modified to prepare the multi-target targeted capture probe; The obtaining of N tether arm sequences comprises: Determining the N connecting arm sequences according to the M original probe sequences; The step of replacing at least one base in the initial probe sequence according to the probe sequence characteristics to generate a target probe sequence comprises: At least one base in the initial probe sequence is replaced according to the probe application information and the probe sequence characteristics to generate the target probe sequence.

2. The method according to claim 1, characterized in that The step of replacing at least one base in the initial probe sequence according to the probe sequence characteristics to generate a target probe sequence comprises: dividing the M original probe sequences into a first number of original probe sequences and a second number of original probe sequences; replacing at least one base of the first number of original probe sequences according to the probe sequence characteristics to generate a corresponding first number of replacement probe sequences; The target probe sequence is generated according to the first number of replacement probe sequences, the second number of original probe sequences, and N connecting arm sequences.

3. The method according to claim 1, characterized in that The step of replacing at least one base in the initial probe sequence according to the probe sequence characteristics to generate a target probe sequence comprises: dividing the N tether sequences into a third number of tether sequences and a fourth number of tether sequences; replacing at least one base of the third number of tether sequences according to the probe sequence characteristics to generate a corresponding third number of replaced tether sequences; The target probe sequence is generated according to the third number of replacement tether sequences, the fourth number of tether sequences, and M original probe sequences.

4. The method according to claim 1, wherein The step of replacing at least one base in the initial probe sequence according to the probe sequence characteristics to generate a target probe sequence comprises: dividing the M original probe sequences into a fifth number of original probe sequences and a sixth number of original probe sequences; replacing at least one base of the fifth number of original probe sequences according to the probe sequence characteristics to generate a corresponding fifth number of replacement probe sequences; dividing the N tether sequences into a seventh number of tether sequences and an eighth number of tether sequences; replacing at least one base of the seventh number of tether sequences according to the probe sequence characteristics to generate a corresponding seventh number of replaced tether sequences; The target probe sequence is generated according to the fifth number of replacement probe sequences, the sixth number of original probe sequences, the seventh number of replacement tether sequences, and the eighth number of tether sequences.

5. The method according to claim 1, wherein The method of modifying the sequence head and / or sequence end of the target probe sequence after the configuration design according to the application scenario information to prepare the multi-target targeted capture probe comprises: Determining the probe carrier type according to the application scenario information; Determine the modifier according to the probe carrier type; The modifier is added to the sequence head and / or sequence end of the target probe sequence after the configuration design to prepare the multi-target targeted capture probe.

6. A device for preparing a multi-target targeted capture probe, characterized in that: include: The first processing module is used to obtain the full-length molecular sequence and molecular identifier of the target molecule; According to a preset annotation database, the full-length molecular sequence is identified to obtain functional domain annotation information and each functional region corresponding to the functional domain annotation information; the molecular identifier is input into a preset region identification database, and the sequence variation region identifier and the structural exposure region identifier corresponding to the full-length molecular sequence are output; the structural exposure region identifier is screened according to the sequence variation region identifier to obtain a screened structural exposure region identifier; Performing structural prediction on the full-length molecular sequence according to the functional domain annotation information to generate intron structure information and exon structure information; Screening the structural exposure regions corresponding to the screened structural exposure region identifiers according to the intron structural information and the exon structural information to determine a plurality of target structural exposure regions; A second processing module is configured to generate M original probe sequences based on target molecule sequences corresponding to M target structure exposure regions among the multiple target structure exposure regions, wherein M is an integer greater than or equal to 3; obtain N tether arm sequences, wherein N is M minus 1; generate an initial probe sequence based on the M original probe sequences and the N tether arm sequences; perform feature analysis on the initial probe sequence to generate a probe sequence feature; and replace at least one base in the initial probe sequence based on the probe sequence feature to generate a target probe sequence; A preparation module, used for preparing multi-target targeted capture probes according to the probe application information and the target probe sequence; When preparing multi-target targeted capture probes based on the probe application information and the target probe sequence, the preparation module is specifically used to: Obtaining molecular configuration information and application scenario information of the target molecule according to the probe application information; Obtaining probe configuration information according to the molecular configuration information and the application scenario information; Performing configuration processing on the target probe sequence according to the probe configuration information to obtain a target probe sequence after configuration design; According to the application scenario information, the sequence head and / or sequence end of the target probe sequence after the configuration design is modified to prepare the multi-target targeted capture probe; When acquiring N tether arm sequences, the second processing module is specifically configured to: Determining the N connecting arm sequences according to the M original probe sequences; When the second processing module replaces at least one base in the initial probe sequence according to the probe sequence characteristics to generate a target probe sequence, it is specifically configured to: At least one base in the initial probe sequence is replaced according to the probe application information and the probe sequence characteristics to generate the target probe sequence.

7. An electronic device, characterized in that: include: a processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 5 when executed by a processor.

9. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 5 when executed by a processor.

Citation Information

Patent Citations

  • Preparation method of nucleic acid targeted capture sequencing library based on long chain molecule inversion probe

    CN108396057A

  • Systems and Computer Program Products for Probe Set Design

    US20080040047A1