A method and device for determining an enrichment probe library

By obtaining microbial genome sequences, identifying conserved regions and generating enrichment probe libraries, the enrichment problem of multiple species or multiple nucleic acid sequences is solved, the detection efficiency is improved and the interference of the host genome is reduced.

CN114267412BActive Publication Date: 2025-10-03SUZHOU GENEWORKS TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111577776.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-22
Publication Date
2025-10-03
Estimated Expiration
2041-12-22

AI Technical Summary

Technical Problem

Existing technologies are unable to efficiently and simultaneously enrich multiple species or multiple nucleic acid sequences, resulting in low detection efficiency and susceptibility to interference from the host genome.

Method used

By obtaining the genome sequence of the microorganism to be enriched, determining the conserved region and using a sliding window to generate base combinations, filtering according to preset conditions, generating an enrichment probe library, and optimizing the number of base combinations to meet detection requirements.

Benefits of technology

It achieves the simultaneous enrichment of target substances of multiple species or multiple nucleic acid sequences, improves detection efficiency and reduces interference from the host genome.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114267412B_ABST
    Figure CN114267412B_ABST
Patent Text Reader

Abstract

The present application provides a method and device for determining an enrichment probe library, wherein the method includes: obtaining the genomic sequences corresponding to the microorganisms to be enriched; determining the gene fragments corresponding to the conserved regions in the genomic sequence of each microorganism to be enriched; for each gene fragment corresponding to the microorganism to be enriched, determining the base combination corresponding to the microorganism to be enriched through a sliding window; filtering all the base combinations of the microorganisms to be enriched according to preset conditions; determining the enrichment probe library according to the identification information corresponding to each microorganism to be enriched and the filtered base combinations. The present application generates an enrichment probe library by generating base combinations according to the genomic sequences of the microorganisms to be enriched, and filtering the base combinations corresponding to multiple microorganisms to be enriched, thereby solving the technical problem of designing an enrichment probe library for multiple species or multiple nucleic acid sequences, and producing the technical effect of enriching multiple species or multiple nucleic acid sequences and removing host gene interference.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to target enrichment in the field of pathogen detection and gene detection, and in particular to a method and device for determining an enrichment probe library. Background Art

[0002] High-throughput sequencing technology is also known as next-generation sequencing technology (NGS). Compared to traditional Sanger capillary electrophoresis sequencing, NGS is less expensive and provides more data, making it widely used in various scientific research. Although sequencing costs are continuously decreasing, first identifying and enriching target species before sequencing and analyzing them is more in line with current practical needs and cost-effectiveness. Targeted enrichment technology has become a common technique in NGS. Targeted enrichment utilizes hybridization technology, using probes to enrich the target fragments of interest from the overall nucleic acid.

[0003] Existing targeted enrichment probe designs are generally used to enrich a single species or a single nucleic acid sequence, and cannot meet the needs of enriching multiple species. When multiple species need to be enriched, multiple single-species enrichments can only be performed on the same sample, which is inefficient. Summary of the Invention

[0004] In view of this, the purpose of the present application is to provide at least a method and device for determining an enrichment probe library, by generating corresponding base combinations for the genome sequences of enriched microorganisms, and filtering the base combinations corresponding to multiple microorganisms to be enriched to generate an enrichment probe library, thereby solving the technical problem of simultaneously enriching multiple species or multiple nucleic acid sequences, and producing the technical effect of improving detection efficiency and removing host gene interference.

[0005] This application mainly includes the following aspects:

[0006] In a first aspect, an embodiment of the present application provides a method for determining an enrichment probe library, the method comprising: obtaining genomic sequences corresponding to the microorganisms to be enriched; determining the gene fragments corresponding to the conserved regions in the genomic sequence of each microorganism to be enriched; for each gene fragment corresponding to the microorganism to be enriched, determining the base combination corresponding to the microorganism to be enriched through a sliding window; filtering all base combinations of the microorganisms to be enriched according to preset conditions; and determining the enrichment probe library based on the identification information corresponding to each microorganism to be enriched and the filtered base combinations.

[0007] Optionally, for each gene fragment corresponding to the microorganism to be enriched, determining the base combination corresponding to the microorganism to be enriched through a sliding window includes: on each gene fragment corresponding to the microorganism to be enriched, sliding the window in sequence with a window spacing, and determining the base corresponding to each window as a base combination.

[0008] Optionally, filtering all base combinations of the microorganisms to be enriched according to preset conditions includes: judging whether the GC content value of each base combination of all the microorganisms to be enriched meets the preset GC content range value; if the GC content value of the base combination meets the preset GC content range value, judging whether the sequences corresponding to the preset initial fragment and the preset end fragment of each base combination are complementary; if the sequences corresponding to the preset initial fragment and the preset end fragment of the base combination are not complementary, determining the base combination of the preset initial fragment and the preset end fragment that are not complementary to each other as the target base combination, and judging whether all the target base combinations are complementary to each other; if all the target base combinations are not complementary to each other, determining all the target base combinations as filtered base combinations; if there are target base combinations that are complementary to each other among all the target base combinations, randomly selecting a target base combination from the complementary target base combinations, and determining it with the remaining target base combinations that are not complementary to each other as the filtered base combinations.

[0009] Optionally, after determining whether all target base combinations are complementary to each other, the method further includes: continuously comparing each filtered base sequence with the base sequence of the host genome to determine whether the comparison result is less than a first threshold; and determining all filtered base combinations that are less than the first threshold as re-filtered base combinations.

[0010] Optionally, after filtering all base combinations of the microorganisms to be enriched according to preset conditions, the method further includes: determining the lengths of the genome sequences corresponding to the microorganisms to be enriched; determining the quantity ratio of the filtered base combinations corresponding to each microorganism to be enriched based on the lengths; doubling the number of base combinations corresponding to each microorganism to be enriched so that the doubled number of base combinations corresponding to each microorganism to be enriched meets the quantity ratio; and determining the base combinations that meet the quantity ratio as the enrichment probe library.

[0011] Optionally, determining the gene fragment corresponding to the conserved region in the genome sequence of each microorganism to be enriched includes: obtaining multiple specific sequences in the genome sequence of the microorganism to be enriched, and generating sequence clusters for the multiple specific sequences; wherein the multiple specific sequences refer to the specific sequences of each microorganism to be enriched and the corresponding subspecies; from the sequence clusters, selecting the conserved region corresponding to the microorganism to be enriched; and determining the gene fragment corresponding to the conserved region according to the base pairing principle.

[0012] Optionally, generating a sequence cluster from multiple specific sequences includes: identifying the length of each specific sequence; determining the specific sequence with the longest length as the target specific sequence; comparing each specific sequence except the target specific sequence with the target specific sequence, and determining whether the comparison result is greater than or equal to a second threshold; if the comparison result is less than the second threshold, determining whether the number of all specific sequences less than the second threshold is greater than 1; if the number of all specific sequences less than the second threshold is greater than 1, determining all specific sequences less than the second threshold as new specific sequences; if the number of all specific sequences less than the second threshold is less than or equal to 1, determining the target specific sequence and the specific sequences less than the second threshold generated in sequence as a sequence cluster.

[0013] In a second aspect, an embodiment of the present application also provides a device for determining an enrichment probe library, which includes: an acquisition module for obtaining the genome sequences corresponding to the microorganisms to be enriched; a first determination module for determining the gene fragments corresponding to the conserved regions in the genome sequences of each microorganism to be enriched; a second determination module for determining, for each gene fragment corresponding to the microorganism to be enriched, the base combination corresponding to the microorganism to be enriched through a sliding window; a filtering module for filtering all base combinations of the microorganisms to be enriched according to preset conditions; and a third determination module for determining the enrichment probe library based on the identification information corresponding to each microorganism to be enriched and the filtered base combination.

[0014] In a third aspect, an embodiment of the present application further provides an electronic device comprising: a processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor and the memory communicate through the bus, and the machine-readable instructions are executed by the processor to execute the steps of the method for determining the enriched probe library in the above-mentioned first aspect or any possible implementation of the first aspect.

[0015] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of determining the enriched probe library in the above-mentioned first aspect or any possible embodiment of the first aspect are executed.

[0016] The embodiment of the present application provides a method and device for determining an enrichment probe library, wherein the method includes: obtaining the genome sequences corresponding to the microorganisms to be enriched; determining the gene fragments corresponding to the conserved regions in the genome sequence of each microorganism to be enriched; for each gene fragment corresponding to the microorganism to be enriched, determining the base combination corresponding to the microorganism to be enriched through a sliding window; filtering all the base combinations of the microorganisms to be enriched according to preset conditions; determining the enrichment probe library according to the identification information corresponding to each microorganism to be enriched and the filtered base combination. By generating corresponding base combinations according to the genome sequences of the microorganisms to be enriched, and filtering the base combinations corresponding to multiple microorganisms to be enriched, an enrichment probe library is generated, which solves the technical problem of being able to enrich multiple species or multiple nucleic acid sequences at the same time, and produces the technical effect of improving detection efficiency and removing host gene interference.

[0017] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.

[0019] Figure 1 A flow chart of a method for determining an enriched probe library provided in an embodiment of the present application is shown.

[0020] Figure 2 A flowchart showing the steps of determining the gene fragments corresponding to the conserved regions in the genome sequence of each microorganism to be enriched provided in the embodiments of the present application is shown.

[0021] Figure 3 A flowchart showing the steps of generating sequence clusters from multiple specific sequences provided in an embodiment of the present application is shown.

[0022] Figure 4 A flow chart showing the steps of filtering all base combinations of microorganisms to be enriched according to preset conditions provided in an embodiment of the present application is shown.

[0023] Figure 5 A flow chart of another method for determining an enriched probe library provided in an embodiment of the present application is shown.

[0024] Figure 6 A functional module diagram of a device for determining an enriched probe library provided in an embodiment of the present application is shown.

[0025] Figure 7 A schematic structural diagram of an electronic device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION

[0026] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. It should be understood that the drawings in the present application only serve the purpose of illustration and description and are not used to limit the scope of protection of the present application. In addition, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate the operations implemented according to some embodiments of the present application. It should be understood that the operations of the flowcharts can be implemented out of sequence, and steps without logical context can be reversed or implemented simultaneously. In addition, those skilled in the art, under the guidance of the contents of this application, can add one or more other operations to the flowchart, or remove one or more operations from the flowchart.

[0027] In addition, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application generally described and shown in the drawings here can be arranged and designed in various configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present application.

[0028] Existing technologies can only enrich a single species or a single nucleic acid sequence, and cannot enrich multiple species or multiple nucleic acid sequences at the same time.

[0029] Based on this, an embodiment of the present application provides a method and device for determining an enrichment probe library, wherein the method includes: obtaining the genome sequences corresponding to the microorganisms to be enriched; determining the gene fragments corresponding to the conserved regions in the genome sequence of each microorganism to be enriched; for each gene fragment corresponding to the microorganism to be enriched, determining the base combination corresponding to the microorganism to be enriched through a sliding window; filtering all the base combinations of the microorganisms to be enriched according to preset conditions; determining the enrichment probe library according to the identification information corresponding to each microorganism to be enriched and the filtered base combinations. By generating corresponding base combinations according to the genome sequences of the microorganisms to be enriched, and filtering the base combinations corresponding to multiple microorganisms to be enriched, an enrichment probe library is generated, which solves the technical problem of being able to enrich multiple species or multiple nucleic acid sequences at the same time, and produces the technical effect of improving detection efficiency and removing host gene interference, as follows:

[0030] See also Figure 1 , Figure 1 Flowchart of a method for determining an enriched probe library provided in an embodiment of the present application. Figure 1 As shown, the embodiment of the present application provides a method for determining an enrichment probe library, comprising the following steps:

[0031] S101. Obtain the genome sequences corresponding to the microorganisms to be enriched.

[0032] Specifically, the microorganisms to be enriched can be viruses, fungi, and bacteria. The genome sequence is the genome sequence of the microorganism to be enriched found in the NCBI biological taxonomy database. The genome sequence of the microorganism to be enriched can be DNA or RNA.

[0033] S102. Determine the gene fragment corresponding to the conserved region in the genome sequence of each microorganism to be enriched.

[0034] Please refer to Figure 2 , Figure 2 Flowchart of the steps for determining the gene fragments corresponding to the conserved regions in the genome sequence of each microorganism to be enriched provided in the embodiment of the present application. Figure 2 As shown, the method provided in the embodiment of the present application for determining the gene fragment corresponding to the conserved region in the genome sequence of each microorganism to be enriched includes the following steps:

[0035] S1021. Acquire multiple specific sequences in the genome sequence of the microorganism to be enriched, and generate sequence clusters for the multiple specific sequences.

[0036] The term "multiple specific sequences" refers to the specific sequences for each microorganism to be enriched and its corresponding subspecies. Specific sequences refer to the CDS coding sequence, UTR, 16S sequence, and 18S sequence corresponding to the microorganism to be enriched. Each microorganism to be enriched may have multiple subspecies, or the sequences obtained by different sequencers may differ slightly.

[0037] For example, there are three genome sequences of bacteria A and its subspecies (or variants of bacteria A), and 16S sequences are extracted from each genome sequence as the specific sequence of bacteria A. That is, there are three specific sequences of bacteria A.

[0038] If the virus does not have a 16S sequence or a 18S sequence, the entire genome sequence of the virus can also be selected as the specific sequence of the virus.

[0039] Please refer to Figure 3 , Figure 3 This is a flow chart of the steps of generating sequence clusters from multiple specific sequences provided in the embodiment of the present application. Figure 3As shown, the embodiment of the present application provides a method for generating sequence clusters from multiple specific sequences, including the following steps:

[0040] S10211. Identify the length of each specific sequence.

[0041] The length refers to the number of bases in a specific sequence. For example, the lengths of each specific sequence of bacteria A are 500 bp, 480 bp, and 620 bp.

[0042] S10212. Determine the longest specific sequence as the target specific sequence.

[0043] From the multiple specific sequences corresponding to the microorganisms to be enriched, the longest specific sequence is determined as the target specific sequence. If there are multiple specific sequences with the longest length, one is randomly selected as the target specific sequence.

[0044] S10213. Compare each specific sequence except the target specific sequence with the target specific sequence, and determine whether the comparison result is greater than or equal to a second threshold.

[0045] Specifically, each specific sequence other than the target specific sequence is compared with the target specific sequence, and the base differences between each specific sequence other than the target specific sequence and the target specific sequence are compared, wherein the second threshold is 97%.

[0046] For example, the length of the target specific sequence is 100 bp. If the length of the specific sequence is 99 bp and there are 2 bases different from the target specific sequence, the similarity between the specific sequence and the target specific sequence is 97%; if the length of the specific sequence is 100 bp and there are 3 bases different from the target specific sequence, the similarity between the specific sequence and the target specific sequence is 97%.

[0047] If the alignment result is greater than or equal to the second threshold, the specific sequence with the alignment result greater than or equal to the second threshold is deleted. If the alignment result of each specific sequence is greater than or equal to the second threshold, only the target specific sequence is used as the sequence cluster, or the second threshold is increased and the sequence cluster is re-determined.

[0048] S10214: Determine whether the number of all specific sequences smaller than the second threshold is greater than 1.

[0049] If the comparison result is less than the second threshold, it is determined whether the number of all specific sequences less than the second threshold is greater than 1.

[0050] S10215. Determine all specific sequences whose values ​​are smaller than a second threshold as new specific sequences.

[0051] Specifically, if the number of all specific sequences smaller than the second threshold is greater than 1 (i.e., the number of all specific sequences smaller than the second threshold is greater than or equal to 2), then all specific sequences smaller than the second threshold are determined as new specific sequences, and the process returns to step S10212, and the specific sequence with the longest length is determined as the target specific sequence.

[0052] S10216: Determine the sequentially generated target specific sequences and specific sequences smaller than a second threshold as sequence clusters.

[0053] If the number of all specific sequences smaller than the second threshold is less than or equal to 1, the target specific sequences and the specific sequences smaller than the second threshold generated in sequence are determined as a sequence cluster.

[0054] For example, if there are five specific sequences, namely specific sequence a, specific sequence b, specific sequence c, specific sequence d, and specific sequence e, specific sequence a is the target specific sequence, and specific sequence b, specific sequence c, specific sequence d, and specific sequence e are aligned with specific sequence a. If the alignment result between specific sequence b and specific sequence c is greater than or equal to a second threshold, specific sequence b and specific sequence c are deleted. If the alignment result between specific sequence d and specific sequence e is less than the second threshold, and specific sequence d has the longest length between specific sequence d and specific sequence e, specific sequence d is the target specific sequence, and specific sequence e is aligned with specific sequence d. If the alignment result between specific sequence e and specific sequence d is greater than or equal to the second threshold, specific sequence e is deleted. Therefore, specific sequence a and specific sequence d form a sequence cluster.

[0055] return Figure 2 , S1022. Select the conserved region corresponding to the microorganism to be enriched from the sequence cluster.

[0056] Specifically, conserved regions corresponding to the microorganisms to be enriched are selected from the sequence clusters, wherein the conserved regions refer to sequences that can only appear in a single species and will not appear in other species.

[0057] For example, if there are sequence a and sequence b in the sequence cluster, where the length of sequence a is 1000 bp and the length of sequence b is 990 bp, and if the conserved regions corresponding to the microorganisms to be enriched are 10 bp-500 bp and 600 bp-900 bp, then the 10 bp-500 bp and 600 bp-900 bp regions of sequence a and the 10 bp-500 bp and 600 bp-900 bp regions of sequence b are selected.

[0058] S1023. Determine the gene fragment corresponding to the conserved region based on the base pairing principle.

[0059] According to the base pairing principle, that is, purine bases pair exclusively with pyrimidine bases, and adenine (A) pairs exclusively with thymine (T) (in RNA molecules, adenine (A) pairs exclusively with uracil (U)); guanine (G) pairs exclusively with cytosine (C), the gene fragments corresponding to the conserved regions are determined.

[0060] return Figure 1 , S103, for each gene fragment corresponding to the microorganism to be enriched, determine the base combination corresponding to the microorganism to be enriched through a sliding window.

[0061] Specifically, on each gene fragment corresponding to the microorganism to be enriched, the window is sequentially slid at intervals of window spacing, and the base corresponding to each window is determined as a base combination.

[0062] Among them, the window spacing is generally 5bp, and the window size is generally 120bp.

[0063] S104. Filter all base combinations of the microorganisms to be enriched according to preset conditions.

[0064] Please refer to Figure 4 , Figure 4 The flowchart of the steps of filtering all base combinations of microorganisms to be enriched according to preset conditions provided in the embodiment of the present application. Figure 4 As shown, the embodiment of the present application provides filtering all base combinations of microorganisms to be enriched according to preset conditions, including the following steps:

[0065] S1041. Determine whether the GC content value of each base combination of all microorganisms to be enriched meets the preset GC content range value.

[0066] The GC content value refers to the ratio of the number of guanine and cytosine to the total number of bases in the genome sequence, and the GC content value generally ranges from 40% to 60%.

[0067] That is, the ratio of the sum of the number of guanines and cytosines in each base combination to the number of bases in the corresponding base combination is calculated to determine whether the ratio meets the preset GC content range.

[0068] S1042: Determine whether the sequences corresponding to the preset initial segment and the preset end segment of each base combination are complementary.

[0069] If the GC content of the base combination meets the preset GC content range, then a determination is made as to whether the sequences corresponding to the preset initial segment and the preset final segment of each base combination are complementary. In other words, a determination is made as to whether each base combination will generate a secondary structure. The preset initial segment and the preset final segment can be 10 bp.

[0070] For example, if the sequence of the base combination is ATTCGTCCACACCTATGATGCATT, the first base pairs with the last base, the second base pairs with the second-to-last base, and so on, the tenth base pairs with the tenth-to-last base, then this base combination will complement each other end to end to generate a secondary structure, and this base combination is deleted.

[0071] If the GC content value of the base combination does not meet the preset GC content range value, the base combination that does not meet the preset GC content range value will be deleted.

[0072] S1043: Determine a base combination that is not complementary to the preset initial segment and the preset end segment as a target base combination, and determine whether all target base combinations are complementary to each other.

[0073] If the sequences corresponding to the preset initial segment and the preset end segment of the base combination are not complementary, the base combination whose preset initial segment and the preset end segment are not complementary is determined as the target base combination, and it is determined whether all target base combinations are complementary to each other.

[0074] Exemplarily, if one target base combination is ATGCGCC and the other target base combination is TACGCGG, the two target base combinations are complementary to each other.

[0075] S1044. Determine all target base combinations as filtered base combinations.

[0076] If all target base combinations are not complementary to each other, all target base combinations are determined as filtered base combinations.

[0077] S1045. Randomly select a target base combination from the mutually complementary target base combinations, and determine it and the remaining target base combinations that are not mutually complementary as the filtered base combination.

[0078] If there are mutually complementary target base combinations among all target base combinations, one target base combination is randomly selected from the mutually complementary target base combinations and is determined as the filtered base combination together with the remaining target base combinations that are not mutually complementary.

[0079] For example, among the complementary target base combinations ATGCGCC and TACGCGG, a target base combination is randomly selected and determined as a filtered base combination together with the remaining target base combinations that are not mutually complementary, and the unselected target base combinations are deleted. That is, there are no mutually complementary base combinations in the filtered base combinations.

[0080] In an optional embodiment, if the enriched probe library generated is used in a detection environment, the base combination needs to be compared with the host genome. The specific method is as follows:

[0081] After determining whether all target base combinations are complementary to each other, the method further comprises:

[0082] Each filtered base sequence is compared with the base sequence of the host genome to determine whether the comparison result is less than a first threshold; all filtered base combinations that are less than the first threshold are determined as re-filtered base combinations.

[0083] The first threshold is generally 80%. The host genome is the genome of the host in which the microorganism to be enriched resides, that is, if a human is to be tested for influenza virus infection, the human is the host and the influenza virus is the microorganism to be enriched.

[0084] That is to say, the base sequence of each filtered base combination and the base sequence of the host genome are obtained, and it is determined whether the continuous complementary part of the base sequence of each filtered base combination and the base sequence of the host genome is less than 80% of the target base combination; if the continuous complementary part of the base sequence of the filtered base combination and the base sequence of the host genome is less than 80% of the filtered base combination, then this filtered base combination is determined as the re-filtered base combination.

[0085] Exemplarily, the length of the filtered base combination is 120bp, of which the base sequence of 1bp-96bp is complementary to the base sequence of the host genome, then this filtered base combination is deleted; if 1bp-90bp and 100bp-106bp are complementary to the host genome, since they are not continuously complementary, then this filtered base combination is determined as the base combination after re-filtering.

[0086] return Figure 1 S105: Determine the enrichment probe library based on the identification information corresponding to each microorganism to be enriched and the filtered base combination.

[0087] Specifically, the identification information corresponding to the microorganism to be enriched may be the name information of the microorganism to be enriched in the NCBI biological classification database.

[0088] That is, the enrichment probe library includes the name information of the microorganisms to be enriched and the corresponding filtered base combinations.

[0089] Specifically, you can also optimize the ratio of base combinations in the enrichment probe library, please refer to Figure 5 , Figure 5 Flowchart of another method for determining an enriched probe library provided in an embodiment of the present application. Figure 5 As shown, another method for determining an enrichment probe library provided in an embodiment of the present application includes the following steps:

[0090] After filtering all base combinations of the microorganisms to be enriched according to preset conditions, the method further includes:

[0091] S201. Determine the lengths of the genome sequences corresponding to the microorganisms to be enriched.

[0092] Specifically, the lengths of the genome sequences of multiple microorganisms to be enriched are determined.

[0093] S202. Determine the ratio of the number of filtered base combinations corresponding to each microorganism to be enriched based on the length.

[0094] For example, if the length of virus X is 1000 bp and the length of virus Y is 500 bp, the ratio of the number of base combinations corresponding to virus X to the number of base combinations corresponding to virus Y in the enriched probe library is 2:1.

[0095] S203. Double the number of base combinations corresponding to each microorganism to be enriched, so that the doubled number of base combinations corresponding to each microorganism to be enriched meets the quantity ratio.

[0096] For example, if the length of virus X is 1000 bp, the number of base combinations corresponding to virus X is 5, the length of virus Y is 500 bp, and the number of base combinations corresponding to virus Y is 2, then the number of base combinations corresponding to virus X in the enriched probe library can be 20 (i.e., the number of base combinations corresponding to virus X is quadrupled), and the number of base combinations corresponding to virus Y can be 10 (i.e., the number of base combinations corresponding to virus Y is quintupled).

[0097] Specifically, if the five base combinations corresponding to virus X are sequence a, sequence b, sequence c, sequence d, and sequence e; if the two base combinations corresponding to virus Y are sequence f and sequence g, then 4 sequences a, 4 sequences b, 4 sequences c, 4 sequences d, and 4 sequences e are placed in the enrichment probe library, and 5 sequences f and 5 sequences g are placed in the enrichment probe library.

[0098] S204. Determine the base combinations that meet the quantity ratio as the enriched probe library.

[0099] That is to say, from a probability perspective, when the length of the genome sequence of the microorganism to be enriched is longer, the more corresponding base combinations there are, the easier it is to capture.

[0100] Based on the same application concept, the embodiments of the present application also provide a device for determining an enriched probe library corresponding to the method for determining an enriched probe library provided in the above embodiments. Since the principle of solving the problem by the device in the embodiments of the present application is similar to the method for determining an enriched probe library in the above embodiments of the present application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be repeated.

[0101] like Figure 6 As shown, Figure 6 This is a functional module diagram of a device for determining an enriched probe library provided in an embodiment of the present application. The device 10 for determining an enriched probe library includes: an acquisition module 101, a first determination module 102, a second determination module 103, a filtering module 104, and a third determination module 105. The acquisition module 101 is used to obtain the genome sequences corresponding to the microorganisms to be enriched; the first determination module 102 is used to determine the gene fragments corresponding to the conserved regions in the genome sequence of each microorganism to be enriched; the second determination module 103 is used to determine the base combinations corresponding to the microorganism to be enriched through a sliding window for each gene fragment corresponding to the microorganism to be enriched; the filtering module 104 is used to filter all base combinations of the microorganisms to be enriched according to preset conditions; and the third determination module 105 is used to determine the enriched probe library based on the identification information corresponding to each microorganism to be enriched and the filtered base combinations.

[0102] The device 10 for determining an enrichment probe library further includes: a fourth determination module for determining the length of the genome sequence corresponding to each microorganism to be enriched; a fifth determination module for determining the number ratio of the filtered base combinations corresponding to each microorganism to be enriched based on the length; a doubling module for doubling the number of base combinations corresponding to each microorganism to be enriched so that the doubled number of base combinations corresponding to each microorganism to be enriched meets the number ratio; and a sixth determination module for determining the base combinations that meet the number ratio as the enrichment probe library.

[0103] Based on the same application concept, see Figure 7 As shown, it is a structural diagram of an electronic device provided in an embodiment of the present application. The electronic device 20 includes: a processor 201, a memory 202 and a bus 203. The memory 202 stores machine-readable instructions executable by the processor 201. When the electronic device 20 is running, the processor 201 and the memory 202 communicate with each other through the bus 203. The machine-readable instructions are executed by the processor 201 when running, as in the steps of the method for determining the enriched probe library provided in the above embodiment.

[0104] Based on the same application concept, the embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by the processor, the steps of the method for determining the enriched probe library provided in the above embodiment are executed. Specifically, the storage medium can be a general storage medium, such as a mobile disk, a hard disk, etc. When the computer program on the storage medium is executed, the above-mentioned method for determining the enriched probe library can be executed, by generating corresponding base combinations for the genome sequences of the enriched microorganisms, and filtering the base combinations corresponding to the multiple microorganisms to be enriched, to generate an enriched probe library, thereby solving the technical problem of being able to enrich multiple species or multiple nucleic acid sequences at the same time, and producing the technical effect of improving detection efficiency and removing host gene interference.

[0105] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the system and device described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here. In the several embodiments provided in this application, it should be understood that the disclosed system, device and method can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0106] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0107] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0108] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a non-volatile computer-readable storage medium that is executable by a processor. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program code.

[0109] The above are only specific embodiments of the present application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A method for determining an enriched probe library, characterized in that: The method comprises: Obtaining the genome sequences corresponding to the microorganisms to be enriched; Determine the gene fragment corresponding to the conserved region in the genome sequence of each microorganism to be enriched, wherein the conserved region refers to a sequence that can only appear in a single species and does not appear in other species; For each gene fragment corresponding to the microorganism to be enriched, the base combination corresponding to the microorganism to be enriched is determined through a sliding window; Filter all base combinations of microorganisms to be enriched according to preset conditions; Determine an enrichment probe library based on the identification information corresponding to each microorganism to be enriched and the filtered base combination; The filtering of all base combinations of microorganisms to be enriched according to preset conditions includes: Determining whether the GC content value of each base combination of all the microorganisms to be enriched meets the preset GC content range value; If the GC content value of the base combination meets the preset GC content range value, determining whether the sequences corresponding to the preset initial segment and the preset end segment of each base combination are complementary; If the sequences corresponding to the preset initial segment and the preset end segment of the base combination are not complementary, the base combination for which the preset initial segment and the preset end segment are not complementary is determined as a target base combination, and it is determined whether all the target base combinations are complementary to each other; If all the target base combinations are not complementary to each other, all the target base combinations are determined as filtered base combinations; If there are mutually complementary target base combinations among all target base combinations, one target base combination is randomly selected from the mutually complementary target base combinations and is determined as the filtered base combination together with the remaining target base combinations that are not mutually complementary.

2. The method according to claim 1, characterized in that For each gene fragment corresponding to the microorganism to be enriched, determining the base combination corresponding to the microorganism to be enriched through the sliding window includes: On each gene fragment corresponding to the microorganism to be enriched, the window is sequentially slid at intervals of window spacing, and the base corresponding to each window is determined as a base combination.

3. The method according to claim 1, characterized in that After determining whether all the target base combinations are complementary to each other, the method further includes: Comparing the base sequence of each filtered base combination with the base sequence of the host genome, and determining whether the comparison result is less than a first threshold; All filtered base combinations that are smaller than the first threshold are determined as re-filtered base combinations.

4. The method according to claim 1, wherein After filtering all base combinations of the microorganisms to be enriched according to preset conditions, the method further comprises: Determine the length of the genome sequence corresponding to each of the microorganisms to be enriched; Determining the ratio of the number of filtered base combinations corresponding to each microorganism to be enriched based on the length; doubling the number of base combinations corresponding to each of the microorganisms to be enriched, so that the doubled number of the base combinations corresponding to each of the microorganisms to be enriched satisfies the quantity ratio; The base combinations satisfying the quantity ratio are determined as the enriched probe library.

5. The method according to claim 1, wherein The step of determining the gene fragment corresponding to the conserved region in the genome sequence of each microorganism to be enriched comprises: Obtaining multiple specific sequences from the genome sequence of the microorganism to be enriched, generating sequence clusters from the multiple specific sequences; wherein the multiple specific sequences refer to specific sequences of each microorganism to be enriched and its corresponding subspecies; selecting conserved regions corresponding to the microorganism to be enriched from the sequence clusters; According to the base pairing principle, the gene fragment corresponding to the conserved region is determined.

6. The method according to claim 5, characterized in that Generating a sequence cluster from the plurality of specific sequences comprises: identifying the length of each of the specific sequences; Determine the specific sequence with the longest length as the target specific sequence; Comparing each of the specific sequences except the target specific sequence with the target specific sequence, and determining whether the comparison result is greater than or equal to a second threshold; If the comparison result is less than a second threshold, determining whether the number of all specific sequences less than the second threshold is greater than 1; If the number of all specific sequences smaller than the second threshold is greater than 1, all specific sequences smaller than the second threshold are determined as new specific sequences; If the number of all the specific sequences smaller than the second threshold is less than or equal to 1, the target specific sequences generated in sequence and the specific sequences smaller than the second threshold are determined as a sequence cluster.

7. A device for determining an enriched probe library, characterized in that: The device comprises: An acquisition module is used to obtain the genome sequences corresponding to the microorganisms to be enriched; The first determination module is used to determine the gene fragment corresponding to the conserved region in the genome sequence of each microorganism to be enriched. The conserved region refers to a sequence that can only appear in a single species and does not appear in other species; The second determination module is used to determine the base combination corresponding to the microorganism to be enriched through a sliding window for each gene fragment corresponding to the microorganism to be enriched; A filtering module is used to filter all base combinations of microorganisms to be enriched according to preset conditions; A third determination module is used to determine the enrichment probe library based on the identification information corresponding to each microorganism to be enriched and the filtered base combination; The filtering module is further used for: Determining whether the GC content value of each base combination of all the microorganisms to be enriched meets the preset GC content range value; If the GC content value of the base combination meets the preset GC content range value, determining whether the sequences corresponding to the preset initial segment and the preset end segment of each base combination are complementary; If the sequences corresponding to the preset initial segment and the preset end segment of the base combination are not complementary, the base combination for which the preset initial segment and the preset end segment are not complementary is determined as a target base combination, and it is determined whether all the target base combinations are complementary to each other; If all the target base combinations are not complementary to each other, all the target base combinations are determined as filtered base combinations; If there are mutually complementary target base combinations among all target base combinations, one target base combination is randomly selected from the mutually complementary target base combinations and is determined as the filtered base combination together with the remaining target base combinations that are not mutually complementary.

8. An electronic device, characterized in that: include: A processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor and the memory communicate through the bus, and the machine-readable instructions are executed by the processor to perform the steps of the method for determining the enriched probe library as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method for determining an enriched probe library according to any one of claims 1 to 6 are executed.

Citation Information

Patent Citations

  • Clustering method and device of gene expression data, computer equipment and storage medium

    CN110827924A

  • Human respiratory virus targeted enrichment capture probe set and application thereof

    CN112342270A