Genome data analysis method and device, computer equipment and storage medium
By identifying the bacterial species and comparing the similarity of the whole genome data to be tested, the time-consuming problem in the existing technology is solved, and the genomic data analysis for efficiently distinguishing Mycobacterium tuberculosis disease from non-tuberculosis mycobacterium disease is achieved, reducing costs and improving analysis efficiency.
Patent Information
- Application Number
- CN202410322573.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-20
- Publication Date
- 2025-09-23
AI Technical Summary
Existing technologies take a long time to analyze genomic data of tuberculosis and non-tuberculosis mycobacteria, making it difficult to effectively distinguish between the two.
The whole genome data of the target bacteria is identified to determine its type, and the similarity is compared with the standard strain data of the target bacteria type. Data analysis is performed based on the similarity, including variation detection, genotype resistance information, genetic evolution tree and other analyses.
It improves data analysis efficiency, reduces analysis costs, and can accurately distinguish between tuberculosis and non-tuberculosis mycobacteria to determine whether further analysis is needed.
Smart Images

Figure CN120690280A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of gene detection technology, and in particular to a genome data analysis method, apparatus, computer equipment, and storage medium. Background Art
[0002] Since the clinical manifestations of non-tuberculous mycobacterial disease and tuberculosis mycobacterial disease are relatively similar, in order to effectively diagnose and analyze mycobacterial disease, existing technologies can perform data analysis on the whole genome data to be tested through laboratory bacterial culture, smear microscopy, isolation and culture, molecular diagnosis and other methods, and realize molecular-level analysis of mycobacterial disease.
[0003] However, the above method for data analysis is time-consuming. Summary of the Invention
[0004] Based on this, it is necessary to provide a genome data analysis method, device, computer equipment and storage medium to address the above technical problems.
[0005] In a first aspect, the present application provides a method for analyzing genomic data. The method comprises:
[0006] Perform bacterial species identification on the whole genome data to be tested to determine the bacterial species type corresponding to the whole genome data to be tested;
[0007] If the bacterial species type is the target Bacillus type, then determine the data similarity between the whole genome data to be tested and the standard strain data corresponding to the target Bacillus type;
[0008] According to the data similarity, data analysis is performed on the whole genome data to be tested to obtain the data analysis results of the whole genome data to be tested.
[0009] In one embodiment, data analysis is performed on the whole genome data to be detected based on data similarity to obtain data analysis results of the whole genome data to be detected, including:
[0010] If the data similarity is less than the similarity threshold, the data similarity is used as the data analysis result of the whole genome data to be tested;
[0011] If the data similarity is greater than or equal to the similarity threshold, data analysis is performed on the whole genome data to be detected to obtain a data analysis result of the whole genome data to be detected.
[0012] In one embodiment, performing data analysis on the whole genome data to be detected to obtain data analysis results of the whole genome data to be detected includes:
[0013] Obtain the data analysis requirements corresponding to the whole genome data to be tested;
[0014] Perform mutation detection on the whole genome data to be tested to obtain gene sequence variation information of the whole genome data to be tested;
[0015] According to data analysis requirements, data analysis is performed on the gene sequence variation information to obtain data analysis results corresponding to the data analysis requirements.
[0016] In one embodiment, if the data analysis requirement includes: determining the genotype resistance information of the whole genome data to be tested; performing data analysis on the gene sequence variation information according to the data analysis requirement, and obtaining the data analysis results corresponding to the data analysis requirement, including:
[0017] The genotype resistance test is performed on the gene sequence variation information to obtain the genotype resistance information of the whole genome data to be tested.
[0018] In one embodiment, if the data analysis requirement includes: determining a target genome sequence of the whole genome data to be tested; performing data analysis on the gene sequence variation information according to the data analysis requirement, and obtaining a data analysis result corresponding to the data analysis requirement, the data analysis requirement includes:
[0019] The gene sequence variation information is assembled with parameters to obtain the target genome sequence of the whole genome data to be tested.
[0020] In one embodiment, if the data analysis requirement further includes: determining a genetic evolution tree of the whole genome data to be tested; performing data analysis on the gene sequence variation information according to the data analysis requirement, and obtaining a data analysis result corresponding to the data analysis requirement, the data analysis result includes:
[0021] From the candidate genome sequences, a reference genome sequence whose consistency with the target genome sequence is greater than a consistency threshold is selected;
[0022] Determine the genetic distance between the target genome sequence and the reference genome sequence;
[0023] Based on the genetic distance and reference genome sequence, a genetic evolutionary tree of the whole genome data to be tested is constructed.
[0024] In one embodiment, the target bacillus type is the Mycobacterium tuberculosis complex.
[0025] In a second aspect, the present application also provides a genomic data analysis device. The device comprises:
[0026] The first determination module is used to perform bacterial species identification on the whole genome data to be detected and determine the bacterial species type corresponding to the whole genome data to be detected;
[0027] The second determination module is configured to determine the data similarity between the whole genome data to be detected and the standard strain data corresponding to the target Bacillus type if the bacterial species type is the target Bacillus type;
[0028] The analysis module is used to perform data analysis on the whole genome data to be detected based on data similarity to obtain data analysis results of the whole genome data to be detected.
[0029] In a third aspect, the present application further provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, the following steps are performed:
[0030] Perform bacterial species identification on the whole genome data to be tested to determine the bacterial species type corresponding to the whole genome data to be tested;
[0031] If the bacterial species type is the target Bacillus type, then determine the data similarity between the whole genome data to be tested and the standard strain data corresponding to the target Bacillus type;
[0032] According to the data similarity, data analysis is performed on the whole genome data to be tested to obtain the data analysis results of the whole genome data to be tested.
[0033] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the following steps:
[0034] Perform bacterial species identification on the whole genome data to be tested to determine the bacterial species type corresponding to the whole genome data to be tested;
[0035] If the bacterial species type is the target Bacillus type, then determine the data similarity between the whole genome data to be tested and the standard strain data corresponding to the target Bacillus type;
[0036] According to the data similarity, data analysis is performed on the whole genome data to be tested to obtain the data analysis results of the whole genome data to be tested.
[0037] The above-mentioned genome data analysis method, device, computer equipment and storage medium, by determining the bacterial species type corresponding to the whole genome data to be detected, and, when the bacterial species type is the target bacillus type, determining the data similarity between the whole genome data to be detected and the standard strain data corresponding to the target bacillus type, and then, according to the data similarity, determining the data analysis result of the whole genome data to be detected. Due to the above process, before the whole genome data to be detected is analyzed, the present application will first perform bacterial species identification on the whole genome data to be detected, so as to distinguish non-tuberculosis mycobacteria from tuberculosis mycobacteria; and, the present application will determine the data analysis result of the whole genome data to be detected based on the data similarity between the whole genome data to be detected and the standard strain data corresponding to the target bacillus type. Compared with the prior art, the present application can not only realize bacterial species identification for the whole genome data to be detected, but also realize the purpose of performing target bacillus type data analysis on the whole genome data to be detected, and the present application effectively improves the efficiency of data analysis on the whole genome data to be detected, reduces the cost of analyzing the whole genome data to be detected, so that the staff can judge whether further data analysis is needed for the whole genome data to be detected based on the data analysis results. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 A diagram illustrating an application environment of a genomic data analysis method provided in an embodiment of the present application;
[0039] Figure 2 A flowchart of a genomic data analysis method provided in an embodiment of the present application;
[0040] Figure 3 A schematic diagram of a process for determining the data analysis results of the whole genome data to be detected provided in an embodiment of the present application;
[0041] Figure 4 A schematic diagram of a process for determining a genetic evolutionary tree according to an embodiment of the present application;
[0042] Figure 5 A flowchart of another genomic data analysis method provided in an embodiment of the present application;
[0043] Figure 6 A structural block diagram of the first genome data analysis device provided in an embodiment of the present application;
[0044] Figure 7 A structural block diagram of a second genomic data analysis device provided in an embodiment of the present application;
[0045] Figure 8 A structural block diagram of a third genomic data analysis device provided in an embodiment of the present application;
[0046] Figure 9 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0047] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0048] It should be understood that the specific embodiments described herein are merely used to explain the present application and are not intended to limit the present application. In the description of the present application, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms are not necessarily directed to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, unless they contradict each other.
[0049] Based on the above situation, the genome data analysis method provided in the embodiment of the present application can be applied to Figure 1 In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as shown in FIG. Figure 1 As shown. The computer device includes a processor, a memory, and a network interface connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data acquired by the genomic data analysis method. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a genomic data analysis method is implemented.
[0050] The present application discloses a genome data analysis method, apparatus, computer equipment and storage medium, which may specifically include the following contents: by determining the bacterial species type corresponding to the whole genome data to be detected, and when the bacterial species type is a target Bacillus type, determining the data similarity between the whole genome data to be detected and the standard strain data corresponding to the target Bacillus type, and then determining the data analysis results of the whole genome data to be detected based on the data similarity.
[0051] In one embodiment, Figure 2 As shown, Figure 2 A flow chart of a genome data analysis method provided in an embodiment of the present application provides a genome data analysis method. Figure 1 The genomic data analysis method performed by the computer device in the method may include the following steps:
[0052] Step 201 : performing bacterial species identification on the whole genome data to be detected to determine the bacterial species type corresponding to the whole genome data to be detected.
[0053] It should be noted that before performing strain identification on the whole genome data to be tested, in order to ensure the accuracy of strain identification, data cleaning and host removal operations are required to ensure that the strain type corresponding to the whole genome data to be tested can be accurately obtained.
[0054] Among them, the host removal operation refers to removing the genome sequence of the human host contained in the whole genome data to be tested.
[0055] It is further explained that when data cleaning is required for the whole genome data to be tested, whole genome data cleaning software can be used to perform data cleaning and host removal operations on the whole genome data to be tested, so as to achieve the purpose of accurately obtaining the bacterial species type corresponding to the whole genome data to be tested.
[0056] In one embodiment of the present application, SOAPnuke (data quality control filtering) software can be used to clean up the whole genome data to be tested, so as to perform cleaning operations on the low-quality and joint-contaminated whole genome data to be tested, thereby achieving the purpose of accurately obtaining the bacterial species type corresponding to the whole genome data to be tested.
[0057] In another embodiment of the present application, the whole genome data to be tested can be subjected to a host removal operation using bowtie2 (host removal operation) software to remove the genome sequence of the human host contained in the whole genome data to be tested, thereby achieving the purpose of accurately obtaining the bacterial species type corresponding to the whole genome data to be tested.
[0058] It is further explained that when it is necessary to perform bacterial species identification on the whole genome data to be tested, the comparison file (i.e., *.bam format file) obtained by comparing the whole genome data to be tested with the whole genome database of the mlstverse (bacterial species identification software) software can be determined by bwa mem (software for short read sequence alignment) software, and then, the comparison file can be sorted and indexed by samtools (tool software for processing bam format) software to obtain a processed comparison file; an R program written based on the mlstverse software and its data package is used to perform bacterial species identification on the whole genome data to be tested according to the processed comparison file, and the content of NTM (nontuberculous mycobacteria) species or MTBC (Mycobacterium tuberculosis complex) species is determined; and the bacterial species type corresponding to the whole genome data to be tested is determined according to the content of NTM species or MTBC species.
[0059] For example, after the whole genome data to be detected is identified using the above process, the bacterial species type corresponding to the whole genome data to be detected is obtained, and the bacterial species type can be shown in Table (1):
[0060] Table (1) Bacteria types
[0061]
[0062] Step 202: If the bacterial species type is the target Bacillus type, determine the data similarity between the whole genome data to be detected and the standard strain data corresponding to the target Bacillus type.
[0063] Among them, the target bacillus type is the Mycobacterium tuberculosis complex type.
[0064] It should be noted that in order to ensure the accuracy of subsequent data analysis of the whole genome data to be tested, the data similarity between the whole genome data to be tested and the standard strain data corresponding to the target Bacillus type can be used as the data basis for data analysis of the whole genome data to be tested. Therefore, when performing data analysis on the whole genome data to be tested, it is necessary to determine the data similarity between the whole genome data to be tested and the standard strain data corresponding to the target Bacillus type.
[0065] In one embodiment of the present application, when it is necessary to determine the data similarity between the whole genome data to be detected and the standard strain data corresponding to the target Bacillus type, the following contents may be specifically included: using BWA (Burrows-WheelerAligner, comparison software) software or minimap2 (sequence alignment software) software, the whole genome data to be detected and the standard strain data corresponding to the target Bacillus type are compared, and the output result of the BWA software or minimap2 software is obtained, and the output result is the data similarity between the whole genome data to be detected and the standard strain data corresponding to the target Bacillus type.
[0066] Step 203 : performing data analysis on the whole genome data to be detected based on the data similarity to obtain a data analysis result of the whole genome data to be detected.
[0067] It should be noted that in order to ensure that the data analysis results of the whole genome data to be tested can be obtained smoothly, it is necessary to ensure that the data similarity is greater than or equal to the similarity threshold; if the data similarity is greater than or equal to the similarity threshold, data analysis is performed on the whole genome data to be tested to obtain the data analysis results of the whole genome data to be tested.
[0068] Among them, if the data similarity is greater than or equal to the similarity threshold, it means that the proportion of the whole genome data of the target bacillus in the whole genome data to be detected exceeds the threshold. At this time, the whole genome data to be detected can be determined to be pure target bacillus whole genome data.
[0069] The above-mentioned genome data analysis method determines the bacterial species type corresponding to the whole genome data to be detected, and, when the bacterial species type is the target bacillus type, determines the data similarity between the whole genome data to be detected and the standard strain data corresponding to the target bacillus type, and then, according to the data similarity, determines the data analysis result of the whole genome data to be detected. Due to the above process, before the whole genome data to be detected is analyzed, the present application will first perform bacterial species identification on the whole genome data to be detected, so as to distinguish non-tuberculosis mycobacteria from tuberculosis mycobacteria; and, according to the data similarity between the whole genome data to be detected and the standard strain data corresponding to the target bacillus type, the present application determines the data analysis result of the whole genome data to be detected. Compared with the prior art, the present application can not only realize bacterial species identification for the whole genome data to be detected, but also realize the purpose of performing target bacillus type data analysis on the whole genome data to be detected, and the present application effectively improves the efficiency of data analysis on the whole genome data to be detected, reduces the cost of analyzing the whole genome data to be detected, so that the staff can judge whether further data analysis is needed for the whole genome data to be detected based on the data analysis results.
[0070] In one embodiment, the whole genome data to be detected is analyzed by laboratory bacterial culture, smear microscopy, isolation culture, molecular diagnosis and other methods to realize the process of genetic data analysis of Mycobacterium tuberculosis. However, the above-mentioned method of data analysis of whole genome data is time-consuming. In order to solve the above technical problems, the computer device of the present application can be used as follows: Figure 3 In the manner shown, data analysis is performed on the whole genome data to be tested according to data similarity to obtain the data analysis results of the whole genome data to be tested. The specific proof process is as follows:
[0071] Step 301: If the data similarity is less than the similarity threshold, the data similarity is used as the data analysis result of the whole genome data to be detected.
[0072] Among them, if the data similarity is less than the similarity threshold, it means that the proportion of the whole genome data of the target bacillus in the whole genome data to be detected is less than the threshold. At this time, it can be determined that the whole genome data to be detected is not pure target bacillus whole genome data.
[0073] It should be noted that if the data similarity is less than the similarity threshold, it is determined that the whole genome data to be tested is not the pure target bacillus whole genome data. Therefore, there is no need to perform data analysis on the whole genome data to be tested, and the data similarity can be used as the data analysis result of the whole genome data to be tested.
[0074] Step 302: If the data similarity is greater than or equal to the similarity threshold, data analysis is performed on the whole genome data to be detected to obtain a data analysis result of the whole genome data to be detected.
[0075] It should be noted that if the data similarity is greater than or equal to the similarity threshold, it means that the proportion of the target bacillus whole genome data in the whole genome data to be tested exceeds the threshold. At this time, the whole genome data to be tested can be determined to be pure target bacillus whole genome data. Therefore, the data analysis requirements corresponding to the whole genome data to be tested can be obtained, and based on the data analysis requirements corresponding to the whole genome data to be tested, the whole genome data to be tested can be analyzed to obtain the data analysis results of the whole genome data to be tested.
[0076] Data analysis requirements may include, but are not limited to, determining genotypic drug resistance information for the whole genome data to be tested, determining the target genome sequence for the whole genome data to be tested, determining the genetic evolutionary tree for the whole genome data to be tested, determining the lineage typing for the whole genome data to be tested, etc. In summary, data analysis requirements can encompass a wide range of content, and are not limited here.
[0077] In one embodiment of the present application, if the data analysis requirement is: pedigree typing of the whole genome data to be detected, therefore, when it is necessary to determine the data analysis results corresponding to the data analysis requirement, the following may be specifically included: using bwamem software, determining the comparison file (i.e., *.bam format file) obtained by comparing the whole genome data to be detected with the whole genome database of mlstverse software, using samtools software, extracting the DR region sequence of the whole genome data to be detected in the comparison file, and generating a fa format file; using blastn (Basic Local Alignment Search Tool, biological macromolecule sequence alignment search tool) software, filtering the DR region sequence to exclude mismatches or low-quality alignment results; counting the depth of each DR region sequence, that is, the number of sequences that match the DR region sequence. Scoring the DR region sequence according to the depth of the DR region sequence, obtaining a binary bit corresponding to the DR region sequence, the depth or score determines whether this binary bit is 0 or 1, and then obtaining a 43-bit binary code corresponding to the DR region sequence. The 43-bit binary code is converted into octal code, and the octal code is compared with the SITVIT2 database to find the best matching typing result. This typing result is the pedigree typing of the whole genome data to be tested.
[0078] Among them, the SITVIT2 database is an internationally widely used Mycobacterium tuberculosis typing database.
[0079] It is further explained that, according to the data analysis requirements corresponding to the whole genome data to be tested, when performing data analysis on the whole genome data to be tested, the following contents may be specifically included: obtaining the data analysis requirements corresponding to the whole genome data to be tested; performing variation detection on the whole genome data to be tested to obtain gene sequence variation information of the whole genome data to be tested; performing data analysis on the gene sequence variation information according to the data analysis requirements to obtain data analysis results corresponding to the data analysis requirements.
[0080] In one embodiment of the present application, freebayes (variation detection software) software can be used to perform variation detection on the whole genome data to be detected, the output results of the freebayes software can be obtained, the output results can be filtered and quality controlled, and the gene sequence variation information of the whole genome data to be detected can be obtained.
[0081] To further illustrate, if the data analysis requirements include: determining the genotype resistance information of the whole genome data to be tested; therefore, according to the data analysis requirements, when performing data analysis on the gene sequence variation information, the following may be specifically included: performing genotype resistance detection on the gene sequence variation information to obtain the genotype resistance information of the whole genome data to be tested.
[0082] In one embodiment of the present application, when it is necessary to perform genotypic drug resistance detection on gene sequence variation information, the drug resistance of the gene sequence variation information can be annotated or predicted based on the drug resistance-related literature to realize genotypic drug resistance detection on the gene sequence variation information and obtain the genotypic drug resistance information of the whole genome data to be tested.
[0083] Among them, genotype resistance information is used to characterize the drug resistance of each mutation site in the gene sequence variation information.
[0084] To further illustrate, if the data analysis requirements include: determining the target genome sequence of the whole genome data to be tested; therefore, according to the data analysis requirements, when performing data analysis on the gene sequence variation information, the following may be specifically included: performing parameter assembly processing on the gene sequence variation information to obtain the target genome sequence of the whole genome data to be tested.
[0085] In one embodiment of the present application, the gene sequence variation information can be subjected to parameter assembly processing by the bcftools consensus software to obtain the output result of the bcftools consensus (a tool for variable calling and operating VCF) software, which is the target genome sequence of the whole genome data to be detected.
[0086] The above-mentioned genomic data analysis method, by determining the gene sequence variation information of the whole genome data to be tested, ensures that the gene sequence variation information can be subsequently analyzed according to the data analysis requirements and obtains the data analysis results corresponding to the data analysis requirements. This ensures that the data analysis results corresponding to the data analysis requirements can be successfully determined.
[0087] In one embodiment, if Figure 4 As shown in Figure 2, core genome multi-locus sequence typing is a gene-by-gene method that uses a large number of candidate genome sequences to compare the target genome sequence. This method has the advantages of high resolution and good repeatability. After typing, this method draws a genetic evolutionary tree to analyze the genetic distance and phylogenetic relationship between strains to achieve bacterial tracing. Therefore, if the data analysis requirements also include: determining the genetic evolutionary tree of the whole genome data to be tested, data analysis can be performed based on core genome multi-locus sequence typing, which may include the following:
[0088] Step 401: Select a reference genome sequence from the candidate genome sequences, the reference genome sequence having a consistency greater than a consistency threshold with the target genome sequence.
[0089] Among them, the candidate genome sequence refers to the core genome multi-site sequence corresponding to the target Bacillus type; therefore, the reference genome sequence is the genome sequence in the core genome multi-site sequence whose consistency with the target genome sequence is greater than the consistency threshold.
[0090] It should be noted that when it is necessary to determine a reference genome sequence whose consistency with the target genome sequence is greater than the consistency threshold, the following may be specifically included: gene annotation of the target genome sequence is performed using prokka (rapid annotation of prokaryotic genomes) software, and blast (Basic Local Alignment Search Tool, biological macromolecule sequence alignment search tool) software is used to select a reference genome sequence whose consistency with the target genome sequence after gene annotation is greater than the consistency threshold from the candidate genome sequences.
[0091] Step 402: Determine the genetic distance between the target genome sequence and the reference genome sequence.
[0092] In one embodiment of the present application, when it is necessary to determine the genetic distance between the target genome sequence and the reference genome sequence, the following may be specifically included: inputting the target genome sequence and the reference genome sequence into muscle (multiple sequence alignment tool) software, and obtaining the output result of the muscle software, which is the genetic distance between the target genome sequence and the reference genome sequence.
[0093] Step 403: construct a genetic evolution tree of the whole genome data to be tested based on the genetic distance and the reference genome sequence.
[0094] In one embodiment of the present application, when it is necessary to construct a genetic evolution tree of the whole genome data to be tested, the following contents may be included: using the reference genome sequence as the node of the genetic evolution tree, and using the distance between each node to represent the genetic distance between the target genome sequence and the reference genome sequence, thereby obtaining a genetic evolution tree of the whole genome data to be tested.
[0095] In another embodiment of the present application, when it is necessary to construct a genetic evolution tree of the whole genome data to be tested, the following contents may also be included: the reference genome sequence is used as the node of the genetic evolution tree, and the reference genome sequences are connected according to the genetic relationship between each other, and the genetic relationship between the reference genome sequence and the target genome sequence, and the genetic distance between the target genome sequence and the reference genome sequence is added as supplementary information between the corresponding nodes.
[0096] The above-mentioned genomic data analysis method constructs a genetic evolutionary tree of the whole genome data to be tested by determining the reference genome sequence and the genetic distance between the target genome sequence and the reference genome sequence. This allows for data analysis of gene sequence variation information while meeting data analysis requirements, resulting in data analysis results that meet those requirements.
[0097] In one embodiment, if Figure 5 As shown, data analysis of the whole genome data to be tested may include the following:
[0098] Step 501 : performing bacterial species identification on the whole genome data to be detected to determine the bacterial species type corresponding to the whole genome data to be detected.
[0099] Step 502: If the bacterial species type is the target Bacillus type, determine the data similarity between the whole genome data to be detected and the standard strain data corresponding to the target Bacillus type.
[0100] Step 503: If the data similarity is greater than or equal to the similarity threshold, the data analysis requirements corresponding to the whole genome data to be detected are obtained.
[0101] Step 504 : performing variation detection on the whole genome data to be detected to obtain gene sequence variation information of the whole genome data to be detected.
[0102] Step 505, if the data analysis requirements further include: determining a genetic evolutionary tree of the whole genome data to be detected; performing parameter assembly processing on the gene sequence variation information to obtain a target genome sequence of the whole genome data to be detected.
[0103] Step 506: Select a reference genome sequence from the candidate genome sequences whose consistency with the target genome sequence is greater than a consistency threshold.
[0104] Step 507: Determine the genetic distance between the target genome sequence and the reference genome sequence.
[0105] Step 508: construct a genetic evolution tree of the whole genome data to be tested based on the genetic distance and the reference genome sequence.
[0106] In one embodiment of the present application, if there are two whole genome data to be detected, the two whole genome data to be detected are: a first genome data numbered R13_IR_13, and a second genome data numbered R8_IR_8; then, the bwamem software is used to determine the comparison file obtained by comparing the two whole genome data to be detected with the whole genome database of the mlstverse software, and then, the comparison file is sorted and indexed by the samtools software to obtain a processed comparison file; an R program written based on the mlstverse software and its data package is used to determine the bacterial species types corresponding to the two whole genome data to be detected. The bacterial species types corresponding to the two whole genome data to be detected are shown in Table (2):
[0107] Table (2) Bacterial species corresponding to the two whole genome data to be tested
[0108]
[0109] According to Table (2), the bacterial species type of the two whole genome data to be detected is determined to be the target Bacillus type, and the data similarity between the whole genome data to be detected and the standard strain data corresponding to the target Bacillus type is determined. The data similarity of the two whole genome data to be detected is shown in Table (3):
[0110] Table (3): Data similarity of two whole genome data to be tested
[0111]
[0112] Perform variation detection on the two whole genome data to be detected to obtain gene sequence variation information of the two whole genome data to be detected. The gene sequence variation information of the two whole genome data to be detected can be shown in Table (4);
[0113] Table (4): Gene sequence variation information of two whole genome data to be tested
[0114] Sample name Total number of mutations R13_IR_13 1073 R8_IR_8 1042
[0115] The genotype resistance test is performed on the two gene sequence variation information to obtain the genotype resistance information of the two whole genome data to be tested. The genotype resistance information of the two whole genome data to be tested can be shown in Table (5);
[0116] Table (5): Genotypic drug resistance information of two whole genome data to be tested
[0117]
[0118] Perform pedigree typing analysis on the two whole genome data to be tested to obtain the pedigree typing of the two whole genome data to be tested. The pedigree typing of the two whole genome data to be tested can be shown in Table (6);
[0119] Table (6): Lineage typing of two whole genome data to be tested
[0120]
[0121] The above-mentioned genome data analysis method determines the bacterial species type corresponding to the whole genome data to be detected, and, when the bacterial species type is the target bacillus type, determines the data similarity between the whole genome data to be detected and the standard strain data corresponding to the target bacillus type, and then, according to the data similarity, determines the data analysis result of the whole genome data to be detected. Due to the above process, before the whole genome data to be detected is analyzed, the present application will first perform bacterial species identification on the whole genome data to be detected, so as to distinguish non-tuberculosis mycobacteria from tuberculosis mycobacteria; and, according to the data similarity between the whole genome data to be detected and the standard strain data corresponding to the target bacillus type, the present application determines the data analysis result of the whole genome data to be detected. Compared with the prior art, the present application can not only realize bacterial species identification for the whole genome data to be detected, but also realize the purpose of performing target bacillus type data analysis on the whole genome data to be detected, and the present application effectively improves the efficiency of data analysis on the whole genome data to be detected, reduces the cost of analyzing the whole genome data to be detected, so that the staff can judge whether further data analysis is needed for the whole genome data to be detected based on the data analysis results.
[0122] It should be understood that, although the steps in the flowcharts of the above embodiments are shown in sequence as indicated by the arrows, these steps are not necessarily performed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be performed in other orders. Moreover, at least a portion of the steps in the flowcharts of the above embodiments may include multiple steps or multiple stages, and these steps or stages are not necessarily performed at the same time, but can be performed at different times. The execution order of these steps or stages is not necessarily to be performed in sequence, but can be performed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0123] Based on the same inventive concept, embodiments of the present application also provide a genomic data analysis device for implementing the aforementioned genomic data analysis method. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations of one or more genomic data analysis device embodiments provided below can be found in the above-described limitations of the genomic data analysis method and will not be further elaborated here.
[0124] In one embodiment, Figure 6 As shown, a genome data analysis device is provided, comprising: a first determination module 10, a second determination module 20 and an analysis module 30, wherein:
[0125] The first determination module 10 is used to perform bacterial species identification on the whole genome data to be detected, and determine the bacterial species type corresponding to the whole genome data to be detected.
[0126] The second determination module 20 is configured to determine the data similarity between the whole genome data to be detected and the standard strain data corresponding to the target Bacillus type if the bacterial species type is the target Bacillus type.
[0127] Among them, the target bacillus type is the Mycobacterium tuberculosis complex type.
[0128] The analysis module 30 is used to perform data analysis on the whole genome data to be detected based on data similarity to obtain data analysis results of the whole genome data to be detected.
[0129] The above-mentioned genome data analysis device determines the bacterial species type corresponding to the whole genome data to be detected, and, when the bacterial species type is the target bacillus type, determines the data similarity between the whole genome data to be detected and the standard strain data corresponding to the target bacillus type, and then determines the data analysis result of the whole genome data to be detected based on the data similarity. Because in the above process, before performing data analysis on the whole genome data to be detected, the present application will first perform bacterial species identification on the whole genome data to be detected, so as to distinguish non-tuberculosis mycobacteria from tuberculosis mycobacteria; and, the present application will determine the data analysis result of the whole genome data to be detected based on the data similarity between the whole genome data to be detected and the standard strain data corresponding to the target bacillus type. Compared with the prior art, the present application can not only realize bacterial species identification for the whole genome data to be detected, but also realize the purpose of performing target bacillus type data analysis on the whole genome data to be detected, and the present application effectively improves the efficiency of data analysis on the whole genome data to be detected, reduces the cost of analyzing the whole genome data to be detected, so that the staff can judge whether further data analysis is needed for the whole genome data to be detected based on the data analysis results.
[0130] In one embodiment, Figure 7 As shown, a genome data analysis device is provided, in which the analysis module 30 includes: a first determination unit 31 and a second determination unit 32, wherein:
[0131] The first determining unit 31 is configured to use the data similarity as a data analysis result of the whole genome data to be detected if the data similarity is less than a similarity threshold.
[0132] The second determining unit 32 is configured to perform data analysis on the whole genome data to be detected if the data similarity is greater than or equal to a similarity threshold, to obtain a data analysis result of the whole genome data to be detected.
[0133] In one embodiment, Figure 8 As shown, a genome data analysis device is provided, in which the second determination unit 32 includes: an acquisition subunit 321, a detection subunit 322 and an analysis subunit 323, wherein:
[0134] The acquisition subunit 321 is used to obtain data analysis requirements corresponding to the whole genome data to be detected.
[0135] The detection subunit 322 is used to perform variation detection on the whole genome data to be detected, and obtain gene sequence variation information of the whole genome data to be detected.
[0136] The analysis subunit 323 is used to perform data analysis on the gene sequence variation information according to data analysis requirements and obtain data analysis results corresponding to the data analysis requirements.
[0137] The analysis subunit is specifically used, if the data analysis requirements include: determining the genotype resistance information of the whole genome data to be tested, performing genotype resistance detection on the gene sequence variation information, and obtaining the genotype resistance information of the whole genome data to be tested.
[0138] The analysis subunit is also specifically used, if the data analysis requirements include: determining the target genome sequence of the whole genome data to be detected, performing parameter assembly processing on the gene sequence variation information, and obtaining the target genome sequence of the whole genome data to be detected.
[0139] The analysis subunit is specifically used if the data analysis requirements also include: determining the genetic evolutionary tree of the whole genome data to be tested; selecting a reference genome sequence from the candidate genome sequences whose consistency with the target genome sequence is greater than a consistency threshold; determining the genetic distance between the target genome sequence and the reference genome sequence; and constructing the genetic evolutionary tree of the whole genome data to be tested based on the genetic distance and the reference genome sequence.
[0140] Each module in the genomic data analysis device described above can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in hardware form, or can be stored in a computer device memory in software form, so that the processor can call and execute the corresponding operations of each module.
[0141] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 9As shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. The processor, the memory and the input / output interface are connected via a system bus, and the communication interface, the display unit and the input device are connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, NFC (near field communication) or other technologies. When the computer program is executed by the processor, a genomic data analysis method is implemented. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device casing, or an external keyboard, touchpad or mouse.
[0142] Those skilled in the art will understand that Figure 9 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0143] In one embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the following steps are implemented:
[0144] Perform bacterial species identification on the whole genome data to be tested to determine the bacterial species type corresponding to the whole genome data to be tested;
[0145] If the bacterial species type is the target Bacillus type, then determine the data similarity between the whole genome data to be tested and the standard strain data corresponding to the target Bacillus type;
[0146] According to the data similarity, data analysis is performed on the whole genome data to be tested to obtain the data analysis results of the whole genome data to be tested.
[0147] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0148] If the data similarity is less than the similarity threshold, the data similarity is used as the data analysis result of the whole genome data to be tested;
[0149] If the data similarity is greater than or equal to the similarity threshold, data analysis is performed on the whole genome data to be detected to obtain a data analysis result of the whole genome data to be detected.
[0150] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0151] Obtain the data analysis requirements corresponding to the whole genome data to be tested;
[0152] Perform mutation detection on the whole genome data to be tested to obtain gene sequence variation information of the whole genome data to be tested;
[0153] According to data analysis requirements, data analysis is performed on the gene sequence variation information to obtain data analysis results corresponding to the data analysis requirements.
[0154] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0155] The genotype resistance test is performed on the gene sequence variation information to obtain the genotype resistance information of the whole genome data to be tested.
[0156] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0157] The gene sequence variation information is assembled with parameters to obtain the target genome sequence of the whole genome data to be tested.
[0158] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0159] From the candidate genome sequences, a reference genome sequence whose consistency with the target genome sequence is greater than a consistency threshold is selected;
[0160] Determine the genetic distance between the target genome sequence and the reference genome sequence;
[0161] Based on the genetic distance and reference genome sequence, a genetic evolutionary tree of the whole genome data to be tested is constructed.
[0162] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0163] The target bacillus type is the Mycobacterium tuberculosis complex type.
[0164] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0165] Perform bacterial species identification on the whole genome data to be tested to determine the bacterial species type corresponding to the whole genome data to be tested;
[0166] If the bacterial species type is the target Bacillus type, then determine the data similarity between the whole genome data to be tested and the standard strain data corresponding to the target Bacillus type;
[0167] According to the data similarity, data analysis is performed on the whole genome data to be tested to obtain the data analysis results of the whole genome data to be tested.
[0168] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0169] If the data similarity is less than the similarity threshold, the data similarity is used as the data analysis result of the whole genome data to be tested;
[0170] If the data similarity is greater than or equal to the similarity threshold, data analysis is performed on the whole genome data to be detected to obtain a data analysis result of the whole genome data to be detected.
[0171] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0172] Obtain the data analysis requirements corresponding to the whole genome data to be tested;
[0173] Perform mutation detection on the whole genome data to be tested to obtain gene sequence variation information of the whole genome data to be tested;
[0174] According to data analysis requirements, data analysis is performed on the gene sequence variation information to obtain data analysis results corresponding to the data analysis requirements.
[0175] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0176] The genotype resistance test is performed on the gene sequence variation information to obtain the genotype resistance information of the whole genome data to be tested.
[0177] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0178] The gene sequence variation information is assembled with parameters to obtain the target genome sequence of the whole genome data to be tested.
[0179] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0180] From the candidate genome sequences, a reference genome sequence whose consistency with the target genome sequence is greater than a consistency threshold is selected;
[0181] Determine the genetic distance between the target genome sequence and the reference genome sequence;
[0182] Based on the genetic distance and reference genome sequence, a genetic evolutionary tree of the whole genome data to be tested is constructed.
[0183] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0184] The target bacillus type is the Mycobacterium tuberculosis complex type.
[0185] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions.
[0186] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processor involved in the various embodiments provided herein may be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic unit, a data processing logic unit based on quantum computing, etc., but are not limited to these.
[0187] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0188] The above embodiments merely illustrate several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art may make various modifications and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. A genome data analysis method, characterized in that: The method comprises: Perform bacterial species identification on the whole genome data to be detected to determine the bacterial species type corresponding to the whole genome data to be detected; If the bacterial species type is the target Bacillus type, determining the data similarity between the whole genome data to be detected and the standard strain data corresponding to the target Bacillus type; According to the data similarity, data analysis is performed on the whole genome data to be detected to obtain a data analysis result of the whole genome data to be detected.
2. The method according to claim 1, characterized in that The step of performing data analysis on the whole genome data to be detected based on the data similarity to obtain a data analysis result of the whole genome data to be detected includes: If the data similarity is less than the similarity threshold, the data similarity is used as the data analysis result of the whole genome data to be detected; If the data similarity is greater than or equal to the similarity threshold, data analysis is performed on the whole genome data to be detected to obtain a data analysis result of the whole genome data to be detected.
3. The method according to claim 2, characterized in that The performing data analysis on the whole genome data to be detected to obtain a data analysis result of the whole genome data to be detected includes: Obtaining data analysis requirements corresponding to the whole genome data to be detected; Performing variation detection on the whole genome data to be detected to obtain gene sequence variation information of the whole genome data to be detected; According to the data analysis requirements, data analysis is performed on the gene sequence variation information to obtain data analysis results corresponding to the data analysis requirements.
4. The method according to claim 3, characterized in that If the data analysis requirement includes: determining the genotype drug resistance information of the whole genome data to be tested; performing data analysis on the gene sequence variation information according to the data analysis requirement to obtain the data analysis results corresponding to the data analysis requirement, including: The gene sequence variation information is subjected to genotype drug resistance detection to obtain genotype drug resistance information of the whole genome data to be detected.
5. The method according to claim 3, characterized in that If the data analysis requirement includes: determining the target genome sequence of the whole genome data to be detected; performing data analysis on the gene sequence variation information according to the data analysis requirement to obtain a data analysis result corresponding to the data analysis requirement, including: The gene sequence variation information is subjected to reference assembly processing to obtain the target genome sequence of the whole genome data to be detected.
6. The method according to claim 5, characterized in that If the data analysis requirement further includes: determining a genetic evolution tree of the whole genome data to be detected; performing data analysis on the gene sequence variation information according to the data analysis requirement to obtain a data analysis result corresponding to the data analysis requirement, including: Selecting a reference genome sequence from the candidate genome sequences whose consistency with the target genome sequence is greater than a consistency threshold; Determining the genetic distance between the target genome sequence and the reference genome sequence; A genetic evolution tree of the whole genome data to be detected is constructed based on the genetic distance and the reference genome sequence.
7. The method according to any one of claims 1 to 6, characterized in that The target bacillus type is a Mycobacterium tuberculosis complex type.
8. A genome data analysis device, characterized in that: The device comprises: The first determination module is used to perform bacterial species identification on the whole genome data to be detected, and determine the bacterial species type corresponding to the whole genome data to be detected; A second determination module is configured to determine the data similarity between the whole genome data to be detected and the standard strain data corresponding to the target Bacillus type if the bacterial species type is the target Bacillus type; The analysis module is used to perform data analysis on the whole genome data to be detected according to the data similarity to obtain a data analysis result of the whole genome data to be detected.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Rapid analysis method and system for genomes of pathogenic microorganisms
CN106886689A
Tubercle bacillus drug resistance detection method and device, computer device and storage medium
CN110706755A
Microbial strain genome analysis method and device and electronic equipment
CN112037847A
Bacteria identification and typing analysis genome database and identification and typing analysis method
CN112863606A
Method and device for creating gene mutation dictionary and method and device for compressing genome data by using gene mutation dictionary
CN114930724A