Drug resistance database construction method and apparatus, drug resistance detection method and apparatus, and device
By screening and building a drug resistance database containing gene dimensions and/or mutation site dimensions, the problems of untimely updates of existing databases and inaccurate information are solved, and the efficiency and accuracy of drug resistance detection are achieved, ensuring the effectiveness of treatment plans.
Patent Information
- Application Number
- PCT/CN2025/078835
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-01
- Filing Date
- 2025-02-24
- Publication Date
- 2025-09-04
AI Technical Summary
The existing drug resistance database is not updated in time, the information is incomplete and inaccurate, resulting in low drug resistance detection efficiency of pathogenic strains and affecting the treatment effect.
By obtaining the preset classification object set corresponding to the target drug and the classification object dimension, using the object weights and sample mutation vector sets in the object feature set for screening, a standard drug resistance mutation information set corresponding to the target drug in the drug resistance database is constructed, including screening of gene dimensions and/or mutation site dimensions.
It improves the efficiency, comprehensiveness and accuracy of the construction of drug resistance databases, supports rapid detection of drug resistance of pathogenic strains, and ensures the effectiveness of treatment plans.
Smart Images

Figure CN2025078835_04092025_PF_FP_ABST
Abstract
Description
A method for constructing a drug resistance database, a drug resistance detection method, a device and an apparatus Technical Field
[0001] The present invention relates to the field of bioinformatics, and in particular to a method for constructing a drug resistance database, a drug resistance detection method, a device and equipment. Background Art
[0002] Patients who miss medication, take medication late, or change or stop medication without authorization during treatment can develop resistance to the drug, rendering the original treatment plan ineffective. Phenotypic resistance testing, used to analyze the resistance of pathogenic strains, often takes several weeks. Waiting for resistance test results before starting medication can significantly delay treatment. Therefore, rapid testing for resistance in pathogenic strains is crucial for disease prevention and control.
[0003] At present, the more commonly used method for detecting drug resistance of pathogenic strains is to use gene sequencing to obtain the mutation site information of the pathogenic strain, compare the mutation site information with the drug resistance mutation information corresponding to various therapeutic drugs compiled in the drug resistance database, and determine whether the pathogenic strain is resistant to a certain therapeutic drug based on the comparison results.
[0004] The above-mentioned drug resistance detection methods have very high requirements for the drug resistance database, while the current drug resistance database relies on manual compilation, and has problems such as untimely updates, incomplete and inaccurate drug resistance mutation information. Summary of the Invention
[0005] The embodiments of the present invention provide a method for constructing a drug resistance database, a method, an apparatus and a device for detecting drug resistance, so as to solve the problem that traditional drug resistance databases need to be manually organized, and improve the efficiency, comprehensiveness and accuracy of constructing drug resistance databases.
[0006] According to one embodiment of the present invention, a method for constructing a drug resistance database is provided, the method comprising:
[0007] Obtaining a preset classification object set corresponding to the target drug and at least one classification object dimension;
[0008] For each classification object dimension, obtaining an object feature set corresponding to each preset classification object in the preset classification object set of the classification object dimension;
[0009] Based on the object weights and sample mutation vector sets in each of the object feature sets, the preset classification object set is screened to obtain a target classification object set;
[0010] Based on at least one target classification object set, constructing a standard drug-resistant mutation information set corresponding to the target drug in a drug resistance database;
[0011] Among them, each of the classification object dimensions includes a gene dimension and / or a mutation point dimension, the sample mutation vector set contains at least two object mutation vectors, the object mutation vector contains at least one mutation point identifier corresponding to the strain sample and the preset classification object, and the mutation point identifier represents whether the reference drug resistance mutation information in the reference drug resistance mutation information set exists in the sample mutation information set corresponding to the strain sample.
[0012] According to one embodiment of the present invention, a method for drug resistance detection is provided, the method comprising:
[0013] Obtaining a mutation information set to be detected of the strain to be detected; wherein the mutation information set to be detected includes at least one mutation information to be detected;
[0014] Obtaining a standard drug resistance mutation information set corresponding to the target drug in a drug resistance database; wherein the standard drug resistance mutation information set includes at least one standard drug resistance mutation information;
[0015] Determining the target drug resistance result of the strain to be tested to the target drug based on the overlapping data corresponding to the mutation information set to be tested and the standard drug resistance mutation information set;
[0016] Wherein, the drug resistance database is obtained by using the method for constructing a drug resistance database as described in any embodiment of the present invention.
[0017] According to another embodiment of the present invention, a device for constructing a drug resistance database is provided, the device comprising:
[0018] A preset classification object set acquisition module is used to obtain a preset classification object set corresponding to the target drug and at least one classification object dimension;
[0019] An object feature set acquisition module is used to acquire, for each classification object dimension, an object feature set corresponding to the target drug and each preset classification object in the preset classification object set of the classification object dimension;
[0020] A preset classification object set screening module is used to screen the preset classification object set to obtain a target classification object set based on the object weight and sample mutation vector set in each object feature set;
[0021] A drug resistance database construction module is used to construct a standard drug resistance mutation information set corresponding to the target drug in the drug resistance database based on at least one target classification object set;
[0022] Among them, each of the classification object dimensions includes a gene dimension and / or a mutation point dimension, the sample mutation vector set contains at least two object mutation vectors, the object mutation vector contains at least one mutation point identifier corresponding to the strain sample and the preset classification object, and the mutation point identifier represents whether the reference drug resistance mutation information in the reference drug resistance mutation information set exists in the sample mutation information set corresponding to the strain sample.
[0023] According to another embodiment of the present invention, a drug resistance detection device is provided, the device comprising:
[0024] A module for acquiring a mutation information set to be detected is used to acquire a mutation information set to be detected of a strain to be detected; wherein the mutation information set to be detected contains at least one mutation information to be detected;
[0025] A standard drug-resistance mutation information set acquisition module is used to acquire a standard drug-resistance mutation information set corresponding to a target drug in a drug-resistance database; wherein the standard drug-resistance mutation information set contains at least one standard drug-resistance mutation information;
[0026] a target drug resistance result determination module, configured to determine the target drug resistance result of the strain to be tested against the target drug based on the overlapping data corresponding to the mutation information set to be tested and the standard drug resistance mutation information set;
[0027] Wherein, the drug resistance database is obtained by using the method for constructing a drug resistance database as described in any embodiment of the present invention.
[0028] According to another embodiment of the present invention, an electronic device is provided, the electronic device including:
[0029] at least one processor; and
[0030] a memory communicatively connected to the at least one processor; wherein,
[0031] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the method for constructing a drug resistance database and / or the drug resistance detection method described in any embodiment of the present invention.
[0032] According to another embodiment of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the method for constructing a drug resistance database and / or the method for detecting drug resistance according to any embodiment of the present invention when executed.
[0033] The technical solution of the embodiment of the present invention is to obtain a preset classification object set corresponding to the target drug and at least one classification object dimension, and for each classification object dimension, obtain an object feature set corresponding to the target drug and each preset classification object in the preset classification object set of the classification object dimension, and based on the object weight and sample mutation vector set in each object feature set, screen the preset classification object set to obtain a target classification object set, and based on at least one target classification object set, construct a standard drug resistance mutation information set corresponding to the target drug in the drug resistance database, wherein each classification object dimension includes a gene dimension and / or a mutation point dimension, which solves the problem that the traditional drug resistance database needs to be manually organized, and improves the construction efficiency, comprehensiveness and accuracy of the drug resistance database.
[0034] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0036] FIG1 is a flow chart of a method for constructing a drug resistance database provided by one embodiment of the present invention;
[0037] FIG2 is a flow chart of a method for screening a preset classification object set provided by one embodiment of the present invention;
[0038] FIG3 is a flowchart of a specific example of a method for constructing a drug resistance database provided by one embodiment of the present invention;
[0039] FIG4 is a flowchart of another method for constructing a drug resistance database provided by one embodiment of the present invention;
[0040] FIG5 is a flowchart of another method for screening a preset classification object set provided by one embodiment of the present invention;
[0041] FIG6 is a flow chart of a drug resistance detection method provided by one embodiment of the present invention;
[0042] FIG7 is a ROC diagram of the drug resistance detection method provided in Example 1 of the present invention for 10% of the test set in Table 1;
[0043] FIG8 is a ROC diagram of the drug resistance detection method provided in Example 2 of the present invention for Table 2;
[0044] FIG9 is a ROC diagram of the 10% test set in Example 1 using the TB-profile software provided in the comparative example;
[0045] FIG10 is a schematic structural diagram of a device for constructing a drug resistance database provided by one embodiment of the present invention;
[0046] FIG11 is a schematic structural diagram of a drug resistance detection device provided by one embodiment of the present invention;
[0047] FIG12 is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0048] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0049] It should be noted that the terms "first", "second", "preset", "target", "reference", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0050] Figure 1 is a flow chart of a method for constructing a drug resistance database according to one embodiment of the present invention. This embodiment is applicable to constructing a drug resistance database for pathogenic strains against one or more therapeutic drugs. The method can be executed by a device for constructing a drug resistance database. The device can be implemented in hardware and / or software and can be configured in a terminal device. As shown in Figure 1, the method includes:
[0051] S110: Obtain a preset classification object set corresponding to the target drug and at least one classification object dimension.
[0052] Specifically, the target drug is associated with the strain belonging to the drug resistance database. That is, the drug resistance database is a database of information on the resistance of a specific strain to the target drug that acts on that strain. This resistance information includes whether the strain is resistant to the target drug, as well as the genetic loci and resistance mutation sites that may be involved in the strain's resistance to the target drug. For a specific strain, the target drug is a drug that can produce drug activity, or at least has produced drug activity in the past, when acting on that strain or the disease caused by that strain. For example, if the strain is Mycobacterium tuberculosis, the target drugs include but are not limited to rifampicin, isoniazid, pyrazinamide, or ethambutol; if the strain is Escherichia coli, the target drugs include but are not limited to amikacin, paracillin-tazobactam, or cefoxitin. The choice of strain and target drug is not limited here, and specific settings can be customized according to actual needs.
[0053] Specifically, the classification object dimension refers to the object dimension for classifying information related to the drug resistance of the strain to the target drug, and the preset classification object set refers to the feature set obtained by classifying the information related to drug resistance for a preset classification object dimension. In the embodiment of the present application, the information related to drug resistance is a reference drug resistance mutation information set, and each classification object dimension may include an object dimension based on drug resistance genes, also referred to as a gene dimension herein; it may also include an object dimension of mutation points that affect drug resistance, also referred to as a mutation point dimension herein; it may also include both a gene dimension and a mutation point dimension. Among them, the reference drug resistance mutation information set contains at least two pre-acquired reference drug resistance mutation information.
[0054] Specifically, drug-resistant mutation information is used to characterize the information on the mutation site in the nucleic acid sequence data of the strain that is resistant to the target drug. For example, drug-resistant mutation information includes but is not limited to the standard base, mutant base, gene name where the mutation site is located, and location of the mutation site on the genome corresponding to the mutation site of the standard nucleic acid sequence of the strain. Drug-resistant mutation information is not limited here and can be set according to actual needs. It should be understood that the standard nucleic acid sequence referred to in the embodiments of the present application refers to the nucleic acid sequence of the strain at the genome where the mutation site is located before the drug-resistant mutation occurs.
[0055] In an optional embodiment, when the classification object dimension is a gene dimension, the preset classification object set is a preset drug-resistance gene set. Specifically, the preset drug-resistance gene set includes at least two preset drug-resistance genes, which are used to characterize genes that are resistant to a target drug in the nucleic acid sequence data of the strain. For example, for a strain of Mycobacterium tuberculosis and a target drug of rifampicin, the preset drug-resistance gene set corresponding to rifampicin includes, but is not limited to, the rpoB gene, the katG gene, the embB gene, and the inhA gene.
[0056] Exemplarily, the nucleic acid sequence data corresponding to strain J contains L preset drug-resistant genes, each preset drug-resistant gene contains one or more drug-resistant mutation sites, and each drug-resistant mutation site corresponds to a drug-resistant mutation information.
[0057] In an optional embodiment, when the classification object dimension is the mutation point dimension, the preset classification object set is a preset drug resistance mutation information set. Specifically, the preset drug resistance mutation information set includes at least two reference drug resistance mutation information sets from a reference drug resistance mutation information set. The reference drug resistance mutation information set is a collection of currently known drug resistance mutation information for at least a portion of the strain to be analyzed. There are no restrictions on how to obtain drug resistance mutation information.
[0058] In an optional embodiment, a preset classification object set corresponding to the target drug and at least one classification object dimension is obtained, including: when the classification object dimension is a mutation point dimension, the reference drug resistance mutation information set is used as the preset drug resistance mutation information set corresponding to the target drug and the mutation point dimension.
[0059] Based on the above embodiment, optionally, when each classification object dimension includes both a gene dimension and a mutation site dimension, each preset classification object set includes a preset drug-resistance gene set and a reference drug-resistance mutation information set. This arrangement has the advantage that the parallel screening dimensions formed by the gene and mutation site dimensions can improve the accuracy of the drug-resistance database.
[0060] In another optional embodiment, when the classification object dimension is the gene dimension, the target classification object set is the target drug-resistant gene set; accordingly, a preset classification object set corresponding to the target drug and at least one classification object dimension is obtained, including: when each classification object dimension includes a gene dimension and a mutation point dimension, based on the target drug-resistant gene set corresponding to the gene dimension, a filtering operation is performed on the reference drug-resistant mutation information set to obtain a preset drug-resistant mutation information set corresponding to the target drug and the mutation point dimension.
[0061] For example, assuming that the target drug-resistant gene set includes drug-resistant gene A and drug-resistant gene B, and the reference drug-resistant mutation information set includes 3 reference drug-resistant mutation information corresponding to drug-resistant gene A, 5 reference drug-resistant mutation information corresponding to drug-resistant gene B, and 2 reference drug-resistant mutation information corresponding to drug-resistant gene C, then the preset drug-resistant mutation information set includes 3 reference drug-resistant mutation information corresponding to drug-resistant gene A and 5 reference drug-resistant mutation information corresponding to drug-resistant gene B, a total of 8 reference drug-resistant mutation information.
[0062] The advantage of this setting is that the serial screening dimension composed of the gene dimension and the mutation point dimension can improve the accuracy of the drug resistance database.
[0063] S120 . For each classification object dimension, obtain an object feature set corresponding to the target drug and each preset classification object in the preset classification object set of the classification object dimension.
[0064] In this embodiment, the object feature set includes object weights and sample mutation vector sets, the sample mutation vector set includes at least two object mutation vectors, the object mutation vector includes at least one mutation point identifier corresponding to the strain sample and the preset classification object, and the mutation point identifier represents whether the reference drug-resistant mutation information in the reference drug-resistant mutation information set exists in the sample mutation information set corresponding to the strain sample.
[0065] Specifically, the object weight is used to characterize the degree of influence of the preset classification object on the drug resistance classification test. In an optional embodiment, when the classification object dimension is the gene dimension, the object weight is the gene weight; when the classification object dimension is the mutation site dimension, the object weight is the mutation site weight. The gene weight is used to characterize the degree of influence of the preset drug resistance gene on the drug resistance classification test, and the mutation site weight is used to characterize the degree of influence of the reference drug resistance mutation site information on the drug resistance classification test.
[0066] Specifically, the sample mutation vector set includes sample mutation vectors corresponding to preset classification objects for at least one strain sample with a drug resistance label of resistant and sample mutation vectors corresponding to preset classification objects for at least one strain sample with a drug resistance label of non-resistant.
[0067] In an optional embodiment, when the classification object dimension is the gene dimension, the object mutation vector is a gene mutation vector; when the classification object dimension is the mutation point dimension, the sample mutation vector is a point mutation vector. Specifically, the gene mutation vector includes the mutation point identifier of at least one reference drug-resistant mutation information corresponding to a preset drug-resistant gene in the reference drug-resistant mutation information set, and the point mutation vector includes the mutation point identifier corresponding to the reference drug-resistant mutation information.
[0068] For example, the mutation point identifier can be represented in the form of a number, text, or graphic. For example, when the mutation point identifier is represented in the form of a number, the mutation point identifier can be 0 or 1; when the mutation point identifier is represented in the form of text, the mutation point identifier can be yes or no; when the mutation point identifier is represented in the form of a graphic, the mutation point identifier can be "○" or "×". The representation format of the mutation point identifier is not limited here and can be customized according to actual needs.
[0069] Specifically, the sample mutation information set includes at least one sample mutation information corresponding to the strain sample, and the sample mutation information may be reference drug-resistant mutation information or other mutation information.
[0070] Taking the object classification dimension as the gene dimension as an example, assuming that the preset drug-resistant gene is drug-resistant gene A, the reference drug-resistant mutation information set contains reference drug-resistant mutation information 1, reference drug-resistant mutation information 2, and reference drug-resistant mutation information 3 corresponding to drug-resistant gene A, and the sample mutation information set corresponding to strain sample I contains reference drug-resistant mutation information 1 and reference drug-resistant mutation information 2, then the gene mutation vector corresponding to strain sample I in the sample mutation vector set corresponding to drug-resistant gene A is [1 1 0].
[0071] Taking the object classification dimension as the mutation point dimension as an example, assuming that the preset classification object is the reference drug-resistant mutation information 1, if the sample mutation information set corresponding to the strain sample I contains the reference drug-resistant mutation information 1, then the point mutation vector corresponding to the strain sample I in the sample mutation information set corresponding to the reference drug-resistant mutation information 1 is [1]; if the sample mutation information corresponding to the strain sample I does not contain the reference drug-resistant mutation information 1, then the point mutation vector corresponding to the strain sample I in the sample mutation information set corresponding to the reference drug-resistant mutation information 1 is [0].
[0072] S130 , based on the object weights and sample mutation vector sets in each object feature set, screening the preset classification object set to obtain a target classification object set.
[0073] In an optional embodiment, based on the object weights and sample mutation vector sets in each object feature set, the preset classification object set is screened to obtain a target classification object set, including: adding the preset classification object with the largest object weight in the preset classification object set as the current classification object to the reference classification object set, and deleting the current classification object from the preset classification object set; based on the sample mutation vector sets corresponding to at least one preset classification object in the reference classification object set, training the first initial model to obtain the current first target model; obtaining the previous first classification performance of the previous first target model in the previous iteration; when the current first classification performance of the current first target model is better than the previous first classification performance, iteratively executing the step of adding the preset classification object with the largest object weight in the preset classification object set as the current classification object to the reference classification object set; until the preset classification object set is an empty set, using the reference classification object set as the target classification object set.
[0074] Specifically, during the iterative screening of the preset classification object set, each preset classification object in the preset classification object set is sequentially added to the reference classification object set as a current classification object according to the order of object weight from large to small.
[0075] Specifically, when the reference classification object set contains a preset classification object, the first initial model is trained based on the sample mutation vector set corresponding to the preset classification object and the drug resistance labels corresponding to each strain sample and the target drug, respectively, to obtain the current first target model. When the reference classification object set contains multiple preset classification objects, for each strain sample, the sample mutation vectors corresponding to the strain sample and each preset classification object are merged into a list mutation vector. Based on each list mutation vector and the drug resistance labels corresponding to each strain sample and the target drug, the first initial model is trained to obtain the current first target model.
[0076] For example, assuming that the reference classification object set contains drug-resistant gene A and drug-resistant gene B, the sample mutation vector set corresponding to drug-resistant gene A contains the gene mutation vector [1 0 0] of strain sample I and the gene mutation vector [0 1 0] of strain sample II, and the sample mutation vector set corresponding to drug-resistant gene B contains the gene mutation vector [1 1 0] of strain sample I and the gene mutation vector [1 0 1] of strain sample II, then the list mutation vectors corresponding to strain sample I and strain sample 2 are [1 0 0 1 1 0] and [0 1 0 1 0 1] respectively.
[0077] Exemplarily, the sample mutation vector set or each list mutation vector is classified to obtain a training set, a validation set and a test set, wherein the training set is used to train the model parameters of the first initial model, the validation set is used to train the hyperparameters of the first initial model, and the test set is used to obtain the current first classification performance of the current first target model.
[0078] Exemplarily, the model architecture of the first initial model includes but is not limited to residual network (ResNet), Transformer network, CNN network (Convolutional Neural Networks), FCN network (Fully Convolutional Networks), DNN network (Deep Neural Networks), RNN network (Recurrent Neural Network) or SVM (Support Vector Machine), etc. The model architecture of the first initial model is not limited here, and can be customized according to actual needs.
[0079] Exemplarily, the performance indicators used for the first classification performance include but are not limited to at least one of accuracy, precision, F1 value, error rate, ROC curve and AUC (area under curve). The performance indicators used for the first classification performance are not limited here and can be customized according to actual needs.
[0080] Specifically, in the first iteration, the previous first classification performance corresponding to the first iteration can be set to a preset value or the maximum object weight in each object feature set. For example, the preset value can be 0. This is not limited here and can be customized according to actual needs.
[0081] In an optional embodiment, the method also includes: if the current accuracy of the current first target model is greater than the previous accuracy of the previous first target model, setting the current first classification performance to be better than the previous first classification performance; if the current accuracy of the current first target model is less than or equal to the previous accuracy of the previous first target model, setting the current first classification performance to be not better than the previous first classification performance.
[0082] In another optional embodiment, the method also includes: obtaining the accuracy difference corresponding to the current accuracy of the current first target model and the previous accuracy of the previous first target model; if the accuracy difference is less than a preset difference threshold, setting the current first classification performance to be better than the previous first classification performance; if the accuracy difference is greater than or equal to the preset difference threshold, setting the current first classification performance to be not better than the previous first classification performance.
[0083] Exemplarily, the preset difference threshold may be 0.01. The preset difference threshold is not limited here and may be customized according to actual needs.
[0084] On the basis of the above embodiment, optionally, based on the object weights and sample mutation vector sets in each object feature set, the preset classification object set is screened to obtain the target classification object set, and the step further includes: when the current first classification performance of the current first target model is not better than the previous first classification performance, deleting the current classification object from the reference classification object set; and iteratively executing the step of adding the preset classification object with the largest object weight in the preset classification object set as the current classification object to the reference classification object set.
[0085] FIG2 is a flowchart of a method for screening a preset classification object set provided by one embodiment of the present invention. FIG2 takes the preset classification object as a preset drug-resistance gene as an example. The preset classification object set includes four drug-resistance genes, represented as Sort_0 = {gene1gene2gene3gene4}. In the i-th iteration, it is determined whether the preset classification object set Sort_i-1 corresponding to the i-1-th iteration is empty. If so, the reference classification object set Volume_i-1 obtained in the i-1-th iteration is used as the target classification object set Target_object. If not, the i-th drug-resistance gene genei in the preset classification object set Sort_0, sorted by gene weight from largest to smallest, is added to the reference classification object set Volume_i-1 to obtain the reference classification object set Volume_i, and the drug-resistance gene genei is deleted from the preset classification object set Sort_i-1 to obtain the preset classification object set Sort_i. Based on the sample mutation vectors corresponding to each drug-resistant gene in the reference classification object set Volume_i and the drug resistance labels of each strain sample, the SVM model is trained to obtain the current first target model SVM_i. Determine whether the current first classification performance of the current first target model SVM_i is better than the current first classification performance of the previous first target model SVM_i-1. If so, set i = i + 1 and start the next iteration. If not, delete the drug-resistant gene genei from the reference classification object set Volume_i, that is, set the reference classification object set Volume_i-1 as the reference classification object set Volume_i, set i = i + 1, and start the next iteration.
[0086] S140. Based on at least one target classification object set, construct a standard drug-resistant mutation information set corresponding to the target drug in the drug resistance database.
[0087] In an optional embodiment, based on at least one target classification object set, a standard drug resistance mutation information set corresponding to the target drug in the drug resistance database is constructed, including: when each classification object dimension only includes a gene dimension, for each target drug resistance gene in the target drug resistance gene set corresponding to the gene dimension, a gene resistance mutation information set consisting of at least one target drug resistance mutation information corresponding to the target drug resistance gene in the reference drug resistance mutation information set is obtained; and each gene resistance mutation information set is used as a standard drug resistance mutation information set corresponding to the target drug in the drug resistance database.
[0088] In another optional embodiment, based on at least one target classification object set, a standard drug resistance mutation information set corresponding to the target drug in the drug resistance database is constructed, including: when each classification object dimension only includes a mutation point dimension, the target drug resistance mutation information set corresponding to the mutation point dimension is used as the standard drug resistance mutation information set corresponding to the target drug in the drug resistance database.
[0089] In another optional embodiment, when the classification object dimension is the gene dimension, the target classification object set is the target drug-resistant gene set; when the classification object dimension is the mutation point dimension, the target classification object set is the target drug-resistant mutation information set; accordingly, based on at least one target classification object set, a standard drug-resistant mutation information set corresponding to the target drug in the drug-resistant database is constructed, including: when each classification object dimension includes a gene dimension and a mutation point dimension in a parallel relationship, for each target drug-resistant gene in the target drug-resistant gene set corresponding to the gene dimension, a gene drug-resistant mutation information set consisting of at least one target drug-resistant mutation information corresponding to the target drug-resistant gene in the reference drug-resistant mutation information set is obtained; based on at least one gene drug-resistant mutation information set and the target drug-resistant mutation information set corresponding to the mutation point dimension, a standard drug-resistant mutation information set corresponding to the target drug in the drug-resistant database is constructed.
[0090] Specifically, a merge operation is performed on each gene resistance mutation information set, and a union operation is performed on the merged set and the target resistance mutation information set to obtain a standard resistance mutation information set corresponding to the target drug in the resistance database.
[0091] In another optional embodiment, based on at least one target classification object set, a standard drug resistance mutation information set corresponding to the target drug in the drug resistance database is constructed, including: when each classification object dimension includes a gene dimension and a mutation point dimension in a serial relationship, the target drug resistance mutation information set corresponding to the mutation point dimension is used as the standard drug resistance mutation information set corresponding to the target drug in the drug resistance database.
[0092] On the basis of the above embodiments, optionally, after constructing a standard drug-resistance mutation information set corresponding to the target drug in the drug-resistance database based on at least one target classification object set, the method further includes: performing a union operation on the standard drug-resistance mutation information set corresponding to the target drug in the drug-resistance database and the traditional drug-resistance mutation information set corresponding to the target drug in the traditional drug-resistance database to obtain a fused drug-resistance database.
[0093] The advantage of this setting is that it expands and filters the traditional drug resistance database, thereby further improving the comprehensiveness and accuracy of the drug resistance database.
[0094] Based on the above embodiment, the standard drug-resistance mutation information set optionally further includes at least one mutation score corresponding to each standard drug-resistance mutation. Specifically, when the standard drug-resistance mutation information is reference drug-resistance mutation information, the mutation score represents the mutation site weight; when the standard drug-resistance mutation information is traditional drug-resistance mutation information, the mutation score represents the credibility rating provided by the traditional drug-resistance database.
[0095] Figure 3 is a flowchart of a specific example of a method for constructing a drug-resistance database provided by one embodiment of the present invention. Specifically, with respect to the gene dimension, each preset drug-resistance gene in the preset drug-resistance gene set is iteratively added to the target drug-resistance gene set based on a descending order of gene weight. The basis for the addition is that the addition of the preset drug-resistance gene to the target drug-resistance gene set results in a 1% improvement in the accuracy of the first target model trained based on the target drug-resistance gene set after the addition. With respect to the mutation point dimension, based on the target drug-resistance gene set, a reference drug-resistance mutation information set is filtered to obtain a preset drug-resistance mutation information set. Each reference drug-resistance mutation information in the preset drug-resistance mutation information set is iteratively added to the target drug-resistance mutation information set based on a descending order of mutation point weight. The basis for the addition is that the addition of the reference drug-resistance mutation information to the target drug-resistance mutation information set results in a 1% improvement in the accuracy of the first target model trained based on the target drug-resistance mutation information set after the addition. The target drug-resistant mutation information set is used as the standard drug-resistant mutation information set, and a union operation is performed on the standard drug-resistant mutation information set and the traditional drug-resistant mutation information set in the traditional drug-resistant database to obtain a fused drug-resistant database.
[0096] The technical solution of this embodiment is to obtain a preset classification object set corresponding to the target drug and at least one classification object dimension, obtain an object feature set corresponding to the target drug and each preset classification object in the preset classification object set of the classification object dimension for each classification object dimension, screen the preset classification object set based on the object weight and sample mutation vector set in each object feature set to obtain a target classification object set, and construct a standard drug resistance mutation information set corresponding to the target drug in the drug resistance database based on at least one target classification object set, wherein each classification object dimension includes a gene dimension and / or a mutation point dimension, thereby solving the problem that the traditional drug resistance database needs to be manually organized, and improving the construction efficiency, comprehensiveness and accuracy of the drug resistance database.
[0097] FIG4 is a flow chart of another method for constructing a drug resistance database provided by one embodiment of the present invention. This embodiment further refines the method for constructing a drug resistance database in the above embodiment. As shown in FIG4 , the method includes:
[0098] S210: Obtain sample nucleic acid sequence data corresponding to each strain sample.
[0099] In an optional embodiment, obtaining sample nucleic acid sequence data corresponding to each strain sample includes: obtaining sample nucleic acid sequence data corresponding to at least two strain samples from a nucleic acid sequence database.
[0100] In another optional embodiment, obtaining sample nucleic acid sequence data corresponding to each strain sample includes: for each strain sample, performing a sequencing operation on the strain nucleic acid of the strain sample to obtain the sample nucleic acid sequence data. Exemplary sequencing tools used in the sequencing operation include, but are not limited to, the Illumina MiSeq tool and the PacBio Sequel II tool. The sequencing tool used in the sequencing operation is not limited herein and can be customized according to actual needs.
[0101] Exemplarily, the nucleic acid sequence data can be whole genome nucleic acid sequence data, or targeted nucleic acid sequence data such as tNGS. The data type of the nucleic acid sequence data is not limited here, and can be customized according to actual needs.
[0102] Based on the above embodiment, optionally, after obtaining the sample nucleic acid sequence data corresponding to each strain sample, the method further includes: performing quality control processing on each sample nucleic acid sequence data. Exemplarily, the quality control processing includes but is not limited to removing adapters, filtering low-quality sequences, and filtering excessively short sequences. The quality control tool used for the quality control processing may be the fastp tool, the Trimmomatic tool, or the FastQC tool. The quality control tool used for the quality control processing is not limited here and can be customized according to actual needs.
[0103] S220 : For each strain sample, perform a mutation processing operation on the sample nucleic acid sequence data corresponding to the strain sample to obtain a sample mutation information set of the strain sample.
[0104] In an optional embodiment, the mutation processing operation includes an alignment operation, a sorting operation, a deduplication operation, a gene mutation point identification operation, a filtering operation, and an annotation operation performed sequentially.
[0105] Specifically, the alignment operation involves comparing the sample nucleic acid sequence data with the standard nucleic acid sequence data of the strain to obtain alignment result data. The alignment result data includes information such as the starting position of the sample nucleic acid sequence data compared to the standard nucleic acid sequence data, the alignment direction, the alignment score, and any mismatches. Exemplary alignment tools used in the alignment operation may include the bowtie2 alignment tool, the BWA-MEM tool, the BWA-MEM2 tool, the SNAP tool, or the Minimap2 tool. The alignment tool used in the alignment operation is not limited herein and can be customized based on actual needs.
[0106] Specifically, the sorting operation refers to sorting the comparison result data according to the comparison coordinates in the comparison result data to obtain sorted comparison result data. Exemplarily, the sorting tool used in the sorting operation can be the SAMtools tool or the sambamba tool. The sorting tool used in the sorting operation is not limited here, and can be customized according to actual needs.
[0107] Specifically, the deduplication operation refers to performing PCR duplication processing on the sorted comparison result data. For example, the repeated redundant parts of the same comparison results aligned to the same comparison coordinate in the sorted comparison result data are removed to obtain deduplicated comparison result data. For example, the deduplication tool used in the deduplication operation can be the GATK tool, the Sambamba tool, the samtools tool or the picard tool. The deduplication tool used in the deduplication operation is not limited here, and can be customized according to actual needs.
[0108] Specifically, the gene mutation point identification operation refers to performing mutation point identification and hard filtering on the deduplicated comparison result data to generate mutation point identification result data. Exemplarily, the gene mutation point identification tool used in the gene mutation point identification operation can be a GATK tool, a varscan tool, a bcftools tool, or a platypus tool. The gene mutation point identification tool used in the gene mutation point identification operation here can be customized according to actual needs.
[0109] Specifically, the filtering operation refers to filtering the mutation site identification result data to remove mutations in the highly variable PE / PPE gene family, repeat regions and mobile elements to obtain filtered mutation site identification result data. Exemplarily, the filtering tool used in the filtering operation can be the VCFtools tool. The filtering tool used in the filtering operation is not limited here, and can be customized according to actual needs.
[0110] Specifically, the annotation operation refers to using an annotation tool to annotate the filtered mutation site identification result data with mutation types, eliminate synonymous mutations, and obtain a sample mutation information set corresponding to the strain sample. Exemplarily, the annotation tool used in the annotation operation can be the ANNOVAR tool, the SnpEff tool, or the Ensembl VEP tool. The annotation tool used in the annotation operation is not limited here, and can be customized according to actual needs.
[0111] S230 , based on the drug resistance labels corresponding to the strain samples, screening the mutation information sets of each sample to obtain at least two initial drug resistance mutation information sets.
[0112] Specifically, the drug resistance label is drug-resistant or drug-inresistant, and the initial drug resistance mutation information set represents the sample mutation information set corresponding to the strain sample whose drug resistance label is drug-resistant.
[0113] S240 , sequentially performing a union operation and a gene filtering operation on each initial drug-resistant mutation information set to obtain a reference drug-resistant mutation information set corresponding to the target drug.
[0114] The gene filtering operation involves filtering the combined resistance information set obtained after the union operation based on a preset resistance gene set to obtain a reference resistance mutation information set. Specifically, for each resistance mutation in the combined resistance information set obtained after the union operation, if the gene containing the resistance mutation is not present in the preset resistance gene set, the resistance mutation is deleted from the combined resistance information set.
[0115] S250: Obtain a preset classification object set corresponding to the target drug and at least one classification object dimension.
[0116] S250 in this embodiment is the same as or similar to S110 shown in FIG. 1 in the above embodiment, and will not be described in detail in this embodiment.
[0117] S260 . For each classification object dimension, obtain an object feature set corresponding to the target drug and each preset classification object in the preset classification object set of the classification object dimension.
[0118] In an optional embodiment, when the classification object dimension is a gene dimension, the preset classification object set is a preset drug-resistant gene set, and the object weight is a gene weight; accordingly, an object feature set corresponding to each preset classification object in the preset classification object set of the target drug and classification object dimension is obtained, including: obtaining a sample mutation vector set corresponding to each preset drug-resistant gene in the preset drug-resistant gene set of the target drug and gene dimension; for each preset drug-resistant gene, based on the sample mutation vector set corresponding to the preset drug-resistant gene, a third initial model is trained to obtain a trained third target model; based on the third classification performance corresponding to the third target model, the gene weight corresponding to the preset drug-resistant gene is determined; and the sample mutation vector set and gene weight corresponding to the preset drug-resistant gene are added to the object feature set corresponding to the preset drug-resistant gene.
[0119] Specifically, the performance parameter value corresponding to the third classification performance is used as the gene weight corresponding to the preset drug-resistant gene. Exemplarily, the model architecture of the third initial model includes, but is not limited to, a Transformer network, a CNN network, an FCN network, a residual network, a DNN network, an RNN network, or a SVM, etc., and the performance indicators used for the third classification performance include, but are not limited to, at least one of accuracy, precision, F1 value, error rate, ROC curve, and AUC. The model architecture of the third initial model and the performance indicators used for the third classification performance are not limited herein and can be customized according to actual needs.
[0120] In an optional embodiment, when the classification object dimension is a mutation point dimension, the preset classification object set is a preset drug resistance mutation information set, the object weight is a mutation point weight, and the object mutation vector is a point mutation vector; accordingly, the object feature set corresponding to each preset classification object in the preset classification object set of the target drug and the classification object dimension is obtained, including: obtaining sample mutation features corresponding to at least two strain samples respectively; wherein the sample mutation features include point mutation vectors corresponding to the strain sample and each reference drug resistance mutation information in the reference drug resistance mutation information set; based on each sample mutation feature, the fourth initial model is trained to obtain a trained fourth target model; for each reference drug resistance mutation information in the preset drug resistance mutation information set, the model weight corresponding to the reference drug resistance mutation information in the fourth target model is used as the mutation point weight; the sample mutation vector set and mutation point weight corresponding to the reference drug resistance mutation information are added to the object feature set corresponding to the reference drug resistance mutation information.
[0121] Specifically, for each reference drug-resistant mutation information in the reference drug-resistant mutation information set, if the sample mutation information set contains the reference drug-resistant mutation information, the point mutation vector corresponding to the reference drug-resistant mutation information in the sample mutation feature is set to 1; if the sample mutation information set does not contain the reference drug-resistant mutation information, the point mutation vector corresponding to the reference drug-resistant mutation information in the sample mutation feature is set to 0.
[0122] In an optional embodiment, the model architecture of the fourth initial model is an SVM model. The SVM model is a two-class classification model, whose basic model is defined as a linear classifier with the largest interval in the feature space. Its learning strategy is to maximize the interval, which can eventually be converted into the solution of a convex quadratic programming problem. In one embodiment, the sample mutation feature is a feature array consisting of 0 and 1, 1 represents that the reference drug-resistant mutation point information is detected in the strain sample, and 0 represents that the reference drug-resistant mutation point information is not detected in the strain sample. Exemplarily, the sample mutation feature of the i-th strain sample can be X i Indicates that drug resistance labels can be Y i express.
[0123] The purpose of the SVM model is to find a maximum segmentation hyperplane to divide the input data into two categories. The linear hyperplane used is defined as Y = (W T X+b), where W contains the model weight corresponding to each point mutation vector in the sample mutation feature. For example, W=[w1, w2, ..., w N ], N represents the number of point mutation vectors, and b represents the intercept. Considering that some strain samples near the boundary may not be well separated by the hyperplane, regularization is used in the SVM model, and the two-classification problem is defined as: i (W T X i +b)≥1-ξ i ,ξ i ≥1, i=1, 2, ..., N, solve the formula under the premise that the conditions are met Established W and b.
[0124] Exemplarily, the performance indicators used for the fourth classification performance include but are not limited to at least one of accuracy, precision, F1 value, error rate, ROC curve and AUC. There is no limitation on the model architecture of the fourth initial model and the performance indicators used for the fourth classification performance, and the specific settings can be customized according to actual needs.
[0125] S270 , based on the object weights and sample mutation vector sets in each object feature set, screening the preset classification object set to obtain a target classification object set.
[0126] In an optional embodiment, S270 is the same as or similar to S130 shown in FIG. 1 in the above embodiment, and is not described again in detail in this embodiment.
[0127] In another optional embodiment, based on the object weights and sample mutation vector sets in each object feature set, the preset classification object set is screened to obtain a target classification object set, including: adding the preset classification object with the largest object weight in the preset classification object set to the current screening classification object set; adding at least one preset classification object that does not exist in the current screening classification object set to the current screening classification object set to obtain at least one current reference classification object set; based on each current reference classification object set and at least two sample mutation vector sets, the preset classification object set is screened to obtain a target classification object set.
[0128] For example, assuming that the preset classification object set includes drug-resistant gene A, drug-resistant gene B, drug-resistant gene C and drug-resistant gene D, the current screening classification object set is [drug-resistant gene A drug-resistant gene B], and accordingly, each reference classification object includes drug-resistant gene C and drug-resistant gene D, then each current reference classification object set includes [drug-resistant gene A drug-resistant gene B drug-resistant gene C] and [drug-resistant gene A drug-resistant gene B drug-resistant gene D].
[0129] On the basis of the above embodiment, optionally, based on each current reference classification object set and at least two sample mutation vector sets, the preset classification object set is screened to obtain a target classification object set, including: for each current reference classification object set, based on the sample mutation vector set corresponding to at least one preset classification object in the current reference classification object set, the second initial model is trained to obtain the current second target model; the previous second classification performance of the previous second target model corresponding to the previous screening classification object set in the previous iteration is obtained; based on the previous second classification performance and the current second classification performance corresponding to at least one current second target model, the next screening classification object set is determined, and the next screening classification object set is used as the current screening classification object set; the step of iteratively adding at least one preset classification object that does not exist in the current screening classification object set to the current screening classification object set to obtain at least one current reference classification object set is executed; until the current second classification performance is not better than the previous second classification performance, the current screening classification object set is used as the target classification object set.
[0130] Specifically, the training method of the second initial model is the same as or similar to the training method of the first initial model in the above embodiment, and will not be repeated here in this embodiment.
[0131] Specifically, based on the previous second classification performance and the current second classification performance corresponding to at least one current second target model, the next screening classification object set is determined, including: taking the optimal current second classification performance as the reference second classification performance; if the reference second classification performance is better than the previous second classification performance, then taking the current reference classification object set corresponding to the reference second classification performance as the next screening classification object set.
[0132] Figure 5 is a flowchart of another method for screening a preset classification object set provided by an embodiment of the present invention. Figure 5 takes the preset classification object as a preset drug-resistant gene as an example. The preset classification object set contains 4 drug-resistant genes, which are expressed as Sort_0 = {gene1gene2gene3gene4}. In the i-th iteration link, k reference drug-resistant genes, namely genei-genej, are added to the screening classification object set Volume_i-1 obtained in the i-1-th iteration link, respectively, to obtain k reference classification object sets, namely Volume_i_1-Volume_i_k. For each reference classification object set, based on at least one sample mutation vector set corresponding to the reference classification object set and the drug resistance label of each strain sample, the SVM model is trained to obtain the current second target model. Obtain k current second target models, namely SVM_i_1-SVM_i_k, and the reference second target model SVM_i_m with the best current second classification performance, and judge whether the reference second classification performance of the reference second target model SVM_i_m is better than the previous second classification performance of the previous second target model SVM_i-1. If not, the screening classification object set Volume_i-1 obtained in the i-1th iteration link is used as the target classification object set Target_object. If so, the reference second target model SVM_i_m is used as the second target model SVM_i corresponding to the i-th iteration link, and the current reference classification object set Volume_i_m corresponding to the reference second target model SVM_i_m is used as the screening classification object set Volume_i corresponding to the i-th iteration link, and set i=i+1 to start the next iteration link.
[0133] S280. Based on at least one target classification object set, construct a standard drug-resistant mutation information set corresponding to the target drug in the drug resistance database.
[0134] S280 in this embodiment is the same as or similar to S140 shown in FIG. 1 in the above embodiment, and will not be described in detail in this embodiment.
[0135] The technical solution of this embodiment obtains the sample nucleic acid sequence data corresponding to each strain sample, performs a mutation processing operation on the sample nucleic acid sequence data corresponding to the strain sample for each strain sample, obtains the sample mutation information set of the strain sample, and based on the drug resistance label corresponding to each strain sample, screens each sample mutation information set to obtain at least two initial drug resistance mutation information sets, performs a union operation and a gene filtering operation on each initial drug resistance mutation information set in turn, and obtains a reference drug resistance mutation information set corresponding to the target drug, which solves the problem of obtaining the reference drug resistance mutation information set and improves the comprehensiveness of the reference drug resistance mutation information set, thereby further improving the construction efficiency, comprehensiveness and accuracy of the drug resistance database.
[0136] Figure 6 is a flow chart of a drug resistance detection method provided by one embodiment of the present invention. This embodiment is applicable to detecting the resistance of pathogenic strains to a certain therapeutic drug. The method can be performed by a drug resistance detection device, which can be implemented in the form of hardware and / or software and can be configured in a terminal device. As shown in Figure 6, the method includes:
[0137] S310: Obtain a mutation information set of the strain to be detected.
[0138] In this embodiment, the mutation information set to be detected includes at least one mutation information to be detected. The method for obtaining the mutation information set to be detected is the same or similar to the method for obtaining the sample mutation information set in the above embodiment, and will not be repeated in this embodiment.
[0139] S320: Obtain a standard drug-resistant mutation information set corresponding to the target drug in the drug-resistant database.
[0140] The drug resistance database used in this embodiment is obtained by the method for constructing a drug resistance database provided in any of the above embodiments, which will not be described in detail here.
[0141] In this embodiment, the standard drug-resistant mutation information set includes at least one piece of standard drug-resistant mutation information.
[0142] S330 , determining the target drug resistance result of the strain to be tested to the target drug based on the overlapping data corresponding to the mutation information set to be tested and the standard drug resistance mutation information set.
[0143] In an optional embodiment, based on the overlapping data corresponding to the mutation information set to be detected and the standard drug-resistant mutation information set, the target drug resistance result of the strain to be detected to the target drug is determined, including: obtaining the number of overlaps corresponding to the mutation information set to be detected and the standard drug-resistant mutation information set; when the number of overlaps is greater than a preset number threshold, the target drug resistance result of the strain to be detected to the target drug is set to resistant; when the number of overlaps is less than or equal to the preset number threshold, the target drug resistance result of the strain to be detected to the target drug is set to non-resistant.
[0144] In another optional embodiment, based on the overlapping data corresponding to the mutation information set to be detected and the standard drug-resistant mutation information set, the target drug resistance result of the strain to be detected to the target drug is determined, including: obtaining the number of overlaps corresponding to the mutation information set to be detected and the standard drug-resistant mutation information set, and taking the ratio of the overlapping number to the standard mutation number corresponding to the standard drug-resistant mutation information set as the overlap rate; when the overlap rate is greater than the overlap rate threshold, the target drug resistance result of the strain to be detected to the target drug is set to drug-resistant; when the overlap rate is less than or equal to the overlap rate threshold, the target drug resistance result of the strain to be detected to the target drug is set to non-drug-resistant.
[0145] In another optional embodiment, the standard drug-resistant mutation information set also includes mutation scores corresponding to each standard drug-resistant mutation information, and the overlapping data includes the overlap rate; accordingly, based on the overlapping data corresponding to the mutation information set to be detected and the standard drug-resistant mutation information set, the target drug resistance result of the strain to be detected for the target drug is determined, including: performing a union operation on the mutation information set to be detected and the standard drug-resistant mutation information set to obtain an overlapping mutation information set; based on the mutation scores corresponding to each standard drug-resistant mutation information in the overlapping mutation information set, the overlap rate of the strain to be detected is determined; based on the overlap rate, the target drug resistance result of the strain to be detected for the target drug is determined.
[0146] Specifically, the statistical value of at least one mutation score corresponding to the overlapping mutation information set is used as the overlap rate of the strain to be tested. Exemplary statistical values include, but are not limited to, sum, maximum, minimum, average, and median values. These statistical values are not limited here and can be customized based on actual needs.
[0147] Specifically, when the overlap rate is greater than the overlap rate threshold, the target drug resistance result of the strain to be tested to the target drug is set to drug-resistant; when the overlap rate is less than or equal to the overlap rate threshold, the target drug resistance result of the strain to be tested to the target drug is set to non-drug-resistant.
[0148] On the basis of the above embodiments, optionally, before obtaining the standard drug-resistant mutation information set corresponding to the target drug in the drug-resistant database, the method also includes: obtaining sample mutation characteristics corresponding to at least two strain samples respectively; wherein the sample mutation characteristics include point mutation vectors corresponding to the strain sample and each reference drug-resistant mutation information in the reference drug-resistant mutation information set; based on each sample mutation characteristic, training at least two fifth initial models respectively to obtain at least two trained fifth target models; based on the mutation information set to be detected and the reference drug-resistant mutation information set, determining the mutation characteristics to be detected corresponding to the strain to be detected; inputting the mutation characteristics to be detected into at least two fifth target models respectively, and obtaining predicted drug resistance results output by each fifth target model; when at least two predicted drug resistance results are the same, using the predicted drug resistance result as the target drug resistance result of the strain to be detected for the target drug.
[0149] Specifically, the model architectures of the fifth initial models are different. In an optional embodiment, the fifth initial models include an SVM model and a CNN network.
[0150] Among them, the CNN network is a new type of artificial neural network method that combines artificial neural networks and deep learning technologies. It has the characteristics of global training that combines local receptive areas, hierarchical structuring, feature extraction and classification processes. Therefore, it has relatively accurate recognition capabilities for local parts of the image. Compared with other image recognition algorithms, the CNN network uses less preprocessing time, which means that the time required for learning is greatly shortened, and the amount of data required to learn free parameters is reduced, thereby reducing the memory requirements for network operation and allowing the construction of more powerful neural networks.
[0151] The embodiment of the present invention folds the sample mutation features into an N*N feature matrix, and fills the feature positions where the sample mutation features are less than N*N with 0 to generate a feature vector similar to the image storage format, which is convenient for subsequent convolution operations. The CNN network constructs a convolution layer to calculate the convolution value of the feature matrix, wherein each convolution layer includes convolution calculation and maximum pooling operation, and then the calculation result is passed to the next convolution layer until the last convolution layer is calculated. After processing through a fully connected layer, it is passed to the final output layer. The output layer uses a softmax activation function to convert the probability value calculated by the fully connected layer into two probability distributions in the range of [0, 1], corresponding to the probability values of the positive and negative classes respectively.
[0152] Exemplarily, the mutation features of each sample are divided into a training set, a test set, and a validation set, and the fifth initial model is trained. The loss function used in the training process may be a cross entropy loss function.
[0153] Based on the above embodiment, the method further includes: when at least two predicted drug resistance results are different, executing the step of obtaining a standard drug resistance mutation information set corresponding to the target drug in the drug resistance database.
[0154] The advantage of this setting is that, on the one hand, multiple deep learning networks perform dual verification of the target drug resistance results, and on the other hand, the deep learning network and the drug resistance database achieve dual prediction of the target drug resistance results. Both aspects help to further improve the accuracy of the target drug resistance results.
[0155] The technical solution of this embodiment is to obtain a mutation information set to be detected of the strain to be detected, wherein the mutation information set to be detected contains at least one mutation information to be detected, obtain a standard resistance mutation information set corresponding to the target drug in the drug resistance database, wherein the standard resistance mutation information set contains at least one standard resistance mutation information, and determine the target resistance result of the strain to be detected for the target drug based on the overlapping data corresponding to the mutation information set to be detected and the standard resistance mutation information set, thereby solving the problem of poor detection performance of resistance detection based on traditional resistance databases, improving the accuracy of target resistance results, and providing data support for subsequent medical tasks.
[0156] The following describes the details in conjunction with specific embodiments.
[0157] Example 1
[0158] The whole genome sequencing data of the MTB strain for which the drug sensitivity test results of four first-line antibacterial drugs, isoniazid (INH), rifampicin (RIF), ethambutol (EMB), and pyrazinamide (PZA) were obtained respectively was used as the sample raw data for training the prediction model of drug resistance of Mycobacterium tuberculosis (MTB) strain to this drug. Specifically, this embodiment collects WGS (whole genome sequencing) data from the NCBI-SRA database as the sample raw data of the present invention. In order to save space, the sequence numbers obtained from the sequencing sequences corresponding to the resistant and non-resistant samples of each drug in the NCBI database are no longer shown.
[0159] Specifically, referring to step S140 above, the drug resistance mutation site database information is obtained based on the sample raw data; the corresponding sample nucleic acid sequence data is obtained based on the drug resistance mutation site database information; the sample nucleic acid sequence data is distributed in an 8:1:1 ratio to obtain a training set, a validation set, and a test set, wherein 80% of the training set and 10% of the validation set are used to train the convolutional neural network model and the SVM model and to construct the drug resistance database. The detection performance of the drug resistance detection method provided by the embodiment of the present invention for the 10% test set in Table 1 is shown in Table 1 below.
[0160] Table 1
[0161] Among them, TP represents the number of positive samples correctly identified, TN represents the number of negative samples correctly identified, FN represents the number of positive samples incorrectly identified, and FP represents the number of negative samples incorrectly identified.
[0162] Figure 7 is a receiver operating characteristic (ROC) plot for a 10% test set using the drug resistance detection method provided in Example 1 of the present invention. Specifically, the abscissa in Figure 7 represents the false positive rate (i.e., 1-specificity), and the ordinate represents sensitivity. The curves from left to right are the ROC curves for isoniazid, rifampicin, ethambutol, and pyrazinamide, respectively. As shown in Figure 7 , the AUCs for isoniazid, rifampicin, ethambutol, and pyrazinamide are 0.98, 0.97, 0.93, and 0.87, respectively.
[0163] Example 2
[0164] It is basically the same as Example 1, except that: the original sample data used is from the article: Chen X, He G, Wang S, et al. Evaluation of Whole-Genome Sequence Method to Diagnose Resistance of 13 Anti-tuberculosis Drugs and Characterize Resistance Genes in Clinical Multi-Drug Resistance Mycobacterium tuberculosis Isolates From China[J].other, 2019.DOI:10.3389 / fmicb.2019.01741, which contains sample nucleic acid sequence data of 424 cases of Mycobacterium tuberculosis and four therapeutic drugs. The sample nucleic acid sequence data is whole genome nucleic acid sequence data, and the last digit represents the drug resistance label. In order to save space, the sample nucleic acid sequence data and its label corresponding to each drug are no longer shown.
[0165] The drug resistance detection performance results of the drug resistance detection method provided by the embodiment of the present invention are shown in Table 2 below. Specifically, the convolutional neural network model, SVM model, and drug resistance database used in the embodiment corresponding to Table 2 were obtained based on the 80% training set and 10% validation set in Table 1.
[0166] Table 2
[0167] Figure 8 is an ROC diagram for a 10% test set using the drug resistance detection method provided in Example 2 of the present invention. Specifically, the abscissa in Figure 8 represents the false positive rate, i.e., 1-specificity, and the ordinate represents sensitivity. The curves from left to right are the ROC curves corresponding to rifampicin, isoniazid, ethambutol, and pyrazinamide, respectively. According to Figure 8, the AUCs corresponding to isoniazid, rifampicin, ethambutol, and pyrazinamide are 0.99, 0.99, 0.86, and 0.81, respectively.
[0168] Comparative Example
[0169] The drug resistance of 10% of the test set in Example 1 was tested using TB-profile software, and the results are shown in Table 5 below.
[0170] Table 3
[0171] As can be seen from Tables 1 and 3, the drug resistance detection method provided by the present invention achieves comparable results for isoniazid and rifampicin resistance monitoring compared to the TB-profile software, while significantly improving performance for ethambutol and pyrazinamide resistance detection. This demonstrates that the drug resistance detection method provided by the present invention has excellent accuracy.
[0172] 2136 Figure 9 is an ROC diagram of drug resistance detection using the drug resistance detection method provided in the comparative example. Specifically, the horizontal axis in Figure 9 represents the false positive rate, that is, 1-specificity, and the vertical axis represents sensitivity. The curves from left to right are the ROC curves corresponding to isoniazid, rifampicin, pyrazinamide and ethambutol, respectively. According to Figure 9, the AUCs corresponding to isoniazid, rifampicin, pyrazinamide and ethambutol are 0.98, 0.97, 0.85 and 0.86, respectively.
[0173] As can be seen from FIG. 7 and FIG. 9 , the drug resistance detection method provided by the embodiment of the present invention has significantly improved AUC performance for ethambutol and pyrazinamide compared to the TB-profile software.
[0174] The following is an embodiment of a device for constructing a drug-resistant database provided in an embodiment of the present invention. This device and the method for constructing a drug-resistant database in the above embodiment belong to the same inventive concept. For details not fully described in the embodiment of the device for constructing a drug-resistant database, please refer to the content of the method for constructing a drug-resistant database in the above embodiment.
[0175] Figure 10 is a schematic diagram of a device for constructing a drug resistance database according to one embodiment of the present invention. As shown in Figure 10 , the device includes a module for acquiring a preset classification object set 410 , an object feature set acquisition module 420 , a module for screening a preset classification object set 430 , and a module for constructing a drug resistance database 440 .
[0176] The preset classification object set acquisition module 410 is used to acquire a preset classification object set corresponding to the target drug and at least one classification object dimension;
[0177] The object feature set acquisition module 420 is used to acquire, for each classification object dimension, an object feature set corresponding to each preset classification object in the preset classification object set of the classification object dimension;
[0178] A preset classification object set screening module 430 is configured to screen the preset classification object set to obtain a target classification object set based on the object weights and sample mutation vector sets in each object feature set;
[0179] A drug resistance database construction module 440 is configured to construct a standard drug resistance mutation information set corresponding to a target drug in a drug resistance database based on at least one target classification object set;
[0180] Among them, each classification object dimension includes a gene dimension and / or a mutation point dimension, and the sample mutation vector set contains at least two object mutation vectors, and the object mutation vector contains at least one mutation point identifier corresponding to the strain sample and the preset classification object. The mutation point identifier represents whether the reference drug-resistant mutation information in the reference drug-resistant mutation information set exists in the sample mutation information set corresponding to the strain sample.
[0181] The technical solution of this embodiment is to obtain a preset classification object set corresponding to the target drug and at least one classification object dimension, and for each classification object dimension, obtain an object feature set corresponding to the target drug and each preset classification object in the preset classification object set of the classification object dimension, and based on the object weights and sample mutation vector sets in each object feature set, screen the preset classification object set to obtain a target classification object set, and based on at least one target classification object set, construct a standard drug resistance mutation information set corresponding to the target drug in the drug resistance database, wherein each classification object dimension includes a gene dimension and / or a mutation point dimension, which solves the problem that the traditional drug resistance database needs to be manually organized, and improves the construction efficiency, comprehensiveness and accuracy of the drug resistance database.
[0182] In an optional embodiment, the preset classification object set screening module 430 includes:
[0183] a first preset classification object set screening unit, configured to add the preset classification object with the largest object weight in the preset classification object set as the current classification object to the reference classification object set, and delete the current classification object from the preset classification object set;
[0184] Based on a sample mutation vector set corresponding to at least one preset classification object in the reference classification object set, the first initial model is trained to obtain a current first target model;
[0185] Obtain the previous first classification performance of the previous first target model in the previous iteration;
[0186] When the current first classification performance of the current first target model is better than the previous first classification performance, iteratively performing the step of adding the preset classification object with the largest object weight in the preset classification object set as the current classification object to the reference classification object set;
[0187] Until the preset classification object set is an empty set, the reference classification object set is used as the target classification object set.
[0188] In an optional embodiment, the first preset classification object set screening unit is further configured to:
[0189] When the current first classification performance of the current first target model is not better than the previous first classification performance, deleting the current classification object from the reference classification object set;
[0190] The step of iteratively executing the step of adding the preset classification object with the largest object weight in the preset classification object set as the current classification object to the reference classification object set.
[0191] In an optional embodiment, the preset classification object set screening module 430 includes:
[0192] A second preset classification object set screening unit, configured to add the preset classification object with the largest object weight in the preset classification object set to the currently screened classification object set;
[0193] adding at least one preset classification object that does not exist in the current screening classification object set to the current screening classification object set to obtain at least one current reference classification object set;
[0194] Based on each current reference classification object set and at least two sample mutation vector sets, the preset classification object set is screened to obtain a target classification object set.
[0195] In an optional embodiment, the second preset classification object set screening unit is specifically configured to:
[0196] For each current reference classification object set, based on a sample mutation vector set corresponding to at least one preset classification object in the current reference classification object set, the second initial model is trained to obtain a current second target model;
[0197] Obtaining the last second classification performance of the last second target model corresponding to the last screened classification object set in the last iteration;
[0198] Determining a next screening classification object set based on the previous second classification performance and the current second classification performance corresponding to the at least one current second target model, and using the next screening classification object set as the current screening classification object set;
[0199] Iteratively executing the step of adding at least one preset classification object that does not exist in the current screening classification object set to the current screening classification object set to obtain at least one current reference classification object set;
[0200] Until each current second classification performance is not better than the previous second classification performance, the current screening classification object set is used as the target classification object set.
[0201] In an optional embodiment, when the classification object dimension is a gene dimension, the target classification object set is a target drug-resistant gene set;
[0202] Accordingly, the preset classification object set acquisition module 410 is specifically used to:
[0203] When each classification object dimension includes a gene dimension and a mutation point dimension, based on the target drug-resistant gene set corresponding to the gene dimension, a filtering operation is performed on the reference drug-resistant mutation information set to obtain a preset drug-resistant mutation information set corresponding to the target drug and mutation point dimension.
[0204] In an optional embodiment, when the classification object dimension is the gene dimension, the target classification object set is the target drug-resistant gene set; when the classification object dimension is the mutation point dimension, the target classification object set is the target drug-resistant mutation information set;
[0205] Accordingly, the drug resistance database construction module 440 is specifically used to:
[0206] When each classification object dimension includes a gene dimension and a mutation point dimension in a parallel relationship, for each target drug-resistant gene in the target drug-resistant gene set corresponding to the gene dimension, obtain a gene drug-resistant mutation information set consisting of at least one target drug-resistant mutation information corresponding to the target drug-resistant gene in the reference drug-resistant mutation information set;
[0207] Based on at least one gene resistance mutation information set and a target resistance mutation information set corresponding to the mutation point dimension, a standard resistance mutation information set corresponding to the target drug in the resistance database is constructed.
[0208] In an optional embodiment, the device further comprises:
[0209] A reference drug resistance mutation information set determination module is used to obtain sample nucleic acid sequence data corresponding to each strain sample;
[0210] For each strain sample, performing a mutation processing operation on the sample nucleic acid sequence data corresponding to the strain sample to obtain a sample mutation information set of the strain sample;
[0211] Based on the drug resistance labels corresponding to each strain sample, each sample mutation information set is screened to obtain at least two initial drug resistance mutation information sets;
[0212] The union operation and gene filtering operation are performed on each initial drug-resistant mutation information set in sequence to obtain the reference drug-resistant mutation information set corresponding to the target drug.
[0213] In an optional embodiment, when the classification object dimension is a gene dimension, the preset classification object set is a preset drug-resistant gene set, and the object weight is a gene weight;
[0214] Accordingly, the object feature set acquisition module 420 includes:
[0215] A first object feature set acquisition unit is used to acquire a sample mutation vector set corresponding to each preset drug-resistant gene in a preset drug-resistant gene set in the target drug and gene dimensions;
[0216] For each preset drug-resistant gene, based on the sample mutation vector set corresponding to the preset drug-resistant gene, the third initial model is trained to obtain a trained third target model;
[0217] Determining gene weights corresponding to preset drug-resistant genes based on the third classification performance corresponding to the third target model;
[0218] The sample mutation vector set and gene weight corresponding to the preset drug-resistant gene are added to the object feature set corresponding to the preset drug-resistant gene.
[0219] In an optional embodiment, when the classification object dimension is a mutation point dimension, the preset classification object set is a preset drug resistance mutation information set, the object weight is a mutation point weight, and the object mutation vector is a point mutation vector;
[0220] Accordingly, the object feature set acquisition module 420 includes:
[0221] The second object feature set acquisition unit is used to obtain sample mutation features corresponding to at least two strain samples respectively; wherein the sample mutation features include point mutation vectors corresponding to the strain sample and each reference drug-resistant mutation information in the reference drug-resistant mutation information set;
[0222] Based on the mutation characteristics of each sample, the fourth initial model is trained to obtain a trained fourth target model;
[0223] For each reference drug-resistant mutation information in the preset drug-resistant mutation information set, the model weight corresponding to the reference drug-resistant mutation information in the fourth target model is used as the mutation point weight;
[0224] The sample mutation vector set and mutation point weight corresponding to the reference drug-resistant mutation information are added to the object feature set corresponding to the reference drug-resistant mutation information.
[0225] The device for constructing a drug resistance database provided in an embodiment of the present invention can execute the method for constructing a drug resistance database provided in any embodiment of the present invention, and has functional modules and beneficial effects corresponding to the execution method.
[0226] The following is an embodiment of a drug resistance detection device provided in an embodiment of the present invention. The device and the drug resistance detection method of the above embodiment belong to the same inventive concept. For details not fully described in the embodiment of the drug resistance detection device, please refer to the content about the drug resistance detection method in the above embodiment.
[0227] FIG11 is a schematic diagram of the structure of a drug resistance detection device provided by an embodiment of the present invention. As shown in FIG11 , the device includes: a to-be-detected mutation information set acquisition module 510 , a standard drug resistance mutation information set acquisition module 520 , and a target drug resistance result determination module 530 .
[0228] The to-be-detected mutation information set acquisition module 510 is configured to acquire the to-be-detected mutation information set of the strain to be detected; wherein the to-be-detected mutation information set includes at least one to-be-detected mutation information;
[0229] The standard drug resistance mutation information set acquisition module 520 is used to obtain the standard drug resistance mutation information set corresponding to the target drug in the drug resistance database; wherein the standard drug resistance mutation information set includes at least one standard drug resistance mutation information;
[0230] The target drug resistance result determination module 530 is used to determine the target drug resistance result of the strain to be tested to the target drug based on the overlapping data corresponding to the mutation information set to be tested and the standard drug resistance mutation information set;
[0231] The drug resistance database is obtained by using the method for constructing a drug resistance database provided in any of the above embodiments.
[0232] The technical solution of this embodiment is to obtain a mutation information set to be detected of the strain to be detected; wherein, the mutation information set to be detected contains at least one mutation information to be detected, obtain a standard resistance mutation information set corresponding to the target drug in the resistance database; wherein, the standard resistance mutation information set contains at least one standard resistance mutation information, and determine the target resistance result of the strain to be detected for the target drug based on the overlapping data corresponding to the mutation information set to be detected and the standard resistance mutation information set, thereby solving the problem of poor detection performance of resistance detection based on traditional resistance databases, improving the accuracy of target resistance results, and providing data support for subsequent medical tasks.
[0233] In an optional embodiment, the standard drug-resistant mutation information set further includes mutation scores corresponding to each standard drug-resistant mutation information, and the overlap data includes the overlap rate;
[0234] Accordingly, the target drug resistance result determination module 530 is specifically configured to:
[0235] Perform a union operation on the mutation information set to be detected and the standard drug-resistant mutation information set to obtain the overlapping mutation information set;
[0236] Based on the mutation scores corresponding to each standard drug-resistant mutation information in the overlapping mutation information set, the overlap rate of the strain to be tested is determined;
[0237] Based on the overlap rate, the target resistance results of the strain to be tested against the target drug are determined.
[0238] In an optional embodiment, the device further comprises:
[0239] The drug resistance prediction result judgment module is used to obtain sample mutation characteristics corresponding to at least two strain samples before obtaining the standard drug resistance mutation information set corresponding to the target drug in the drug resistance database; wherein the sample mutation characteristics include point mutation vectors corresponding to the strain sample and each reference drug resistance mutation information in the reference drug resistance mutation information set;
[0240] Based on the mutation characteristics of each sample, the at least two fifth initial models are trained to obtain at least two trained fifth target models;
[0241] Determine the mutation characteristics corresponding to the strain to be tested based on the mutation information set to be tested and the reference drug-resistant mutation information set;
[0242] Inputting the mutation characteristics to be detected into at least two fifth target models respectively, and obtaining the predicted drug resistance results output by each fifth target model;
[0243] In the case where at least two predicted drug resistance results are the same, the predicted drug resistance result is used as the target drug resistance result of the strain to be tested against the target drug.
[0244] The drug resistance detection device provided in the embodiment of the present invention can execute the drug resistance detection method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0245] FIG12 is a schematic diagram of the structure of an electronic device provided by one embodiment of the present invention. The electronic device 10 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.
[0246] As shown in FIG12 , the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 and a random access memory (RAM) 13, that is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor 11. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the ROM 12 or loaded from the storage unit 18 into the RAM 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0247] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information or data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0248] The processor 11 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors that run machine learning model algorithms, digital signal processors (DSP), and any appropriate processors, controllers, microcontrollers, etc. The processor 11 executes the various methods and processes described above, such as the method for constructing a drug resistance database and / or the drug resistance detection method provided in the above embodiments.
[0249] In some embodiments, the method for constructing a drug resistance database and / or the method for detecting drug resistance provided in the above embodiments may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the method for constructing a drug resistance database and / or the method for detecting drug resistance described above may be executed. Alternatively, in other embodiments, the processor 11 may be configured to execute the method for constructing a drug resistance database and / or the method for detecting drug resistance in any other appropriate manner (for example, by means of firmware).
[0250] Various embodiments of the systems and techniques described herein can be implemented in the following systems or combinations thereof: digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard parts (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0251] The computer programs for implementing the methods for constructing a drug resistance database and / or drug resistance detection methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the computer program is executed by the processor, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer program can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0252] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. A computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable storage medium. Examples of machine-readable storage media can include an electrical connection based on at least one line, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0253] To provide interaction with a user, the systems and techniques described herein can be implemented on a terminal device having: a display device (e.g., a cathode ray tube (CRT) or a liquid crystal display (LCD) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball), through which the user can provide input to the terminal device. Other types of devices can also provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0254] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0255] A computing system may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. The client-server relationship arises through computer programs running on the respective computers and establishing a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, a host product within a cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosts and virtual private server (VPS) services.
[0256] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.
[0257] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. A method for constructing a drug resistance database, characterized in that: include: Obtaining a preset classification object set corresponding to the target drug and at least one classification object dimension; For each classification object dimension, obtaining an object feature set corresponding to each preset classification object in the preset classification object set of the classification object dimension; Based on the object weights and sample mutation vector sets in each of the object feature sets, the preset classification object set is screened to obtain a target classification object set; Based on at least one target classification object set, constructing a standard drug-resistant mutation information set corresponding to the target drug in a drug resistance database; Among them, each of the classification object dimensions includes a gene dimension and / or a mutation point dimension, the sample mutation vector set contains at least two object mutation vectors, the object mutation vector contains at least one mutation point identifier corresponding to the strain sample and the preset classification object, and the mutation point identifier represents whether the reference drug resistance mutation information in the reference drug resistance mutation information set exists in the sample mutation information set corresponding to the strain sample.
2. The method according to claim 1, characterized in that The step of screening the preset classification object set to obtain the target classification object set based on the object weight and the sample mutation vector set in each of the object feature sets includes: Adding the preset classification object with the largest object weight in the preset classification object set as the current classification object to the reference classification object set, and deleting the current classification object from the preset classification object set; Based on a sample mutation vector set corresponding to at least one preset classification object in the reference classification object set, training the first initial model to obtain a current first target model; Obtain the previous first classification performance of the previous first target model in the previous iteration; In a case where the current first classification performance of the current first target model is better than the previous first classification performance, iteratively performing the step of adding the preset classification object with the largest object weight in the preset classification object set as the current classification object to the reference classification object set; Until the preset classification object set is an empty set, the reference classification object set is used as the target classification object set.
3. The method according to claim 2, characterized in that The step of screening the preset classification object set to obtain the target classification object set based on the object weight and the sample mutation vector set in each of the object feature sets further includes: If the current first classification performance of the current first target model is not better than the previous first classification performance, deleting the current classification object from the reference classification object set; The step of iteratively executing the step of adding the preset classification object with the largest object weight in the preset classification object set as the current classification object to the reference classification object set.
4. The method according to claim 1, wherein The step of screening the preset classification object set to obtain the target classification object set based on the object weight and the sample mutation vector set in each of the object feature sets includes: Adding the preset classification object with the largest object weight in the preset classification object set to the current screening classification object set; adding at least one preset classification object that does not exist in the current screening classification object set to the current screening classification object set to obtain at least one current reference classification object set; Based on each of the current reference classification object sets and at least two sample mutation vector sets, the preset classification object set is screened to obtain a target classification object set.
5. The method according to claim 4, characterized in that The step of screening the preset classification object set to obtain the target classification object set based on each of the current reference classification object sets and at least two sample mutation vector sets includes: For each current reference classification object set, based on a sample mutation vector set corresponding to at least one preset classification object in the current reference classification object set, the second initial model is trained to obtain a current second target model; Obtaining the last second classification performance of the last second target model corresponding to the last screened classification object set in the last iteration; Determining a next screening classification object set based on the previous second classification performance and the current second classification performance corresponding to the at least one current second target model, and using the next screening classification object set as the current screening classification object set; Iteratively executing the step of adding at least one preset classification object that does not exist in the current screening classification object set to the current screening classification object set to obtain at least one current reference classification object set; Until the current second classification performances are all not better than the previous second classification performances, the current screened classification object set is used as the target classification object set.
6. The method according to any one of claims 1 to 5, characterized in that When the classification object dimension is a gene dimension, the target classification object set is a target drug-resistant gene set; Accordingly, the step of obtaining a preset classification object set corresponding to the target drug and at least one classification object dimension includes: When each of the classification object dimensions includes a gene dimension and a mutation point dimension, based on the target drug-resistant gene set corresponding to the gene dimension, a filtering operation is performed on the reference drug-resistant mutation information set to obtain a preset drug-resistant mutation information set corresponding to the target drug and the mutation point dimension.
7. The method according to any one of claims 1 to 5, characterized in that When the classification object dimension is the gene dimension, the target classification object set is the target drug-resistant gene set; when the classification object dimension is the mutation point dimension, the target classification object set is the target drug-resistant mutation information set; Accordingly, the method of constructing a standard drug-resistant mutation information set corresponding to the target drug in the drug-resistant database based on at least one target classification object set includes: When each of the classification object dimensions includes a gene dimension and a mutation point dimension in a parallel relationship, for each target drug-resistant gene in the target drug-resistant gene set corresponding to the gene dimension, obtaining a gene drug-resistant mutation information set consisting of at least one target drug-resistant mutation information corresponding to the target drug-resistant gene in the reference drug-resistant mutation information set; Based on at least one gene drug resistance mutation information set and a target drug resistance mutation information set corresponding to the mutation point dimension, a standard drug resistance mutation information set corresponding to the target drug in the drug resistance database is constructed.
8. The method according to claim 1, characterized in that The method further comprises: Obtaining sample nucleic acid sequence data corresponding to each of the strain samples; For each strain sample, performing a mutation processing operation on the sample nucleic acid sequence data corresponding to the strain sample to obtain a sample mutation information set of the strain sample; Based on the drug resistance labels corresponding to the strain samples, the mutation information sets of the samples are screened to obtain at least two initial drug resistance mutation information sets; A union operation and a gene filtering operation are sequentially performed on each of the initial drug-resistant mutation information sets to obtain a reference drug-resistant mutation information set corresponding to the target drug.
9. The method according to claim 1, characterized in that When the classification object dimension is a gene dimension, the preset classification object set is a preset drug-resistant gene set, and the object weight is a gene weight; Accordingly, the obtaining of the object feature set corresponding to each preset classification object in the preset classification object set of the classification object dimension includes: Obtaining a sample mutation vector set corresponding to the target drug and each preset drug-resistant gene in the preset drug-resistant gene set in the gene dimension; For each preset drug-resistant gene, based on the sample mutation vector set corresponding to the preset drug-resistant gene, the third initial model is trained to obtain a trained third target model; Determining a gene weight corresponding to the preset drug-resistant gene based on a third classification performance corresponding to the third target model; The sample mutation vector set and gene weight corresponding to the preset drug-resistant gene are added to the object feature set corresponding to the preset drug-resistant gene.
10. The method according to claim 1, characterized in that When the classification object dimension is a mutation point dimension, the preset classification object set is a preset drug resistance mutation information set, the object weight is a mutation point weight, and the object mutation vector is a point mutation vector; Accordingly, the obtaining of the object feature set corresponding to each preset classification object in the preset classification object set of the classification object dimension includes: Obtaining sample mutation signatures corresponding to at least two strain samples, wherein the sample mutation signatures include point mutation vectors corresponding to the strain sample and each reference drug-resistant mutation information in the reference drug-resistant mutation information set; Based on the mutation characteristics of each sample, the fourth initial model is trained to obtain a trained fourth target model; For each reference drug-resistant mutation information in the preset drug-resistant mutation information set, using the model weight corresponding to the reference drug-resistant mutation information in the fourth target model as the mutation point weight; The sample mutation vector set and mutation point weight corresponding to the reference drug-resistant mutation information are added to the object feature set corresponding to the reference drug-resistant mutation information.
11. A method for detecting drug resistance, characterized in that: include: Obtaining a mutation information set to be detected of the strain to be detected; wherein the mutation information set to be detected includes at least one mutation information to be detected; Obtaining a standard drug resistance mutation information set corresponding to the target drug in a drug resistance database; wherein the standard drug resistance mutation information set includes at least one standard drug resistance mutation information; Determining the target drug resistance result of the strain to be tested to the target drug based on the overlapping data corresponding to the mutation information set to be tested and the standard drug resistance mutation information set; Wherein, the drug resistance database is obtained by using the method for constructing a drug resistance database according to any one of claims 1 to 10.
12. The method according to claim 11, characterized in that The standard drug-resistant mutation information set also includes mutation scores corresponding to each of the standard drug-resistant mutation information, and the overlap data includes the overlap rate; Accordingly, determining the target drug resistance result of the strain to be detected to the target drug based on the overlapping data corresponding to the mutation information set to be detected and the standard drug resistance mutation information set includes: Performing a union operation on the mutation information set to be detected and the standard drug-resistant mutation information set to obtain an overlapping mutation information set; Determining the coincidence rate of the strain to be tested based on the mutation scores corresponding to each of the standard drug-resistant mutation information in the coincident mutation information set; Based on the overlap rate, the target drug resistance result of the strain to be tested to the target drug is determined.
13. The method according to claim 11, characterized in that Before obtaining the standard drug-resistant mutation information set corresponding to the target drug in the drug-resistant database, the method further includes: Obtaining sample mutation features corresponding to at least two strain samples, wherein the sample mutation features include point mutation vectors corresponding to the strain samples and each reference drug-resistant mutation information in the reference drug-resistant mutation information set; Based on the mutation characteristics of each sample, respectively training at least two fifth initial models to obtain at least two trained fifth target models; Determining the mutation characteristics to be detected corresponding to the strain to be detected based on the mutation information set to be detected and the reference drug-resistant mutation information set; Inputting the mutation characteristics to be detected into at least two fifth target models respectively, and obtaining the predicted drug resistance results output by each of the fifth target models; In the case where at least two predicted drug resistance results are the same, the predicted drug resistance result is used as the target drug resistance result of the strain to be tested to the target drug.
14. A device for constructing a drug resistance database, characterized in that: include: A preset classification object set acquisition module is used to obtain a preset classification object set corresponding to the target drug and at least one classification object dimension; An object feature set acquisition module is used to acquire, for each classification object dimension, an object feature set corresponding to the target drug and each preset classification object in the preset classification object set of the classification object dimension; A preset classification object set screening module is used to screen the preset classification object set to obtain a target classification object set based on the object weight and sample mutation vector set in each object feature set; A drug resistance database construction module is used to construct a standard drug resistance mutation information set corresponding to the target drug in the drug resistance database based on at least one target classification object set; Among them, each of the classification object dimensions includes a gene dimension and / or a mutation point dimension, the sample mutation vector set contains at least two object mutation vectors, the object mutation vector contains at least one mutation point identifier corresponding to the strain sample and the preset classification object, and the mutation point identifier represents whether the reference drug resistance mutation information in the reference drug resistance mutation information set exists in the sample mutation information set corresponding to the strain sample.
15. A drug resistance detection device, characterized in that: include: A module for acquiring a mutation information set to be detected is used to acquire a mutation information set to be detected of a strain to be detected; wherein the mutation information set to be detected contains at least one mutation information to be detected; A standard drug-resistance mutation information set acquisition module is used to acquire a standard drug-resistance mutation information set corresponding to a target drug in a drug-resistance database; wherein the standard drug-resistance mutation information set contains at least one standard drug-resistance mutation information; a target drug resistance result determination module, configured to determine the target drug resistance result of the strain to be tested against the target drug based on the overlapping data corresponding to the mutation information set to be tested and the standard drug resistance mutation information set; Wherein, the drug resistance database is obtained by using the method for constructing a drug resistance database according to any one of claims 1 to 10.
16. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the method for constructing a drug resistance database according to any one of claims 1 to 10, and / or the method for detecting drug resistance according to any one of claims 11 to 13.
17. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the method for constructing a drug resistance database according to any one of claims 1 to 10, and / or the method for detecting drug resistance according to any one of claims 11 to 13 when executed.
Citation Information
Patent Citations
Construction of a comparative database and identification of virulence factors through comparison of polymorphic regions in clinical isolates of infectious organisms
CN101421415A
Method and system for determining a bacterial resistance to an antibiotic drug
CN105593865A
Training method and application of mycobacterium tuberculosis drug resistance prediction model based on hierarchical attention neural network
CN114566209A
Microbial drug resistance sequence data acquisition method, device, equipment and medium
CN115810391A
Method for automized remittance using blockchain
KR1020200125534A