An ncRNA gene mutation interpretation method, a storage medium and a terminal
By constructing a scoring model for the harmfulness of ncRNA gene variants and the similarity of disease phenotypes, the problem of low efficiency in interpreting ncRNA gene mutations was solved, and efficient and accurate automated interpretation and reporting were achieved.
Patent Information
- Application Number
- CN202310653276.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-02
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2043-06-02
AI Technical Summary
Existing technologies lack systematic methods and tools for efficiently screening and interpreting ncRNA gene mutations, resulting in low interpretation efficiency and inconsistent reporting, making it impossible to achieve automated interpretation and reporting of ncRNA gene mutations.
We constructed a standard dataset of pathogenic and benign mutations in ncRNA genes, and used various supervised machine learning algorithms to establish a mutation harmfulness scoring model and a disease phenotype similarity scoring model. Through logistic regression modeling, we achieved high-throughput intelligent interpretation and reporting of ncRNA gene mutations.
It has enabled standardized, automated, and high-throughput interpretation and reporting of ncRNA gene mutations, improving interpretation efficiency and accuracy and reducing interpretation inconsistencies.
Smart Images

Figure CN116825192B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of genetic technology, and in particular to an ncRNA gene mutation interpretation method, a storage medium and a terminal. BACKGROUND
[0002] Non-coding RNA (ncRNA) is a kind of RNA that does not encode protein in cells, which can be divided into long non-coding RNA, microRNA, circular RNA, piRNA, small nuclear RNA and small nucleolar RNA according to its physical and physiological characteristics. Recent research has found that ncRNA gene mutation is closely related to the occurrence and development of various human major diseases, and has become a potential marker for disease diagnosis, treatment and prognosis. At present, a typical clinical whole genome sequencing can find millions of genomic variations, and most of the variations fall in the non-coding region of the genome, involving various ncRNA genes. Compared with the interpretation of coding gene variation, there is currently a lack of systematic screening and identification methods to realize the ncRNA gene mutation method and tool system in the whole genome range, which can only rely on manual knowledge and experience to interpret the report based on the published literature and ncRNA disease knowledge base, resulting in very low efficiency of interpretation report. Due to the lack of a unified interpretation report professional standard for ncRNA gene mutation, different interpretation reports may appear for the same detection data. For example, there are currently a variety of ncRNA disease knowledge bases, and different ncRNA disease knowledge bases have strong heterogeneity in disease and ncRNA naming, lacking the use of unified ontology terms for annotation.
[0003] Although some tools for screening and identifying pathogenic genomic variations assisted by artificial intelligence have been developed by international researchers at present, these tools can only screen and identify pathogenic variations on coding genes and their regulatory regions in the theoretical level, and cannot effectively screen and identify pathogenic variations on ncRNA genes, nor can they realize automatic interpretation and report of ncRNA mutations.
[0004] Therefore, the prior art still needs to be further improved and improved. SUMMARY
[0005] In view of the above deficiencies of the prior art, the purpose of the present application is to provide an ncRNA gene mutation interpretation method, which can perform high-throughput screening, interpretation and report of ncRNA pathogenic mutations.
[0006] To achieve this purpose, the present application adopts the following technical solutions:
[0007] In a first aspect, an ncRNA gene mutation interpretation method is provided, which comprises:
[0008] construct a ncRNA gene pathogenic mutation standard dataset, a ncRNA gene benign variation standard dataset and an experimentally verified ncRNA gene variation harmfulness standard dataset;
[0009] Using the ncRNA gene pathogenic mutation standard dataset and the ncRNA gene benign variation standard dataset, model training is respectively performed on two or more supervised machine learning algorithms to obtain a scoring model for specifically evaluating the harmfulness of ncRNA gene variation;
[0010] The scoring model for specifically evaluating the harmfulness of ncRNA gene variation is used to calculate the harmfulness of the variation of a candidate ncRNA gene, and a harmfulness scoring file of the variation site of the ncRNA gene is obtained.
[0011] A standard dataset of ncRNA gene and human disease phenotype association is constructed, and a ncRNA gene related disease phenotype similarity scoring model is constructed based on the standard dataset. The disease phenotype similarity of the candidate ncRNA gene is calculated using the ncRNA gene related disease phenotype similarity scoring model to obtain a disease phenotype similarity scoring file of the candidate ncRNA gene.
[0012] The harmfulness scoring file and the disease phenotype similarity scoring file of the ncRNA gene containing the same number of pathogenic and benign ncRNA gene variation sites are subjected to logistic regression modeling by using data mining software to obtain an algorithm model for specifically screening and identifying ncRNA gene pathogenic mutations.
[0013] The algorithm model is optimized and evaluated by using the experimentally verified ncRNA gene variation harmfulness standard dataset, and the optimized and evaluated algorithm model is used to interpret and report ncRNA gene mutations.
[0014] By constructing a ncRNA gene variation harmfulness standard dataset, based on the standard dataset, an accurate and specific harmfulness scoring model for evaluating ncRNA gene variation is established by using a plurality of advanced supervised machine learning algorithms. Further, an accurate and specific disease phenotype similarity scoring model for evaluating ncRNA gene disease phenotype similarity is established by using a plurality of advanced phenotype similarity algorithms, and then an intelligent screening interpretation and reporting system for ncRNA pathogenic mutations is established by using an advanced machine learning algorithm to fuse the ncRNA gene variation harmfulness scoring model and the ncRNA gene disease phenotype similarity scoring model. The standardization, automation, intelligence and high-throughput clinical interpretation and reporting of ncRNA gene mutations are realized, which can greatly improve the interpretation efficiency and greatly increase the accuracy of interpretation, and overcome the inconsistency of interpretation caused by subjective factors.
[0015] The following are preferred technical solutions of the present application, but not as a limitation of the technical solutions provided by the present application. Through the following preferred technical solutions, the purposes and beneficial effects of the present application can be better achieved and implemented.
[0016] As a preferred technical solution, the ncRNA gene mutation interpretation method, wherein the construction of the ncRNA gene pathogenic mutation standard dataset comprises:
[0017] Obtain various genomic variation data from the disease-related genomic variation database, including ncRNA gene pathogenic mutations and benign variations, and standardize the variation reference genome chromosome position using the USCS liftover tool;
[0018] Obtain various ncRNA genome annotation files, which include the chromosome position information of ncRNA genes, and standardize the names of the various ncRNA genes using professional term resources;
[0019] Based on the disease-related genomic variation and the ncRNA gene chromosome position information, map the ncRNA genomic variation annotation file to the ncRNA gene to obtain the ncRNA gene pathogenic mutation standard dataset.
[0020] As a preferred technical solution, the ncRNA gene mutation interpretation method, wherein the construction of the ncRNA gene benign variation standard dataset comprises:
[0021] Obtain various benign genomic variation data from the healthy population genomic variation database, including ncRNA gene benign variations, and standardize the variation reference genome chromosome position using the USCS liftover tool;
[0022] Obtain various ncRNA genome annotation files, which include the chromosome position information of ncRNA genes, and standardize the names of the various ncRNA genes using professional term resources;
[0023] Based on the genomic benign variation and the ncRNA gene chromosome position information, map the ncRNA genomic variation annotation file to the ncRNA gene to obtain the ncRNA gene benign variation standard dataset.
[0024] As a preferred technical solution, the ncRNA gene mutation interpretation method, wherein the obtaining of the scoring model for evaluating the harmfulness of ncRNA gene variation comprises:
[0025] The ncRNA gene pathogenic mutation standard dataset and the ncRNA gene benign variation standard dataset are used to train and 3-7 times cross-validation of support vector machine and random forest model, and the mean value of the model parameters set by cross-validation is calculated as the final prediction parameter of the support vector machine and random forest model.
[0026] The support vector machine and random forest model are weighted integrated to obtain a weighted integrated prediction model.
[0027] The weighted integrated prediction model is compared with the support vector machine and random forest model to obtain a scoring model for evaluating the harmfulness of ncRNA gene variation.
[0028] As a preferred technical solution, the ncRNA gene mutation interpretation method, wherein the construction of the ncRNA gene and human disease phenotype association standard dataset specifically includes:
[0029] The data associated with various ncRNAs and diseases is obtained by downloading and integrating annotations from disease association databases, and the names of the various ncRNAs are standardized using professional term resources.
[0030] The disease phenotype ontology term database is used to standardize the annotations of various disease phenotype names to obtain the ncRNA gene and human disease phenotype association standard dataset.
[0031] As a preferred technical solution, the ncRNA gene mutation interpretation method, wherein the construction of the ncRNA gene related disease phenotype similarity scoring model specifically includes:
[0032] Based on the ncRNA gene and human disease phenotype association standard dataset, the Phenomizer phenotype similarity algorithm and the Phrank phenotype similarity algorithm are used to establish ncRNA related disease phenotype similarity scoring models, and the two scoring models are compared to obtain the ncRNA gene related disease phenotype similarity scoring model.
[0033] The second aspect is an ncRNA gene mutation interpretation system, which includes:
[0034] The ncRNA gene variation harmfulness dataset construction module is used to construct the ncRNA gene pathogenic mutation standard dataset, the ncRNA gene benign variation standard dataset, and the experimentally verified ncRNA gene variation harmfulness standard dataset.
[0035] The scoring model construction module for specifically evaluating the harmfulness of ncRNA gene variation is used for training two or more supervised machine learning algorithms respectively by using the ncRNA gene pathogenic mutation standard dataset and the ncRNA gene benign variation standard dataset, and obtaining a scoring model for specifically evaluating the harmfulness of ncRNA gene variation; and the scoring model for specifically evaluating the harmfulness of ncRNA gene variation is used to obtain a candidate ncRNA gene variation harmfulness scoring file.
[0036] The ncRNA gene related disease phenotype similarity scoring model construction module is used for constructing a standard dataset of ncRNA gene and human disease phenotype association; and based on the standard dataset, an ncRNA gene related disease phenotype similarity scoring model is constructed, and through disease phenotype association, a pathogenic scoring file of a candidate ncRNA gene set is obtained.
[0037] The algorithm model construction module for ncRNA gene pathogenic mutation screening and identification is used for performing logistic regression modeling on the harmfulness scoring file and the disease phenotype similarity scoring file of ncRNA containing the same number of pathogenic and benign ncRNA variation sites by using data mining software, and obtaining an algorithm model for specifically evaluating the pathogenic mutation of ncRNA gene.
[0038] The ncRNA gene mutation interpretation module is used for optimizing and evaluating the algorithm model by using the experimentally verified ncRNA gene variation harmfulness standard dataset, and interpreting and reporting the ncRNA gene mutation by using the algorithm model after optimization and evaluation.
[0039] As a preferred technical solution, the ncRNA gene mutation interpretation system further comprises an artificial review module for reviewing the result interpreted by the trained algorithm model.
[0040] In a third aspect, a computer readable storage medium stores one or more programs, which are executable by one or more processors to implement the steps of the ncRNA gene mutation interpretation method described above.
[0041] In a fourth aspect, a terminal device comprises a processor, a memory and a communication bus; the memory stores a computer readable program executable by the processor.
[0042] The communication bus realizes the connection and communication between the processor and the memory.
[0043] The processor implements the steps in the ncRNA gene mutation interpretation method as described above when executing the computer readable program.
[0044] Beneficial effects: compared with the prior art, the ncRNA gene mutation interpretation method provided by the application, by constructing a ncRNA gene variation harmfulness scoring model and a disease phenotype similarity scoring model for ncRNA genes, respectively obtaining ncRNA gene variation harmfulness scoring files and pathogenicity scoring files of candidate ncRNA gene sets, and performing logistic regression modeling on the harmfulness scoring files and the disease phenotype similarity scoring files containing the same number of pathogenic and benign ncRNA variation sites. Obtain the algorithm model, optimize and evaluate the algorithm model, and use the optimized and evaluated algorithm model to realize standardized, automatic, intelligent and high-throughput clinical interpretation report for ncRNA gene mutation, greatly improving the interpretation efficiency and accuracy. BRIEF DESCRIPTION OF DRAWINGS
[0045] Figure 1 The ncRNA gene mutation interpretation method flowchart in the application.
[0046] Figure 2 The ncRNA gene variation harmfulness standard data set construction flowchart in the application.
[0047] Figure 3 The ncRNA gene variation harmfulness scoring model establishment method in the application.
[0048] Figure 4 The ncRNA gene variation harmfulness scoring model establishment method in the application.
[0049] Figure 5 The ncRNA gene variation harmfulness scoring model establishment method in the application.
[0050] Figure 6 The ncRNA gene variation harmfulness scoring model establishment method in the application.
[0051] Figure 7 The ncRNA gene variation harmfulness scoring model establishment method in the application.
[0052] Figure 8 The terminal structure in the application. DETAILED DESCRIPTION
[0053] The application provides an ncRNA gene mutation interpretation method, an interpretation system, a computer readable storage medium and a terminal.
[0054] Those skilled in the art can understand that, unless specifically stated otherwise, the singular forms "one", "a", "said" and "the" used herein also include the plural forms. It should be further understood that the phrase "comprising" used in the specification of the present application means that the features, integers, steps, operations, elements and / or components exist, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to the other element, or there can be an intermediate element. In addition, "connected" or "coupled" used herein can include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any single unit and all combinations of the associated listed items.
[0055] Those skilled in the art can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as that generally understood by those skilled in the art to which the present application belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have meanings consistent with those in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined as such.
[0056] It should be understood that the sequence numbers and sizes of the steps in the embodiments do not mean the order of execution, and the execution order of the processes is determined by their functions and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0057] The application content will be further described by the description of the embodiments in combination with the drawings.
[0058] The present embodiment provides an ncRNA gene mutation interpretation method, as shown in the method, the method comprises: Figure 1
[0059] S10, constructing an ncRNA gene pathogenic mutation standard data set, an ncRNA gene benign variation standard data set and an experimentally verified ncRNA gene variation harmfulness standard data set.
[0060] Specifically, the first step is to download and integrate various genomic variation information data from databases such as ncRNAVar, MSDD, GWAS Catalog, COSMIC (Catalogue of Somatic Mutations in Cancer), and IGSR (The International Genome Sample Resource), involving genomic variation names, their chromosomal locations, and pathogenicity information, including ncRNA gene pathogenic mutations and benign mutations, and further using the USCS liftover tool to map the chromosomal location information of genomic variation annotations on different reference genomes (A, C, and D in the Figure 2 Figure 2 B in the middle); the second step is to download the latest ncRNA genomic annotation files from databases such as GENCODE, miRBase, circAtlas, circBase, and piRbase, involving ncRNA gene names and their chromosomal location information (B in the Figure 2 Figure 2 D in the middle); the third step is to map the genomic variations obtained in the first step to the ncRNA genes obtained in the second step based on the chromosomal location information of the genomic variations and ncRNA genes, i.e., to establish ncRNA gene variation harmfulness data, including ncRNA gene pathogenic mutation standard data sets and ncRNA gene benign variation standard data sets (D in the Figure 2
[0061] Feature annotation of ncRNA gene variation harmfulness data (D in the Figure 2
[0062] Systematically integrate from databases such as ncRNAVar, MSDD, miRNASNP, and lncRNASNP, and use professional term resources such as RNAcentral, piRBase, miRBase, circBase, Ensembl, and NONCODE to standardize the names of various ncRNAs, to establish experimentally verified ncRNA gene variation harmfulness data sets (E and D in the Figure 2
[0063] S20, using the ncRNA gene pathogenic mutation standard data set and the ncRNA gene benign variation standard data set, model training is performed on two or more supervised machine learning algorithms respectively to obtain a scoring model for evaluating the harmfulness of ncRNA gene variation.
[0064] Specifically, the ncRNA gene pathogenic mutation standard dataset and the ncRNA gene benign variation standard dataset established by S10 are used to train the model and perform 3-7 times cross-validation on supervised machine learning algorithms such as support vector machine (SVM) and random forest (RF), and the average value of the model parameters set by cross-validation is calculated as the final prediction parameter of the respective individual model. Figure 3 ) Subsequently, the two individual prediction models can be further integrated by weighting to establish a weighted integrated prediction model, and finally the performance of the weighted integrated prediction model is compared with that of the two individual final prediction models (including accuracy and specificity, etc.), to obtain a scoring model for evaluating the harmfulness of ncRNA gene variation ( Figure 3 and Figure 5 ①).
[0065] S30, using the scoring model for evaluating the harmfulness of ncRNA gene variation to calculate the harmfulness of the variation site of the candidate ncRNA gene, to obtain a harmfulness scoring file of the variation site of the ncRNA gene.
[0066] S40, constructing a standard dataset of ncRNA gene and human disease phenotype association; based on the standard dataset, constructing an ncRNA gene related disease phenotype similarity scoring model, using the ncRNA gene related disease phenotype similarity scoring model to calculate the disease phenotype similarity of the candidate ncRNA gene, to obtain a disease phenotype similarity scoring file of the candidate ncRNA gene.
[0067] Specifically, as shown in Figure 4 , first, various ncRNA and disease association data are downloaded and standardized and integrated for annotation from high-quality public databases such as ncRPheno, RNADisease, ncRNAVar, CircR2Disease and piRBase. Subsequently, not only the names of various ncRNAs are standardized using professional term resources such as RNAcentral, piRBase, miRBase, circBase, Ensembl and NONCODE, but also the names of various disease phenotypes are standardized and annotated using disease phenotype ontology term databases such as Human Phenotype Ontology (HPO), OMIMI, Disease Ontology (DO) and Experimental Factor Ontology (EFO), to establish a high-quality standard dataset of ncRNA gene and human disease phenotype association.
[0068] Furthermore, based on a standard dataset linking ncRNAs to human disease phenotypes, phenotype similarity scoring models for ncRNA-related diseases can be established using Phenomizer (a novel method for ranking disease phenotype similarity by searching clinical phenotypes) and Phrank (a novel method for ranking disease phenotype similarity inspired by information theory). By comparing their actual performance (including accuracy and specificity), the optimal disease phenotype similarity scoring model for ncRNA genes can be obtained. Figure 5 (②), thereby associating the disease phenotype with the candidate ncRNA pathogenic gene set and achieving pathogenicity scoring of the candidate gene set.
[0069] S50. Using data mining software, logistic regression modeling is performed on the harmfulness scoring files and the disease phenotype similarity scoring files of ncRNA genes containing the same number of pathogenic and benign ncRNA gene variant sites to obtain an algorithm model for screening and identifying pathogenic mutations of ncRNA genes specifically.
[0070] Specifically, such as Figure 5 As shown, using high-quality pathogenic or functional ncRNA gene variant sites standardized and integrated in S10 and a considerable number of benign ncRNA gene variant sites from the 1000 Genomes Project, logistic regression modeling was performed using the Weka data mining suite on harmfulness scores and disease phenotype similarity scores of ncRNAs containing the same number of pathogenic and benign ncRNA variant sites. The model was trained and tested using 3-7 fold cross-validation, and the average values of the parameters during the cross-validation process were calculated. This established an AI screening and identification algorithm model for accurately and specifically assessing the pathogenicity of ncRNA gene mutations, and scored and classified the clinical pathogenicity levels of the identified ncRNA mutations. Figure 5 (Middle ③)
[0071] S60. The algorithm model is optimized and evaluated using the experimentally verified standard dataset of harmful ncRNA gene mutations, and the optimized algorithm model is used to interpret and report ncRNA gene mutations.
[0072] Specifically, the performance of the artificial intelligence screening and identification algorithm model was optimized and evaluated using the experimentally validated ncRNA gene variant harmfulness dataset established using S10. Figure 5 (④). Finally, based on the ncRNA and ncRNA mutation-disease phenotype association knowledge bases established in S10 and S40 respectively, computer technology is used to realize the automated clinical reporting function of pathogenic ncRNA mutations.
[0073] Based on the ncRNA gene mutation interpretation method described above, the embodiment provides an ncRNA gene mutation interpretation system, as shown in Figure 6 The device comprises:
[0074] A dataset construction module 100 is configured to construct an ncRNA gene pathogenic mutation standard dataset, an ncRNA gene benign variation standard dataset, and an ncRNA gene variation harmfulness standard dataset verified by experiments.
[0075] A harmfulness scoring model construction module 200 is configured to train two or more supervised machine learning algorithms respectively by using the ncRNA gene pathogenic mutation standard dataset and the ncRNA gene benign variation standard dataset, to obtain a specificity evaluation ncRNA gene variation harmfulness scoring model; and to calculate a candidate ncRNA gene variation harmfulness scoring file by using the specificity evaluation ncRNA gene variation harmfulness scoring model.
[0076] A disease phenotype similarity scoring model construction module 300 is configured to construct a standard dataset of ncRNA gene and human disease phenotype association; to construct an ncRNA gene related disease phenotype similarity scoring model based on the standard dataset; and to obtain a pathogenicity scoring file of a candidate ncRNA gene set by associating the disease phenotype to the candidate ncRNA pathogenic gene set.
[0077] An algorithm model construction module 400 is configured to perform logistic regression modeling on the harmfulness scoring and the disease phenotype similarity scoring file of the ncRNA containing the same number of pathogenic and benign ncRNA variation sites by using data mining software, to obtain a screening identification algorithm model for specifically evaluating ncRNA gene pathogenicity mutation.
[0078] An interpretation module 500 is configured to train the algorithm model by using the ncRNA gene variation harmfulness standard dataset verified by experiments, and to interpret and report the ncRNA gene mutation by using the trained algorithm model.
[0079] Specifically, as shown in Figure 7 The sequencing analysis module involves sequencing data quality control (such as total base number, total number of aligned reads, and number of uniquely aligned reads statistics, and sequencing depth analysis), alignment analysis with a reference genome sequence after filtering low-quality reads, and assembly of the consistent sequence of the individual genome. Further, the maximum likelihood genotype of each base site is detected by using a Bayesian statistical model and a genotype likelihood value, and linkage disequilibrium or inference technology is used for optimizing the accuracy of genome variation recognition detection. Finally, a high-confidence genome variation dataset is screened and its distribution in the genome is counted.
[0080] The ncRNA gene variant annotation module is related to the screening and identification of ncRNA gene variants, and further annotates the identified ncRNA variants according to reference genome information, wherein the annotation for different ncRNA types annotates the corresponding genomic features and potential impact information on the structure and function of the ncRNA located therein, including gene variant frequency information in various populations, variant types (such as in promoter regions, pre-miRNA regions, and miRNA seed regions, and splice sites, etc.), RNAfold and other tools to predict the impact of variants on ncRNA structure, whether it is a GWAS-related variant, CADD and GWAVA and other software to predict functional deleteriousness scores, etc.
[0081] The electronic case clinical phenotype extraction module can automatically extract clinical phenotypes from the clinical electronic cases of patients using natural language processing algorithms, and annotate the clinical phenotypes using ontology terms such as human disease phenotype ontology terms; the ncRNA gene related disease phenotype similarity scoring module takes the automatically extracted clinical phenotype ontology terms as input, and based on the established ncRNA gene disease phenotype similarity scoring method, realizes the systematic screening and pathogenicity scoring of the candidate ncRNA pathogenic gene set; the ncRNA gene variant functional deleteriousness scoring module takes the ncRNA variant data identified and annotated by the system as input, and based on the above-mentioned established ncRNA gene variant functional deleteriousness scoring method, realizes the scoring and sorting of candidate ncRNA gene pathogenic mutations.
[0082] The ncRNA gene variant pathogenicity scoring and clinical classification module is a weighted integration and fusion of the ncRNA gene related disease phenotype similarity scoring and the ncRNA gene variant functional deleteriousness scoring based on the above-mentioned artificial intelligence method model, which realizes the pathogenicity scoring and clinical classification of ncRNA gene mutations, wherein the pathogenicity score can be normalized (such as normalized to 0-1), and the pathogenicity of the variant is divided into 5 clinical grades according to the scoring results, including benign, suspected benign, clinical significance unknown, suspected pathogenic and pathogenic, etc.
[0083] The ncRNA gene disease phenotype knowledge base is established by the above-mentioned method. This knowledge base systematically integrates and standardizes the annotation of ncRNA mutations and their related diseases and supporting clinical and experimental evidence information, involving specific ncRNA names and function descriptions, disease names and descriptions and their treatment and pathogenesis information, ncRNA mutation types and pathogenicity information, and allele frequency information in various population databases, etc.
[0084] The ncRNA gene mutation clinical interpretation report module is a functional module for standardizing, automatically generating, intelligently and high-throughput generating the ncRNA gene mutation clinical interpretation report based on the ncRNA gene mutation pathogenicity scoring and clinical classification module and the ncRNA gene disease phenotype knowledge base.
[0085] The expert manual review module can notify and allow manual experts to log in the system to review the ncRNA gene mutation clinical interpretation report automatically generated by the review module, and has the authority of approving whether the report passes, returning and correcting, etc. The gene detection report online display and delivery module can enable the customer to log in the online system to query the report and view the real-time status of the report. At the same time, when the manual expert reviews the report and passes, the module will automatically send the gene detection report to the customer's receiving email and notify the customer to check the gene detection report information through the way of mobile phone short message. In addition, the module also supports the function of online consultation and feedback of the detection report by the customer on mobile clients such as mobile phones and computers.
[0086] Based on the ncRNA gene mutation interpretation method described above, the embodiment provides a computer readable storage medium, the computer readable storage medium stores one or more programs, the one or more programs can be executed by one or more processors to implement the steps in the ncRNA gene mutation interpretation method as described in the above embodiment.
[0087] Based on the ncRNA gene mutation interpretation method described above, the present application also provides a terminal device, as shown in Figure 8 The terminal device includes at least one processor 20, a display screen 21, and a memory 22, and can also include a communications interface 23 and a bus 24. The processor 20, the display screen 21, the memory 22, and the communications interface 23 can communicate with each other through the bus 24. The display screen 21 is configured to display a user guide interface preset in an initial setting mode. The communications interface 23 can transmit information. The processor 20 can call the logic instructions in the memory 22 to execute the method in the above embodiment.
[0088] In addition, the logic instructions in the memory 22 described above can be implemented in the form of a software functional unit and sold or used as an independent product when used, and can be stored in a computer readable storage medium.
[0089] The memory 22, as a computer readable storage medium, can be configured to store software programs, computer executable programs, such as program instructions or modules corresponding to the method in the embodiments of the present disclosure. The processor 20 executes the functions of the application and data processing by running the software programs, instructions or modules stored in the memory 22, that is, implements the method in the above embodiments.
[0090] The memory 22 can include a program storage area and a data storage area, wherein the program storage area can store an operating system and at least one application required by a function; the data storage area can store data created according to the use of the terminal device, etc. In addition, the memory 22 can include a high-speed random access memory, and can also include a non-volatile memory. For example, a variety of media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc., can also be a transitory storage medium.
[0091] In addition, the specific process of the memory medium and the plurality of instructions in the terminal device loaded and executed by the processor has been described in detail in the above method, and will not be repeated here.
[0092] Finally, it should be pointed out that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for interpreting ncRNA gene mutations, characterized in that, include: We constructed a standard dataset of pathogenic mutations in ncRNA genes, a standard dataset of benign mutations in ncRNA genes, and a standard dataset of harmful mutations in ncRNA genes for experimental validation. Using the aforementioned ncRNA gene pathogenic mutation standard dataset and ncRNA gene benign mutation standard dataset, two or more supervised machine learning algorithms were trained to obtain a scoring model that specifically assesses the harmfulness of ncRNA gene mutations. The harmfulness of candidate ncRNA gene variants is calculated using the scoring model specifically used to assess the harmfulness of ncRNA gene variants, resulting in a harmfulness scoring file for the variant sites of ncRNA genes. A standard dataset is constructed to associate ncRNA genes with human disease phenotypes. Based on this standard dataset, a similarity scoring model for ncRNA gene-related disease phenotypes is constructed. The similarity of the disease phenotypes of the candidate ncRNA genes is calculated using the ncRNA gene-related disease phenotype similarity scoring model to obtain the disease phenotype similarity scoring file of the candidate ncRNA genes. Logistic regression modeling was performed on a harm score file containing the same number of pathogenic and benign ncRNA gene variant sites and a disease phenotype similarity score file of ncRNA genes using data mining software to obtain an algorithm model for screening and identifying pathogenic mutations in ncRNA genes that specifically assesses their pathogenicity. The algorithm model was evaluated and its performance was optimized using the experimentally validated standard dataset of harmful ncRNA gene variants. The optimized algorithm model was then used to interpret and report ncRNA gene mutations. The acquisition of the scoring model for specifically assessing the harmfulness of ncRNA gene variants includes: Using the aforementioned standard datasets of pathogenic mutations and benign variants of ncRNA genes, support vector machines and random forest models were trained and cross-validated at a ratio of 3-7. The mean values of the model parameters set for cross-validation were calculated and used as the final prediction parameters for the support vector machine and random forest models. The support vector machine and random forest models are weighted and integrated to obtain a weighted prediction model; The weighted integrated prediction model is compared with the support vector machine and random forest models respectively to obtain the scoring model for specifically assessing the harmfulness of ncRNA gene variants; The construction of the ncRNA gene-related disease phenotype similarity scoring model specifically includes: Based on the standard dataset of associations between the ncRNA gene and human disease phenotypes, ncRNA-related disease phenotype similarity scoring models were established using the Phenomizer phenotype similarity algorithm and the Phrank phenotype similarity algorithm, respectively. The two scoring models were then integrated and compared to obtain the ncRNA gene-related disease phenotype similarity scoring model.
2. The method for interpreting ncRNA gene mutations according to claim 1, characterized in that, The construction of the ncRNA gene pathogenic mutation standard dataset includes: Various genomic variation data, including pathogenic and benign mutations in ncRNA genes, were obtained from disease-related genomic variation databases, and the chromosomal locations of the variations were standardized using the USCS liftover tool; Obtain various ncRNA genome annotation files, which include chromosomal location information of ncRNA genes, and standardize the names of various ncRNA genomes using professional terminology resources; Based on disease-related genomic variations and ncRNA gene chromosomal location information, the ncRNA genomic variation annotation file is mapped to the ncRNA gene to obtain the pathogenic mutation standard dataset of the ncRNA gene.
3. The method for interpreting ncRNA gene mutations according to claim 1, characterized in that, The construction of the standard dataset for benign variants of the ncRNA gene includes: We obtained various benign genomic variant data from a gene variant database of healthy individuals, including benign ncRNA gene variants, and standardized the variant reference chromosomal locations in the genome using the USCS liftover tool. Obtain various ncRNA genome annotation files, which include chromosomal location information of ncRNA genes, and standardize the names of the various ncRNA genes using professional terminology resources; Based on benign genomic variants and ncRNA gene chromosomal location information, the ncRNA genomic variant annotation file is mapped to the ncRNA gene to obtain the standard dataset of benign ncRNA gene variants.
4. The method for interpreting ncRNA gene mutations according to claim 1, characterized in that, The standard dataset for constructing associations between ncRNA genes and human disease phenotypes specifically includes: Data on the association between various ncRNAs and diseases were downloaded from disease-associated databases and integrated and annotated. The names of these various ncRNAs were standardized using specialized terminology resources. By using a human disease phenotype ontology terminology database, standardized annotations were performed on various disease phenotype names to obtain a standardized dataset that associates ncRNA genes with human disease phenotypes.
5. A system for interpreting ncRNA gene mutations, characterized in that, include: The dataset construction module is used to construct standard datasets of pathogenic mutations in ncRNA genes, standard datasets of benign mutations in ncRNA genes, and experimentally validated standard datasets of harmful ncRNA gene mutations. The module for constructing a scoring model for the harmfulness of mutations is used to train two or more supervised machine learning algorithms using the ncRNA gene pathogenic mutation standard dataset and the ncRNA gene benign mutation standard dataset to obtain a scoring model that specifically assesses the harmfulness of ncRNA gene mutations; and to calculate candidate ncRNA gene mutation harmfulness scoring files using the scoring model that specifically assesses the harmfulness of ncRNA gene mutations. The disease phenotype similarity scoring model building module is used to construct a standard dataset that associates ncRNA genes with human disease phenotypes; Based on the standard dataset, a similarity scoring model for ncRNA gene-related disease phenotypes is constructed. By associating disease phenotypes with a set of candidate ncRNA pathogenic genes, a pathogenicity scoring file for the candidate ncRNA gene set is obtained. The algorithm model building module is used to perform logistic regression modeling on files containing the same number of pathogenic and benign ncRNA variant sites and ncRNA disease phenotype similarity scores, using data mining software, to obtain an algorithm model for screening and identifying pathogenic mutations in ncRNA genes that specifically assesses their pathogenicity. The interpretation module is used to optimize and evaluate the algorithm model using the experimentally validated standard dataset of harmful ncRNA gene variations, and to interpret and report ncRNA gene mutations using the optimal algorithm model after optimization and evaluation. The acquisition of the scoring model for specifically assessing the harmfulness of ncRNA gene variants includes: Using the aforementioned standard datasets of pathogenic mutations and benign variants of ncRNA genes, support vector machines and random forest models were trained and cross-validated at a ratio of 3-7. The mean values of the model parameters set for cross-validation were calculated and used as the final prediction parameters for the support vector machine and random forest models. The support vector machine and random forest models are weighted and integrated to obtain a weighted prediction model; The weighted integrated prediction model is compared with the support vector machine and random forest models respectively to obtain the scoring model for specifically assessing the harmfulness of ncRNA gene variants; The construction of the ncRNA gene-related disease phenotype similarity scoring model specifically includes: Based on the standard dataset of associations between the ncRNA gene and human disease phenotypes, ncRNA-related disease phenotype similarity scoring models were established using the Phenomizer phenotype similarity algorithm and the Phrank phenotype similarity algorithm, respectively. The two scoring models were then integrated and compared to obtain the ncRNA gene-related disease phenotype similarity scoring model.
6. The ncRNA gene mutation interpretation system according to claim 5, characterized in that, It also includes a manual review module, which is used to review the results interpreted by the trained algorithm model.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores one or more programs, which can be executed by one or more processors to implement the steps in the method for interpreting ncRNA gene mutations as described in any one of claims 1-4.
8. A terminal device, characterized in that, include: Processor, memory, and communication bus; the memory stores a computer-readable program that can be executed by the processor; The communication bus enables communication between the processor and the memory; When the processor executes the computer-readable program, it implements the steps in the method for interpreting ncRNA gene mutations as described in any one of claims 1-4.
Citation Information
Patent Citations
Gene mutation pathogenicity detection method and system based on neural network and medium
CN111063392A