Tumor driving gene bioinformatics method based on hierarchical constraint automatic learning
Through the automatic learning method of stratified constraints, combining genomic data and clinical feature information, the driver genes of tumor samples are identified, which solves the problem of identification in the existing technology, and achieves rapid and accurate detection direction prompts and efficiency improvements.
Patent Information
- Application Number
- CN202510307270.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-16
- Publication Date
- 2025-07-08
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing technology cannot quickly and accurately identify the driver oncogene corresponding to the tumor sample, resulting in the inability to intelligently prompt the next key direction of detection, and increases the total time required for the gene detection process.
Using a method based on stratified constraint automatic learning, by obtaining genomic data of tumor samples and normal tissue samples, initial abnormal genes are identified and gene regulation networks are established, and multi-dimensional correlation analysis is carried out to identify genes closely related to tumor occurrence and development.
It realizes the rapid and accurate identification of the driver genes for tumor samples, reduces the overall demand for gene detection, provides a clear detection direction, and improves detection efficiency.
Smart Images

Figure CN120279996A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular, to a tumor driver gene bioinformatics method based on hierarchical constraint automatic learning. Background Art
[0002] Worldwide, lung cancer ranks first among the causes of cancer death. There is an urgent need for more effective treatment strategies today. Gene dependence refers to the phenomenon that the occurrence, survival, and proliferation of tumor cells depend on activated oncogenes, and these abnormal molecular signals are also called driver oncogenes. Cancer cells require driver oncogenes to continuously function, while normal cells do not. Therefore, oncogenes can be used as clear therapeutic targets, enabling targeted drugs to specifically kill tumor cells without damaging normal cells.
[0003] There are numerous lung cancer driver genes, and there are complex correlations among the interaction relationships and signaling pathways of each driver gene. In the era of multi-target precision medicine, it is necessary to determine the relationship between driver genes and disease prognosis through comprehensive methods. Currently, traditional reinforcement learning methods face the curse of dimensionality, that is, when the environment is relatively complex or the task is relatively difficult, the number of parameters to be learned and the required storage space increase rapidly, and it is difficult for traditional reinforcement learning methods to achieve ideal results. As a result, it is impossible to quickly and accurately identify the driver oncogenes corresponding to tumor samples as bioinformatics results, and thus it is impossible to intelligently prompt the next key detection direction based on the bioinformatics results, nor can it reduce the total time required for the gene detection process.
[0004] Therefore, there is an urgent need to design a technical solution to solve at least one of the above technical problems. Summary of the Invention
[0005] The main purpose of the embodiments of the present invention is to provide a tumor driver gene bioinformatics method based on hierarchical constraint automatic learning, aiming to solve the problems in the related art that it is impossible to quickly and accurately identify the driver oncogenes corresponding to tumor samples, and thus it is impossible to intelligently prompt the next key detection direction based on the bioinformatics results, nor can it reduce the total time required for the gene detection process.
[0006] In a first aspect, the embodiments of the present invention provide a tumor driver gene bioinformatics method based on hierarchical constraint automatic learning, including:
[0007] Obtain the first genomic data corresponding to the tumor sample of the target patient and the second genomic data corresponding to the normal tissue sample of the target patient from the database;
[0008] Determine the biological development information and clinical characteristic information corresponding to the tumor sample from the database;
[0009] Performing abnormal gene identification based on the first genomic data and the second genomic data to obtain the corresponding initial abnormal gene data in the tumor sample and the initial abnormal types corresponding to the initial abnormal gene data;
[0010] Performing correlation analysis on the initial abnormal gene data according to the biological development information and the initial abnormal types to obtain the first correlation relationship between the initial abnormal gene data;
[0011] Performing correlation analysis on the initial abnormal gene data according to the clinical characteristic information to obtain the second correlation relationship between the initial abnormal gene data;
[0012] Establishing a gene regulatory network for the initial abnormal gene data, and performing correlation analysis on the initial abnormal gene data according to the gene regulatory network to obtain the third correlation relationship between the initial abnormal genes;
[0013] Based on hierarchical constraints, combining the first correlation relationship, the second correlation relationship, and the third correlation relationship to identify the genes closely related to tumor occurrence and development in the tumor sample, and obtaining the target tumor driver gene data as the bioinformatics result.
[0014] In a second aspect, an embodiment of the present invention provides a tumor driver gene bioinformatics system based on hierarchical constraint automatic learning, including:
[0015] A data acquisition module, configured to acquire the first genomic data corresponding to the tumor sample of a target patient and the second genomic data corresponding to the normal tissue sample of the target patient from a database;
[0016] A data determination module, configured to determine the biological development information and clinical characteristic information corresponding to the tumor sample from the database;
[0017] An abnormal identification module, configured to perform abnormal gene identification according to the first genomic data and the second genomic data to obtain the corresponding initial abnormal gene data in the tumor sample and the initial abnormal types corresponding to the initial abnormal gene data;
[0018] A first analysis module, configured to perform correlation analysis on the initial abnormal gene data according to the biological development information and the initial abnormal types to obtain the first correlation relationship between the initial abnormal gene data;
[0019] A second analysis module, configured to perform correlation analysis on the initial abnormal gene data according to the clinical characteristic information to obtain the second correlation relationship between the initial abnormal gene data;
[0020] A third analysis module, configured to establish a gene regulatory network for the initial abnormal gene data, and perform correlation analysis on the initial abnormal gene data according to the gene regulatory network to obtain a third correlation relationship between the initial abnormal genes;
[0021] A gene recognition module, configured to identify genes closely related to tumorigenesis and development in the tumor sample based on hierarchical constraints in combination with the first correlation relationship, the second correlation relationship, and the third correlation relationship, and obtain target tumor driver gene data as a bioinformatics result.
[0022] In a third aspect, an embodiment of the present invention further provides a terminal device, which includes a processor, a memory, a computer program stored on the memory and executable by the processor, and a data bus for realizing connection communication between the processor and the memory. When the computer program is executed by the processor, the steps of any one of the tumor driver gene bioinformatics methods based on hierarchical constraint automatic learning provided in the specification of the present invention are implemented.
[0023] In a fourth aspect, an embodiment of the present invention further provides a storage medium for computer-readable storage, characterized in that the storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of any one of the tumor driver gene bioinformatics methods based on hierarchical constraint automatic learning provided in the specification of the present invention.
[0024] An embodiment of the present invention provides a tumor driver gene bioinformatics method based on hierarchical constraint automatic learning. The method includes: obtaining first genomic data corresponding to a tumor sample of a target patient and second genomic data corresponding to a normal tissue sample of the target patient from a database; determining biological development information and clinical feature information corresponding to the tumor sample from the database; performing abnormal gene identification based on the first genomic data and the second genomic data to obtain initial abnormal gene data corresponding to the tumor sample and an initial abnormal type corresponding to the initial abnormal gene data, so as to identify initial abnormal genes and their types by comparing the genomic data of the tumor sample and the normal tissue sample, and clearly reveal the differences between tumor cells and normal cells at the genomic level; performing correlation analysis on the initial abnormal gene data according to the biological development information and the initial abnormal type to obtain a first correlation relationship between the initial abnormal gene data; performing correlation analysis on the initial abnormal gene data according to the clinical feature information to obtain a second correlation relationship between the initial abnormal gene data; establishing a gene regulatory network for the initial abnormal gene data, and performing correlation analysis on the initial abnormal gene data according to the gene regulatory network to obtain a third correlation relationship between the initial abnormal genes, so that there are significant differences in genomic characteristics, biological development, and clinical characteristics among tumor samples of different patients. Using multi-dimensional data such as biological development information and clinical feature information for correlation analysis can more comprehensively understand the heterogeneity of tumors, and then identify genes closely related to tumor occurrence and development in tumor samples based on hierarchical constraints combined with the first correlation relationship, the second correlation relationship, and the third correlation relationship, and obtain target tumor driver gene data as the bioinformatics result. The determination of the target tumor driver gene data as the bioinformatics result can provide a clear detection direction for subsequent tumor detection and reduce the detection of irrelevant genes. It also solves the problem in the related art that the driver oncogene corresponding to the tumor sample cannot be quickly and accurately identified, and thus the next detection key direction cannot be intelligently prompted according to the bioinformatics result, nor can the total time required for the gene detection link be reduced. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0026] Figure 1 It is a schematic flowchart of a tumor driver gene bioinformatics method based on hierarchical constraint automatic learning provided by an embodiment of the present invention;
[0027] Figure 2Schematic diagram of the module structure of a tumor driver gene bioinformatics system based on hierarchical constraint automatic learning provided by an embodiment of the present invention;
[0028] Figure 3 Schematic block diagram of the structure of a terminal device provided by an embodiment of the present invention. Detailed implementation manners
[0029] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0030] The flowchart shown in the accompanying drawings is only an example illustration, and does not necessarily include all contents and operations / steps, nor does it necessarily need to be executed in the described order. For example, some operations / steps can be decomposed, combined, or partially merged, so the actual execution order may change according to the actual situation.
[0031] It should be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless otherwise clearly specified in the context, the singular forms "a", "an", and "the" are intended to include the plural forms.
[0032] An embodiment of the present invention provides a tumor driver gene bioinformatics method based on hierarchical constraint automatic learning. Among them, the tumor driver gene bioinformatics method based on hierarchical constraint automatic learning can be applied to a terminal device, and the terminal device can be an electronic device such as a tablet computer, a notebook computer, a desktop computer, a personal digital assistant, and a wearable device. The terminal device can be a server or a server cluster.
[0033] Next, some embodiments of the present invention will be described in detail with reference to the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0034] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a tumor driver gene bioinformatics method based on hierarchical constraint automatic learning provided by an embodiment of the present invention.
[0035] As Figure 1 shown, the tumor driver gene bioinformatics method based on hierarchical constraint automatic learning includes steps S101 to S107.
[0036] Step S101: Obtain the first genomic data corresponding to the tumor sample of the target patient and the second genomic data corresponding to the normal tissue sample of the target patient from the database.
[0037] Exemplarily, determine the genomic data content, such as whole-genome sequencing data, exome sequencing data, sequencing data of specific gene regions, etc., and then collect the unique identifier of the historical patient (such as medical record number, ID number, etc.), and then store it in the database according to the unique identifier and the genomic data content.
[0038] Exemplarily, obtain the target identifier corresponding to the target patient and the characteristics of the database, then determine the query conditions according to the target identifier and the query strategy corresponding to the database, and then perform a query operation in the database according to the query conditions to obtain the first genomic data corresponding to the tumor sample of the target patient and the second genomic data corresponding to the normal tissue sample of the target patient.
[0039] Exemplarily, the database can also be used to determine the RNA-seq data of lung adenocarcinoma and normal lung tissue through the Cancer Genome Atlas (TCGA) database and the Genotype-Tissue Expression (GTEx) database.
[0040] For example, after completing account registration on platforms such as the TCGA database, GTEx database, Immune Single-Cell Center, and Human Protein Atlas, collect the clinical laboratory indicators of lung adenocarcinoma patients in the corresponding databases respectively, such as blood biochemical indicators (liver and kidney function, tumor markers, etc.), imaging examination results (CT, PET-CT, etc.), and pathological diagnosis reports. For example, enter the search interface of the TCGA database, and use keywords "Lung Adenocarcinoma (LUAD)" and "RNA-seq" for screening. Set appropriate filters, such as selecting gene expression quantification for the data type and tumor tissue sample for the sample type, then access the GTEx database, and enter "lung tissue" and "RNA-seq" in the search box. Screen out the RNA-seq data of normal lung tissue, and then download the data to the local according to the platform guidance, and then organize the RNA-seq data downloaded from TCGA and GTEx to obtain the first genomic data corresponding to the tumor sample of the target patient and the second genomic data corresponding to the normal tissue sample of the target patient.
[0041] Step S102: Determine the biological development information and clinical characteristic information corresponding to the tumor sample from the database.
[0042] Exemplarily, obtain the biological development information and clinical characteristic information corresponding to the tumor sample from the database according to the target identifier corresponding to the target patient.
[0043] Exemplarily, the biological development information mainly focuses on information related to various changes and development processes of tumor cells at the biological level, including but not limited to the proliferation rate of tumor cells, proliferation index, degree of tumor cell differentiation, invasion and metastasis ability of tumor cells, characteristics of the tumor microenvironment, and tumor heterogeneity.
[0044] Exemplarily, the clinical feature information focuses on various symptoms, signs shown by patients clinically, and information related to disease diagnosis, treatment, and prognosis, including but not limited to basic information of patients such as age, gender, etc., symptoms and signs such as pain, cough, hemoptysis, weight loss, etc., disease stage, and treatment history.
[0045] Step S103: Identify abnormal genes based on the first genomic data and the second genomic data to obtain the corresponding initial abnormal gene data in the tumor sample and the initial abnormal type corresponding to the initial abnormal gene data.
[0046] Exemplarily, according to the length of the gene sequence, it is equally divided into several segments of a fixed length; alternatively, the first gene data can be segmented according to factors such as the functional region of the gene, chromosomal location, etc. to obtain the first segmented data, and its characteristics are analyzed and the corresponding first distribution information is obtained. The first distribution information is the occurrence frequency of this sequence pattern.
[0047] Exemplarily, the second gene data is segmented using the same or similar segmentation rules as those for processing the first gene data to obtain the second segmented data. Ensure the consistency of the segmentation method for subsequent effective comparison. Similarly, for each second segmented data, its corresponding second distribution information is obtained, providing a basis for subsequent similarity calculation and abnormal judgment.
[0048] Exemplarily, the base sequence similarity between each first segmented data and all second segmented data is calculated to measure the similarity between them to obtain the target similarity value. Then, based on the calculated target similarity value, the second segmented data that is most similar to each first segmented data, that is, the nearest segmented data, is found from all second segmented data. The second segmented data with the largest similarity value can be selected as the nearest segmented data by sorting the target similarity values.
[0049] Exemplarily, the third distribution information corresponding to the nearest segmented data is extracted from the second distribution information. Then, the first distribution information of the first segmented data is compared with the third distribution information of the nearest segmented data, and the distribution difference between them is calculated. And a preset value for abnormal judgment is set in advance. When the calculated distribution difference is greater than or equal to this preset value, it can be considered that the first segmented data has a large difference from the normal situation, and it is determined as the initial abnormal gene data.
[0050] Exemplarily, analyze the differences in gene sequences between the two to determine the possible types of gene changes, such as single nucleotide variations (base substitutions), indel variations (insertions or deletions of fragments), copy number variations, etc. Then, based on the determined types of gene changes and combined with existing biological knowledge and relevant research, judge the initial abnormal type corresponding to the initial abnormal gene data. For example, if the type of gene change is a single nucleotide variation in a key functional region and this variation may affect the structure and function of the protein, then the corresponding initial abnormal type may be a pathogenic mutation; if it is a small indel in some non-critical regions, it may be a benign variation, etc.
[0051] In some embodiments, the obtaining of the initial abnormal gene data corresponding to the tumor sample and the initial abnormal type corresponding to the initial abnormal gene data by identifying abnormal genes according to the first genomic data and the second genomic data includes: determining a target window, and performing data segmentation on the first genomic data according to the target window to obtain a plurality of first segmentation data; performing data segmentation on the second genomic data according to the target window to obtain a plurality of second segmentation data; obtaining any two data from the first segmentation data and determining them as the first data and the second data, and obtaining the first associated data corresponding to the first data and the second associated data corresponding to the second data from the second segmentation data; performing an intersection calculation on the first associated data and the second data to obtain a first result, and performing an intersection calculation on the second associated data and the first data to obtain a second result; when the first result is not an empty set and the second result is not an empty set, then determining the second data as the initial associated data corresponding to the first data; performing abnormal analysis on the first data according to the initial associated data to obtain the target outlier value corresponding to the first data; determining the initial abnormal gene data corresponding to the tumor sample according to the target outlier value; obtaining the associated gene data corresponding to the initial abnormal gene data from the second genomic data, and determining the initial abnormal type corresponding to the initial abnormal gene data according to the associated gene data.
[0052] Exemplarily, determine the length corresponding to the target window according to a preset gene sequence length range. Then, taking the target window as a unit, segment the first genomic data segment by segment, and divide it into a plurality of first segmentation data. The length of each first segmentation data is equal to the size of the target window. Use the same target window to segment the second genomic data to obtain a plurality of second segmentation data.
[0053] Exemplarily, any two data are randomly selected from multiple first segmentation data, and they are respectively determined as the first data and the second data. Among multiple second segmentation data, multiple second segmentation data with similar characteristics or positional relationships to the first data are found according to the similarity of gene sequences, and they are determined as the first associated data; similarly, the second associated data corresponding to the second data is found. For example, the second segmentation data corresponding to when the similarity is less than a preset value is determined as the first associated data, or the second segmentation data corresponding to when the similarity is sorted among the top N data is determined as the first associated data, where N represents the target quantity and is a positive integer.
[0054] Exemplarily, the first associated data and the second data are subjected to an intersection calculation to obtain a first result, that is, to determine whether the second data is within the first associated data of the first data, and then the second associated data and the first data are subjected to an intersection calculation to obtain a second result. That is, to determine whether the first data is within the second associated data of the second data.
[0055] Exemplarily, if the first result is not an empty set and the second result is not an empty set, it indicates that the second data is within the first associated data of the first data and the first data is within the second associated data of the second data. Then, the second data is determined as the initial associated data corresponding to the first data, and the remaining data in the first segmentation data and the first data also perform the above steps to obtain all the initial associated data corresponding to the first data.
[0056] Exemplarily, taking the initial associated data as a reference, an outlier analysis is performed on the first data. By comparing the differences between the first data and the initial associated data in terms of gene sequences, expression levels, etc., a value measuring the degree of deviation of the first data from the normal situation is calculated, that is, the target outlier value.
[0057] Exemplarily, a threshold for the target outlier value is preset in advance, and this threshold can be determined according to a large amount of experimental data or experience. The calculated target outlier value is compared with the threshold. If the target outlier value exceeds the threshold, it is considered that the gene segment corresponding to the first data is abnormal, and it is determined as the initial abnormal gene data corresponding to the tumor sample.
[0058] Exemplarily, in the second genomic data, gene data associated with the initial abnormal gene data in terms of position, function, etc. is found and determined as the associated gene data. Thus, by comparing the differences between the initial abnormal gene data and the associated gene data, combined with existing biological knowledge and research results, the initial abnormal type corresponding to the initial abnormal gene data is determined. For example, if there is a deletion of a certain gene in the initial abnormal gene data while the gene is normal in the associated gene data, the abnormal type can be determined as gene deletion.
[0059] Specifically, by using the target window for data segmentation, the genomic data is divided into comparable segments, avoiding the interference caused by differences in data length and structure, and making subsequent analysis more accurate. By using intersection calculation and determination of associated data, the internal relationship between tumor samples and normal samples can be fully considered, reducing the possibility of misjudgment and improving the accuracy of abnormal gene recognition. And by performing abnormal analysis with the initial associated data as a reference and calculating the target outlier, a clear standard and basis for abnormal judgment are obtained. This analysis method based on data comparison can more objectively reflect the abnormal situation of gene data and enhance the reliability of abnormal analysis.
[0060] In some embodiments, the obtaining of the target outlier corresponding to the first data by performing abnormal analysis on the first data according to the initial associated data includes: performing intersection calculation on the first associated data and the second associated data to obtain the associated intersection data corresponding to the first data and the second data; obtaining third data from the first segmented data and obtaining the third associated data corresponding to the third data from the second segmented data; determining the first intersection corresponding to the third associated data and the first data, and obtaining the first inverse correlation data corresponding to the first data according to the first intersection; determining the second intersection corresponding to the third associated data and the second data, and obtaining the second inverse correlation data corresponding to the second data according to the second intersection; performing intersection calculation according to the first inverse correlation data and the second inverse correlation data to determine the first shared data corresponding to the first data and the second data; performing intersection calculation according to the associated intersection data and the first shared data to determine the second shared data corresponding to the first data and the second data; performing intersection calculation according to the second shared data and the initial associated data to obtain the target associated data corresponding to the first data; and performing abnormal analysis on the first data according to the target associated data to obtain the target outlier corresponding to the first data.
[0061] Exemplarily, performing intersection calculation on the first associated data and the second associated data is to find out the gene information they commonly contain and thus obtain the associated intersection feature. For example, if the first associated data has genes A, B, C and the second associated data has genes B, C, D, then the associated intersection data obtained after intersection calculation is genes B and C. This associated intersection data reflects the common feature of the first data and the second data at the associated level.
[0062] Exemplarily, select any remaining data other than the first data and the second data from the multiple first segmentation data obtained previously as the third data. Then, find the one in the multiple second segmentation data that has a similar feature or corresponding relationship with the third data, and determine it as the third associated data. This process is similar to finding the first associated data and the second associated data before, and is to prepare for the subsequent calculation of inverse correlation data.
[0063] Exemplarily, determine the intersection of the third associated data and the first data, that is, judge whether the first data is in the third associated data corresponding to the third data to obtain the first intersection. If the first intersection is not empty, that is, the first data is in the third associated data corresponding to the third data, then determine the third data as the first inverse correlation data corresponding to the first data.
[0064] Using the same method, determine the intersection of the third associated data and the second data to obtain the second intersection. Then obtain the second inverse correlation data corresponding to the second data according to the second intersection.
[0065] Exemplarily, perform an intersection calculation on the first inverse correlation data and the second inverse correlation data. If the first inverse correlation data contains genes E, F, G, and the second inverse correlation data contains genes F, G, H, then the first shared data obtained after the intersection calculation is genes F and G. The first shared data reflects the common features of the first data and the second data at the inverse correlation level.
[0066] Exemplarily, perform an intersection calculation on the previously obtained associated intersection data and the first shared data. Assume that the associated intersection data is genes B and C, and the first shared data is genes F and G. If they have no common genes, then the intersection is an empty set; if there are common genes, then these common genes form the second shared data. The second shared data further synthesizes the common features at the associated and inverse correlation levels.
[0067] Exemplarily, perform an intersection calculation on the second shared data and the initial associated data, and the result obtained is the target associated data corresponding to the first data. The target associated data synthesizes information from multiple levels and can more comprehensively reflect the association between the first data and other data.
[0068] Exemplarily, taking the target associated data as a reference, perform an anomaly analysis on the first data. It can be compared from multiple aspects such as gene sequence, gene expression level, and gene function. For example, compare the expression quantity differences of some key genes in the first data and the target associated data, and then obtain the degree to which the first data deviates from the normal situation represented by the target associated data according to the expression difference value, that is, obtain the target outlier value.
[0069] Exemplarily, the target associated data is obtained through multiple steps of calculation and comprehensive analysis. It synthesizes information from the associated level and the inverse correlation level, and can better represent the normal situation related to the first data. Performing anomaly analysis with such comprehensive and reliable reference data makes the calculation of the target outlier more reliable, providing a solid foundation for subsequent work such as abnormal gene identification.
[0070] In some embodiments, the obtaining of the target outlier corresponding to the first data by performing anomaly analysis on the first data according to the target associated data includes: calculating the data distance between the first data and the target associated data, and calculating the first data volume corresponding to the second shared data; determining the first similarity between the first data and the target associated data according to the first data volume and the data distance; counting the second data volume corresponding to the target associated data, and determining the second similarity corresponding to the first data under a preset range according to the first similarity and the second data volume; performing the same steps on each sub-data in the target associated data and the first data to obtain the third similarity corresponding to each sub-data under the preset range; determining the target outlier corresponding to the first data according to the second similarity and the third similarity; wherein, the target outlier is obtained according to the following formula:
[0071]
[0072] wherein, sim 1ij represents the first similarity between the i-th first data and the j-th target associated data, count1 represents the first data volume, dis(num 1i , num 2j ) represents the data distance between the i-th first data num 1i and the j-th target associated data num 2j , α represents a constant, sim 2i represents the second similarity corresponding to the i-th first data under the preset range, count2 represents the second data volume, outlier i represents the target outlier corresponding to the i-th first data, sim 3j represents the third similarity corresponding to the j-th sub-data under the preset range, and the calculation methods of the third similarity and the second similarity are the same.
[0073] Exemplarily, for the first data and the target associated data, the degree of difference between them needs to be considered from multiple dimensions to calculate the corresponding data distance. For example, in gene data, the differences in aspects such as the base arrangement order of gene sequences and gene expression levels can be compared. Taking gene expression level as an example, the Euclidean distance, Manhattan distance, etc. between the numerical values of the gene expression amounts of the two can be calculated to obtain the data distance reflecting their difference magnitudes.
[0074] Exemplarily, count the number of data elements included in the second shared data, and this number is the first data volume. Then, based on the first data volume and the data distance calculated previously, determine the first similarity between the first data and the target associated data according to the following formula:
[0075]
[0076] where, sim 1ij represents the first similarity between the i-th first data and the j-th target associated data, count1 represents the first data volume, and dis(num 1i , num 2j ) represents the data distance between the i-th first data num 1i and the j-th target associated data num 2j , and α represents a constant.
[0077] Exemplarily, count the number of data elements included in the target associated data to obtain the second data volume, and then, in combination with the first similarity and the second data volume calculated previously, considering the influence of the preset range, determine the second similarity corresponding to the first data under the preset range. The preset range is also the local space. Thus, according to
[0078]
[0079] where, sim 1ij represents the first similarity between the i-th first data and the j-th target associated data, sim 2i represents the second similarity corresponding to the i-th first data under the preset range, and count2 represents the second data volume.
[0080] Exemplarily, for each sub-data in the target associated data, repeat the above steps of calculating the first similarity and the second similarity. That is, first calculate the data distance and the corresponding first data volume between each sub-data and the first data to determine the first similarity, and then, in combination with the second data volume corresponding to the sub-data, obtain the third similarity corresponding to each sub-data under the preset range, and further obtain the target outlier according to the following formula:
[0081]
[0082] where, sim 2i represents the second similarity corresponding to the i-th first data under a preset range, count2 represents the second data volume, and outlier i represents the target outlier corresponding to the i-th first data, and sim 3j represents the third similarity corresponding to the j-th sub-data under a preset range, and moreover, the calculation methods of the third similarity and the second similarity are the same.
[0083] Exemplarily, the above formula comprehensively considers the similarity between the first data and the overall target associated data as well as the similarity between the first data and each sub-data in the target associated data, and obtains a value that can reflect the deviation degree of the first data through specific mathematical operations.
[0084] Specifically, by calculating multiple similarities (the first similarity, the second similarity, the third similarity) and finally obtaining the target outlier, the deviation degree of the first data relative to the target associated data can be measured more precisely. This step-by-step and progressive analysis method helps to more accurately identify the true abnormal data and reduce the situations of misjudgment and missed judgment. Calculating the similarity for each sub-data in the target associated data fully considers the internal structure and diversity of the target associated data. Different sub-data may represent different biological meanings or data characteristics. By calculating the similarity with each sub-data, the relationship between the first data and the target associated data can be analyzed more meticulously, thereby more accurately judging the abnormal situation.
[0085] Step S104: Perform an association relationship analysis on the initial abnormal gene data according to the biological development information and the initial abnormal type to obtain a first association relationship between the initial abnormal gene data.
[0086] Exemplarily, according to the biological development information, association rules are formulated. For example, if biological research shows that certain genes have a cooperative effect in a specific biological process, then when these genes are abnormal simultaneously, it can be considered that there is an association between them. Another example is that certain gene mutations will cause abnormal expression of downstream genes, and association rules can be set based on this causal relationship.
[0087] Exemplarily, the association rules are refined considering the initial abnormal type. For example, for the situation of abnormal gene expression levels, it can be set that when the expression levels of two genes increase or decrease simultaneously and the change range is within a certain range, it is considered that there is an association between them; for the situation of gene mutations, the association between genes can be judged according to factors such as the position and type of the mutation sites.
[0088] Exemplarily, the initial abnormal gene data is matched with the biological development information and the initial abnormal type. For each abnormal gene, relevant information such as functions and regulatory relationships is searched for in the biological development information, and it is determined whether it conforms to the association rule in combination with the initial abnormal type. Through information matching, the association relationship between the initial abnormal gene data is determined. If two or more abnormal genes meet the association rule, it is considered that there is a first association relationship between them, such as causal association, cooperative association, regulatory association, etc.
[0089] In some embodiments, the obtaining of the first association relationship between the initial abnormal gene data by performing an association relationship analysis on the initial abnormal gene data according to the biological development information and the initial abnormal type includes: obtaining first development information corresponding to the initial abnormal gene from the biological development information, and determining a first control group according to the initial abnormal gene and the first development information; replacing the initial abnormal gene data in the first control group with the initial abnormal type to obtain a second control group; obtaining second development information from the second control group, and determining the relevant abnormal type corresponding to the second development information; performing frequent item mining according to the relevant initial abnormal type to obtain the associated abnormal type corresponding to the second development information; comparing the associated abnormal type with the initial abnormal type corresponding to the initial abnormal gene data to obtain the first association relationship between the initial abnormal gene data.
[0090] Exemplarily, a part corresponding to the initial abnormal gene is screened out from the biological development information to obtain the first development information, and then the initial abnormal gene and the corresponding first development information are combined together to form a first control group.
[0091] Exemplarily, the initial abnormal gene data in the first control group is replaced with the initial abnormal type, that is, the initial abnormal gene is replaced with the corresponding initial abnormal type, so as to obtain a second control group. This process is to integrate the information of the abnormal type into the combination of the original gene data and the development information for subsequent analysis of the association between abnormal types.
[0092] Exemplarily, second development information is obtained from the second control group, and the relevant abnormal types corresponding to all the second development information are obtained from the second control group.
[0093] Exemplarily, the method of frequent item mining is used to analyze the relevant initial abnormal types, and then the combination of abnormal types that often appear simultaneously in the data set is found. For example, in a large number of gene data, it is found that the two abnormal types of abnormal increase in gene expression level and gene mutation often appear simultaneously, then these two abnormal types may constitute an associated abnormal type. Through this mining, the potential association relationship between abnormal types under the second development information can be found.
[0094] Exemplarily, the associated abnormal types obtained by mining are compared with the initial abnormal types corresponding to the initial abnormal gene data. Through comparative analysis, similarities, inclusion relationships, or other logical connections between them are found, so as to determine the first association relationship between the initial abnormal gene data. For example, if most of the abnormal conditions in the initial abnormal types are included in the associated abnormal types, it can be considered that there is a strong association relationship between these initial abnormal gene data.
[0095] Specifically, by combining biological development information and initial abnormal types for association relationship analysis, a variety of information is fully integrated, and the associated abnormal types are mined through frequent items, which can discover potential association patterns between abnormal types that are easily overlooked in conventional analysis. These association patterns may reveal the cooperative mechanism of gene abnormalities in biological processes, providing new clues for studying the pathogenesis, development process, etc. of diseases.
[0096] Step S105: Perform association relationship analysis on the initial abnormal gene data according to the clinical feature information to obtain the second association relationship between the initial abnormal gene data.
[0097] Exemplarily, according to the characteristics of the clinical feature information and the initial abnormal gene data, the dimensions of the association analysis are determined. For example, the analysis can be carried out from dimensions such as the association between gene expression levels and symptoms, and the association between gene mutations and disease diagnosis. Then, statistical analysis methods (such as correlation analysis, chi-square test, etc.) and machine learning methods (such as decision trees, random forests, etc.) are used to match the clinical feature information and the initial abnormal gene data to ensure that the clinical feature information and gene data of each target patient correspond to each other. Furthermore, the correlation coefficient between the clinical feature information and the initial abnormal gene data is calculated, and then according to the set significance threshold, genes and clinical features with significant associations are screened out.
[0098] Exemplarily, according to the selected significant associations, an association network between the initial abnormal gene data is constructed. In the network, nodes represent genes, and edges represent the association relationships between genes. The association relationships can be weighted according to the clinical feature information, such as the association strength, association type, etc., so as to analyze the association network and find the association patterns between genes. For example, it is found that certain genes form a tight association module under specific clinical features, and these modules may jointly participate in a certain biological process or disease mechanism. Then, according to the association network and the association patterns, the second association relationship between the initial abnormal gene data is determined. The second association relationship can be a direct association (a direct association is established between two genes through clinical features) or an indirect association (an association is established through other genes or clinical features).
[0099] In some embodiments, performing association analysis on the initial abnormal gene data according to the clinical feature information to obtain the second association relationship between the initial abnormal gene data includes: performing cluster analysis on the clinical feature information to obtain a target clustering result corresponding to the clinical feature information; obtaining, from the initial abnormal gene data, related abnormal genes corresponding to each cluster cluster in the target clustering result; performing frequent item mining according to the related abnormal genes to obtain associated abnormal genes, and obtaining the second association relationship between the initial abnormal gene data according to the associated abnormal genes.
[0100] For example, the clinical feature information should first be cleaned to remove erroneous data, duplicate data, and data with many missing values. The data should be standardized so that different types of clinical features (such as numerical blood pressure and heart rate, categorical symptom manifestations, etc.) are comparable. For example, the Z-score standardization method can be used for numerical data.
[0101] Exemplarily, hierarchical clustering or K-means clustering is used to perform clustering operations on the preprocessed clinical characteristic information to obtain target clustering results corresponding to the clinical characteristic information. Each cluster represents a group of samples with similar clinical characteristics.
[0102] Exemplarily, the initial abnormal gene data is matched with the target clustering result, that is, the abnormal gene data corresponding to the samples in each cluster is determined. The matching can be performed by the unique identifier of the sample (such as the patient number).
[0103] Exemplarily, for each cluster, the abnormal gene data of the samples therein are analyzed to find abnormal genes that appear frequently in the cluster or have a potential association with the clinical characteristics of the cluster, and these genes are determined as relevant abnormal genes corresponding to the cluster.
[0104] Exemplarily, the related abnormal genes corresponding to each cluster are regarded as a transaction, and the related abnormal gene sets of all clusters constitute a transaction data set, and then frequent item sets are mined using frequent item mining algorithms such as Apriori algorithm and FP-growth algorithm, so that the abnormal genes corresponding to the frequent item sets are determined as associated abnormal genes.
[0105] Exemplarily, the initial abnormal gene data is used as nodes, and the association relationship between the associated abnormal genes is used as edges to construct an association relationship graph. The second association relationship between the initial abnormal gene data can be intuitively displayed through the association relationship graph.
[0106] Specifically, by performing cluster analysis on clinical feature information, different clinical feature combination patterns can be discovered, reflecting the heterogeneity of the disease. Further identifying the relevant abnormal genes and associated abnormal genes corresponding to each cluster helps to deeply understand the gene characteristics of different subtypes of the disease.
[0107] Step S106: Establish a gene regulatory network for the initial abnormal gene data, and perform association relationship analysis on the initial abnormal gene data according to the gene regulatory network to obtain the third association relationship among the initial abnormal genes.
[0108] Exemplarily, initial abnormal gene data are obtained, which usually include gene expression levels, gene sequences, gene mutation information, etc. At the same time, relevant biological background information, such as gene function annotations, signal pathway information, etc., also needs to be collected to provide support for subsequent construction of the gene regulatory network.
[0109] Exemplarily, algorithms such as Pearson correlation coefficient and Spearman correlation coefficient are used to calculate the correlation coefficients between the initial abnormal gene data, and then the regulatory relationships between genes are determined according to the magnitudes of the correlation coefficients. Gene pairs with higher correlations may have direct or indirect regulatory relationships. Thus, the initial abnormal genes are used as nodes of the network, the edges represent the regulatory relationships between genes, the direction of the edges represents the direction of regulation (such as positive regulation or negative regulation), and the weights of the edges can represent the intensity of regulation, that is, the correlation coefficients between the initial abnormal gene data.
[0110] Exemplarily, regulatory paths between genes are searched in the gene regulatory network. If there is one or more regulatory paths between two genes, it indicates that there is an association relationship between them. Methods such as the shortest path algorithm can be used to find the shortest regulatory paths between genes, and then the gene regulatory network is divided into modules, dividing the network into multiple functional modules. Genes within each module usually have similar functions or participate in the same biological process. Analyze the distribution of the initial abnormal genes in different modules and the connection relationships between the modules to determine the third association relationship between genes.
[0111] In some embodiments, the obtaining of the third association relationship between the initial abnormal gene data according to the gene regulatory network includes: determining current abnormal gene data from the initial abnormal gene data, and removing the current abnormal gene data from the initial abnormal gene data to obtain remaining abnormal gene data; obtaining first path information from the remaining abnormal gene data to the current abnormal gene data according to the gene regulatory network; obtaining first similar abnormal gene data corresponding to the current abnormal gene data from the remaining abnormal gene data according to the first path information; analyzing the first similar abnormal gene data according to the gene regulatory network to obtain second path information corresponding to each sub-gene data in the first similar abnormal gene data; determining second similar abnormal gene data corresponding to each sub-gene data in the first similar abnormal gene data according to the second path information; and performing an association relationship analysis on the current abnormal gene data and the sub-gene data according to the first similar abnormal gene data and the second similar abnormal gene data to obtain the third association relationship.
[0112] Exemplarily, a gene data is randomly selected from the initial abnormal gene data as the current abnormal gene data, and then the determined current abnormal gene data is removed from the initial abnormal gene data, and the remaining gene data is the remaining abnormal gene data.
[0113] Exemplarily, based on the constructed gene regulatory network, a graph search algorithm (such as breadth-first search, depth-first search, etc.) is used to find the regulatory paths from each gene in the remaining abnormal gene data to the current abnormal gene data. Here, the paths represent the regulatory connections between genes, which may pass through multiple intermediate genes, and then the relevant information of these found paths is sorted out to form the first path information, including but not limited to the path length, the intermediate genes passed through, etc.
[0114] Exemplarily, according to the first path information, screening conditions are set, such as the path length is less than a certain threshold, the regulatory intensity on the path is greater than a certain value, etc., so as to screen out the gene data that meet the conditions from the remaining abnormal gene data according to the set criteria, and these gene data are the first similar abnormal gene data corresponding to the current abnormal gene data.
[0115] Exemplarily, for each sub-gene data in the first similar abnormal gene data, the gene regulatory network is used again, and a graph search algorithm is used to find the regulatory paths from other genes (here can be the genes in the initial abnormal gene data except the current abnormal gene data and the sub-gene data) to the sub-gene data, and then the relevant information of these paths is summarized to form the second path information corresponding to each sub-gene data.
[0116] Exemplarily, similar to screening the first similar abnormal gene data, appropriate screening criteria are set according to the second path information, and gene data similar to each sub-gene data are screened out from the relevant gene data, that is, the second similar abnormal gene data.
[0117] Exemplarily, the first similar abnormal gene data and the second similar abnormal gene data are combined to analyze the association between the current abnormal gene data and each sub-gene data. Consider factors such as the position of the gene in the regulatory network, the characteristics of the path (such as path length, regulatory direction, etc.), and the function of the gene. Thus, according to the results of the comprehensive analysis, the type of association relationship (such as direct association, indirect association, etc.) and the strength of the association (such as strong association, weak association, etc.) between the current abnormal gene data and the sub-gene data are determined, so as to obtain the third association relationship.
[0118] Specifically, by gradually screening similar abnormal gene data and analyzing path information, the association relationship between the initial abnormal genes can be mined more accurately. Genes that seemingly have no direct correlation but have an indirect association through the regulatory network can be discovered, which helps to deeply understand the complex interactions between genes.
[0119] Step S107: Based on hierarchical constraints, identify the genes closely related to tumor occurrence and development in the tumor sample by combining the first association relationship, the second association relationship, and the third association relationship, and obtain the target tumor driver gene data as the bioinformatics result.
[0120] Exemplarily, the first association relationship, the second association relationship, and the third association relationship obtained from the previous analysis are collected and summarized. These association relationships contain the interaction information between genes and are important bases for identifying tumor driver genes.
[0121] Exemplarily, according to biological knowledge, constraint conditions are set at different biological levels. For example, at the gene level, the functional classification of genes can be considered, such as genes related to signal pathways, transcription factor genes, etc.; at the cell level, the role of genes in processes such as the cell cycle and apoptosis is considered; at the tissue level, the characteristics of the tumor tissue are combined, such as the grade and stage of the tumor.
[0122] Exemplarily, constraint conditions are set according to the strength of the association relationship. For the first association relationship, the second association relationship, and the third association relationship, different association strength thresholds can be set respectively, and only the association relationships that meet a certain strength can be included in the subsequent analysis. For example, a correlation coefficient threshold between genes or a confidence threshold for regulatory relationships is set.
[0123] Exemplarily, according to the set hierarchical constraint rules, the genes in the tumor sample are preliminarily screened. From the perspective of the biological level, the genes involved in the biological processes closely related to tumorigenesis and development are screened out; from the perspective of the association strength, the genes that meet the association strength threshold are screened out. Thus, on the basis of the preliminary screening, iterative screening is carried out. Combining the constraint conditions at different levels, the gene range is gradually narrowed. For example, first, the genes related to signal pathways are screened out at the gene function level, and then combined with the association strength constraint, the genes with a higher association strength with other genes are further screened out. Furthermore, the screened genes are comprehensively analyzed in combination with the first association relationship, the second association relationship, and the third association relationship. Considering the direct and indirect associations between genes, a gene association network is constructed to intuitively display the interactions between genes, and then the association patterns between genes are mined, such as the cooperative action pattern and the upstream and downstream regulation pattern between genes. By analyzing the association patterns, the gene combinations that play key roles in the process of tumorigenesis and development are found.
[0124] Exemplarily, according to the results of the comprehensive association relationship analysis, the genes that are in the core position in the gene association network, have strong associations with multiple other genes, and meet the biological level constraints are identified. These genes are very likely to be the key genes closely related to tumorigenesis and development. Thus, the identified key genes are sorted into the target tumor driver gene data, and the driver oncogenes corresponding to the tumor sample are quickly and accurately identified as the bioinformatics results. Furthermore, according to the bioinformatics results, the key directions for the next detection are intelligently prompted, reducing the total time required for the gene detection link and winning more time for the patients.
[0125] Specifically, by analyzing through hierarchical constraints combined with multiple association relationships, the genes can be screened and evaluated from multiple perspectives, avoiding the limitations of a single method or a single association relationship analysis. Thus, the genes closely related to tumorigenesis and development can be more accurately identified. Considering the biological information at different levels and multiple gene association relationships comprehensively helps to deeply reveal the molecular mechanism of tumorigenesis and development. Discover the complex interactions and regulatory networks between genes, and thus quickly and accurately identify the driver oncogene data.
[0126] In addition, after the present application obtains the corresponding initial abnormal gene data in the tumor sample and the initial abnormal type corresponding to the initial abnormal gene data by identifying abnormal genes based on the first genomic data and the second genomic data, the present application also uses an online bioinformatics analysis tool (such as GEPIA2, etc.) to upload the first genomic data and the second genomic data for differential expression analysis. Genes that are significantly upregulated or downregulated in lung adenocarcinoma tissues are identified, the functions and related pathways of the differentially expressed genes are analyzed. If it is found that certain genes are related to immune regulation, the next step is to focus on immune-related verification and analysis, so as to perform correlation analysis between the obtained differentially expressed gene results and the laboratory indicators of the target patient. For example, check whether there is a correlation between the expression levels of certain genes and the content of tumor markers. If it is found that certain genes are closely related to laboratory indicators, the next step can focus on studying the specific mechanism of action of these genes in tumorigenesis and development, and then continue to perform correlation analysis on the initial abnormal gene data based on biological development information and the initial abnormal type to obtain the first correlation relationship between the initial abnormal gene data; perform correlation analysis on the initial abnormal gene data according to the clinical characteristic information to obtain the second correlation relationship between the initial abnormal gene data; establish a gene regulatory network for the initial abnormal gene data, and perform correlation analysis on the initial abnormal gene data according to the gene regulatory network to obtain the third correlation relationship between the initial abnormal genes; identify genes closely related to tumorigenesis and development in the tumor sample based on hierarchical constraints combined with the first correlation relationship, the second correlation relationship and the third correlation relationship to obtain the target tumor driver gene data.
[0127] If it is found that certain genes are not closely related to laboratory indicators, the immune single-cell center TISCH can be logged in to search for single-cell sequencing data of lung adenocarcinoma. Select data similar to the previous analysis sample type and research direction, so as to compare the gene expression in the TISCH single-cell sequencing data with the initial abnormal gene data. Check whether the expression patterns of these genes are consistent at the single-cell level. If the expression patterns of most genes are consistent, further focus on the gene subset with special expression patterns at the single-cell level as the key point of the next step of research.
[0128] If it is found that the expression patterns of these genes are inconsistent at the single-cell level, query the immunohistochemistry data, input the names of the previously screened initial abnormal gene data, and query the immunohistochemistry images and expression data of these genes in lung tissues. Compare the immunohistochemistry results with the previous RNA-seq and single-cell sequencing results to verify the expression consistency of the genes at the RNA level and the protein level. Based on all verification results, determine genes or biomarkers with important biological significance and clinical application potential. Provide a basis for the selection of the treatment plan and prognosis assessment of the patient.
[0129] Therefore, in the above step-by-step biological information sequencing process, the results of each step should be summarized and analyzed in a timely manner, and the research focus of the next step should be flexibly adjusted according to the results to improve the research efficiency and reduce the total time-consuming of the detection link.
[0130] Please refer to Figure 2 , Figure 2 A tumor driver gene bioinformatics system 200 based on hierarchical constraint automatic learning provided by an embodiment of the present application. The tumor driver gene bioinformatics system 200 based on hierarchical constraint automatic learning includes a data acquisition module 201, a data determination module 202, an anomaly recognition module 203, a first analysis module 204, a second analysis module 205, a third analysis module 206, and a gene recognition module 207. Among them, the data acquisition module 201 is used to obtain first genomic data corresponding to a tumor sample of a target patient and second genomic data corresponding to a normal tissue sample of the target patient from a database; the data determination module 202 is used to determine biological development information and clinical feature information corresponding to the tumor sample from the database; the anomaly recognition module 203 is used to perform anomaly gene recognition based on the first genomic data and the second genomic data to obtain initial anomaly gene data corresponding to the tumor sample and an initial anomaly type corresponding to the initial anomaly gene data; the first analysis module 204 is used to perform correlation analysis on the initial anomaly gene data according to the biological development information and the initial anomaly type to obtain a first correlation relationship between the initial anomaly gene data; the second analysis module 205 is used to perform correlation analysis on the initial anomaly gene data according to the clinical feature information to obtain a second correlation relationship between the initial anomaly gene data; the third analysis module 206 is used to establish a gene regulatory network for the initial anomaly gene data and perform correlation analysis on the initial anomaly gene data according to the gene regulatory network to obtain a third correlation relationship between the initial anomaly genes; the gene recognition module 207 is used to identify genes closely related to tumor occurrence and development in the tumor sample based on hierarchical constraints in combination with the first correlation relationship, the second correlation relationship, and the third correlation relationship to obtain target tumor driver gene data as a bioinformatics result.
[0131] In some embodiments, the tumor driver gene bioinformatics system 200 based on hierarchical constraint automatic learning can be applied to a terminal device.
[0132] It should be noted that those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working process of the above-described tumor driver gene bioinformatics system 200 based on hierarchical constraint automatic learning can refer to the corresponding process in the foregoing embodiment of the tumor driver gene bioinformatics method based on hierarchical constraint automatic learning, and will not be described herein again.
[0133] Please refer to Figure 3 , Figure 3 which is a schematic block diagram of a terminal device provided by an embodiment of the present invention.
[0134] As Figure 3 shown, the terminal device 300 includes a processor 301 and a memory 302, and the processor 301 and the memory 302 are connected through a bus 303, which is, for example, an I2C (Inter - integrated Circuit) bus.
[0135] Specifically, the processor 301 is used to provide computing and control capabilities to support the operation of the entire terminal device. The processor 301 can be a central processing unit (CPU), and the processor 301 can also be other general - purpose processors, digital signal processors (DSPs), application - specific integrated circuits (ASICs), field - programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general - purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.
[0136] Specifically, the memory 302 can be a Flash chip, read - only memory (ROM), magnetic disk, optical disc, USB flash drive, or mobile hard disk, etc.
[0137] Those skilled in the art can understand that Figure 3 the structure shown in
[0138] is only a block diagram of a part of the structure related to the solution of the embodiment of the present invention, and does not constitute a limitation on the terminal device to which the solution of the embodiment of the present invention is applied. The specific server may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0139] In one embodiment, the processor is used to run a computer program stored in the memory, and when executing the computer program, the following steps are implemented:
[0140] Obtain the first genomic data corresponding to the tumor sample of the target patient and the second genomic data corresponding to the normal tissue sample of the target patient from the database;
[0141] Determine the biological development information and clinical characteristic information corresponding to the tumor sample from the database;
[0142] Perform abnormal gene identification based on the first genomic data and the second genomic data to obtain the corresponding initial abnormal gene data in the tumor sample and the initial abnormal type corresponding to the initial abnormal gene data;
[0143] Perform correlation analysis on the initial abnormal gene data according to the biological development information and the initial abnormal type to obtain the first correlation relationship between the initial abnormal gene data;
[0144] Perform correlation analysis on the initial abnormal gene data according to the clinical characteristic information to obtain the second correlation relationship between the initial abnormal gene data;
[0145] Establish a gene regulatory network for the initial abnormal gene data, and perform correlation analysis on the initial abnormal gene data according to the gene regulatory network to obtain the third correlation relationship between the initial abnormal genes;
[0146] Based on hierarchical constraints, combine the first correlation relationship, the second correlation relationship, and the third correlation relationship to identify the genes closely related to tumorigenesis and development in the tumor sample, and obtain the target tumor driver gene data as the bioinformatics result.
[0147] It should be noted that those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process of the above-described terminal device can refer to the corresponding process in the embodiment of the tumor driver gene bioinformatics method based on hierarchical constraint automatic learning, which will not be elaborated here.
[0148] The embodiment of the present invention also provides a storage medium for computer-readable storage. The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of any one of the tumor driver gene bioinformatics methods provided in the specification of the embodiment of the present invention.
[0149] Among them, the storage medium may be the internal storage unit of the terminal device described in the foregoing embodiments, such as the hard disk or memory of the terminal device. The storage medium may also be an external storage device of the terminal device, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the terminal device.
[0150] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations. In the hardware embodiments, the division between the functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, one physical component may have multiple functions, or one function or step may be executed by several physical components in cooperation. Some or all of the physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, communication media typically contains computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.
[0151] It should be understood that the term "and / or" used in the specification and appended claims of the present invention refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations. It should be noted that in this text, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or system comprising a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or system. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of another identical element in the process, method, article or system comprising the element.
[0152] The serial numbers of the above embodiments of the present invention are only for description and do not represent the superiority or inferiority of the embodiments. The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A tumor driver gene bioinformatics method based on hierarchical constraint automatic learning, characterized in that The method includes: Obtaining first genomic data corresponding to a tumor sample of a target patient and second genomic data corresponding to a normal tissue sample of the target patient from a database; Determining biological development information and clinical characteristic information corresponding to the tumor sample from the database; Performing abnormal gene identification based on the first genomic data and the second genomic data to obtain initial abnormal gene data corresponding to the tumor sample and an initial abnormal type corresponding to the initial abnormal gene data; Performing correlation analysis on the initial abnormal gene data according to the biological development information and the initial abnormal type to obtain a first correlation relationship between the initial abnormal gene data; Performing correlation analysis on the initial abnormal gene data according to the clinical characteristic information to obtain a second correlation relationship between the initial abnormal gene data; Establishing a gene regulatory network for the initial abnormal gene data, and performing correlation analysis on the initial abnormal gene data according to the gene regulatory network to obtain a third correlation relationship between the initial abnormal genes; Based on hierarchical constraints, combining the first correlation relationship, the second correlation relationship, and the third correlation relationship to identify genes closely related to tumor occurrence and development in the tumor sample, and obtaining target tumor driver gene data as a bioinformatics result.
2. The method according to claim 1, wherein The performing abnormal gene identification based on the first genomic data and the second genomic data to obtain initial abnormal gene data corresponding to the tumor sample and an initial abnormal type corresponding to the initial abnormal gene data includes: Determining a target window, and performing data segmentation on the first genomic data according to the target window to obtain a plurality of first segmentation data; Performing data segmentation on the second genomic data according to the target window to obtain a plurality of second segmentation data; Obtaining any two data from the first segmentation data and determining them as first data and second data, and obtaining first associated data corresponding to the first data and second associated data corresponding to the second data from the second segmentation data; Performing an intersection calculation on the first associated data and the second data to obtain a first result, and performing an intersection calculation on the second associated data and the first data to obtain a second result; When the first result is not an empty set and the second result is not an empty set, determining the second data as the initial associated data corresponding to the first data; Performing abnormal analysis on the first data according to the initial associated data to obtain a target outlier value corresponding to the first data; Determining the initial abnormal gene data corresponding to the tumor sample according to the target outlier value; Obtaining associated gene data corresponding to the initial abnormal gene data from the second genomic data, and determining the initial abnormal type corresponding to the initial abnormal gene data according to the associated gene data.
3. The method according to claim 2, characterized in that The performing abnormal analysis on the first data according to the initial associated data to obtain a target outlier value corresponding to the first data includes: Performing an intersection calculation based on the first associated data and the second associated data to obtain the associated intersection data corresponding to the first data and the second data; Obtaining third data from the first segmentation data and obtaining third associated data corresponding to the third data from the second segmentation data; Determining a first intersection between the third associated data and the first data, and obtaining first inverse correlation data corresponding to the first data according to the first intersection; Determining a second intersection between the third associated data and the second data, and obtaining second inverse correlation data corresponding to the second data according to the second intersection; Performing an intersection calculation based on the first inverse correlation data and the second inverse correlation data to determine first shared data corresponding to the first data and the second data; Performing an intersection calculation based on the associated intersection data and the first shared data to determine second shared data corresponding to the first data and the second data; Performing an intersection calculation based on the second shared data and the initial associated data to obtain target associated data corresponding to the first data; Performing anomaly analysis on the first data according to the target associated data to obtain the target outlier value corresponding to the first data.
4. The method according to claim 3, wherein The performing anomaly analysis on the first data according to the target associated data to obtain the target outlier value corresponding to the first data includes: Calculating the data distance between the first data and the target associated data, and calculating the first data volume corresponding to the second shared data; Determining a first similarity between the first data and the target associated data according to the first data volume and the data distance; Counting the second data volume corresponding to the target associated data, and determining a second similarity corresponding to the first data within a preset range according to the first similarity and the second data volume; Performing the same steps on each sub-data in the target associated data and the first data to obtain a third similarity corresponding to each sub-data within a preset range; Determining the target outlier value corresponding to the first data according to the second similarity and the third similarity.
5. The method according to claim 1, characterized in that, The performing an association relationship analysis on the initial abnormal gene data according to the biological development information and the initial abnormal type to obtain a first association relationship between the initial abnormal gene data includes: Obtaining first development information corresponding to the initial abnormal gene from the biological development information, and determining a first control group according to the initial abnormal gene and the first development information; Replacing the initial abnormal gene data in the first control group with the initial abnormal type to obtain a second control group; Obtaining second development information from the second control group, and determining the relevant abnormal type corresponding to the second development information; Performing frequent item mining according to the relevant initial abnormal type to obtain an associated abnormal type corresponding to the second development information; Comparing the associated abnormal type with the initial abnormal type corresponding to the initial abnormal gene data to obtain the first association relationship between the initial abnormal gene data.
6. The method according to claim 1, characterized in that Performing an association relationship analysis on the initial abnormal gene data according to the clinical feature information to obtain a second association relationship between the initial abnormal gene data includes: Performing a clustering analysis on the clinical feature information to obtain a target clustering result corresponding to the clinical feature information; Obtaining relevant abnormal genes corresponding to each clustering cluster in the target clustering result from the initial abnormal gene data; Performing frequent item mining according to the relevant abnormal genes to obtain associated abnormal genes, and obtaining the second association relationship between the initial abnormal gene data according to the associated abnormal genes.
7. The method according to claim 1, characterized in that, Performing an association relationship analysis on the initial abnormal gene data according to the gene regulatory network to obtain a third association relationship between the initial abnormal genes includes: Determining current abnormal gene data from the initial abnormal gene data, and removing the current abnormal gene data from the initial abnormal gene data to obtain remaining abnormal gene data; Obtaining first path information from the remaining abnormal gene data to the current abnormal gene data according to the gene regulatory network; Obtaining first similar abnormal gene data corresponding to the current abnormal gene data from the remaining abnormal gene data according to the first path information; Performing an analysis on the first similar abnormal gene data according to the gene regulatory network to obtain second path information corresponding to each sub-gene data in the first similar abnormal gene data; Determining second similar abnormal gene data corresponding to each sub-gene data in the first similar abnormal gene data according to the second path information; Performing an association relationship analysis on the current abnormal gene data and the sub-gene data according to the first similar abnormal gene data and the second similar abnormal gene data to obtain the third association relationship.
8. A tumor driver gene bioinformatics system based on hierarchical constraint automatic learning, characterized in that, Including: A data acquisition module, configured to acquire first genomic data corresponding to a tumor sample of a target patient and second genomic data corresponding to a normal tissue sample of the target patient from a database; A data determination module, configured to determine biological development information and clinical feature information corresponding to the tumor sample from the database; An abnormality recognition module, configured to perform abnormal gene recognition according to the first genomic data and the second genomic data to obtain initial abnormal gene data corresponding to the tumor sample and an initial abnormal type corresponding to the initial abnormal gene data; A first analysis module, configured to perform an association relationship analysis on the initial abnormal gene data according to the biological development information and the initial abnormal type to obtain a first association relationship between the initial abnormal gene data; A second analysis module, configured to perform an association relationship analysis on the initial abnormal gene data according to the clinical feature information to obtain a second association relationship between the initial abnormal gene data; A third analysis module, configured to establish a gene regulatory network for the initial abnormal gene data, and perform an association relationship analysis on the initial abnormal gene data according to the gene regulatory network to obtain a third association relationship between the initial abnormal genes; A gene recognition module, configured to recognize genes closely related to tumor occurrence and development in the tumor sample based on hierarchical constraints in combination with the first association relationship, the second association relationship, and the third association relationship, and obtain target tumor driver gene data as bioinformatics results.
9. A terminal device, characterized in that, The terminal device includes a processor and a memory; The memory is used for storing computer programs; The processor is configured to execute the computer programs and, when executing the computer programs, implement the tumor driver gene bioinformatics method based on hierarchical constraint automatic learning according to any one of claims 1 to 7.
10. A computer storage medium for computer storage, characterized in that, The computer storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the tumor driver gene bioinformatics method based on hierarchical constraint automatic learning according to any one of claims 1 to 7.