A method, system, electronic device, and medium for retrieving publicly available omics data
By extracting data from open source omics databases and standardizing it with biomedical knowledge graphs and large models, the problem of difficulty in screening and high learning cost of omics data retrieval is solved, and efficient and accurate omics data retrieval is achieved.
Patent Information
- Application Number
- CN202510600869.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-05-12
AI Technical Summary
Existing omics data retrieval programs face the problems of difficulty in screening, high learning costs and low search efficiency. Especially in multiple scattered databases, it is difficult for users to effectively match research needs and understand unstructured text introductions.
Omics data is extracted from open source omics database, split the annotated data through text segmentation algorithm, standardize and group it with biomedical knowledge graphs and large models, integrate the data using weighted mixed search methods, and decompose user problems into sub-problems through thinking graphs to assist users in generating interpretable search results.
Reduces the complexity of cross-border search, improves retrieval efficiency and accuracy, and users can quickly judge the appropriateness of retrieval data, reducing learning and time costs.
Smart Images

Figure CN120123532B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of bioinformatics processing, and in particular, to a method, a system, an electronic device, and a medium for retrieving omics data. Background Art
[0002] With the rapid development of the field of life science research, the quantity and variety of omics data (including genomics, transcriptomics, proteomics, etc.) have shown an explosive growth trend. This trend has greatly promoted the progress of biological research, but at the same time, it has also brought unprecedented challenges, especially in the management and retrieval of omics data. Existing omics data retrieval solutions usually rely on specific database platforms, such as NCBI GEO (Gene Expression Omnibus), ArrayExpress, etc. These platforms store a large amount of experimental data, and usually use Boolean logic operators such as AND, OR, NOT, etc. inside the platform to perform retrievals through various combinations of conditions.
[0003] However, with the expansion of the dataset scale, the amount of data samples in the database has become extremely large, and each sample is usually attached with a long description text. When a user tries to perform a retrieval through keywords, thousands of records may be obtained. For example, when searching for the keyword "breast cancer" in the GEO database, a large number of records containing various types of samples and experimental conditions may be returned. Therefore, the user has to add more conditions to further narrow down the search scope, which undoubtedly greatly increases the difficulty of screening suitable datasets. For users, how to effectively match their research needs with the information in the database has become a major problem. Even more complicated is that since omics data is scattered in multiple different databases, each database has its own unique data organization method and retrieval logic, which requires users to spend a lot of time learning the operation methods of each database before using these resources. And due to the continuous growth of data types, the learning cost of the database has gradually increased. At the same time, since it is difficult for users to fully understand what is in the database, they cannot directly determine the retrieval conditions during retrieval and need to perform multiple retrievals to obtain the list of omics data they need. In addition, the detailed introduction of omics data is usually recorded in unstructured text summaries, and the content of these summaries is provided by different data producers, with different formats and details, which results in that even if users can find potentially relevant datasets, they need to spend extra time reading and understanding the specific meaning of each record to determine whether the dataset truly meets their research needs.
[0004] In view of the existence of the above problems, it is necessary to develop a brand-new retrieval method to enable users to reduce the retrieval time and cost, and improve the accuracy and efficiency of retrieval. Summary of the Invention
[0005] An object of the present invention is to provide a method, system, electronic device, and medium for retrieving public omics data in view of the deficiencies of the prior art.
[0006] According to the first aspect of this specification, a method for retrieving public omics data is provided. The method includes:
[0007] Extract omics data from an open-source omics database, where the omics data includes omics ontology data and omics annotation data;
[0008] Split the omics annotation data into multiple text blocks based on a text segmentation algorithm to form an omics annotation text set;
[0009] Perform standard preprocessing on the omics ontology data, group samples according to the omics annotation data using a large model, extract key information, and generate multiple paragraphs of an omics ontology text set for describing the omics ontology data through the large model;
[0010] Integrate the omics annotation text set and the omics ontology text set to obtain an omics text collection containing multiple omics text blocks;
[0011] Through a weighted hybrid search method that combines biological weights, BM25 retrieval, and semantic retrieval, integrate the omics text blocks in the omics text collection with the nodes of the biomedical knowledge graph, and then integrate the omics data with the biomedical knowledge graph;
[0012] Use a mind map to decompose the user's complex problem into sub-problems represented by the relationships of the biomedical knowledge graph, and then use a large model to answer the sub-problems. Control the large model generation process through the mind map to achieve an interpretable generation result.
[0013] Furthermore, the construction of the omics ontology text set is specifically as follows: Perform standard preprocessing on the omics ontology data. According to the omics annotation data, use a large model to group samples according to experimental conditions and / or biological characteristics. Then, according to the grouping results, extract a gene set with significant changes between groups, and use the large model to convert the analysis results into natural language descriptions to generate multiple paragraphs of an omics ontology text set for describing the omics ontology data.
[0014] Furthermore, the expression of the weighted hybrid search method is as follows:
[0015]
[0016]
[0017]
[0018] Wherein, is the weighted mixed search score, is the semantic similarity score, is the keyword score, is the biological weight vector, is the biological weight set according to the type importance of the nodes in the biomedical knowledge graph, is the vector form of the omics text block and is the combined text of the node and its definition, is in vector form, means to calculate the cosine similarity of vectors as the semantic similarity score, means to calculate the keyword score through the BM25 algorithm, which is used to adjust the proportion of the semantic similarity score and the keyword score.
[0019] Furthermore, the user's graph query plan is optimized and designed as follows: The user inputs the original query, and the large model identifies the entities required by the user in the original query according to the knowledge graph, obtains the set of entity node labels related to the query, expands the entity table according to the entity type and relationship type in the knowledge graph, obtains all the entities on the shortest path between entities, and then converts the queried entities into sub-questions to obtain a list of sub-questions ; The large model organizes the queried entities and relationships into natural language form and returns them to the user to assist the user in judging whether the query plan is correct; The user judges whether the query result is appropriate and whether the sub-questions need to be optimized according to the requirements. If optimization is required, a new list of sub-questions is regenerated ; According to the entity nodes and the associated relationships between the nodes finally confirmed by the user, a set of node labels closely related to the original query is obtained.
[0020] Furthermore, the quality of sub-questions is evaluated by optimizing the scoring function. If the sub-question score is lower than the set threshold, the path is re-analyzed to obtain new sub-questions. The optimized scoring function expression is as follows:
[0021]
[0022] where is the score value calculated by the optimized scoring function, is the number of sub-questions, is the sub-question and the original query cosine similarity, is the degree of association between sub-questions, that is, whether the nodes corresponding to sub-questions can be connected to the same omics data, is initially 0 and is used to judge the newly generated list of sub-questions after user feedback The similarity with the original sub-question list , that is The number of omics data shared with ; , and are weight parameters, which are dynamically adjusted according to user feedback.
[0023] Furthermore, after obtaining the node label set closely associated with the original query, calculate the Jaccard similarity between the node label set and each node label corresponding to each omics data associated therewith, generate an omics data list according to the Jaccard similarity ranking, analyze the omics text set corresponding to the omics data, and the large model explains the reason for selecting specific omics data in natural language form and elaborates the reasons for the selection in the returned results; the large model comprehensively summarizes the content of the selected omics data so that users can quickly judge whether the omics data meets the requirements.
[0024] According to the second aspect of this specification, there is provided an open omics data retrieval system, which is used to implement the above-mentioned open omics data retrieval method. The system includes:
[0025] An omics data information extraction module, which is used to extract omics data from an open-source omics database. The omics data includes omics ontology data and omics annotation data; split the omics annotation data into multiple text blocks based on a text segmentation algorithm to form an omics annotation text set; perform standardized preprocessing on the omics ontology data, group samples using the large model according to the omics annotation data, extract key information, and generate an omics ontology text set consisting of multiple segments for describing the omics ontology data through the large model; integrate the omics annotation text set and the omics ontology text set to obtain an omics text set containing multiple omics text blocks;
[0026] An omics data integrated biomedical knowledge graph module, which is used to integrate the omics text blocks in the omics text set with the nodes of the biomedical knowledge graph through a weighted hybrid search method combining biological weights, BM25 retrieval, and semantic retrieval, and then integrate the omics data with the biomedical knowledge graph;
[0027] A large model reasoning module based on a mind map, which uses the mind map to decompose the complex problems of users into sub-problems represented by the relationships of the biomedical knowledge graph, and then uses the large model to answer the sub-problems, and controls the large model generation process through the mind map to achieve an interpretable generation result.
[0028] According to a third aspect of this specification, there is provided an electronic device, including a memory and a processor, the memory being coupled to the processor; wherein, the memory is used for storing program data, and the processor is used for executing the program data to implement the above-mentioned public omics data retrieval method.
[0029] According to a fourth aspect of this specification, there is provided a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned public omics data retrieval method is implemented.
[0030] According to a fifth aspect of this specification, there is provided a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the above-mentioned public omics data retrieval method is implemented.
[0031] Compared with the prior art, the beneficial effects of the present invention are:
[0032] 1. Aiming at the problems of high screening difficulty and high learning cost, the present invention obtains descriptive information from open-source omics databases, fuses descriptions of multiple databases, and reduces the problem of cross-database retrieval. The present invention integrates omics data and knowledge graphs through a large model, uses multiple search logics based on the knowledge graph, and automatically corresponds to the omics data page.
[0033] 2. Aiming at the problem of low retrieval efficiency, the present invention uses mind maps to assist users in retrieving, decomposes and expands users' questions, uses a large model as an assistant to assist users in thinking about the retrieval process, and improves retrieval efficiency. Based on the large model, the present invention explains the reason for selecting a certain omics data in the form of natural language, and comprehensively summarizes the content of the omics data page, so that users can quickly judge whether the retrieved omics data meets their needs. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the following-described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0035] Figure 1 It is a flowchart of the public omics data retrieval method shown in an exemplary embodiment;
[0036] Figure 2 It is a flowchart of integrating biomedical knowledge graphs of omics data shown in an exemplary embodiment;
[0037] Figure 3 It is a flowchart of large model reasoning based on a mind map shown in an exemplary embodiment;
[0038] Figure 4 Structural diagram of the disclosed omics data retrieval system shown for an exemplary embodiment;
[0039] Figure 5 Structural diagram of an electronic device shown for an exemplary embodiment. Detailed implementation manners
[0040] For a better understanding of the technical solutions of this application, the embodiments of this application will be described in detail below with reference to the accompanying drawings.
[0041] It should be clear that the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts belong to the scope of protection of this application.
[0042] The terms used in the embodiments of this application are only for the purpose of describing specific embodiments, and are not intended to limit this application. The singular forms "a", "the" and "said" used in the embodiments of this application and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise.
[0043] The present invention provides an open omics data retrieval method, as Figure 1 shown, the method includes parts such as omics data information extraction, omics data integration into a biomedical knowledge graph, and large model reasoning based on a mind map. The specific implementation processes of each part will be elaborated in detail below.
[0044] S1. Omics data information extraction.
[0045] Extract information from semi-structured open-source omics databases (such as GEO database, TCGA database). The selected types of omics data include genomic data, transcriptomic data, and proteomic data. The extracted omics data s includes omics ontology data and omics annotation data . The omics ontology data is mainly sequencing information, mainly represented by sequences or numbers, and cannot be directly used for subsequent large model reasoning. Therefore, the present invention adopts a text alignment method for multi-omics data.
[0046] The omics annotation data is the descriptive information of omics data, mainly represented by unstructured multi-segment long texts. The present invention uses a recursive segmentation algorithm to split the original omics annotation data into multiple text blocks to form an omics annotation text set , expressed as , where is the text segmentation algorithm, and the number of text blocks is not fixed in different texts.
[0047] Omics ontology data Adopt unified standardized preprocessing, with genes as the core, and extract key information from omics data. First, for the omics ontology data perform standardized preprocessing to generate standard data . Subsequently, according to the omics annotation data , use a large model to intelligently group samples according to experimental conditions (such as before / after treatment) and / or biological characteristics (such as disease status). Then, according to the grouping, extract the gene set G with significant changes between groups through an algorithm for finding significant changes in the biological field. Taking transcriptome data as an example, use the Limma algorithm to perform differential expression analysis on the samples and extract the gene set G. Finally, use a large model to convert the complex analysis results into natural language descriptions in the format of grouping information + analysis results, and generate a multi-segment omics ontology text set for describing the omics ontology data . For example: "After drug treatment, gene A has a high expression phenomenon compared with the untreated group."
[0048] The present invention integrates the omics annotation text set and the omics ontology text set , and converts the omics data s into an omics text collection , which contains multiple omics text blocks . Finally, through OpenAI's text embedding model text-embedding-3, each omics text block is converted into a vector form , and stored in the Weaviate vector database for subsequent retrieval.
[0049] S2. Integrate omics data with the biomedical knowledge graph
[0050] Through step S1, the omics ontology data information is obtained, but the associations between omics are lacking. For example, we know that there is an upregulation of gene A in the omics data of a certain drug, but we cannot directly obtain other drug information associated with gene A. Therefore, the present invention introduces the biomedical knowledge graph of the biomedical information ontology system, which is a comprehensive biomedical knowledge graph constructed based on biomedical text data and includes 26 node types such as diseases, drugs, genes, cells, treatment or prevention procedures, medical terms, laboratory procedures, etc. The present invention first integrates the omics text blocks with the nodes of the biomedical knowledge graph respectively , where is the set of nodes of the biomedical knowledge graph, and then integrates the omics data with the biomedical knowledge graph
[0051] Specifically, the present invention proposes a weighted hybrid search method to integrate the omics text blocks Integrate with nodes of the biomedical knowledge graph to obtain a set of node labels that are most similar to the omics text block . The weighted hybrid search method combines biological weights, BM25 retrieval, and semantic retrieval.
[0052] Use the biological weight vector to adjust the node scores to reflect the importance of each specific type of node. Each biological weight is set according to the type of entity nodes in the biomedical knowledge graph . Considering that drug and gene labels appear more frequently in the omics text block, initially the weight of the drug label is set to 5, the weight of the gene label is set to 10, and the weight of disease or other type labels is set to 3.
[0053] Semantic retrieval is achieved by calculating the semantic similarity score. Since the node labels themselves are relatively short compared to the omics text block, the present invention uses the combined text of the label and its definition when calculating the semantic similarity score to provide richer semantic information.
[0054] In summary, the weighted hybrid search score of the weighted hybrid search method can be expressed by the following formula:
[0055]
[0056]
[0057]
[0058] where represents calculating the cosine similarity of vectors as the semantic similarity score, is the normalized keyword score (calculated by the BM25 algorithm), is in vector form, is in vector form, and r is used to adjust the proportion of the semantic similarity score and the keyword score. The larger r is, the more it tends to the keyword score. Since the omics database contains a large number of proper nouns, r is set to 0.8 in this embodiment. During the implementation process, the biological weight vector W can also be optimized by combining experimental verification and user feedback.
[0059] Based on the integration result of the above omics text block and the nodes of the biomedical knowledge graph , all nodes related to the omics text set are obtained, forming a node label set containing n nodes . The basic process of integrating omics data with the biomedical knowledge graph is asFigure 2 as shown
[0060] S3, large model reasoning based on the mind map
[0061] The present invention uses a mind map (GoT) to decompose the user's complex problem into sub-problems represented by the relationships of a biomedical knowledge graph, and then uses a large model to answer the sub-problems. By using the mind map, the generation process of the large model can be controlled to achieve more interpretable generation results. As Figure 3 shown, the main process of this step is as follows
[0062] 3.1 Graph query plan design optimization: The user inputs the original query Q, and the large model identifies the entities required by the user in the original query Q based on the knowledge graph, and obtains the entity node label set related to the query Q . Subsequently, the entity table is expanded according to the entity type and relationship type of the knowledge graph, and all entities on the shortest path between entities are obtained using the A* shortest path algorithm. Then, the queried entities are transformed into sub-problems, that is, knowledge graph triples, which correspond to the entities and relationships of the knowledge graph, and the returned result is a list of triple sub-problems in json format , is the number of sub-problems, and each sub-problem is a triple. The present invention uses the large model to organize the queried entities and relationships into natural language form and returns them to the user as output to assist the user in judging whether the query plan is correct. The user makes a choice according to their own needs, judges whether the query result is appropriate and whether the sub-problems need to be optimized. If optimization is required, a new list of sub-problems is regenerated , and the above process is repeated
[0063] The present invention sets an optimization scoring function to evaluate the quality of sub-problems. If the score of a sub-problem is low, the path is re-analyzed to obtain new sub-problems. The expression of the optimization scoring function is as follows
[0064]
[0065] wherein is the score value calculated by the optimization scoring function is the cosine similarity between the sub-problem and the original query Q is the degree of association between sub-problems, that is, whether the nodes corresponding to the sub-problems can be connected to the same omics data. The present invention defines it as the number of omics data shared by all sub-problems divided by the number of all omics data involved in the sub-problems is initially 0 and is used to judge whether the newly generated list of sub-problems is the same as the original list of sub-problems after user feedback has a high similarity, that is, the newly generated sub-question list and the original sub-question list The number of omics data shared. If and has a high similarity, the score is reduced. 、 and are weight parameters, which are dynamically adjusted according to user feedback. According to the entity nodes and the association relationships between nodes finally confirmed by the user, a set of node labels closely associated with the query Q is obtained .
[0066] 3.2 Obtaining the omics data page: Obtain all the omics data s associated with the set of node labels as alternatives for the return result. Subsequently, calculate the Jaccard similarity score between the set of node labels and each node label corresponding to each associated omics data s, and sort the return results according to the Jaccard similarity score. The calculation method is as follows:
[0067]
[0068] where represents the Jaccard similarity.
[0069] Specifically, first generate an ordered list of omics data pages based on the Jaccard similarity score. Subsequently, by analyzing the set of omics text associated with all the omics data s related to the set of node labels , the large model will explain in natural language the reason for selecting a specific omics data s and elaborate on the reasons for the selection in the returned results. Next, the large model will comprehensively summarize the content of the selected omics data s so that users can quickly judge whether the omics data meets the requirements, and can also provide the original link of the omics data s for users to directly access and view. Finally, users can decide whether this query result meets the requirements based on the provided summary; if not satisfied, users can choose to query again to obtain more matching results.
[0070] The following gives an inference case of a large model based on a mind map:
[0071] Prompt: You are an expert in biomedical knowledge graphs, and you are good at splitting questions into multiple instructions according to the knowledge graph. You will receive a query, and this query needs to be split into the specific relationships between nodes in the knowledge graph.
[0072] The following is the query you received: Find RNA sequencing data related to the action of drugs for rheumatoid arthritis.
[0073] The biomedical knowledge graph contains the following node types: {Chemical or Drug, Gene or Genome, Therapeutic or Preventive Procedure, Disease, Syndrome or Pathologic Function, Anatomical Structure}. The biomedical knowledge graph contains the following relationship types: {is a, reverse is a, interacts with, may be treated by, may treat, has adverse effect, is adverseeffect of, may be diagnosed by, may diagnose, contraindication, may cause, may becaused by}. Please do not mention relationships not included in the knowledge graph.
[0074] You need to return the answer in JSON format. The following is an example of the return:
[0075] Query: Which drugs act on psoriasis?
[0076] Return instruction: [{'Node or Node Type: psoriasis,'relation':' may betreatedby', 'Node or Node Type': 'Drug'}].
[0077] Answer from the large model:
[0078] [{"Node or Node Type": "Rheumatoid Arthritis",
[0079] "relation": "may be treated by",
[0080] "Node or Node Type": "Chemical or Drug"},
[0081] {"Node or Node Type": "Gene or Genome",
[0082] "relation": "may be diagnosed by",
[0083] "Node or Node Type": "RNA Sequencing Data"}]。
[0084] On the other hand, the embodiments of the present invention provide an open omics data retrieval system for implementing the above-mentioned open omics data retrieval method, such as Figure 4 shown, the system includes:
[0085] An omics data information extraction module, which is used to extract omics data from an open-source omics database, where the omics data includes omics ontology data and omics annotation data; split the omics annotation data into multiple text blocks based on a text segmentation algorithm to form an omics annotation text set; perform standardized preprocessing on the omics ontology data, group samples using a large model according to the omics annotation data, extract key information, and generate multiple segments of omics ontology text sets for describing the omics ontology data through the large model; integrate the omics annotation text set and the omics ontology text set to obtain an omics text collection containing multiple omics text blocks;
[0086] An omics data integrated biomedical knowledge graph module, which is used to integrate the omics text blocks in the omics text collection with the nodes of the biomedical knowledge graph through a weighted hybrid search method that combines biological weights, BM25 retrieval, and semantic retrieval, and then integrate the omics data with the biomedical knowledge graph;
[0087] A large model inference module based on a mind map uses the mind map to decompose the user's complex problem into sub-problems represented by the relationships of the biomedical knowledge graph, and then uses the large model to answer the sub-problems, and controls the large model generation process through the mind map to achieve an interpretable generation result.
[0088] Regarding the system in the above embodiments, the specific ways in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0089] For the system embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial descriptions of the method embodiments. The system embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present application. Those of ordinary skill in the art can understand and implement it without creative work.
[0090] Correspondingly, the present application further provides an electronic device, including: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the above-mentioned public omics data retrieval method. As Figure 5 shown, it is a hardware structure diagram of any device with data processing capabilities where the public omics data retrieval method provided by the embodiment of the present invention is located. In addition to Figure 5 the processors, memory, and network interfaces shown, any device with data processing capabilities where the device in the embodiment is located usually includes other hardware according to the actual functions of the device with data processing capabilities, which will not be elaborated here.
[0091] Correspondingly, the present application further provides a computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed by a processor, the above-mentioned public omics data retrieval method is implemented. The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium may also be an external storage device, such as a plug-in hard disk, a smart media card (SMC), an SD card, a flash card, etc. equipped on the device. Further, the computer-readable storage medium may also include both the internal storage unit of any device with data processing capabilities and the external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and may also be used to temporarily store the data that has been output or will be output.
[0092] Those skilled in the art will readily think of other implementation schemes of the present application after considering the specification and practicing the content disclosed herein. The present application aims to cover any variations, uses, or adaptive changes of the present application, and these variations, uses, or adaptive changes follow the general principles of the present application and include the common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and embodiments are only regarded as exemplary.
[0093] It should be understood that the present application is not limited to the exact structure described above and shown in the drawings, and various modifications and changes can be made without departing from its scope.
[0094] The above are only the preferred embodiments of the present invention. Although the present invention has been disclosed above in preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make many possible changes and modifications to the technical solution of the present invention, or modify it into equivalent embodiments with equivalent changes, without departing from the scope of the technical solution of the present invention. Therefore, any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the present invention without departing from the content of the technical solution of the present invention still fall within the scope of protection of the technical solution of the present invention.
Claims
1. A method for retrieving publicly available omics data, characterized in that, The method includes: Extracting omics data from an open-source omics database, where the omics data includes omics ontology data and omics annotation data; Splitting the omics annotation data into multiple text blocks based on a text segmentation algorithm to form an omics annotation text set; Performing standardized preprocessing on the omics ontology data, grouping samples according to the omics annotation data using a large model, extracting key information, and generating multiple segments of an omics ontology text set for describing the omics ontology data through the large model; Integrating the omics annotation text set and the omics ontology text set to obtain an omics text collection containing multiple omics text blocks; Integrating the omics text blocks in the omics text collection with the nodes of the biomedical knowledge graph through a weighted hybrid search method that combines biological weights, BM25 retrieval, and semantic retrieval, and then integrating the omics data with the biomedical knowledge graph; Using a mind map to decompose the user's complex problem into sub-problems represented by the relationships of the biomedical knowledge graph, and then using a large model to answer the sub-problems. Controlling the large model generation process through the mind map to achieve an interpretable generation result.
2. The method according to claim 1, wherein The construction of the omics ontology text set is specifically as follows: Performing standardized preprocessing on the omics ontology data, grouping samples according to the omics annotation data using a large model according to experimental conditions and / or biological characteristics, and then extracting a gene set with significant changes between groups based on the grouping results. Using the large model to convert the analysis result into a natural language description to generate multiple segments of an omics ontology text set for describing the omics ontology data.
3. The method according to claim 1, wherein The expression of the weighted hybrid search method is as follows: ; ; ; Among them, is the weighted mixed search score, is the semantic similarity score, is the keyword score, is the biological weight vector, is the biological weight set according to the type importance of the node in the biomedical knowledge graph, is the vector form of the omics text block , is the combined text of the node and its definition, is 's vector form, represents calculating the cosine similarity of vectors as the semantic similarity score, represents calculating the keyword score through the BM25 algorithm, is used to adjust the proportion of the semantic similarity score and the keyword score.
4. The method according to claim 1, wherein Optimize the design of the user's graph query plan, specifically: the user inputs an original query, the large model identifies the entities required by the user in the original query based on the knowledge graph, obtains the set of entity node labels related to the query, expands the entity table according to the entity types and relationship types in the knowledge graph, obtains all the entities on the shortest path between entities, and then converts the queried entities into sub-problems to obtain a list of sub-problems ; Use a large model to organize the retrieved entities and relationships into natural language form and return them to the user to assist the user in judging whether the query plan is correct; the user determines whether the query results are appropriate and whether sub-problems need to be optimized according to the requirements. If optimization is required, a list of sub-problems is regenerated ; Obtain a set of node labels that are closely related to the original query based on the entity nodes and the associated relationships between the nodes finally confirmed by the user 5. The method according to claim 4, characterized in that Evaluating the quality of sub-problems by optimizing the scoring function. If the score of a sub-problem is lower than the set threshold, re-analyzing the path to obtain a new sub-problem. The expression for optimizing the scoring function is as follows: ; Among them, is the score value obtained by optimizing the scoring function, is the number of sub-problems, is the sub-problem and the cosine similarity of the original query ; is the degree of association between sub-problems, that is, whether the nodes corresponding to the sub-problems can be connected to the same omics data, is initially 0 and is used to judge the similarity between the newly generated sub-problem list and the original sub-problem list after user feedback, that is, the number of omics data shared by ; , and are weight parameters and are dynamically adjusted according to user feedback.
6. The method according to claim 4, characterized in that, After obtaining a set of node labels closely associated with the original query, calculating the Jaccard similarity between the set of node labels and each node label corresponding to each associated omics data, sorting according to the Jaccard similarity to generate a list of omics data. By analyzing the omics text collection corresponding to the omics data, the large model explains the reason for selecting specific omics data in natural language form and elaborates on the reasons for the selection in the returned result; the large model comprehensively summarizes the content of the selected omics data.
7. An open omics data retrieval system, characterized in that, For implementing the open omics data retrieval method described in any one of claims 1-6, the system includes: An omics data information extraction module for extracting omics data from an open-source omics database, where the omics data includes omics ontology data and omics annotation data; splitting the omics annotation data into multiple text blocks based on a text segmentation algorithm to form an omics annotation text set; performing standardized preprocessing on the omics ontology data, grouping samples according to the omics annotation data using a large model, extracting key information, and generating multiple segments of an omics ontology text set for describing the omics ontology data through the large model; integrating the omics annotation text set and the omics ontology text set to obtain an omics text collection containing multiple omics text blocks; The omics data integration biomedical knowledge graph module is used to integrate the omics text blocks in the omics text set with the nodes of the biomedical knowledge graph through a weighted hybrid search method that combines biological weights, BM25 retrieval, and semantic retrieval, and then integrate the omics data with the biomedical knowledge graph; The large model reasoning module based on the mind map decomposes the complex problems of users into sub-problems represented by the relationships of the biomedical knowledge graph using the mind map, and then uses the large model to answer the sub-problems, controlling the large model generation process through the mind map to achieve an interpretable generation result.
8. An electronic device, comprising a memory and a processor, characterized in that, The memory is coupled to the processor; wherein, the memory is used to store program data, and the processor is used to execute the program data to implement the disclosed omics data retrieval method according to any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the disclosed omics data retrieval method according to any one of claims 1-6.
10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the disclosed omics data retrieval method according to any one of claims 1-6.
Citation Information
Patent Citations
Recommendation method based on knowledge graph and attention mechanism
CN119474557A
Method and system for predicting biological entities
WO2024236317A1