Interpretable Biological Information Data Processing Method, System and Electronic Device
The method enhances biological data processing by using a trained model and verification system to generate and validate data analysis flows, improving efficiency and interpretability of results.
Patent Information
- Application Number
- CN202211642032.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-20
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-12-20
AI Technical Summary
Existing machine learning methods have problems in the processing of biological information data, such as poor data processing effects and difficult to explain analysis results, especially in deep learning applications, which lack transparency and understanding.
By constructing an interpretable biological information data processing method, a data analysis execution flow is generated using the trained data analysis decision model, and the confidence verification model is used to verify the prediction results. It is divided into a set of high confidence, medium confidence and low confidence results, and the target overall model is constructed and verified.
It improves the efficiency of data processing and interpretability of results, ensures the transparency and accuracy of analysis results, and enhances the generalization ability of the model and the credibility of the results.
Smart Images

Figure CN116130007B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of bioinformatics technology, and in particular, to an interpretable bioinformatics data processing method, system, and electronic device. Background Art
[0002] In the post-genomic era, through high-throughput sequencing and other technical means, more and more bioinformatics data can be obtained, including sequence data of nucleic acids, proteins, etc. With the rapid growth of data volume and the continuous development of artificial intelligence and big data technologies, how to effectively process and analyze bioinformatics data has gradually become a very important requirement, and this requirement is particularly obvious in the current computational biology era.
[0003] Currently, machine learning technology has been rapidly applied to current bioinformatics data processing. Although machine learning, especially deep learning, has advantages such as end-to-end, due to the black-box characteristics of most machine learning and the fact that the current understanding of life phenomena is not comprehensive, when applying bioinformatics data processing based on machine learning, especially deep learning methods, to specific application scenarios, there are problems of poor data processing effects and difficult-to-interpret analysis results. Summary of the Invention
[0004] The purpose of the present invention is to provide an interpretable bioinformatics data processing method, system, and electronic device to alleviate the technical problems of poor data processing effects and unexplainable analysis results existing in the prior art.
[0005] In a first aspect, an embodiment of the present invention provides an interpretable bioinformatics data processing method, which includes:
[0006] According to the user's data processing requirements, use the trained data analysis decision model to make a prediction and generate a data analysis execution flow; the above data analysis execution flow includes several bioinformatics data analysis links and their corresponding order; the above data analysis links include a data processing link and a modeling link;
[0007] Based on the above modeling link of the above data analysis execution flow, perform corresponding model construction to generate a target overall model;
[0008] Use the above target overall model to execute the corresponding data processing link based on the above data analysis execution flow to generate a prediction result;
[0009] Use the trained result verification model to verify the prediction result of the above target overall model, and determine the confidence set to which the above prediction result belongs; the above confidence set includes: a high-confidence result set, a medium-confidence result set, and a low-confidence result set; among them, the above high-confidence result set and the above medium-confidence result set are the final prediction results.
[0010] In some possible embodiments, before the step of generating the data analysis execution flow by predicting using the trained data analysis decision-making model according to the user data processing requirements, the above method further includes: generating an analysis decision training database based on a pre-constructed biological information data processing knowledge base; the above analysis decision training database is used to train the data analysis decision-making model to generate a trained data analysis decision-making model; wherein, the framework of the above biological information data processing knowledge base includes several first-level categories and multiple sub-level categories; the latter sub-level category is obtained by subdividing the previous sub-level category.
[0011] In some possible embodiments, after the step of generating the data analysis execution flow by predicting using the pre-generated data analysis decision-making model according to the user data processing requirements, the above method further includes: generating an interaction frequency parameter between the execution server and the core server using the pre-generated recommendation decision-making model; the above recommendation decision-making model includes predefined rules and a machine learning model trained based on server operation data.
[0012] The step of generating the target overall model by performing corresponding model construction based on the above modeling link of the data analysis execution flow includes: each execution server performs corresponding model training for the above modeling link, and if the above interaction frequency parameter requirements are met during the training process, it sends the current model parameters to the core server; the above core server generates overall model parameters based on the received current model parameters and sends them to each of the above execution servers; each of the above execution servers performs model training based on the above overall model parameters, and repeats the process of parameter interaction until the above model meets the corresponding standard, and determines the current model as the target overall model.
[0013] In some possible embodiments, before the step of verifying the prediction result of the above target overall model using the trained result verification model and determining the confidence set to which the above prediction result belongs, the above method further includes: constructing a result verification knowledge base based on a pre-constructed biological information data processing knowledge base; the result verification knowledge base includes: a relationship knowledge base and a causal relationship knowledge base; constructing a result verification training database based on the above causal relationship knowledge base; wherein, the above result verification training database is used to train the above result verification model to generate a trained result verification model; the data in the above result verification training database includes: object combinations, result confidence levels, and relationship path sequences; the object combinations are composed of coding sequences of different entity objects.
[0014] In some possible implementation manners, the step of constructing a result verification knowledge base based on a pre-constructed bioinformatics data processing knowledge base includes: determining various entity objects with reference to the bioinformatics data processing knowledge base; determining a relationship knowledge base according to the relationships between the various entity objects in the bioinformatics data processing knowledge base; and determining a causal relationship knowledge base according to the causal relationships between the various entity objects in the bioinformatics data processing knowledge base.
[0015] In some possible implementation manners, the step of using a trained result verification model to verify the prediction result of the target overall model and determining the confidence set to which the prediction result belongs includes: using the target overall model to make a prediction and generate a prediction result; the prediction result is the output object of the target overall model; the input object of the target overall model is the specific input content on the selected execution server; verifying the prediction result of the target overall model based on the result verification knowledge base and outputting a verification result; the verification result includes: a low-confidence result and a high-confidence result; the high-confidence result is used to generate a high-confidence result set; inputting the object combination corresponding to the low-confidence result into the trained result verification model for verification and outputting a verification model prediction result; the verification model prediction result includes: a low-confidence prediction result and a high-confidence prediction result; generating a medium-confidence result set according to the high-confidence prediction result; generating a low-confidence result set according to the low-confidence prediction result; and using the medium-confidence result set and the high-confidence result set to determine the final prediction result.
[0016] In some possible implementation manners, the step of verifying the prediction result of the target overall model based on the result verification knowledge base and outputting a verification result includes: determining whether there is a connection path between the input object and the output object of the target overall model in the relationship knowledge base; if not, determining that the prediction result belongs to the low-confidence result set; if so, determining whether the connection path exists in the causal relationship knowledge base and whether the direction of the causal relationship is consistent; if so, determining that the prediction result belongs to the high-confidence result; if not, determining that the prediction result belongs to the low-confidence result.
[0017] In a second aspect, an embodiment of the present invention provides an interpretable bioinformatics data processing system, which includes:
[0018] A data analysis execution flow generation module, configured to make a prediction using a trained data analysis decision model according to user data processing requirements and generate a data analysis execution flow; the data analysis execution flow includes several bioinformatics data analysis links and corresponding sequences; the data analysis links include a data processing link and a modeling link;
[0019] A target overall model construction module for performing corresponding model construction based on the above-mentioned modeling link of the data analysis execution flow to generate a target overall model; a prediction result verification module for verifying the prediction result of the above-mentioned target overall model by using the trained result verification model to determine the confidence set to which the above-mentioned prediction result belongs; the above-mentioned confidence set includes: a high-confidence result set, a medium-confidence result set, and a low-confidence result set; wherein, the above-mentioned high-confidence result set and the above-mentioned medium-confidence result set are the final prediction results.
[0020] In a third aspect, an embodiment of the present invention provides an electronic device, including a memory and a processor. A computer program that can run on the processor is stored in the memory. When the processor executes the computer program, the steps of the method described in any one of the above first aspects are implemented.
[0021] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium. The computer-readable storage medium stores machine-executable instructions. When the computer-executable instructions are called and run by a processor, the computer-executable instructions cause the processor to run the method described in any one of the above first aspects.
[0022] The present invention provides an interpretable biological information data processing method, system, and electronic device. The method includes: first, according to the user data processing requirements, using the trained data analysis decision model to make a prediction to generate a data analysis execution flow; then performing corresponding model construction based on the modeling link of the data analysis execution flow to generate a target overall model; and then using the trained result verification model to verify the prediction result of the target overall model to determine the confidence set to which the prediction result belongs, so as to obtain the final prediction result. This method alleviates the technical problems of poor data processing effect and difficult-to-interpret analysis results, and achieves the technical effects of improving data processing efficiency and improving the interpretability of processing results. Description of the Drawings
[0023] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required to be used in the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0024] Figure 1 It is a schematic flowchart of an interpretable biological information data processing method provided by an embodiment of the present invention;
[0025] Figure 2Schematic diagram of the specific process for verifying the prediction result of the model in an interpretable biological information data processing method provided by an embodiment of the present invention;
[0026] Figure 3 Schematic diagram of the structure of an interpretable biological information data processing system provided by an embodiment of the present invention;
[0027] Figure 4 Schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Components of the embodiments of the present invention generally described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.
[0029] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0030] It should be noted that like reference numerals and letters denote like items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. Some embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Without conflict, the following embodiments and features in the embodiments can be combined with each other.
[0031] In the post-genomic era, through technical means such as high-throughput sequencing, more and more biological information data can be obtained, including sequence data of nucleic acids, proteins, etc. With the rapid growth of the data volume and the continuous development of artificial intelligence and big data technologies, how to effectively process and analyze biological information data has gradually become a very important requirement, and this requirement is particularly obvious in the current computational biology era. Applications of biological information data analysis and processing include region recognition, structure and function prediction, molecular interaction prediction, binding site recognition, drug development, etc. The basic analysis links can be divided into, for example, preprocessing, feature extraction, statistical analysis, modeling and other links.
[0032] At present, machine learning technology has been rapidly applied to current bioinformatics data processing. Although machine learning, especially deep learning, has advantages such as end-to-end, due to the black-box characteristics of most machine learning and the fact that the current understanding of life phenomena is incomplete, when bioinformatics data processing based on machine learning, especially deep learning methods, is applied to specific application scenarios, there are problems of poor data processing effects and difficult-to-interpret analysis results. Based on this, the embodiments of the present invention provide an interpretable bioinformatics data processing method, system, and electronic device to alleviate the above problems.
[0033] To facilitate the understanding of this embodiment, first, a detailed introduction to an interpretable bioinformatics data processing method disclosed in the embodiments of the present invention is provided. Refer to Figure 1 The flowchart of an interpretable bioinformatics data processing method shown. This method can be executed by an electronic device and mainly includes the following steps S110 to step S130:
[0034] S110: According to the user's data processing requirements, use the trained data analysis decision model for prediction to generate a data analysis execution flow;
[0035] Among them, the data analysis execution flow includes several bioinformatics data analysis links and their corresponding order; the data analysis links include a data processing link and a modeling link. As a specific example, the data processing link may include: preprocessing, feature extraction, statistical analysis, visualization, result annotation, processing strategies, etc.; the modeling link may include: machine learning modeling, deep learning modeling, etc.
[0036] In this embodiment, the relevant content of the user's data processing requirements may include: data processing purpose, parameters related to the data to be processed (i.e., the data itself), data processing expectations, storage and servers participating in the calculation. For security reasons, the servers and storage that can perform calculations are all specified in advance, and the data has been saved in the servers and storage.
[0037] In one embodiment, after the data analysis decision model in S110 predicts and generates a data analysis execution flow, it can be customized and confirmed by the user, and then the confirmed data analysis execution flow (excluding machine learning modeling and deep learning modeling) is started on each execution server to automatically complete the data analysis process (such as generally including data preprocessing, feature extraction, statistical analysis, etc.).
[0038] S120: Based on the modeling link of the data analysis execution flow, perform corresponding model construction to generate a target overall model;
[0039] That is to say, after generating the data analysis execution flow using the data analysis decision model, and the data processing links included in the confirmed data analysis execution flow (excluding machine learning modeling and deep learning modeling), each execution server and the core server jointly perform collaborative modeling to construct the target overall model, and the target overall model can also be sent to each execution server to process its own data accordingly.
[0040] Correspondingly, the step of S120 constructing the corresponding model based on the modeling link of the data analysis execution flow to generate the target overall model may include the following processes:
[0041] (1) Each execution server performs corresponding model training for the modeling link. If the interaction frequency parameter requirements are met during the training process, the current model parameters are sent to the core server; (2) The core server generates overall model parameters based on the received current model parameters and sends them to each execution server; (3) Each execution server performs model training based on the overall model parameters, and repeats the process of parameter interaction until the model meets the corresponding standards, and determines the current model as the target overall model.
[0042] As a specific example, after generating and customizing and confirming the data analysis execution flow using the data analysis decision model, each execution server starts to execute the system modeling process (for example: machine learning modeling, deep learning modeling): according to the pre-specified division of training, validation, and test data sets, the machine learning (or deep learning) algorithm determined in the data analysis execution flow is used for model training; after obtaining the intermediate results of the training model (completing one or several epochs) during the process, if the interaction frequency requirements determined by the above recommendation are met, the execution server sends the current model parameters to the core server; the core server comprehensively processes the model parameters of the received execution servers to obtain the overall model parameters, and then the core server sends the new overall model parameters to each execution server; then each execution server uses the new overall model parameters as the baseline and restarts the above modeling process until the model meets the corresponding standards, and then stops the modeling process, and the current model is the target overall model.
[0043] It should be noted that the data transmission between the core server and the execution server can be encrypted by methods such as symmetric, asymmetric, and homomorphic encryption.
[0044] In this example, the process by which the core server synthesizes the model parameters received from the execution servers to obtain the overall model parameters is as follows: Determine the weights of the execution servers. First, obtain the number of training samples X (vector) on each execution server. Then, calculate the proportion (between 0 and 1, Y - vector) of the number of training samples (SUM) on each execution server to the total number of training samples. According to the interval positions of the proportions of the number of training samples on each execution server to the total number of training samples (for example, 0 - 0.2 is set to 0.1, 0.2 - 0.4 is set to 0.3, etc.), set the weight ratios Z (vector) for each execution server respectively; For the corresponding parameters in the models of each execution server, multiply them by their respective weights and then sum and average to obtain the overall synthesized model parameters.
[0045] S130: Use the trained result verification model to verify the prediction result of the target overall model, and determine the confidence set to which the prediction result belongs;
[0046] Among them, the confidence set includes: a high - confidence result set, a medium - confidence result set, and a low - confidence result set; Among them, the high - confidence result set and the medium - confidence result set are the final prediction results.
[0047] The present invention provides an interpretable biological information data processing method. The method includes: First, according to the user's data processing requirements, use the trained data analysis decision model for prediction to generate a data analysis execution flow; Then, perform corresponding model construction based on the modeling link of the data analysis execution flow to generate a target overall model; Then, use the trained result verification model to verify the prediction result of the target overall model, and determine the confidence set to which the prediction result belongs, so as to obtain the final prediction result. This method alleviates the technical problems of poor data processing effect and difficult - to - interpret analysis results, and achieves the technical effects of improving data processing efficiency and the interpretability of processing results.
[0048] In one embodiment, when constructing the data analysis decision model, use the analysis decision training database for training, and the analysis decision training database can be generated according to the pre - constructed biological information data processing knowledge base.
[0049] Among them, the framework of the biological information data processing knowledge base includes several first - level categories and multiple sub - level categories; Each subsequent sub - level category is obtained by subdividing the previous sub - level category.
[0050] That is to say, to achieve interpretable biological information data processing calculations, first construct the framework of the biological information data processing knowledge base, and use a multi - level and multi - link method for construction. As a specific example, the category relationship of the knowledge base framework includes several first - level categories, and the first - level categories include: processing objectives, object data, analysis links, files, tasks; Each first - level category can be subdivided into several second - level categories, and the specific inclusion relationships are as follows:
[0051] The categories under the processing target include: diseases, concerns, efficiency, performance, etc. The categories under the object data include: original, conversion, intermediate results, final results, etc. The categories under the analysis process include: preprocessing, feature extraction, statistical analysis, machine learning modeling, deep learning modeling, visualization, result annotation, auxiliary tools, processing strategies, etc. The categories under the file include: configuration files, data files, log files, etc. The categories under the task include: file transfer tasks, data processing tasks, etc.
[0052] The third-level categories are obtained by subdividing the second-level categories. For example, the categories under diseases can include: dementia, tumors, etc. The categories under concerns can include: DNA, RNA, protein, metabolism, drugs, pathways, etc. The categories under efficiency can include operation time, occupied space, etc. The categories under performance can include accuracy, recall rate, significance level, etc. The categories under the original data include: demographic characteristics, groups, categories, platform devices, sequences, additional features, related data, etc. The categories under conversion / intermediate results include: steps, inputs, algorithms, results, etc. The categories under the final results include: statistical analysis results (including subjects / groups, algorithms, results), model modeling results (including subjects / groups, algorithms, models), etc.; The categories under preprocessing include format conversion, adapter removal, primer removal, duplicate removal, sequence alignment, Barcode recognition, filling, screening, assembly, etc. The categories under feature extraction include: nucleic acid features, protein features, metabolic features, drug features, feature selection and dimensionality reduction, etc. The categories under statistical analysis can include: differences (expression, mutation, snp), correlations, regressions, etc. The categories under machine learning modeling can include: classification, clustering, ensemble, etc. The categories under deep learning can include: cnn-based models, rnn-based models, attention-based models, ensemble models, etc. The categories under result annotation can include annotation methods, annotation algorithms, etc. The categories under auxiliary tools can include format conversion, anonymization, etc. The categories under processing strategies can include analysis strategies, calculation strategies, etc.; The categories under data files can include original files, conversion files, intermediate process files, calculation result files, etc.; The categories under data processing tasks can include: data conversion tasks, preprocessing tasks, feature extraction tasks, statistical analysis tasks, machine learning tasks, deep learning tasks, etc.
[0053] The fourth-level, fifth-level, and other categories are obtained by successive downward decomposition. For example, the DNA category can be further subdivided into: region, structure, function, mutation, heredity, susceptibility, pathogenicity, sensitivity, prognosis, etc. The RNA category can be further subdivided into: region, structure, function, mutation, non-coding RNA recognition, heredity, susceptibility, pathogenicity, sensitivity, prognosis, etc. The protein category can be further subdivided into: sequence, structure, function, etc. The drug category can be further subdivided into: target protein, binding site, affinity, water solubility, toxicity, etc. The nucleic acid feature category can be further subdivided into kmer, RCkmer (Reverse Compliment Kmer), NAC (Nucleic Acid Composition), DNC (Di-Nucleotide Composition), TNC (Tri-Nucleotide Composition), etc. The protein feature category can be further subdivided into: AAC (Amino Acid composition), EAAC (Enhanced Amino Acid Composition), CKSAAP (composition of k-spaced Amino Acid Pairs), etc. The analysis strategy category can be further subdivided into single analysis, voting analysis, multi-dimensional analysis, etc. And so on until it cannot be further divided.
[0054] In addition, each subcategory under the analysis link category includes detailed parameters such as input, output, parameters, algorithms, tools, steps, templates (if any), etc. Each subcategory under the object data category includes subcategories such as total data volume, format, location, proportion of small files, etc.
[0055] The relationship categories in the knowledge base framework mainly include is-a and attribute relationships, etc. According to the above knowledge base framework, using the evidence-based literature approach, the processing objectives, object data, analysis links, etc. with clear and high-quality evidence support are recorded to establish a knowledge base (the knowledge content can be obtained as a whole using information extraction methods or manual collation methods, and the constructed knowledge base is essentially a knowledge graph).
[0056] Referring to the above knowledge base framework, the content of the data processing objectives required in this embodiment mainly includes: the targeted disease (which can be empty), the focus, etc. The content related to the data itself includes: data size, format, modality, current location (as well as the directory hierarchy - such as the subject, group, etc.), attributes (training, testing, validation), etc. The content related to the expectations of data processing includes: the expected data processing time, the expected occupied space, the expected performance indicators, etc. The storage and server content participating in the calculation includes the storage name, type, location, capacity, IP, as well as the name, type, location, IP of the execution server, etc. (Overall, the servers performing the calculation are divided into two categories: core servers and execution servers. The former is unique, and the latter are multiple. The training, validation, and test data sets required during the modeling process are also specified here).
[0057] To effectively perform interpretable bioinformatics data processing, considering the possible differences in the data of each data owner, it is necessary to first perform relevant inspection and processing on the data. It mainly includes: (1) Data format conversion: For the data objects in the storage, check according to the pre-defined data organization method (such as modifying according to the neuroimaging bids standard). If it conforms to the agreed format, it passes the format check; if the data does not conform to the agreed format, convert the data object according to the agreed format; (2) Data selection: Mainly check whether the bioinformatics data participating in the calculation meets the required types, platform devices, conventional quality requirements, etc.; (3) Data anonymization: Process the personal information in the data to avoid the existence of names and other privacy information. Thus, the data inspection work is completed.
[0058] The unified data format, data quality, etc. after the data inspection here lay the foundation for the subsequent flexible selection, automatic execution of the data processing flow, and efficient data processing modeling.
[0059] All of the above work can be automatically executed by a unified script located on each execution server.
[0060] To obtain a data analysis decision-making model, it is necessary to first construct an analysis decision-making training database. The training data comes from the content of the bioinformatics data processing knowledge base and has been confirmed by experts. The specific data mainly includes analysis requirements (mainly including diseases, concerns - standardized to form a coded text sequence, and computational efficiency and performance requirements - corresponding the efficiency and performance requirements to options such as extremely high, relatively high, high, medium, etc. according to pre-determined rules - and forming a coded text sequence) and the data analysis execution flow; the data analysis execution flow consists of one or more components, and the components mainly refer to preprocessing, feature extraction, statistical analysis, machine learning modeling, deep learning modeling, visualization, result annotation, processing strategies, etc.; the components consist of one or more modules, and the modules refer to specific algorithms applied to specific calculations, such as the secondary screening algorithm (there can be multiple) in preprocessing; from this perspective, the data analysis execution flow is a directed graph, the nodes are modules, and the edges are sequential relationships; the text sequences of the data analysis requirements and the data analysis execution flow (both after standardization) are represented in the following way; and the training data is divided into training, validation, and test sets according to the ratio of 7:1.5:1.5.
[0061] The expression of the coded text sequence of the analysis requirements is composed of: where the diseases, concerns, etc. are standardized to form a coded text sequence, and the computational efficiency, performance requirements, etc. correspond the efficiency and performance requirements to options such as extremely high, relatively high, high, medium, etc. according to pre-determined rules and form a coded text sequence, and then the text sequences are combined together to form a unified text sequence (the connection symbol is a comma,), for example: (Alzheimer's disease, DNA, high efficiency, relatively high performance). The expression of the text sequence of the data analysis execution flow is composed of: the main part is the component, and the component includes the module, and the specific algorithms and relevant parameter settings are recommended (connected by commas), represented in a hierarchical method, the sequential relationship between components is represented by "->", the sequential relationship between modules is represented by "=>", and the parallel relationship is represented by a comma ",", for example: (
[0063] (Preprocessing (
[0065] (Secondary screening, algorithm =, mincell = 3, …) => (…) ) )
[0068] ->(Feature extraction (…) => …)
[0069] ->(…) )
[0071] Correspondingly, the input of the data analysis decision-making model is the relevant requirements of data analysis, that is, a text sequence expression composed of diseases, concerns, computational efficiency, performance requirements, etc. (after standardization). The output of the data analysis decision-making model is a text sequence expression composed of components (i.e., preprocessing, feature extraction, statistical analysis, machine learning modeling, deep learning modeling, visualization, result annotation, processing strategies, etc., and specific modules that make up the components). The model adopts a sequence-to-sequence generation model, such as an rnn combined with attention, or T5, etc.
[0072] From the data analysis decision-making model, according to the user's data analysis requirements, the data analysis execution flow can be obtained. The software-related preset parameters in the data analysis execution flow are preset according to evidence-based literature and the sources are indicated; the preset parameters related to data objects are detected in advance by the software (including the part preset by the user).
[0073] In addition, since the interaction frequency between each execution server and the core server participating in the calculation will affect parameters such as the effect and duration of data analysis modeling, the interaction frequency is also automatically recommended.
[0074] In one embodiment, after predicting using the pre-generated data analysis decision-making model according to the user data processing requirements and generating the data analysis execution flow in step S110, the above method may further include: generating the interaction frequency parameter between the execution server and the core server using the pre-generated recommendation decision-making model; the recommendation decision-making model includes predefined rules and a machine learning model trained based on server operation data, where the server includes an execution server and a core server.
[0075] As a specific example, the recommendation decision-making model includes two parts: predefined rules and a machine learning model. First, establish the rules between the interaction frequency parameter, the relevant requirements of the user's processed data, the number of servers, and the data analysis execution flow determined by the above recommendation according to relevant expert knowledge, and make relevant decisions according to the rules; then, based on the relevant data during the server operation, use machine learning methods to construct the relevant models between the interaction frequency parameter, the relevant requirements of the user's processed data, the number of servers, and the data analysis execution flow determined by the above recommendation (the training data comes from the operation data confirmed by expert annotation). After the model effect meets the relevant requirements, execute the decision recommendation according to the model; among them, the algorithms used include the integration based on neural networks; the input of the model is the relevant requirements of the user's processed data, the number of servers, and the data analysis execution flow determined by the above recommendation, and the output of the model is the interaction frequency; the loss function of the model takes into account objectives such as the minimum execution time and the optimal model effect. It should be noted that the interaction frequency parameter predicted and output by the recommendation decision-making model can be customized (that is, the output result can be selected and modified).
[0076] In one embodiment, before verifying the prediction result in step S130 above, the target overall model obtained in step S120 can be evaluated first. Through different evaluation parameters, the performance status of the model itself can be objectively evaluated to obtain a high-performance target overall model. Among them, the evaluation parameters can include the significance level, correlation coefficient size, regression coefficient size, model accuracy, recall rate, F1, ROC, AUC, etc. in statistical analysis. In practical applications, in combination with specific application scenarios, a high-performance model with relatively high-quality corresponding evaluation parameters can be selected as the final target overall model for subsequent result verification.
[0077] In this embodiment, when constructing the result verification model, it is necessary to first construct a result verification training database, and use this result verification training database to train the result verification model to generate a trained result verification model. Among them, the result verification training database can be constructed based on the causal relationship knowledge base, and the specific data therein includes: object combination, result confidence level, and relationship path sequence; the object combination is composed of the coding sequences of different entity objects.
[0078] As a specific embodiment, the data in the result verification training database usually comes from the content in the causal relationship knowledge base and is confirmed by expert annotation. The specific data mainly includes object combinations (independent variables) composed of different entity objects such as DNA, RNA, and proteins (forming a coding sequence after standardization), result confidence levels, and relationship path sequences (dependent variables, the result confidence levels are divided into high, low -1, and 0. For the object combinations with the same expression direction in the path relationship formed by at least two connections between the selected object combinations in the causal relationship knowledge base, the result confidence level is high, otherwise it is low; the path relationship is represented in the following manner after standardization to form a relationship path sequence).
[0079] Therefore, before the step of using the trained result verification model to verify the prediction result of the target overall model in S130 and determining the confidence set to which the prediction result belongs, the method can further include: constructing a result verification knowledge base based on a pre-constructed bioinformatics data processing knowledge base; the result verification knowledge base includes: a relationship knowledge base and a causal relationship knowledge base.
[0080] In one embodiment, the step of constructing a result verification knowledge base based on a pre-constructed bioinformatics data processing knowledge base includes determining various entity objects with reference to the bioinformatics data processing knowledge base; determining the relationship knowledge base according to the relationships between various entity objects in the bioinformatics data processing knowledge base; and determining the causal relationship knowledge base according to the causal relationships between various entity objects in the bioinformatics data processing knowledge base.
[0081] That is to say, a result verification knowledge base can be constructed first; according to the basic relationships and principles such as DNA-RNA-protein-metabolism, epigenetics, and signal pathways, a result verification knowledge base including dimensions such as diseases, drugs, DNA, RNA, proteins, and metabolism is constructed. Then, the prediction results of the model are verified based on the result verification knowledge base. The basic logic of the verification is that the logical relationships between objects should conform to the basic principles of DNA-RNA-protein-metabolism. For example, for the problem of finding drug targets that may be effective for a certain disease, in the connection path between the disease and the protein, the expression direction (high or low) of the protein of the specific disease object in the experimental group (compared with the control group) should be the same as the expression direction of the RNA in the path. Finally, the prediction results are verified based on the result verification model. The basic logic of the verification is to judge the prediction results through the verification model.
[0082] In one embodiment, the steps of using the trained result verification model to verify the prediction results of the target overall model and determining the confidence set to which the prediction results belong include: (1) Using the target overall model for prediction to generate prediction results; the prediction results are the output objects of the target overall model; the input objects of the target overall model are the specific input contents on the selected execution server (after the same processing calculations as in the data processing link); (2) Verifying the prediction results of the target overall model based on the result verification knowledge base and outputting verification results; the verification results include: low-confidence results and high-confidence results; the high-confidence results are used to generate a high-confidence result set; (3) Inputting the object combinations corresponding to the low-confidence results into the trained result verification model for verification and outputting the verification model prediction results; the verification model prediction results include: low-confidence prediction results and high-confidence prediction results; (4) Generating a medium-confidence result set according to the high-confidence prediction results; generating a low-confidence result set according to the low-confidence prediction results; (5) Using the medium-confidence result set and the high-confidence result set to determine the final prediction results.
[0083] Among them, the process of verifying the prediction results of the target overall model based on the result verification knowledge base in step (2) above specifically includes: determining whether there is a connection path between the input object and the output object of the target overall model in the relationship knowledge base; if not, determining that the prediction results belong to the low-confidence result set.
[0084] Among them, a high-performance target overall model is obtained through model evaluation. The specific input content on the selected execution server is input into the model (after the same processing calculations as in the data processing link) to obtain the output result of the model, that is, the prediction result, and then the prediction result is verified. For example: the input object is a specific disease, and the output object is the predicted target.
[0085] In one embodiment, the step of using the trained result verification model to verify the prediction result of the target overall model and determining the confidence set to which the prediction result belongs further includes: If it exists, determine whether the connection path exists in the causal relationship knowledge base and whether the direction of the causal relationship is consistent; If so, determine that the prediction result belongs to a high-confidence result; If not, determine that the prediction result belongs to a low-confidence result.
[0086] As a specific example, refer to Figure 2 As shown, the steps for verifying the prediction result of the model may include the following process: (S210) Construct a result verification knowledge base; (S220) Verify the model prediction result based on the result verification knowledge base; (S230) Construct a result verification model; (S240) Verify the prediction result based on the result verification model.
[0087] Among them, step (S210) specifically includes: According to different categories such as DNA, RNA, protein, etc., determine the specific entity objects of each category; Extract part of the relationships between entity objects of each category from the pre-constructed biological information data processing knowledge base to obtain a relationship knowledge base; Extract part of the causal relationships (positive effects, negative effects) between entity objects of each category from the pre-constructed biological information data processing knowledge base to obtain a causal relationship knowledge base.
[0088] The process of step (S220) verifying the model prediction result based on the result verification knowledge base is as follows:
[0089] (S021) For the input object and output object of the model, find all connection paths between the two in the relationship knowledge base. If a connection path exists, execute (S022), otherwise execute (S023);
[0090] (S022) For the connection path between the input and output objects, determine whether it exists in the causal relationship knowledge base and whether the causal direction relationship therein is consistent (for example, in the connection path between a disease and a protein regarding a drug target that may be effective for a certain disease, the protein expression direction (high, low) and the RNA expression direction of the specific disease object in the experimental group should be the same). Put the model output results with consistent causal relationships into the high-confidence result set, otherwise put this model output result into the low-confidence result set;
[0091] (S023) For the model output results without a connection path, put them into the low-confidence result set.
[0092] Step (S230) for constructing the result verification model specifically includes:
[0093] First, construct a training database. The training data comes from the content in the above causal relationship knowledge base and has been confirmed by expert annotation. The specific data mainly includes object combinations (independent variables) composed of different entity objects such as DNA, RNA, and proteins (forming a coding sequence after standardization), and result confidence levels and relationship path sequences (dependent variables, where the result confidence levels are divided into high, low -1, and 0. For the object combinations with the same expression direction in the path relationship formed by at least two connections between the selected object combinations in the causal relationship knowledge base, their result confidence levels are high, otherwise low; the path relationships are represented in the following manner after standardization to form relationship path sequences). For example, for the problem of potential drug targets for a certain disease mentioned above, the input of the training data is a combination of the disease and the protein. If in the connection path between the disease and the protein, the protein expression direction (high, low) and the RNA expression direction of the specific disease object in the experimental group should be the same, then the dependent variable result is a high confidence level, otherwise it is a low confidence level.
[0094] For high-confidence data, simultaneously obtain the causal relationship path sequences between object combinations, which are represented in the following manner after standardization; thus construct a training database, and divide the training data into training, validation, and test sets according to the ratio of 7:1.5:1.5. The training model is constructed using the random forest algorithm and sequence-to-sequence generation models such as RNN with attention, or T5, etc.
[0095] Among them, the expression formula of the object combination sequence is: different entity objects such as DNA, RNA, proteins, and diseases form a coded text sequence after standardization (the connection symbol is a comma,), for example: (Alzheimer's disease, APOE4). The result verification model confidence level and relationship path sequence (the confidence level is divided into high, low -1, and 0, and the path relationship forms a relationship path sequence after being represented by a text sequence after standardization). The expression formula of the relationship path sequence is composed of: the main part is the standardized entity, which is represented by a hierarchical method, and the order relationship between entities is represented by "->" (negative effect "= >"), the branches in the path are represented by parentheses "()", and the parallel relationship is represented by a comma ",", for example: (
[0097] (APOE4)
[0098] ->(ERK1)
[0099] ->(Amyloidβ-protein)
[0100] ->(Alzheimer’s disease) )
[0102] Correspondingly, the input of the result verification model is a combined sequence expression of different objects such as DNA, RNA, and protein, and the output result of the model is a confidence level. For the object combination with a high confidence level, the output result also includes the relational path sequence expression.
[0103] The process of verifying the prediction result based on the result verification model in step (S240) is as follows: For the object combination that has been verified by the result verification knowledge base and then placed in the low-confidence result set, input it into the result verification model for verification; (S041) For those with a high confidence level in the output result of the result verification model, use this part of the result as the medium-confidence result set of the final result, and at the same time output its relational path sequence (as a possible path relationship explanation for the relationship between objects); (S042) For those with a low confidence level in the output result of the result verification model, use this part of the result as the low-confidence result set of the final result.
[0104] Based on this, the model prediction results are finally divided into a high-confidence result set, a medium-confidence result set, and a low-confidence result set, and the high- and medium-confidence result sets are used as the output result set of the final prediction model, thus narrowing the range of the output results of the prediction model and improving the accuracy of the output results. At the same time, combined with the above result verification method, the output results of the model can be logically explained (the causal relationship path between the model input object and the output object is the logical relationship between the two).
[0105] To maximize the convenience for users to process bioinformatics-related data, the embodiment of this method can also provide a graphical interface for users to express the relevant requirements for bioinformatics data processing. The content for users to express data processing-related requirements in the graphical interface can be divided into four categories: namely, data processing purpose, data itself, data processing expectation, and storage and servers participating in the calculation (for security reasons, the servers and storage that can perform calculations are all specified in advance, and the data has been saved in the servers and storage).
[0106] After data calculation and analysis processes such as data analysis, model evaluation, and result verification are completed, the modeling results, prediction results, etc. (as needed) can be viewed through the graphical interface provided by this embodiment, and further applications can be carried out using the results.
[0107] It should be noted that the intermediate result files, final result files, model results of data analysis, model evaluation, prediction results, etc. during the data analysis process are all saved in the location specified by the user according to a pre-determined format (such as modified with reference to neuroimaging bids).
[0108] The interpretable bioinformatics data processing method provided by the embodiments of the present invention configures the objectives, data-related situations, etc. of bioinformatics data processing through an interface, and can select and automatically complete the whole process from preprocessing, feature extraction, statistical analysis, machine learning modeling, deep learning modeling, etc. in one stop; for the model prediction results, based on the constructed result verification knowledge base and model, filtering and verification are carried out according to the principles of bioinformatics correlation, and they are divided into result sets with different confidence levels, narrowing the range of the final model output results and improving the accuracy of the output results; for the model prediction results, based on the constructed result verification knowledge base, logical explanations can be provided for the final output results of the model, improving the interpretability; based on the evidence-based medical knowledge base and the standardized text sequence expression method, it can support the automatic recommendation of the bioinformatics data analysis process and the automatic recommendation of the parameter settings in the analysis process, reducing the analysis workload and difficulty; on the basis of secure computing, efficient computing of bioinformatics data is realized, which can complete system modeling while ensuring that the data is available but invisible, and at the same time improves the effect of system modeling and enhances the generalization ability of the model.
[0109] In addition, the embodiments of the present invention also provide an interpretable bioinformatics data processing system. As shown in Figure 3 the figure, the system includes:
[0110] A data analysis execution flow generation module 310, configured to generate a data analysis execution flow by using a trained data analysis decision model for prediction according to user data processing requirements; the data analysis execution flow includes several bioinformatics data analysis links and corresponding sequences; the data analysis links include a data processing link and a modeling link;
[0111] A target overall model construction module 320, configured to perform corresponding model construction based on the modeling link of the data analysis execution flow to generate a target overall model;
[0112] A prediction result verification module 330, configured to verify the prediction result of the target overall model by using a trained result verification model to determine the confidence level set to which the prediction result belongs; the confidence level set includes: a high-confidence result set, a medium-confidence result set, and a low-confidence result set; among them, the high-confidence result set and the medium-confidence result set are the final prediction results.
[0113] The interpretable bioinformatics data processing system provided by the embodiments of the present application can be specific hardware on a device, or software or firmware installed on the device, etc. The device provided by the embodiments of the present application has the same implementation principle and the same technical effects as those of the foregoing method embodiments. For the sake of brief description, for the parts not mentioned in the device embodiments, reference may be made to the corresponding content in the foregoing method embodiments. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can all refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein. The interpretable bioinformatics data processing system provided by the embodiments of the present application has the same technical features as the interpretable bioinformatics data processing method provided by the foregoing embodiments, so it can also solve the same technical problems and achieve the same technical effects.
[0114] The embodiments of the present application also provide an electronic device. Specifically, the electronic device includes a processor and a storage device; a computer program is stored on the storage device, and the computer program executes the method according to any one of the above-mentioned implementation manners when being run by the processor.
[0115] Figure 4 FIG. 7 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. The electronic device 400 includes: a processor 40, a memory 41, a bus 42, and a communication interface 43. The processor 40, the communication interface 43, and the memory 41 are connected through the bus 42; the processor 40 is configured to execute an executable module stored in the memory 41, such as a computer program.
[0116] Among them, the memory 41 may include a high-speed random access memory (RAM, Random Access Memory), and may also include a non-volatile memory, such as at least one disk memory. Through at least one communication interface 43 (which may be wired or wireless), a communication connection is realized between the system network element and at least one other network element, and the Internet, a wide area network, a local area network, a metropolitan area network, etc. can be used.
[0117] The bus 42 may be an ISA bus, a PCI bus, an EISA bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 4 only a bidirectional arrow is used in FIG. 7, but it does not mean that there is only one bus or one type of bus.
[0118] Among them, the memory 41 is used to store a program. After receiving an execution instruction, the processor 40 executes the program. The method executed by the device defined by the flow process disclosed in any one of the foregoing embodiments of the present invention can be applied to the processor 40 or implemented by the processor 40.
[0119] The processor 40 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 40 or the instructions in the form of software. The above-mentioned processor 40 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field-programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present invention can be directly embodied as being executed and completed by the hardware decoding processor, or can be executed and completed by the combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 41, and the processor 40 reads the information in the memory 41 and combines its hardware to complete the steps of the above method.
[0120] Corresponding to the above method, an embodiment of the present application further provides a computer-readable storage medium, where the computer-readable storage medium stores machine-executable instructions, and when the computer-executable instructions are called and run by a processor, the computer-executable instructions cause the processor to run the steps of the above method.
[0121] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other may be through some communication interfaces, and the indirect coupling or communication connection of the devices or units may be in an electrical, mechanical, or other form.
[0122] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0123] In addition, each functional unit in the embodiments provided in this application may be integrated in a processing unit, may exist separately as individual physical units, or two or more units may be integrated in one unit.
[0124] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, an electronic device, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.
[0125] It should be noted that similar reference numerals and letters indicate similar items in the drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. In addition, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0126] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the various embodiments of the present invention.
Claims
1. An interpretable biological information data processing method, characterized in that Including: Generating an analysis decision training database based on a pre-constructed biological information data processing knowledge base; The analysis decision training database is used to train a data analysis decision model to generate a trained data analysis decision model; wherein, the framework of the biological information data processing knowledge base includes several first-level categories and multiple sub-level categories; the latter sub-level category is obtained by subdividing the previous sub-level category; According to the user data processing requirements, using the trained data analysis decision model for prediction to generate a data analysis execution flow; the data analysis execution flow includes several biological information data analysis links and corresponding sequences; the data analysis links include data processing links and modeling links; Generating an interaction frequency parameter between the execution server and the core server using a pre-generated recommendation decision model; the recommendation decision model includes predefined rules and a machine learning model trained based on server operation data; Performing corresponding model construction based on the modeling link of the data analysis execution flow to generate a target overall model; Based on a pre-constructed biological information data processing knowledge base, constructing a result verification knowledge base; the result verification knowledge base includes: a relationship knowledge base and a causal relationship knowledge base; Constructing a result verification training database based on the causal relationship knowledge base; wherein, the result verification training database is used to train the result verification model to generate a trained result verification model; the data in the result verification training database includes: object combinations, result confidence levels, and relationship path sequences; the object combinations are composed of coding sequences of different entity objects; Using the trained result verification model to verify the prediction result of the target overall model to determine the confidence set to which the prediction result belongs; the confidence set includes: a high-confidence result set, a medium-confidence result set, and a low-confidence result set; wherein, the high-confidence result set and the medium-confidence result set are the final prediction results; The steps of performing corresponding model construction based on the modeling link of the data analysis execution flow to generate a target overall model include: each execution server performs corresponding model training for the modeling link, and if the interaction frequency parameter requirements are met during the training process, it sends the current model parameters to the core server; the core server generates overall model parameters based on the received current model parameters and sends them to each execution server; each execution server performs model training based on the overall model parameters, and repeats the process of parameter interaction until the model meets the corresponding standards and determines the current model as the target overall model.
2. The interpretable biological information data processing method according to claim 1, wherein The steps of constructing a result verification knowledge base based on a pre-constructed biological information data processing knowledge base include: Determining various entity objects with reference to the biological information data processing knowledge base; Determining a relationship knowledge base according to the relationships between the various entity objects in the biological information data processing knowledge base; Determining a causal relationship knowledge base according to the causal relationships between the various entity objects in the biological information data processing knowledge base.
3. The interpretable biological information data processing method according to claim 2, wherein The steps of using the trained result verification model to verify the prediction result of the target overall model and determining the confidence set to which the prediction result belongs include: Use the target overall model to make a prediction and generate a prediction result; the prediction result is the output object of the target overall model; the input object of the target overall model is the specific input content on the selected execution server; Verify the prediction result of the target overall model based on the result verification knowledge base and output a verification result; the verification result includes: a low-confidence result and a high-confidence result; the high-confidence result is used to generate a high-confidence result set; Input the object combination corresponding to the low-confidence result into the trained result verification model for verification and output a verification model prediction result; the verification model prediction result includes: a low-confidence prediction result and a high-confidence prediction result; Generate a medium-confidence result set according to the high-confidence prediction result; generate a low-confidence result set according to the low-confidence prediction result; Use the medium-confidence result set and the high-confidence result set to determine the final prediction result.
4. The interpretable biological information data processing method according to claim 3, wherein The steps of verifying the prediction result of the target overall model based on the result verification knowledge base and outputting a verification result include: Determine whether there is a connection path between the input object and the output object of the target overall model in the relationship knowledge base; If not, determine that the prediction result belongs to the low-confidence result; If so, determine whether the connection path exists in the causal relationship knowledge base and whether the direction of the causal relationship is consistent; If so, determine that the prediction result belongs to the high-confidence result; if not, determine that the prediction result belongs to the low-confidence result.
5. An interpretable biological information data processing system, characterized in that, The system includes: A training module for generating an analysis and decision training database based on a pre-constructed biological information data processing knowledge base; the analysis and decision training database is used to train an analysis and decision data model to generate a trained analysis and decision data model; wherein, the framework of the biological information data processing knowledge base includes several first-level categories and multiple sub-level categories; the latter sub-level category is obtained by subdividing the former sub-level category; An analysis data execution flow generation module for making a prediction using the trained analysis and decision data model according to the user data processing requirement and generating an analysis data execution flow; the analysis data execution flow includes several biological information data analysis links and corresponding sequences; the data analysis link includes a data processing link and a modeling link; A generation module for generating an interaction frequency parameter between the execution server and the core server using a pre-generated recommendation decision model; the recommendation decision model includes predefined rules and a machine learning model trained based on server operation data; A target overall model construction module for performing corresponding model construction based on the modeling link of the analysis data execution flow to generate a target overall model; A building module for: constructing a result verification knowledge base based on a pre-constructed bioinformatics data processing knowledge base; the result verification knowledge base includes: a relationship knowledge base and a causal relationship knowledge base; constructing a result verification training database based on the causal relationship knowledge base; wherein, the result verification training database is used to train the result verification model to generate a trained result verification model; the data in the result verification training database includes: object combinations, result confidence levels, and relationship path sequences; the object combinations are composed of coding sequences of different entity objects. A prediction result verification module for using the trained result verification model to verify the prediction result of the target overall model and determining the confidence set to which the prediction result belongs; the confidence set includes: a high-confidence result set, a medium-confidence result set, and a low-confidence result set; wherein, the high-confidence result set and the medium-confidence result set are the final prediction results. The target overall model construction module is further used for: performing corresponding model construction based on the modeling link of the data analysis execution flow to generate the target overall model, and the steps include: each execution server performs corresponding model training for the modeling link, and if the interaction frequency parameter requirement is met during the training process, it sends the current model parameters to the core server; the core server generates overall model parameters based on the received current model parameters and sends them to each execution server; each execution server performs model training based on the overall model parameters, and repeats the process of parameter interaction until the model meets the corresponding standard, and determines the current model as the target overall model.
6. An electronic device, comprising a memory and a processor, wherein a computer program capable of running on the processor is stored in the memory, characterized in that When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 4 above.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores machine-executable instructions, and when the computer-executable instructions are called and run by the processor, the computer-executable instructions cause the processor to run the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Data processing method and device based on knowledge graph, equipment and storage medium
CN113779272A