Knowledge graph-based industry large model training method, equipment and medium
Through a knowledge graph-based method, customized training of large models is carried out for specific industries, which solves the problem of poor performance of general large models when dealing with industry professional knowledge, and achieves higher industry adaptability and training efficiency.
Patent Information
- Application Number
- CN202510106021.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-23
AI Technical Summary
The existing general model performs poorly when dealing with the professional knowledge, terms, norms and timeliness requirements of specific industries, resulting in limited effectiveness in industry intelligent applications.
Using a knowledge graph-based method, we determine the target industry, obtain relevant unstructured data, use natural language processing algorithms to generate industry knowledge graphs, and use graph neural networks to embed the graphs into large models, combining knowledge injection and multi-task learning algorithms for training.
It improves the adaptability and generalization ability of large models in specific industries, improves the accuracy and practicality of models in industry applications, and reduces training costs.
Smart Images

Figure CN120031072A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of artificial intelligence, and in particular to a method, device and medium for training a large industry model based on a knowledge graph. Background Art
[0002] With the rapid development of artificial intelligence technology, large models based on deep learning have made breakthrough progress in natural language processing, machine translation, image recognition and other fields. However, when applying general large models to specific industries (such as medical, financial, legal, etc.), people find that these models are often difficult to adapt to the industry-specific knowledge system and needs. Existing large models, such as GPT and BERT, although they perform well in processing general text tasks, are unable to cope with industry-specific terminology, specifications, timeliness requirements, etc. due to the lack of knowledge understanding and reasoning capabilities for specific industries. This limitation seriously restricts the actual effect of general large models in industry intelligent applications.
[0003] Specifically, industry knowledge is highly professional, complex, and time-sensitive. For example, in the medical field, information such as disease names, drug efficacy, and treatment plans are constantly updated; in the financial field, market data, policies and regulations are also changing rapidly. Due to the lack of effective integration and understanding of these industry-specific knowledge, general large models often find it difficult to accurately answer industry-related questions or provide intelligent decision-making support.
[0004] Therefore, how to build a large model that can be customized for specific industries to improve the performance of the large model in industry applications has become a technical problem that needs to be solved urgently. Summary of the invention
[0005] The embodiments of the present application provide a method, device and medium for training an industry big model based on a knowledge graph to solve the following technical problem: how to build a big model that can be customized for a specific industry to improve the performance of the big model in industry applications.
[0006] In a first aspect, an embodiment of the present application provides 1. a method for training an industry big model based on a knowledge graph, characterized in that the method comprises: determining the target industry required for the big model to be trained, and obtaining unstructured data related to the target industry; wherein the unstructured data comprises at least one of the following: text, image, audio and video, and the big model to be trained is a big model capable of realizing general tasks, and the general tasks comprise at least one of the following: question answering, text generation and image generation; processing the unstructured data based on a preset natural language processing algorithm to generate an industry knowledge graph related to the target industry; wherein the natural language processing algorithm comprises at least one of the following: a named entity recognition algorithm, a relationship extraction algorithm and a semantic analysis algorithm; embedding the industry knowledge graph as input data into the big model to be trained based on a preset graph neural network algorithm; training the big model to be trained based on the industry knowledge graph and preset general knowledge using knowledge injection and multi-task learning algorithms; and outputting the big model when the big model to be trained meets preset conditions.
[0007] In one implementation of the present application, unstructured data is processed based on a preset natural language processing algorithm to generate an industry knowledge graph related to the target industry, specifically including: identifying entities related to the target industry in unstructured data based on a named entity recognition algorithm; extracting relationships between entities from unstructured data based on a relationship extraction algorithm; constructing a preliminary industry knowledge graph based on entities and relationships; processing the preliminary industry knowledge graph based on a semantic analysis algorithm to supplement the industry knowledge graph to generate an industry knowledge graph.
[0008] In one implementation of the present application, a preliminary industry knowledge graph is processed based on a semantic analysis algorithm to supplement the industry knowledge graph to generate an industry knowledge graph, specifically including: analyzing the contextual relationships between entities based on a semantic analysis algorithm to identify implicit associations and attributes, and generating semantic analysis results; based on the semantic analysis results, adding missing entities, relationships, and attributes to the preliminary industry knowledge graph to supplement the preliminary industry knowledge graph.
[0009] In one implementation of the present application, the industry knowledge graph is embedded as input data into the large model to be trained based on a preset graph neural network algorithm, specifically including: representing the industry knowledge graph as graph structure data; wherein entities are used as nodes and relationships are used as edges; performing feature extraction on the graph structure data based on the graph neural network algorithm, and converting the attribute information of the nodes and the relationship information of the edges into training features represented by high-dimensional vectors through node embedding and edge embedding; and embedding the training features as input data into the large model to be trained.
[0010] In one implementation of the present application, feature extraction is performed on graph structure data based on a graph neural network algorithm, and the attribute information of the node and the relationship information of the edge are converted into training features represented by high-dimensional vectors through node embedding and edge embedding. Specifically, the node features are aggregated layer by layer based on a preset multi-layer graph convolutional network to fuse the information of neighboring nodes; the edge features are encoded based on the attention mechanism to integrate the edge information into the node embedding; the node embedding and edge embedding are processed through a preset graph aggregation function to generate graph-level training features.
[0011] In one implementation of the present application, based on the industry knowledge graph and preset general knowledge training, knowledge injection and multi-task learning algorithms are used to train the large model to be trained, specifically including: integrating the entities, relationships and attributes in the industry knowledge graph into the training of the large model to be trained based on the knowledge injection algorithm; constructing a multi-task learning algorithm, and combining the industry knowledge graph and general knowledge for joint training.
[0012] In one implementation of the present application, when the large model to be trained meets the preset conditions, the large model is output: a model performance evaluation indicator is constructed; wherein the performance evaluation indicator of the large model to be trained includes at least one of the following: accuracy, recall rate, and industry-specific evaluation indicators; a performance evaluation is performed on the model to be trained based on a validation set in the input data, and whether the model meets the preset conditions is judged according to the evaluation results; when the model meets the preset conditions, the trained large model is output, and deployed and applied; when the model does not meet the preset conditions, the parameters of the large model to be trained are adjusted and the model training continues until the conditions are met.
[0013] In one implementation of the present application, the method also includes: when the industry knowledge graph is updated, updating the training features based on the graph neural network algorithm; integrating the updated training features into the large model, and updating the parameters of the large model to be trained based on a preset incremental learning algorithm.
[0014] In a second aspect, an embodiment of the present application further provides an industry big model training device based on a knowledge graph, the device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by at least one processor, and the instructions are executed by at least one processor so that the at least one processor can: determine the target industry required for the big model to be trained, and obtain unstructured data related to the target industry; wherein the unstructured data includes at least one of the following: text, image, audio and video, and the big model to be trained is a big model capable of achieving general tasks, and the general tasks include at least one of the following: question answering, text generation and image generation; based on a preset natural language processing algorithm, the unstructured data is processed to generate an industry knowledge graph related to the target industry; wherein the natural language processing algorithm includes at least one of the following: a named entity recognition algorithm, a relationship extraction algorithm and a semantic analysis algorithm; based on a preset graph neural network algorithm, the industry knowledge graph is embedded as input data into the big model to be trained; based on the industry knowledge graph and the preset general knowledge training, the big model to be trained is trained using knowledge injection and multi-task learning algorithms; when the big model to be trained meets the preset conditions, the big model is output.
[0015] In a third aspect, an embodiment of the present application further provides a non-volatile computer storage medium for training an industry big model based on a knowledge graph, storing computer executable instructions, characterized in that the computer executable instructions are set to: determine the target industry required for the big model to be trained, and obtain unstructured data related to the target industry; wherein the unstructured data includes at least one of the following: text, image, audio and video, and the big model to be trained is a big model that can realize general tasks, and the general tasks include at least one of the following: question answering, text generation and image generation; based on a preset natural language processing algorithm, the unstructured data is processed to generate an industry knowledge graph related to the target industry; wherein the natural language processing algorithm includes at least one of the following: named entity recognition algorithm, relationship extraction algorithm and semantic analysis algorithm; based on a preset graph neural network algorithm, the industry knowledge graph is embedded as input data into the big model to be trained; based on the industry knowledge graph and the preset general knowledge training, the knowledge injection and multi-task learning algorithms are used to train the big model to be trained; when the big model to be trained meets the preset conditions, the big model is output.
[0016] The embodiment of the present application provides a method, device and medium for training a large industry model based on a knowledge graph, which at least includes the following technical effects:
[0017] Improve the industry adaptability of the model: By identifying the target industry and obtaining relevant unstructured data, the trained large model can better adapt to the needs of specific industries and improve the accuracy and practicality of the model in industry applications;
[0018] Enhance the generalization ability of the model: Generate industry knowledge graphs based on natural language processing algorithms, integrate industry knowledge into the model in a structured form, enhance the model's ability to understand and apply industry knowledge, thereby improving the generalization ability of the model and enabling it to handle more diverse tasks;
[0019] Improve model training efficiency: By embedding the industry knowledge graph into the large model to be trained through the graph neural network algorithm, the model training process is accelerated and the training cost is reduced. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used in the present application and do not constitute an improper limitation on the present application. In the drawings:
[0021] Figure 1 A flow chart of a method for training a large industry model based on a knowledge graph provided in an embodiment of the present application;
[0022] Figure 2 A schematic diagram of the internal structure of an industry large model training device based on a knowledge graph provided in an embodiment of the present application. DETAILED DESCRIPTION
[0023] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the present application will be clearly and completely described below in combination with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.
[0024] The embodiments of the present application provide a method, device and medium for training an industry big model based on a knowledge graph to solve the following technical problem: how to build a big model that can be customized for a specific industry to improve the performance of the big model in industry applications.
[0025] The technical solution proposed in the embodiments of the present application is described in detail below with reference to the accompanying drawings.
[0026] Figure 1 A flowchart of industry large model training based on knowledge graph is provided in the embodiment of this application. Figure 1 As shown, the industry large model training method based on the knowledge graph provided in the embodiment of the present application specifically includes the following steps:
[0027] Step 1. Determine the target industry required for the large model to be trained and obtain unstructured data related to the target industry; the unstructured data includes at least one of the following: text, image, audio and video, and the large model to be trained is a large model that can achieve general tasks, and the general tasks include at least one of the following: question answering, text generation and image generation.
[0028] First, we need to identify the specific industry field that the large model to be trained is targeting, that is, the target industry. Taking the medical industry as an example, the medical industry covers multiple sub-fields such as disease diagnosis, drug development, health management, and medical consultation.
[0029] After determining that the target industry is medical, obtain unstructured data related to the industry. Unstructured data refers to data that does not have a fixed format or predefined structure. It exists in a natural form and is difficult to use directly in traditional databases or data analysis tools. In the medical industry, unstructured data includes at least one or more of the following types:
[0030] Text data: such as medical literature, clinical records, doctors’ notes, patient feedback, etc.
[0031] Image data: such as medical images (X-rays, CT scans, MRI images, etc.), pathological slice images, etc.
[0032] Audio data: such as recordings of doctor-patient conversations, heart sound records, etc.
[0033] Video data: such as surgical videos, patient recovery process records, etc.
[0034] Ways to obtain these unstructured data include collecting them from hospitals, medical institutions, medical databases, online medical platforms and other channels, while complying with relevant data privacy and regulatory requirements.
[0035] In the embodiment of this application, a large medical model is trained, which needs to be able to handle general tasks such as medical question answering, generating medical text, and generating medical image descriptions. First, the target industry is determined to be medical. Then, a large amount of unstructured data is collected from multiple hospitals and medical institutions, including:
[0036] Text data: Medical literature, clinical records, and physician notes from the past decade were collected.
[0037] Image data: Millions of medical images, including X-rays, CT scans, and MRI images, are acquired.
[0038] Audio data: Thousands of hours of audio recordings of doctor-patient conversations and heart sound recordings.
[0039] Video data: Hundreds of surgical videos and patient recovery process records were collected.
[0040] Step 2: Process the unstructured data based on a preset natural language processing algorithm to generate an industry knowledge graph related to the target industry; wherein the natural language processing algorithm includes at least one of the following: a named entity recognition algorithm, a relationship extraction algorithm, and a semantic analysis algorithm.
[0041] Step 21: Identify entities related to the target industry in unstructured data based on a named entity recognition algorithm.
[0042] Named Entity Recognition (NER) algorithm: Named entity recognition is a natural language processing technology used to identify entities with specific meanings from text, such as names of people, places, and organizations.
[0043] In the embodiment of the present application, the NER algorithm identifies entities such as disease names, drug names, symptom names, etc. that are closely related to the target industry.
[0044] Apply pre-trained NER models or NER models customized for the medical industry to scan unstructured text data, identify and annotate all medical-related entities. These entities will become the basic elements for subsequent knowledge graph construction.
[0045] Step 22: Extract the relationship between entities from the unstructured data based on the relationship extraction algorithm.
[0046] Relation extraction algorithm: Relation extraction is used to identify associations or relationships between entities from text. In the medical field, these relationships include "disease-symptoms", "drug-treatment", "disease-cause", etc.
[0047] Through the relationship extraction algorithm, the identified medical entities are further analyzed to extract various relationships between medical entities.
[0048] Step 23: Build a preliminary industry knowledge graph based on entities and relationships.
[0049] Preliminary industry knowledge graph construction: After obtaining the relationship between medical entities, a preliminary medical industry knowledge graph is constructed. This graph will display various elements in the medical field in the form of entities and connect them through relationships to form a structured knowledge system.
[0050] Using graph building tools or programming languages, the identified entities and relationships are organized in a predetermined format to form a preliminary medical industry knowledge graph. This graph can be a directed or undirected graph, where nodes represent entities and edges represent relationships.
[0051] Step 24: Process the preliminary industry knowledge graph based on the semantic analysis algorithm to supplement the industry knowledge graph to generate the industry knowledge graph.
[0052] The initially constructed knowledge graph is deepened and improved through semantic analysis algorithms to capture more implicit information and connections.
[0053] Step 241: Analyze the contextual relationship between entities based on a semantic analysis algorithm to identify implicit associations and attributes, and generate a semantic analysis result.
[0054] Semantic analysis algorithm: Semantic analysis is used to understand the deep meaning and context of text. In the medical industry, semantic analysis can help identify implicit associations and attributes between entities, such as the severity of a disease, the side effects of a drug, etc.
[0055] The semantic analysis algorithm is used to analyze the initially constructed medical industry knowledge graph to identify the implicit associations and attributes between entities. These analysis results will be output in a structured form to provide a basis for subsequent knowledge graph supplementation.
[0056] Step 242: Based on the semantic analysis results, add missing entities, relationships, and attributes to the preliminary industry knowledge graph to supplement the preliminary industry knowledge graph.
[0057] Supplementing the knowledge graph: Based on the results of semantic analysis, we can find possible missing parts in the initially constructed knowledge graph, such as missing entities, relationships, or attributes. These missing parts need to be supplemented by adding new nodes and edges.
[0058] According to the results of semantic analysis, the missing entities, relationships, and attributes are manually or automatically added to the preliminary medical industry knowledge graph. It should be noted that this process may require manual review and confirmation to ensure the accuracy and reliability of the added information.
[0059] In an embodiment of the present application, the NER algorithm of the medical industry is applied to scan the collected text data to identify and extract all medical-related entities, such as disease names, drug names, symptom descriptions, etc. Based on the relationship extraction algorithm, the identified medical entities are further analyzed to extract various relationships between them, such as "disease-symptoms" and "drug-treatment". The identified entities and relationships are organized in a predetermined format to form a preliminary medical industry knowledge graph. This graph displays various elements in the medical field in the form of nodes and is connected by edges. The semantic analysis algorithm is applied to analyze the preliminary constructed medical industry knowledge graph to identify implicit associations and attributes between entities, such as disease transmission routes, drug interactions, etc. According to the results of the semantic analysis, the missing entities, relationships and attributes are manually or automatically added to the preliminary medical industry knowledge graph to improve the structure and content of the knowledge graph.
[0060] Step 3: Based on the preset graph neural network algorithm, the industry knowledge graph is embedded as input data into the large model to be trained.
[0061] Based on the graph neural network (GNN) algorithm, the constructed industry knowledge graph is effectively embedded into the large model to be trained, so that the model can fully rely on the structured knowledge in the graph to improve its performance on specific industry tasks.
[0062] Step 31: Represent the industry knowledge graph as graph structure data, wherein entities are regarded as nodes and relationships are regarded as edges.
[0063] Graph structured data: Graph structured data is a data format used to represent entities and their relationships, where entities are represented as nodes and relationships are represented as edges. This data structure can intuitively display the complex relationships between entities and is the main object of graph neural network processing.
[0064] In the medical industry knowledge graph, entities such as diseases, drugs, and symptoms are used as nodes, and relationships such as "disease-symptoms" and "drug-treatment" are used as edges to form graph structure data. This representation method can retain the structured characteristics of medical knowledge and provide a basis for subsequent graph neural network processing.
[0065] Step 32: Extract features of the graph structure data based on the graph neural network algorithm, and convert the attribute information of the nodes and the relationship information of the edges into training features represented by high-dimensional vectors through node embedding and edge embedding.
[0066] Through the graph neural network algorithm, deep feature extraction is performed on the graph structure data, and the attribute information of the nodes and the relationship information of the edges are converted into training features represented by high-dimensional vectors so that the large model can understand and be based on these structured knowledge.
[0067] Node embedding and edge embedding: Node embedding is the process of converting the attribute information of a node into a high-dimensional vector, while edge embedding is the process of converting the relationship information of an edge into a vector. Through embedding, entities and relationships in graph structured data can be converted into numerical forms that can be processed by the model.
[0068] Step 321: Based on a preset multi-layer graph convolutional network, node features are aggregated layer by layer to fuse the information of neighboring nodes.
[0069] Multi-layer graph convolutional network: Graph convolutional network (GCN) is a neural network specially designed for processing graph structure data. It can update the feature representation of nodes by aggregating the information of neighbor nodes layer by layer. Multi-layer GCN can capture multi-order neighbor relationships in the graph structure, thereby extracting richer features.
[0070] In the medical industry knowledge graph, multi-layer GCN is applied to aggregate node features layer by layer. For each node (such as a disease node), GCN considers the information of its neighboring nodes (such as related symptom nodes) and fuses the information of neighboring nodes into the feature representation of the current node through weighted summation and other methods. In this way, the feature representation of each node contains the information of its neighboring nodes, thereby capturing the local associations in the graph structure.
[0071] Step 322: Encode edge features based on the attention mechanism to incorporate edge information into node embedding.
[0072] Attention mechanism: Attention mechanism is a mechanism that can dynamically assign different weights to different input information. In graph neural networks, the attention mechanism can be used to encode edge features, that is, to assign different weights to edges according to their importance or relevance.
[0073] In the medical industry knowledge graph, the attention mechanism is applied to encode edge features. For each edge (such as the "disease-symptom" relationship), its importance or relevance score is calculated, and then a weight is assigned to the edge based on the score. In the subsequent node embedding or graph-level feature extraction process, important edges will receive more attention, thereby improving the model's ability to capture key information.
[0074] Step 323: Process node embedding and edge embedding through a preset graph aggregation function to generate graph-level training features.
[0075] Graph aggregation function: A graph aggregation function is a function used to integrate node embeddings and edge embeddings into graph-level features. The graph aggregation function can aggregate the feature representations of all nodes in the graph or the feature representations of a specific subgraph into an overall feature vector for subsequent model training or prediction.
[0076] In the medical industry knowledge graph, graph aggregation functions are applied to process node embedding and edge embedding. Specifically, the feature representations of all nodes can be averaged or weighted summed to obtain a graph-level feature vector. This feature vector contains information about the entire medical industry knowledge graph and can be used as input data for large models for training or prediction.
[0077] Step 4: Based on the industry knowledge graph and preset general knowledge training, use knowledge injection and multi-task learning algorithms to train the large model to be trained.
[0078] By combining industry knowledge graphs and general knowledge, using knowledge injection and multi-task learning algorithms, the large model to be trained is trained to enhance the model's understanding and application capabilities in specific industry fields.
[0079] Step 41: Based on the knowledge injection algorithm, the entities, relationships and attributes in the industry knowledge graph are integrated into the training of the large model to be trained.
[0080] Knowledge injection algorithm: The knowledge injection algorithm is a technology that incorporates structured information from external knowledge sources (such as knowledge graphs) into the model training process. Through this algorithm, the model can learn rich information such as entities, relationships, and attributes in the knowledge graph, thereby enhancing its performance on specific tasks.
[0081] Entity integration: Entities in the industry knowledge graph (such as diseases, drugs, symptoms, etc. in the medical field) are used as one of the input features for model training. By associating these entities with words or concepts in the model, the model can understand and recognize the appearance of these entities in text or dialogue.
[0082] Relationship integration: Convert the relationships in the industry knowledge graph (such as "disease-symptoms", "drug-treatment", etc.) into semantic relationships that the model can understand. This can be achieved by introducing relationship constraints or regularization terms during model training, so that the model can capture the relationships between entities while learning entity features.
[0083] Attribute integration: Attributes in the industry knowledge graph (such as the severity of the disease, the side effects of the drug, etc.) are used as additional information or features of the model. These attributes can provide the model with a more fine-grained understanding and help it make more accurate judgments on specific tasks.
[0084] Step 42: Build a multi-task learning algorithm and perform joint training by combining industry knowledge graph and general knowledge.
[0085] Multi-task learning algorithm: Multi-task learning is a machine learning method that shares representations and improves learning efficiency by learning multiple related tasks at the same time. In the training that combines industry knowledge graphs and general knowledge, the multi-task learning algorithm can fully leverage the correlation between different tasks to improve the overall performance of the model.
[0086] Task definition: Based on the application scenarios of the medical big model, multiple related tasks are defined, such as medical question answering, disease diagnosis, drug recommendation, etc. Each task has its specific input and output, but shares the same model representation or feature space.
[0087] Joint training: During model training, the objectives and loss functions of multiple tasks are considered simultaneously. By optimizing the weighted sum of these objectives and loss functions, the model can achieve better performance on multiple tasks. At the same time, additional information and constraints are provided to the model based on industry knowledge graphs and general knowledge to help the model better understand and deal with problems in the medical field.
[0088] Knowledge fusion: In the multi-task learning process, the structured information in the industry knowledge graph is fused with the unstructured information in the general knowledge. This can be achieved by adding a knowledge graph embedding layer or an attention mechanism to the model, enabling the model to reason and make decisions based on both types of knowledge at the same time.
[0089] Step 5: When the large model to be trained meets the preset conditions, output the large model.
[0090] By building a scientific model performance evaluation system, a comprehensive performance evaluation is conducted on the large model to be trained, and a decision is made based on the evaluation results whether to output the trained large model.
[0091] Step 51: construct model performance evaluation indicators; wherein the large model performance evaluation indicators of the model to be trained include at least one of the following: accuracy, recall rate, and industry-specific evaluation indicators.
[0092] Model performance evaluation metrics: Model performance evaluation metrics are standards for measuring how well a model performs on a specific task.
[0093] Accuracy: Accuracy is an indicator that measures the consistency between the model's prediction results and the actual results. In large medical models, accuracy can reflect the model's predictive ability for tasks such as disease diagnosis and drug recommendation.
[0094] Recall: Recall is a measure of the proportion of positive samples that the model can correctly identify to all actual positive samples.
[0095] Industry-specific evaluation indicators: In addition to the general accuracy and recall rate, the medical field may also have some specific evaluation indicators, such as F1 score (the harmonic mean of accuracy and recall rate), AUC-ROC curve (area under the receiver operating characteristic curve), etc. These indicators can more comprehensively reflect the performance of the model in medical tasks.
[0096] Step 52: Perform a performance evaluation on the model to be trained based on the validation set in the input data, and determine whether the model meets the preset conditions based on the evaluation results.
[0097] Validation set: The validation set is a portion of data separated from the original dataset and used for performance evaluation during model training to verify the generalization ability of the model.
[0098] Select a validation set: Select a representative portion of data from the medical dataset as the validation set to ensure that the validation set covers various situations in the medical field.
[0099] Performance evaluation: Use the constructed model performance evaluation indicators to evaluate the model prediction results on the validation set. Specifically, you can calculate indicators such as the accuracy, recall, F1 score, and AUC-ROC curve of the model on the validation set.
[0100] Determine model performance: Compare the evaluation results with the preset conditions to determine whether the model meets the expected performance standards. If the model performance meets the preset conditions, proceed to the next step; otherwise, adjust the model parameters and continue training.
[0101] Step 53: When the model meets the preset conditions, the trained large model is output and deployed and applied.
[0102] When the performance evaluation results of the model on the validation set meet the preset conditions, the trained large model is output to a deployable format. The output large model is deployed to actual medical applications, such as online medical consultation platforms, intelligent auxiliary diagnosis systems, etc., to provide intelligent services in the medical field.
[0103] Step 54: When the model does not meet the preset conditions, adjust the parameters of the large model to be trained and continue to train the model until the conditions are met.
[0104] When the performance evaluation results of the model on the validation set do not meet the preset conditions, it is necessary to conduct an in-depth analysis of the model to find out the performance bottleneck. This may be caused by unreasonable model structure, improper parameter settings, or insufficient training data. According to the performance analysis results, adjust the parameters of the model. For example, you can increase the number of model layers, adjust the learning rate, optimize the loss function, etc. to improve the performance of the model. Use the adjusted parameters to continue training the model until the performance evaluation results of the model on the validation set meet the preset conditions.
[0105] It is understandable that the unstructured data about the target industry will be updated, so the big model needs to be updated, which includes the following steps:
[0106] A1. When the industry knowledge graph is updated, the training features are updated based on the graph neural network algorithm.
[0107] Establish a monitoring mechanism to check the updates of the industry knowledge graph regularly or in real time.
[0108] When an update to the knowledge graph is detected, the subsequent feature update process is triggered.
[0109] Based on the graph neural network algorithm, new node features, edge features and global structure information are extracted from the updated knowledge graph, and the extracted new knowledge features are fused with the original training features to form an updated training feature set.
[0110] A2. Integrate the updated training features into the large model, and update the parameters of the large model to be trained based on the preset incremental learning algorithm.
[0111] Incremental learning algorithm: Incremental learning is a method of updating model parameters by gradually adding new data or new knowledge during model training. It can effectively reduce the cost and time of retraining based on the existing model.
[0112] Incorporate the updated training feature set in step A1 into the training process of the large model. Select a suitable incremental learning algorithm based on the specific situation and training requirements of the large model. Based on the selected incremental learning algorithm, update the parameters of the large model according to the updated training features. After updating the model parameters, perform performance verification on the updated large model. By comparing the accuracy, recall and other indicators of the model before and after the update, evaluate the effect of the new knowledge on the performance of the model.
[0113] The above is an embodiment of the method proposed in this application. Based on the same inventive concept, the embodiment of this application also provides an industry large model training device based on a knowledge graph, and its structure is as follows Figure 2 shown.
[0114] Figure 2 A schematic diagram of the internal structure of a large industry model training device based on a knowledge graph provided in an embodiment of the present application. Figure 2 As shown, the device includes:
[0115] at least one processor 201;
[0116] and, a memory 202 communicatively connected to the at least one processor;
[0117] The memory 202 stores instructions that can be executed by at least one processor, and the instructions are executed by at least one processor 201 to enable at least one processor 201 to:
[0118] Determine the target industry required for the large model to be trained, and obtain unstructured data related to the target industry; wherein the unstructured data includes at least one of the following: text, image, audio and video, and the large model to be trained is a large model that can achieve general tasks, and the general tasks include at least one of the following: question answering, text generation and image generation; process the unstructured data based on a preset natural language processing algorithm to generate an industry knowledge graph related to the target industry; wherein the natural language processing algorithm includes at least one of the following: named entity recognition algorithm, relationship extraction algorithm and semantic analysis algorithm; embed the industry knowledge graph as input data into the large model to be trained based on a preset graph neural network algorithm; based on the industry knowledge graph and preset general knowledge training, use knowledge injection and multi-task learning algorithms to train the large model to be trained; when the large model to be trained meets the preset conditions, output the large model.
[0119] Some embodiments of the present application provide corresponding Figure 1 A non-volatile computer storage medium for training a large industry model based on a knowledge graph, storing computer executable instructions, wherein the computer executable instructions are set as:
[0120] Determine the target industry required for the large model to be trained, and obtain unstructured data related to the target industry; wherein the unstructured data includes at least one of the following: text, image, audio and video, and the large model to be trained is a large model that can achieve general tasks, and the general tasks include at least one of the following: question answering, text generation and image generation; process the unstructured data based on a preset natural language processing algorithm to generate an industry knowledge graph related to the target industry; wherein the natural language processing algorithm includes at least one of the following: named entity recognition algorithm, relationship extraction algorithm and semantic analysis algorithm; embed the industry knowledge graph as input data into the large model to be trained based on a preset graph neural network algorithm; based on the industry knowledge graph and preset general knowledge training, use knowledge injection and multi-task learning algorithms to train the large model to be trained; when the large model to be trained meets the preset conditions, output the large model.
[0121] Each embodiment in this application is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the IoT device and medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.
[0122] The system and medium provided in the embodiments of the present application correspond one-to-one to the method. Therefore, the system and medium also have similar beneficial technical effects to the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the system and medium will not be repeated here.
[0123] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.
[0124] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0125] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0126] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0127] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0128] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0129] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0130] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0131] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.
Claims
1. A method for training a large industry model based on a knowledge graph, characterized in that: The method comprises: Determine the target industry required for the large model to be trained, and obtain unstructured data related to the target industry; wherein the unstructured data includes at least one of the following: text, image, audio, and video; the large model to be trained is a large model capable of implementing general tasks, and the general tasks include at least one of the following: question answering, text generation, and image generation; Processing the unstructured data based on a preset natural language processing algorithm to generate an industry knowledge graph related to the target industry; wherein the natural language processing algorithm includes at least one of the following: a named entity recognition algorithm, a relationship extraction algorithm, and a semantic analysis algorithm; Based on a preset graph neural network algorithm, the industry knowledge graph is embedded as input data into the large model to be trained; Based on the industry knowledge graph and preset general knowledge training, the large model to be trained is trained using knowledge injection and multi-task learning algorithms; When the large model to be trained meets a preset condition, the large model is output.
2. According to claim 1, a method for training a large industry model based on a knowledge graph is characterized in that: The unstructured data is processed based on a preset natural language processing algorithm to generate an industry knowledge graph related to the target industry, specifically including: Identifying entities related to the target industry in the unstructured data based on the named entity recognition algorithm; Extracting the relationship between the entities from the unstructured data based on the relationship extraction algorithm; Based on the entities and relationships, construct a preliminary industry knowledge graph; The preliminary industry knowledge graph is processed based on the semantic analysis algorithm to supplement the industry knowledge graph to generate the industry knowledge graph.
3. According to claim 2, a method for training a large industry model based on a knowledge graph is characterized in that: Processing the preliminary industry knowledge graph based on the semantic analysis algorithm to supplement the industry knowledge graph to generate the industry knowledge graph specifically includes: Analyzing the contextual relationships between the entities based on the semantic analysis algorithm to identify implicit associations and attributes, and generating semantic analysis results; Based on the semantic analysis results, missing entities, relationships, and attributes are added to the preliminary industry knowledge graph to supplement the preliminary industry knowledge graph.
4. According to claim 1, a method for training a large industry model based on a knowledge graph is characterized in that: Based on the preset graph neural network algorithm, the industry knowledge graph is embedded as input data into the large model to be trained, specifically including: The industry knowledge graph is represented as graph structure data, wherein the entities are used as nodes and the relationships are used as edges; Based on the graph neural network algorithm, the graph structure data is featured extracted. By embedding nodes and edges, the attribute information of nodes and the relationship information of edges are converted into training features represented by high-dimensional vectors. The training features are embedded as input data into the large model to be trained.
5. According to claim 4, a method for training a large industry model based on a knowledge graph is characterized in that: Based on the graph neural network algorithm, the graph structure data is feature extracted. By embedding nodes and edges, the attribute information of nodes and the relationship information of edges are converted into training features represented by high-dimensional vectors. Specifically, it includes: Based on the preset multi-layer graph convolutional network, node features are aggregated layer by layer to fuse the information of neighboring nodes; Encoding edge features based on an attention mechanism to incorporate edge information into the node embedding; The node embedding and edge embedding are processed by a preset graph aggregation function to generate graph-level training features.
6. The industry large model training method based on knowledge graph according to claim 1 is characterized in that: Based on the industry knowledge graph and the preset general knowledge training, the large model to be trained is trained using knowledge injection and multi-task learning algorithms, specifically including: Integrate the entities, relationships and attributes in the industry knowledge graph into the training of the large model to be trained based on the knowledge injection algorithm; Construct the multi-task learning algorithm and perform joint training in combination with the industry knowledge graph and the general knowledge.
7. The industry large model training method based on knowledge graph according to claim 6 is characterized in that: When the large model to be trained meets the preset conditions, the large model is output: Constructing model performance evaluation indicators; wherein the large model performance evaluation indicators of the model to be trained include at least one of the following: accuracy, recall rate, and industry-specific evaluation indicators; Perform performance evaluation on the training model based on the validation set in the input data, and determine whether the model meets the preset conditions based on the evaluation results; When the model meets the preset conditions, the trained large model is output and deployed and applied; When the model does not meet the preset conditions, the parameters of the large model to be trained are adjusted and the model is continued to be trained until the conditions are met.
8. According to claim 1, a method for training a large industry model based on a knowledge graph is characterized in that: The method further comprises: When the industry knowledge graph is updated, updating the training features based on the graph neural network algorithm; The updated training features are integrated into the large model, and the parameters of the large model to be trained are updated based on a preset incremental learning algorithm.
9. A knowledge graph-based industry large model training device, characterized in that: The device comprises: at least one processor; and, a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to: Determine the target industry required for the large model to be trained, and obtain unstructured data related to the target industry; wherein the unstructured data includes at least one of the following: text, image, audio, and video; the large model to be trained is a large model capable of implementing general tasks, and the general tasks include at least one of the following: question answering, text generation, and image generation; Processing the unstructured data based on a preset natural language processing algorithm to generate an industry knowledge graph related to the target industry; wherein the natural language processing algorithm includes at least one of the following: a named entity recognition algorithm, a relationship extraction algorithm, and a semantic analysis algorithm; Based on a preset graph neural network algorithm, the industry knowledge graph is embedded as input data into the large model to be trained; Based on the industry knowledge graph and preset general knowledge training, the large model to be trained is trained using knowledge injection and multi-task learning algorithms; When the large model to be trained meets a preset condition, the large model is output.
10. A non-volatile computer storage medium for training a large industry model based on a knowledge graph, storing computer executable instructions, characterized in that: The computer executable instructions are configured to: Determine the target industry required for the large model to be trained, and obtain unstructured data related to the target industry; wherein the unstructured data includes at least one of the following: text, image, audio, and video; the large model to be trained is a large model capable of implementing general tasks, and the general tasks include at least one of the following: question answering, text generation, and image generation; Processing the unstructured data based on a preset natural language processing algorithm to generate an industry knowledge graph related to the target industry; wherein the natural language processing algorithm includes at least one of the following: a named entity recognition algorithm, a relationship extraction algorithm, and a semantic analysis algorithm; Based on a preset graph neural network algorithm, the industry knowledge graph is embedded as input data into the large model to be trained; Based on the industry knowledge graph and preset general knowledge training, the large model to be trained is trained using knowledge injection and multi-task learning algorithms; When the large model to be trained meets a preset condition, the large model is output.