Industrial exclusive customized model generation method based on large model

By building a knowledge graph and using supervision fine-tuning and reinforcing learning algorithms to optimize the large language model, the problem of insufficient accuracy and relevance of question-and-answer in the professional field of large language models is solved, and more accurate and relevant question-and-answer capabilities are achieved.

CN120542597AActive Publication Date: 2025-08-26CHENGDU MINGTU TECH CO LTD

Patent Information

Application Number
CN202510612026.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-08-26
Estimated Expiration
2045-05-13

AI Technical Summary

Technical Problem

The existing large language models are difficult to provide accurate industry knowledge in highly professional industries such as medical care, law, and finance, resulting in low accuracy and relevance of Q&A.

Method used

By collecting multimodal data from the target industry, building a knowledge graph, performing knowledge representation, extraction and fusion, using supervised fine-tuning algorithms and reinforcement learning algorithms to optimize models, and combining the search enhancement technology of the knowledge graph to improve the relevance and accuracy of the question and answer.

Benefits of technology

It significantly improves the accuracy and relevance of Q&A in the target industry field of the large language model, generates answers that are more in line with user preferences and industry specifications, and enhances the application value of the model in the professional field.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120542597A_ABST
    Figure CN120542597A_ABST
Patent Text Reader

Abstract

The invention discloses an industry exclusive customization model generation method based on a large model, and the method comprises the following key steps: firstly, collecting structured and semi-structured multi-modal data containing key entities, concepts and relationships in a target industry; then, constructing a knowledge graph through a deep learning method; in a model construction stage, carefully selecting a pre-trained large language model, and carrying out personalized adjustment on feature engineering and a model architecture according to industry characteristics; on the basis of the constructed knowledge graph, a retrieval enhancement technology based on the knowledge graph is adopted, the answer content of the large language model is optimized, and target industry domain knowledge is injected; and finally, by using a reinforcement learning algorithm, optimizing model output by training a reward model. According to the method, specialized customization of the industry exclusive model is realized, the question and answer ability of the large language model in the target industry field is remarkably improved, and high accuracy and high correlation of the model when the model answers related questions of the target industry are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of industry big models and natural language processing technology, and specifically to a method for generating industry-specific customized models based on big models. This method constructs a knowledge graph for the target industry field, supervises and fine-tunes a customized big language model, injects the knowledge graph into the customized big language model using knowledge graph retrieval enhancement technology, and uses a reinforcement learning algorithm to optimize the customized model output. Background Art

[0002] Amid the rapid development of artificial intelligence, large language models, a key technology in natural language processing, have become capable of accurately answering questions. However, their training model, based on general-purpose datasets, has exposed significant limitations, particularly in highly specialized industries such as healthcare, law, and finance. General-purpose models often struggle to provide precise industry knowledge, resulting in low accuracy and relevance in answers for these applications. Ensuring that models output highly relevant answers to target industries during the question-and-answer generation process is a pressing issue. Summary of the Invention

[0003] The purpose of this invention is to propose an industry-specific customized model generation method based on a large model. The method collects and integrates multimodal data sets of the target industry, uses deep learning technology to build a knowledge graph, and includes three steps: knowledge representation, extraction and fusion; uses a supervised fine-tuning algorithm to build a customized model for the pre-trained large language model; uses a retrieval enhancement technology based on the knowledge graph to optimize the model answer; and finally optimizes the model output through a reinforcement learning algorithm to improve the relevance and accuracy of questions and answers in the target industry.

[0004] The purpose of the present invention is achieved through the following technical solutions:

[0005] A method for generating an industry-specific customized model based on a large model includes the following steps:

[0006] Step S1: Collect structured and semi-structured multimodal data such as documents and reports containing key entities, concepts, and relationships in the target industry;

[0007] Step S2: Construct a knowledge graph through deep learning methods, including three steps: knowledge representation, knowledge extraction, and knowledge fusion;

[0008] Step S3: During the model building phase, we select a pre-trained large language model, perform feature engineering and personalize the model architecture based on industry characteristics, and design a customized model that meets the characteristics of the target industry through meticulous parameter tuning and strict monitoring of the training process.

[0009] Step S4: Based on the constructed knowledge graph, the knowledge graph-based retrieval enhancement technology is used to optimize the answer content of the large language model and inject target industry domain knowledge;

[0010] Step S5: Utilize the reinforcement learning algorithm to optimize the model output by training the reward model and provide feedback on the best question-answering results.

[0011] Furthermore, the step S1 specifically includes:

[0012] Step S101: Obtain reports and public data sets published by target industry analysis institutions, build a structured relational database for the target industry, select an appropriate data set, and convert the structured data in the data set into entities and relationships between entities in the knowledge graph;

[0013] Step S102: Obtain semi-structured data and unstructured text corpus data of the target industry from web pages, online encyclopedia Wikipedia, online books, and literature. The specific process is as follows: obtain website links belonging to the target industry and collect them into a list, obtain unstructured text corpus data for each link in turn until the list is empty, filter, deduplicate and aggregate the data, and finally store it in the database.

[0014] Furthermore, the step S2 specifically includes:

[0015] Step S201: Use deep learning methods to perform knowledge representation on the multimodal data of the target industry obtained in S1. Deep learning algorithms can effectively integrate multimodal information, enhance the complementarity of multi-source heterogeneous information and the effectiveness of knowledge representation. Use a Transformer-based bidirectional encoder to pre-train the target industry dataset and generate word embeddings;

[0016] Step S202: Use the word embedding vector generated by the Transformer-based bidirectional encoder as the input of the bidirectional long short-term memory conditional random field model to perform entity recognition. The specific steps are as follows:

[0017] (1) The word vector output by the Transformer's bidirectional encoder model is used as the input of the bidirectional long short-term memory network at each time step;

[0018] (2) At each time step t, the forward long short-term memory network layer processes the sequence from time step 1 to time step t to obtain the forward hidden state sequence

[0019] (3) Use the reverse long short-term memory network layer to process the sequence from time step t to time step 1 to obtain the reverse hidden state sequence

[0020] (4) Concatenate the forward and reverse hidden states of each time step to obtain a comprehensive hidden state sequence

[0021] (5) Use the mapping matrix Q of the linear output layer to transform the hidden state sequence H t Mapped to n-dimensional space, n represents the number of label types, the mapping matrix Q = (Q1, Q2, ..., Q i )∈R i*n , Q i Represents the score of the i-th category relative to n labels, and R represents a real number;

[0022] (6) The conditional random field decoder is used to learn the dependency between entity labels and the relationship between adjacent entity labels is used to obtain the optimal entity recognition sequence, thereby reducing the probability of unreasonable entity recognition sequences appearing in the prediction sequence.

[0023] The specific working principle of the above conditional random field decoder is as follows:

[0024] The conditional random field decoder constrains the rationality of the label sequence by learning the transition probability between labels, and uses the overall information of the label sequence to optimize the output of the model. The role of the conditional random field decoder in the model is to convert the features output by the bidirectional long short-term memory network layer into the optimal entity sequence; the input sentence is X = (x1, x2, ..., x n ), the label sequence Y corresponding to X=(y1,y2,…,y n ), the conditional random field decoder calculates the label sequence path score formula as follows:

[0025]

[0026] Among them, n represents the length of the sequence, y i represents the label of the i-th position in the sequence, is the transition probability matrix from label y i Move to label y i+1 The probability of Indicates that the i-th position in the output probability matrix of the bidirectional long short-term memory network layer is marked as label y i probability;

[0027] After the conditional random field decoder calculates the path scores of all possible label sequences, it selects the label sequence with the highest score as the final recognition result. The calculation formula for the best label sequence is as follows:

[0028] Y*=argmax{Score(X,Y)}

[0029] The conditional random field decoder constrains the rationality of the label sequence by learning the transition probability between labels, learns the label sequence restriction rules, and reduces the probability of unreasonable sequences.

[0030] Step S203: Based on entity recognition, perform relationship extraction on the entity;

[0031] The specific principle of entity relationship extraction in step S203 is:

[0032] For a given sentence S={c1,c2,…,c n There are two target entities e1 and e2. The formula for calculating the sentence vector is:

[0033] M0′=W0·tanh(M0)+b0

[0034] Among them, M0 represents the final hidden layer vector of the Transformer's bidirectional encoder model, W0 is the weight matrix, and b0 is the bias vector;

[0035] Calculate the entity vector for each target entity and take the average value of the word vector in the entity vector. The calculation formula is:

[0036]

[0037] Among them, l s , l e Indicates the start and end position of the entity, M i Represents word vector, i≥1;

[0038] Finally, the two entity vectors and the sentence vector are concatenated and then input into a fully connected layer to obtain the output vector M' as follows:

[0039] M″=W1·concat(M0′,M i ′,M j ′)+b1

[0040] Where W1∈R p*3d , R represents a real number, where p is the total number of relation types, d represents the hidden layer dimension of each word vector in the Transformer encoder; M0' represents the final vector of sentence S, M i ' and M j 'represents the target entity e i and e j The final entity vector of , b1 is the bias vector;

[0041] Finally, M' is input into a SoftMax layer to classify the relationship type and obtain the relationship between the target entities. The formula is as follows:

[0042] R(ei , e j ) = SoftMax(M″)

[0043] Step S204: Based on entity recognition, extract events from multi-modal data, and then construct an event database and store it in the knowledge graph database.

[0044] Furthermore, step S3 specifically includes:

[0045] Step S301: Select a suitable pre-trained large language model according to the actual tasks of the target industry;

[0046] Step S302: Inject the knowledge graph in the knowledge graph database obtained in step S2 into the large language model for supervised fine-tuning to adapt to the Q&A of the target industry;

[0047] Step S303: Through the supervised fine-tuning algorithm, without changing the original weights of the pre-trained model, quickly adjust the model through a low-rank matrix to adapt to the new task, reduce the number of parameters that need to be fine-tuned, speed up the training speed and reduce the demand for computing resources, and adapt to the specific target industry.

[0048] Furthermore, step S303 specifically includes:

[0049] (1) Obtain the attention weight matrix W of the pre-trained large language model q , W k , W[[ID=​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​2 F +||A|| 2 F )

[0055] Among them, Loss t represents the target task loss, λ represents the regularization coefficient, ||B|| 2 F Indicates L2 regularization of B, ||A|| 2 F Indicates L2 regularization of A;

[0056] Define the target loss function Loss t , as shown below:

[0057]

[0058] Among them, j represents the input question, y t Represents the word vector of the model answer at time step t, P(y t |j) represents each word y in the target answer y for question j t The predicted probability of Loss t represents the sum of the negative logarithmic probabilities of each word being correctly predicted, and T represents the total number of time steps;

[0059] (4) Use the optimizer to update the parameters of the large language model to adapt to the target industry task.

[0060] Furthermore, the step S4 specifically includes:

[0061] Step S401: On the supervised fine-tuned large language model, use the knowledge graph-based retrieval enhancement technology to optimize the large language model's answer content and inject target industry domain knowledge. The retrieval enhancement technology uses the structured knowledge in the knowledge graph and uses the retrieval module to search the knowledge graph for entity information that matches the user query. The retrieval module accurately locates relevant entities, attributes, and relationships in the knowledge graph based on the specific content of the user query, generating more accurate answers that are more relevant to the target industry. The specific steps of the knowledge graph-based retrieval enhancement technology include:

[0062] (1) When processing multiple consecutive rounds of dialogue, the large language model first semantically integrates the historical dialogue records with the current question to form a composite question containing contextual elements. The integrated question not only covers the core issues in the evolution of the dialogue, but also fully preserves the user's intention, the keywords in the current question, and their contextual associations.

[0063] (2) The large language model converts the complex question into a search instruction for the knowledge graph and locates the relevant subgraph in the knowledge graph through the relationship path tracing algorithm; the search results include entity nodes, attribute parameters, and relationship paths between entities that are directly related to the conversation history and the current question;

[0064] (3) By semantically fusing the retrieved knowledge graph subgraph information with the original compound question, a new knowledge-enhanced question is reconstructed using the large language model; the question supplemented with graph information is used as a new input instruction and resubmitted to the large language model for parsing;

[0065] (4) The large language model performs dual parsing simultaneously during processing, namely: semantic reasoning on the question; calling the structured data in the knowledge graph in real time through the interface; and obtaining the final response content, which includes the objective fact nodes in the knowledge graph and the logical deduction content generated by the model based on semantic understanding;

[0066] The specific steps of the relationship path tracing algorithm are as follows:

[0067] (1) Semantic anchor point extraction and path initialization: Input the composite question J = Σj, where j represents the specific question. The anchor point set A is extracted through the large language model as shown below:

[0068] A={(E s ,r,E o )∣E s ∈E,r∈R,E o ∈E}

[0069] Among them, A represents the anchor point set, E s is the subject entity, r represents the relationship, E o is the object entity, E is the knowledge graph entity set, and R is the knowledge graph relationship set;

[0070] Initialize the entity neighbor relationship queue:

[0071]

[0072] Among them, N1(e i ) represents entity e i The set of 1-degree adjacency relationships directly connected to entity e i The associated entity, Qu, represents the entity neighborhood relationship queue to be expanded;

[0073] (2) Multi-stage path expansion and weight evaluation: The breadth-first search algorithm is used to expand the path of entity relationships. The specific method is as follows:

[0074]

[0075] Among them, Pathn+1 Represents the set of n+1 hop paths after path expansion, path represents a path, is the path join operator, r new Indicates a new relationship, e new Represents a new entity, Path n represents the set of n-hop paths, e last Path n Path end entity;

[0076] After expanding the entity, the path comprehensive weight of each path is calculated. The calculation formula of the path comprehensive weight is as follows:

[0077] W(path)=α*S(path,Q)+β*C(path)-γ*L(path)+c*H(path)

[0078] Among them, α, β, γ, and c are hyperparameters. S(·) represents the semantic matching function, which is used to measure the semantic similarity between the path and the question; C(·) represents the statistical confidence function calculated based on the co-occurrence frequency of the path in the knowledge graph; L(·) represents the path length penalty term. The longer the path, the lower the score; H(·) represents the entropy measurement function, which is used to measure the consistency of different nodes and relationships in the path.

[0079] (3) Dynamic pruning and path optimization:

[0080] According to the comprehensive weight of the path obtained in step 2, the path set Path is expanded n Perform dynamic pruning and path optimization. The specific operations are as follows:

[0081] 1) Filter low-quality paths. Given a path scoring threshold, if the path score is lower than the threshold, the path is discarded; if the path score is higher than the threshold, the path is retained.

[0082] 2) Detect logical conflicts and further optimize the path set. If there is a semantic conflict in the path, remove the logical conflict path;

[0083] 3) Enhance the weight of paths that include core entities and increase the weight of paths directly related to core entities;

[0084] 4) Adopt an adaptive strategy to dynamically adjust the path weight hyperparameters.

[0085] Furthermore, the step S5 specifically includes:

[0086] Step S501: Collect a public dataset of knowledge questions and answers in the target industry. The dataset comes from question-and-answer conversations in online consultation and knowledge question-and-answer systems. Use the questions in the collected question-and-answer dataset as input to a customized model specifically for the target industry. Sample multiple answers from the model, use the real answers as a benchmark, and use evaluation indicators to generate preference data.

[0087] The preference data generation process in step S501 is specifically performed as follows:

[0088] (1) The model samples multiple answers to the same question and uses the character matching rate (CMR) indicator to evaluate the quality of the candidate answers. The specific calculation formula is as follows:

[0089] CMR=μ*e Σw*log{T}

[0090]

[0091] Among them, μ represents the penalty factor, which aims to prevent short texts from getting too high a score; w represents the weight of the short sequence; T represents the short sequence matching rate; u represents the number of matched short sequences, where a short sequence represents a sequence of n consecutive characters in the text; v represents the total number of short sequences;

[0092] (2) Generate preference data based on the calculated character matching rate. The higher the character matching rate, the higher the similarity with the benchmark answer, and the higher the quality of the candidate answer. The answer with the highest character matching rate is selected as the best preferred answer, and the benchmark answer is used as the standard answer. The specific format of the preference data is as follows:

[0093] (j,k b ,k z )

[0094] Among them, j represents the question input, k b represents the best preferred answer, k z Indicates the benchmark answer;

[0095] Step S502 optimizes the question-answer output through the reward model:

[0096] A reward model is trained using preference data. The reward model provides reward values ​​for the target industry-specific customized model. The specific reward model is trained and generated by minimizing the following loss function:

[0097] Reward(j,k b ,k z )=min(Loss r )

[0098] Loss r =-E[log{Sigmoid(reward(j,k z)-reward(j,k b ))}]

[0099] Among them, j represents the problem of customizing the model input for the target industry, reward(j,k z ) represents the reward value for the benchmark answer to input question j, reward(j,k b ) represents the reward value of the best preferred answer to the input question j, Sigmoid(·) represents the sigmoid function; E[·] represents the average loss value of each sample in the entire training dataset.

[0100] The beneficial effects of the present invention include:

[0101] (1) The present invention uses deep learning technology to construct a knowledge graph, which includes three steps: knowledge representation, extraction, and fusion. By constructing a target industry knowledge graph, a large amount of target industry expertise is injected into the large language model, enabling the customized model to more accurately understand industry-specific terms, concepts, and logic, thereby significantly improving the large language model's question-answering capabilities in the target industry field;

[0102] (2) The present invention uses a conditional random field decoder to learn the dependencies between entity labels and uses the relationship between adjacent entity labels to obtain the optimal entity recognition sequence. The conditional random field decoder constrains the rationality of the label sequence by learning the transition probability between labels and uses the overall information of the label sequence to reduce the probability of unreasonable entity recognition sequences appearing in the prediction sequence.

[0103] (3) The present invention uses a supervised fine-tuning algorithm to quickly adjust the model to adapt to new tasks through a low-rank matrix without changing the original weights of the pre-trained model, thereby reducing the number of parameters that need to be fine-tuned, speeding up training and reducing the demand for computing resources to adapt to specific target industries; using the supervised fine-tuning algorithm to build a customized model ensures high accuracy and high relevance of the model when answering questions related to the target industry;

[0104] (4) This invention constructs a reward model and introduces a reinforcement learning algorithm to guide the large language model to continuously optimize the output results during the content generation process, generating answers that are more in line with user preferences and industry standards. This mechanism effectively enhances the consistency between the model and the real question-and-answer scenarios of the target industry, significantly improves the relevance and accuracy of the output of the customized model, and further enhances the application value and intelligence level of the customized model in the target industry.

[0105] (5) The present invention uses a knowledge graph-based retrieval enhancement technology to optimize model answers. The retrieval module retrieves entity information that matches the user query in the knowledge graph. The retrieval module accurately locates relevant entities, attributes, and relationships in the knowledge graph based on the specific content of the user query, generating more accurate answers that are more relevant to the target industry and generating question-and-answer results that meet the user's preferences.

[0106] (6) The present invention combines the collaborative mechanism of knowledge graph retrieval and large language model reasoning, which significantly improves the domain adaptability and factual accuracy of large model question answering; through the complementarity of the precise knowledge positioning capability of the knowledge graph and the semantic understanding capability of the large language model, it not only ensures the reliability of professional knowledge, but also maintains the fluency of natural dialogue responses, thereby achieving dual optimization of quality and user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0107] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be further described in detail below with reference to the accompanying drawings, in which:

[0108] Figure 1 This is a flow chart of a method for generating an industry-specific customized model based on a large model according to the present invention;

[0109] Figure 2 This is the overall framework and main module diagram of an industry-specific customized model based on a large model of the present invention;

[0110] Figure 3 Obtain multi-source heterogeneous data flow diagram for the present invention;

[0111] Figure 4 This is the internal structure diagram of the bidirectional long short-term memory network-conditional random field of the present invention;

[0112] Figure 5 Constructing a process diagram for the industry knowledge graph of the present invention;

[0113] Figure 6 、 Figure 7 This is an actual question-and-answer effect diagram of an industry-specific customized model based on a large model in the present invention. DETAILED DESCRIPTION

[0114] The preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the preferred embodiments are only for illustrating the present invention, and are not intended to limit the scope of protection of the present invention.

[0115] As a form of expression for structured semantic knowledge, knowledge graphs use nodes to represent entities and edges to represent relationships between entities, thereby constructing a multi-dimensional graph-structured data model. Their primary function is to enhance the semantic understanding of data and increase the efficiency of information organization and retrieval. In natural language processing and intelligent question-answering, knowledge graphs can effectively implement entity disambiguation, relationship reasoning, and contextual semantic enhancement, improving the system's intelligence and decision-making accuracy. Injecting target industry domain knowledge from knowledge graphs into industry-specific customized models can effectively enhance the industry relevance of answers provided by these models.

[0116] Supervised fine-tuning is a core technology in the current large-scale language model customization process. Its basic concept is to introduce manually annotated or high-quality datasets based on pre-trained language models, and further optimize model parameters in a supervised manner to improve the model's question-answering capabilities for specific tasks or domains. Supervised fine-tuning can significantly reduce the number of parameters in industry-specific customized models, shortening training time. By fine-tuning certain parameters, the model can be made more suitable for question-answering in the target industry.

[0117] Reinforcement learning, a machine learning method centered on interactive learning, has demonstrated significant value in the customization and optimization of large language models in recent years. Through the dynamic interaction between an agent and its environment, reinforcement learning can continuously adjust strategies based on user feedback or reward signals in scenarios lacking supervisory signals, thereby guiding industry-specific customized models to generate outputs that better align with the preferences of the user group.

[0118] The present invention uses deep learning technology to construct a knowledge graph, realizes the structured modeling and semantic enhancement of the target industry domain knowledge; uses the supervised fine-tuning algorithm to efficiently adapt to the large language question answering task in the target industry by adjusting a small number of parameters; uses the reinforcement learning algorithm to optimize the output of the model, and finally forms a customized model for the target industry with high accuracy and high industry relevance. Figure 1 As shown, a method for generating an industry-specific customized model based on a large model in this embodiment includes the following steps:

[0119] Step S1: Collect structured and semi-structured multimodal data such as documents and reports containing key entities, concepts, and relationships in the target industry;

[0120] Step S2: Construct a knowledge graph through deep learning methods, including three steps: knowledge representation, knowledge extraction, and knowledge fusion;

[0121] Step S3: During the model building phase, we select a pre-trained large language model, perform feature engineering and personalize the model architecture based on industry characteristics, and design a customized model that meets the characteristics of the target industry through meticulous parameter tuning and strict monitoring of the training process.

[0122] Step S4: Based on the constructed knowledge graph, the knowledge graph-based retrieval enhancement technology is used to optimize the answer content of the large language model and inject target industry domain knowledge;

[0123] Step S5: Utilize the reinforcement learning algorithm to optimize the model output by training the reward model and provide feedback on the best question-answering results.

[0124] The present invention proposes an industry-specific customized model overall framework and main modules based on a large model. Figure 2 shown.

[0125] In this embodiment, step S1 specifically includes the following steps:

[0126] Step S101: Obtain reports and public data sets such as MySQL and Oracle published by target industry analysis institutions, build a structured relational database for the target industry, select a suitable data set, and convert the structured data in the data set into entities and relationships between entities in the knowledge graph;

[0127] Step S102: Obtain the target industry semi-structured data and unstructured text corpus data from web pages, online encyclopedia Wikipedia, online books and documents through network technology. The specific process is as follows: obtain the website links of the target industry and collect them into a list, obtain unstructured text corpus data for each link in turn until the list is empty, filter, remove duplicates and aggregate the data, and finally store it in the neo4j database; the flowchart of the present invention for obtaining multi-source heterogeneous data is as follows: Figure 3 shown.

[0128] In this embodiment, step S2 is described in detail below with a specific example:

[0129] Step S201: Use deep learning methods to perform knowledge representation on the multimodal data of the target industry obtained in S1. Deep learning algorithms can effectively integrate multimodal information, enhance the complementarity of multi-source heterogeneous information and the effectiveness of knowledge representation. Use a Transformer-based bidirectional encoder to pre-train the target industry dataset and generate word embeddings;

[0130] Step S202: Use the word embedding vector generated by the Transformer-based bidirectional encoder as the input of the bidirectional long short-term memory conditional random field model to perform entity recognition. The specific steps are as follows:

[0131] (1) The word vector output by the Transformer's bidirectional encoder model is used as the input of the bidirectional long short-term memory network at each time step;

[0132] (2) At each time step t, the forward long short-term memory network layer processes the sequence from time step 1 to time step t to obtain the forward hidden state sequence

[0133] (3) Use the reverse long short-term memory network layer to process the sequence from time step t to time step 1 to obtain the reverse hidden state sequence

[0134] (4) Concatenate the forward and reverse hidden states of each time step to obtain a comprehensive hidden state sequence

[0135] (5) Use the mapping matrix Q of the linear output layer to transform the hidden state sequence H t Mapped to n-dimensional space, n represents the number of label types, the mapping matrix Q = (Q1, Q2, ..., Q i )∈R i*n , Q i Represents the score of the i-th category relative to n labels, and R represents a real number;

[0136] (6) The conditional random field decoder is used to learn the dependency between entity labels, and the relationship between adjacent entity labels is used to obtain the optimal entity recognition sequence, thereby reducing the probability of unreasonable entity recognition sequences appearing in the prediction sequence; the internal structure diagram of the bidirectional long short-term memory network-conditional random field of the present invention is shown in FIG. Figure 4 shown.

[0137] The specific working principle of the above conditional random field decoder is as follows:

[0138] The conditional random field decoder constrains the rationality of the label sequence by learning the transition probability between labels, and uses the overall information of the label sequence to optimize the output of the model. The role of the conditional random field decoder in the model is to convert the features output by the bidirectional long short-term memory network layer into the optimal entity sequence; the input sentence is X = (x1, x2, ..., x n ), the label sequence Y corresponding to X=(y1,y2,…,y n ), the conditional random field decoder calculates the label sequence path score formula as follows:

[0139]

[0140] Among them, n represents the length of the sequence, y i represents the label of the i-th position in the sequence, is the transition probability matrix from label y i Move to label y i+1 The probability of Indicates that the i-th position in the output probability matrix of the bidirectional long short-term memory network layer is marked as label y i probability;

[0141] After the conditional random field decoder calculates the path scores of all possible label sequences, it selects the label sequence with the highest score as the final recognition result. The calculation formula for the best label sequence is as follows:

[0142] Y*=argmax{Score(X,Y)}

[0143] The conditional random field decoder constrains the rationality of the label sequence by learning the transition probability between labels, learns the label sequence restriction rules, and reduces the probability of unreasonable sequences.

[0144] The following example uses medical industry data to illustrate the working principle of the conditional random field decoder:

[0145] Suppose the following entity types need to be identified from electronic medical records:

[0146] Diseases (such as high blood pressure), medications (such as aspirin), symptoms (such as stomach pain), tests (such as an electrocardiogram).

[0147] The BIO tag method is used for the input statement, as shown below:

[0148] Input sentence: The patient has high blood pressure and diabetes and develops stomach pain after taking aspirin.

[0149] Target tag: OO{B-disease}{I-disease}OO{B-drug}OO{B-symptom}O

[0150] Among them, B- indicates the beginning of an entity, I- indicates the inside of an entity, and O- indicates a non-entity.

[0151] The conditional random field decoder can be used to impose reasonable constraints on label dependencies. In medical texts, the transfer of entity labels needs to be logical, for example:

[0152] Drugs are often followed by symptoms, i.e., stomach pain (B-symptom) after taking aspirin (B-drug) is a reasonable association;

[0153] The conditional random field decoder will learn a high transition probability from B-drug to B-symptom;

[0154] The disease may be followed by other diseases or examinations, i.e. hypertension (B-disease) and diabetes (B-disease) or hypertension (B-disease) requires an electrocardiogram (B-examination);

[0155] The conditional random field decoder will learn a high transition probability of B-disease → B-exam.

[0156] Step S203: Based on entity recognition, perform relationship extraction on the entity;

[0157] The specific principle of entity relationship extraction in step S203 is:

[0158] For a given sentence S={c1,c2,…,c n There are two target entities e1 and e2. The formula for calculating the sentence vector is:

[0159] M0′=W0·tanh(M0)+b0

[0160] Among them, M0 represents the final hidden layer vector of the Transformer's bidirectional encoder model, W0 is the weight matrix, and b0 is the bias vector;

[0161] Calculate the entity vector for each target entity and take the average value of the word vector in the entity vector. The calculation formula is:

[0162]

[0163] Among them, l s , l e Indicates the start and end position of the entity, M i Represents word vector, i≥1;

[0164] Finally, the two entity vectors and the sentence vector are concatenated and then input into a fully connected layer to obtain the output vector M' as follows:

[0165] M″=W1·concat(M0′,M i ′,M j ′)+b1

[0166] Where W1∈R p*3d , R represents a real number, where p is the total number of relation types, d represents the hidden layer dimension of each word vector in the Transformer encoder; M0' represents the final vector of sentence S, M i ' and M j 'represents the target entity e i and e j The final entity vector of , b1 is the bias vector;

[0167] Finally, M' is input into a SoftMax layer to classify the relationship type and obtain the relationship between the target entities. The formula is as follows:

[0168] R(e i ,e j )=SoftMax(M″)

[0169] Step S204: Based on entity recognition, extract events from multi-modal data, and then construct an event database and store it in the knowledge graph database; The process diagram of the industry knowledge graph construction of the present invention is as Figure 5 shown.

[0170] In this embodiment, step S3 includes:

[0171] Step S301: Select a suitable pre-trained large language model according to the actual tasks of the target industry;

[0172] Step S302: Inject the knowledge graph in the knowledge graph database obtained in step S2 into the large language model for supervised fine-tuning to adapt to the question and answer of the target industry;

[0173] Step S303: Through the supervised fine-tuning algorithm, without changing the original weights of the pre-trained model, quickly adjust the model through a low-rank matrix to adapt to the new task, reduce the number of parameters that need to be fine-tuned, speed up the training speed, and reduce the demand for computing resources, and adapt to a specific target industry.

[0174] As a further improvement, step S303 specifically includes:

[0175] (1) Obtain the attention weight matrix W of the pre-trained large language model q , W k , W v ;

[0176] (2) For the selected weight matrix W q (The operations of W k and W v are the same), perform low-rank decomposition on its updated matrix ΔW, which is expressed as:

[0177] W q +ΔW = W q +B·A

[0178] where B ∈ R d×r , A ∈ R r×k , R represents real numbers, and r << min(d, k), k represents the dimension of the input layer features, d represents the dimension of the output layer features, r represents the rank of the intermediate hidden layer of the supervised fine-tuning algorithm, R d×r represents a matrix with d rows and r columns, and R r×k represents a matrix with r rows and k columns;

[0179] (3) Define the total loss function Loss as follows:

[0180] Loss = Loss t +λ*(||B|| 2 F +||A||2 F )

[0181] Among them, Loss t represents the target task loss, λ represents the regularization coefficient, ||B|| 2 F Indicates L2 regularization of B, ||A|| 2 F Indicates L2 regularization of A;

[0182] Define the target loss function Loss t , as shown below:

[0183]

[0184] Among them, j represents the input question, y t Represents the word vector of the model answer at time step t, P(y t |j) represents each word y in the target answer y for question j t The predicted probability of Loss t represents the sum of the negative logarithmic probabilities of each word being correctly predicted, and T represents the total number of time steps;

[0185] (4) Use the Adam optimizer to update the parameters of the large language model to adapt to the target industry task.

[0186] In this embodiment, step S4 includes:

[0187] Step S401: On the supervised fine-tuned large language model, use the knowledge graph-based retrieval enhancement technology to optimize the large language model's answer content and inject target industry domain knowledge. The retrieval enhancement technology uses the structured knowledge in the knowledge graph and uses the retrieval module to search the knowledge graph for entity information that matches the user query. The retrieval module accurately locates relevant entities, attributes, and relationships in the knowledge graph based on the specific content of the user query, generating more accurate answers that are more relevant to the target industry. The specific steps of the knowledge graph-based retrieval enhancement technology include:

[0188] (1) When processing multiple consecutive rounds of dialogue, the large language model first semantically integrates the historical dialogue records with the current question to form a composite question containing contextual elements. The integrated question not only covers the core issues in the evolution of the dialogue, but also fully preserves the user's intention, the keywords in the current question, and their contextual associations.

[0189] (2) The large language model converts the complex question into a search instruction for the knowledge graph and locates the relevant subgraph in the knowledge graph through the relationship path tracing algorithm; the search results include entity nodes, attribute parameters, and relationship paths between entities that are directly related to the conversation history and the current question;

[0190] (3) By semantically fusing the retrieved knowledge graph subgraph information with the original compound question, a new knowledge-enhanced question is reconstructed using the large language model; the question supplemented with graph information is used as a new input instruction and resubmitted to the large language model for parsing;

[0191] (4) The large language model performs dual parsing simultaneously during processing, namely: semantic reasoning on the question; calling the structured data in the knowledge graph in real time through the interface; and obtaining the final response content, which includes the objective fact nodes in the knowledge graph and the logical deduction content generated by the model based on semantic understanding.

[0192] The specific steps of the relationship path tracing algorithm are as follows:

[0193] (1) Semantic anchor point extraction and path initialization: Input the composite question J = Σj, where j represents the specific question. The anchor point set A is extracted through the large language model as shown below:

[0194] A={(E s ,r,E o )∣E s ∈E,r∈R,E o ∈E}

[0195] Among them, A represents the anchor point set, E s is the subject entity, r represents the relationship, E o is the object entity, E is the knowledge graph entity set, and R is the knowledge graph relationship set;

[0196] Initialize the entity neighbor relationship queue:

[0197]

[0198] Among them, N1(e i ) represents entity e i The set of 1-degree adjacency relationships directly connected to entity e i The associated entity, Qu, represents the entity neighborhood relationship queue to be expanded;

[0199] (2) Multi-stage path expansion and weight evaluation: The breadth-first search algorithm is used to expand the path of entity relationships. The specific method is as follows:

[0200]

[0201] Among them, Pathn+1 Represents the set of n+1 hop paths after path expansion, path represents a path, is the path join operator, r new Indicates a new relationship, e new Represents a new entity, Path n represents the set of n-hop paths, e last Path n Path end entity;

[0202] After expanding the entity, the path comprehensive weight of each path is calculated. The calculation formula of the path comprehensive weight is as follows:

[0203] W(path)=α*S(path,Q)+β*C(path)-γ*L(path)+c*H(path)

[0204] Among them, α, β, γ, and c are hyperparameters. S(·) represents the semantic matching function, which is used to measure the semantic similarity between the path and the question; C(·) represents the statistical confidence function calculated based on the co-occurrence frequency of the path in the knowledge graph; L(·) represents the path length penalty term. The longer the path, the lower the score; H(·) represents the entropy measurement function, which is used to measure the consistency of different nodes and relationships in the path.

[0205] (3) Dynamic pruning and path optimization:

[0206] According to the comprehensive weight of the path obtained in step 2, the path set Path is expanded n Perform dynamic pruning and path optimization. The specific operations are as follows:

[0207] 1) Filter low-quality paths. Given a path scoring threshold, if the path score is lower than the threshold, the path is discarded; if the path score is higher than the threshold, the path is retained.

[0208] 2) Detect logical conflicts and further optimize the path set. If there is a semantic conflict in the path, remove the logical conflict path;

[0209] 3) Enhance the weight of paths that include core entities and increase the weight of paths directly related to core entities;

[0210] 4) Adopt an adaptive strategy to dynamically adjust the path weight hyperparameters.

[0211] The following example illustrates the process of entity relationship multi-stage path expansion and path evaluation:

[0212] (1) Semantic anchor extraction and path initialization:

[0213] Question: Patient A has symptoms of fever and cough. What medicines might the doctor prescribe?

[0214] Anchor point extraction: A = {(patient A, present, fever), (patient A, present, cough)}

[0215] Neighborhood expansion: Path = {(fever, indication, flu), (cough, indication, pneumonia)}

[0216] (2) Multi-stage path expansion and weight evaluation:

[0217] Path 1 = [Patient A → Presentation → Fever → Indication → Flu → Prescription → Oseltamivir]

[0218] Path 2 = [Patient A → Presentation → Cough → Indication → Pneumonia → Prescription → Parovert]

[0219] Score the two paths above:

[0220] W(Path1)=0.5×0.85+0.3×0.75-0.2×2+0.1×0.5=0.25

[0221] W(Path2)=0.5×0.92+0.3×0.80-0.2×2+0.1×0.4=0.34

[0222] (3) Dynamic pruning and path optimization:

[0223] After calculation in step (2), it can be seen that: W(Path1) is lower than the minimum threshold of 0.3, so the path is discarded; W(Path2) is higher than the minimum threshold, so the path is retained.

[0224] In this embodiment, step S5 includes:

[0225] Step S501: Collect a public dataset of knowledge questions and answers in the target industry. The dataset comes from question-and-answer conversations in online consultation and knowledge question-and-answer systems. Use the questions in the collected question-and-answer dataset as input to a customized model specifically for the target industry. Sample multiple answers from the model, use the real answers as a benchmark, and use evaluation indicators to generate preference data.

[0226] The preference data generation process in step S501 is specifically performed as follows:

[0227] (1) The model samples multiple answers to the same question and uses the character matching rate (CMR) indicator to evaluate the quality of the candidate answers. The specific calculation formula is as follows:

[0228] CMR=μ*e Σw*log{T}

[0229]

[0230] Among them, μ represents the penalty factor, which aims to prevent short texts from getting too high a score; w represents the weight of the short sequence; T represents the short sequence matching rate; u represents the number of matched short sequences, where a short sequence represents a sequence of n consecutive characters in the text; v represents the total number of short sequences;

[0231] (2) Generate preference data based on the calculated character matching rate. The higher the character matching rate, the higher the similarity with the benchmark answer, and the higher the quality of the candidate answer. The answer with the highest character matching rate is selected as the best preferred answer, and the benchmark answer is used as the standard answer. The specific format of the preference data is as follows:

[0232] (j,k b ,k z )

[0233] Among them, j represents the question input, k b represents the best preferred answer, k z Indicates the benchmark answer;

[0234] Step S502 optimizes the question-answer output through the reward model:

[0235] A reward model is trained using preference data. The reward model provides reward values ​​for the target industry-specific customized model. The specific reward model is trained and generated by minimizing the following loss function:

[0236] Reward(j,k b ,k z )=min(Loss r )

[0237] Loss r =-E[log{Sigmoid(reward(j,k z )-reward(j,k b ))}]

[0238] Among them, j represents the problem of customizing the model input for the target industry, reward(j,k z ) represents the reward value for the benchmark answer to input question j, reward(j,k b ) represents the reward value of the best preferred answer to the input question j, Sigmoid(·) represents the sigmoid function; E[·] represents the average loss value of each sample in the entire training dataset;

[0239] The following example illustrates the process of generating preference data:

[0240] Question input j: What is hypertension?

[0241] Sample multiple responses to the same question into the model:

[0242] k1: Hypertension is a chronic disease characterized by persistently high arterial blood pressure. It usually has no obvious symptoms but can cause damage to organs such as the heart, brain, and kidneys.

[0243] k2: Hypertension is a high blood pressure disease that can cause a variety of complications.

[0244] k3: A disease of high blood pressure is called hypertension.

[0245] Extracting benchmark answers from real public datasets:

[0246] k z Hypertension: Hypertension is a common chronic disease that refers to a persistent increase in arterial blood pressure, which may lead to cardiovascular and cerebrovascular complications.

[0247] The similarity with the benchmark answer is calculated using the character matching rate (CMR), and the parameters are defined as follows: n = 5, μ = 0.9, w = 1.5;

[0248] Calculate the character matching rates CMR of k1, k2, and k3 respectively and write them into the table, as shown in Table 1:

[0249] Table 1

[0250] answer Number of matching short sequences u Total number of short sequences v Short sequence matching rate T Character matching rate CMR <![CDATA[k1]]> 35 68 0.5147 0.332 <![CDATA[k2]]> 11 24 0.4583 0.279 <![CDATA[k3]]> 4 13 0.3077 0.153

[0251] As can be seen from the table, answer k1 is the best candidate answer with the highest character matching rate.

[0252] Finally, the embodiment shows the real question-answering results of an industry-specific customized model based on a large model of the present invention, such as Figure 6 、 Figure 7 shown.

[0253] It should be noted that the present invention is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the process in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0254] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0255] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0256] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and it should be understood that the scope of protection of the present invention is not limited to such specific descriptions and embodiments. Those skilled in the art can make various other specific variations and combinations based on the technical teachings disclosed in the present invention without departing from the essence of the present invention, and such variations and combinations are still within the scope of protection of the present invention.

Claims

1. A method for generating an industry-specific customized model based on a large model, characterized by: The following steps are involved: Step S1: Collect structured and semi-structured multimodal data containing key entities, concepts and relationships in the target industry; Step S2: Construct a knowledge graph through deep learning methods, including three steps: knowledge representation, knowledge extraction, and knowledge fusion; Step S3: During the model building phase, we select a pre-trained large language model, perform feature engineering and personalize the model architecture based on industry characteristics, and design a customized model that meets the characteristics of the target industry through meticulous parameter tuning and strict monitoring of the training process. Step S4: Based on the constructed knowledge graph, the knowledge graph-based retrieval enhancement technology is used to optimize the answer content of the large language model and inject target industry domain knowledge; Step S5: Utilize the reinforcement learning algorithm to optimize the model output by training the reward model and provide feedback on the best question-answering results.

2. The method for generating an industry-specific customized model based on a large model according to claim 1, characterized in that: The step S1 specifically includes: Step S101: Obtain reports and public data sets published by target industry analysis institutions, build a structured relational database for the target industry, select an appropriate data set, and convert the structured data in the data set into entities and relationships between entities in the knowledge graph; Step S102: Obtain the semi-structured data and unstructured text corpus data of the target industry. The specific process is as follows: obtain the website links of the target industry and collect them into a list, obtain unstructured text corpus data for each link in turn until the list is empty, filter, deduplicate and aggregate the data, and finally store it in the database.

3. The method for generating an industry-specific customized model based on a large model according to claim 1, characterized in that: The step S2 specifically includes: Step S201: Use deep learning methods to perform knowledge representation on the multimodal data of the target industry obtained in S1. Deep learning algorithms can effectively integrate multimodal information, enhance the complementarity of multi-source heterogeneous information and the effectiveness of knowledge representation. Use a Transformer-based bidirectional encoder to pre-train the target industry dataset and generate word embeddings; Step S202: Use the word embedding vector generated by the Transformer-based bidirectional encoder as the input of the bidirectional long short-term memory conditional random field model to perform entity recognition. The specific steps are as follows: (1) The word vector output by the Transformer's bidirectional encoder model is used as the input of the bidirectional long short-term memory network at each time step; (2) At each time step t, the forward long short-term memory network layer processes the sequence from time step 1 to time step t to obtain the forward hidden state sequence (3) Use the reverse long short-term memory network layer to process the sequence from time step t to time step 1 to obtain the reverse hidden state sequence (4) Concatenate the forward and reverse hidden states of each time step to obtain a comprehensive hidden state sequence (5) Use the mapping matrix Q of the linear output layer to transform the hidden state sequence H t Mapped to n-dimensional space, n represents the number of label types, the mapping matrix Q = (Q1, Q2, ..., Q i )∈R i*n , Q i Represents the score of the i-th category relative to n labels, and R represents a real number; (6) Using the conditional random field decoder to learn the dependency between entity labels, the relationship between adjacent entity labels is used to obtain the optimal entity recognition sequence, which reduces the probability of unreasonable entity recognition sequences appearing in the prediction sequence; Step S203: Based on entity recognition, perform relationship extraction on the entity; Step S204: Based on entity recognition, event extraction is performed on the multimodal data, and then an event database is constructed and stored in the knowledge graph database.

4. The method for generating an industry-specific customized model based on a large model according to claim 1, characterized in that: The step S3 specifically includes: Step S301: Select an appropriate pre-trained large language model based on the actual tasks of the target industry; Step S302: injecting the knowledge graph in the knowledge graph database obtained in step S2 into the large language model for supervised fine-tuning to adapt to the question answering of the target industry; Step S303: Through the supervised fine-tuning algorithm, without changing the original weights of the pre-trained model, the model is quickly adjusted to adapt to the new task through the low-rank matrix, reducing the number of parameters that need to be fine-tuned, speeding up the training speed and reducing the demand for computing resources to adapt to specific target industries.

5. The method for generating an industry-specific customized model based on a large model according to claim 1, characterized in that: The step S4 specifically includes: On the supervised fine-tuned large language model, we use knowledge graph-based retrieval enhancement technology to optimize the large language model's answers and inject target industry domain knowledge. Retrieval enhancement technology leverages the structured knowledge within the knowledge graph and uses the retrieval module to retrieve entity information matching the user's query within the knowledge graph. Based on the specific content of the user's query, the retrieval module accurately locates relevant entities, attributes, and relationships within the knowledge graph, generating more accurate answers that are more relevant to the target industry. The specific steps of knowledge graph-based retrieval enhancement technology include: (1) When processing multiple consecutive rounds of dialogue, the large language model first semantically integrates the historical dialogue records with the current question to form a composite question containing contextual elements. The integrated question not only covers the core issues in the evolution of the dialogue, but also fully preserves the user's intention, the keywords in the current question, and their contextual associations. (2) The large language model converts the complex question into a search instruction for the knowledge graph and locates the relevant subgraph in the knowledge graph through the relationship path tracing algorithm; the search results include entity nodes, attribute parameters, and relationship paths between entities that are directly related to the conversation history and the current question; (3) By semantically fusing the retrieved knowledge graph subgraph information with the original compound question, a new knowledge-enhanced question is reconstructed using the large language model; the question supplemented with graph information is used as a new input instruction and resubmitted to the large language model for parsing; (4) The large language model performs dual parsing simultaneously during processing, namely: semantic reasoning on the question; calling the structured data in the knowledge graph in real time through the interface; and obtaining the final response content, which includes the objective fact nodes in the knowledge graph and the logical deduction content generated by the model based on semantic understanding.

6. The method for generating an industry-specific customized model based on a large model according to claim 1, characterized in that: The step S5 comprises: Step S501: Collect a public dataset of knowledge questions and answers in the target industry. The dataset comes from question-and-answer conversations in online consultation and knowledge question-and-answer systems. Use the questions in the collected question-and-answer dataset as input to a customized model specifically for the target industry. Sample multiple answers from the model, use the real answers as a benchmark, and use evaluation indicators to generate preference data. Step S502: The preference data generated in S501 is used as the input for training the reward model. The reward model calculates the reward value of the preference data and provides an evaluation criterion for the preference strategy. The preference strategy chooses how to answer the next question based on the evaluation criterion to maximize the long-term cumulative reward. The target industry-specific customized model optimizes the question and answer output through the reward model.

7. The method for generating an industry-specific customized model based on a large model according to claim 2, characterized by: The specific working method of the conditional random field decoder in step S201 is as follows: The conditional random field decoder constrains the rationality of the label sequence by learning the transition probability between labels, and uses the overall information of the label sequence to optimize the output of the model. The role of the conditional random field decoder in the model is to convert the features output by the bidirectional long short-term memory network layer into the optimal entity sequence; the input sentence is X = (x1, x2, ..., x n ), the label sequence Y corresponding to X=(y1,y2,…,y n ), the conditional random field decoder calculates the label sequence path score formula as follows: Among them, n represents the length of the sequence, y i represents the label of the i-th position in the sequence, is the transition probability matrix from label y i Move to label y i+1 The probability of Indicates that the i-th position in the output probability matrix of the bidirectional long short-term memory network layer is marked as label y i probability; After the conditional random field decoder calculates the path scores of all possible label sequences, it selects the label sequence with the highest score as the final recognition result. The calculation formula for the best label sequence is as follows: Y*=argmax{Score(X,Y)} The conditional random field decoder constrains the rationality of the label sequence by learning the transition probability between labels, learns the label sequence restriction rules, and reduces the probability of unreasonable sequences.

8. The method for generating an industry-specific customized model based on a large model according to claim 2, characterized in that: The method for extracting relationships between entities in step S203 is: For a given sentence S={c1,c2,…,c n There are two target entities e1 and e2. The formula for calculating the sentence vector is: M′0=W0·tanh(M0)+b0 Among them, M0 represents the final hidden layer vector of the Transformer's bidirectional encoder model, W0 is the weight matrix, and b0 is the bias vector; Calculate the entity vector for each target entity and take the average value of the word vector in the entity vector. The calculation formula is: Among them, l s , l e Indicates the start and end position of the entity, M i Represents word vector, i≥1; Concatenate the two entity vectors and the sentence vector, and then input them into a fully connected layer to obtain the output vector M' as follows: M″=W1·concat(M′0,M′ i ,M′ j )+b1 Where W1∈R p*3d , R represents a real number, p is the total number of relation types, d represents the hidden layer dimension of each word vector in the Transformer encoder; M0' represents the final vector of sentence S, M i ' and M j 'represents the target entity e i and e j The final entity vector of , b1 is the bias vector; Finally, M' is input into a SoftMax layer to classify the relationship type and obtain the relationship between the target entities. The formula is as follows: Re i ,and j )=SoftMax(M″)。 9. The method for generating an industry-specific customized model based on a large model according to claim 3, characterized in that: The specific process of the supervised fine-tuning algorithm in step S303 is as follows: (1) Obtain the attention weight matrix W of the pre-trained large language model q , W k , W v ; (2) For the selected weight matrix W q (W k and W v The update matrix ΔW is decomposed into low rank, which is expressed as: IN q +ΔW=W q +B A where B ∈ R d×r , A ∈ R r×k , R represents the set of real numbers, and r << min(d, k), where k represents the dimension of the input layer features, d represents the dimension of the output layer features, and r represents the rank of the intermediate hidden layer of the supervised fine-tuning algorithm, R d×r represents a matrix of d rows and r columns, and R r×k represents a matrix of r rows and k columns; (3) Define the total loss function Loss as follows: Loss=Loss t +λ*(||B|| 2 F +||A|| 2 F ) Among them, Loss t represents the target task loss, λ represents the regularization coefficient, ||B|| 2 F Indicates L2 regularization of B, ||A|| 2 F Indicates L2 regularization of A; Define the target loss function Loss t , as shown below: Among them, j represents the input question, y t Represents the word vector of the model answer at time step t, P(y t |j) represents each word y in the target answer y for question j t The predicted probability of Loss t represents the sum of the negative logarithmic probabilities of each word being correctly predicted, and T represents the total number of time steps; (4) Use the optimizer to update the parameters of the large language model to adapt to the target industry task.

10. The method for generating an industry-specific customized model based on a large model according to claim 6, characterized in that: The preference data generation process in step S501 is specifically performed as follows: (1) The model samples multiple answers to the same question and uses the character matching rate (CMR) indicator to evaluate the quality of the candidate answers. The specific calculation formula is as follows: CMR=μ*e Σw*log{T} Among them, μ represents the penalty factor, which aims to prevent short texts from getting too high a score; w represents the weight of the short sequence; T represents the short sequence matching rate; u represents the number of matched short sequences, where a short sequence represents a sequence of n consecutive characters in the text; v represents the total number of short sequences; (2) Generate preference data based on the calculated character matching rate. The higher the character matching rate, the higher the similarity with the benchmark answer, and the higher the quality of the candidate answer. The answer with the highest character matching rate is selected as the best preferred answer, and the benchmark answer is used as the standard answer. The specific format of the preference data is as follows: (j,k b ,k z ) Among them, j represents the question input, k b represents the best preferred answer, k z Indicates the benchmark answer; Step S502 optimizes the question-answer output through the reward model: A reward model is trained using preference data. The reward model provides reward values ​​for the target industry-specific customized model. The specific reward model is trained and generated by minimizing the following loss function: Reward(j,k b ,k z )=min(Loss r ) Loss r =-E[log{Sigmoid(reward(j,k z )-reward(j,k b ))}] Among them, j represents the problem of customizing the model input for the target industry, reward(j,k z ) represents the reward value for the benchmark answer to input question j, reward(j,k b ) represents the reward value of the best preferred answer to the input question j, Sigmoid(·) represents the sigmoid function; E[·] represents the average loss value of each sample in the entire training dataset.

Citation Information

Patent Citations

  • Reinforcement learning knowledge graph reasoning method based on generative adversarial imitation learning

    CN115269861A

  • Knowledge graph question-answering method for sub-graph retrieval optimization

    CN117149974A

  • Learning path recommendation method based on attention knowledge tracking

    CN117494059A

  • Intelligent question answering method and system based on domain knowledge graph

    CN117648984A

  • Data intelligent question and answer method and system fusing domain knowledge

    CN118779438A

Cited By

  • Knowledge base-based reasoning comparison method

    CN121119142A

  • Multi-modal large model training method and related device

    CN121257750A