Large language model output detection method and device, electronic equipment and storage medium

By obtaining the output text of the large language model and its target knowledge graph in its field, knowledge extraction and detection are carried out, the problem of error knowledge output by large language model is solved, and the accuracy of the output text is judged and error information is avoided.

CN120257970APending Publication Date: 2025-07-04SHENZHEN INTELLIFUSION TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311870542.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-30
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

There are knowledge errors in the content output by large language models, and it is difficult to confirm through technical means whether it exists in the training data and is accurate.

Method used

By obtaining the output text of the large language model and its target knowledge graph in its field, knowledge extraction and detection are performed, the target knowledge graph is used to match the detected knowledge data to determine whether there are knowledge errors in the output text.

Benefits of technology

It effectively avoids the knowledge of output errors in large language models to users, and improves the accuracy and reliability of output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120257970A_ABST
    Figure CN120257970A_ABST
Patent Text Reader

Abstract

The invention provides an output detection method of a large language model, which comprises the following steps: obtaining an output text of the large language model for a question text, and obtaining a target knowledge graph of a field to which the question text belongs, the target knowledge graph comprising knowledge data of the field to which the question text belongs; performing knowledge extraction on the output text to obtain to-be-detected knowledge data of the output text; and detecting the to-be-detected knowledge data through the target knowledge graph to obtain a knowledge detection result of the output text. The to-be-detected knowledge data of the output text is obtained by performing knowledge extraction on the output text of the large language model, and the to-be-detected knowledge data is detected by using the target knowledge graph of the belonging field, so that whether knowledge errors exist in the output text or not can be judged, and wrong knowledge is prevented from being output to a user by the large language model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and particularly to an output detection method, device, electronic device, and storage medium for a large language model. Background Art

[0002] In the process of the rapid development of general artificial intelligence technology, large language models have a first-mover advantage. There are three reasons for this. First, compared with data in other modalities (such as image, video, audio, etc.), text has the characteristics of high knowledge content and data sparsity. Second, with the advantage of the Internet, text data has been digitized globally. Third, text knowledge is organized by language and has a certain semantic structure, which can be used for representation learning through deep learning models (such as Transformer).

[0003] Based on Transformer and a vast amount of text knowledge, large language models adopt auto-regressive modeling to achieve rich multilingual semantic learning. Auto-regressive modeling has many advantages. One is that parallel training can be started based on the server. The other is that it can quickly learn the semantic information of the language. However, auto-regressive modeling also brings the hallucination problem. Specifically, the hallucination problem refers to the fact that the large language model outputs knowledge that does not exist in the training data, and this knowledge is incorrect. The difficulty of the hallucination problem lies in that it is difficult to confirm through technical means whether the output knowledge exists in the training data, and it is also difficult to judge whether the knowledge is accurate, resulting in knowledge errors in the content output by the large language model. Summary of the Invention

[0004] An embodiment of the present invention provides an output detection method for a large language model, aiming to solve the problem of knowledge errors in the content output by the large language model. By extracting knowledge from the output text of the large language model, the knowledge data to be detected of the output text is obtained, and the target knowledge graph in the relevant field is used to detect the knowledge data to be detected, so as to determine whether there are knowledge errors in the output text and prevent the large language model from outputting incorrect knowledge to users.

[0005] In a first aspect, an embodiment of the present invention provides an output detection method for a large language model, the method comprising:

[0006] Obtaining the output text of the large language model for the question text, and obtaining the target knowledge graph in the field to which the question text belongs, the target knowledge graph including the knowledge data in the field to which the question text belongs;

[0007] Performing knowledge extraction on the output text to obtain the knowledge data to be detected of the output text;

[0008] Detect the to-be-detected knowledge data through the target knowledge graph to obtain the knowledge detection result of the output text.

[0009] Optionally, before obtaining the target knowledge graph of the field to which the question text belongs, the method further includes:

[0010] Obtain entity tokens and relationship tokens in different fields;

[0011] For the entity tokens and the relationship tokens in one field, annotate between different entity tokens through the relationship tokens to obtain knowledge data, where the knowledge data includes at least two entity tokens and at least one relationship token that associates at least two entity tokens;

[0012] Construct knowledge graphs for different fields according to the knowledge data in different fields.

[0013] Optionally, the annotating between different entity tokens through the relationship tokens to obtain knowledge data includes:

[0014] Divide the entity tokens into subject tokens and object tokens;

[0015] Based on the relationship tokens, annotate the subject tokens and the object tokens to obtain knowledge data.

[0016] Optionally, before extracting knowledge from the output text to obtain the to-be-detected knowledge data of the output text, the method further includes:

[0017] Obtain a natural language processing model to be trained and a training data set, where the training data set includes sample texts and knowledge labels of the sample texts, the input of the natural language processing model to be trained is constructed as the sample texts, and the output of the natural language processing model to be trained is constructed as predicted knowledge data;

[0018] Input the sample texts into the natural language processing model to obtain the predicted knowledge data of the sample texts;

[0019] Calculate the loss value between the predicted knowledge data of the sample texts and the knowledge labels of the sample texts;

[0020] Based on the loss value, adjust the parameters of the natural language processing model to be trained, and iterate the parameter adjustment process. After training is completed, obtain a trained natural language processing model, and the trained natural language processing model is used to extract knowledge from the output text.

[0021] Optionally, before performing knowledge extraction on the output text to obtain the to-be-detected knowledge data of the output text, the method further includes:

[0022] Obtain a natural language processing model to be trained and a training data set, where the training data set includes sample texts, entity labels of the sample texts, and relationship labels, the input of the natural language processing model to be trained is constructed as the sample texts, and the output of the natural language processing model to be trained is constructed as predicted entities and predicted relationships;

[0023] Input the sample texts into the natural language processing model to obtain the predicted entity data and predicted relationship data of the sample texts;

[0024] Calculate a first loss value between the predicted entity data of the sample texts and the entity labels of the sample texts, and a second loss value between the predicted relationship data of the sample texts and the relationship labels of the sample texts;

[0025] Based on the first loss value and the second loss value, adjust the parameters of the natural language processing model to be trained, and iterate the parameter adjustment process. After training is completed, obtain a trained natural language processing model, and the trained natural language processing model is used to perform knowledge extraction on the output text.

[0026] Optionally, performing knowledge extraction on the output text to obtain the to-be-detected knowledge data of the output text includes:

[0027] Perform knowledge extraction on the output text through the trained natural language processing model to obtain the to-be-detected entities and to-be-detected relationships of the output text;

[0028] Combine the to-be-detected entities and to-be-detected relationships to obtain the knowledge data of the output text, and the to-be-detected knowledge data includes to-be-detected entities and to-be-detected relationships.

[0029] Optionally, the detecting the to-be-detected knowledge data through the target knowledge graph to obtain the knowledge detection result of the output text includes:

[0030] Determine whether the to-be-detected entities and the to-be-detected relationships exist in the target knowledge graph;

[0031] If the to-be-detected entities exist in the target knowledge graph and the to-be-detected relationships also exist in the target knowledge graph, determine that the knowledge detection of the output text is correct;

[0032] If the entity to be detected does not exist in the target knowledge graph, and / or the relationship to be detected does not exist in the target knowledge graph, it is determined that there is a knowledge detection error in the output text.

[0033] In a second aspect, an embodiment of the present invention further provides an output detection device for a large language model. The output detection device for the large language model includes:

[0034] A first acquisition module, configured to acquire the output text of the large language model for the question text, and acquire the target knowledge graph of the field to which the question text belongs. The target knowledge graph includes knowledge data of the field to which the question text belongs;

[0035] An extraction module, configured to perform knowledge extraction on the output text to obtain the knowledge data to be detected in the output text;

[0036] A detection module, configured to detect the knowledge data to be detected through the target knowledge graph to obtain the knowledge detection result of the output text.

[0037] In a third aspect, an embodiment of the present invention provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps in the output detection method for the large language model provided by the embodiment of the present invention are implemented.

[0038] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the output detection method for the large language model provided by the embodiment of the invention are implemented.

[0039] In the embodiment of the present invention, the output text of the large language model for the question text is acquired, and the target knowledge graph of the field to which the question text belongs is acquired. The target knowledge graph includes knowledge data of the field to which the question text belongs; knowledge extraction is performed on the output text to obtain the knowledge data to be detected in the output text; the knowledge data to be detected is detected through the target knowledge graph to obtain the knowledge detection result of the output text. By performing knowledge extraction on the output text of the large language model to obtain the knowledge data to be detected in the output text, and using the target knowledge graph of the relevant field to detect the knowledge data to be detected, it is possible to determine whether there is a knowledge error in the output text, and avoid the large language model from outputting incorrect knowledge to users. Description of the Drawings

[0040] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0041] Figure 1 is a flowchart of an output detection method for a large language model provided by an embodiment of the present invention;

[0042] Figure 2 is a flowchart of another output detection method for a large language model provided by an embodiment of the present invention;

[0043] Figure 3 is a flowchart of a knowledge graph detection method provided by an embodiment of the present invention;

[0044] Figure 4 is a flowchart of a secondary detection method provided by an embodiment of the present invention;

[0045] Figure 5 is a schematic structural diagram of an output detection device for a large language model provided by an embodiment of the present invention;

[0046] Figure 6 is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Specific Embodiments

[0047] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0048] As Figure 1 shown, Figure 1 is a flowchart of an output detection method for a large language model provided by an embodiment of the present invention. The output detection method for the large language model includes the steps of:

[0049] 101. Obtain the output text of the large language model for the question text, and obtain the target knowledge graph of the field to which the question text belongs.

[0050] In an embodiment of the present invention, the above-mentioned large language model (LLM) can be a large language model obtained by autoregressive modeling, or any large language model that can generate output hallucinations. The above-mentioned question text is the text input by the user into the large language model to obtain the desired answer. The large language model performs semantic understanding and knowledge search based on the question text, and assembles the retrieved knowledge into a natural language text for output to obtain the output text. The above-mentioned target knowledge graph includes knowledge data in the field to which the question text belongs.

[0051] It should be noted that the problem of hallucinations inevitably occurs in large language models based on autoregressive modeling. During the implementation of large language models, users require the model to have a certain degree of innovation while also demanding accuracy. Innovation requires the large language model to process and understand the training data, ultimately forming some new knowledge, and this knowledge is likely to be obtained through reasoning. Accuracy requires that the large language model cannot fabricate and synthesize based on the training data, resulting in incorrect knowledge output. Innovation and accuracy are mutually restrictive and require a balance. After training and learning based on a large amount of text data, large language models often exhibit a certain degree of innovation, bringing about the problem of hallucinations. For all correct knowledge, an abstract mathematical set K can be defined. The knowledge contained in a large amount of text data is defined as X. Generally speaking, it can be assumed that the text knowledge X ≤ K. Suppose a large language model f is trained based on X. Given a specific input x, the large language model has an output y = f(x). When it can be considered that the large model has generated hallucinations. It should be noted that as the boundaries of human knowledge expand, K is a dynamically growing set. By taking the set K at a certain point in time, it can be determined whether the output y = f(x) of the large model is a hallucination. However, in practical applications, it is impossible to construct a domain-complete set K. In addition, it is very resource-consuming to determine from X based on the output of the large model because X often involves hundreds of billions of tokens and the knowledge therein is difficult to retrieve. Therefore, it is a difficult problem to actually detect whether the output of the large model is a hallucination.

[0052] In an embodiment of the present invention, the knowledge graph is constructed according to different fields, avoiding the collection of cross-domain knowledge. There is no need to consider the completeness of knowledge at the full-domain level, and only the completeness of knowledge in each field needs to be considered. The above-mentioned fields can be secondary education, corporate finance, surgical medicine, etc. It can be understood that the finer the field is divided, the lower the construction difficulty of the knowledge graph, and the smaller the constructed knowledge graph, the easier it is to achieve a knowledge accuracy of 100%.

[0053] The above knowledge data is encapsulated according to a preset knowledge structure, such as encapsulation by entity and relationship. Specifically, it can be encapsulated by relationship between entity and entity to obtain. The above entities and relationships can be different types of tokens. The classification of entity tokens can be personal names, referring nouns, organizations, locations, time, etc. The classification of relationship tokens can be inheritance relationship, synonym relationship, attribution relationship, correlation relationship, causal relationship, etc. By combining two entity tokens through relationship tokens, a piece of knowledge data is obtained. The knowledge data in the knowledge graphs of each field needs to be reviewed by professionals in the field to ensure that the knowledge data in the knowledge graph is 100% correct.

[0054] Different knowledge graphs correspond to different fields. The large language model can output the output text of the question text and the knowledge graph interface API. The knowledge graph interface API corresponds to the knowledge graph of the field to which the question text belongs. The corresponding knowledge graph can be called through the knowledge graph interface API. Different knowledge graphs correspond to different knowledge graph interface APIs. The target knowledge graph of the field to which the question text belongs can be loaded or called through the knowledge graph interface API. Of course, it is also possible to directly perform natural language processing on the question text to determine the field to which the question text belongs, and load or call the target knowledge graph of the field to which the question text belongs according to the field to which the question text belongs, without determining the field to which the question text belongs through the large language model. In the embodiment of the present invention, it is preferably to output the output text of the question text and the knowledge graph interface API through the large language model. In this way, only a prompt word needs to be added to the question text. For example, add "and output the knowledge graph interface API corresponding to the field of the question", and there is no need to train an additional field classifier.

[0055] 102. Perform knowledge extraction on the output text to obtain the knowledge data to be detected of the output text.

[0056] In the embodiment of the present invention, the above output text is obtained by the large language model processing and outputting according to the user's question text. The above output text will not be output to the user as the final output text, but needs to be knowledge detected, and on the basis of passing the knowledge detection, the output text will be output to the user as the final output text.

[0057] Entity and relationship extraction can be performed on the output text through sequence labeling technology, and through the combination between entities and relationships, the knowledge data of the output text is obtained as the knowledge data to be detected.

[0058] It is also possible to train a knowledge data extraction model, and use the trained knowledge data extraction model to perform knowledge data extraction processing on the output text to obtain the knowledge data to be detected of the output text. The structure of the knowledge data to be detected is the same as the structure of the knowledge data in the knowledge graph.

[0059] In an output text, the knowledge data to be detected can be one or more.

[0060] 103. Detect the knowledge data to be detected through the target knowledge graph to obtain the knowledge detection result of the output text.

[0061] In an embodiment of the present invention, after obtaining the knowledge data to be detected, match the knowledge data to be detected with all the knowledge data in the target knowledge graph. If all the knowledge data to be detected are matched with the knowledge data in the knowledge graph, it indicates that there is no incorrect knowledge in the output text, and then determine that the knowledge detection result of the output text is knowledge detection passed in the next step. If there is knowledge data to be detected that is not matched with the knowledge data in the knowledge graph, it indicates that there is incorrect knowledge in the output text, and then determine that the knowledge detection result of the output text is knowledge detection failed in the next step.

[0062] After obtaining the knowledge detection result of the output text, the output text with knowledge detection passed can be used as the final output text and output to the user. If the knowledge detection of the output text fails, the large language model can be prompted to regenerate a new output text according to the question text until the knowledge detection of the new output text passes, or until the number of times of knowledge detection failure reaches the preset number of times, and the user is prompted to re-describe the question text.

[0063] In an embodiment of the present invention, obtain the output text of the large language model for the question text, and obtain the target knowledge graph of the field to which the question text belongs. The target knowledge graph includes the knowledge data of the field to which the question text belongs; perform knowledge extraction on the output text to obtain the knowledge data to be detected of the output text; detect the knowledge data to be detected through the target knowledge graph to obtain the knowledge detection result of the output text. By performing knowledge extraction on the output text of the large language model to obtain the knowledge data to be detected of the output text, and using the target knowledge graph of the relevant field to detect the knowledge data to be detected, it is possible to determine whether there is knowledge error in the output text, and avoid the large language model from outputting incorrect knowledge to the user.

[0064] Optionally, before the step of obtaining the target knowledge graph of the field to which the question text belongs, entity tokens and relationship tokens in different fields can also be obtained; for the entity tokens and relationship tokens in a field, label different entity tokens through the relationship tokens to obtain knowledge data, where the knowledge data includes at least two entity tokens and at least one relationship token that associates at least two entity tokens; construct knowledge graphs for different fields according to the knowledge data in different fields.

[0065] In the embodiments of the present invention, for each field, a knowledge graph will be constructed. The entity tokens and relationship tokens in each field can be collected and sorted out, and the entity tokens and relationship tokens are associated according to the structure of the knowledge data to obtain the knowledge data.

[0066] Generally speaking, a knowledge graph is a node relationship graph, which represents knowledge through nodes and connection relationships. Each node corresponds to an entity, and the relationship between entities is the connection relationship between the corresponding nodes. Entities include subjects and objects. A piece of knowledge data in the knowledge graph is formed by entities and relationships. For example, "cat" is an entity token, "animal" is an entity token, and the relationship token is "belongs to", and the knowledge it represents is that a cat belongs to an animal; "grandma" is an entity token, "dad" is an entity token, "mom" is an entity token, the relationship token between grandma and dad is "is", and the relationship token between dad and mom is "of", and the knowledge it represents is that dad's mom is grandma. It should be noted that due to the complexity of the combination of entities and relationships, the knowledge graphs in the embodiments of the present invention are all domain- or scenario-oriented. In practical applications, a specific scenario is specified to construct a knowledge graph. For example, the scenarios are middle school education, corporate finance, surgical medical treatment, etc., and the corresponding knowledge graphs of middle school education, corporate finance, surgical medical treatment, etc. are constructed respectively. The construction of the knowledge graph can be obtained by directly annotating entities and relationships manually, or by generating a model and then verifying it manually. After the knowledge graph is constructed, its accuracy needs to be strictly checked. In order to achieve the application effect, the accuracy of the knowledge graph is required to be 100%.

[0067] Optionally, in the step of obtaining knowledge data by annotating different entity tokens with relationship tokens, the entity tokens can be divided into subject tokens and object tokens; the subject tokens and object tokens are annotated based on the relationship tokens to obtain knowledge data.

[0068] In the embodiments of the present invention, the above-mentioned subject token can represent the subject in a piece of knowledge data, and the above-mentioned object token can represent the object in a piece of knowledge data. A piece of knowledge data can be composed of a subject token, an object token, and a relationship token. For example, for a cat as the subject, an animal as the object, and the relationship is belongs to. The knowledge it represents is that a cat belongs to an animal.

[0069] The subject tokens and object tokens can be annotated with relationship tokens by manual annotation to obtain knowledge data, or knowledge data is generated by a model and then verified manually to obtain verified knowledge data.

[0070] By dividing the entity tokens into subject tokens and object tokens, the knowledge data can be made clearer and more standardized, the matchability of the knowledge graph is improved, and the matching speed of the knowledge data is increased.

[0071] Optionally, before the step of performing knowledge extraction on the output text to obtain the knowledge data to be detected of the output text, it is also possible to obtain a natural language processing model to be trained and a training data set. The training data set includes sample texts and knowledge labels of the sample texts. The input of the natural language processing model to be trained is constructed as the sample texts, and the output of the natural language processing model to be trained is constructed as predicted knowledge data; input the sample texts into the natural language processing model to obtain the predicted knowledge data of the sample texts; calculate the loss value between the predicted knowledge data of the sample texts and the knowledge labels of the sample texts; based on the loss value, adjust the parameters of the natural language processing model to be trained, and iterate the parameter adjustment process. After the training is completed, a trained natural language processing model is obtained, and the trained natural language processing model is used to perform knowledge extraction on the output text.

[0072] In an embodiment of the present invention, the trained natural language processing model can be used to perform knowledge extraction on the output text to obtain the knowledge data to be detected of the output text.

[0073] Specifically, a natural language processing model to be trained and a training data set are constructed, and the natural language processing model to be trained is subjected to supervised training through the training data set. After the training is completed, a trained natural language processing model is obtained.

[0074] The above-mentioned natural language processing model to be trained can be constructed based on RNN, LSTM, Transformer, etc., such as it can be a Bert, TF-IDF model, etc. The input of the natural language processing model to be trained is constructed as the text output, and the output of the natural language processing model to be trained is constructed as predicted knowledge data, that is, when a text is input into the natural language processing model to be trained, the output is the predicted knowledge data, and the number of predicted knowledge data can be one or more.

[0075] The above-mentioned training data set includes sample texts and knowledge labels of the sample texts. The knowledge labels can include one or more correct knowledge data, and the correct knowledge data is obtained based on manual annotation.

[0076] During the training process, sample texts can be input into the natural language processing model to be trained. The natural language processing model to be trained processes the sample texts to obtain the predicted knowledge data of the sample texts. The loss value between the predicted knowledge data of the sample texts and the knowledge labels of the sample texts is calculated through a preset loss function. Taking the minimization of the loss value as the optimization objective, the model parameters of the natural language processing model to be trained are updated and adjusted. The above update and adjustment process is iterated until the number of iterations reaches the preset number or the loss value converges at the minimum loss value, and the training is completed to obtain the trained natural language processing model. The above loss function can be a cross-entropy loss function, a mean squared error loss function, a logarithmic loss function, etc.

[0077] After the training is completed, it is also necessary to detect the trained natural language processing model to detect the accuracy of the trained natural language processing model, and take the accuracy of the trained natural language processing model that meets the accuracy as the final trained natural language processing model. For example, the knowledge data extraction accuracy of the trained natural language processing model needs to reach more than 95% before it can be applied in practice. If the accuracy requirement is not met, the accuracy is improved by adding training data to obtain the finally trained natural language processing model that can be applied in practice.

[0078] Optionally, before the step of extracting the knowledge of the output text to obtain the knowledge data to be detected of the output text, the natural language processing model to be trained and the training data set can also be obtained. The training data set includes sample texts, entity labels and relationship labels of the sample texts. The input of the natural language processing model to be trained is constructed as the sample text, and the output of the natural language processing model to be trained is constructed as the predicted entity and the predicted relationship; the sample text is input into the natural language processing model to obtain the predicted entity data and predicted relationship data of the sample text; the first loss value between the predicted entity data of the sample text and the entity label of the sample text, and the second loss value between the predicted relationship data of the sample text and the relationship label of the sample text are calculated; the parameters of the natural language processing model to be trained are adjusted based on the first loss value and the second loss value, and the parameter adjustment process is iterated. After the training is completed, the trained natural language processing model is obtained, and the trained natural language processing model is used to extract the knowledge of the output text.

[0079] In the embodiment of the present invention, the trained natural language processing model can be used to extract the knowledge of the output text to obtain the knowledge data to be detected of the output text.

[0080] Specifically, a natural language processing model to be trained and a training data set are constructed, and the natural language processing model to be trained is supervised by the training data set. After the training is completed, the trained natural language processing model is obtained.

[0081] The natural language processing model to be trained can be constructed based on RNN, LSTM, Transformer, etc., such as Bert, TF-IDF model, etc. The input of the natural language processing model to be trained is constructed as a text output, and the output of the natural language processing model to be trained is constructed as predicted entity data and predicted relationship data, that is, a text is input into the natural language processing model to be trained, and the output is predicted entity data and predicted relationship data. The number of predicted entity data and predicted relationship data can both be one or more.

[0082] The above training dataset includes sample texts, as well as entity labels and relationship labels of the sample texts. The entity labels and relationship labels are obtained based on manual annotation.

[0083] During the training process, the sample text can be input into the natural language processing model to be trained, and the sample text is processed by the natural language processing model to be trained to obtain the predicted entity data and predicted relationship data of the sample text. The first loss value between the predicted entity data of the sample text and the entity label of the sample text is calculated through a preset loss function, and the second loss value between the predicted relationship data of the sample text and the relationship label of the sample text is calculated. The first loss value and the second loss value are added together to obtain the total loss value. The model parameters of the natural language processing model to be trained are updated and adjusted with the goal of minimizing the total loss value. The above update and adjustment process is iterated until the number of iterations reaches the preset number, or the total loss value converges at the minimum loss value, and the training is completed to obtain the trained natural language processing model. The above loss function can be a cross-entropy loss function, a mean squared error loss function, a logarithmic loss function, etc.

[0084] After the training is completed, it is also necessary to detect the trained natural language processing model to detect the accuracy of the trained natural language processing model, and the accuracy of the trained natural language processing model that meets the accuracy is used as the final trained natural language processing model. For example, the knowledge data extraction accuracy of the trained natural language processing model needs to reach more than 95% before it can be applied in practice. If the accuracy requirement is not met, the accuracy is improved by adding training data to obtain the finally trained natural language processing model that can be applied in practice.

[0085] Optionally, in the step of extracting knowledge from the output text to obtain the knowledge data to be detected of the output text, the trained natural language processing model can be used to extract knowledge from the output text to obtain the entities to be detected and relationships to be detected of the output text; the entities to be detected and relationships to be detected are combined to obtain the knowledge data of the output text, and the knowledge data to be detected includes entities to be detected and relationships to be detected.

[0086] In the embodiments of the present invention, based on the output text of the large language model, relationship and entity recognition can be performed through a trained natural language processing model. Entity recognition is classified into person names, referring nouns, organizations, locations, times, etc. Relationship recognition is classified into inheritance relationships, synonymous relationships, attribution relationships, correlation relationships, causal relationships, etc.

[0087] In one embodiment, through sequence labeling technology, entity extraction can be performed, and recognition and classification are carried out through Bert. Relationship classification is directly based on the extracted entities and is classified based on Bert. Based on a Bert model with a qualified accuracy, the output text of the large language model is obtained, and the output text of the large model is used as the input text of Bert. After obtaining the output of Bert, all entity and relationship combinations in the output text can be formed.

[0088] Optionally, in the step of detecting the knowledge data to be detected through the target knowledge graph and obtaining the knowledge detection result of the output text, it can be determined whether the entity to be detected and the relationship to be detected exist in the target knowledge graph; if the entity to be detected exists in the target knowledge graph and the relationship to be detected also exists in the target knowledge graph, it is determined that the knowledge detection of the output text is correct; if the entity to be detected does not exist in the target knowledge graph and / or the relationship to be detected does not exist in the target knowledge graph, it is determined that the knowledge detection of the output text is incorrect.

[0089] In the embodiments of the present invention, based on graph matching technology, knowledge secondary inspection of the knowledge graph can be achieved through node and relationship matching. Specifically, all entity and relationship combinations are input into the target graph. First, it is determined whether the entity to be detected exists in the target graph. If any entity to be detected does not exist in the target graph, it is determined that the knowledge detection of the output text is incorrect, and the knowledge detection error is returned. If all entities to be detected exist in the target graph, then the next step is to determine whether the relationship to be detected exists in the target graph. If any relationship to be detected does not exist in the target graph, it is determined that the knowledge detection of the output text is incorrect, and the knowledge detection error is returned. If all relationships to be detected exist in the target graph, it is determined that the knowledge detection of the output text is correct, and the knowledge detection correct is returned.

[0090] Specifically, in combination with Figure 2 it is described as follows. Figure 2 is a flowchart of another output detection method for the large language model provided by the embodiments of the present invention. In Figure 2 it includes the following steps:

[0091] 201. Text input.

[0092] The above text input can be the question text of the user.

[0093] 202. Large language model LLM.

[0094] Process the text input through a large language model.

[0095] 203. Text output y.

[0096] The above text output y is the output text obtained from the large language model.

[0097] 204. Knowledge graph detection.

[0098] Perform knowledge detection on the output text through the target knowledge graph. If the detection fails, go to step 205; if the detection passes, go to step 206.

[0099] 205. Reject to answer.

[0100] 206. Secondary confirmation output.

[0101] Further, in combination with Figure 3 for illustration, Figure 3 is a flowchart of a knowledge graph detection method provided by an embodiment of the present invention. In Figure 3 it includes the following steps:

[0102] 301. Relationship and entity extraction based on Bert.

[0103] Extract the relationships and entities in the text output y based on Bert to obtain the entity-relationship pair set C of the text output.

[0104] 302. For any element c ∈ C, check whether c ∈ G.

[0105] Detect whether any element c in the set C is also in the target graph G.

[0106] 303. Return the knowledge graph detection result.

[0107] If any element c ∈ C satisfies c ∈ G, it can be determined that the detection passes; otherwise, it is determined that the detection fails.

[0108] Further, based on a Bert model with qualified accuracy, obtain the output of the large language model and use the output of the large model as the Bert input. After obtaining the output of Bert, all entity and relationship combinations can be formed. That is, the output is a series of "subject-relationship-object" knowledge points. Based on the graph matching technology, perform secondary knowledge detection of the knowledge graph through node and relationship matching. In combination with Figure 4 for illustration, Figure 4 is a flowchart of a secondary detection method provided by an embodiment of the present invention. In Figure 4 it includes the following steps:

[0109] 401. Input all combinations of entities and relationships.

[0110] Use all combinations of entities and relationships output by Bert as input to perform node and relationship matching with the target knowledge graph G.

[0111] 402. Check if the matched entity is in G.

[0112] If the matched entity is in the target knowledge graph G, go to step 403; if the matched entity is not in the target knowledge graph G, go to step 406.

[0113] 403. Check if the matched relationship is in G.

[0114] If the matched relationship is in the target knowledge graph G, go to step 404; if the matched relationship is not in the target knowledge graph G, go to step 406.

[0115] 404. Return that the knowledge detection is correct.

[0116] 405. Return that the knowledge detection is incorrect.

[0117] Through the secondary check of the knowledge graph, the hallucination problem in the output of the large language model is alleviated, the implementation effect of the large language model is improved, and its usability is increased.

[0118] As Figure 5 shown, an embodiment of the present invention provides an output detection device for a large language model. The output detection device for the large language model includes:

[0119] A first acquisition module 501, configured to acquire the output text of the large language model for the question text, and acquire the target knowledge graph of the field to which the question text belongs, where the target knowledge graph includes knowledge data of the field to which the question text belongs;

[0120] An extraction module 502, configured to perform knowledge extraction on the output text to obtain the knowledge data to be detected of the output text;

[0121] A detection module 503, configured to detect the knowledge data to be detected through the target knowledge graph to obtain the knowledge detection result of the output text.

[0122] Optionally, the device further includes:

[0123] A second acquisition module, configured to acquire entity lemmas and relationship lemmas in different fields.

[0124] The first knowledge processing module is used to label different entity tokens in a field through the relationship tokens, so as to obtain knowledge data, where the knowledge data includes at least two entity tokens and at least one relationship token that associates at least two entity tokens.

[0125] The second knowledge processing module is used to construct knowledge graphs for different fields according to the knowledge data of different fields.

[0126] Optionally, the first knowledge processing module is further used to divide the entity tokens into subject tokens and object tokens; and label the subject tokens and the object tokens based on the relationship tokens to obtain knowledge data.

[0127] Optionally, the device further includes:

[0128] The third acquisition module is used to acquire a natural language processing model to be trained and a training data set, where the training data set includes sample texts and knowledge labels of the sample texts, the input of the natural language processing model to be trained is constructed as the sample texts, and the output of the natural language processing model to be trained is constructed as predicted knowledge data;

[0129] The first processing module is used to input the sample texts into the natural language processing model to obtain the predicted knowledge data of the sample texts;

[0130] The second processing module is used to calculate the loss value between the predicted knowledge data of the sample texts and the knowledge labels of the sample texts;

[0131] The third processing module is used to adjust the parameters of the natural language processing model to be trained based on the loss value, and iterate the parameter adjustment process. After training is completed, a trained natural language processing model is obtained, and the trained natural language processing model is used to extract knowledge from the output text.

[0132] Optionally, the device further includes:

[0133] The fourth acquisition module is used to acquire a natural language processing model to be trained and a training data set, where the training data set includes sample texts and entity labels and relationship labels of the sample texts, the input of the natural language processing model to be trained is constructed as the sample texts, and the output of the natural language processing model to be trained is constructed as predicted entities and predicted relationships;

[0134] The fourth processing module is used to input the sample texts into the natural language processing model to obtain the predicted entity data and predicted relationship data of the sample texts;

[0135] A fifth processing module, configured to calculate a first loss value between the predicted entity data of the sample text and the entity label of the sample text, and a second loss value between the predicted relationship data of the sample text and the relationship label of the sample text;

[0136] A sixth processing module, configured to adjust the parameters of the natural language processing model to be trained based on the first loss value and the second loss value, and iterate the parameter adjustment process. After the training is completed, a trained natural language processing model is obtained, and the trained natural language processing model is used to extract knowledge from the output text.

[0137] Optionally, the extraction module 502 is further configured to extract knowledge from the output text through the trained natural language processing model to obtain the entities to be detected and relationships to be detected in the output text; combine the entities to be detected and relationships to be detected to obtain the knowledge data of the output text, where the knowledge data to be detected includes entities to be detected and relationships to be detected.

[0138] Optionally, the detection module 503 is further configured to determine whether the entities to be detected and the relationships to be detected exist in the target knowledge graph; if the entities to be detected exist in the target knowledge graph and the relationships to be detected also exist in the target knowledge graph, it is determined that the knowledge detection of the output text is correct; if the entities to be detected do not exist in the target knowledge graph and / or the relationships to be detected do not exist in the target knowledge graph, it is determined that the knowledge detection of the output text is incorrect.

[0139] The output detection device of the large language model provided by the embodiments of the present invention can implement each process implemented by the output detection method of the large language model in the above method embodiments, and can achieve the same beneficial effects. To avoid repetition, it will not be elaborated here.

[0140] See Figure 6 , Figure 6 is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. As Figure 6 shown, it includes: a memory 602, a processor 601, and a computer program of the output detection method of the large language model stored on the memory 602 and executable on the processor 601, where:

[0141] The processor 601 is configured to call the computer program stored in the memory 602 and execute the following steps:

[0142] Obtain the output text of the large language model for the question text, and obtain the target knowledge graph of the field to which the question text belongs, where the target knowledge graph includes the knowledge data of the field to which the question text belongs;

[0143] Extract knowledge from the output text to obtain the knowledge data to be detected of the output text;

[0144] Detect the knowledge data to be detected through the target knowledge graph to obtain the knowledge detection result of the output text.

[0145] Optionally, before obtaining the target knowledge graph of the field to which the question text belongs, the method executed by the processor 601 further includes:

[0146] Obtain entity tokens and relationship tokens in different fields;

[0147] For the entity tokens and the relationship tokens in one field, label between different entity tokens through the relationship tokens to obtain knowledge data, where the knowledge data includes at least two entity tokens and at least one relationship token that associates at least two entity tokens;

[0148] Construct knowledge graphs for different fields according to the knowledge data in different fields.

[0149] Optionally, the method executed by the processor 601 to label between different entity tokens through the relationship tokens to obtain knowledge data includes:

[0150] Divide the entity tokens into subject tokens and object tokens;

[0151] Label the subject tokens and the object tokens based on the relationship tokens to obtain knowledge data.

[0152] Optionally, before extracting knowledge from the output text to obtain the knowledge data to be detected of the output text, the method executed by the processor 601 further includes:

[0153] Obtain a natural language processing model to be trained and a training data set, where the training data set includes sample texts and knowledge labels of the sample texts, the input of the natural language processing model to be trained is constructed as the sample texts, and the output of the natural language processing model to be trained is constructed as predicted knowledge data;

[0154] Input the sample texts into the natural language processing model to obtain the predicted knowledge data of the sample texts;

[0155] Calculate the loss value between the predicted knowledge data of the sample texts and the knowledge labels of the sample texts;

[0156] Based on the loss value, adjust the parameters of the natural language processing model to be trained, and iterate the parameter adjustment process. After training is completed, a trained natural language processing model is obtained, and the trained natural language processing model is used to extract knowledge from the output text.

[0157] Optionally, before extracting the knowledge of the output text to obtain the knowledge data to be detected of the output text, the method executed by the processor 601 further includes:

[0158] Obtain a natural language processing model to be trained and a training data set. The training data set includes sample texts, entity labels, and relationship labels of the sample texts. The input of the natural language processing model to be trained is constructed as a sample text, and the output of the natural language processing model to be trained is constructed as predicted entities and predicted relationships;

[0159] Input the sample text into the natural language processing model to obtain the predicted entity data and predicted relationship data of the sample text;

[0160] Calculate a first loss value between the predicted entity data of the sample text and the entity label of the sample text, and a second loss value between the predicted relationship data of the sample text and the relationship label of the sample text;

[0161] Based on the first loss value and the second loss value, adjust the parameters of the natural language processing model to be trained, and iterate the parameter adjustment process. After training is completed, a trained natural language processing model is obtained, and the trained natural language processing model is used to extract knowledge from the output text.

[0162] Optionally, the processor 601 executes the extraction of the knowledge of the output text to obtain the knowledge data to be detected of the output text, including:

[0163] Extract the knowledge of the output text through the trained natural language processing model to obtain the entities to be detected and relationships to be detected of the output text;

[0164] Combine the entities to be detected and relationships to be detected to obtain the knowledge data of the output text. The knowledge data to be detected includes entities to be detected and relationships to be detected.

[0165] Optionally, the processor 601 executes the detection of the knowledge data to be detected through the target knowledge graph to obtain the knowledge detection result of the output text, including:

[0166] Determine whether the entities to be detected and the relationships to be detected exist in the target knowledge graph;

[0167] If the entity to be detected exists in the target knowledge graph and the relationship to be detected also exists in the target knowledge graph, it is determined that the knowledge detection of the output text is correct;

[0168] If the entity to be detected does not exist in the target knowledge graph and / or the relationship to be detected does not exist in the target knowledge graph, it is determined that the knowledge detection of the output text is incorrect.

[0169] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements each process of the output detection method of the large language model provided by the embodiment of the present invention and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0170] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the computer-readable storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0171] The above-disclosed are only the preferred embodiments of the present invention. Of course, the scope of the rights of the present invention cannot be limited thereby. Therefore, equivalent changes made according to the claims of the present invention still fall within the scope covered by the present invention.

Claims

1. An output detection method for a large language model, characterized in that, Including: Obtaining the output text of the large language model for the query text, and obtaining the target knowledge graph of the field to which the query text belongs, where the target knowledge graph includes knowledge data of the field to which the query text belongs; Performing knowledge extraction on the output text to obtain the knowledge data to be detected of the output text; Detecting the knowledge data to be detected through the target knowledge graph to obtain the knowledge detection result of the output text.

2. The method according to claim 1, characterized in that, Before obtaining the target knowledge graph of the field to which the query text belongs, the method further includes: Obtaining entity tokens and relationship tokens in different fields; For the entity tokens and the relationship tokens in a field, annotating different entity tokens through the relationship tokens to obtain knowledge data, where the knowledge data includes at least two entity tokens and at least one relationship token that associates the at least two entity tokens; Constructing knowledge graphs of different fields according to the knowledge data of different fields.

3. The method according to claim 2, wherein The annotating different entity tokens through the relationship tokens to obtain knowledge data includes: Dividing the entity tokens into subject tokens and object tokens; Annotating the subject tokens and the object tokens based on the relationship tokens to obtain knowledge data.

4. The method according to claim 1, wherein Before performing knowledge extraction on the output text to obtain the knowledge data to be detected of the output text, the method further includes: Obtaining a natural language processing model to be trained and a training data set, where the training data set includes sample texts and knowledge labels of the sample texts, the input of the natural language processing model to be trained is constructed as the sample texts, and the output of the natural language processing model to be trained is constructed as predicted knowledge data; Inputting the sample texts into the natural language processing model to obtain the predicted knowledge data of the sample texts; Calculating the loss value between the predicted knowledge data of the sample texts and the knowledge labels of the sample texts; Adjusting the parameters of the natural language processing model to be trained based on the loss value and iterating the parameter adjustment process. After training is completed, a trained natural language processing model is obtained, and the trained natural language processing model is used to perform knowledge extraction on the output text.

5. The method according to claim 1, characterized in that, Before performing knowledge extraction on the output text to obtain the knowledge data to be detected of the output text, the method further includes: Obtaining a natural language processing model to be trained and a training data set, where the training data set includes sample texts, entity labels, and relationship labels of the sample texts, the input of the natural language processing model to be trained is constructed as the sample texts, and the output of the natural language processing model to be trained is constructed as predicted entities and predicted relationships; Inputting the sample texts into the natural language processing model to obtain the predicted entity data and predicted relationship data of the sample texts; Calculating a first loss value between the predicted entity data of the sample texts and the entity labels of the sample texts, and a second loss value between the predicted relationship data of the sample texts and the relationship labels of the sample texts; Based on the first loss value and the second loss value, adjust the parameters of the natural language processing model to be trained, and iterate the parameter adjustment process. After the training is completed, obtain a trained natural language processing model, which is used to extract knowledge from the output text.

6. The method according to claim 5, characterized in that, The extracting knowledge from the output text to obtain the knowledge data to be detected of the output text includes: Extract knowledge from the output text through the trained natural language processing model to obtain the entities to be detected and the relationships to be detected in the output text; Combine the entities to be detected and the relationships to be detected to obtain the knowledge data of the output text, where the knowledge data to be detected includes the entities to be detected and the relationships to be detected.

7. The method according to claim 6, wherein The detecting the knowledge data to be detected through the target knowledge graph to obtain the knowledge detection result of the output text includes: Determine whether the entities to be detected and the relationships to be detected exist in the target knowledge graph; If the entity to be detected exists in the target knowledge graph and the relationship to be detected also exists in the target knowledge graph, it is determined that the knowledge detection of the output text is correct; If the entity to be detected does not exist in the target knowledge graph and / or the relationship to be detected does not exist in the target knowledge graph, it is determined that the knowledge detection of the output text is incorrect.

8. An output detection device for a large language model, characterized in that, The output detection device of the large language model includes: A first acquisition module, configured to acquire the output text of the large language model for the question text, and acquire the target knowledge graph of the field to which the question text belongs, where the target knowledge graph includes the knowledge data of the field to which the question text belongs; An extraction module, configured to extract knowledge from the output text to obtain the knowledge data to be detected of the output text; A detection module, configured to detect the knowledge data to be detected through the target knowledge graph to obtain the knowledge detection result of the output text.

9. An electronic device, characterized in that, including: A memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, the steps in the output detection method of the large language model according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by the processor, the steps in the output detection method of the large language model according to any one of claims 1 to 7 are implemented.