Large model reasoning process detection method and device, medium and equipment

By obtaining training samples marked with output results, determining the activation status of LLM neurons, and training the detection classifier, the problem of LLM output hallucination is solved, and efficient and accurate detection of LLM output results is achieved, saving costs and suitable for a variety of LLMs.

CN120541230AActive Publication Date: 2025-08-26ZHEJIANG ANT MISUAN TECHNOLOGY CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511037213.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-08-26
Estimated Expiration
2045-07-25

AI Technical Summary

Technical Problem

When performing tasks, large language models (LLMs) may cause problems such as lack of knowledge, knowledge errors or expired knowledge due to the limitations of training data, resulting in the output of incorrect results (illusions). The existing solutions are costly and difficult to balance the effects.

Method used

By obtaining training samples marked with the correctness of the output results, determine the neuron activation status of the LLM to be detected, train the detection classifier, and use the neuron activation status to detect the correctness of the LLM output results, avoiding attention to the LLM inference process and output results.

Benefits of technology

It realizes the recognition of illusions from the bottom of the model when LLM performs tasks, avoids outputting error results, and saves training and labor costs. It is suitable for different LLMs and has better versatility and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120541230A_ABST
    Figure CN120541230A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a large model output result detection method, and the method comprises the steps: determining the activation state of neurons of a to-be-detected large model LLM in a reasoning process based on a training sample marked with whether the output result is correct or not, and obtaining the relation between whether the output result is correct or not and the activation state of the neurons, the method comprises the following steps: training a detection classifier by using a neural network to obtain an activation state of a neuron when a to-be-detected LLM executes task input data, and detecting the correctness of an LLM output result through the detection classifier, namely whether illusion appears or not. According to the method, whether the illusion occurs in the LLM or not is identified by associating the activation state of the LLM neurons with the illusion occurring in the LLM, and whether the illusion occurs in the LLM or not is identified from the bottom layer of the model without paying attention to the LLM reasoning process and the output result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of computer technology, and in particular to a large model reasoning process detection method, device, storage medium and equipment. Background Art

[0002] With the development of artificial intelligence (AI) technology, large language models (LLM) have been widely used in various fields.

[0003] However, due to the limitations of training data, LLM may suffer from problems such as lack of knowledge, knowledge errors, or outdated knowledge. These problems may cause LLM to experience "hallucinations" when performing tasks, thereby reducing the accuracy and reliability of LLM's task execution.

[0004] Based on this, detecting the "hallucination" of LLM during the reasoning process and preventing LLM from outputting wrong answers has become an urgent problem to be solved. Therefore, this manual provides a large model reasoning process detection method. Summary of the Invention

[0005] The embodiments of this specification provide a large model output result detection method, device, storage medium and electronic device to partially solve the problems existing in the above-mentioned prior art.

[0006] The embodiments of this specification adopt the following technical solutions: This specification provides a method for detecting large model output results, the method comprising: Obtaining training samples for training the large model, wherein the training samples include annotations indicating whether the model output results are correct; Determine the activation state of a neuron in the large model to be detected based on the input of the training sample as the activation state corresponding to the training sample; Taking the activation state corresponding to the training sample as input and labeling whether the output result is correct, a detection classifier is trained; In response to the activation state of the neurons of the large model to be detected executing the business process, the correctness of the output result of the large model to be detected is detected by the detection classifier.

[0007] This specification provides a large model output result detection device, the device comprising: An acquisition module is used to obtain training samples for training the large model, wherein the training samples include a label indicating whether the model output result is correct; An inference module, configured to determine, based on the input of the training sample, an activation state of a neuron in the large model to be tested as the activation state corresponding to the training sample; A training module is used to take the activation state corresponding to the training sample as input, mark whether the output result is correct, and train the detection classifier; The detection module is used to detect the correctness of the output result of the large model to be detected through the detection classifier in response to the activation state of the neurons in the business process executed by the large model to be detected.

[0008] This specification provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned large model output result detection method.

[0009] This specification provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the above-mentioned large model output result detection method is implemented.

[0010] At least one of the above technical solutions adopted in the embodiments of this specification can achieve the following beneficial effects: The embodiments of this specification disclose a method for detecting the output results of a large model. This method is based on training samples marked with whether the output results are correct, and determines the activation state of neurons in the large model LLM to be detected during the reasoning process. The relationship between the correctness of the output result and the activation state of the neurons is obtained to train a detection classifier. When the LLM to be detected executes the task input data, the activation state of the neurons is obtained, and the correctness of the LLM output result, that is, whether "hallucination" occurs, is detected by the detection classifier. This method identifies whether the LLM has hallucinated by linking the activation state of the LLM neurons with the relationship between the LLM's hallucination. It does not need to pay attention to the LLM's reasoning process and output results. It identifies whether the LLM has hallucinated from the bottom of the model to avoid executing business based on the output result of "hallucination". BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The drawings described herein are used to provide a further understanding of this specification and constitute a part of this specification. The exemplary embodiments and descriptions of this specification are used to explain this specification and do not constitute an improper limitation of this specification. In the drawings: Figure 1 A flow chart of a large model output result detection provided in the embodiments of this specification; Figure 2 A schematic diagram of recording neuron activation states provided in an embodiment of this specification; Figure 3 A schematic diagram of the large model output result detection process provided in the embodiments of this specification; Figure 4A schematic diagram of a large model output result detection device provided in an embodiment of this specification; Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this specification. DETAILED DESCRIPTION

[0012] To make the objectives, technical solutions, and advantages of this specification more clear, the following will clearly and completely describe the technical solutions of this specification in conjunction with the specific embodiments of this specification and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this specification.

[0013] A large language model (LLM), also known as a large model, performs logical reasoning on natural language input and outputs user-understandable natural language descriptions. While LLMs can generate fluent and logical content, their output is not always based on facts, a phenomenon known as "hallucination."

[0014] Existing methods for addressing hallucinations in LLMs include embedding knowledge into the LLM through supervised fine-tuning (SFT) to identify the occurrence of hallucinations. Alternatively, by labeling training samples that cause hallucinations, the LLM learns the conditions under which hallucinations occur and then corrects the output results that produce hallucinations in actual business scenarios. However, due to the differences between different LLMs, the effects of training samples on different LLMs may vary, necessitating adjustments to the training samples.

[0015] However, this also requires a large number of targeted training samples for training, which not only increases training costs but also consumes a considerable amount of time to fine-tune the LLM. Furthermore, fine-tuning the LLM can cause conflicts between training tasks, making it difficult to balance training results. For example, there could be a conflict between training tasks that require the LLM to produce accurate results and training tasks that require the LLM to produce "illusions."

[0016] The technical solutions provided by the embodiments of this specification are described in detail below with reference to the accompanying drawings.

[0017] Figure 1 A flow chart for detecting a large model output result provided in an embodiment of this specification specifically includes the following steps: S100: Acquire training samples for training a large model, wherein the training samples include annotations indicating whether the model output result is correct.

[0018] In the embodiments of this specification, Figure 1 The device for performing the large model output result detection method shown can be any electronic device, such as a computer, a server, or a server cluster composed of multiple servers, etc. For the convenience of description, the following description only takes the server as an example.

[0019] In order to be able to detect whether the LLM output results have "hallucinations", the server needs to obtain training samples that are marked with "hallucinations" and those that do not. Specifically, the server can first obtain training samples for training LLM from an open source database. The training samples include: model input, output results, and annotations of whether "hallucinations" have occurred. Model input refers to the data input into LLM. Output results refer to the response that LLM should output for the model input. The annotation of whether "hallucinations" have occurred refers to the situation where LLM has model "hallucinations" when performing logical reasoning based on the model input and determining the output results of the training samples. Of course, since the number of training samples marked with "hallucinations" may be small, the training samples required for the embodiments of this specification can also be obtained by modifying the samples that are not marked with "hallucinations".

[0020] For example, a large number of LLM training samples are obtained and annotations are added for the samples that do not have “hallucinations”. The training samples are then fed into the LLM to be trained, and the incorrect output of the LLM to be trained is used as the output of the training samples, and annotations for the samples that have “hallucinations” are added.

[0021] Of course, in the embodiments of this specification, there is no limitation on the form of annotation for whether "hallucination" occurs. It can be text described in natural language, or represented by a different identifier. For example, 0 indicates no "hallucination" occurs, and 1 indicates the presence of "hallucination." The specific setting can be determined as needed and is not limited in this specification. In addition, the model input and output results included in the above training samples are all text described in natural language.

[0022] S102: Determine the activation state of the neurons in the large model to be detected based on the input of the training sample as the activation state corresponding to the training sample.

[0023] After the server obtains the training sample, it can input it into the LLM to be tested. It then records the activation state of each neuron in the LLM as it infers and produces its output, using this as the activation state corresponding to the training sample. The LLM can be pre-deployed on the server or other devices. If deployed on other devices, the server can send the business logic process to the other devices, allowing them to input the received business logic process into the LLM. The following description uses the LLM deployed on the server as an example.

[0024] Of course, the neurons in the LLM described in the embodiments of this specification are mathematical functions in the LLM, that is, the basic computing units that constitute the neural network. The existing neural network model, that is, the basic computing principle of the LLM, is to achieve complex pattern recognition and reasoning capabilities by simulating the "input-processing-output" behavior of biological neurons. Of course, from the perspective of the model structure, each basic computing function can be regarded as a neuron, and a computing set composed of multiple basic mathematical functions can be regarded as a neuron. This specification does not impose specific restrictions on the structure of neurons. Neurons can be determined according to the model structure designed when the LLM is constructed, or the most basic mathematical function can be used as a neuron and set as needed. Generally speaking, the neurons designed when constructing the LLM are sufficient to implement the embodiments of this specification, and this example will be used for explanation later.

[0025] For different model inputs, the LLM's internal neuron activation states differ when it infers the output. Recording the neuron activation states in the LLM corresponding to different training samples can characterize how the LLM arrives at its output. Subsequently, different inference representations can be correlated with whether the output exhibits hallucinations, thereby determining under what inference representations the LLM under test exhibits hallucinations. For example, when the LLM cannot arrive at the correct output, it will often rely on various methods to generate outputs that "look like the answer," thus exhibiting hallucinations. The neuron activation states that produce these "look like the answer" outputs inevitably share commonalities. Therefore, by identifying these common neuron activation states in these hallucinations, subsequent steps can provide a means to identify hallucinations.

[0026] Furthermore, in the embodiments of this specification, for LLM neurons, regardless of the data input to the neuron, the corresponding activation state is generally discrete, typically being activated and continuing to output data to the next neuron, or being inactivated and not outputting data. Therefore, the LLM neuron activation states recorded by the server can simply record whether each neuron is activated. This simpler information represents the LLM reasoning process, making it easier to identify common points in neuron activation states across different training samples.

[0027] Alternatively, the server can fully record the specific activation states of neurons. For example, the inputs that triggered them, or the inputs that prevented them, can be used to record the activation states of neurons. This allows subsequent steps to draw deeper connections between neuron activation states and hallucinations based on richer information, thereby increasing the accuracy of detecting hallucinations in LLM.

[0028] Figure 2 This is a schematic diagram of recording the activation state of neurons provided in an embodiment of this specification. Figure 2 The figure below is a schematic diagram of the structure of a neural network model, which shows the data transmission and neuron status after a training sample is input into the model. Among them, dark dots represent activated neurons, and light dots represent inactivated neurons. After the model outputs the result, record which neurons are activated, such as Figure 2 In the figure, a through d represent a four-layer model structure, with neurons represented by numbers from top to bottom, i.e., a1 through a4, b1 through b6, and so on. The neuron activation states can be determined as (a1, a2, a3, a4, b1, b2, b5, c4, c5, d1, d2) as the activation states corresponding to the training samples.

[0029] S104: Taking the activation state corresponding to the training sample as input, marking whether the output result is correct, and training a detection classifier.

[0030] After obtaining the neuron activation states corresponding to each training sample, the server can train a detection classifier to classify the LLM output results based on the neuron activation states. The detection classifier is trained using the activation states corresponding to the training samples as input, and the correctness of the output results in the training sample annotations as the classification results. The detection classifier trained using the classification samples does not care about the LLM output results or the LLM model inputs. It only needs to obtain the neuron activation states in the LLM to classify whether the LLM neuron activation states are correct, that is, to detect whether "hallucinations" occur.

[0031] In the embodiments of this specification, the content generated by the LLM after performing logical reasoning based on the model input is called the output result. The output results are divided into correct output results and incorrect output results. When the LLM experiences "hallucination", the LLM outputs an incorrect output result. Therefore, the output of the detection classifier is the classification of whether the LLM output result is correct or incorrect.

[0032] In addition, in the embodiments of this specification, the type of detection classifier is not limited. A classifier based on random forest, or an optimized distributed gradient boosting library (eXtreme Gradient Boosting, XGBoost), a gradient boosting decision tree (Gradient Boosting Decision Tree, GBDT) and other different classifier model structures can be selected, and the specific settings can be made according to needs.

[0033] Of course, in the end, the server trains the detection classifier through the generated classification samples, which is a classifier for the LLM to be detected, because the activation state corresponding to the training samples contained in the classification samples is obtained through the LLM to be detected. This specification does not exclude the application of the detection classifier on other LLMs, but it should be noted that the neurons and model structures of other LLMs need to be consistent with the LLM to be detected, otherwise the accuracy of the detection classifier is difficult to guarantee. For example, if a detection classifier trained for an LLM with 1000 neurons is used for an LLM containing 3000 neurons, it is obvious that the detection classifier is difficult to give a reliable classification result.

[0034] S106: In response to the activation state of the neurons in the business process executed by the large model to be detected, the correctness of the output result of the large model to be detected is detected by the detection classifier.

[0035] In the embodiments of this specification, the detection classifier trained through the aforementioned steps can be used to detect the correctness of the LLM output results. Of course, the detection process is usually performed by a server, but this specification does not limit whether it is the server that trained the detection classifier or another server.

[0036] Furthermore, since it is necessary to obtain the activation state of the neurons in the LLM to be tested, the server used to detect the correctness of the output result can be the same server as the server that performs the service using the LLM to be tested. That is, after the server uses the LLM to be tested to perform the service and determines the output result, it inputs the activation state of the neurons in the LLM to be tested into the detection classifier to determine whether the output result is correct. If the server performing the test and the server performing the service are different devices, the server performing the service records the activation state of the neurons in the LLM to be tested and sends it to the server performing the test. The server performing the test uses the detection classifier to determine whether the output result is correct, and then returns it to the server performing the service. The server performing the service determines whether to perform the service based on the output result based on the feedback result. In other words, when the output result is incorrect, the service is not executed based on the output result. When the output result is correct, the service is executed based on the output result.

[0037] Specifically, first, in response to the input data of the large model to be tested, the server can determine the activation state of the neurons in the large model to be tested, which serves as the test data, that is, the input of the detection classifier. The process of the server obtaining the activation state of the neurons is consistent with the process described in step S102, and is therefore not further described.

[0038] Then, the server inputs the detection data into the trained detection classifier, and detects the correctness of the output results of the large model to be tested based on the output of the detection classifier.

[0039] In the embodiments of this specification, if the LLM to be detected is used to execute a service, then in the service scenario, the LLM to be detected is a service model. Then, when the server determines that the output result of the service model is correct based on the output of the detection classifier, it can continue to execute the service based on the output result.

[0040] When the server determines, based on the output of the detection classifier, that the business model's output is erroneous and that the business cannot be executed based on that output, the server then determines subsequent actions based on the business process and continues executing the business. The specific business process specifies how to handle errors in the LLM output, i.e., when "hallucinations" occur. This specification does not impose any restrictions and can be configured as needed. For example, the user may be asked to modify the text they entered, and based on the text re-entered by the user after the modifications, the LLM to be tested is used to obtain the output result again, and the test is then repeated.

[0041] It should be noted that in step S102, the above text only emphasizes the role of the wrong output results and the activation state of the neurons in training the detection classifier. But on the other hand, when LLM outputs the correct output result, the activation state of the neurons corresponding to it must be different from the activation state of the neurons corresponding to the wrong output result output by LLM. Therefore, even if there are only a small number of training samples or no training samples marked with "hallucinations", they can be used for the subsequent training of the detection classifier to determine whether the LLM output result is correct. It is equivalent to the detection classifier trained in the subsequent steps, which only needs to determine whether the activation state of the neurons corresponds to the correct output result. Therefore, in the embodiments of this specification, the need for training samples marked with "hallucinations" is not a necessary condition, but the training samples are all high-quality samples, and the output results they contain are all correct.

[0042] based on Figure 1 The large model output result detection method shown in the figure is based on training samples labeled with whether the output result is correct. It also determines the activation state of the neurons in the large model (LLM) to be tested during the inference process. The relationship between the correct output result and the neuron activation state is obtained to train a detection classifier. When the LLM to be tested performs a task input data, the neuron activation state is obtained, and the detection classifier is used to detect the correctness of the LLM output result, that is, whether "hallucination" occurs. This method identifies whether the LLM is hallucinating by linking the activation state of the LLM neurons with the occurrence of hallucinations in the LLM. This method does not focus on the LLM's inference process and output results, but identifies whether the LLM is hallucinating from the bottom of the model.

[0043] Compared to existing methods for addressing model hallucinations, such as manually removing false data from training datasets—data that causes model hallucinations—to maintain the accuracy of LLM training, this method relies on human experience and judgment, which can be influenced by human subjectivity, resulting in less objective and accurate data filtering. Furthermore, manual filtering of training datasets requires significant time and human resources. Rule-based methods for retrieving training datasets rely on expert experience to determine rules. Furthermore, these rules struggle to cover all types of false data and still require high labor costs, making it impossible to completely eliminate problematic false data from training datasets. While SFT can significantly improve the LLM's ability to output accurate results, it relies heavily on fine-tuning the LLM with large amounts of high-quality, labeled data. This type of data is difficult to obtain, and the fine-tuning process requires significant hardware resources.

[0044] However, the large-model output result detection method provided in this specification does not require retraining or fine-tuning the LLM, which is more convenient. It also does not require manual elimination of false data in the training dataset, which can save a lot of manpower and time costs. Similarly, it does not rely on expert experience and rules, and the above steps can be implemented on different LLMs to achieve the same effect. Compared with the situation where the search rules may only apply to a single LLM, this method has greater versatility.

[0045] In the embodiments of this specification, the training samples obtained in step S100 should match the LLM to be tested. Because the knowledge required by LLMs used in different business scenarios may vary, and the orientation of the output results may also differ, the LLM used in general business scenarios needs to be fine-tuned through the corresponding SFT process. For example, assuming the business scenario is a legal consulting scenario, a general LLM is first obtained. Then, based on legal knowledge, training samples are determined to perform SFT on the LLM. In this specification, the LLM fine-tuned based on the business scenario is referred to as the business model.

[0046] Obviously, there are certain differences between the general LLM and the business model, and these differences are not limited to differences in the activation states of neurons for the same input. Therefore, in order to improve the accuracy of the trained detection classifier, the server can also first train a screening classifier for the general LLM, which is used to filter the training samples in step S100 from the general samples used to train the general LLM.

[0047] Specifically, the server can first obtain universal samples for training the LLM. These universal samples include annotations indicating whether the model output is correct. Of course, the LLM used to train the universal samples here is not the LLM to be tested, but rather a universal LLM. Furthermore, when the detection classifier described in step S104 is applied to other LLMs, it can be applied to other LLMs with the same structure as the LLM to be tested. Based on the same principle, when selecting universal samples, universal samples can also be selected based on the model structure of the LLM to be tested to train a universal LLM with the same model structure as the LLM to be tested. Model structure consistency here can, on the one hand, mean consistent model architecture, such as consistent number of neurons and neuron connectivity, and on the other hand, mean consistent neuron types. Of course, the latter requirement for consistency is more stringent and can be used selectively. For example, universal samples of universal LLMs with consistent model architectures can be selected first. If the number of universal samples is sufficient, universal samples of universal LLMs with consistent neuron types can be further selected. The universal samples of universal LLMs referred to in this specification refer to universal samples used to train universal LLMs.

[0048] Secondly, the activation state of neurons in the preset universal large model based on the input universal sample is determined as the universal state corresponding to the universal sample; Then, the general state corresponding to the general sample is used as input, and the label of whether the output result is correct is output to train the screening classifier.

[0049] The above two steps are similar to the process of inputting training samples into the LLM to be detected, determining the activation state, and training the detection classifier in step S104. However, the input data is changed to universal samples, activating neurons in the universal LLM, and recording the activation state of neurons in the universal LLM (referred to as the universal state corresponding to the universal sample). The training is based on a screening classifier for the universal LLM. For relevant details, please refer to the corresponding descriptions of steps S102 and S104, and this specification will not elaborate on them here.

[0050] Finally, the server may screen out training samples for training the detection classifier from the general samples based on the second activation states of neurons corresponding to the general samples in the large model to be detected and the screening classifier.

[0051] By using the performance of common samples on the LLM to be tested, the screening classifier is used to filter out common samples that still produce hallucinations from the common samples and use them as training samples. This way, even if the business scenario corresponding to the LLM to be tested does not have training samples pre-labeled as having produced hallucinations, common samples from the common scenario and the screening classifier can be used to obtain samples that still produce hallucinations in the LLM to be tested and use them as training samples.

[0052] Specifically, to select training samples from the general sample, the server first determines the activation state of neurons in the large model to be tested based on the input general sample, and uses this as the specific state corresponding to the general sample. Specifically, the general sample is input into the LLM to be tested, and the activation state of neurons in the LLM to be tested when inferring the general sample is recorded as the specific state of the general sample.

[0053] In the second step, the server determines the correctness of the output results of the large model to be tested by screening the classifier according to the specific state corresponding to the general sample, and determines the training sample at least based on the correctness of the output results.

[0054] In an embodiment of the present specification, the server takes a specific state as input and inputs it into a screening classifier to obtain a classification result of whether the output result of the LLM to be detected is correct. That is, the screening classifier determines whether the LLM to be detected produces a "hallucination". If it is determined that the output result is wrong, that is, a "hallucination" is produced, it is determined that the general sample can be used as a training sample for training the detection classifier. Due to the differences in business scenarios and model parameters, directly using general samples to train the detection classifier may cause inappropriate general samples to contaminate the detection classifier. Through the above-mentioned screening process, the training samples obtained can be determined to be samples that cause "hallucinations" on the LLM to be detected, which can avoid the pollution problem caused by inappropriate samples.

[0055] Furthermore, in the embodiments of this specification, the server can use only generic samples that produce hallucinations as training samples. Specifically, based on the correctness of the output results, the server can identify generic samples with incorrect output results, label them as incorrect, and update the label of whether the output results of the included model are correct, thereby obtaining training samples. In other words, any generic sample, regardless of whether it causes hallucinations in the generic LLM, can be used as a training sample to train the detection classifier as long as it causes hallucinations in the LLM to be detected.

[0056] Furthermore, in the embodiments of this specification, when determining training samples based on the correctness of the output results, the server may also determine, from the common samples, samples whose annotated output results and the output results obtained by the screening classifier are consistent, as training samples. In other words, the training samples determined are those that induce "hallucinations" in both the common LLM and the LLM to be tested. This ensures the stability of the training samples and ensures that the resulting training samples are sufficient to induce hallucinations in the LLM.

[0057] Of course, if the above method is used to screen and obtain the training samples, then the step S102 of determining the activation state of the training samples can be replaced by the above process.

[0058] Figure 3 This is a schematic diagram of the large model output result detection process provided in the embodiment of this specification. Figure 3 For combination Figure 1 The steps and the process of screening the general samples to obtain the training samples are described above. The contents of each step have been described above and will not be repeated here.

[0059] In the embodiment of this specification, the server may also send the common samples with erroneous output results to the user terminal according to the correctness of the output results through a human-computer interaction process, and determine the training samples in response to at least the annotations returned by the user terminal.

[0060] Specifically, a common sample that causes hallucinations on the LLM to be tested is sent to the user terminal, prompting the user to label whether the common sample causes hallucinations. The user terminal then returns the labeling of the common sample, which is manual labeling, and uses the common sample labeled as an error in the output as a training sample.

[0061] Because the above screening process significantly reduces the number of common samples that the server can send to the user terminal, manual labeling can focus on only those common samples that the screening classifier determines are incorrect. Hallucination data typically accounts for a small proportion, significantly improving manual labeling efficiency.

[0062] The above is a large model output result detection method provided in the embodiment of this specification. Based on the same idea, this specification also provides corresponding devices, storage media and electronic devices.

[0063] Figure 4 A schematic diagram of a large model output result detection device provided in an embodiment of this specification, the device comprising: An acquisition module 201 is used to acquire training samples for training a large model, wherein the training samples include a label indicating whether the model output result is correct; An inference module 202 is configured to determine, based on the input of the training sample, an activation state of a neuron in the large model to be tested as the activation state corresponding to the training sample; A training module 203 is configured to take the activation state corresponding to the training sample as input, mark whether the output result is correct, and train a detection classifier; The detection module 204 is used to detect the correctness of the output result of the large model to be detected through the detection classifier in response to the activation state of the neurons in the business process executed by the large model to be detected.

[0064] Optionally, the detection module 204 is used to determine the activation state of neurons in the large model to be detected as detection data in response to the input data of the large model to be detected; input the detection data into a trained detection classifier, and detect the correctness of the output result of the large model to be detected based on the output of the detection classifier.

[0065] Optionally, the acquisition module 201 is used to obtain a general sample for training a large model, and the general sample includes an annotation of whether the model output result is correct; determine the activation state of neurons in a preset general large model based on the input of the general sample as the general state corresponding to the general sample; train a screening classifier with the general state corresponding to the general sample as input and the annotation of whether the output result is correct; based on the activation state of neurons corresponding to the general sample in the large model to be detected and the screening classifier, filter out training samples for training a detection classifier from the general sample.

[0066] Optionally, the acquisition module 201 is used to determine the activation state of neurons in the large model to be tested based on the input of the general sample as the specific state corresponding to the general sample; according to the specific state, the correctness of the output result of the large model to be tested is determined through the screening classifier, and the training sample is determined at least based on the correctness of the output result.

[0067] Optionally, the acquisition module 201 is configured to update the label of whether the output result of the general sample is correct according to the correctness of the output result, so as to obtain a training sample.

[0068] Optionally, the acquisition module 201 is configured to determine, from the common samples, samples having the same output results as those obtained by the screening classifier and the labeled output results, as training samples.

[0069] Optionally, the acquisition module 201 is configured to send the general sample with an erroneous output result to a user terminal according to the correctness of the output result; and determine the training sample at least in response to the annotation returned by the user terminal.

[0070] Optionally, the large model to be detected is a business model; The device further comprises: The execution module 205 is configured to execute the business corresponding to the business model based on the output result when it is determined that the output result of the business model is correct according to the output of the detection classifier.

[0071] This specification also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it can be used to execute the large model output result detection method provided above.

[0072] based on Figure 1 The large model output result detection method shown in the embodiment of this specification also provides Figure 5 The structural diagram of the electronic device shown in FIG. Figure 5At the hardware level, the electronic device includes a processor, an internal bus, a network interface, memory, and non-volatile storage, and may also include other hardware required for its operations. The processor reads the corresponding computer program from the non-volatile storage into the memory and then runs it to implement the aforementioned large model output result detection method.

[0073] The foregoing is merely an example of the present invention and is not intended to limit the present invention. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be included within the scope of the claims of the present invention.

Claims

1. A method for detecting a large model output result, the method comprising: Obtaining training samples for training the large model, wherein the training samples include annotations indicating whether the model output results are correct; Determine the activation state of a neuron in the large model to be detected based on the input of the training sample as the activation state corresponding to the training sample; Taking the activation state corresponding to the training sample as input and labeling whether the output result is correct, a detection classifier is trained; In response to the activation state of the neurons of the large model to be detected executing the business process, the correctness of the output result of the large model to be detected is detected by the detection classifier.

2. The method according to claim 1, wherein, in response to the activation state of the neurons in the business process executed by the large model to be tested, the correctness of the output result of the large model to be tested is detected by the detection classifier, specifically comprising: In response to the input data of the large model to be detected, determining the activation state of neurons in the large model to be detected as detection data; The detection data is input into the trained detection classifier, and the correctness of the output result of the large model to be detected is detected based on the output of the detection classifier.

3. The method according to claim 1, wherein obtaining training samples for training a large model comprises: Obtaining a general sample for training a large model, wherein the general sample includes a label indicating whether the model output result is correct; Determine, based on the input of the universal sample, an activation state of neurons in a preset universal large model as a universal state corresponding to the universal sample; Taking the general state corresponding to the general sample as input and marking whether the output result is correct, a screening classifier is trained; According to the activation state of the neurons corresponding to the general samples in the large model to be detected and the screening classifier, training samples for training the detection classifier are screened from the general samples.

4. The method of claim 3, wherein the method further comprises: selecting a training sample for training a detection classifier from the general sample based on the activation state of the neuron corresponding to the general sample in the large model to be detected and the screening classifier; Determining, based on the input of the universal sample, an activation state of a neuron in the large model to be detected as a specific state corresponding to the universal sample; According to the specific state, the correctness of the output result of the large model to be tested is determined through the screening classifier, and the training sample is determined at least based on the correctness of the output result.

5. The method according to claim 4, wherein determining the training sample based on the correctness of the output result specifically comprises: According to the correctness of the output result, the label of whether the output result of the general sample is correct is updated to obtain a training sample.

6. The method according to claim 4, wherein determining the training sample based on the correctness of the output result specifically comprises: From the common samples, samples having the same output results as those obtained by the screening classifier and the marked output results are determined as training samples.

7. The method according to claim 4, wherein determining the training sample based on the correctness of the output result specifically comprises: According to the correctness of the output result, the universal sample whose output result is an error is sent to the user terminal; A training sample is determined at least in response to the annotation returned by the user terminal.

8. The method according to claim 2, wherein the large model to be detected is a business model; The method further comprises: When it is determined that the output result of the business model is correct according to the output of the detection classifier, the business corresponding to the business model is executed based on the output result.

9. A large model output result detection device, the device comprising: An acquisition module is used to obtain training samples for training the large model, wherein the training samples include a label indicating whether the model output result is correct; An inference module, configured to determine, based on the input of the training sample, an activation state of a neuron in the large model to be tested as the activation state corresponding to the training sample; A training module is used to take the activation state corresponding to the training sample as input, mark whether the output result is correct, and train the detection classifier; The detection module is used to detect the correctness of the output result of the large model to be detected through the detection classifier in response to the activation state of the neurons in the business process executed by the large model to be detected.

10. A computer-readable storage medium storing a computer program, wherein the computer program implements the method according to any one of claims 1 to 8 when executed by a processor.

11. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 8 when executing the program.

Citation Information

Patent Citations

  • Target crowd mining method, device, server, and readable storage medium

    CN109087145A

  • Method and device for processing capital reservation intention of user, service equipment and storage medium

    CN116956223A

  • Model illusion detection method and device, equipment, storage medium and program product

    CN118887474A

  • Enterprise expert talent prediction method and device based on big data

    CN119338034A

  • Hidden layer activation-based prejudice illusion detection method

    CN119829962A