Information detection method and system and electronic equipment

By monitoring and analyzing the internal state information of the dialogue model, obtaining the target text characteristics, and constructing diagnosing degree measurement indicators, the accuracy of reliability detection of large-scale model output reply information is solved, and the reliability detection of the information output of dialogue model is realized.

CN120372341APending Publication Date: 2025-07-25HANGZHOU ALICLOUD FEITIAN INFORMATION TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202410099131.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-23
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The prior art cannot effectively detect the reliability of reply information output by large models, especially due to inaccurate detection results caused by uncertainty in sentence length and diversity of expression forms.

Method used

By monitoring the input information, the target text features of the reply information are obtained using the internal state information of the dialogue model, and the reliability of the reply information is determined based on these features, including the call-up of the dialogue model matching the scene task in the question-and-answer system for analysis, and the construction of a covariance matrix based on the target text features is carried out to calculate the illusion degree metric.

Benefits of technology

The reliability detection of the large model output reply information is realized, the accuracy of the detection is improved, and the problem of inaccurate detection results in the prior art is solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372341A_ABST
    Figure CN120372341A_ABST
Patent Text Reader

Abstract

The invention discloses an information detection method and system and electronic equipment. The method comprises the following steps: monitoring input information; inputting the input information into a dialogue model for analysis to obtain at least one piece of reply information matched with the input information; target text features of the reply information are obtained from internal state information of the dialogue model, the internal state information is used for representing a rule of analyzing the input information by the dialogue model, and the target text features are used for representing semantics of the reply information; and based on the target text feature of the reply information, determining a detection result of the reply information, the detection result being used for representing whether the reply information output by the dialogue model is reliable. The technical problem that reliability detection cannot be effectively carried out on reply information output by a large model is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical fields of large model technology and large language model detection technology. Specifically, it relates to a method, system, and electronic device for detecting information. Background Art

[0002] Currently, with the rapid development of large models, the requirements for the authenticity and reliability of the reply information output by large models are also constantly increasing. How to detect whether the reply information output by large models is reliable has become particularly important.

[0003] In related technologies, an illusion detection method based on perplexity or an illusion detection method based on self-check (SelfCheckGPT) is usually used to determine whether the reply information output by a large model is reliable. Among them, the illusion detection method based on perplexity mainly determines the uncertainty of each word in the reply information output by the large model word by word, and then multiplies the uncertainties of each word to obtain the uncertainty of the reply information output by the large model. However, due to inaccurate estimation of the uncertainty of the sentence length of the reply information output by the large model and the diversity of the expression forms of the sentences in the reply information output by the large model, the detection result of the reliability of the reply information output by the large model may be inaccurate.

[0004] Therefore, the above method has the technical problem that it cannot effectively detect the reply information output by the large model.

[0005] In response to the above problems, no effective solution has been proposed yet. Summary of the Invention

[0006] Embodiments of this application provide a method, system, and electronic device for detecting information to at least solve the technical problem of being unable to effectively detect the reliability of the reply information output by a large model.

[0007] According to one aspect of the embodiments of this application, a method for detecting information is provided. The method may include: monitoring input information; inputting the input information into a dialogue model for analysis to obtain at least one reply information that matches the input information; obtaining, from the internal state information of the dialogue model, the target text feature of the reply information, where the internal state information is used to represent the rules for the dialogue model to analyze the input information, and the target text feature is used to represent the semantics of the reply information; and determining the detection result of the reply information based on the target text feature of the reply information, where the detection result is used to represent whether the reply information output by the dialogue model is reliable.

[0008] According to another aspect of the embodiments of the present application, there is also provided a method for generating information, which is applied to a question-and-answer system deployed in a scenario task. The method may include: monitoring input information in the scenario task on the operation interface of the question-and-answer system; retrieving a dialogue model matching the scenario task, inputting the input information into the dialogue model for analysis, and obtaining at least one reply information matching the input information; obtaining, from the internal state information of the dialogue model, the target text feature of the reply information, where the internal state information is used to represent the rule for the dialogue model to analyze the input information, and the target text feature is used to represent the semantics of the reply information in the scenario task; determining the detection result of the reply information based on the target text feature of the reply information; outputting the reply information in response to the detection result indicating that the reply information output by the dialogue model is reliable in the scenario task; and outputting the corresponding prompt information in response to the detection result indicating that the reply information output by the dialogue model is unreliable in the scenario task.

[0009] According to another aspect of the embodiments of the present application, there is also provided a method for detecting information. The method may include: monitoring input information on the dialogue interface; displaying, on the dialogue interface, at least one reply information matching the input information, where the reply information is obtained by analyzing the input information using a dialogue model; in response to an information detection operation on the dialogue interface, displaying the detection result of the reply information on the dialogue interface, where the detection result is used to indicate whether the reply information output by the dialogue model is reliable, and is determined based on the target text feature of the reply information, the target text feature is used to represent the semantics of the reply information, and is obtained from the internal state information of the dialogue model, and the internal state information is used to represent the rule for the dialogue model to analyze the input information.

[0010] According to one aspect of the embodiments of the present application, there is provided a device for detecting information. The device may include: a first monitoring unit, configured to monitor input information; a first analysis unit, configured to input the input information into a dialogue model for analysis, and obtain at least one reply information matching the input information; a first obtaining unit, configured to obtain, from the internal state information of the dialogue model, the target text feature of the reply information, where the internal state information is used to represent the rule for the dialogue model to analyze the input information, and the target text feature is used to represent the semantics of the reply information; and a first determining unit, configured to determine the detection result of the reply information based on the target text feature of the reply information, where the detection result is used to indicate whether the reply information output by the dialogue model is reliable.

[0011] According to another aspect of the embodiments of the present application, there is also provided an information generation device, which may include: a second monitoring unit, configured to monitor input information in a scenario task on an operation interface of a question-and-answer system; a second analysis unit, configured to retrieve a dialogue model matching the scenario task, input the input information into the dialogue model for analysis, and obtain at least one reply information matching the input information; a second acquisition unit, configured to obtain target text features of the reply information from internal state information of the dialogue model, where the internal state information is used to represent the rules for the dialogue model to analyze the input information, and the target text features are used to represent the semantics of the reply information in the scenario task; a second determination unit, configured to determine a detection result of the reply information based on the target text features of the reply information; a first output unit, configured to output the reply information in response to the detection result indicating that the reply information output by the dialogue model is reliable in the scenario task; and a second output unit, configured to output corresponding prompt information in response to the detection result indicating that the reply information output by the dialogue model is unreliable in the scenario task.

[0012] According to another aspect of the embodiments of the present application, there is also provided an information detection device, which may include: a third monitoring unit, configured to monitor input information on a dialogue interface; a first display unit, configured to display at least one reply information matching the input information on the dialogue interface, where the reply information is obtained by analyzing the input information using the internal state information of the dialogue model; and a second display unit, configured to display a detection result of the reply information on the dialogue interface in response to an information detection operation on the dialogue interface, where the detection result is used to indicate whether the reply information output by the dialogue model is reliable, and is determined based on the target text features of the reply information, the target text features are used to represent the semantics of the reply information, and are obtained from the internal state information of the dialogue model, and the internal state information is used to represent the rules for the dialogue model to analyze the input information.

[0013] According to another aspect of the embodiments of the present application, there is also provided an information detection system, which may include: an information input end, configured to monitor input information; an information detection end, configured to input the input information into a dialogue model for analysis, and obtain at least one reply information matching the input information; obtain target text features of the reply information from the internal state information of the dialogue model, where the internal state information is used to represent the rules for the dialogue model to analyze the input information, and the target text features are used to represent the semantics of the reply information; determine a detection result of the reply information based on the target text features of the reply information, where the detection result is used to indicate whether the reply information output by the dialogue model is reliable; an information output end, configured to output the reply information in response to the detection result indicating that the reply information output by the dialogue model is reliable; and output corresponding prompt information in response to the detection result indicating that the reply information output by the dialogue model is unreliable.

[0014] According to another aspect of the embodiments of the present application, an electronic device is further provided, including a memory and a processor. The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, and the computer-executable instructions are executed by the processor to perform the steps of the information detection method.

[0015] According to another aspect of the embodiments of the present application, a computer-readable storage medium is further provided. The computer-readable storage medium includes a stored program, wherein when the program processor runs, it controls the device where the computer-readable storage medium is located to perform the steps of the information detection method.

[0016] In the embodiments of the present application, input information is monitored; the input information is input into a dialogue model for analysis to obtain at least one reply information that matches the input information; from the internal state information of the dialogue model, the target text feature of the reply information is obtained, where the internal state information is used to represent the rules for the dialogue model to analyze the input information, and the target text feature is used to represent the semantics of the reply information; based on the target text feature of the reply information, the detection result of the reply information is determined, where the detection result is used to represent whether the reply information output by the dialogue model is reliable. That is to say, in the present application, based on the analysis of the input information by the dialogue model, at least one reply information that matches the input information can be obtained. Since the internal state information of the dialogue model is used to represent the rules for the dialogue model to analyze the input information, the target text feature of the reply information can be obtained from the internal state information of the dialogue model, thereby better mining and utilizing the semantic features of the reply information. The detection result is determined using the target text feature, and this detection result can characterize the overall uncertainty degree of the reply information, that is, it reflects whether the reply information output by the dialogue model is reliable, achieving the technical effect of effectively detecting the reliability of the reply information output by the dialogue model, and thus solving the technical problem of being unable to effectively detect the reliability of the reply information output by the large model.

[0017] It is easy to note that the above general description and the following detailed description are only for exemplifying and explaining the present application, and do not constitute a limitation to the present application. Description of the Drawings

[0018] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:

[0019] Figure 1 is a schematic diagram of an application scenario of an information detection method according to an embodiment of the present application;

[0020] Figure 2 is a flowchart of an information detection method according to an embodiment of the present application;

[0021] Figure 3 is a flowchart of another method for generating information according to an embodiment of the present application;

[0022] Figure 4 is a flowchart of another method for detecting information according to an embodiment of the present application;

[0023] Figure 5 is a schematic diagram of a system for detecting an information according to an embodiment of the present application;

[0024] Figure 6 is a flowchart of another method for detecting information according to an embodiment of the present application;

[0025] Figure 7 is a schematic diagram of a feature response of a text feature set according to an embodiment of the present application;

[0026] Figure 8 is a schematic diagram of detecting knowledge hallucination of a large model according to an embodiment of the present application;

[0027] Figure 9 is a schematic diagram of a device for detecting an information according to an embodiment of the present application;

[0028] Figure 10 is a schematic diagram of another device for generating information according to an embodiment of the present application;

[0029] Figure 11 is a schematic diagram of another device for detecting information according to an embodiment of the present application;

[0030] Figure 12 is a structural block diagram of a computer terminal according to an embodiment of the present application. Detailed implementation manners

[0031] In order to enable those skilled in the art of the present technology to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0032] It should be noted that the terms "first", "second", etc. in the description, claims and the above-mentioned drawings of this application are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products or devices.

[0033] The technical solution provided by this application is mainly implemented using large model technology. Here, a large model refers to a deep learning model with a large number of model parameters, usually including hundreds of millions, tens of billions, hundreds of billions, trillions or even more than one quadrillion model parameters. A large model can also be called a foundation model. Through large-scale pre-training of the large model with unlabeled corpora, a pre-trained model with more than one billion parameters is produced. This model can adapt to a wide range of downstream tasks and has good generalization ability. For example, large language models (LLMs), multi-modal pre-training models, etc.

[0034] It should be noted that when a large model is actually applied, the pre-trained model can be fine-tuned with a small number of samples so that the large model can be applied to different tasks. For example, large models can be widely applied in fields such as natural language processing (NLP), computer vision, and speech processing. Specifically, they can be applied to tasks in the field of computer vision such as visual question answering (VQA), image captioning (IC), and image generation. They can also be widely applied to tasks in the field of natural language processing such as text-based sentiment classification, text summary generation, and machine translation. Therefore, the main application scenarios of large models include but are not limited to digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc. In the embodiments of this application, the data processing by a large language model in the knowledge hallucination detection scenario is used as an example for explanation.

[0035] First, some nouns or terms that appear in the process of describing the embodiments of this application are applicable to the following explanations:

[0036] Large language models, also known as large-scale language models, are a type of language model represented by the Transformer model with a very large number of parameters (e.g., above 1B) and good natural language generation capabilities;

[0037] Knowledge hallucination refers to the situation where, when the large model outputs content that "does not conform to facts" or "creates something out of nothing", it is considered that the output of the large model has knowledge hallucination;

[0038] The internal state information (Internal States) of the large model refers to the information related to the weights and features inside the large model;

[0039] Feature Clipping (abbreviated as FC) refers to clipping abnormal features in the hidden layer of the model;

[0040] Token embedding means that different words correspond to a feature representation in the output layer;

[0041] Perplexity is a measure of output confidence calculated based on the logit of the large model output.

[0042] Embodiment 1

[0043] According to an embodiment of the present application, a method for detecting information is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0044] Considering that the large model has a huge number of model parameters and the computing resources of mobile terminals are limited, the above-mentioned method for detecting information provided by the embodiments of the present application can be applied to Figure 1 the application scenarios shown, but not limited thereto. Figure 1 is a schematic diagram of the application scenario of a method for detecting information according to an embodiment of the present application. In Figure 1 the application scenario shown, the large model is deployed in the server 10. The server 10 can be connected to one or more client devices 20 through a local area network connection, a wide area network connection, an Internet connection, or other types of data networks. Here, the client devices 20 can include but are not limited to: smart phones, tablet computers, laptop computers, palmtop computers, personal computers, smart home devices, vehicle-mounted devices, etc. The client device 20 can interact with the user through a graphical user interface to call the large model, thereby implementing the method provided by the embodiments of the present application.

[0045] In the embodiments of the present application, the system composed of a client device and a server may perform the following steps: The client device executes input information, which may be of types such as inquiry information, query information, etc. The server executes step S101 to monitor the input information; step S102 to input the input information into a dialogue model for analysis to obtain at least one reply information that matches the input information; step S103 to obtain the target text features of the reply information from the internal state information of the dialogue model; step S104 to determine the detection result of the reply information based on the target text features of the reply information, where the detection result is used to indicate whether the reply information output by the dialogue model is reliable. It should be noted that in the case where the operating resources of the client device can meet the deployment and operating conditions of the large model, the embodiments of the present application can be carried out in the client device, and the comparison model adopted by the present application can be a generative dialogue model.

[0046] Under the above operating environment, the present application provides a method for detecting information as follows Figure 2 shown. Figure 2 is a flowchart of a method for detecting information according to an embodiment of the present application. As Figure 2 shown, the method may include the following steps:

[0047] Step S201, monitor the input information.

[0048] In the technical solution provided in step S201 of the present application above, the input information may be multimodal information, and the types of multimodal information include at least one of the following: text information containing character information, video frame information containing frame image information, audio information. For example, inquiry information or chatting information, etc., and no specific limitation is made here. The question-and-answer system can monitor the input information.

[0049] For example, taking the input information as text information, the user can input multiple text questions in the question-and-answer system, and the multiple text questions are the input information. The question-and-answer system can monitor the multiple questions in real time and then generate corresponding reply information according to the multiple questions.

[0050] It should be noted that when the input information is non-text information, for example, when the input information is video information or audio information, etc., the video information or audio information can be converted into text information for processing.

[0051] Step S202, input the input information into a dialogue model for analysis to obtain at least one reply information that matches the input information.

[0052] In the technical solution provided in step S202 of the present application, the dialogue model can be a large model capable of understanding and generating natural language texts. For example, a generative dialogue model. For example, the dialogue model is a model pre-trained using natural language processing techniques and machine learning algorithms for generating dialogue content, answering questions, providing suggestions, etc. Based on this, after the input information is monitored according to step S201, the input information can be input into the dialogue model for analysis, and then at least one reply information matching the input information can be obtained. Among them, the reply information can be multimodal information, and the types of the reply information can include at least one of the following: text information, image information, video information, and voice information.

[0053] In this embodiment, the dialogue model is used to analyze the input information of the user to generate at least one reply information matching the input information. Based on this, after the question-and-answer system monitors the input information, the input information can be directly input into the dialogue model to analyze the input information with the help of the dialogue model, and then at least one reply information matching the input information can be generated.

[0054] Step S203, obtain the target text feature of the reply information from the internal state information of the dialogue model.

[0055] In the technical solution provided in step S203 of the present application, after at least one reply information matching the input information is obtained according to step S202, the target text feature of the reply information can be obtained from the internal state information of the dialogue model. Among them, the internal state information is used to represent the rules for the dialogue model to analyze the input information, and the target text feature is used to represent the semantics of the reply information. For example, the target text feature can be the feature of the sentence corresponding to the reply information.

[0056] In this embodiment, the internal state information of the dialogue model can include the parameters, features, weights, learning rules, etc. of the model. The internal state information of the dialogue model can affect the processing method and processing result of the dialogue model for the input information. When processing the input information, the dialogue model can analyze and reason about the input information according to the internal state information, so as to generate corresponding output results.

[0057] For example, the internal state information of the dialogue model can include the rules for multiple network layers inside the dialogue model to analyze the input information. Among them, the multiple network layers of the dialogue model can include: a decoder layer (Decoder), a fully connected layer (FC Layer), etc. The multiple network layers are used to perform different processing operations on the input information, and then obtain the output result. The target text feature of the reply information can be obtained from the output results of the multiple network layers.

[0058] For example, after multiple network layers in a dialogue model process the input information, reply information matching the input information can be obtained. Among them, the reply information can include multiple characters, and each of the multiple characters corresponds to a feature vector in different network layers of the dialogue model. Since in a large model, the feature of the last character in the generated reply information can often represent the feature of the entire reply information, based on this, according to the arrangement order of the multiple characters in the reply information, the feature vector corresponding to the last character in the reply information in the middle layer of the multiple network layers of the dialogue model can be determined, and then the feature vector of the middle layer corresponding to the last character among the multiple characters of the reply information is determined as the target text feature of the reply information. Among them, taking the Large Language Model Meta-7B (abbreviated as llama-7B) as an example, assuming that the llama-7B model contains 33 network layers, after using the llama-7B model to process the input information, the feature of the last character output by the 17th network layer (middle layer) can be selected as the target text feature of the reply information output by the llama-7B model.

[0059] Step S204, determine the detection result of the reply information based on the target text feature of the reply information.

[0060] In the technical solution provided in step S204 of the present application, after obtaining the target text feature of the reply information according to step S203, the detection result of the reply information can be determined based on the target text feature of the reply information, where the detection result is used to indicate whether the reply information output by the dialogue model is reliable.

[0061] In this embodiment, after obtaining the target text feature of the reply information, a covariance matrix can be constructed according to the target text feature of the reply information, and then a metric index of the reply information output by the dialogue model can be calculated according to the covariance matrix, where the metric index can be a hallucination degree metric score corresponding to the reply information output by the dialogue model. After obtaining the metric index, it can be further determined whether the reply information output by the dialogue model is reliable according to the metric index and the metric index threshold.

[0062] For example, assuming that the metric index is the hallucination degree metric score corresponding to the reply information output by the dialogue model, the metric index threshold can be a score threshold. In this case, the hallucination degree metric score can be compared with the score threshold. When the hallucination degree metric score is greater than the score threshold, it indicates that there is a hallucination in the output of the dialogue model, that is, the reply information output by the dialogue model is unreliable. When the hallucination degree metric score is not greater than the score threshold, it indicates that there is no hallucination in the output of the dialogue model, that is, the reply information output by the dialogue model is reliable.

[0063] Based on steps S201 to S204 of the above embodiment, by analyzing the input information based on the dialogue model, at least one reply information matching the input information can be obtained. Since the internal state information of the dialogue model is used to represent the rules for the dialogue model to analyze the input information, the target text features of the reply information can be obtained from the internal state information of the dialogue model, thereby better mining and utilizing the semantic features of the reply information. Using the target text features to determine the detection result of the reply information, this detection result can characterize the overall uncertainty degree of the reply information, that is, it reflects whether the reply information output by the dialogue model is reliable, achieving the technical effect of effectively detecting the reliability of the reply information output by the dialogue model, and thus solving the technical problem of being unable to effectively detect the reply information output by the model.

[0064] The above method of this embodiment will be further introduced below.

[0065] As an optional implementation manner, in step S203, to obtain the target text features corresponding to the reply information from the internal state information of the dialogue model, it includes: obtaining the text feature set of the reply information in the network layer of the dialogue model, where the internal state information includes the text feature set in the network layer, and the text feature set in the network layer includes the text features of the text units constituting the reply information; determining the target text features from the text feature set in the network layer.

[0066] In this embodiment, the dialogue model may include multiple network layers. The multiple network layers can respectively output the text features of each text unit in the reply information, and the text features of each text unit can constitute the text feature set of the reply information in the network layer of the dialogue model. Based on this, the text feature set of the reply information in the network layer of the dialogue model can be obtained.

[0067] For example, the text feature set of the reply information in the network layer of the dialogue model can be represented as Tokenembedding, where the text unit of the text feature included in this text feature set can be represented as token. By obtaining the text features of each text unit in the reply information output by each network layer of the dialogue model, and then combining the text features of each text unit, the text feature set of the reply information in each network layer of the dialogue model can be obtained.

[0068] In this embodiment, directly extracting the text feature set from the reply information output by the network layer of the dialogue model without using an additional model to extract the text features of the reply information output by the dialogue model saves computational overhead and improves computational efficiency.

[0069] Optionally, after determining the text feature set of the reply information in each network layer of the dialogue model, the target text features can be determined from the text feature set in the network layer.

[0070] As an alternative implementation, determining the target text feature from the text feature set in the network layer includes: determining the text feature corresponding to the target text unit in the reply information from the text feature set in the network layer, where the target text unit includes the semantics of the reply information; and determining the text feature corresponding to the target text unit as the target text feature.

[0071] In this embodiment, since the text feature set includes the text features of each text unit constituting the reply information, and the target text unit includes the semantics of the reply information, based on this, the text feature corresponding to the target text unit in the reply information can be determined from the text feature set in the network layer, and then the text feature corresponding to the target text unit is determined as the target text feature.

[0072] For example, since in a large model, the text feature of the last character of a sentence often contains the semantic information of the whole sentence, based on this, the last text unit in the reply information can be used as the target text unit of the reply information, that is, the last character in the sentence corresponding to the reply information is used as the target text unit of the reply information. The text feature corresponding to the last text unit can be determined from the text feature set in the network layer according to the arrangement order of each text unit in the reply information, and then the text feature corresponding to the last text unit is determined as the text feature corresponding to the target text unit in the reply information.

[0073] As an alternative implementation, obtaining the text feature set of the reply information in the network layer of the dialogue model includes: obtaining the text feature set of the reply information in the output network layer of the dialogue model; and determining the text feature set of the target hidden network layer corresponding to the text feature set of the output network layer in the target hidden network layer of the dialogue model.

[0074] In this embodiment, the network layer of the dialogue model may include an output network layer and a target hidden network layer, etc. Among them, the output network layer may be the penultimate layer in the network layer of the dialogue model, and the target hidden network layer may be the middle layer in the network layer of the dialogue model.

[0075] For example, from the foregoing introduction, it can be known that the text feature set is composed of the text features of each text unit constituting the reply information. Based on this, the text features of each text unit output by the output network layer of the dialogue model can be obtained, and then the text features of each text unit are combined to form the text feature set of the output network layer of the dialogue model, where this text feature set can be called token embedding.

[0076] Optionally, since the target hidden network layer of the dialogue model can be the intermediate layer of multiple network layers of the dialogue model, based on this, in the target hidden network layer of the dialogue model, the text feature set of the target hidden network layer corresponding to the text feature set of the output network layer can be determined. Next, the process of determining the text feature set of the target hidden network layer corresponding to the text feature set of the output network layer in the target hidden network layer of the dialogue model will be further introduced.

[0077] As an alternative implementation, determining the text feature set of the target hidden network layer corresponding to the text feature set of the output network layer in the target hidden network layer of the dialogue model includes: updating the text feature set of the output network layer to obtain an updated text feature set, where the updated text feature set does not include noise features; in the target hidden network layer, determining the text feature set of the target hidden network layer corresponding to the updated text feature set.

[0078] In this embodiment, there may be a large number of extreme feature responses in the text features included in the text feature set of the output network layer, and such extreme feature responses can be considered noise features. For example, some text features with too large or too small response amplitudes included in the text feature set are of this type, and this type of text feature is likely to cause the dialogue model to output incorrect classification results. Based on this, after obtaining the text feature set of the output network layer of the dialogue model, the text feature set can be updated to obtain an updated text feature set.

[0079] For example, the text feature set can be updated by adjusting the response amplitudes of the text features included in the text feature set of the output network layer. For example, the text features in the text feature set can be adjusted within the response amplitude threshold range according to the response amplitude threshold range corresponding to the text feature set of the output network layer to update the text feature set in the output network layer.

[0080] Optionally, the text feature set can be updated by performing a feature clipping operation on the text feature set of the output network layer. For example, the text features with too large or too small response amplitudes in the text feature set are clipped to obtain a clipped text feature set, and the clipped text feature set is determined as the updated text feature set to remove the noise features in the text feature set. Among them, performing a clipping operation on the text feature set is only an exemplary example of updating the text feature set, and the method for updating the text feature set is not limited here.

[0081] In this embodiment, performing an update operation on the text feature set of the output network layer can remove the noise features in the text feature set, which helps to improve the output quality of the dialogue model, and further avoid the hallucination output with a high consistency caused by the oversaturated classification results output by the dialogue model. Next, the process of updating the text feature set of the output network layer will be further introduced.

[0082] As an alternative implementation, updating the text feature set of the output network layer to obtain an updated text feature set includes: in response to the response amplitude of the text features in the text feature set of the output network layer being less than the first response amplitude threshold, adjusting the response amplitude to the first response amplitude threshold; in response to the response amplitude of the text features in the text feature set of the output network layer being greater than or equal to the first response amplitude threshold and less than or equal to the second response amplitude threshold, keeping the response amplitude, where the second response amplitude threshold is greater than or equal to the first response amplitude threshold; in response to the response amplitude of the text features in the text feature set of the output network layer being greater than the second response amplitude threshold, adjusting the response amplitude to the second response amplitude threshold; and determining the response amplitudes adjusted to the first response amplitude threshold, the kept response amplitudes, and the response amplitudes adjusted to the second response amplitude threshold as the updated text feature set.

[0083] In this embodiment, the first response amplitude threshold and the second response amplitude threshold are used to measure the response amplitude of the text features in the text feature set. For example, the first response amplitude threshold may represent the minimum response amplitude of the text feature, and the second response amplitude threshold may represent the maximum response amplitude of the text feature. Based on this, the response amplitudes of the respective text features included in the text feature set of the output network layer are compared with the first response amplitude threshold and the second response amplitude threshold respectively, and then the response amplitude to be adjusted in the text feature set is determined. If there is a text feature in the text feature set of the output network layer whose response amplitude is less than the first response amplitude threshold, it indicates that the response amplitude of this text feature is too low. In this case, the response amplitude of the text feature in the text feature set can be adjusted to the first response amplitude threshold; similarly, if in response to the fact that there is a text feature in the text feature set of the output network layer whose response amplitude is greater than the second response amplitude threshold, it indicates that the response amplitude of this text feature is too high. In this case, the response amplitude of the text feature in the text feature set can be adjusted to the second response amplitude threshold. If in response to the fact that the text features in the text feature set of the output network layer are greater than or equal to the first response amplitude threshold and less than or equal to the second response amplitude threshold, it indicates that the response amplitudes of the text features in the text feature set are within the normal range. In this case, the response amplitudes of the text features in the text feature set do not need to be processed. After determining the first response amplitude threshold and the second response amplitude threshold corresponding to the text feature set of the output network layer, the text features within the range of this first response amplitude threshold and the second response amplitude threshold can be determined as the text features included in the updated text feature set, that is, the text features included in the updated text feature set are all within the response amplitude threshold range.

[0084] For example, a piecewise function can be used to determine whether the response amplitude of the text features in the text feature set is within the normal range, and the text features with a response amplitude less than the first response amplitude threshold and the text features with a response amplitude greater than the second response amplitude threshold are removed. Among them, the piecewise function is shown in the following formula (1):

[0085]

[0086] Among them, FC(h) can be used to represent the piecewise function; h min can be used to represent the first response amplitude threshold; h can be used to represent the response amplitude of the text features within the normal range in the text feature set; h max can be used to represent the second response amplitude threshold. It can be seen from the above formula that the response amplitude of the text features in the text feature set is greater than or equal to the first response amplitude threshold and less than or equal to the second response amplitude threshold. Based on this, when the response amplitude h of the text features in the text feature set is less than the first response amplitude threshold h minWhen it is possible, the response amplitude can be truncated to the first response amplitude threshold h min ; when the response amplitude of the text feature in the text feature set is greater than the second response amplitude threshold h max When it is possible, the response amplitude can be truncated to the second response amplitude threshold h max , to ensure that the response amplitudes of the text features in the text feature set are all within the response amplitude threshold range.

[0087] Optionally, the first response amplitude threshold h min and the second response amplitude threshold h max can be obtained by using the dynamic storage feature response of a memory bank. Among them, the Memory Bank is a first-in-first-out feature memory that can store the feature responses of 3000 text units, sort the feature responses of 3000 text units, and set h max to the 0.2% percentile of the maximum response amplitude, and set h min to the 0.2% percentile of the minimum response amplitude. This is only an exemplary example here, and it does not limit the determination method of the first response amplitude threshold h min and the second response amplitude threshold h max .

[0088] As an optional implementation manner, the detection method of the reply information further includes: respectively obtaining the arrangement order of multiple hidden network layers in the dialogue model; determining a target hidden network layer among the multiple hidden network layers based on the arrangement order.

[0089] In this embodiment, since the dialogue model includes multiple hidden network layers, the multiple hidden network layers help the generative dialogue model learn more complex semantic and syntactic rules, so as to generate more accurate reply information. The multiple hidden network layers have a sequential order in the dialogue model. Based on this, the arrangement order of the multiple hidden network layers in the dialogue model can be respectively obtained, and then the target hidden network layer can be determined among the multiple hidden network layers according to the arrangement order. Among them, the target hidden network layer can be the middle layer of the multiple hidden network layers. This is only an exemplary example here, and it does not limit the specific position of the target hidden network layer in the dialogue model.

[0090] Next, the process of determining the target hidden network layer among the multiple hidden network layers based on the arrangement order of the multiple hidden network layers will be further introduced.

[0091] As an optional implementation manner, determining the target hidden network layer among the multiple hidden network layers based on the arrangement order includes: determining the hidden network layer located in the middle arrangement order as the target hidden network layer.

[0092] In this embodiment, after determining the arrangement order of multiple hidden network layers in the dialogue model, the hidden network layer with the middle arrangement order among the multiple hidden network layers can be determined as the target hidden network layer. That is to say, the target hidden network layer can be the middle layer among the multiple hidden network layers.

[0093] For example, assume the dialogue model is the llama-7B model. Assume the llama-7B model contains 33 hidden network layers. Based on this, the 17th hidden network layer (the middle layer) can be selected as the target hidden network layer.

[0094] As an alternative implementation, in step S204, determining the detection result of the reply information based on the target text features of the reply information includes: determining the metric of the reply information based on the target text features of the reply information, where the metric is used to represent the reliability degree of the reply information output by the dialogue model; determining the detection result of the reply information based on the metric.

[0095] In this embodiment, after determining the target text features of the reply information, a feature covariance matrix of the reply information can be constructed according to the target text features of the reply information, and then the metric of the reply information can be determined according to the feature covariance matrix, where the metric is used to represent the reliability degree of the reply information output by the dialogue model. For example, the metric can be the hallucination degree metric score of the dialogue model, and the hallucination degree metric score can better measure the inconsistency of the reply information output by the dialogue model. This is only an exemplary example here, and the specific content of the metric is not limited.

[0096] For example, assume the metric is the hallucination degree metric score. In this case, after obtaining the feature covariance matrix of the reply information, the metric can be calculated using the following formula (2).

[0097]

[0098] Among them, E can be used to represent the metric, K can be used to represent the number of reply information output by multiple network layers of the dialogue model, Σ can be used to represent the feature covariance matrix, and αI K is a relatively small regularization term used to prevent the covariance matrix from being a non-full rank matrix, and I K can be used to represent the identity matrix with dimension K, and α is a feature parameter that can take 0.001.

[0099] Optionally, since the determinant of a matrix can be obtained by calculating the eigenvalues, the metric in the above formula (2) can be represented by the following formula (3):

[0100]

[0101] where λ = {λ1, λ2, ……, λ K} represents the K eigenvalues of the matrix Σ + αI K .

[0102] Optionally, after determining the metric for the response information, the detection result of the response information can be determined based on the metric, where the detection result can be used to indicate whether the response information output by the dialogue model is reliable.

[0103] Optionally, after determining the target text features of the response information, the target text features of the response information can also be evaluated using a machine learning model to determine the detection result of the response information, or a text similarity algorithm can be used to measure the similarity between the response information and the input information, thereby determining the reliability of the response information. This is only an exemplary example and does not limit the specific method for determining the detection result of the response information based on the target text features of the response information.

[0104] As an alternative implementation, determining the detection result of the response information based on the metric includes: in response to the metric being greater than the metric threshold, determining that the detection result is that the response information output by the dialogue model is unreliable; in response to the metric being less than or equal to the metric threshold, determining that the detection result is that the response information output by the dialogue model is reliable.

[0105] In this embodiment, since there is a knowledge hallucination problem in the response information output by the dialogue model, that is, the response information output by the dialogue model is unreliable. Based on this, the metric threshold can be used to measure the metric of the response information output by the dialogue model, so as to determine whether the response information output by the dialogue model is reliable. For example, the metric of the response information output by the dialogue model can be compared with the metric threshold, and then according to the comparison result, it can be determined whether the response information output by the dialogue model is reliable.

[0106] For example, if the metric is greater than the metric threshold, it means that the response information output by the dialogue model may have a relatively high degree of hallucination. Based on this, it can be determined that the detection result of the response information is that the response information output by the dialogue model is unreliable; on the contrary, if the metric is less than or equal to the metric threshold, it means that the degree of hallucination in the response information output by the dialogue model is relatively low. In this case, it can be determined that the detection result of the response information is that the response information output by the dialogue model is reliable.

[0107] As an alternative implementation, the detection method of the response information further includes: in the case where the detection result is that the response information output by the dialogue model is unreliable, prohibiting the output of the response information; in the case where the detection result is that the response information output by the dialogue model is reliable, outputting the response information.

[0108] In this embodiment, when the reply information output by the dialogue model is unreliable, the dialogue model can be prohibited from outputting the reply information. Only when the reply information output by the dialogue model is reliable, the dialogue model is allowed to output the reply information, so as to ensure that only the verified reliable reply information will be output, avoiding the spread of misleading or incorrect information.

[0109] As an alternative implementation, based on the target text features of the reply information, a metric index of the reply information is determined, including: constructing a feature covariance matrix by using at least one target text feature of at least one reply information; and determining the metric index of the reply information based on the feature covariance matrix.

[0110] In this embodiment, as can be seen from the foregoing introduction, the target text feature is the text feature corresponding to the target text unit of the reply information, and the target text unit includes the semantics of the reply information. Based on this, when determining the metric index of the reply information based on the target text features of the reply information, a feature covariance matrix can be constructed according to the target text features of the reply information, and then the metric index of the reply information can be determined based on the feature covariance matrix.

[0111] For example, when there are multiple reply information in at least one reply information, since the multiple reply information respectively correspond to target text features, in this case, the feature covariance matrix can be constructed by the following formula (4).

[0112] ∑=Z T ·J d ·Z (4)

[0113] Where Z can be used to represent the target text feature matrix corresponding to the reply information, and Z T can be used to represent the transpose matrix of the target text feature matrix corresponding to the reply information, and J d can be used to represent the centering matrix, where where 1 K can be used to represent a column vector of all 1s with a dimension of K.

[0114] Optionally, the target text feature matrix Z corresponding to the reply information can be obtained by splicing the target text features corresponding to multiple reply information. For example, Z = [Z1, Z2,... Z i …, Z k , where Z i can be used to represent the target text feature h T corresponding to the i-th reply information, that is, Z i = h T , where the target text feature h T can be the text feature of the last text unit of the i-th reply information.

[0115] Optionally, when there is only one response message in at least one response message, K in the above formula can be set to 1, that is, let K = 1. Then, construct the covariance matrix with reference to the method introduced above, and further use the covariance matrix to determine the metric index of the response message, which will not be elaborated here.

[0116] As an alternative implementation, based on the feature covariance matrix, determine the metric index of the response message, including: determining the logarithmic determinant of the feature covariance matrix; determining the metric index corresponding to the logarithmic determinant.

[0117] In this embodiment, since the feature covariance matrix is a matrix used to describe the correlation and variance between different features, based on this, after determining the feature covariance matrix, the logarithmic determinant of the feature covariance matrix can be further determined. The logarithmic determinant can be used to measure the degree of correlation and variance between features, and then determine the metric index corresponding to the logarithmic determinant according to the logarithmic determinant to measure the reliability of the response message output by the model.

[0118] In the above steps, by calculating the metric index of the response message output by the dialogue model, and then directly using the metric index to determine the reliability degree of the response message output by the dialogue model, it can better represent the uncertainty degree of the entire output sentence corresponding to the response message, thereby reflecting the knowledge hallucination degree of the response message, that is, the reliability degree of the response message.

[0119] Under the above operating environment, the present application also provides a method for generating information as shown in Figure 3 which is applied to a question-and-answer system deployed in a scenario task, for example, a generative question-and-answer system. Figure 3 It is a flowchart of another method for generating information according to an embodiment of the present application. As shown in Figure 3 it, the method may include the following steps.

[0120] Step S301, on the operation interface of the question-and-answer system, monitor the input information in the scenario task.

[0121] In the technical solution provided in step S301 of the present application above, the question-and-answer system can be deployed in a scenario task, and the question-and-answer system can generate corresponding response messages according to the input information proposed by the user. The operation interface can be an interface for the user to interact with the question-and-answer system. The user can input input information related to the scenario task on the operation interface. Among them, the input information can be any question proposed by the user, and the input information in the scenario task can be monitored on the operation interface of the question-and-answer system.

[0122] Step S302: Retrieve the dialogue model that matches the scenario task, input the input information into the dialogue model for analysis, and obtain at least one reply message that matches the input information.

[0123] In the technical solution provided in step S302 of the present application above, after detecting the input information in the scenario task, a dialogue model that matches the scenario task can be retrieved. Among them, the dialogue model can be a large model. Then, the input information is input into the dialogue model for analysis to obtain at least one reply message that matches the input information.

[0124] In this embodiment, the dialogue model is used to analyze the input information to generate at least one reply message that matches the input information. Based on this, after the question-and-answer system detects the input information in the scenario task, a dialogue model that matches the scenario task can be retrieved. Since the internal state information of the dialogue model can help the large model better understand the user's input information, based on this, the internal state information of the dialogue model can be used to analyze the input information so that the large model generates at least one reply message that matches the input information.

[0125] Step S303: Obtain the target text feature of the reply message from the internal state information of the dialogue model.

[0126] In the technical solution provided in step S303 of the present application above, the internal state information of the dialogue model is used to represent the rule for the dialogue model to analyze the input information. Based on this, after analyzing the input information using the internal state information of the dialogue model, the target text feature of the reply message can be obtained from the internal state information of the dialogue model. Among them, the target text feature is used to represent the semantics of the reply message in the scenario task.

[0127] In this embodiment, the internal state information of the dialogue model may include the rules for multiple network layers inside the dialogue model to analyze the input information. Among them, the multiple network layers of the dialogue model may include: a decoding layer (Decoder), a fully connected layer (FC Layer), etc. The multiple network layers are used to perform different processing operations on the input information, and then an output result is obtained. The target text feature of the reply message can be obtained from the output results of the multiple network layers. The method for obtaining the target text feature of the reply message can refer to the introduction in the foregoing step S203 and will not be elaborated here.

[0128] Step S304: Determine the detection result of the reply message based on the target text feature of the reply message.

[0129] In the technical solution provided in step S304 of the present application, after obtaining the target text features of the response information, the detection result of the response information can be determined based on the target text features of the response information, where the detection result is used to indicate whether the response information output by the dialogue model is reliable.

[0130] In this embodiment, after obtaining the target text features of the response information, a covariance matrix can be constructed according to the target text features of the response information, and then a metric index of the response information output by the dialogue model can be calculated based on the covariance matrix. Among them, the metric index can be the hallucination degree metric score corresponding to the response information output by the dialogue model. After obtaining the metric index, it can be further determined whether the response information output by the dialogue model is reliable according to the metric index and the metric index threshold. The specific method for determining whether the response information output by the dialogue model is reliable according to the metric index and the metric index threshold can refer to the introduction in the foregoing step S204, which will not be elaborated here.

[0131] Step S305, in response to the detection result indicating that the response information output by the dialogue model is reliable in the scenario task, output the response information.

[0132] In the technical solution provided in step S305 of the present application, after determining the detection result of the response information, it can be determined whether to output the response information according to whether the response information output by the dialogue model indicated by the detection result is reliable in the scenario task.

[0133] In this embodiment, if the detection result indicates that the response information output by the dialogue model is reliable in the scenario task, it means that the accuracy of the response information output by the dialogue model is relatively high. In this case, the response information can be output.

[0134] Step S306, in response to the detection result indicating that the response information output by the dialogue model is unreliable in the scenario task, output the corresponding prompt information.

[0135] In the technical solution provided in step S306 of the present application, different from step S305, if the detection result indicates that the response information output by the dialogue model is unreliable in the scenario task, it means that the accuracy of the response information output by the dialogue model is relatively low. In this case, the response information can be not output, but the corresponding prompt information can be output to remind the user that the response information output by the dialogue model does not match the input information.

[0136] Based on steps S301 to S306 of the above embodiments, by using the internal state information of the dialogue model to obtain the target text features of the reply information output by the dialogue model, the semantic features of the reply information can be better mined and utilized. Furthermore, the detection result of the reply information is determined using the target text features. This detection result can indicate whether the reply information is reliable in the scenario task. When the reply information is reliable in the scenario task, the reply information is output. When the reply information is not reliable in the scenario task, a prompt message is output to ensure that only the reply information verified as reliable is output, achieving the technical effect of effectively detecting the reliability of the reply information output by the dialogue model, and thus solving the technical problem of being unable to effectively detect the reply information output by the model.

[0137] According to an embodiment of the present application, a method for detecting information is also provided from the human-computer interaction side. Figure 4 It is a flowchart of another method for detecting information according to an embodiment of the present application. As Figure 4 shown, the method may include the following steps:

[0138] Step S401, monitor the input information on the dialogue interface.

[0139] In the technical solution provided in step S401 of the present application above, the dialogue interface can be an interface for the user to interact with the question-and-answer system. The user can input input information on the dialogue interface. Among them, the input information can be multimodal information, and the types of multimodal information include at least one of the following: text information containing character information, video frame information containing frame image information, and audio information. The question-and-answer system can monitor the input information on the dialogue interface.

[0140] Step S402, display at least one reply information matching the input information on the dialogue interface.

[0141] In the technical solution provided in step S402 of the present application above, after the input information is monitored on the dialogue interface, the dialogue model can be called to analyze the input information, and then at least one reply information matching the input information is obtained and displayed on the dialogue interface. Among them, the reply information can be multimodal information, and the types of the reply information can include at least one of the following: text information, image information, video information, and voice information, and the reply information is obtained by analyzing the input information using the internal state information of the dialogue model. The process of analyzing the input information using the internal state information of the dialogue model can refer to the introduction in step S202 above and will not be elaborated here.

[0142] Step S403, in response to an information detection operation on the dialogue interface, display the detection result of the reply information on the dialogue interface.

[0143] In the technical solution provided in step S403 of the present application above, when detecting an operation on the information on the dialogue interface, the detection result of the reply information can be displayed on the dialogue interface, where the detection result is used to indicate whether the reply information output by the dialogue model is reliable, and is determined based on the target text feature of the reply information, the target text feature is used to represent the semantics of the reply information, and is obtained from the internal state information of the dialogue model, and the internal state information is used to represent the rule for the dialogue model to analyze the input information. Among them, the method for determining whether the reply information output by the dialogue model is reliable can refer to the introduction in the foregoing steps S203 and S204, which will not be elaborated here.

[0144] Based on steps S401 to S403 of the above embodiment, monitor the input information on the dialogue interface; display at least one reply information matching the input information on the dialogue interface; in response to the information detection operation on the dialogue interface, display the detection result of the reply information on the dialogue interface, where the detection result is used to indicate whether the reply information output by the dialogue model is reliable. That is, the detection result of the reply information can be intuitively displayed through the dialogue interface, and the user can intuitively understand whether the reply information output by the dialogue model is reliable, and the user can quickly obtain the required information, improving the user experience. In addition, the display of the detection result can help the user judge the reliability of the reply information, which helps to improve the accuracy and credibility of the human-computer dialogue.

[0145] Figure 5 It is a schematic diagram of an information detection system according to an embodiment of the present application. As Figure 5 shown, the information detection system 500 includes: an information input end 501, an information detection end 502, and an information output end 503.

[0146] The information input end 501 is used to monitor the input information.

[0147] In this embodiment, the user can input information at the information input end, and then the information input end can monitor the user's input information, where the input information can be multimodal information. For example, the input information can be text information containing characters, video frame information containing frame image information, audio information, etc., and no specific limitation is made here.

[0148] The information detection end 502 is used to input the input information into the dialogue model for analysis, obtain at least one reply information that matches the input information, and obtain the target text feature of the reply information from the internal state information of the dialogue model, where the internal state information is used to represent the rules for the dialogue model to analyze the input information, and the target text feature is used to represent the semantics of the reply information; based on the target text feature of the reply information, determine the detection result of the reply information, where the detection result is used to indicate whether the reply information output by the dialogue model is reliable.

[0149] In this embodiment, after the information input end monitors the input information, it can transmit the input information to the information detection end. After the information detection end receives the input information, it can call the dialogue model, and then use the internal state information of the dialogue model to analyze the input information to obtain at least one reply information that matches the input information. And based on the target text feature of the reply information, determine the detection result of the reply information, and then determine whether the reply information output by the dialogue model is reliable according to the detection result. Among them, the step of using the internal state information of the dialogue model to analyze the input information to obtain at least one reply information that matches the input information can refer to the introduction of the foregoing step S202 and will not be elaborated here. The step of obtaining the target text feature of the reply information from the internal state information of the dialogue model can refer to the introduction of the foregoing step S203 and will not be elaborated here. The step of determining the detection result of the reply information based on the target text feature of the reply information can refer to the introduction of the foregoing step S204 and will not be elaborated here.

[0150] The information output end 503 is used to output the reply information in response to the detection result indicating that the reply information output by the dialogue model is reliable; and output the corresponding prompt information in response to the detection result indicating that the reply information output by the dialogue model is unreliable.

[0151] In this embodiment, after the information detection end determines the detection result of the reply information, it can transmit the detection result to the information output end. After the information output end receives the detection result, it can determine whether the reply information output by the dialogue model is reliable according to the detection result. If the detection result indicates that the reply information output by the dialogue model is reliable, the information output end outputs the reply information. On the contrary, if the detection result indicates that the reply information output by the dialogue model is unreliable, in this case, the information output end can output a prompt information to prompt the user that the reply information output by the dialogue model is unreliable. Among them, the prompt information can be a text prompt information. For example, "The reply information output by the dialogue model is unreliable", or image information, and the specific content of the prompt information is not limited here.

[0152] Next, the technical solutions of the embodiments of the present application will be further introduced by way of preferred embodiments.

[0153] Currently, the requirements for the authenticity and reliability of the response information output by large models are also constantly increasing. When the large model outputs content that "does not conform to facts" or "fabricates out of nothing", it is considered that there is a knowledge hallucination in the output of the large model. General large models will inevitably output incorrect and unfounded hallucinatory outputs. This potential uncertainty and unreliability pose serious challenges to the commercial implementation and application of large models. Therefore, it has become particularly important to detect whether the output of large models is reliable and accurate.

[0154] In one implementation, a hallucination detection method based on Preplexity or a hallucination detection method based on SelfCheckGPT can be used to determine whether the response information output by the large model is reliable. Among them, the hallucination detection method based on Preplexity mainly judges the uncertainty of each word in the response information output by the large model word by word, and then multiplies the uncertainties of each word to obtain the uncertainty of the response information output by the large model. However, due to the inaccurate estimation of the uncertainty of the sentence length of the response information output by the large model, and the diversity of the expression forms of the sentences in the response information output by the large model, it may lead to inaccurate detection results of the reliability of the response information output by the large model. The hallucination detection method based on SelfCheckGPT judges the reliability of the response information output by the large model by measuring the consistency of the response information output by the large model multiple times. However, this method requires relying on an additional large model to calculate the consistency of the response information output multiple times, resulting in a large computational time overhead, and this method can only detect the self-contradictory hallucinations output by the large model and cannot detect the hallucinations with high consistency output by the large model. Therefore, the above methods all have the technical problem of being unable to effectively detect the response information output by the large model.

[0155] However, the embodiments of the present application provide a method for detecting information. By obtaining the input information entered by the user and enabling the large model to output multiple response information based on the input information entered by the user; obtaining the text feature set of the response information in the penultimate layer of the internal output layer of the large model; using a dynamic feature pruning scheme to prune the text features in the text feature set that exceed the threshold to remove abnormal feature responses; obtaining the text features of the last text unit in the intermediate layer of different output hidden states of the large model; constructing a feature covariance matrix for the text features of the last text unit in different output layers of the large model; calculating a metric according to the covariance matrix; and judging whether there is an illusion in the response information output by the large model according to the metric. Wherein, when the metric is greater than the metric threshold, it is determined that there is a knowledge illusion in the response information output by the large model, that is, the response information output by the large model is unreliable. When the metric is not greater than the metric threshold, it is determined that there is no knowledge illusion in the response information output by the large model, that is, the response information output by the large model is reliable. That is to say, in the embodiments of the present application, the metric can be directly used to measure the inconsistency of the multiple response information output by the large model, which is equivalent to measuring the continuous entropy of semantics in the feature space. Using the metric can better represent the uncertainty degree of the response information output by the large model, thereby reflecting the degree of knowledge illusion of the response information, that is, reflecting whether the response information output by the large model is reliable, achieving the technical effect of effectively detecting the reliability of the response information output by the large model, and further solving the technical problem of being unable to effectively detect the reliability of the response information output by the large model.

[0156] Next, the information detection method in the embodiments of the present application will be further introduced.

[0157] Figure 6 It is a flowchart of an information detection method according to an embodiment of the present application. As Figure 6 shown, the method may include the following steps:

[0158] Step S601, obtain the user's question.

[0159] In this embodiment, the user can enter any question on the operation interface, and the large model can monitor the user's question entered on the operation interface, and then obtain the user's question.

[0160] Step S602, based on the user's question, output 10 response information.

[0161] In this embodiment, after the large model obtains the question entered by the user, it can output 10 response information according to the user's question, and the 10 response information is the information for answering the user's question.

[0162] Step S603, obtain the text feature set output by the penultimate layer of the internal output layer of the large model.

[0163] In this embodiment, the composition of the large model can be divided into two parts: the Decoder layer + the final fully connected layer (FC Layer). The penultimate layer is used to indicate the output of the Decoder layer and the input of the fully connected layer (FC Layer). The output of the Decoder layer is the text feature representation of the text unit token in the response message. The output of the Decoder layer can be obtained as the text feature of the text unit token in the response message to form a text feature set.

[0164] Step S604: Use a dynamic feature pruning scheme to prune the part of the text feature set that exceeds the threshold.

[0165] In this embodiment, a dynamic pruning scheme can be used to prune the abnormal feature responses in the text feature set that exceed the threshold, so as to remove the abnormal feature responses in the text feature set.

[0166] Figure 7 It is a schematic diagram of the feature response of a text feature set according to an embodiment of the present application. As Figure 7 shown, it shows the feature distribution of the text feature set of a large model in the penultimate layer. It can be seen from Figure 7 that the text feature set contains a large number of extreme feature responses. Such extreme feature responses are likely to cause the large model to output feature classification results with too high confidence. Based on this, in order to alleviate the hallucination output with high consistency caused by oversaturated classification, a feature pruning method can be used to prune the abnormal feature responses in the text feature set.

[0167] Optionally, feature pruning can be performed using the following piecewise function. When the feature response h exceeds the maximum threshold h max , the response is truncated to the maximum threshold h max ; when the feature response h is less than the minimum threshold h min , the response is truncated to the minimum threshold h min . Among them, the piecewise function is as shown in the foregoing formula (1), which will not be elaborated here.

[0168] Optionally, the minimum threshold h min and the maximum threshold h max can be obtained by dynamically storing the feature responses in a Memory Bank. Among them, the Memory Bank is a first-in-first-out feature memory that can store the feature responses of 3000 text units. Sort the feature responses of 3000 text units, and set h max to the 0.2% percentile of the maximum response amplitude, and set h min to the 0.2% percentile of the minimum response amplitude. This is only an exemplary example here, and does not limit the minimum threshold h minand the maximum threshold h max is defined.

[0169] Step S605: Obtain the text features of the last text unit in the intermediate layer of different output hidden states of the large model.

[0170] In this embodiment, for each question of the user, the model can output 10 answers (10 sentence outputs), and each answer contains multiple text units, that is, each answer contains multiple words (tokens). Each word (token) of the 10 sentences corresponds to a feature vector in different layers of the large model. Since in the large model, the feature of the last word (token) of a sentence often contains the information of the entire sentence, therefore, the feature of the last word (token) is commonly used as the feature of the entire sentence. Based on this, the feature of the last word (token) in the intermediate layer can be used as the feature of the entire sentence. For example, the llama7B model has 33 layers, and the feature of the last word (token) in the 17th layer can be selected as the feature of the entire sentence. That is, for the 10 sentence outputs, the intermediate layer features of the last word (token) of each sentence are obtained as the features of these 10 sentences respectively.

[0171] Step S606: Construct a sentence feature covariance matrix for the text features of the last text unit of different outputs of the large model.

[0172] In this embodiment, after obtaining the text features of the last text unit of different outputs of the large model, the sentence feature covariance matrix can be constructed using these text features. The feature representation of the i-th answer of the model can be Z i = h T . Then for the K generated answers, the covariance matrix of the sentences can be represented by the following formula (5):

[0173] ∑ = Z T · J d · Z (5)

[0174] where Z is used to represent the sentence feature matrix, and Z T can be used to represent the transpose matrix of the sentence feature matrix, and J d can be used to represent the centering matrix, where, where 1 K can be used to represent a column vector of all 1s with a dimension of K.

[0175] Optionally, the sentence feature matrix Z can be obtained by concatenating the text features of the last text unit of different answers of the large model. For example, Z = [Z1, Z2,... Z i …, Z k , where Zi It can be used to represent the text features corresponding to the i-th answer of the large model.

[0176] Step S607, calculate the hallucination degree metric according to the covariance matrix.

[0177] In this embodiment, after determining the covariance matrix, the hallucination degree metric can also be calculated according to the covariance matrix, where the hallucination degree metric can be the hallucination degree metric score.

[0178] For example, after obtaining the covariance matrix, the hallucination degree metric can be calculated by taking the logarithm determinant of the covariance matrix through the following formula (6).

[0179]

[0180] Where, E can be used to represent the hallucination degree metric, K can be used to represent the number of answers output by the large model, Σ can be used to represent the covariance matrix, and αI K is a small regularization term used to prevent the covariance matrix from being a non-full rank matrix, and I K can be used to represent the identity matrix of dimension K, α is a characteristic parameter and can take 0.001.

[0181] Optionally, since the determinant of a matrix can be obtained by calculating the eigenvalues, the metric in the above formula can be expressed as:

[0182]

[0183] Where, λ = {λ1, λ2,..., λ K} represents the K eigenvalues of the matrix Σ + αI K .

[0184] Step S608, judge whether there is knowledge hallucination in the output of the large model according to the hallucination degree metric.

[0185] In this case, after obtaining the hallucination degree metric, the hallucination degree metric can be compared with the metric threshold. If the hallucination degree metric is greater than the metric threshold, it means that there is knowledge hallucination in the output of the large model. If the hallucination degree metric is not greater than the metric threshold, it means that there is no knowledge hallucination in the output of the large model.

[0186] In the above steps S601 to S608, by obtaining the question input by the user and having the large model output multiple reply messages according to the question input by the user; obtaining the text feature set of the reply message of the penultimate layer of the output layer inside the large model; using a dynamic feature pruning scheme to prune the text features in the text feature set that exceed the threshold to remove abnormal feature responses; obtaining the text features of the last text unit of the intermediate layer of different output hidden states of the large model; constructing a feature covariance matrix for the text features of the last text unit of different output layers of the large model; calculating a metric according to the covariance matrix; judging whether there is an hallucination in the reply message output by the large model according to the metric, wherein when the metric is greater than the metric threshold, it is determined that there is a knowledge hallucination in the reply message output by the large model, that is, the reply message output by the large model is unreliable, and when the metric is not greater than the metric threshold, it is determined that there is no knowledge hallucination in the reply message output by the large model, that is, the reply message output by the large model is reliable. That is to say, in the embodiment of the present application, the metric can be directly used to measure the inconsistency of the multiple reply messages output by the large model, which is equivalent to measuring the continuous entropy of semantics in the feature space. Using the metric can better represent the uncertainty degree of the reply messages output by the large model, thereby reflecting the knowledge hallucination degree of the reply messages, that is, reflecting whether the reply messages output by the large model are reliable, achieving the technical effect of effectively detecting the reliability of the reply messages output by the large model, and further solving the technical problem of being unable to effectively detect the reliability of the reply messages output by the large model.

[0187] Figure 8 is a schematic diagram of detecting knowledge hallucination of a large model according to an embodiment of the present application. As Figure 8 shown, when the question input by the user is "What is the specific time when XXX first landed on the moon in 1969?", after receiving the question input by the user, the large language model can use multiple internal network layers to process the question, and then output K different answers. By analyzing the K different answers, K hallucination degree measurement scores are obtained, where the K hallucination degree measurement scores are respectively used to measure the reliability of the K answers. Among them, when the hallucination degree measurement score of a certain answer is higher than the score threshold, the answer corresponding to the hallucination degree measurement score of the answer is output. If there is no hallucination degree measurement score of an answer higher than the score threshold, only a prompt message is output. For example, the answer to this question is not supported.

[0188] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject.

[0189] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0190] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the technical solution of this application, in essence, or the part that makes a contribution to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions to enable a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of this application.

[0191] Embodiment 2

[0192] According to an embodiment of this application, there is also provided an information detection device for implementing the information detection method described above. Figure 9 It is a schematic diagram of an information detection device according to an embodiment of this application, as Figure 9 shown. The information detection device 900 includes: a first monitoring unit 901, a first analysis unit 902, a first acquisition unit 903, and a first determination unit 904.

[0193] The first monitoring unit 901 is used to monitor input information.

[0194] The first analysis unit 902 is used to input the input information into a dialogue model for analysis to obtain at least one reply information that matches the input information.

[0195] The first acquisition unit 903 is configured to acquire the target text features of the reply information from the internal state information of the dialogue model, where the internal state information is used to represent the rules for the dialogue model to analyze the input information, and the target text features are used to represent the semantics of the reply information.

[0196] The first determination unit 904 is configured to determine the detection result of the reply information based on the target text features of the reply information, where the detection result is used to indicate whether the reply information output by the dialogue model is reliable.

[0197] It should be noted here that the above first monitoring unit 901, first analysis unit 902, first acquisition unit 903, and first determination unit 904 correspond to steps S201 to S204 in Embodiment 1. The functions implemented by the four units and the corresponding steps are the same in terms of examples and application scenarios, but are not limited to the content disclosed in the above Embodiment 1. It should be noted that the above modules or units may be hardware components or software components stored in a memory (for example, memory 104) and processed by one or more processors (for example, processors 102a, 102b,..., 102n). The above modules may also be part of a device and can run in the computer terminal 10 provided in Embodiment 1.

[0198] According to an embodiment of the present application, there is also provided an information generation device for implementing the above information generation method. Figure 10 It is a schematic diagram of another information generation device according to an embodiment of the present application. As Figure 10 described, the information generation device 1000 may include: a second monitoring unit 1001, a second analysis unit 1002, a second acquisition unit 1003, a second determination unit 1004, a first output unit 1005, and a second output unit 1006.

[0199] The second monitoring unit 1001 is configured to monitor the input information in the scenario task on the operation interface of the question-and-answer system.

[0200] The second analysis unit 1002 is configured to retrieve a dialogue model that matches the scenario task, input the input information into the dialogue model for analysis, and obtain at least one reply information that matches the input information.

[0201] The second acquisition unit 1003 is configured to acquire the target text features of the reply information from the internal state information of the dialogue model, where the internal state information is used to represent the rules for the dialogue model to analyze the input information, and the target text features are used to represent the semantics of the reply information in the scenario task.

[0202] The second determination unit 1004 is configured to determine the detection result of the reply information based on the target text features of the reply information.

[0203] A first output unit 1005, configured to output reply information in response to the detection result indicating that the reply information output by the dialogue model is reliable in the scenario task.

[0204] A second output unit 1006, configured to output corresponding prompt information in response to the detection result indicating that the reply information output by the dialogue model is unreliable in the scenario task.

[0205] It should be noted here that the above-mentioned second monitoring unit 1001, second analysis unit 1002, second acquisition unit 1003, second determination unit 1004, first output unit 1005, and second output unit 1006 correspond to steps S301 to S306 in Embodiment 1. The instances and application scenarios implemented by the six units and the corresponding steps are the same, but are not limited to the content disclosed in the above-mentioned Embodiment 1. It should be noted that the above-mentioned module or unit may be a hardware component or software component stored in a memory (for example, memory 104) and processed by one or more processors (for example, processors 102a, 102b,..., 102n). The above-mentioned module may also be part of a device and may run in the computer terminal 10 provided in Embodiment 1.

[0206] According to an embodiment of the present application, there is also provided an information detection device for implementing the above-mentioned information detection method. Figure 11 It is a schematic diagram of another information detection device according to an embodiment of the present application, as Figure 11 shown. The information detection device 1100 may include: a third monitoring unit 1101, a first display unit 1102, and a second display unit 1103.

[0207] The third monitoring unit 1101 is configured to monitor input information on the dialogue interface.

[0208] The first display unit 1102 is configured to display at least one reply information matching the input information on the dialogue interface, where the reply information is obtained by analyzing the input information using the internal state information of the dialogue model.

[0209] The second display unit 1103 is configured to display the detection result of the reply information on the dialogue interface in response to an information detection operation on the dialogue interface, where the detection result is used to indicate whether the reply information output by the dialogue model is reliable and is determined based on the target text feature of the reply information. The target text feature is used to represent the semantics of the reply information and is obtained from the internal state information of the dialogue model, and the internal state information is used to represent the rule for the dialogue model to analyze the input information.

[0210] It should be noted that the above third monitoring unit 1101, first display unit 1102, and second display unit 1103 correspond to steps S401 to S403 in Embodiment 1. The instances and application scenarios realized by the three units and the corresponding steps are the same, but are not limited to the content disclosed in the above Embodiment 1. It should be noted that the above modules or units can be hardware components or software components stored in a memory (for example, memory 104) and processed by one or more processors (for example, processors 102a, 102b,..., 102n). The above modules can also be part of a device and can run in the computer terminal 10 provided in Embodiment 1.

[0211] It should be noted that the preferred implementation schemes involved in the above embodiments of the present application are the same as the schemes, application scenarios, and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0212] Embodiment 3

[0213] An embodiment of the present application can provide a computer terminal, and the computer terminal can be any computer terminal device in a computer terminal group. Optionally, in this embodiment, the above computer terminal can also be replaced with a terminal device such as a mobile terminal.

[0214] Optionally, in this embodiment, the above computer terminal can be located in at least one of multiple network devices in a computer network.

[0215] In this embodiment, the above computer terminal can execute program codes of the following steps in an information detection method: monitoring input information; inputting the input information into a dialogue model for analysis to obtain at least one reply information matching the input information; obtaining, from the internal state information of the dialogue model, a target text feature of the reply information, where the internal state information is used to represent the rules for the dialogue model to analyze the input information, and the target text feature is used to represent the semantics of the reply information; determining a detection result of the reply information based on the target text feature of the reply information, where the detection result is used to represent whether the reply information output by the dialogue model is reliable.

[0216] Optionally, Figure 12 is a structural block diagram of a computer terminal according to an embodiment of the present application. As Figure 12 shown, the computer terminal A may include: one or more (only one is shown in the figure) processors 1202, a memory 1204, a storage controller, and a peripheral interface, where the peripheral interface is connected to a radio frequency module, an audio module, and a display.

[0217] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the detection method and device of reply information in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, to implement the above-mentioned detection method of reply information. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory may further include a memory remotely provided with respect to the processor, and these remote memories can be connected to terminal A through a network. Examples of the above network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and their combinations.

[0218] The processor can call the information and application programs stored in the memory through the transmission device to execute the following steps: monitor the input information; input the input information into the dialogue model for analysis to obtain at least one reply information that matches the input information; obtain the target text feature of the reply information from the internal state information of the dialogue model, where the internal state information is used to represent the rule for the dialogue model to analyze the input information, and the target text feature is used to represent the semantics of the reply information; based on the target text feature of the reply information, determine the detection result of the reply information, where the detection result is used to represent whether the reply information output by the dialogue model is reliable.

[0219] Optionally, the above processor may further execute the program code of the following steps: obtain the text feature set of the reply information in the network layer of the dialogue model, where the internal state information includes the text feature set in the network layer, and the text feature set in the network layer includes the text features of the text units constituting the reply information; determine the target text feature from the text feature set in the network layer.

[0220] Optionally, the above processor may further execute the program code of the following steps: determine the text feature corresponding to the target text unit in the reply information from the text feature set in the network layer, where the target text unit includes the semantics of the reply information; determine the text feature corresponding to the target text unit as the target text feature.

[0221] Optionally, the above processor may further execute the program code of the following steps: obtain the text feature set of the reply information in the output network layer of the dialogue model; determine the text feature set of the target hidden network layer corresponding to the text feature set of the output network layer in the target hidden network layer of the dialogue model.

[0222] Optionally, the above-mentioned processor may also execute program code for the following steps: updating the text feature set of the output network layer to obtain an updated text feature set, where the updated text feature set does not include noise features; in the target hidden network layer, determining the text feature set of the target hidden network layer corresponding to the updated text feature set.

[0223] Optionally, the above-mentioned processor may also execute program code for the following steps: in response to the response amplitude of the text features in the text feature set of the output network layer being less than the first response amplitude threshold, adjusting the response amplitude to the first response amplitude threshold; in response to the response amplitude of the text features in the text feature set of the output network layer being greater than or equal to the first response amplitude threshold and less than or equal to the second response amplitude threshold, keeping the response amplitude, where the second response amplitude threshold is greater than or equal to the first response amplitude threshold; in response to the response amplitude of the text features in the text feature set of the output network layer being greater than the second response amplitude threshold, adjusting the response amplitude to the second response amplitude threshold; determining the adjusted first response amplitude threshold, the kept response amplitude, and the adjusted second response amplitude threshold as the updated text feature set.

[0224] Optionally, the above-mentioned processor may also execute program code for the following steps: respectively obtaining the arrangement order of multiple hidden network layers in the dialogue model; based on the arrangement order, determining the target hidden network layer among the multiple hidden network layers.

[0225] Optionally, the above-mentioned processor may also execute program code for the following steps: determining the hidden network layer with the middle arrangement order as the target hidden network layer.

[0226] Optionally, the above-mentioned processor may also execute program code for the following steps: determining a metric index of the reply information based on the target text features of the reply information, where the metric index is used to represent the reliability degree of the reply information output by the dialogue model; determining the detection result of the reply information based on the metric index.

[0227] Optionally, the above-mentioned processor may also execute program code for the following steps: in response to the metric index being greater than the metric index threshold, determining that the detection result is that the reply information output by the dialogue model is unreliable; in response to the metric index being less than or equal to the metric index threshold, determining that the detection result is that the reply information output by the dialogue model is reliable.

[0228] Optionally, the above-mentioned processor may also execute program code for the following steps: in the case where the detection result is that the reply information output by the dialogue model is unreliable, prohibiting the output of the reply information; in the case where the detection result is that the reply information output by the dialogue model is reliable, outputting the reply information.

[0229] Optionally, the above-mentioned processor may also execute program code of the following steps: constructing a feature covariance matrix by using multiple target text features of multiple reply messages; determining a metric index of the reply message based on the feature covariance matrix.

[0230] By adopting the embodiment of the present application, a method for detecting information is provided. By using the internal state information of the dialogue model to obtain the target text features of the reply message output by the dialogue model, the semantic features of the reply message can be better mined and utilized. Furthermore, the detection result of the reply message is determined by using the target text features. The detection result can characterize the overall uncertainty degree of the reply message, that is, it reflects whether the reply message output by the dialogue model is reliable, achieving the technical effect of effectively detecting the reliability of the reply message output by the dialogue model, and further solving the technical problem that the reliability of the reply message output by the large model cannot be effectively detected.

[0231] Those of ordinary skill in the art can understand that Figure 12 The structure shown is only schematic. The computer terminal may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a palm computer, and a mobile Internet device (Mobile Internet Devices, MID), a PAD and other terminal devices. Figure 12 It does not limit the structure of the above-mentioned electronic device. For example, computer terminal A may further include more or fewer components (such as a network interface, a display device, etc.) than those shown in Figure 12 or have a different configuration from that shown in Figure 12 shown.

[0232] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the relevant hardware of the terminal device through a program. The program can be stored in a computer-readable storage medium. The storage medium may include: a flash drive, a read-only memory (Read-Only Memory, ROM), a random access memory (Random Access Memory, RAM), a magnetic disk or an optical disc, etc.

[0233] Embodiment 4

[0234] The embodiment of the present application also provides a computer-readable storage medium. Optionally, in this embodiment, the above storage medium may be used to save the program code executed by the detection method of the reply message provided in the first embodiment above.

[0235] Optionally, in this embodiment, the above storage medium may be located in any one of the computer terminals in the computer terminal group in the computer network, or in any one of the mobile terminals in the mobile terminal group.

[0236] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: monitoring input information; inputting the input information into a dialogue model for analysis to obtain at least one reply information that matches the data information; obtaining, from the internal state information of the dialogue model, the target text feature of the reply information, where the internal state information is used to represent the rules for the dialogue model to analyze the input information, and the target text feature is used to represent the semantics of the reply information; determining the detection result of the reply information based on the target text feature of the reply information, where the detection result is used to indicate whether the reply information output by the dialogue model is reliable.

[0237] An embodiment of the present application also provides a computer program product, including computer instructions that, when executed by a processor, implement the information detection method provided by the embodiment of the present application.

[0238] In this embodiment, the above computer instructions may be stored in a read-only memory (ROM), or may be computer instructions loaded from a storage unit into a random access memory (RAM) to be executed by the processor for various appropriate actions and processes in the state detection method of the database product.

[0239] In some embodiments, part or all of the above computer instructions may be loaded and / or installed onto an electronic device via a read-only memory and / or a communication unit. When the computer instructions are loaded into the random access memory and executed by a computing unit, one or more steps in the information detection method described above may be executed.

[0240] The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages or disadvantages of the embodiments.

[0241] In the above embodiments of the present application, the descriptions of the respective embodiments have their own focuses. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0242] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other may be through some interfaces, and the indirect couplings or communication connections of units or modules may be in an electrical or other form.

[0243] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0244] In addition, each functional unit in various embodiments of the present application may be integrated in a processing unit, may exist separately as individual physical units, or two or more units may be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0245] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical disks and other various media that can store program codes.

[0246] The above is only the preferred embodiment of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A method for detecting information, characterized in that, Including: Monitoring input information; Inputting the input information into a dialogue model for analysis to obtain at least one reply information matching the input information; Obtaining, from the internal state information of the dialogue model, a target text feature of the reply information, where the internal state information is used to represent the rules for the dialogue model to analyze the input information, and the target text feature is used to represent the semantics of the reply information; Determining a detection result of the reply information based on the target text feature of the reply information, where the detection result is used to indicate whether the reply information output by the dialogue model is reliable.

2. The method according to claim 1, characterized in that Obtaining, from the internal state information of the dialogue model, a target text feature corresponding to the reply information, including: Obtaining a text feature set of the reply information in the network layer of the dialogue model, where the internal state information includes the text feature set in the network layer, and the text feature set in the network layer includes text features of text units constituting the reply information; Determining the target text feature from the text feature set in the network layer.

3. The method according to claim 2, characterized in that, Determining the target text feature from the text feature set in the network layer, including: Determining, from the text feature set in the network layer, a text feature corresponding to a target text unit in the reply information, where the target text unit includes the semantics of the reply information; Determining the text feature corresponding to the target text unit as the target text feature.

4. The method according to claim 2, wherein Obtaining the text feature set of the reply information in the network layer of the dialogue model, including: Obtaining the text feature set of the reply information in the output network layer of the dialogue model; Determining, in a target hidden network layer of the dialogue model, a text feature set of the target hidden network layer corresponding to the text feature set of the output network layer.

5. The method according to claim 4, characterized in that, Determining, in a target hidden network layer of the dialogue model, a text feature set of the target hidden network layer corresponding to the text feature set of the output network layer, including: Updating the text feature set of the output network layer to obtain an updated text feature set, where the updated text feature set does not include noise features; Determining, in the target hidden network layer, a text feature set of the target hidden network layer corresponding to the updated text feature set.

6. The method according to claim 5, characterized in that, Updating the text feature set of the output network layer to obtain an updated text feature set, including: In response to the response amplitude of a text feature in the text feature set of the output network layer being less than a first response amplitude threshold, adjusting the response amplitude to the first response amplitude threshold; In response to the response amplitude of a text feature in the text feature set of the output network layer being greater than or equal to the first response amplitude threshold and less than or equal to a second response amplitude threshold, maintaining the response amplitude, where the second response amplitude threshold is greater than or equal to the first response amplitude threshold; In response to the response amplitude of a text feature in the text feature set of the output network layer being greater than the second response amplitude threshold, adjusting the response amplitude to the second response amplitude threshold; Determine the adjusted first response amplitude threshold, the maintained response amplitude, and the adjusted second response amplitude threshold as the updated text feature set.

7. The method according to claim 4, wherein The method further includes: Obtain the arrangement order of multiple hidden network layers in the dialogue model respectively; Based on the arrangement order, determine the target hidden network layer among the multiple hidden network layers.

8. The method according to claim 7, wherein Based on the arrangement order, determining the target hidden network layer among the multiple hidden network layers includes: Determine the hidden network layer with the middle arrangement order as the target hidden network layer.

9. The method according to claim 1, wherein Based on the target text feature of the reply information, determining the detection result of the reply information includes: Based on the target text feature of the reply information, determine the metric of the reply information, where the metric is used to represent the reliability of the dialogue model outputting the reply information; Determine the detection result of the reply information based on the metric.

10. The method according to claim 9, characterized in that, Based on the metric, determining the detection result of the reply information includes: In response to the metric being greater than the metric threshold, determine that the detection result is that the reply information output by the dialogue model is unreliable; In response to the metric being less than or equal to the metric threshold, determine that the detection result is that the reply information output by the dialogue model is reliable.

11. The method according to claim 10, wherein The method further includes: In the case where the detection result is that the reply information output by the dialogue model is unreliable, prohibit the output of the reply information; In the case where the detection result is that the reply information output by the dialogue model is reliable, output the reply information.

12. The method according to claim 9, wherein Based on the target text feature of the reply information, determining the metric of the reply information includes: Use at least one target text feature of at least one reply information to construct a feature covariance matrix; Based on the feature covariance matrix, determine the metric of the reply information.

13. A method for generating information, characterized in that, Applied to a question-and-answer system deployed in a scenario task, the method includes: On the operation interface of the question-and-answer system, monitor the input information in the scenario task; Retrieve a dialogue model matching the scenario task, input the input information into the dialogue model for analysis, and obtain at least one reply information matching the input information; Obtain the target text feature of the reply information from the internal state information of the dialogue model, where the internal state information is used to represent the rule for the dialogue model to analyze the input information, and the target text feature is used to represent the semantics of the reply information in the scenario task; Based on the target text feature of the reply information, determine the detection result of the reply information; In response to the detection result being that the reply information output by the dialogue model is reliable in the scenario task, output the reply information; In response to the detection result being that the reply information output by the dialogue model is unreliable in the scenario task, output the corresponding prompt information.

14. A method for detecting information, characterized in that, Includes: Monitor the input information on the dialogue interface; On the dialogue interface, at least one reply message matching the input message is displayed, where the reply message is obtained by analyzing the input message using a dialogue model. In response to an information detection operation on the dialogue interface, a detection result of the reply message is displayed on the dialogue interface, where the detection result is used to indicate whether the reply message output by the dialogue model is reliable and is determined based on the target text feature of the reply message, the target text feature is used to represent the semantics of the reply message, and is obtained from the internal state information of the dialogue model, and the internal state information is used to represent the rule for the dialogue model to analyze the input message.

15. The method according to claim 14, characterized in that, The input message and the reply message are multi-modal messages, and the types of the multi-modal messages include at least one of the following: text information including character information, video frame information including frame image information, and audio information, and the types of the reply message include at least one of the following: text information, image information, video information, and voice information.

16. An information detection system, characterized in that, It includes: An information input end for monitoring the input message. An information detection end for inputting the input message into a dialogue model for analysis to obtain at least one reply message matching the input message. Obtain the target text feature of the reply message from the internal state information of the dialogue model, where the internal state information is used to represent the rule for the dialogue model to analyze the input message, and the target text feature is used to represent the semantics of the reply message; determine the detection result of the reply message based on the target text feature of the reply message, where the detection result is used to indicate whether the reply message output by the dialogue model is reliable. An information output end for outputting the reply message in response to the detection result indicating that the reply message output by the dialogue model is reliable; and outputting a corresponding prompt message in response to the detection result indicating that the reply message output by the dialogue model is unreliable.

17. An electronic device, characterized in that, It includes: A memory and a processor. The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, and when the computer-executable instructions are executed by the processor, the steps of the method according to any one of claims 1 to 15 are implemented.

Citation Information

Cited By

  • Large model illusion detection method, system and device based on PCA contribution rate and medium

    CN120781171A

  • Large-scale hallucination detection method, system, equipment, and media based on PCA contribution rate

    CN120781171B

  • Plan validation

    US20260187063A1