Retrieval method, device and equipment based on attention guidance and medium
Through the attention-guided search method, the utilization of word elements in the intelligent question-and-answer system is detected and corrected, and the problems of "context illusion" phenomena and waste of computing resources in the prior art are solved, achieving higher interpretability of the generation process and accuracy of the search results.
Patent Information
- Application Number
- CN202510337947.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-06-20
AI Technical Summary
The prior art has limited effect in alleviating the phenomenon of "contextual illusion" in intelligent question-and-answer systems, and there are problems of insufficient interpretability and waste of computing resources, which limits the effectiveness and reliability of large language models in practical applications.
The attention-guided search method is adopted to obtain the attention matrix in the search model, construct the attention feature vector of the word element, and use the logistic regression classifier to detect the utilization of word element, and dynamically correct the probability distribution of word element to control the search results.
It improves the interpretability of the generation process, reduces the consumption of computing resources, and effectively reduces the occurrence of "context illusion" phenomena, and improves the accuracy and credibility of search results.
Smart Images

Figure CN120179878A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of artificial intelligence, finance, and medical and health technologies, and in particular, to a retrieval method, device, equipment, and medium based on attention guidance. Background Art
[0002] In intelligent customer service and question-and-answer systems, in the face of the rapid changes in financial business and the diversification of customer needs, how to ensure the accuracy and timeliness of answers is a key issue. In the field of medical and health, with the continuous increase in the demand for online medical consultations, how to accurately and quickly answer users' medical questions is also becoming increasingly important.
[0003] Currently, the mainstream method usually adopts the Retrieval-augmented Generation (RAG) technology, which is widely used in intelligent question-and-answer systems. By combining the information in the external knowledge base, high-quality answers are generated, that is, the context information related to the user's query is retrieved from the external and input together with the user's query into the Large Language Models (LLMs) for solution. However, during the generation process, the model often has a large degree of uncertainty, resulting in the phenomenon of "context hallucination" in the generated answers, that is, the generated content does not match or is irrelevant to the retrieved context information. This problem seriously affects the reliability and effectiveness of the model in practical applications.
[0004] In response to the above problems, the mainstream solutions include methods such as Contrastive Decoding (CAD) and Contextual Information Entropy Constrained Decoding (COIECD), but these solutions still have the following two main problems:
[0005] (1) Poor interpretability: These methods are mostly based on simple assumptions, and there is a lack of sufficient verification of the assumptions, and they fail to provide a convincing theoretical basis or result explanation, which limits the transparency and credibility of their practical applications.
[0006] (2) Excessive consumption of computing resources: Methods such as CAD and COIECD usually require multiple decoding processes, resulting in a large consumption of computing resources. This high computing cost is particularly disadvantageous for large-scale deployment scenarios, especially in application scenarios with high concurrent requests and high requirements for real-time performance.
[0007] It can be seen that the existing technologies have limited effects in alleviating context hallucinations, and there are also problems of insufficient method interpretability and waste of computing resources, which greatly restrict the effectiveness and reliability of LLMs in actual scenarios, and thus limit the wide promotion and value realization of large language models in practical applications. Summary of the Invention
[0008] In view of the above, it is necessary to provide a retrieval method, device, equipment and medium based on attention guidance, aiming to solve the problem of inaccurate retrieval results.
[0009] A retrieval method based on attention guidance, the retrieval method based on attention guidance includes:
[0010] Obtain a retrieval model, and obtain the attention matrix corresponding to each attention head in the retrieval model;
[0011] Construct an attention feature vector for each specified token according to each attention matrix;
[0012] Obtain the relevance of each pre-labeled specified token to the retrieval question and the retrieval answer;
[0013] Construct training data according to the relevance of each specified token to the retrieval question and the retrieval answer, and the attention feature vector of each specified token;
[0014] Use the training data to train a logistic regression classifier to obtain a detector for detecting whether a token is utilized;
[0015] When performing retrieval using the retrieval model, obtain the input data of the retrieval model;
[0016] Use the detector to real-time detect the tokens being utilized in the input data as each target token;
[0017] Obtain the feature coefficients of the detector, and use the feature coefficients to calculate the utilization probability of each target token;
[0018] Modify the original probability distribution of each target token according to the utilization probability of each target token to obtain the target output probability of each target token;
[0019] Control the retrieval model to output retrieval results according to the target output probability of each target token.
[0020] A retrieval device based on attention guidance, the retrieval device based on attention guidance includes:
[0021] An obtaining unit, configured to obtain a retrieval model, and obtain the attention matrix corresponding to each attention head in the retrieval model;
[0022] A building unit, configured to construct an attention feature vector for each specified token according to each attention matrix;
[0023] The obtaining unit is further configured to obtain the relevance between each pre-labeled specified token and the retrieval question and the retrieval answer;
[0024] The building unit is further configured to construct training data according to the relevance between each specified token and the retrieval question and the retrieval answer, and the attention feature vector of each specified token;
[0025] A training unit, configured to train a logistic regression classifier using the training data to obtain a detector for detecting whether a token is utilized;
[0026] The obtaining unit is further configured to obtain the input data of the retrieval model when performing retrieval using the retrieval model;
[0027] A detection unit, configured to use the detector to detect in real time the tokens being utilized in the input data as each target token;
[0028] A calculation unit, configured to obtain the feature coefficients of the detector and calculate the utilization probability of each target token using the feature coefficients;
[0029] A correction unit, configured to correct the original probability distribution of each target token according to the utilization probability of each target token to obtain the target output probability of each target token;
[0030] A control unit, configured to control the retrieval model to output a retrieval result according to the target output probability of each target token.
[0031] A computer device, comprising:
[0032] A memory, storing at least one instruction; and
[0033] A processor, configured to execute the instruction stored in the memory to implement the attention-guided retrieval method.
[0034] A computer-readable storage medium, in which at least one instruction is stored, and the at least one instruction is executed by a processor in a computer device to implement the attention-guided retrieval method.
[0035] As can be seen from the above technical solutions, the present invention constructs training data based on the relevance of each specified token to the retrieval question and the retrieval answer, as well as the attention feature vector of each specified token, and trains a logistic regression classifier using the training data to obtain a detector for detecting whether a token is utilized, which can utilize a lightweight detector to identify the utilized tokens in real time during the inference process, thereby achieving a higher degree of explanation for the generation process; when using a retrieval model for retrieval, the detector is used to detect in real time the tokens being utilized in the input data as each target token, the original probability distribution of each target token is corrected according to the utilization probability of each target token to obtain the target output probability of each target token, and the retrieval result is controlled to be output by the retrieval model according to the target output probability of each target token, dynamically enhancing the output probability of the utilized tokens, thereby dynamically focusing the attention mechanism of the model on the context information highly relevant to the user query, and fundamentally reducing the occurrence of the "context hallucination" phenomenon. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 is a flowchart of a preferred embodiment of the retrieval method based on attention guidance of the present invention.
[0037] Figure 2 is a functional module diagram of a preferred embodiment of the retrieval device based on attention guidance of the present invention.
[0038] Figure 3 is a schematic structural diagram of a computer device of a preferred embodiment for implementing the retrieval method based on attention guidance of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0039] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0040] As Figure 1 shown, it is a flowchart of a preferred embodiment of the retrieval method based on attention guidance of the present invention. According to different requirements, the order of steps in this flowchart can be changed, and some steps can be omitted.
[0041] The retrieval method based on attention guidance is applied to one or more computer devices, and the computer device is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0042] The computer device can be any electronic product that can interact with users. For example, a personal computer, a tablet computer, a smart phone, a personal digital assistant (PDA), a game console, an Internet Protocol Television (IPTV), a smart wearable device, etc.
[0043] The computer device may further include a network device and / or a user device. Among them, the network device includes, but is not limited to, a single network server, a server group composed of multiple network servers, or a cloud composed of a large number of hosts or network servers based on cloud computing.
[0044] The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0045] Among them, artificial intelligence (AI) is a theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.
[0046] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0047] The network where the computer device is located includes, but is not limited to, the Internet, wide area network, metropolitan area network, local area network, virtual private network (VPN), etc.
[0048] S10. Obtain a retrieval model and obtain the attention matrix corresponding to each attention head in the retrieval model.
[0049] In this embodiment, the retrieval model can be an artificial intelligence model with a question-and-answer function such as retrieval-augmented language models (RALMs).
[0050] In this embodiment, the retrieval model may be based on a multi-head attention mechanism, and each attention head corresponds to an attention matrix.
[0051] S11. Construct an attention feature vector for each specified token according to each attention matrix.
[0052] In this embodiment, the constructing an attention feature vector for each specified token according to each attention matrix includes:
[0053] Extract the elements corresponding to each specified token from the elements of the last row of each attention matrix as the attention weights of each specified token under the corresponding attention head;
[0054] Obtain the attention weights of each specified token under each attention head to construct the attention feature vector of each specified token.
[0055] Among them, the attention weight of each specified token under the corresponding attention head is used to reflect the attention degree of each attention head relative to each specified token.
[0056] Among them, the specified token may be any token in the external knowledge cited from an external database. Such as keywords "patient", "medicine", etc. in the field of medical and health, and "bank", "amount", etc. in the financial field.
[0057] For example: for any token, if there are three attention matrices, the attention weights corresponding to this token in the three attention matrices are 7, 8, and 9 respectively, then the attention feature vector corresponding to this token is (7, 8, 9).
[0058] In the above embodiment, since the last row of the attention matrix in the attention mechanism represents the degree of attention of the model to the context in the current inference step, therefore, extract the elements corresponding to each specified token from the elements of the last row of each attention matrix as the attention weights of each specified token under the corresponding attention head, so that the subsequent model inference can fully consider the context of the model.
[0059] S12. Obtain the relevance between each pre-labeled specified token and the retrieval question and the retrieval answer.
[0060] In this embodiment, each token may be labeled according to the influence degree of each token on the retrieval answer during actual retrieval.
[0061] For example: if the influence degree is high, the corresponding token is labeled as 1, indicating that this token is utilized during retrieval; if the influence degree is low, the corresponding token is labeled as 0, indicating that this token is not utilized during retrieval.
[0062] S13. Construct training data based on the relevance of each specified token to the retrieval question and the retrieval answer, as well as the attention feature vector of each specified token.
[0063] In this embodiment, the constructing training data based on the relevance of each specified token to the retrieval question and the retrieval answer, as well as the attention feature vector of each specified token includes:
[0064] When there is a first specified token in each specified token that is relevant to the retrieval question and the retrieval answer, determine the attention feature vector of the first specified token as a positive sample;
[0065] When there is a second specified token in each specified token that is not relevant to the retrieval question and the retrieval answer, determine the attention feature vector of the second specified token as a negative sample;
[0066] Combine the positive samples and the negative samples to obtain the training data.
[0067] Among them, the ratio of the positive samples to the negative samples can be 1:1.
[0068] In the above embodiment, based on the constructed positive samples and negative samples, the detector obtained by subsequent training can accurately identify the tokens that play an important role in the output of the retrieval model, thereby improving the accuracy of the output of the retrieval model.
[0069] S14. Use the training data to train a logistic regression classifier to obtain a detector for detecting whether a token is utilized.
[0070] In this embodiment, the output of the detector can include 1 or 0.
[0071] Specifically, when the output of the detector is 1, it indicates that the corresponding token is being utilized; when the output of the detector is 0, it indicates that the corresponding token is not being utilized.
[0072] S15. When performing retrieval using the retrieval model, obtain the input data of the retrieval model.
[0073] In this embodiment, the input data can include data cited from an external knowledge base, as well as the user's question content, etc.
[0074] S16. Use the detector to real-time detect the tokens being utilized in the input data as each target token.
[0075] In this embodiment, the using the detector to real-time detect the tokens being utilized in the input data as each target token includes:
[0076] After inputting the input data into the retrieval model, the attention matrix corresponding to each attention head in the retrieval model is obtained in real time to construct the attention feature vector of each token in the input data;
[0077] Input the attention feature vector of each token into the detector;
[0078] When the output of the detector corresponding to the first token is the first output data, it is determined that the first token is being utilized, and the first token is determined to be the target token.
[0079] Among them, the first output data can be customized. For example: the first output data can be 1.
[0080] S17. Obtain the feature coefficients of the detector, and calculate the utilization probability of each target token using the feature coefficients.
[0081] In this embodiment, the feature coefficients of the detector are parameters obtained through model training, which are used to reflect the influence direction and degree of each feature on the target variable, and to evaluate the contribution of different attention heads during the retrieval process.
[0082] Among them, in terms of the influence direction, the positive or negative of the feature coefficient can reflect the relationship between the feature and the target variable. A positive coefficient indicates that as the feature value increases, the probability of the target variable taking the positive class (such as 1) increases; a negative coefficient means that as the feature value increases, the probability of the positive class decreases. For example: in the financial field, when predicting whether to purchase a product, a positive coefficient of the "income" feature indicates that the higher the income, the greater the purchase possibility; a negative coefficient of the "product negative review rate" indicates that the higher the negative review rate, the smaller the purchase possibility.
[0083] Among them, in terms of the influence degree, the absolute value of the feature coefficient reflects the influence degree. The larger the absolute value, the more significant the influence of the feature on the target variable. For example: in the medical and health field, when judging whether a patient is ill, if the absolute value of the feature coefficient of a "certain key indicator" is large, it indicates that this indicator plays a key role in the illness judgment. However, the dimensions of different features may affect the absolute value of the coefficient, and usually, the data needs to be standardized before comparison.
[0084] In this embodiment, calculating the utilization probability of each target token using the feature coefficients includes:
[0085] Obtain the attention weights of each target token under each attention head in real time;
[0086] Calculate the comprehensive utilization score of each target token according to the attention weights of each target token under each attention head and the feature coefficients;
[0087] Calculate the cumulative sum of the comprehensive utilization scores of each target token;
[0088] Calculate the quotient of the comprehensive utilization score of each target token and the cumulative sum to obtain the utilization probability of each target token.
[0089] Specifically, the calculating the comprehensive utilization score of each target token according to the attention weight of each target token under each attention head and the feature coefficient includes:
[0090] Use the following formula to calculate the feature coefficient corresponding to each attention head after normalization:
[0091]
[0092] where, w i represents the feature coefficient corresponding to the i-th attention head after normalization; c i represents the original feature coefficient corresponding to the i-th attention head; i is a positive integer; I represents the total number of attention heads;
[0093] Use the following formula to calculate the comprehensive utilization score of each target token according to the feature coefficient corresponding to each attention head after normalization and the attention weight of each target token under each attention head:
[0094]
[0095] where, s j represents the comprehensive utilization score of the j-th target token; a ij represents the attention weight of the j-th target token under the i-th attention head; when f(v j ) = 1, it means the j-th target token v j is being utilized; when f(v j ) = 0, it means the j-th target token v j is not being utilized; j is a positive integer.
[0096] In the above embodiment, since it is difficult to determine the specific utilization degree of each token by the retrieval model only through the output of the detector, therefore, based on the weighted method of the detector feature coefficient, the attention weights are weighted and summed, and a comprehensive utilization score can be calculated for each token. This comprehensive utilization score can quantify the utilization degree of the token by the retrieval model and be used to dynamically adjust the model output process.
[0097] S18, correct the original probability distribution of each target token according to the utilization probability of each target token to obtain the target output probability of each target token.
[0098] In this embodiment, modifying the original probability distribution of each target token according to the utilization probability of each target token to obtain the target output probability of each target token includes:
[0099] Obtaining the original output probability of each target token from the original probability distribution;
[0100] Calculating the sum of the utilization probability of each target token and the corresponding original output probability to obtain the target output probability of each target token.
[0101] Through the above embodiment, the probability of the tokens utilized in the reasoning process can be made more prominent, so that the retrieval model outputs an answer that is more context - compliant through dynamic context decoding.
[0102] S19. Controlling the retrieval model to output a retrieval result according to the target output probability of each target token.
[0103] In this embodiment, through the dynamic change of the internal attention weights of the retrieval model, the tokens being utilized by the model are identified in real - time during the reasoning process, and the output probabilities of these tokens are dynamically enhanced, so that a high - quality answer that is more context - compliant can be generated.
[0104] Specifically, this embodiment mainly has the following advantages:
[0105] (1) Stronger interpretability: By training a small white - box model (i.e., the logistic regression classifier) in this embodiment, the keyword tokens utilized during the generation process can be effectively detected. Moreover, with the help of the feature coefficients of the logistic regression classifier, the features that play a core role in the generation result can be identified and quantified, thereby achieving a higher degree of explanation for the retrieval process. This method significantly improves the transparency of the system and facilitates in - depth understanding of the generation mechanism.
[0106] (2) Saving computing resources: Compared with traditional dual - decoding methods (such as CAD, etc.), this embodiment proposes an efficient reasoning method that uses a lightweight detector to identify the tokens being utilized in real - time during the reasoning process and dynamically enhance the probabilities of the outputs of these tokens. In this way, without multiple decodings, the usage cost of computing resources is significantly reduced. This embodiment not only improves the quality of the generated content but also shows obvious performance advantages in real - time processing scenarios and large - scale application scenarios.
[0107] (3) Reducing the "context hallucination" phenomenon: This embodiment introduces a reasoning intervention technique to ensure that the model gives priority to the input context information during the generation process, fundamentally reducing the occurrence of the "context hallucination" phenomenon. Especially in complex knowledge domains, this embodiment can effectively avoid the problem of the model generating inaccurate or incorrect content, thereby significantly improving the credibility and applicability of the generated results.
[0108] In summary, in this embodiment, by guiding the attention mechanism of the retrieval model to dynamically focus on the context information highly relevant to the user query, the context processing ability is optimized, thereby enhancing the relevance and accuracy of the generated content. It is not only applicable to fields with extremely high requirements for knowledge accuracy, such as finance and healthcare, but also widely applicable to application scenarios such as intelligent customer service and automated question-answering systems. By strengthening the model's ability in the context understanding and reasoning process, especially the accuracy in knowledge update and dynamic response, it can ensure the authenticity, reliability of the generated content and its conformity with the latest external knowledge, greatly improving the practical application value of the intelligent system.
[0109] It can be seen from the above technical solutions that the present invention constructs training data based on the relevance of each designated token to the retrieval question and the retrieval answer, as well as the attention feature vector of each designated token, and uses the training data to train a logistic regression classifier to obtain a detector for detecting whether a token is utilized, so as to be able to use a lightweight detector to identify the utilized tokens in real time during the inference process, thereby achieving a higher degree of interpretation of the generation process; when using the retrieval model for retrieval, the detector is used to detect in real time the tokens being utilized in the input data as each target token, correct the original probability distribution of each target token according to the utilization probability of each target token to obtain the target output probability of each target token, and control the retrieval model to output the retrieval result according to the target output probability of each target token, dynamically enhancing the output probability of the utilized tokens, thereby dynamically focusing the attention mechanism of the model on the context information highly relevant to the user query, and fundamentally reducing the occurrence of the "context hallucination" phenomenon.
[0110] As Figure 2 shown, it is a functional module diagram of a preferred embodiment of the retrieval device based on attention guidance of the present invention. The retrieval device 11 based on attention guidance includes an acquisition unit 110, a construction unit 111, a training unit 112, a detection unit 113, a calculation unit 114, a correction unit 115, and a control unit 116. The modules / units referred to in the present invention refer to a series of computer program segments that can be executed by a processor and can complete fixed functions, and are stored in a memory. In this embodiment, the functions of each module / unit will be described in detail in subsequent embodiments.
[0111] The acquisition unit 110 is configured to acquire a retrieval model and acquire the attention matrix corresponding to each attention head in the retrieval model.
[0112] In this embodiment, the retrieval model may be an artificial intelligence model with a question-answering function such as Retrieval-Augmented Language Models (RALMs).
[0113] In this embodiment, the retrieval model may be based on a multi-head attention mechanism, and each attention head corresponds to an attention matrix.
[0114] The construction unit 111 is configured to construct an attention feature vector for each specified token according to each attention matrix.
[0115] In this embodiment, the construction unit 111 constructing an attention feature vector for each specified token according to each attention matrix includes:
[0116] Extracting the elements corresponding to each specified token from the elements of the last row of each attention matrix as the attention weight of each specified token under the corresponding attention head;
[0117] Obtaining the attention weights of each specified token under each attention head to construct an attention feature vector for each specified token.
[0118] Among them, the attention weight of each specified token under the corresponding attention head is used to reflect the attention degree of each attention head relative to each specified token.
[0119] Among them, the specified token may be any token in the external knowledge cited from an external database. Such as keywords in the field of medical and health, such as "patient", "medicine", etc., and "bank", "amount", etc. in the financial field.
[0120] For example: for any token, if there are three attention matrices, the attention weights corresponding to this token in the three attention matrices are 7, 8, and 9 respectively, then the attention feature vector corresponding to this token is (7, 8, 9).
[0121] In the above embodiment, since the last row of the attention matrix in the attention mechanism represents the degree of attention of the model to the context in the current inference step, therefore, extracting the elements corresponding to each specified token from the elements of the last row of each attention matrix as the attention weight of each specified token under the corresponding attention head enables the subsequent model inference to fully consider the context of the model.
[0122] The acquisition unit 110 is further configured to acquire the relevance between each pre-labeled specified token and the retrieval question and the retrieval answer.
[0123] In this embodiment, each token may be labeled according to the influence degree of each token on the retrieval answer during actual retrieval.
[0124] For example: if the influence degree is high, the corresponding token is labeled as 1, indicating that the token is utilized during retrieval; if the influence degree is low, the corresponding token is labeled as 0, indicating that the token is not utilized during retrieval.
[0125] The building unit 111 is further configured to construct training data according to the relevance of each specified token to the retrieval question and the retrieval answer, and the attention feature vector of each specified token.
[0126] In this embodiment, the building unit 111 constructs training data according to the relevance of each specified token to the retrieval question and the retrieval answer, and the attention feature vector of each specified token, including:
[0127] When there is a first specified token in each specified token that is relevant to the retrieval question and the retrieval answer, the attention feature vector of the first specified token is determined as a positive sample;
[0128] When there is a second specified token in each specified token that is not relevant to the retrieval question and the retrieval answer, the attention feature vector of the second specified token is determined as a negative sample;
[0129] The positive samples and the negative samples are combined to obtain the training data.
[0130] Wherein, the ratio of the positive samples to the negative samples can be 1:1.
[0131] In the above embodiment, based on the constructed positive samples and negative samples, the detector obtained by subsequent training can accurately identify the tokens that play an important role in the output of the retrieval model, thereby improving the accuracy of the output of the retrieval model.
[0132] The training unit 112 is configured to train a logistic regression classifier using the training data to obtain a detector for detecting whether a token is utilized.
[0133] In this embodiment, the output of the detector may include 1 or 0.
[0134] Specifically, when the output of the detector is 1, it indicates that the corresponding token is being utilized; when the output of the detector is 0, it indicates that the corresponding token is not being utilized.
[0135] The obtaining unit 110 is further configured to obtain the input data of the retrieval model when performing a retrieval using the retrieval model.
[0136] In this embodiment, the input data may include data cited from an external knowledge base, and the user's question content, etc.
[0137] The detection unit 113 is configured to use the detector to detect in real time the tokens being utilized in the input data as each target token.
[0138] In this embodiment, the detection unit 113 uses the detector to detect in real time the tokens being utilized in the input data as each target token, including:
[0139] After inputting the input data into the retrieval model, the attention matrices corresponding to each attention head in the retrieval model are obtained in real time to construct the attention feature vectors of each token in the input data;
[0140] The attention feature vectors of each token are input into the detector;
[0141] When the output of the detector corresponding to the first token is the first output data, it is determined that the first token is being utilized, and the first token is determined as the target token.
[0142] Wherein, the first output data can be customized. For example: the first output data can be 1.
[0143] The calculation unit 114 is configured to obtain the feature coefficients of the detector and calculate the utilization probability of each target token using the feature coefficients.
[0144] In this embodiment, the feature coefficients of the detector are parameters obtained through model training, used to reflect the influence direction and degree of each feature on the target variable, and to evaluate the contributions of different attention heads during the retrieval process.
[0145] Among them, in terms of the influence direction, the positive or negative of the feature coefficient can reflect the relationship between the feature and the target variable. A positive coefficient indicates that as the feature value increases, the probability of the target variable taking the positive class (such as 1) increases; a negative coefficient means that as the feature value increases, the probability of the positive class decreases. For example: in the financial field, when predicting whether to purchase a product, a positive coefficient of the "income" feature indicates that the higher the income, the greater the purchase possibility; a negative coefficient of the "product negative review rate" indicates that the higher the negative review rate, the smaller the purchase possibility.
[0146] Among them, in terms of the influence degree, the absolute value of the feature coefficient reflects the influence degree. The larger the absolute value, the more significant the influence of the feature on the target variable. For example: in the medical and health field, when judging whether a patient is ill, if the absolute value of the feature coefficient of a "certain key indicator" is large, it indicates that this indicator plays a key role in the illness judgment. However, the dimensions of different features may affect the absolute value of the coefficient, and usually, the data needs to be standardized before comparison.
[0147] In this embodiment, the calculation unit 114 calculates the utilization probability of each target token using the feature coefficients, including:
[0148] Obtaining in real time the attention weights of each target token under each attention head;
[0149] Calculate the comprehensive utilization score of each target token according to the attention weight of each target token under each attention head and the feature coefficient;
[0150] Calculate the cumulative sum of the comprehensive utilization scores of each target token;
[0151] Calculate the quotient of the comprehensive utilization score of each target token and the cumulative sum to obtain the utilization probability of each target token.
[0152] Specifically, the calculating the comprehensive utilization score of each target token according to the attention weight of each target token under each attention head and the feature coefficient includes:
[0153] Use the following formula to calculate the feature coefficient corresponding to each attention head after normalization:
[0154]
[0155] where, w i represents the feature coefficient corresponding to the i-th attention head after normalization; c i represents the original feature coefficient corresponding to the i-th attention head; i is a positive integer; I represents the total number of attention heads;
[0156] Use the following formula to calculate the comprehensive utilization score of each target token according to the feature coefficient corresponding to each attention head after normalization and the attention weight of each target token under each attention head:
[0157]
[0158] where, s j represents the comprehensive utilization score of the j-th target token; a ij represents the attention weight of the j-th target token under the i-th attention head; when f(v j ) = 1, it means that the j-th target token v j is being utilized; when f(v j ) = 0, it means that the j-th target token v j is not being utilized; j is a positive integer.
[0159] In the above embodiment, since it is difficult to determine the specific utilization degree of each token by the retrieval model only through the output of the detector, therefore, based on the weighted method of the detector feature coefficient, the attention weights are weighted and summed, and a comprehensive utilization score can be calculated for each token. This comprehensive utilization score can quantify the utilization degree of the tokens by the retrieval model and be used to dynamically adjust the model output process.
[0160] The correction unit 115 is configured to correct the original probability distribution of each target token according to the utilization probability of each target token, so as to obtain the target output probability of each target token.
[0161] In this embodiment, the correction unit 115 corrects the original probability distribution of each target token according to the utilization probability of each target token, and obtaining the target output probability of each target token includes:
[0162] Obtain the original output probability of each target token from the original probability distribution;
[0163] Calculate the sum of the utilization probability of each target token and the corresponding original output probability to obtain the target output probability of each target token.
[0164] Through the above embodiments, the probability of the tokens utilized in the inference process can be made more prominent, so that the retrieval model outputs an answer that is more context - compliant through dynamic context decoding.
[0165] The control unit 116 is configured to control the retrieval model to output a retrieval result according to the target output probability of each target token.
[0166] In this embodiment, through the dynamic change of the internal attention weights of the retrieval model, the tokens being utilized by the model are identified in real - time during the inference process, and the output probabilities of these tokens are dynamically enhanced, so as to be able to generate a high - quality answer that is more context - compliant.
[0167] Specifically, this embodiment mainly has the following advantages:
[0168] (1) Stronger interpretability: In this embodiment, by training a small white - box model (i.e., the logistic regression classifier), the keyword tokens utilized during the generation process can be effectively detected. Moreover, with the help of the feature coefficients of the logistic regression classifier, the features that play a core role in the generation result can be identified and quantified, thereby achieving a higher degree of interpretation of the retrieval process. This method significantly improves the transparency of the system and facilitates in - depth understanding of the generation mechanism.
[0169] (2) Saving computing resources: Compared with traditional dual - decoding methods (such as CAD, etc.), this embodiment proposes an efficient inference method that uses a lightweight detector to identify the tokens being utilized in real - time during the inference process and dynamically enhances the probabilities of the outputs of these tokens. In this way, without multiple decodings, the usage cost of computing resources is significantly reduced. This embodiment not only improves the quality of the generated content but also shows obvious performance advantages in real - time processing scenarios and large - scale application scenarios.
[0170] (3) Reducing the phenomenon of "context hallucination": In this embodiment, an inference intervention technique is introduced to ensure that the model gives priority to the input context information during the generation process, fundamentally reducing the occurrence of the "context hallucination" phenomenon. Especially in complex knowledge domains, this embodiment can effectively avoid the problem of the model generating inaccurate or incorrect content, thus significantly improving the credibility and applicability of the generated results.
[0171] In summary, in this embodiment, by guiding the attention mechanism of the retrieval model to dynamically focus on the context information highly relevant to the user query and optimizing the context processing ability, the relevance and accuracy of the generated content are enhanced. It is not only applicable to fields with extremely high requirements for knowledge accuracy, such as finance and healthcare, but also widely applicable to application scenarios such as intelligent customer service and automated question-and-answer systems. By strengthening the model's ability in the context understanding and reasoning process, especially the accuracy in knowledge update and dynamic response, it can ensure the authenticity, reliability, and conformity with the latest external knowledge of the generated content, greatly enhancing the practical application value of the intelligent system.
[0172] As can be seen from the above technical solutions, the present invention constructs training data according to the relevance of each specified token to the retrieval question and the retrieval answer, as well as the attention feature vector of each specified token, and uses the training data to train a logistic regression classifier to obtain a detector for detecting whether a token is utilized, which can utilize a lightweight detector to identify the utilized tokens in real time during the inference process, thereby achieving a higher degree of interpretation of the generation process; when using the retrieval model for retrieval, the detector is used to detect in real time the tokens being utilized in the input data as each target token, correct the original probability distribution of each target token according to the utilization probability of each target token to obtain the target output probability of each target token, and control the retrieval model to output the retrieval result according to the target output probability of each target token, dynamically enhancing the output probability of the utilized tokens, thereby fundamentally reducing the occurrence of the "context hallucination" phenomenon by guiding the attention mechanism of the model to dynamically focus on the context information highly relevant to the user query.
[0173] As Figure 3 shown, it is a schematic structural diagram of a computer device of a preferred embodiment for implementing the retrieval method based on attention guidance of the present invention.
[0174] The computer device 1 may include a memory 12, a processor 13, and a bus, and may also include a computer program stored in the memory 12 and executable on the processor 13, such as a retrieval program based on attention guidance.
[0175] Those skilled in the art can understand that the schematic diagram is only an example of the computer device 1, and does not constitute a limitation on the computer device 1. The computer device 1 can be either a bus structure or a star structure. The computer device 1 can also include more or fewer other hardware or software than shown in the figure, or different component arrangements. For example, the computer device 1 can also include input / output devices, network access devices, etc.
[0176] It should be noted that the computer device 1 is only an example. Other existing or future electronic products that can be adapted to the present invention should also be included within the protection scope of the present invention and are hereby incorporated by reference.
[0177] Among them, the memory 12 includes at least one type of readable storage medium. The readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disk, etc. The memory 12 can be an internal storage unit of the computer device 1 in some embodiments, such as the mobile hard disk of the computer device 1. The memory 12 can also be an external storage device of the computer device 1 in other embodiments, such as a plug-in mobile hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the computer device 1. Further, the memory 12 can also include both the internal storage unit and the external storage device of the computer device 1. The memory 12 can be used not only to store the application software installed on the computer device 1 and various types of data, such as the code of the retrieval program based on attention guidance, etc., but also to temporarily store the data that has been output or will be output.
[0178] The processor 13 can be composed of integrated circuits in some embodiments. For example, it can be composed of a single packaged integrated circuit, or can be composed of multiple integrated circuits with the same or different functions packaged together, including the combination of one or more Central Processing Units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips, etc. The processor 13 is the control core (Control Unit) of the computer device 1, connecting various components of the entire computer device 1 through various interfaces and circuits. By running or executing the programs or modules stored in the memory 12 (such as executing the retrieval program based on attention guidance, etc.), and calling the data stored in the memory 12, it can execute various functions of the computer device 1 and process data.
[0179] The processor 13 executes the operating system of the computer device 1 and various installed application programs. The processor 13 executes the application programs to implement the steps in the above-mentioned various embodiments of the attention-guided retrieval method, such as Figure 1 the steps shown.
[0180] Exemplarily, the computer program may be divided into one or more modules / units. The one or more modules / units are stored in the memory 12 and executed by the processor 13 to complete the present invention. The one or more modules / units may be a series of computer-readable instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the computer device 1. For example, the computer program may be divided into an acquisition unit 110, a construction unit 111, a training unit 112, a detection unit 113, a calculation unit 114, a correction unit 115, and a control unit 116.
[0181] The integrated units implemented in the form of software function modules can be stored in a computer-readable storage medium. The above-mentioned software function modules stored in a storage medium include several instructions for causing a computer device (which may be a personal computer, a computer device, or a network device, etc.) or a processor to execute a part of the attention-guided retrieval method described in various embodiments of the present invention.
[0182] If the integrated module / unit of the computer device 1 is implemented in the form of a software function unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present invention, it can also be completed by a computer program instructing relevant hardware devices. The computer program can be stored in a computer-readable storage medium, and when the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented.
[0183] Among them, the computer program includes computer program code, and the computer program code may be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM, Read-Only Memory), a random access memory, etc.
[0184] Further, the computer-readable storage medium may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function, etc.; the data storage area may store data created according to the use of the blockchain node, etc.
[0185] The blockchain referred to in the present invention is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm. Blockchain, in essence, is a decentralized database, a string of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of the information (anti-counterfeiting) and generate the next block. The blockchain may include a blockchain underlying platform, a platform product service layer, an application service layer, etc.
[0186] The bus may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. This bus may be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, in Figure 3 it is only represented by a straight line, but it does not mean that there is only one bus or one type of bus. The bus is arranged to realize the connection and communication between the memory 12 and at least one processor 13, etc.
[0187] Although not shown, the computer device 1 may further include a power supply (such as a battery) for supplying power to each component. Preferably, the power supply may be logically connected to the at least one processor 13 through a power management device, so as to realize functions such as charging management, discharging management, and power consumption management through the power management device. The power supply may further include any components such as one or more DC or AC power supplies, a recharge device, a power failure detection circuit, a power converter or an inverter, and a power status indicator. The computer device 1 may further include a variety of sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.
[0188] Further, the computer device 1 may further include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is usually used to establish a communication connection between the computer device 1 and other computer devices.
[0189] Optionally, the computer device 1 may further include a user interface, which may be a display, an input unit (such as a keyboard), and optionally, the user interface may also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, which is used to display the information processed in the computer device 1 and to display a visual user interface.
[0190] It should be understood that the above embodiments are only for illustrative purposes and are not limited by this structure in the scope of the patent application.
[0191] Those skilled in the art can understand that Figure 3 the structure shown does not constitute a limitation on the computer device 1, and it may include fewer or more components than shown in the figure, or combine some components, or have a different component arrangement.
[0192] In combination with Figure 1 , the memory 12 in the computer device 1 stores a plurality of instructions to implement a retrieval method based on attention guidance, and the processor 13 can execute the plurality of instructions to implement:
[0193] Obtain a retrieval model and obtain the attention matrix corresponding to each attention head in the retrieval model;
[0194] Construct an attention feature vector for each specified token according to each attention matrix;
[0195] Obtain the relevance of each pre-labeled specified token to the retrieval question and the retrieval answer;
[0196] Construct training data according to the relevance of each specified token to the retrieval question and the retrieval answer, and the attention feature vector of each specified token;
[0197] Use the training data to train a logistic regression classifier to obtain a detector for detecting whether a token is being utilized;
[0198] When performing retrieval using the retrieval model, obtain the input data of the retrieval model;
[0199] Use the detector to real-time detect the tokens being utilized in the input data as each target token;
[0200] Obtain the feature coefficients of the detector and calculate the utilization probability of each target token using the feature coefficients;
[0201] Modify the original probability distribution of each target token according to the utilization probability of each target token to obtain the target output probability of each target token;
[0202] Control the retrieval model to output a retrieval result according to the target output probability of each target token.
[0203] Specifically, for the specific implementation method of the above instructions by the processor 13, reference can be made to Figure 1 the description of the relevant steps in the corresponding embodiments, which will not be elaborated here.
[0204] It should be noted that all the data involved in this case are legally obtained. The non-company software tools or components that appear in the embodiments of this application are only for illustrative introduction and do not represent actual use.
[0205] In several embodiments provided by the present invention, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation.
[0206] The present invention can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, small computers, large computers, distributed computing environments including any of the above systems or devices, and so on. The present invention can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0207] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0208] In addition, in each embodiment of the present invention, each functional module can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware, or in the form of a combination of hardware and software functional modules.
[0209] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms.
[0210] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be regarded as limiting the claims involved.
[0211] In addition, it is obvious that the term "including" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units or devices described in the present invention can also be implemented by one unit or device through software or hardware. The terms such as "first" and "second" are used to represent names and do not represent any specific order.
[0212] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A retrieval method based on attention guidance, characterized in that: The attention-guided retrieval method comprises: Obtain a retrieval model, and obtain an attention matrix corresponding to each attention head in the retrieval model; Construct an attention feature vector for each specified word based on each attention matrix; Obtaining the relevance of each pre-marked specified word to the search question and the search answer; Construct training data according to the relevance of each specified word unit to the retrieval question and the retrieval answer, and the attention feature vector of each specified word unit; Using the training data to train a logistic regression classifier to obtain a detector for detecting whether a word unit is used; When searching using the retrieval model, obtaining input data of the retrieval model; Using the detector to detect in real time the word element being used in the input data as each target word element; Acquire a characteristic coefficient of the detector, and use the characteristic coefficient to calculate the utilization probability of each target word; The original probability distribution of each target word is modified according to the utilization probability of each target word to obtain the target output probability of each target word; The retrieval model is controlled to output retrieval results according to the target output probability of each target word.
2. The attention-guided retrieval method according to claim 1, characterized in that: The step of constructing the attention feature vector of each specified word according to each attention matrix includes: Extract the elements corresponding to each specified word from the elements in the last row of each attention matrix as the attention weight of each specified word under the corresponding attention head; Get the attention weight of each specified word under each attention head to construct the attention feature vector of each specified word.
3. The attention-guided retrieval method according to claim 1, characterized in that: The step of constructing training data based on the relevance of each designated word unit to the search question and the search answer, and the attention feature vector of each designated word unit includes: When a first designated word-element in each designated word-element is relevant to the search question and the search answer, determining the attention feature vector of the first designated word-element as a positive sample; When a second designated word-unit is not relevant to the retrieval question and the retrieval answer in each designated word-unit, determining the attention feature vector of the second designated word-unit as a negative sample; The positive samples and the negative samples are combined to obtain the training data.
4. The attention-guided retrieval method according to claim 1, characterized in that: The step of using the detector to detect in real time the word element being used in the input data as each target word element comprises: After the input data is input into the retrieval model, an attention matrix corresponding to each attention head in the retrieval model is obtained in real time to construct an attention feature vector for each word in the input data; Inputting the attention feature vector of each word unit into the detector; When the detector output corresponding to the first word-gram is the first output data, it is determined that the first word-gram is being used, and the first word-gram is determined to be the target word-gram.
5. The attention-guided retrieval method according to claim 1, characterized in that: The step of calculating the utilization probability of each target word by using the characteristic coefficient includes: Obtain the attention weight of each target word under each attention head in real time; Calculate the comprehensive utilization score of each target word according to the attention weight of each target word under each attention head and the feature coefficient; Calculate the cumulative sum of the comprehensive utilization scores of each target word; The quotient of the comprehensive utilization score of each target word and the accumulated sum is calculated to obtain the utilization probability of each target word.
6. The attention-guided retrieval method according to claim 5, characterized in that: The calculating of the comprehensive utilization score of each target word according to the attention weight of each target word under each attention head and the characteristic coefficient includes: The following formula is used to calculate the normalized feature coefficient of each attention head: Among them, w i represents the feature coefficient corresponding to the i-th attention head after normalization; c i represents the original feature coefficient corresponding to the i-th attention head; i is a positive integer; I represents the total number of attention heads; The following formula is used to calculate the comprehensive utilization score of each target word based on the normalized feature coefficient corresponding to each attention head and the attention weight of each target word under each attention head: Among them, s j represents the comprehensive utilization score of the jth target word; a ij represents the attention weight of the j-th target word under the i-th attention head; when f(v j )=1, it means the jth target word v j is being used; when f(v j )=0, it means the jth target word v j Not used; j is a positive integer.
7. The attention-guided retrieval method according to claim 1, characterized in that: The step of modifying the original probability distribution of each target word unit according to the utilization probability of each target word unit to obtain the target output probability of each target word unit includes: Obtaining the original output probability of each target word from the original probability distribution; The sum of the utilization probability of each target word and the corresponding original output probability is calculated to obtain the target output probability of each target word.
8. A retrieval device based on attention guidance, characterized in that: The attention-guided retrieval device comprises: An acquisition unit, used to acquire a retrieval model and an attention matrix corresponding to each attention head in the retrieval model; A construction unit, used to construct an attention feature vector for each specified word according to each attention matrix; The acquisition unit is further used to acquire the relevance of each pre-marked designated word element with the search question and the search answer; The construction unit is further used to construct training data according to the relevance of each designated word unit to the search question and the search answer, and the attention feature vector of each designated word unit; A training unit, used to train a logistic regression classifier using the training data to obtain a detector for detecting whether a word unit is used; The acquisition unit is further used to acquire input data of the retrieval model when the retrieval model is used for retrieval; A detection unit, used to detect in real time using the detector the word element being used in the input data as each target word element; A calculation unit, used to obtain a characteristic coefficient of the detector and calculate the utilization probability of each target word using the characteristic coefficient; A correction unit, used to correct the original probability distribution of each target word unit according to the utilization probability of each target word unit, so as to obtain the target output probability of each target word unit; A control unit is used to control the retrieval model to output retrieval results according to the target output probability of each target word.
9. A computer device, characterized in that: The computer device comprises: a memory storing at least one instruction; and A processor executes instructions stored in the memory to implement the attention-guided retrieval method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one instruction, and the at least one instruction is executed by a processor in a computer device to implement the attention-guided retrieval method as described in any one of claims 1 to 7.
Citation Information
Cited By
Dynamic retrieval enhancement generation method and system based on attention attribution
CN121029954A
Data processing method and device based on artificial intelligence, computer equipment and medium
CN121233745A
Data processing methods, devices, computer equipment, and media based on artificial intelligence
CN121233745B