Attention bias matrix injection method for multi-agent cognitive alignment
Patent Information
- Application Number
- CN202610603349.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-30
- Publication Date
- 2026-09-11
AI Technical Summary
现有LLM智能体多通过提示或对话协作,却易在高噪声长上下文下产生显著性偏差,导致协作失效
[0021] Fifthly, this application proposes a program product comprising at least one of a program and instructions, wherein when the program and instructions are executed by an electronic device, they implement the steps of the method described in the first aspect.
Smart Images

Figure CN122735892A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of natural language processing and large language model technology, and in particular to an attention bias matrix injection method, apparatus, device and storage medium for multi-agent cognitive alignment. Background Technology
[0002] In related technologies, system logs and alarm streams in scenarios such as mine safety monitoring and industrial control are characterized by high significant noise and low significant truth values. Existing LLM agents mostly collaborate through prompts or dialogue, but they are prone to significant biases in high-noise, long-term contexts, leading to collaboration failure. Fine-tuning models are costly, slow to iterate, have poor adaptability, and are also subject to data and compliance constraints. Summary of the Invention
[0003] This application aims to at least partially address one of the technical problems in the related art.
[0004] In a first aspect, this application proposes an attention bias matrix injection method for multi-agent cognitive alignment. The method includes: obtaining the task rule context of a model inference task; generating a cognitive concept vector based on the task rule context; obtaining the embedding vector corresponding to the input text lexical; obtaining an activation similarity score based on the cognitive concept vector and the embedding vector; obtaining an additive bias corresponding to the lexical based on the activation similarity; generating an attention mask matrix based on the additive bias; and injecting the attention mask matrix into the model's self-attention calculation formula.
[0005] In one implementation, generating a cognitive concept vector based on the task rule context includes: obtaining multiple sets of cognitive keywords based on the task rule context; for each keyword in the cognitive keyword set, splitting the keyword into corresponding keyword units; performing vector mapping on the keyword units to obtain a keyword vector; and for each keyword set, aggregating the average of the keyword vectors corresponding to each keyword in the keyword set to obtain the corresponding cognitive concept vector.
[0006] In one implementation, obtaining the activation similarity based on the cognitive concept vector and the embedding vector includes: obtaining the vector similarity between the embedding vector and the cognitive concept vector; and performing similarity-gated activation based on the vector similarity to obtain the activation similarity.
[0007] In one implementation, the activation similarity score includes an amplified concept similarity score and a suppressed concept similarity score, and the additive bias corresponding to the lexical unit is obtained based on the activation similarity using the following formula:
[0008]
[0009] in, The additive bias for the i-th lexical unit, To amplify the intensity parameter, To suppress the intensity parameter, The similarity score of the amplified concept in the activation similarity is... Let be the inhibition concept similarity score between the i-th lexical unit and the cognitive concept vector.
[0010] In one implementation, generating the attention mask matrix based on the additive bias includes: injecting the additive bias into the attention matrix used for model inference to obtain a cognitive bias vector; obtaining a cognitive bias matrix based on the cognitive bias vector; and fusing the cognitive bias matrix with a causal mask to obtain the attention mask matrix.
[0011] In one implementation, the task rule context includes at least one of the following: task objective, rule text, and context history memory.
[0012] Secondly, this application proposes an attention bias matrix injection device for multi-agent cognitive alignment. The device includes: an acquisition module for acquiring the task rule context of a model inference task; a first processing module for generating a cognitive concept vector based on the task rule context; a second processing module for acquiring the embedding vectors corresponding to input text lexical units; a third processing module for acquiring an activation similarity score based on the cognitive concept vector and the embedding vector; a fourth processing module for acquiring the additive bias corresponding to the lexical unit based on the activation similarity; a fifth processing module for generating an attention mask matrix based on the additive bias; and a sixth processing module for injecting the attention mask matrix into the model's self-attention calculation formula.
[0013] In one implementation, the first processing module can be used to: obtain multiple sets of cognitive keywords based on the task rule context; for each keyword in the set of cognitive keywords, split the keyword into corresponding keyword units; perform vector mapping on the keyword units to obtain keyword vectors; and for each set of keywords, perform mean aggregation on the keyword vectors corresponding to each keyword in the set of keywords to obtain the corresponding cognitive concept vector.
[0014] In one implementation, the third processing module can be used to: obtain the vector similarity between the embedded vector and the cognitive concept vector; perform similarity-gated activation based on the vector similarity to obtain the activation similarity.
[0015] In one implementation, the activation similarity score includes an amplified concept similarity score and a suppressed concept similarity score, and the additive bias corresponding to the lexical unit is obtained based on the activation similarity using the following formula:
[0016] in, The additive bias for the i-th lexical unit, To amplify the intensity parameter, To suppress the intensity parameter, The similarity score of the amplified concept in the activation similarity is... Let be the inhibition concept similarity score between the i-th lexical unit and the cognitive concept vector.
[0017] In one implementation, the fifth processing module can be used to: inject the additive bias into the attention matrix used for model inference to obtain a cognitive bias vector; obtain a cognitive bias matrix based on the cognitive bias vector; and fuse the cognitive bias matrix with a causal mask to obtain the attention mask matrix.
[0018] In one implementation, the task rule context includes at least one of the following: task objective, rule text, and context history memory.
[0019] Thirdly, this application proposes an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the attention bias matrix injection method for multi-agent cognitive alignment as described in the first aspect.
[0020] Fourthly, this application proposes a storage medium storing instructions that, when executed on an electronic device, cause the electronic device to perform the method described in the first aspect.
[0021] Fifthly, this application proposes a program product comprising at least one of a program and instructions, wherein when the program and instructions are executed by an electronic device, they implement the steps of the method described in the first aspect.
[0022] The attention bias matrix injection method, apparatus, device, and storage medium provided in this application for multi-agent cognitive alignment can generate cognitive concept vectors based on the task rule context of the model reasoning task, calculate activation similarity scores by combining the word embedding vectors of the input text, obtain the additive bias corresponding to the word, and generate an attention mask matrix. Finally, the mask is injected into the model's self-attention calculation formula for reasoning intervention. It can achieve noise suppression and key information amplification, significantly improving the controllability and accuracy of model reasoning.
[0023] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description
[0024] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1 This is a flowchart illustrating an attention bias matrix injection method for multi-agent cognitive alignment provided in an embodiment of this application. Figure 2 This is a flowchart illustrating another attention bias matrix injection method for multi-agent cognitive alignment provided in this application embodiment; Figure 3 This is a schematic diagram of an attention bias matrix injection scheme for multi-agent cognitive alignment provided in an embodiment of this application; Figure 4 This is a schematic diagram illustrating the comparison of target token probabilities under baseline and intervention conditions, provided in an embodiment of this application. Figure 5 This is a schematic diagram of an attention bias matrix injection device for multi-agent cognitive alignment provided in an embodiment of this application; Figure 6 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0025] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.
[0026] The attention bias matrix injection method and apparatus for multi-agent cognitive alignment according to embodiments of this application are described below with reference to the accompanying drawings.
[0027] Figure 1This is a flowchart illustrating an attention bias matrix injection method for multi-agent cognitive alignment provided in an embodiment of this application. Figure 1 As shown, the method may include, but is not limited to, the following steps: S110: Obtain the task rule context for the model inference task.
[0028] In the embodiments of this application, the task rule context includes at least one of the following: task objective, rule text, and context history memory.
[0029] S120: Generate cognitive concept vectors based on task rule context.
[0030] For example, keywords are obtained by parsing the task rule context, and cognitive concept vectors are obtained by vector mapping based on the keywords.
[0031] In some embodiments, see Figure 2 , Figure 2 This is a flowchart illustrating another attention bias matrix injection method for multi-agent cognitive alignment provided in an embodiment of this application. Figure 2 As shown, step S120 may include the following steps: S1201: Obtain multiple sets of cognitive keywords based on the context of task rules.
[0032] In the embodiments of this application, the aforementioned keywords may include amplifying keywords and suppressing keywords.
[0033] For example, the set of amplified keywords and the set of suppressed keywords are extracted and divided from the context of the task rules.
[0034] S1202: For each set of cognitive keywords, the keywords are broken down into corresponding keyword units.
[0035] For example, each keyword in the amplified keyword set and the suppressed keyword set is segmented into tokens that the model can process, thus obtaining the tokens corresponding to each keyword.
[0036] S1203: Perform vector mapping on keyword lexical units to obtain keyword vectors.
[0037] For example, the keyword tokens obtained from the splitting are fed into the model's embedding layer, and each token is converted into a corresponding embedding vector through an embedding lookup table; then, the average of all token vectors for the same keyword is aggregated to obtain the keyword vector for that keyword.
[0038] S1204: For each keyword set, the keyword vectors corresponding to each keyword in the keyword set are averaged to obtain the corresponding cognitive concept vector.
[0039] For example, for each set of cognitive keywords, the average of the keyword vectors corresponding to all keywords in the set is applied to obtain the cognitive concept vector corresponding to that set.
[0040] S130: Obtain the embedding vectors corresponding to the words in the input text.
[0041] For example, the input text is segmented into tokens by a tokenizer, and then the resulting tokens are fed into the model's embedding layer. We perform an embedding lookup operation to obtain the embedding vector corresponding to each input text word.
[0042] S140: Obtain activation similarity scores based on cognitive concept vectors and embedding vectors.
[0043] In one implementation, activation similarity is obtained based on the cognitive concept vector and the embedding vector, including: obtaining the vector similarity between the embedding vector and the cognitive concept vector; and performing similarity-gated activation based on the vector similarity to obtain activation similarity.
[0044] For example, for the embedding vector of each input word, the cosine similarity is calculated with the amplified cognitive concept vector and the suppressed cognitive concept vector respectively to obtain the original similarity score; then, a preset similarity threshold is used for filtering, and only positive scores exceeding the threshold are retained to obtain the activation similarity.
[0045] As an example, the activated amplified concept vector can be represented as:
[0046] in, To amplify the concept cosine similarity score of the i-th lexical after activation, Let be the original cosine similarity between the embedding vector of the i-th text word and the amplified concept vector. This is the similarity activation threshold.
[0047] As an example, the activated inhibition concept vector can be represented as:
[0048] in, Let cosine similarity score be the sum of the activation of the i-th lexical unit and the suppressed concept vector. Let be the original cosine similarity between the embedding vector of the i-th text word and the suppressed concept vector. This is the similarity activation threshold.
[0049] S150: Obtain the additive bias corresponding to the lexical unit based on activation similarity.
[0050] In one implementation, the activation similarity score includes an amplified concept similarity score and a suppressed concept similarity score, and the additive bias corresponding to the lexical unit is obtained based on the activation similarity using the following formula:
[0051] in, The additive bias for the i-th lexical unit. To amplify the intensity parameter, To suppress the intensity parameter, To activate the amplified concept similarity score in the similarity analysis, Let be the inhibition concept similarity score between the i-th lexical unit and the cognitive concept vector.
[0052] S160: Generate attention mask matrix based on additive bias.
[0053] In one implementation, the above-mentioned generation of the attention mask matrix based on additive bias may include the following steps: injecting the additive bias into the attention matrix used for model inference to obtain a cognitive bias vector; obtaining a cognitive bias matrix based on the cognitive bias vector; and fusing the cognitive bias matrix with a causal mask to obtain the attention mask matrix.
[0054] For example, the additive bias of each lexical is calculated, and the bias is injected into the target row (e.g., the last row) of the attention matrix to obtain the cognitive bias vector; the cognitive bias vector is expanded into a cognitive bias matrix of the same dimension as the causal mask by padding and broadcasting; the cognitive bias matrix and the causal mask matrix are matrix-added to obtain the cognitively aligned attention mask matrix.
[0055] As an example, the additive bias for each lexical can be calculated using the following formula:
[0056] in, The additive bias corresponding to the i-th input term. To amplify the intensity hyperparameter, The amplified similarity score after activating the i-th word. To suppress intensity hyperparameters, denoted as the inhibition similarity score after activation of the i-th lexical unit.
[0057] As an example, the attention mask matrix can be represented as follows:
[0058] S170: Inject the attention mask matrix into the model's self-attention calculation formula.
[0059] For example, the above steps can be represented as follows:
[0060] It should be noted that, in the embodiments of this application, the attention bias matrix injection method for multi-agent cognitive alignment provided in any embodiment of this application can be executed multiple times.
[0061] By implementing the embodiments of this application, cognitive concept vectors can be generated based on the task rule context of the model reasoning task. Activation similarity scores are then calculated by combining these vectors with the word embedding vectors of the input text, thereby obtaining the additive bias corresponding to the word and generating an attention mask matrix. Finally, the mask is injected into the model's self-attention calculation formula for reasoning intervention. This approach can achieve noise suppression and amplification of key information, significantly improving the controllability and accuracy of model reasoning.
[0062] As an example, please see Figure 3 , Figure 3 This is a schematic diagram of an attention bias matrix injection scheme for multi-agent cognitive alignment provided in an embodiment of this application, as shown below. Figure 3 As shown, Agent A first extracts suppression keywords and amplification keywords according to the task rules, synthesizes suppression concept vectors and amplification concept vectors respectively, and transmits them to Agent B; Agent B calculates the cosine similarity between the token sequence of the input text and the two types of concept vectors, constructs a cognitive bias vector after activation processing, and then fuses the standard causal mask to obtain a cognitive alignment attention mask. This mask is then injected into the model to complete inference, and finally outputs the result after cognitive intervention.
[0063] As an example, please see Figure 4 , Figure 4 This is a schematic diagram illustrating the comparison of target token probabilities under baseline and intervention conditions, provided in an embodiment of this application. Figure 4 As shown, the method of this application can effectively suppress noise tokens and amplify the generation probability of key signal tokens.
[0064] Please see Figure 5 , Figure 5 This is a schematic diagram of an attention bias matrix injection device for multi-agent cognitive alignment provided in an embodiment of this application. Figure 5As shown, the device 500 includes: an acquisition module 501 for acquiring the task rule context of the model inference task; a first processing module 502 for generating cognitive concept vectors based on the task rule context; a second processing module 503 for acquiring the embedding vectors corresponding to the input text lexical units; a third processing module 504 for acquiring activation similarity scores based on the cognitive concept vectors and embedding vectors; a fourth processing module 505 for acquiring the additive biases corresponding to the lexical units based on the activation similarity; a fifth processing module 506 for generating an attention mask matrix based on the additive biases; and a sixth processing module 507 for injecting the attention mask matrix into the model's self-attention calculation formula.
[0065] In one implementation, the first processing module 502 can be used to: obtain multiple sets of cognitive keywords based on the context of task rules; for each set of cognitive keywords, split the keywords into corresponding keyword units; perform vector mapping on the keyword units to obtain keyword vectors; and for each set of keywords, perform mean aggregation on the keyword vectors corresponding to each keyword in the set to obtain the corresponding cognitive concept vector.
[0066] In one implementation, the third processing module 504 can be used to: obtain the vector similarity between the embedded vector and the cognitive concept vector; perform similarity-gated activation based on the vector similarity to obtain the activation similarity.
[0067] In one implementation, the activation similarity score includes an amplified concept similarity score and a suppressed concept similarity score, and the additive bias corresponding to the lexical unit is obtained based on the activation similarity using the following formula:
[0068] in, The additive bias for the i-th lexical unit. To amplify the intensity parameter, To suppress the intensity parameter, To activate the amplified concept similarity score in the similarity analysis, Let be the inhibition concept similarity score between the i-th lexical unit and the cognitive concept vector.
[0069] In one implementation, the fifth processing module 506 can be used to: inject additive bias into the attention matrix used for model inference to obtain a cognitive bias vector; obtain a cognitive bias matrix based on the cognitive bias vector; and fuse the cognitive bias matrix with a causal mask to obtain an attention mask matrix.
[0070] In one implementation, the task rule context includes at least one of the following: task objective, rule text, and context history.
[0071] The apparatus of this application embodiment can generate cognitive concept vectors based on the task rule context of the model reasoning task, calculate activation similarity scores by combining them with the word embedding vectors of the input text, obtain additive biases corresponding to the words, and generate an attention mask matrix. Finally, the mask is injected into the model's self-attention calculation formula for reasoning intervention. This can achieve noise suppression and amplification of key information, significantly improving the controllability and accuracy of model reasoning.
[0072] It should be noted that the foregoing explanation of the embodiment of the attention bias matrix injection method for multi-agent cognitive alignment also applies to the attention bias matrix injection device for multi-agent cognitive alignment in this embodiment, and will not be repeated here.
[0073] To implement the above embodiments, this application also proposes an electronic device. Please see [link to relevant documentation]. Figure 6 , Figure 6 This is a schematic diagram of the structure of the electronic device provided in an embodiment of this application. For example... Figure 6 As shown, the electronic device 600 includes: a processor 601, and a memory 602 communicatively connected to the processor 601; the memory 602 stores computer-executable instructions; the processor 601 executes the computer-executable instructions stored in the memory to implement the method provided in the foregoing embodiments.
[0074] To implement the above embodiments, this application also proposes a storage medium storing instructions that, when executed on an electronic device, cause the electronic device to perform the methods provided in the foregoing embodiments.
[0075] To implement the above embodiments, this application also proposes a program product, including at least one of a program and instructions, wherein when the program and instructions are executed by an electronic device, they implement the steps of the method provided in the foregoing embodiments.
[0076] It should be noted that the acquisition, transmission, storage, use, and processing of data in this application comply with the relevant provisions of national laws and regulations and do not violate public order and good morals.
[0077] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, data stored, data displayed, etc.) and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0078] It is worth noting that in the embodiments of this application, certain software, components, models and other existing solutions in the industry may be mentioned. These should be regarded as exemplary and are only intended to illustrate the feasibility of implementing the technical solution of this application. However, they do not mean that the applicant has used or necessarily used the solution.
[0079] In the description of this application, unless otherwise stated, " / " means "or", for example, A / B can mean A or B; "and / or" in this document is merely a description of the relationship between related objects, indicating that there can be three relationships, for example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone.
[0080] In the foregoing descriptions of the embodiments, the terms "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0081] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0082] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.
[0083] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.
[0084] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0085] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.
[0086] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0087] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.
Claims
1. A method for injecting attention bias matrix for multi-agent cognitive alignment, characterized in that, include: Obtain the task rule context for the model inference task; Generate cognitive concept vectors based on the task rule context; Obtain the embedding vectors corresponding to the words in the input text; Based on the cognitive concept vector and the embedding vector, the activation similarity score is obtained; Based on the activation similarity, the additive bias corresponding to the lexical unit is obtained; An attention mask matrix is generated based on the additive bias; The attention mask matrix is injected into the model's self-attention calculation formula.
2. The method according to claim 1, characterized in that, The generation of cognitive concept vectors based on the task rule context includes: Multiple sets of cognitive keywords are obtained based on the task rule context; For each keyword in the cognitive keyword set, the keyword is broken down into corresponding keyword units; The keyword lexical units are vectorized to obtain the keyword vector; For each set of keywords, the keyword vectors corresponding to each keyword in the set are averaged to obtain the corresponding cognitive concept vector.
3. The method according to claim 1, characterized in that, The step of obtaining activation similarity based on the cognitive concept vector and the embedding vector includes: Obtain the vector similarity between the embedded vector and the cognitive concept vector; Similarity-gated activation is performed based on the vector similarity to obtain the activation similarity.
4. The method according to claim 1, characterized in that, The activation similarity score includes an amplified concept similarity score and a suppressed concept similarity score. The additive bias corresponding to the lexical unit is obtained based on the activation similarity using the following formula: in, The additive bias for the i-th lexical unit, To amplify the intensity parameter, To suppress the intensity parameter, The similarity score of the amplified concept in the activation similarity is... denoted as the inhibition concept similarity score between the i-th lexical unit and the cognitive concept vector.
5. The method according to claim 1, characterized in that, The generation of the attention mask matrix based on the additive bias includes: The additive bias is injected into the attention matrix used for model inference to obtain the cognitive bias vector; The cognitive bias matrix is obtained based on the cognitive bias vector; The cognitive bias matrix is fused with the causal mask to obtain the attention mask matrix.
6. The method according to claim 1, characterized in that, The task rule context includes at least one of the following: task objective, rule text, and context history.
7. An attention bias matrix injection device for multi-agent cognitive alignment, characterized in that, include: The acquisition module is used to acquire the task rule context for the model inference task; The first processing module is used to generate cognitive concept vectors based on the task rule context; The second processing module is used to obtain the embedding vectors corresponding to the words in the input text. The third processing module is used to obtain the activation similarity score based on the cognitive concept vector and the embedding vector; The fourth processing module is used to obtain the additive bias corresponding to the lexical based on the activation similarity; The fifth processing module is used to generate an attention mask matrix based on the additive bias; The sixth processing module is used to inject the attention mask matrix into the model's self-attention calculation formula.
8. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method as described in any one of claims 1 to 6.
9. A storage medium storing instructions, characterized in that, When the instructions are executed on an electronic device, the electronic device causes the electronic device to perform the method of any one of claims 1 to 6.
10. A program product comprising at least one of a program and instructions, characterized in that, When at least one of the program or instructions is executed by an electronic device, it implements the steps of the method according to any one of claims 1 to 6.