Large model generation content credible repair method based on representation engineering self-adaptive guidance
By constructing anchor vectors and an adaptive repair strategy library, the applicability and flexibility of representation engineering algorithms in the credibility repair of large language models are solved, achieving efficient and flexible credibility enhancement, improving the model's authenticity, fairness and security, while maintaining generality.
Patent Information
- Application Number
- CN202511658251.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-12
- Publication Date
- 2026-02-10
AI Technical Summary
Existing representation engineering algorithms lack a unified framework for the credibility repair of large language models, have varying applicability and insufficient flexibility, resulting in high repair costs and low efficiency. Furthermore, traditional methods rely on manually setting the intensity of intervention, making it difficult to achieve flexible and efficient credibility repair.
By integrating various representation engineering algorithms, an anchor vector and adaptive repair strategy library are constructed. The optimal algorithm and intervention intensity are selected based on activation differences and matching degrees to achieve adaptive repair and avoid damage to the overall model capability from intervention.
It achieves greater flexibility and credibility in the generation of content from large models, enhances realism, fairness and security, maintains versatility, has fast reasoning speed and strong adaptability, and supports the credibility requirements of different fields.
Smart Images

Figure CN121503558A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, specifically to a reliable content repair method for large model generation based on adaptive guidance of representation engineering. Background Technology
[0002] Large Language Models (LLMs) have fundamentally transformed the field of natural language processing, demonstrating unprecedented language understanding and generation capabilities, and have been widely applied in various scenarios. However, with the continued expansion of their deployment in critical domains, the inherent reliability issues of large language models (such as illusions in generated content, model decision bias, and the risk of system breaches) pose a significant obstacle to their secure use in high-risk areas such as healthcare, finance, and automation systems. The widespread and diverse nature of these issues necessitates efficient, reliable, and scalable reliability remediation strategies.
[0003] Current mainstream methods for credibility restoration of large language models face many challenges: On the one hand, post-training methods (such as supervised fine-tuning SFT and reinforcement learning based on human feedback RLHF) are highly dependent on intensive computing resources, which not only leads to high restoration costs but also slows down the iteration cycle; on the other hand, prompt engineering requires human effort to design and optimize prompt statements, and it also has the problems of limited generalization ability and robustness, which reduces its actual effectiveness in credibility restoration.
[0004] Currently, representation engineering is emerging as a promising paradigm for guiding the behavior of large-scale language models. This paradigm guides the behavior of LLMs by injecting target concept representations into the reasoning process. [1,2] Existing research suggests that representation engineering can alleviate hallucinations in large language models. [3,4] , remove bias [5,6] and enhanced security [7-9] It demonstrates effectiveness in improving credibility and shows potential for efficient, reliable and scalable adaptive repair.
[0005] Therefore, those skilled in the art are dedicated to developing a reliable content repair method for large model generation based on adaptive guidance of representation engineering.
[0006] References: [1]Zou A, Phan L, Chen S, et al. Representation Engineering: A Top-Down Approach to AI Transparency[A]. arXiv, 2025. [2]Turner A M, Thiergart L, Leech G, et al. Activation Addition:Steering Language Models Without Optimization[A]. arXiv, 2024. [3]Wang T, Jiao X, Zhu Y, et al. Adaptive Activation Steering: ATuning-Free LLM Truthfulness Improvement Method for Diverse HallucinationsCategories[C] / / Proceedings of the ACM on Web Conference 2025. 2025: 2562-2578. [4]Li K, Patel O, Viégas F, et al. Inference-time intervention:Eliciting truthful answers from a language model[J]. Advances in NeuralInformation Processing Systems, 2023, 36: 41451-41530. [5]Adila D, Zhang S, Han B, et al. Discovering Bias in Latent Space:An Unsupervised Debiasing Approach[A]. arXiv, 2024. [6]Qiu Y, Zhao Z, Ziser Y, et al. Spectral editing of activations forlarge language model alignment[C] / / Globerson A, Mackey L, Belgrave D, et al.Advances in neural information processing systems: Vol. 37. CurranAssociates, Inc., 2024: 56958-56987. [7]Cao Z, Yang Y, Zhao H. SCANS: Mitigating the Exaggerated Safetyfor LLMs via Safety-Conscious Activation Steering[J]. Proceedings of the AAAIConference on Artificial Intelligence, 2025, 39(22): 23523-23531. [8]Lee B W, Padhi I, Ramamurthy K N, et al. Programming Refusal withConditional Activation Steering[C] / / The Thirteenth International Conferenceon Learning Representations, ICLR 2025, Singapore, April 24-28, 2025.OpenReview.net, 2025. [9]Ghosh S, Bhattacharjee A, Ziser Y, et al. SafeSteer: InterpretableSafety Steering with Refusal-Evasion in LLMs[A]. arXiv, 2025.
[10] Rimsky N, Gabrieli N, Schulz J, et al. Steering Llama 2 viaContrastive Activation Addition[C] / / Ku L W, Martins A, Srikumar V.Proceedings of the 62nd Annual Meeting of the Association for ComputationalLinguistics (Volume 1: Long Papers). Bangkok, Thailand: Association forComputational Linguistics, 2024: 15504-15522.
[11] Hegazy A, Elhoushi M, Alanwar A. Guiding Giants: LightweightControllers for Weighted Activation Steering in LLMs[A]. arXiv, 2025.
[12] Im S, Li Y. A Unified Understanding and Evaluation of SteeringMethods[A]. arXiv, 2025.
[13] Tigges C, Hollinsworth OJ, Geiger A, et al. LinearRepresentations of Sentiment in Large Language Models[A]. arXiv, 2023. Summary of the Invention
[0007] In view of the above-mentioned deficiencies of the prior art, the present invention at least solves the following technical problems: 1. Various algorithmic strategies exist for quickly repairing the credibility problem of LLMs using representation engineering (such as mean difference, logistic regression, PCA, K-means, etc.), and there are different applicability among them. However, the existing technology lacks a unified framework for in-depth analysis and cannot help to adopt the optimal strategy when applying it. 2. Existing representation engineering algorithms can often only obtain unit vectors of the target concept repair representation. The specific intervention range depends on user customization, which lacks flexibility and limits the application and promotion.
[0008] This invention discloses a method for trustworthy content restoration of large model-generated content based on adaptive guidance of representation engineering. First, it analyzes the applicability of mainstream representation engineering algorithms and assigns anchor vectors to different algorithms accordingly. Based on this, it matches the optimal algorithm during inference to achieve adaptive selection of different algorithms and leverage their respective advantages. Second, it predefines differentiated guidance strengths based on the projection of the target restoration representation in the guidance direction, reducing the debugging cost for users. In summary, this method can achieve flexible adaptive restoration of trustworthy requirements such as realism, fairness, and security, thereby enhancing the trustworthiness of large model-generated content. Specifically, the method includes the following steps: S1: Construct a trusted repair sample dataset in the form of an A / B test based on a publicly available trusted evaluation dataset; S2: Input the sample data into the large model and extract the activation features of positive and negative samples in each Transformer decoder layer; S3: It integrates multiple representation engineering algorithms to calculate the repair vectors corresponding to each layer, and selects the optimal repair layer based on the alignment between activation differences and repair vectors; S4: Divide the adaptive sample set according to the alignment degree between the sample and each algorithm, calculate the anchor vector and default intervention intensity of each algorithm, and build an adaptive repair strategy library. S5: During inference, extract the token activation of the user query in the optimal repair layer, determine whether to repair and select the optimal algorithm by matching the degree with the anchor vector, and implement targeted intervention to achieve reliable repair; Furthermore, in step S1, the construction of the AB test-style trusted repair sample dataset is specifically as follows: each sample contains one question and two answers, the two answers being a positive answer corresponding to the expected behavior and a negative answer corresponding to the unexpected behavior, and the positive answers are randomly and evenly distributed among the two answers of all samples; Furthermore, in step S2, the extraction of activation features specifically involves: concatenating the questions in the samples with positive and negative answers to form input text; inputting the input text into the large model, and extracting the activation vector of the first token in the answer part of each Transformer decoder layer, which are used as the activation features of the positive and negative samples in the corresponding layers. Furthermore, in step S2, the positive and negative samples are in the first... The activation features of the layers are respectively represented by the activation matrix. and It means, and , ,in, This represents the number of positive and negative samples. The activation dimension of the model; Furthermore, in step S3, the selection of the optimal repair layer based on the alignment between the activation difference and the repair vector specifically includes: S31: Calculate the difference vector of activation features of positive and negative samples in each layer. ,in, For positive samples in the 1st The activation matrix of the layer, For negative samples in the 1st The activation matrix of the layer; S32: Calculate the cosine similarity between each difference vector and the corresponding repair vector in each layer; S33: Statistical analysis of the proportion of weakly aligned samples at each layer The weakly aligned samples are those where the cosine similarity between the difference vector and all repaired vectors is below the threshold. ; S34: Select the weakly aligned sample ratio The smallest layer is the optimal repair layer. ; Furthermore, in step S4, the calculation process of the anchor vector is as follows: for each representation engineering algorithm, determine the suitable sample set for that algorithm. This refers to the set of samples that best align with the vectors repaired by the algorithm; the adapted sample set is then calculated. All negative samples in the optimal repair layer The mean of the activation features is used as the anchor vector corresponding to this algorithm. ; Furthermore, in step S4, the calculation process for the default intervention intensity is as follows: For each representation engineering algorithm, an adaptation sample set is used... Calculate the projection of the activation difference vector of each sample onto the algorithm's repair vector; then calculate the mean of all projection values, which is used as the default intervention strength of the algorithm. ; Furthermore, in step S5, the process of determining whether to repair based on the matching degree with the anchor point vector includes: S51: Extract user queries in the optimal repair layer Each token activation vector ; S52: Calculate each activation vector anchor vectors of each algorithm The cosine similarity is used to determine if the token is trustworthy. If all cosine similarities are less than 0, the token is considered trustworthy. S53: Statistical analysis of the percentage of tokens without trust issues. ,like If ≥0.5, no repair is needed; if If the value is less than 0.5, then repair is required. Furthermore, in step S5, the process of selecting the optimal algorithm is as follows: query the users who are determined to need repair, and count the highest alignment degree between each token activation vector and each algorithm anchor vector; select the algorithm with the highest alignment degree frequency as the optimal algorithm; Furthermore, in step S3, the injection position of the repair vector is after the feedforward network module of each Transformer decoder layer of the large model, and after the repair vector is injected, the activation vector of the corresponding layer is adjusted according to the formula. Calculate, where, The activation vector after injecting the repair vector. This is the activation vector output by the attention module of this layer. It is a feedforward network. This is the repair vector for the corresponding layer. The intensity of intervention.
[0009] This invention achieves at least the following technical effects: 1. In terms of technological advantages: 1) Optimal Algorithm Matching Strategy Based on Anchor Vectors: Breaking through the limitations of traditional "one-size-fits-all" repair algorithms, this strategy constructs a precise "problem-algorithm" adaptation mechanism using anchor vectors. Anchor vectors are generated based on the negative activation mean of each algorithm's adaptation sample set, accurately capturing representative activation patterns of different credible problem types (such as activation differences in hallucinations and biases). During inference, token-level similarity matching quickly locates credible problems and selects the optimal algorithm, avoiding damage to the model's original capabilities caused by overall intervention, while significantly improving repair flexibility. 2) Projection-based default intervention intensity: This addresses the pain point of traditional activation interventions where "intensity depends on manual definition." The intervention intensity is calculated by projecting the mean of sample activation differences onto the repair vector. Its physical meaning is "the necessary magnitude to correct unreliable activations," and it is entirely based on statistical data rather than empirical settings. This intensity forms a synergistic closed loop with the repair vector and anchor vector, avoiding both over-intervention leading to rigid generation and under-intervention causing repair failure, thus further improving repair robustness. 2. Regarding performance indicators: 1) Credibility repair effect: On the three core credibility issues of authenticity (TruthfulQA dataset), fairness (BBQ dataset), and security (SafeEdit dataset), the test results for the LLaMA-3.1-8B-Chat model show that the overall relative performance improvement of this invention is 15.36%, and the repair gain follows the "diminishing returns pattern" (the repair effect is more significant for models with weaker initial performance).
[0010] 2) General capabilities preserved: By dynamically selecting the optimal algorithm and targeted intervention, unnecessary activation modifications are avoided. While improving reliability, the model's general capabilities such as MMLU (Multi-task Language Understanding) and AlpacaEval (Instruction Follow-up) are not compromised, and a slight improvement is achieved in some scenarios. 3) Processing speed and adaptability: The inference stage requires no complex calculations, only a lightweight operation of "activation vector similarity matching" is added, and the user interaction delay is negligible; it supports large models with different architectures and has good adaptability to the credibility requirements of different domains. 3. Regarding production implementation: 1) System Architecture: The system adopts a highly modular design. The core modules include the "Repair Vector Calculation Module", "Repair Layer Selection Module", and "Adaptive Repair Strategy Construction Module". The modules are loosely coupled and support independent iteration and customized adjustments. 2) Model training and deployment: No additional training is required for large models. The policy library is built solely from sample data, making it lightweight and supporting algorithm expansion. It can be deployed in the cloud or local environment and provides API interfaces for easy integration into existing LLM development systems. Attached Figure Description
[0011] Figure 1 This is a schematic diagram of the architecture of the large model generation content credibility repair method based on characterization engineering adaptive guidance of the present invention; Figure 2 This is a schematic diagram of the adaptive repair strategy construction process of the present invention; Figure 3 This is a schematic diagram of the trusted repair application process during the inference phase of this invention; Figure 4 This is a schematic diagram of the experimental results verifying the effectiveness of the adaptive intervention intensity of the present invention. Detailed Implementation
[0012] The following description, with reference to the accompanying drawings, illustrates several preferred embodiments of the present invention to make its technical content clearer and easier to understand. The present invention can be embodied in many different forms, and the scope of protection of the present invention is not limited to the embodiments mentioned herein.
[0013] In the accompanying drawings, components with the same structure are indicated by the same numerical designation, and components with similar structures or functions are indicated by similar numerical designations. The dimensions and thicknesses of each component shown in the drawings are arbitrary, and the present invention does not limit the dimensions and thicknesses of each component. To make the illustrations clearer, the thickness of some components has been appropriately exaggerated in the drawings.
[0014] This invention is aimed at having A large model with a Transformer decoder layer By injecting repair vectors during the inference stage, the credibility of the model-generated content can be improved without additional training, while retaining its general reasoning and language generation capabilities.
[0015] The Transformer decoder layer contains two core residual blocks, and the basic activation calculation logic is as follows: Attention module residual block: ,in The MHA (Multi-Head Self-Attention) module is used to capture contextual dependencies between tokens, activating the output of the previous layer. Feedforward network residual block: The FFN (Feedforward Network) module refines token activation through nonlinear transformation, providing abstract guidance for subsequent token prediction.
[0016] The injection point for the repair vector is after the feedforward network module. After injection, the activation vector of this layer is updated according to the following formula:
[0017] in, The activation vector after injecting the repair vector. For the first Layer repair vector, The intensity of intervention is used to control the repair effort, such as... Figure 1 As shown, the reliable content repair method for large model generation based on adaptive guidance of representation engineering proposed in this invention includes the following steps: S1: Construct a trusted repair sample dataset in the form of an A / B test based on a publicly available trusted evaluation dataset; S2: Input the sample data into the large model and extract the activation features of positive and negative samples in each Transformer decoder layer; S3: It integrates multiple representation engineering algorithms to calculate the repair vectors corresponding to each layer, and selects the optimal repair layer based on the alignment between activation differences and repair vectors; S4: Divide the adaptive sample set according to the alignment degree between the sample and each algorithm, calculate the anchor vector and default intervention intensity of each algorithm, and build an adaptive repair strategy library. S5: During inference, extract the token activation of the user query in the optimal repair layer, determine whether to repair and select the optimal algorithm by matching the degree with the anchor vector, and implement targeted intervention to achieve reliable repair.
[0018] The implementation steps of the present invention are described below through specific embodiments, including as follows: Figure 2 The build process shown and as Figure 3 The following is the reasoning application process: 1. Construction of A / B Test Data for Trusted Repair Repair Vector It is an abstract representation in the model activation space aligned with a specific target concept. To extract effective target concept representations, a corresponding sample dataset needs to be constructed. Given a repair target, the sample set typically includes positive cue words with the desired behavior. and negative cue words that indicate undesirable behavior When m=n, these samples generally appear in pairs.
[0019] Based on publicly available, reliable evaluation datasets, samples were standardized into A / B test format sets. That is, a sample consists of a question and two answers, A and B, one of which is a positive answer and the other is a negative answer, denoted as . ,in, For the problematic text in the scenario to be repaired, A positive response that meets the requirements of credibility A negative answer indicating a question of credibility.
[0020] To ensure fairness and effectiveness, positive responses are randomly and evenly distributed across the two responses (A and B) in all samples to avoid interference from sample bias in subsequent calculations, and the number of positive and negative responses in the sample set remains consistent (let's say n).
[0021] 2. Representation Extraction Related to Trustworthy Concepts
[0022] For the model It is necessary to iterate through each sample. Positive and negative activations are extracted as the basis for subsequent computation of credible conceptual representations.
[0023] Therefore, using samples For example, this embodiment will address the problem. With positive answers Concatenate the responses, using the activation of the first token in the response as the positive activation for that sample, and concatenate the corresponding negative responses. Extract negative activation.
[0024] In general, for the model ,make and Representing the positive and negative samples respectively Layer activation, where This represents the activation dimension of the model.
[0025] 3. Calculate the repair vector and select the optimal repair layer.
[0026] Based on existing repair vector calculation methods, corresponding repair vectors are obtained respectively. Subsequently, they are organically integrated in an innovative way, and multiple algorithms are brought together into an adaptive strategy library.
[0027] Therefore, this embodiment collects an algorithm library consisting of k algorithms. For each algorithm and each layer Calculate a repair vector Therefore, each layer Generate a set of repair vectors Each corresponds to one algorithm. .
[0028] Traditional methods select intervention layers based on evaluation results, which leads to inference computation costs that increase linearly with model depth. This embodiment evaluates the suitability of intervention at the appropriate layer by directly measuring the alignment between the difference in positive and negative activations and the repair vector. At each layer... Activate differences for:
[0029] Each row It is the first Activation differences among samples.
[0030] This embodiment uses each With repair vector set The cosine similarity between the samples is used to measure the degree of alignment. If the activation difference of a sample is less than a threshold aligned with all the repair vectors, the alignment is considered complete. If the sample is weakly aligned, it is considered a weakly aligned sample at that layer. The weakly aligned sample ratio of a layer is defined as:
[0031] Selecting weakly aligned sample ratio The smallest layer is the optimal repair intervention layer. This ensures maximum alignment and robust consistency in guiding the construction of repair strategies.
[0032] 4. Construction of Adaptive Trusted Repair Strategy
[0033] At the optimal intervention layer Different algorithms produce repair vectors that exhibit varying degrees of suitability and effectiveness. To maximize overall repair performance, samples are assigned to the algorithm with the highest alignment to each algorithm based on the degree of alignment between each sample and the different algorithms (i.e., the cosine similarity between the sample activation difference in the previous step and the repair vector of a specific algorithm).
[0034] For each algorithm Calculate an anchor vector This refers to the mean of the negative activation vectors in the sample set adapted by the corresponding algorithm. This vector represents the algorithm's... The representative activation patterns of samples with credible problems that need repair can serve as a reference for inference time matching. Algorithm Default intervention intensity Defined as the projection of sample activation differential onto the repair vector The mean of . The formula is defined as follows:
[0035]
[0036] in, Indicates applicable to the algorithm The sample set.
[0037] The final adaptive repair strategy library is Each tuple describes the algorithm. The complete repair parameters are encapsulated in this strategy, which includes the repair vector, anchor vector, and intervention strength, enabling precise and effective intervention during inference.
[0038] 5. Targeted intervention and repair during the reasoning stage
[0039] During the inference application phase, it is necessary to assess whether intervention and repair are needed based on actual user input. If so, an appropriate repair vector should be selected to ensure the credibility of the LLM-generated content.
[0040] For any user query input ,Model During reasoning, in the first The layer obtains the activation of each token. Activation at each location Calculate the alignment degree between its activation and each anchor vector. If all are less than 0, it is considered that the token at that position does not have any related trust issues. Otherwise, the repair vector with the highest alignment degree needs to be obtained as one of the candidates.
[0041] Analyze the matching results for all positions and select the most suitable repair algorithm (if most tokens do not have trust issues, no repair is needed). The percentage of tokens without trust issues is also considered. If the value is below the threshold of 0.5, the query input needs to be intervened and repaired. The relevant formula is as follows:
[0042] If repair is needed, the corresponding repair vector and intervention intensity are obtained, and the standard formula of the characterization engineering is applied to intervene and repair, effectively guiding the model. Fix the credibility issue.
[0043] The experimental setup for this embodiment is as follows: 1. Dataset aspect To evaluate its effectiveness on the three core trust issues of LLM—truthfulness, fairness, and security—this embodiment uses three widely used, publicly authoritative datasets: 1) Truthful QA contains 817 questions across 38 categories, used to measure whether a language model can avoid mimicking human cognitive errors and maintain authenticity when generating answers; 2) BBQ assesses social biases and stereotypes in language models through real-world scenarios and related questions; 3) SafeEdit covers nine unsafe categories, including legal and ethical ones, and is used to study the feasibility of knowledge editing technology for removing harmful content from large language models. For each type of trust issue, in addition to focusing on practical effectiveness, this invention also evaluates whether intervention and remediation have a negative impact on the general capabilities of the model.
[0044] MMLU is a large-scale multi-task language understanding benchmark proposed in 2021. It contains 15,908 multiple-choice questions from 57 subject areas to evaluate the extensive knowledge and problem-solving capabilities acquired by the model. AlpacaEval is a fully automated evaluation tool developed in 2024. It quickly and cost-effectively evaluates the instruction following and overall alignment quality of large language models by calculating the model's win rate in various tasks.
[0045] 2. Regarding the comparison methods: This example compares the performance of each algorithm in the algorithm library when applied individually. The descriptions of each algorithm are as follows: 1) CAA / MD: This method calculates the average difference between positive and negative activations as the repair vector.
[0046] 2) ITI / LR: This method uses cross-entropy loss to train a simple binary classifier to distinguish between positive and negative activations, and then uses the normal vector of the decision boundary, i.e. the classifier weight vector, as the repair vector.
[0047] 3) RepE / PCA: This method extracts the first principal component as the repair vector by performing PCA on the contrast activation difference set.
[0048] 4) Kmeans: This method uses K=2 KMeans to cluster the positive and negative activation combination set, and the repair vector is defined as the difference between the two cluster centers.
[0049] 3. Evaluation indicators: All assessments were rephrased as A / B test multiple-choice questions, and this embodiment actually compares the overall performance average accuracy (ACC).
[0050] 4. Experimental Environment
[0051] CentOS 7.9 system, Intel Xeon E5-2630 v4 CPU (12 cores 2.2GHz), PyTorch 2.0 framework, GPU is a single A6000 GPU (48GB).
[0052] To comprehensively evaluate the technical advancement and practical value of this invention, verification will be conducted from two dimensions: overall detection capability and innovative value. The positioning and objectives of each verification module are as follows: 1) Effectiveness experiment: Verify overall performance This embodiment evaluates the invention on a representative LLM, namely LLaMA-3.1-8B-Chat, and performs an overall performance evaluation on multiple standard benchmarks covering realism, fairness, and security. In addition to evaluating improvements for specific problems, its impact on general capabilities is also examined. As shown in Table 1, the invention consistently outperforms all baseline methods. Notably, compared to other methods, it enhances performance on specific problems while improving general capabilities; its main findings are summarized below: Table 1. Experimental Results Comparison of Overall Model Performance in Reliability and General Ability Assessment
[0053] Repair gain varies with the initial performance of the model, and performance improvements under the characterization engineering paradigm follow a diminishing returns pattern: weaker models benefit more.
[0054] Repairability and intervention side effects differ in their credibility. Fairness is easier to control than realism because realism has a broader range of factors; the baseline approach directly improves fairness by at least 4.63%, while realism only improves by 0.45%. While safety improvements benefit harmlessness, overly stringent interventions may increase rejection rates, sometimes even at the expense of general usability. Conversely, improvements in realism and fairness contribute to better overall model performance.
[0055] This invention improves repair performance while maintaining general performance, with the advantage of dynamically selecting the optimal strategy from multiple algorithms during inference. The method offers significant improvements on the target problem while preserving general capabilities by avoiding unnecessary interventions. In contrast, baseline methods lack this adaptability, often leading to reduced robustness. For example, Rep E ranks second in fairness and safety but falls short of the baseline model in overall capability.
[0056] 2) Validation of Innovation: The Importance of Adaptive Intervention Intensity
[0057] To evaluate the effectiveness of the adaptive intervention strength assigned to each strategy, regarding authenticity, such as... Figure 4 As shown, two size adjustment methods are compared: I. Ignoring sample features (see...) Figure 4 (a) Apply a fixed intensity α to all repair vectors; II. Adjust the adaptive intensity using a global sensitivity factor β (see...) Figure 4 (b))). This comparison reveals the benefits of adaptive intervention magnitude for credibility restoration.
[0058] Overall, the adaptive intervention intensity of this invention maintains overall performance while providing a more stable improvement in reliability. Specifically, a fixed intervention intensity between 1 and 6 maintains general capability and gradually improves reliability, reaching a peak at α=4.5 with a gain of 4.65%. In contrast, using a global scaling factor β in the range of (0.7~1.8) achieves better results, with a performance gain as high as 8.90%.
[0059] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.
Claims
1. A method for reliable content repair in large model generation based on adaptive guidance of representation engineering, characterized in that, The method includes the following steps: S1: Construct a trusted repair sample dataset in the form of an A / B test based on a publicly available trusted evaluation dataset; S2: Input the sample data into the large model and extract the activation features of positive and negative samples in each Transformer decoder layer; S3: It integrates multiple representation engineering algorithms to calculate the repair vectors corresponding to each layer, and selects the optimal repair layer based on the alignment between activation differences and repair vectors; S4: Divide the adaptive sample set according to the alignment degree between the sample and each algorithm, calculate the anchor vector and default intervention intensity of each algorithm, and build an adaptive repair strategy library. S5: During inference, extract the token activation of the user query in the optimal repair layer, determine whether to repair and select the optimal algorithm by matching the degree with the anchor vector, and implement targeted intervention to achieve reliable repair.
2. The method for reliable content repair of large model generation based on adaptive guidance of representation engineering as described in claim 1, characterized in that, In step S1, the construction of the AB test-style trusted repair sample dataset is as follows: each sample contains a question and two answers, the two answers being a positive answer corresponding to the desired behavior and a negative answer corresponding to the undesired behavior, and the positive answers are randomly and evenly distributed among the two answers of all samples.
3. The method for reliable content repair of large model generation based on adaptive guidance of representation engineering as described in claim 1, characterized in that, In step S2, the extraction of activation features specifically involves: concatenating the questions in the samples with the positive and negative answers to form the input text; after inputting the input text into the large model, extracting the activation vector of the first token in the answer part of each Transformer decoder layer, which is used as the activation features of the positive and negative samples in the corresponding layers.
4. The method for reliable content repair of large model generation based on adaptive guidance of representation engineering as described in claim 1, characterized in that, In step S2, the positive and negative samples are in the first... The activation features of the layers are respectively represented by the activation matrix. and It means, and , ,in, This represents the number of positive and negative samples. This represents the activation dimension of the model.
5. The method for reliable content repair of large model generation based on adaptive guidance of representation engineering as described in claim 1, characterized in that, In step S3, the selection of the optimal repair layer based on the alignment between the activation difference and the repair vector specifically includes: S31: Calculate the difference vector of activation features of positive and negative samples in each layer. ,in, For positive samples in the 1st The activation matrix of the layer, For negative samples in the 1st The activation matrix of the layer; S32: Calculate the cosine similarity between each difference vector and the corresponding repair vector in each layer; S33: Statistical analysis of the proportion of weakly aligned samples at each layer The weakly aligned samples are those where the cosine similarity between the difference vector and all repaired vectors is below the threshold. The sample; S34: Select the weakly aligned sample ratio The smallest layer is the optimal repair layer. .
6. The method for reliable content repair of large model generation based on adaptive guidance of representation engineering as described in claim 1, characterized in that, In step S4, the calculation process of the anchor vector is as follows: for each representation engineering algorithm, determine the suitable sample set for that algorithm. This refers to the set of samples that best align with the vectors repaired by the algorithm; the adapted sample set is then calculated. All negative samples in the optimal repair layer The mean of the activation features is used as the anchor vector corresponding to this algorithm. .
7. The method for reliable content repair of large model generation based on adaptive guidance of representation engineering as described in claim 1, characterized in that, In step S4, the calculation process for the default intervention intensity is as follows: For each representation engineering algorithm, an adaptation sample set is used... Calculate the projection of the activation difference vector of each sample onto the algorithm's repair vector; then calculate the mean of all projection values, which is used as the default intervention strength of the algorithm. .
8. The method for reliable content repair of large model generation based on adaptive guidance of representation engineering as described in claim 5, characterized in that, In step S5, the process of determining whether to repair based on the matching degree with the anchor point vector includes: S51: Extract user queries in the optimal repair layer Each token activation vector ; S52: Calculate each activation vector anchor vectors of each algorithm The cosine similarity is used to determine if the token is trustworthy. If all cosine similarities are less than 0, the token is considered trustworthy. S53: Statistical analysis of the percentage of tokens without trust issues. ,like If ≥0.5, no repair is needed; if If the value is less than 0.5, then repair is required.
9. The method for reliable content repair of large model generation based on adaptive guidance of representation engineering as described in claim 1, characterized in that, In step S5, the process of selecting the optimal algorithm is as follows: query the users who are determined to need repair, and count the highest alignment degree between each token activation vector and each algorithm anchor vector; select the algorithm with the highest alignment degree frequency as the optimal algorithm.
10. The method for reliable content repair of large model generation based on adaptive guidance of representation engineering as described in claim 1, characterized in that, In step S3, the injection position of the repair vector is after the feedforward network module of each Transformer decoder layer of the large model, and after the repair vector is injected, the activation vector of the corresponding layer is adjusted according to the formula. Calculate, where, The activation vector after injecting the repair vector. This is the activation vector output by the attention module of this layer. It is a feedforward network. This is the repair vector for the corresponding layer. The intensity of intervention.