Large model security neuron screening method and device

By adopting a large-model safe neuron discovery method based on activation comparison in large language models, the problem of difficulty in scaling to large models in the existing technology is solved, and the ability to discover safe neurons in large models is realized, providing a scalable safe neuron discovery method.

CN119990203APending Publication Date: 2025-05-13NAT COMP NETWORK & INFORMATION SECURITY MANAGEMENT CENT +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411812196.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-10
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art is difficult to expand in large language models, is limited by task type, and is mostly experimented on small models, making it difficult to effectively solve the security challenges of large models.

Method used

A large-model security neuron discovery method based on activation comparison is used to calculate the neuron activation differences between the safe alignment model and the basic model to determine the safe neurons.

Benefits of technology

The ability to discover safe neurons in large models is realized, breaks through the limitations of task form and model scale, and provides a scalable safe neuron discovery method that is not limited by task form.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990203A_ABST
    Figure CN119990203A_ABST
Patent Text Reader

Abstract

The invention provides a large model safety neuron screening method and device, and the method comprises the steps: carrying out the safety alignment of a basic large model, and obtaining a safety alignment model; calculating a neuron activation difference between the security alignment model and the basic large model; and based on the neuron activation difference, determining a safe neuron when safe alignment is performed on the basic large model. The method starts from the internal properties of the model, is not limited by task forms, is easy to expand, is suitable for finding safe neurons in the large model, and provides a scheme for further researching the safety mechanism of the large model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of large models, and in particular to a large model safety neuron screening method and device. Background Art

[0002] Large Language Models (LLMs) have achieved remarkable achievements in the field of natural language processing, but also face security challenges such as bias, discrimination, and adversarial attacks.

[0003] At present, the research on the safety mechanism of large models includes neuron-based methods and neural circuit-based methods. Among them, the neuron-based method mainly calculates the predictability of the activation of each neuron to the task label, and finds skill neurons related to specific natural language processing tasks in the pre-trained language model. Only using the activation of these neurons can achieve performance similar to that of the complete model in the corresponding task. And pruning the language model based on this discovery can significantly improve the running speed of the model. The neural circuit-based method mainly replaces the activation of each component in the model with activations that are useless for completing the task, and observes the changes in the performance of the model on the task, and finally determines the neural circuit in the model. Each component in the circuit has a causal effect on the output of the model. Destroying the circuit will cause the model to be unable to complete the task. Keeping only the circuit can restore most of the model performance.

[0004] The above methods are limited by the task type and are mostly experimented on small models, making it difficult to expand to large models. Summary of the invention

[0005] The present invention provides a large-model safe neuron discovery method and device based on activation comparison, which are used to solve the defects of the prior art that it is limited by the task type and experiments are mostly conducted on small models, which are difficult to expand to large models, and realize safe neuron discovery in large models.

[0006] The present invention provides a large-model safe neuron screening method, comprising the following steps: Performing safety alignment on the basic large model to obtain a safety alignment model; Calculating the difference in neuron activation between the secure alignment model and the base large model; Based on the neuron activation differences, safe neurons for safely aligning the basic large model are determined.

[0007] According to a large model security neuron screening method provided by the present invention, the steps of performing security alignment on the basic large model to obtain a security alignment model specifically include: Securely align the underlying large model using a parameter-efficient fine-tuning method.

[0008] According to a large model safety neuron screening method provided by the present invention, the step of calculating the neuron activation difference between the safety alignment model and the basic large model specifically includes: Based on a preset specific corpus, respectively counting responses generated by the security alignment model and the basic large model for the specific corpus; Based on the two generated responses, the difference in neuron activations between the secure alignment model and the base large model is calculated.

[0009] According to a large model safety neuron screening method provided by the present invention, the step of calculating the neuron activation difference between the safety alignment model and the basic large model based on two generated responses specifically includes: Assume the basic model is , the safety alignment model is ; For a given input , respectively and express and Generate a response; In the full conversation Perform forward calculation and collect The corresponding position of the neuron is activated, and we get Neuron activation when generating a response , where the complete conversation is and The splicing of , Indicates that the input No. The word unit, No. Layer The activation of neurons; On the complete conversation Perform forward calculation and collect The corresponding position of the neuron is activated, and we get Approximate neuron activations when generating responses , Indicates that the input No. The word unit, No. Layer The activation of neurons; calculate and The difference in neuron activation when generating a response is: - .

[0010] According to a large model safety neuron screening method provided by the present invention, the step of calculating the neuron activation difference between the safety alignment model and the basic large model based on two generated responses specifically includes: Assume the basic model is , the safety alignment model is ; For a given input , respectively and express and Generate a response; In the full conversation Perform forward calculation and collect The corresponding position of the neuron is activated, and we get Approximate neuron activations when generating responses , where the complete conversation is and The splicing of , Indicates that the input No. The word unit, No. Layer The activation of neurons; On the complete conversation Perform forward calculation and collect The corresponding position of the neuron is activated, and we get Neuron activation when generating a response ,in, Indicates that the input No. The word unit, No. Layer The activation of neurons; calculate and The difference in neuron activations when generating a response is: - .

[0011] According to a large model safety neuron screening method provided by the present invention, based on the neuron activation difference, the step of determining the safety neuron when performing safety alignment on the basic large model specifically includes: Calculating a change score for each neuron based on the neuron activation difference; Neurons are ranked based on their change scores, and neurons with the highest change scores are selected as safe neurons based on their causal effects on the safety of the model.

[0012] According to a large model safety neuron screening method provided by the present invention, the process of calculating the change score of each neuron based on the neuron activation difference specifically includes: make Indicates that the input No. The word unit, No. Layer The activation of neurons; For a given data set , defined based on Change score For middle and The root mean square of the differences in neuron activations when generating responses: , Where q = 1 or 2, correspondingly, Pick or .

[0013] According to a large model safety neuron screening method provided by the present invention, based on the causal effect of neurons on model safety, the process of selecting neurons with the highest change score as safety neurons specifically includes: S1, select different numbers of neurons with the largest change scores as candidate safe neurons; S2. Perform a secure alignment model on a given input The forward computation process of and caches the activations of candidate safety neurons; S3. Execute the basic large model on the given input In the forward calculation process, the activation of a given neuron is replaced by the activation of the cached safe neuron in the previous step, and the forward calculation continues; S4: After the forward calculation is completed, the prediction of the next word is obtained, and the input is expanded with the predicted word and steps S2 to S3 are repeated until the generation is completed; S5. The minimum set of neurons with causal effects similar to those of all neurons is taken as safe neurons.

[0014] The present invention also provides a large-model safe neuron screening device, comprising the following modules: A safety alignment module is used to safely align the basic large model to obtain a safety alignment model; A difference calculation module, used for calculating the difference in neuron activation between the secure alignment model and the basic large model; A determination module is used to determine the safe neurons when performing safe alignment on the basic large model based on the neuron activation difference.

[0015] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the large-model safety neuron screening method as described above is implemented.

[0016] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the large-model safety neuron screening methods described above.

[0017] The present invention also provides a computer program product, comprising a computer program, which, when executed by a processor, implements any of the large-model safety neuron screening methods described above.

[0018] The large model safety neuron screening method and device provided by the present invention obtains a safety alignment model by safety alignment of the basic large model; calculates the neuron activation difference between the safety alignment model and the basic large model; and determines the safety neuron when the basic large model is safety aligned based on the neuron activation difference. The present invention starts from the internal properties of the model itself, is not limited by the task form, is easy to expand, is suitable for the discovery of safety neurons in large models, and provides a solution for further studying the safety mechanism of large models. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0020] Figure 1 This is one of the flow charts of the large-model safe neuron screening method provided by the present invention.

[0021] Figure 2 It is a schematic flow chart of the dynamic activation repair method provided by the present invention.

[0022] Figure 3 This is the second flow chart of the large-model safe neuron screening method provided by the present invention.

[0023] Figure 4 It is a schematic diagram of the structure of the large-model safe neuron screening device provided by the present invention.

[0024] Figure 5It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0025] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0026] The present invention is described in detail below in conjunction with the accompanying drawings of the specification. The specific operating method in the method embodiment can also be applied to the device embodiment or the system embodiment. In the description of the present invention, unless otherwise specified, "at least one" includes one or more. "Multiple" refers to two or more. For example, at least one of A, B and C includes: A exists alone, B exists alone, A and B exist at the same time, A and C exist at the same time, B and C exist at the same time, and A, B and C exist at the same time. In the present invention, " / " means or, for example, A / B can mean A or B; "and / or" in this article is only a kind of association relationship that describes the associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone.

[0027] As an important development in the field of natural language processing, large language models (LLMs) have made remarkable achievements in recent years. These models are trained with large-scale data and can show excellent performance in various language tasks, including but not limited to text generation, translation, question answering, text summarization, etc. With the increase of computing power and data resources, the ability of large models in language understanding and generation has been continuously improved, becoming one of the hot spots in artificial intelligence research.

[0028] Although large models have brought many innovative applications, they also face many security challenges. These challenges include (1) bias and discrimination: the bias and discrimination learned by large models from data during training may be reflected in the generated results, resulting in unfair impacts on certain groups in practical applications; (2) adversarial attacks (jailbreaking): large models may be subject to adversarial attacks, that is, by inputting specific malicious data to induce the model to produce incorrect or harmful outputs. Such attacks pose a serious threat to the security of the model.

[0029] Studying the safety mechanism of large models is a fundamental way to address these challenges. At present, the mechanism research of large models can be divided into two categories: neuron-based methods and neural circuit-based methods.

[0030] The paper "Finding Skill Neurons in Pre-trained Transformer-based Language Models" written by Xiaozhi Wang et al. is a neuron-based method that finds skill neurons related to specific natural language processing tasks in pre-trained language models by calculating the predictability of each neuron's activation to the task label. Using only the activation of these neurons can achieve performance similar to that of the complete model in the corresponding task. Based on this discovery, pruning the language model can significantly improve the model's running speed.

[0031] The paper "Interpretability in the Wild: A Circuit for Indirect Object Identification in GPT-2 Small" written by Kevin Wang et al. is a neural circuit-based method. By replacing the activation of each component in the model with activations that are useless for completing the task and observing the changes in the model's performance on the task, the neural circuit in the model is ultimately determined. Each component in this circuit has a causal effect on the model's output. Destroying this circuit will cause the model to be unable to complete the task, while retaining only this circuit can restore most of the model's performance.

[0032] The method proposed by Xiaozhi Wang et al. is based on the predictability of neuron activation to task labels, so it can only be used for classification tasks. Security, as a property of the model, is more reflected in open-ended generation tasks. At the same time, their method has only been experimented on smaller models such as RoBERTa, and it is still unknown whether it can be migrated to today's large models.

[0033] The method proposed by Kevin Wang et al. is applicable to tasks where the target output of the model is known given the input, such as the indirect object recognition task studied in the paper, while the model generates safe responses but there is no fixed answer. At the same time, their method was only experimented on smaller models such as GPT-2 Small, and is difficult to expand to models of today's scale.

[0034] In view of this, the present invention provides a large-model safe neuron screening method and device to solve the above-mentioned problems.

[0035] The present invention will be described in detail below in conjunction with specific implementation modes.

[0036] In some specific embodiments of the present invention, Figure 1 As shown, this scheme provides a large-model safe neuron screening method, including: Step 100: align the basic large model securely to obtain a securely aligned model; Step 200, calculating the neuron activation difference between the secure alignment model and the basic large model; Step 300: Determine the safe neurons for safe alignment of the basic large model based on the neuron activation differences.

[0037] It should be noted that the existing safe neuron screening schemes have task limitations and model size limitations, and cannot be widely applied to all tasks, especially open generation tasks, and are difficult to expand to large models.

[0038] Therefore, the present invention obtains the security alignment model, determines the security neurons of the large model when performing security alignment according to the difference in neuron activation between the security alignment model and the basic large model, and designs a specific neuron discovery method that is not restricted by the task form and is scalable based on the internal properties of the model itself, so that it can be used for security neuron discovery in the large model, providing a preliminary solution for further studying the security mechanism of the large model.

[0039] The above steps are described in detail below through specific embodiments.

[0040] Step 100: align the basic large model securely to obtain a securely aligned model; In some possible implementations of the present invention, the step of performing secure alignment on the basic large model to obtain a secure alignment model specifically includes: Securely align the underlying large model using a parameter-efficient fine-tuning method.

[0041] Specifically, this embodiment provides an implementation method for securely aligning a basic large model, by securely aligning the basic large model and ensuring that model parameters related to neurons remain unchanged during the alignment process by efficient parameter fine-tuning.

[0042] It is worth mentioning that Parameter Efficient Fine-tuning (PEFT) is a technique for large pre-trained models that aims to reduce the amount of parameters that need to be adjusted during fine-tuning while maintaining or improving the performance of the model.

[0043] In a possible embodiment, the following specific parameter efficient fine-tuning methods may be used to achieve safe alignment: IA3 (Infused Adapter by Inhibiting and Amplifying Inner Activations), IA3 is a parameter-efficient fine-tuning technique that improves fine-tuning efficiency by reducing the number of trainable parameters in the model.

[0044] Specifically, the implementation process of IA3 includes: Injection of learned vectors: IA3 rescales internal activations using learned vectors. These vectors are injected into the feed-forward modules of the Transformer-based architecture, while the original weights remain frozen.

[0045] These learned vectors are the only trainable parameters during fine-tuning, and IA3 can significantly reduce the number of trainable parameters compared to methods that learn low-rank weight matrices.

[0046] Fine-tuning process: IA3 can be applied to any subset of the weight matrix in a neural network to reduce the number of trainable parameters. According to the author's implementation, IA3 weights are added to the Feedforward layer of the Transformer model.

[0047] Specifically, in each Transformer block, the IA3 weights are added to the input of the second feed-forward layer.

[0048] Code implementation: The implementation of IA3 involves suppressing or amplifying certain activation layers of the model, that is, weighting part of the model's parameters by dot-multiplying a vector.

[0049] At the code level, initialization and application of IA3 involves updating specific layers, such as the Feedforward layer, and applying these learned vectors in the forward propagation.

[0050] Model Performance and Efficiency: Models fine-tuned with IA3 perform on par with fully fine-tuned models without increasing any inference latency, as the adapter weights can be merged with the base large model.

[0051] Through the above steps, IA3 achieves efficient parameter fine-tuning, so that the model can adapt to specific downstream tasks by introducing a small number of trainable parameters while keeping the pre-trained weights unchanged, thereby improving the efficiency and performance of fine-tuning.

[0052] LoRA (Low-Rank Adaptation): LoRA adjusts the behavior of the model by adding small, low-rank matrices in the key layers of the model, rather than directly changing the structure of the entire model.

[0053] Adapter Tuning: Adapter Tuning adds adapter layers to the pre-trained model to achieve fine-tuning for specific tasks.

[0054] Prompt Tuning: Prompt Tuning adds learnable embedding vectors as prompts to the input of the pre-trained language model. By introducing task-specific prompts instead of updating all parameters of the entire model, the model can be efficiently fine-tuned.

[0055] Prefix-Tuning: Prefix-Tuning affects the output of the model by adding trainable prefix embeddings before the model input layer.

[0056] BitFit: The BitFit method only fine-tunes the bias terms of the model without changing the weight matrix, thereby achieving efficient and effective fine-tuning.

[0057] It can be understood that the present invention can adapt the model to new specific tasks or fields while maintaining the powerful feature extraction capability of the pre-trained model through these efficient parameter fine-tuning techniques, and significantly reduce training time and cost.

[0058] Step 200, calculating the neuron activation difference between the secure alignment model and the basic large model; In some possible implementations of the present invention, the step of calculating the neuron activation difference between the secure alignment model and the basic large model specifically includes: Based on a preset specific corpus, respectively counting responses generated by the security alignment model and the basic large model for the specific corpus; Based on the two generated responses, the difference in neuron activations between the secure alignment model and the base large model is calculated.

[0059] Specifically, this embodiment provides an implementation method for calculating the difference in neuron activation between a security alignment model and a basic large model, by obtaining generated responses for a specific corpus, and determining the difference in neuron activation based on the difference in generated responses between the security alignment model and the basic large model.

[0060] In some possible implementations of the present invention, the step of calculating the neuron activation difference between the secure alignment model and the basic large model based on the two generated responses specifically includes: Assume the basic model is , the safety alignment model is ; For a given input , respectively and express and Generate a response; In the full conversation Perform forward calculation and collect The corresponding position of the neuron is activated, and we get Neuron activation when generating a response , where the complete conversation is and The splicing of , Indicates that the input No. The word unit, No. Layer The activation of neurons; On the complete conversation Perform forward calculation and collect The corresponding position of the neuron is activated, and we get Approximate neuron activations when generating responses , Indicates that the input No. The word unit, No. Layer The activation of neurons; calculate and The difference in neuron activation when generating a response is: - .

[0061] Specifically, this embodiment provides an implementation method for calculating the difference in neuron activation of two models based on two generated responses. The generated responses of the basic large model are spliced ​​with the input into a complete conversation, and the neuron activation of the two models when generating responses on the complete conversation is calculated respectively. According to the difference between the two, the difference in neuron activation of the two models is obtained.

[0062] In some possible implementations of the present invention, the step of calculating the neuron activation difference between the secure alignment model and the basic large model based on the two generated responses specifically includes: Assume the basic model is , the safety alignment model is ; For a given input , respectively and express and Generate a response; In the full conversation Perform forward calculation and collect The corresponding position of the neuron is activated, and we get Approximate neuron activations when generating responses , where the complete conversation is and The splicing of , Indicates that the input No. The word unit, No. Layer The activation of neurons; On the complete conversation Perform forward calculation and collect The corresponding position of the neuron is activated, and we get Neuron activation when generating a response ,in, Indicates that the input No. The word unit, No. Layer The activation of neurons; calculate and The difference in neuron activations when generating a response is: - .

[0063] Specifically, this embodiment provides another implementation method for calculating the difference in neuron activations of two models based on two generated responses. The generated responses of the security alignment model are concatenated with the input into a complete conversation, and the neuron activations of the two models when generating responses on the complete conversation are calculated respectively. The difference between the neuron activations of the two models is obtained based on the difference between the two.

[0064] Step 300: Determine the safe neurons for safe alignment of the basic large model based on the neuron activation differences.

[0065] In some possible implementations of the present invention, the step of determining the safe neurons when safely aligning the basic large model based on the neuron activation difference specifically includes: Calculating a change score for each neuron based on the neuron activation difference; Neurons are ranked based on their change scores, and neurons with the highest change scores are selected as safe neurons based on their causal effects on the safety of the model.

[0066] Specifically, this embodiment provides an implementation method for determining safe neurons based on neuron activation differences, sorting neurons by calculating the change score of each neuron, and determining safe neurons based on the sorting results to meet the needs of safe neuron screening.

[0067] In some possible implementations of the present invention, the process of calculating the change score of each neuron based on the neuron activation difference specifically includes: make Indicates that the input No. The word unit, No. Layer The activation of neurons; For a given data set , defined based on Change score For middle and The root mean square of the differences in neuron activations when generating responses: (1), Where q = 1 or 2, correspondingly, Pick or .

[0068] Specifically, this embodiment provides an implementation method for calculating the change score of each neuron based on the difference in neuron activation. The change score based on the basic large model is obtained by calculating the root mean square of the difference in neuron activation when the two models generate responses.

[0069] In a possible embodiment, the reply of the basic large model can be concatenated with the input to obtain a complete conversation, and the two models are forward calculated on the complete conversation, and the neuron activations at corresponding positions of the two models are collected. Then, the difference in the neuron activations at corresponding positions of the two models is used as the neuron activation difference of the two models. Finally, the neuron with the largest activation difference between the unaligned basic large model and the safe alignment model is used as a candidate for a safe neuron.

[0070] It is worth noting that the complete conversation in the above embodiment can also be obtained by concatenating the reply of the security alignment model with the input.

[0071] For example, we take the example of concatenating the response of the basic large model with the input to obtain a complete conversation. For the above two LLMs, and , for a given input ,use and express and Generated response. The build-time activation can be done by entering the complete dialog ( and The splicing of ) and collect The corresponding position of the neuron activation is obtained. Indicates that the input No. The word unit, No. Layer The activation of neurons, for a given dataset , we define based on Change score For middle and The root mean square of the activation difference is generated as shown in formula (1).

[0072] It is worth noting that the change score of the neuron reflects the relevance of each neuron to the model security, but this does not mean that the activation of these neurons is actually used when the model is generated. Therefore, the present invention can further screen the neurons through the following method.

[0073] In some possible implementations of the present invention, the process of selecting a neuron with the highest change score as a safety neuron based on the causal effect of the neuron on the model security specifically includes: S1, select different numbers of neurons with the largest change scores as candidate safe neurons; S2. Perform a secure alignment model on a given input The forward computation process of and caches the activations of candidate safety neurons; S3. Execute the basic large model on the given input In the forward calculation process, the activation of a given neuron is replaced by the activation of the cached safe neuron in the previous step, and the forward calculation continues; S4: After the forward calculation is completed, the prediction of the next word is obtained, and the input is expanded with the predicted word and steps S2 to S3 are repeated until the generation is completed; S5. The minimum set of neurons with causal effects similar to those of all neurons is taken as safe neurons.

[0074] Specifically, this embodiment provides an implementation method for determining safe neurons. After sorting the neurons according to the change scores, candidate safe neurons are obtained, and then the causal effect of the neurons on the output is further verified through a dynamic activation patching method, thereby determining the safe neurons that play a role in the model security.

[0075] Specifically, Figure 2 As shown, the dynamic activation patching method includes the following steps: Step 210, candidate safety neurons: In this step, the neurons are sorted according to their change scores, and different numbers of neurons with the largest change scores are selected as candidate safe neurons; Step 220: Cache activation: In this step, the security alignment model is executed on the given input The forward computation process of and caches the activations of candidate safety neurons; Step 230, forward calculation: In this step, the base large model is executed on the same input In the forward computation process, the activation of a given neuron is replaced by the activation cached in the previous step, and the forward computation continues; Step 240: Prediction: In this step, after completing the forward calculation, the prediction of the next word is obtained, the input is extended with the predicted word, and the above steps are repeated until the generation is completed.

[0076] In general, the present invention selects different numbers of neurons with maximum change scores as safe neuron candidates, and verifies the causal effects of the candidates through dynamic activation patching. The minimum set of neurons with causal effects similar to those of all neurons is the safe neuron.

[0077] In some possible embodiments of the present invention, Figure 3 As shown, the large model safety neuron screening method provided by the present invention may include the following steps: Step 310: Use an efficient parameter fine-tuning method to securely align the basic large model to obtain a securely aligned model; Step 320: using the difference in neuron activation between the specific corpus statistics-based large model and the security alignment model when generating responses, and calculating the change score of each neuron based on the difference; Step 330: Sort the neurons based on the change scores, and select some neurons with the highest change scores as safe neurons by observing the causal effect of neurons on the safety of the model.

[0078] Through the above steps, the present invention first performs a secure alignment on the basic large model, and ensures that the model parameters related to the neurons remain unchanged during the alignment process by means of efficient parameter fine-tuning. Therefore, the neurons with the largest activation difference between the unaligned basic large model and the secure alignment model will be used as candidates for secure neurons.

[0079] The present invention proposes for the first time the problem of studying the safety mechanism of large models and designs a method for screening neurons with large model safety, which breaks through the task limitations of previous research on the mechanism of large models. This method can also be used for neuron discovery of other properties of large models, such as usefulness.

[0080] By using the large model security neuron screening method provided by the present invention, security neurons accounting for about 5% of the total number of neurons were screened and found in the mainstream large model. By manipulating the activation of these security neurons, 90% of the security of the complete model can be restored, and consistent results have been achieved on multiple red team test benchmarks, including Beavertails, RedTeam, HarmBench, and JailBreakLLMs. In addition, only about 1,500 security neurons need to be activated to predict with a high accuracy (about 80%) whether the large model will generate harmful content before generating a response, which can be used to build a security fence (filter harmful output) for the large model.

[0081] The large-model safety neuron screening device provided by the present invention is described below. The large-model safety neuron screening device described below and the large-model safety neuron screening method described above can be referenced to each other.

[0082] In some specific embodiments of the present invention, Figure 4 As shown, the present invention provides a large-scale model safety neuron screening device, which includes: A safety alignment module 41 is used to perform safety alignment on the basic large model to obtain a safety alignment model; A difference calculation module 42, used for calculating the difference in neuron activation between the secure alignment model and the basic large model; The determination module 43 is used to determine the safe neurons when performing safe alignment on the basic large model based on the neuron activation difference.

[0083] The large-model safe neuron screening device provided in an embodiment of the present invention has similar implementation principles and beneficial effects to the implementation principles and beneficial effects of the large-model safe neuron screening method shown in the above embodiment. Please refer to the implementation principles and beneficial effects of the large-model safe neuron screening method shown in the above embodiment, and no further details will be given here.

[0084] Figure 5 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 5As shown, the electronic device may include: a processor 510, a communication interface 520, a memory 530 and a communication bus 540, wherein the processor 510, the communication interface 520 and the memory 530 communicate with each other through the communication bus 540. The processor 510 may call the logic instructions in the memory 530 to execute the large model security neuron screening method, which includes: performing security alignment on the basic large model to obtain a security alignment model; calculating the neuron activation difference between the security alignment model and the basic large model; and determining the security neuron when the basic large model is securely aligned based on the neuron activation difference.

[0085] In addition, the logic instructions in the above-mentioned memory 530 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0086] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the large model safety neuron screening method provided by the above methods, and the method includes: safely aligning the basic large model to obtain a safety alignment model; calculating the neuron activation difference between the safety alignment model and the basic large model; based on the neuron activation difference, determining the safety neurons when the basic large model is safely aligned.

[0087] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, it is implemented to execute the large model safety neuron screening method provided by the above-mentioned methods. The method includes: performing safety alignment on the basic large model to obtain a safety alignment model; calculating the neuron activation difference between the safety alignment model and the basic large model; and determining the safety neurons when performing safety alignment on the basic large model based on the neuron activation difference.

[0088] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0089] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0090] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A large model safety neuron screening method, characterized in that: include: Performing safety alignment on the basic large model to obtain a safety alignment model; Calculating the difference in neuron activation between the secure alignment model and the base large model; Based on the neuron activation differences, safe neurons for safely aligning the basic large model are determined.

2. The large model safety neuron screening method according to claim 1, characterized in that: The steps of safely aligning the basic large model to obtain the safe alignment model specifically include: Securely align the underlying large model using a parameter-efficient fine-tuning method.

3. The large model safety neuron screening method according to claim 1, characterized in that: The step of calculating the neuron activation difference between the secure alignment model and the basic large model specifically includes: Based on a preset specific corpus, respectively counting responses generated by the security alignment model and the basic large model for the specific corpus; Based on the two generated responses, the difference in neuron activations between the secure alignment model and the base large model is calculated.

4. The large model safety neuron screening method according to claim 3, characterized in that: The step of calculating the neuron activation difference between the secure alignment model and the basic large model based on the two generated responses specifically includes: Assume the basic model is , the safety alignment model is ; For a given input , respectively and express and Generate a response; In the full conversation Perform forward calculation and collect The corresponding neuron is activated, and we get Neuron activation when generating a response , where the complete conversation is and The splicing of , Indicates that the input No. The word unit, No. Layer The activation of neurons; On the complete conversation Perform forward calculation and collect The corresponding neuron is activated, and we get Approximate neuron activations when generating responses , Indicates that the input No. The word unit, No. Layer The activation of neurons; calculate and The difference in neuron activation when generating a response is: - .

5. The large model safety neuron screening method according to claim 3, characterized in that: The step of calculating the neuron activation difference between the secure alignment model and the basic large model based on the two generated responses specifically includes: Assume the basic model is , the safety alignment model is ; For a given input , respectively and express and Generate a response; In the full conversation Perform forward calculation and collect The corresponding neuron is activated, and we get Approximate neuron activations when generating responses , where the complete conversation is and The splicing of , Indicates that the input No. The word unit, No. Layer The activation of neurons; On the complete conversation Perform forward calculation and collect The corresponding neuron is activated, and we get Neuron activation when generating a response ,in, Indicates that the input No. The word unit, No. Layer The activation of neurons; calculate and The difference in neuron activations when generating a response is: - .

6. The large model safety neuron screening method according to claim 4 or 5, characterized in that: The step of determining the safe neurons for safe alignment of the basic large model based on the neuron activation difference specifically includes: Calculating a change score for each neuron based on the neuron activation difference; Neurons are ranked based on their change scores, and neurons with the highest change scores are selected as safe neurons based on their causal effects on the safety of the model.

7. The large model safety neuron screening method according to claim 6, characterized in that: The process of calculating the change score of each neuron based on the neuron activation difference specifically includes: make Indicates that the input No. The word unit, No. Layer The activation of neurons; For a given data set , defined based on Change score For middle and The root mean square of the differences in neuron activations when generating responses: , Where q = 1 or 2, correspondingly, Pick or .

8. The large model safety neuron screening method according to claim 7, characterized in that: Based on the causal effect of neurons on the safety of the model, the process of selecting neurons with the highest change scores as safe neurons includes: S1, select different numbers of neurons with the largest change scores as candidate safe neurons; S2. Perform a secure alignment model on a given input The forward computation process of and caches the activations of candidate safety neurons; S3. Execute the basic large model on the given input In the forward calculation process, the activation of a given neuron is replaced by the activation of the cached safe neuron in the previous step, and the forward calculation continues; S4: After the forward calculation is completed, the prediction of the next word is obtained, and the input is expanded with the predicted word and steps S2 to S3 are repeated until the generation is completed; S5. The minimum set of neurons with causal effects similar to those of all neurons is taken as safe neurons.

9. A large-scale model safety neuron screening device, characterized in that: include: A safety alignment module is used to safely align the basic large model to obtain a safety alignment model; A difference calculation module, used for calculating the difference in neuron activation between the secure alignment model and the basic large model; A determination module is used to determine the safe neurons when performing safe alignment on the basic large model based on the neuron activation difference.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the large-model safe neuron screening method as described in any one of claims 1 to 8 is implemented.