Large model security management method, system, device and program medium
By implementing security management during the input, generation, and output phases of large models, and utilizing intervention databases and security detection models, security issues throughout the entire lifecycle of large models are resolved, enabling effective risk interception and secure content filtering, thereby improving the user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-05
- Publication Date
- 2026-06-05
AI Technical Summary
Existing technologies lack comprehensive security management measures throughout the entire lifecycle of large models, leading to missed calls and a decline in user experience.
Security management is implemented in three stages of the large model: input, generation, and output. This includes matching detection, input security detection, risk level label identification, response generation, and output security checks. Intervention databases and security detection models are used for risk interception and security filtering of response answers.
This significantly reduces the risk of missed calls, improves user experience, and ensures the security and reliability of the content output by the large model.
Smart Images

Figure CN122153926A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of model security technology, specifically to a method, system, device, and program medium for large model security management. Background Technology
[0002] Currently, large-scale models are experiencing rapid development. With continuous technological advancements, large-scale models are becoming increasingly larger, with the number of parameters growing exponentially, enabling them to handle more complex and diverse tasks. Simultaneously, the application areas of large-scale models are constantly expanding. From natural language processing and computer vision to healthcare, finance, education, and many other fields, large-scale models are playing a vital role.
[0003] However, behind the booming development of large-scale models, the importance of their security throughout their entire lifecycle is becoming increasingly prominent. During the data acquisition phase, large models require massive amounts of data. If the data source is unreliable, or if the data is tampered with during acquisition, it will severely impact the model's accuracy and reliability. The model training phase also faces numerous security challenges. Malicious attacks or data contamination may occur during training, leading to training interruptions or corrupted training results. When large-scale models are deployed, content security is even more crucial. It ensures that the model's output complies with relevant regulations and ethical standards. For example, in social media content recommendation models, if inappropriate or harmful information is output, it could negatively impact public opinion.
[0004] In conclusion, the security of large-scale model content throughout its entire lifecycle is the cornerstone for ensuring the effective, reliable, and legal operation of large-scale models, and it is of paramount importance to the stability and development of individuals, enterprises, and society.
[0005] Existing technologies mostly perform security checks during the application of large models, which is a relatively efficient security measure for large models. However, security checks alone are far from sufficient. On the one hand, there is the risk of missed recalls, and on the other hand, false identifications can also affect the user experience. Summary of the Invention
[0006] The purpose of this application is to provide a large model security management method, system, device, and program medium to provide security protection for each stage of the large model, reduce risk of missed recalls, and improve user experience.
[0007] To achieve the above objectives, the first aspect of this application provides a large-scale model security management method, the method comprising:
[0008] Perform matching and detection on the input question;
[0009] If the matching detection passes, an input security check is performed on the question to obtain the risk level label of the question;
[0010] The risk level label and the question are input into the response model to generate a response answer;
[0011] Perform an output security check on the response answer to obtain the output check result;
[0012] The response answer will be pushed based on the output check results.
[0013] Based on the aforementioned technical means, this method performs matching detection on the input questions, eliminating data that fails the matching detection, thus achieving security management of the input questions. Questions that pass the matching detection then undergo input security detection to determine their risk level. The response model then provides a safe response to risky questions based on the risk level label, achieving security management during response generation. Finally, the response undergoes an output security check before being pushed out based on the output check results. This final check before the large model's generated results are provided to the user further ensures security management during response output. Therefore, by managing the input, generation, and output stages of the large model's question answering system, risky missed calls and improved user experience are significantly reduced.
[0014] In some feasible embodiments, the matching detection of the input question includes:
[0015] The input question is matched against intervention resources in a pre-built intervention database. If the question does not match any intervention resources in the database, the matching test is considered successful.
[0016] Based on the aforementioned technical means, the input questions are matched by intervening resources in the intervention database, and risky inputs are quickly intercepted, providing the first layer of security protection for the large model inference application stage.
[0017] In some feasible embodiments, the input security detection of the question includes:
[0018] The question is input into a security detection model for risk level identification to obtain a risk level label for the question; the security detection model is trained based on risk data and the corresponding risk level label.
[0019] Based on the above technical means, the risk level of the question is identified by the input security detection model. The risk level label of the question can help the response model avoid risky output when generating the response answer and optimize the generated content.
[0020] In some feasible embodiments, the step of inputting the risk level label and the question into the response model to generate a response includes:
[0021] The response strategy for the question is determined based on the risk level label.
[0022] When the response strategy is to generate an answer, a response answer is generated based on the question.
[0023] When the response strategy is to refuse to answer, no response answer will be generated.
[0024] Based on the aforementioned technical means, it is possible to determine whether a response can be generated according to the risk level label. This allows for the rejection of generating response responses for questions with high risk levels, thereby improving the security of the model's response.
[0025] In some feasible embodiments, determining the response strategy for the question based on the risk level label includes:
[0026] Determine whether the risk level label is lower than the preset level label. If so, determine that the response strategy for the question is to generate an answer; otherwise, determine that the response strategy for the question is to refuse to answer.
[0027] Based on the aforementioned technical means, a preset level label is set as the boundary for whether to generate an answer. If the risk level label is higher than the preset level label, it indicates that the risk level of the corresponding question is very high, which will bring greater risks to the large model. In order to avoid this risk, a response strategy of refusing to answer is adopted. Conversely, if the risk level of the corresponding question is lower, it indicates that the risk level of the corresponding question is low. In combination with the large model, a safe response can avoid the risk. Therefore, in order to improve the user experience, an answer generation strategy is adopted.
[0028] In some feasible embodiments, the output check result includes a risk label; the step of pushing a response answer based on the output check result includes:
[0029] Determine whether there is any risk in the response answer based on the risk label;
[0030] If the response is deemed risky, the risky content segments in the response are filtered out, and the response after filtering out the risky content segments is pushed out.
[0031] Alternatively, if the given response is risky, a new response can be generated based on the question.
[0032] If the response answer poses no risk, the response answer will be sent directly.
[0033] Based on the aforementioned technical means, the generated response can be checked before being provided to the user, ensuring the security of the data provided to the user. If the response poses a risk, risk avoidance can be achieved by directly filtering risky content segments or by regenerating the response. Directly filtering risky content segments before pushing them to the user is a simple and quick method. Regenerating the response ensures the security and readability of the content provided to the user.
[0034] In some feasible embodiments, the training method of the response model includes:
[0035] Obtain the first sample data, which is secure data;
[0036] The first sample data is input into the initial model for training to obtain the first model;
[0037] Obtain second sample data, which includes risk issues and corresponding security responses;
[0038] The second sample data is input into the first model for training to obtain the second model;
[0039] Obtain third sample data, which includes risk issues, corresponding security responses, and risk responses;
[0040] The third sample is input into the second model for training to obtain the response model.
[0041] Based on the aforementioned technical methods, training with secure data can minimize the amount of risky data in the training data and reduce its impact on model training. Training with risky questions and their corresponding secure responses allows the large model to respond to risky inputs according to the constructed secure response method, ensuring the security of the large model's responses and quickly and easily resolving security issues. Direct preference training using risky questions, their corresponding secure responses, and risky responses allows the large model to learn the content of secure responses while avoiding the content of risky responses, making the large model more inclined to output secure responses, ensuring the security of the large model's responses while also having higher generalization ability.
[0042] In some feasible embodiments, the step of inputting the third sample into the second model for training to obtain the response model includes:
[0043] The safe responses in the third sample are used as preferred data, and the risky responses in the third sample are used as non-preferred data;
[0044] The third sample is input into the second model for training.
[0045] By using the aforementioned technical means, treating safe responses as preferred data and risky responses as unpreferred data, the model can learn user preferences during training, thereby acquiring the ability to respond safely.
[0046] In some feasible embodiments, the step of inputting the third sample into the second model for training to obtain the response model includes:
[0047] A strategy model and a reference model are constructed based on the second model;
[0048] The third sample is input into the strategy model and the reference model respectively to obtain the first probability corresponding to the safe response in the strategy model, the second probability corresponding to the risk response in the strategy model, the third probability corresponding to the safe response in the reference model, and the fourth probability corresponding to the risk response in the reference model.
[0049] Calculate the loss function based on the first probability, second probability, third probability, and fourth probability;
[0050] Adjust the parameters of the policy model according to the loss function;
[0051] The strategy model obtained by iterating a preset number of times or until the loss function is less than the target value is used as the response model.
[0052] Based on the above technical means, the direct preference optimization method is used for training, which makes the training process faster and more efficient, and the trained model has universal safety response capability.
[0053] A second aspect of this application provides a large-scale model security management system, the system comprising:
[0054] Intervention unit, used to perform matching detection on the input question;
[0055] An input security detection unit is used to perform input security detection on the question if the matching detection passes, and obtain the risk level label of the question;
[0056] The response generation unit is used to input the risk level label and the question into the response model to generate a response answer;
[0057] An output security check unit is used to perform an output security check on the response answer and obtain the output check result.
[0058] The push unit is used to push response answers based on the output check results.
[0059] Based on the aforementioned technical methods, the intervention unit in this system performs matching detection on the input questions, eliminating data that fails the matching detection, thus achieving security management of the input questions. Questions that pass the matching detection are then checked by the input security detection unit to determine their risk level. The response generation unit then provides a safe response to risky questions based on the risk level label, achieving security management during response generation. Finally, the response is checked by the output security check unit before being pushed by the push unit based on the output check results. This final check before the large model's generated results are provided to the user further ensures security management during response output. Therefore, by managing the input, generation, and output stages of the large model's question answering system, risky missed calls and improved user experience are significantly reduced.
[0060] A third aspect of this application provides an electronic device, comprising:
[0061] The memory is configured to store instructions; and
[0062] The processor is configured to retrieve the instructions from the memory and, when executing the instructions, to implement the large-scale model security management method.
[0063] A fourth aspect of this application provides a machine-readable storage medium storing instructions that cause a machine to perform the large-scale model security management method.
[0064] The fifth aspect of this application provides a computer program product, including a computer program that, when executed by a processor, implements the large-scale model security management method.
[0065] Through the above technical solution, this method performs matching detection on the input questions, eliminating data that fails the matching detection, thus achieving security management of the input questions. Questions that pass the matching detection then undergo input security detection to determine their risk level. The response model then provides a safe response to risky questions based on the risk level label, achieving security management during response generation. Finally, the response undergoes an output security check before being pushed out based on the output check results. This final check before the large model's generated results are provided to the user further ensures security management during response output. Therefore, by managing the input, generation, and output stages of the large model's question answering system, risky missed calls and high recalls are significantly reduced, improving the user experience.
[0066] Other features and advantages of the embodiments of this application will be described in detail in the following detailed description section. Attached Figure Description
[0067] The accompanying drawings are provided to further illustrate the embodiments of this application and form part of the specification. They are used together with the following detailed description to explain the embodiments of this application, but do not constitute a limitation on the embodiments of this application. In the drawings:
[0068] Figure 1 The illustration shows a flowchart of a large-scale security management method according to an embodiment of this application;
[0069] Figure 2 This illustration schematically shows a large-scale model security management method according to an embodiment of the present application.
[0070] Figure 3 The schematic diagram illustrates the structure of a large-scale security management device according to an embodiment of this application. Detailed Implementation
[0071] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only for illustration and explanation of the embodiments of this application and are not intended to limit the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0072] It should be noted that if the embodiments of this application involve directional indicators (such as up, down, left, right, front, back, etc.), the directional indicators are only used to explain the relative positional relationship and movement of the components in a certain specific posture (as shown in the figure). If the specific posture changes, the directional indicators will also change accordingly.
[0073] Furthermore, if the embodiments of this application involve descriptions such as "first" or "second," these descriptions are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, features defined with "first" or "second" may explicitly or implicitly include at least one of those features. Additionally, the technical solutions of various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. If the combination of technical solutions is contradictory or impossible to implement, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed in this application.
[0074] Figure 1This illustration schematically depicts a flowchart of a large-scale model security management method according to an embodiment of this application. This method can be applied to terminal devices such as in-vehicle systems, smartphones, PDAs, tablets, laptops, all-in-one computers, and autonomous driving devices. It is understood that the large-scale model security management method provided in this disclosure can also be applied to other scenarios. For example... Figure 1 As shown in the figure, this application provides a large model security management method, which may include the following steps.
[0075] S1: Perform matching detection on the input question. In some feasible embodiments, the question refers to content obtained from external sources during the question-and-answer process by the large model. This could be a dialogue question input by the user's voice, or a text question input by the user through an interactive interface, etc. Matching detection refers to matching the input question with resources in the intervention database. Specifically:
[0076] The input question is matched against intervention resources in a pre-built intervention database. If the question does not match any intervention resource in the database, the matching test is considered successful. If the question matches an intervention resource in the database, the matching test is considered unsuccessful.
[0077] In some feasible embodiments, the intervention database can be a database composed of risk terms, risk statements, etc. Upon receiving a user-inputted question, the question is first matched against the data in the intervention database. If a preset matching condition is met, the question is determined to have matched an intervention resource in the database, and the question is then intercepted; otherwise, the question is allowed. The preset matching condition is set according to requirements. It can be a perfect match, meaning the question must match a specific data point in the intervention database with 100% accuracy. Alternatively, it can be a fuzzy match based on the question, meaning the question must match a specific data point in the intervention database with a preset value. Intercepted questions will not proceed to the next processing step.
[0078] In some feasible embodiments, such as Figure 2 As shown, the intervention platform can be used to perform matching and detection on the input data.
[0079] By intervening in the database, operations can quickly resolve unexpected issues (bad cases) that arise in large online models, including false positives and false negatives. Therefore, by intervening in the database to match input questions, risky inputs can be quickly intercepted, providing the first layer of security for large model inference applications.
[0080] S2: In the case where the matching detection passes, perform input security detection on the question sentence to obtain the risk level label of the question sentence. Input security detection is another detection of the question sentence. Matching detection mainly detects situations where the large model does not meet expectations, while input security detection is a risk identification detection for the question sentence itself.
[0081] In some feasible embodiments, the performing input security detection on the question sentence includes:
[0082] Input the question sentence into a security detection model for risk level identification to obtain the risk level label of the question sentence; the security detection model is trained based on risk data and corresponding risk level labels. The risk level refers to the security degree of the question sentence, and the risk level label is a sign used to represent the risk level. For example, in some feasible embodiments, the risk level can be divided into three levels: high risk, medium risk, and low risk. Among them, high risk represents the lowest security degree, and low risk represents the highest security degree. The risk level labels corresponding to each risk level can be set according to requirements. For example, the Chinese character "high" can be used as the label for high risk, and the Chinese character "low" can be used as the label for low risk. Other symbols, pictures, etc. can also be used as the risk level labels for each risk level. Thus, based on the input security detection model to identify the risk level of the question sentence, the obtained risk level label of the question sentence can help the reply model avoid risk output when generating a reply answer and optimize the generated content.
[0083] S3: Input the risk level label and the question sentence into a reply model to generate a reply answer. The reply model can perform a safe reply to a risky question sentence according to the risk level label. In the embodiment of the present application, the reply answer refers to the content generated by the reply model according to the question sentence.
[0084] Currently, it is impossible to recall 100% of the risks in input security detection, and there will still be problems of missed identification. Then, the large model itself needs to have a strong ability to reply safely. This requires training the large model's ability to reply safely to risk problems in the supervised fine-tuning stage, so that the large model can generate safe reply content for risk problems. In some feasible embodiments, after the large model is trained by PTST (a security training method that does not use prompts during training and uses prompts during inference), it can both identify the risks in the input and has the ability to directly generate a safe reply for a risky question sentence (query).
[0085] In some feasible embodiments, the inputting the risk level label and the question sentence into a reply model to generate a reply answer includes:
[0086] The response strategy for the question is determined based on the risk level label. In some feasible embodiments, for questions with a risk level label of medium to high risk, the response model may refuse to answer; for low-risk questions, the response model may avoid outputting risky content when generating a response.
[0087] When the response strategy is to generate an answer, a response answer is generated based on the question.
[0088] When the response strategy is to refuse to answer, no response answer will be generated.
[0089] Therefore, it is possible to determine whether to generate a response based on the risk level label. This allows us to refuse to generate responses for questions with high risk levels, thereby improving the safety of the model's response.
[0090] In some feasible embodiments, determining the response strategy for the question based on the risk level label includes:
[0091] The system determines whether the risk level label is lower than a preset risk level label. If so, the response strategy for the question is determined to generate an answer; otherwise, the response strategy is determined to refuse to answer. In some feasible embodiments, the preset risk level label is set according to security requirements. Taking three risk levels—high, medium, and low—as an example, if the security requirement is high, the medium risk level can be set as the preset risk level label. An answer can only be generated when the risk level label of the question is lower than the medium risk level; that is, only questions with a low risk level label can generate an answer, while questions with medium or high risk level labels cannot. If the security requirement is low, the high risk level can be set as the preset risk level label. An answer can be generated when the risk level label of the question is lower than the high risk level; that is, questions with low or medium risk level labels can generate an answer, while only questions with a high risk level label cannot generate an answer.
[0092] Based on the aforementioned technical means, a preset level label is set as the boundary for whether to generate an answer. If the risk level label is higher than the preset level label, it indicates that the risk level of the corresponding question is very high, which will bring greater risks to the large model. In order to avoid this risk, a response strategy of refusing to answer is adopted. Conversely, if the risk level of the corresponding question is lower, it indicates that the risk level of the corresponding question is low. In combination with the large model, a safe response can avoid the risk. Therefore, in order to improve the user experience, an answer generation strategy is adopted.
[0093] S4: Perform an output security check on the response answer to obtain the output check result. In this embodiment of the application, the output security check is the last check before outputting the response answer, and is used to check whether there are any risky content segments in the response answer.
[0094] Although the question has gone through the previous three stages, the creativity of the large model's generated results is a double-edged sword. It may produce amazing results, but it may also bring unexpected risks. In order to ensure the overall security of the system, it is still necessary to perform security checks on the responses generated by the large model.
[0095] In some feasible embodiments, the response answer can be checked sentence by sentence to determine whether there are any risky content segments. If risky content segments are found, the response answer is determined to be risky, and the output of the check result should indicate that the response answer is risky. If no risky content segments are found, the response answer is determined to be risk-free, and the output of the check result should indicate that the response answer is risk-free. In some feasible embodiments, risk labels can be used to indicate whether a response answer is risky. Risk labels can be text, images, etc. For example, the English character "S" can be used as a risk label for no risk, and the English character "D" can be used as a risk label for risky situations.
[0096] In some feasible embodiments, a trained inspection model can be used to inspect each sentence generated by the large model. This inspection model is trained on a large amount of labeled data. Therefore, the generated response can be checked before being provided to the user, ensuring the security of the data delivered to the user.
[0097] S5: Push the reply answer based on the output check results.
[0098] In some feasible embodiments, the output check result includes a risk label; the step of pushing a response answer based on the output check result includes:
[0099] The presence of risk in a response answer is determined based on risk labels. This can be achieved by establishing a correspondence between risk labels and the presence of risk. For example, assuming risk label "S" represents no risk and risk label "D" represents risk, then if risk label "S" is present in the output check results, the current response answer is considered risk-free; if risk label "D" is present, the current response answer is considered risky.
[0100] If a response contains risky content, the risky segments are filtered out, and the filtered response is pushed out. Directly filtering risky segments is simple and quick; the filtered response no longer poses a risk. However, this reduces the overall coherence and readability of the response.
[0101] Alternatively, if the given response poses a risk, a new response can be generated based on the question. Regenerating the response involves using the response model to re-generate an answer for the question. Since the large model's generation results are creative, the regenerated response may or may not be risky. If the regeneration option is chosen, the process should continue until the generated response is risk-free. This ensures that the content provided to users is both safe and readable.
[0102] If the response answer poses no risk, the response answer will be sent directly.
[0103] Based on the aforementioned technical means, this method performs matching detection on the input questions, eliminating data that fails the matching detection, thus achieving security management of the input questions. Questions that pass the matching detection then undergo input security detection to determine their risk level. The response model then provides a safe response to risky questions based on the risk level label, achieving security management during response generation. Finally, the response undergoes an output security check before being pushed out based on the output check results. This final check before the large model's generated results are provided to the user further ensures security management during response output. Therefore, by managing the input, generation, and output stages of the large model's question answering system, risky missed calls and improved user experience are significantly reduced.
[0104] To achieve more comprehensive security management for large models, security management is also performed during the offline training phase of the response model in this embodiment.
[0105] The offline training phase of the response model in this application includes a pre-training phase, a supervised-finetuning (SFT) training phase, and a preference alignment training phase. The pre-training phase typically refers to training a large model on an unlabeled text corpus to learn general language representations. The supervised-finetuning phase further refines the pre-trained large model using labeled task-specific data, giving it specific capabilities. Preference alignment training is a model optimization method used in the post-training phase, aiming to directly optimize the language model to conform to human preferences. A large model typically refers to a machine learning model with a massive number of parameters (usually over a billion) and complex computational structures. It is typically capable of processing massive amounts of data and performing various complex tasks, such as natural language processing and image recognition. In this embodiment, the large model refers to a Large Language Model (LLM). In this embodiment, the response model is a large model.
[0106] In some feasible embodiments, such as Figure 2As shown, the training method for the response model includes:
[0107] During the pre-training phase:
[0108] 1) Obtain first sample data, which is secure data. In some feasible embodiments, the first sample data may be data filtered by a security detection model, which is trained using risky data. In some feasible embodiments, the security detection model may be a classification model trained using risky data. When filtering the training data, the training data is input into the security detection model, which performs classification and identification to filter out risky training data, thus obtaining higher-quality training data. Therefore, a security detection model trained using risky data is used to quickly filter the training data, minimizing the risky data in the training data and reducing its impact on model training.
[0109] In other feasible embodiments, data quality can be further improved from the training data source, for example, by purchasing and collecting data from sources that are clearly identified and subject to oversight, thus avoiding data contamination.
[0110] 2) Input the first sample data into the initial model for training to obtain the first model. The first sample data is mainly used to train the initial model's generation ability, such as how to generate semantically clear sentences and the part-of-speech collocation of each word in the generated sentences. The less risk data in the first sample data, the less risk content the first model will learn.
[0111] In other feasible embodiments, the ability of the large model to identify risk data can be trained during the pre-training phase. For example, in the antenna risk identification task within the large model, the model can be trained based on risk data and corresponding labels to enable it to identify risk data. In some feasible embodiments, risk data refers to data that poses a safety hazard. During training, data selected by the safety detection model can be used as risk data. After adding corresponding labels to this data, it can be used for training the large model. If the amount of selected data is small, risk data can be collected directly for training.
[0112] During the supervised fine-tuning phase:
[0113] 3) Obtain second sample data, which includes risk issues and corresponding security responses. In some feasible embodiments, the security response does not involve a response to a security issue.
[0114] 4) Input the second sample data into the first model for training to obtain the second model. During the fine-tuning phase, a certain number of corresponding response methods are provided for risky questions. These response methods do not involve security issues. The large model is trained during the fine-tuning phase to respond according to these corresponding response methods, ensuring that the responses provided by the large model always do not involve security issues when dealing with risky questions. Therefore, the large model can be trained to respond to risky questions according to the constructed safe response methods, ensuring the security of the large model's responses and quickly and easily resolving security problems.
[0115] During the preference alignment training phase:
[0116] 5) Obtain third sample data, which includes risk issues, corresponding security responses, and risk responses. Security responses are responses that do not involve security issues, while risk responses are responses that involve security issues.
[0117] 6) Input the third sample into the second model for training to obtain the response model. During the preference alignment training process, safe responses are taken as human preferences, which can train the large model to learn safe responses while avoiding risky responses. This makes the large model more inclined to output safe responses, ensuring the safety of the large model's responses while having higher generalization ability.
[0118] In some feasible embodiments, the step of inputting the third sample into the second model for training to obtain the response model includes:
[0119] The safe responses in the third sample are used as preferred data, and the risky responses in the third sample are used as non-preferred data;
[0120] The third sample is input into the second model for training.
[0121] By using the aforementioned technical means, treating safe responses as preferred data and risky responses as unpreferred data, the model can learn user preferences during training, thereby acquiring the ability to respond safely.
[0122] In some feasible embodiments, the step of inputting the third sample into the second model for training to obtain the response model includes:
[0123] A strategy model and a reference model are constructed based on the second model. In the first round of iterations, the strategy model and the reference model are the same, and both are the same as the second model. In subsequent iterations, the parameters in the strategy model are adjusted as the iterations proceed, while the parameters in the reference model remain unchanged.
[0124] The third sample is input into the strategy model and the reference model respectively to obtain the first probability corresponding to the safe response in the strategy model, the second probability corresponding to the risk response in the strategy model, the third probability corresponding to the safe response in the reference model, and the fourth probability corresponding to the risk response in the reference model.
[0125] The loss function is calculated based on the first probability, second probability, third probability, and fourth probability. In the embodiments of this application, the first probability, second probability, third probability, and fourth probability are substituted into the loss function calculation formula of the standard direct preference optimization method to calculate the loss function.
[0126] The parameters of the policy model are adjusted according to the loss function; the goal of adjusting the parameters of the policy model is to make the loss function smaller and smaller.
[0127] The strategy model obtained by iterating a preset number of times or until the loss function is less than the target value is used as the response model.
[0128] Based on the above technical means, the direct preference optimization method is used for training, which requires less computation during the training process. The loss function of the standard direct preference optimization method converges faster, making the training process faster and more efficient. The trained model has universal safety recovery capabilities.
[0129] The second aspect of this application provides a large-scale model security management system, such as... Figure 3 As shown, the system includes:
[0130] The intervention unit is used to perform matching detection on the input question. In some feasible embodiments, the question refers to content obtained from external sources during the question-and-answer process of the large model. This could be a dialogue question input by the user's voice, or a text question input by the user through an interactive interface, etc. Matching detection refers to matching the input question with resources in the intervention database.
[0131] An input security detection unit is used to perform input security detection on the question if the matching detection passes, and obtain the risk level label of the question.
[0132] The response generation unit is used to input the risk level label and the question into the response model to generate a response answer. The response model can provide a safe response to the risk question based on the risk level label.
[0133] An output security check unit is used to perform an output security check on the response answer and obtain the output check result.
[0134] The push unit is used to push response answers based on the output check results.
[0135] Based on the aforementioned technical methods, the intervention unit in this system performs matching detection on the input questions, eliminating data that fails the matching detection, thus achieving security management of the input questions. Questions that pass the matching detection are then checked by the input security detection unit to determine their risk level. The response generation unit then provides a safe response to risky questions based on the risk level label, achieving security management during response generation. Finally, the response is checked by the output security check unit before being pushed by the push unit based on the output check results. This final check before the large model's generated results are provided to the user further ensures security management during response output. Therefore, by managing the input, generation, and output stages of the large model's question answering system, risky missed calls and improved user experience are significantly reduced.
[0136] In some feasible embodiments, the intervention unit is specifically used for:
[0137] The input question is matched against intervention resources in a pre-built intervention database. If the question does not match any intervention resource in the database, the matching detection is considered successful. Thus, by matching the input question against intervention resources in the database, risky inputs are quickly intercepted, providing the first layer of security for large-scale model inference applications.
[0138] In some feasible embodiments, the input security detection unit is specifically used to: input the question into a security detection model for risk level identification to obtain a risk level label for the question; the security detection model is trained based on risk data and the corresponding risk level label. Thus, by identifying the risk level of the question based on the input security detection model, the obtained risk level label can help the response model avoid risky outputs and optimize the generated content when generating response answers.
[0139] In some feasible embodiments, the response generation unit is further configured to:
[0140] The response strategy for the question is determined based on the risk level label.
[0141] When the response strategy is to generate an answer, a response answer is generated based on the question.
[0142] When the response strategy is to refuse to answer, no response answer will be generated.
[0143] Based on the aforementioned technical means, it is possible to determine whether a response can be generated according to the risk level label. This allows for the rejection of generating response responses for questions with high risk levels, thereby improving the security of the model's response.
[0144] In some feasible embodiments, determining the response strategy for the question based on the risk level label includes:
[0145] Determine whether the risk level label is lower than the preset level label. If so, determine that the response strategy for the question is to generate an answer; otherwise, determine that the response strategy for the question is to refuse to answer.
[0146] Based on the aforementioned technical means, a preset level label is set as the boundary for whether to generate an answer. If the risk level label is higher than the preset level label, it indicates that the risk level of the corresponding question is very high, which will bring greater risks to the large model. In order to avoid this risk, a response strategy of refusing to answer is adopted. Conversely, if the risk level of the corresponding question is lower, it indicates that the risk level of the corresponding question is low. In combination with the large model, a safe response can avoid the risk. Therefore, in order to improve the user experience, an answer generation strategy is adopted.
[0147] In some feasible embodiments, the output inspection result includes a risk label; the push unit is further configured to:
[0148] Determine whether there is any risk in the response answer based on the risk label;
[0149] If the response is deemed risky, the risky content segments in the response are filtered out, and the response after filtering out the risky content segments is pushed out.
[0150] Alternatively, if the given response is risky, a new response can be generated based on the question.
[0151] If the response answer poses no risk, the response answer will be sent directly.
[0152] Based on the aforementioned technical means, the generated response can be checked before being provided to the user, ensuring the security of the data provided to the user. If the response poses a risk, risk avoidance can be achieved by directly filtering risky content segments or by regenerating the response. Directly filtering risky content segments before pushing them to the user is a simple and quick method. Regenerating the response ensures the security and readability of the content provided to the user.
[0153] A third aspect of this application provides an electronic device, comprising:
[0154] The memory is configured to store instructions; and
[0155] The processor is configured to retrieve the instructions from the memory and, when executing the instructions, to implement the large-scale model security management method.
[0156] A fourth aspect of this application provides a machine-readable storage medium storing instructions that cause a machine to perform the large-scale model security management method.
[0157] The fifth aspect of this application provides a computer program product, including a computer program that, when executed by a processor, implements the large-scale model security management method.
[0158] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0159] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0160] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0161] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0162] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0163] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0164] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0165] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0166] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A large-scale model security management method, characterized in that, The method includes: Perform matching and detection on the input question; If the matching detection passes, an input security check is performed on the question to obtain the risk level label of the question; The risk level label and the question are input into the response model to generate a response answer; Perform an output security check on the response answer to obtain the output check result; The response answer will be pushed based on the output check results.
2. The method according to claim 1, characterized in that, The matching and detection of the input question includes: The input question is matched against intervention resources in a pre-built intervention database. If the question does not match any intervention resources in the database, the matching test is considered successful.
3. The method according to claim 1, characterized in that, The input security check on the question includes: The question is input into a security detection model for risk level identification to obtain a risk level label for the question; the security detection model is trained based on risk data and the corresponding risk level label.
4. The method according to any one of claims 1-3, characterized in that, The step of inputting the risk level label and the question into the response model to generate a response includes: The response strategy for the question is determined based on the risk level label. When the response strategy is to generate an answer, a response answer is generated based on the question. When the response strategy is to refuse to answer, no response answer will be generated.
5. The method according to claim 4, characterized in that, The step of determining the response strategy for the question based on the risk level label includes: Determine whether the risk level label is lower than the preset level label. If so, determine that the response strategy for the question is to generate an answer; otherwise, determine that the response strategy for the question is to refuse to answer.
6. The method according to any one of claims 1-5, characterized in that, The output inspection results include risk labels; The step of pushing out response answers based on the output check results includes: Determine whether there is any risk in the response answer based on the risk label; If the response is deemed risky, the risky content segments in the response are filtered out, and the response after filtering out the risky content segments is pushed out. Alternatively, if the given response is risky, a new response can be generated based on the question. If the response answer poses no risk, the response answer will be sent directly.
7. The method according to any one of claims 1-6, characterized in that, The training method for the response model includes: Obtain the first sample data, which is secure data; The first sample data is input into the initial model for training to obtain the first model; Obtain second sample data, which includes risk issues and corresponding security responses; The second sample data is input into the first model for training to obtain the second model; Obtain third sample data, which includes risk issues, corresponding security responses, and risk responses; The third sample is input into the second model for training to obtain the response model.
8. The method according to claim 7, characterized in that, The step of inputting the third sample into the second model for training to obtain the response model includes: The safe responses in the third sample are used as preferred data, and the risky responses in the third sample are used as non-preferred data; The third sample is input into the second model for training.
9. The method according to claim 7, characterized in that, The step of inputting the third sample into the second model for training to obtain the response model includes: A strategy model and a reference model are constructed based on the second model; The third sample is input into the strategy model and the reference model respectively to obtain the first probability corresponding to the safe response in the strategy model, the second probability corresponding to the risk response in the strategy model, the third probability corresponding to the safe response in the reference model, and the fourth probability corresponding to the risk response in the reference model. Calculate the loss function based on the first probability, second probability, third probability, and fourth probability; Adjust the parameters of the policy model according to the loss function; The strategy model obtained by iterating a preset number of times or until the loss function is less than the target value is used as the response model.
10. A large-scale model safety management system, characterized in that, The system includes: Intervention unit, used to perform matching detection on the input question; An input security detection unit is used to perform input security detection on the question if the matching detection passes, and obtain the risk level label of the question; The response generation unit is used to input the risk level label and the question into the response model to generate a response answer; An output security check unit is used to perform an output security check on the response answer and obtain the output check result. The push unit is used to push response answers based on the output check results.
11. An electronic device, characterized in that, include: The memory is configured to store instructions; as well as A processor is configured to retrieve the instructions from the memory and, when executing the instructions, to implement the large model security management method of any one of claims 1 to 9.
12. A machine-readable storage medium, characterized in that, The machine-readable storage medium stores instructions for causing the machine to perform the large model security management method as described in any one of claims 1 to 9.
13. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the large model security management method as described in any one of claims 1 to 9.