Corpus construction method and device, equipment, storage medium and product
By acquiring a set of risky questions and using a large response model to generate safe responses, a safe training corpus is constructed. This solves the problem of low efficiency in manual annotation of large language models, enables rapid acquisition of safe training corpus, and improves the model training progress and output security.
Patent Information
- Application Number
- CN202510308346.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-03-14
AI Technical Summary
When existing large language models generate risky content such as politically sensitive and socially ethical conflicts in open environments, the efficiency of manually annotating safe corpora is low, resulting in slow training progress and an inability to obtain enough safe corpora for iteration in a timely manner.
By acquiring a set of risky questions, generating original responses using a large response model, and then constructing a large model from the corpus to obtain safe responses, a safe training corpus is built. The automated process improves the efficiency of corpus construction and reduces the cost of manual annotation.
It enables the rapid acquisition of large amounts of secure training data, improves model training progress, reduces manpower costs, enhances resource utilization efficiency, and strengthens the security and reliability of model output.
Smart Images

Figure CN120234389B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and in particular to a corpus construction method and device, equipment, storage medium and product. BACKGROUND
[0002] As a leading model in the field of generative models, large language models have shown excellent performance in natural language processing, dialogue generation and many other tasks. However, when it runs in an open environment, it also exposes some security risks. For example, the model may generate risky content such as political sensitive and social ethical conflicts under various risk problem attacks.
[0003] To effectively address this problem, the large language model can be fine-tuned using artificially annotated safety corpus to help the model generate more secure and compliant output. However, the efficiency of artificially annotating safety corpus is extremely low, resulting in slow growth of safety corpus available for training, so that the large language model cannot obtain enough new safety corpus for iteration in time, and the training progress is seriously slowed down.
[0004] The above content is only used to assist in understanding the technical solutions of the present application, and does not represent the acknowledgement of the above content as prior art. SUMMARY
[0005] The main purpose of the present application is to provide a corpus construction method, device, equipment, storage medium and product, which can greatly improve the efficiency of constructing model training corpus.
[0006] To achieve the above purpose, the present application provides a corpus construction method, which comprises:
[0007] Obtain a set of risk problems, the set of risk problems comprising a plurality of risk problems, the risk problems being used to guide the reply large model to generate risky content;
[0008] Generate original replies corresponding to each risk problem through the reply large model;
[0009] Obtain safety replies corresponding to each risk problem through a corpus construction large model based on each risk problem and the original reply corresponding to each risk problem;
[0010] Construct safety training corpus of the reply large model based on each risk problem and the safety reply corresponding to each risk problem.
[0011] Optionally, the corpus construction large model comprises a paraphrase large model.
[0012] The corpus construction large model obtains safety replies corresponding to each risk problem based on each risk problem and the original reply corresponding to each risk problem, comprising:
[0013] By rewriting the large model, the original responses corresponding to each risk issue are rewritten based on each risk issue to obtain the security responses corresponding to each risk issue.
[0014] Optionally, the training process of rewriting the large model includes:
[0015] Obtain training samples, which include sample risk issues and corresponding sample safety responses to the sample risk issues;
[0016] The original responses to the sample risk issues are generated using the aforementioned response model.
[0017] By rewriting the large model and based on the sample risk issue, the original response of the sample is rewritten to obtain the rewritten response of the sample;
[0018] While keeping the parameters of the large response model unchanged, the parameters of the large rewrite model are adjusted based on the sample rewritten response and the sample safe response to reduce the difference between the sample rewritten response and the sample safe response.
[0019] Optionally, the number of original responses corresponding to each risk question generated by the response big model is multiple, and the corpus-based big model includes a reward big model;
[0020] The process involves constructing a large model from a corpus, and based on each risk question and its corresponding original response, obtaining the corresponding safe response for each risk question, including:
[0021] Using the aforementioned reward model, preference scores are assigned to multiple original responses corresponding to each risk question, resulting in preference scores for each original response.
[0022] The original response with the highest preference score is selected from the multiple original responses corresponding to each risk question as the safety response.
[0023] Optionally, the training process of the large reward model includes:
[0024] The large response model is used to sequentially generate multiple responses corresponding to sample risk questions.
[0025] Obtain the preference scores labeled for each of the multiple responses;
[0026] The reward model is trained based on the risk question, the multiple responses, and the preference scores labeled for each of the multiple responses.
[0027] Optionally, the corpus is used to construct a large model, including a security assessment model;
[0028] The process involves constructing a large model from a corpus, and based on each risk question and its corresponding original response, obtaining the corresponding safe response for each risk question, including:
[0029] The security assessment model described above is used to conduct a security assessment on the original response corresponding to any risk issue.
[0030] If the evaluation results indicate that the content of the original response is safe, the original response will be identified as the safe response;
[0031] Alternatively, if the assessment result indicates that the content of the original response is risky, the response corresponding to the risky issue is regenerated using the response big model, and a security assessment is performed on the newly generated response until the response generated by the response big model is assessed as a safe response by the security assessment big model.
[0032] Optionally, the training process of the large security assessment model includes:
[0033] The large response model is used to sequentially generate multiple responses corresponding to sample risk questions.
[0034] Obtain the risk assessment level and the reason for the risk assessment for each response;
[0035] The security assessment model is trained based on the sample risk questions, the multiple responses to the sample risk questions, and the risk assessment level and reason for each response.
[0036] Optionally, the security training corpus for constructing the large-scale response model based on each risk issue and its corresponding security response includes:
[0037] The security responses corresponding to each risk issue are used as training labels for each risk issue.
[0038] The supervised training corpus of the response model is constructed based on each risk issue and its corresponding training label.
[0039] Optionally, the security training corpus for constructing the large-scale response model based on each risk issue and its corresponding security response includes:
[0040] The preference training corpus of the response model is constructed based on each risk question, the original response to each risk question, and the safe response. The original response to each risk question and the safe response to each risk question are different.
[0041] Optionally, the method further includes:
[0042] In response to a dialogue instruction, extract the target question from the dialogue instruction;
[0043] Search the reference whitelist for reference questions that match the target question;
[0044] If a reference question matching the target question is found from the reference whitelist, the response to the target question is generated based on the security response corresponding to the reference question in the whitelist using the response big model.
[0045] Optionally, the method further includes:
[0046] If no reference question matching the target question can be found in the reference whitelist, a reference question matching the target question will be searched in the reference blacklist.
[0047] If a reference question matching the target question is found from the reference blacklist, a default response is generated, the content of which indicates that the target question is refused to be answered.
[0048] Optionally, the method further includes:
[0049] If no reference question matching the target question can be found in the reference blacklist, the target question is subjected to risk detection using a large risk detection model to obtain the risk detection result.
[0050] When the risk detection result indicates that the target problem has a specified type of security risk, a response corresponding to the target problem is generated by a security response model, which is obtained by training the response model with the security training corpus.
[0051] Furthermore, to achieve the above objectives, this application also proposes a corpus construction apparatus, the apparatus comprising:
[0052] The question set acquisition module is used to acquire a risk question set, which includes multiple risk questions. These risk questions are used to guide the response model to generate risk content.
[0053] The original response generation module is used to generate original responses for each risk issue based on the large response model.
[0054] The security response acquisition module is used to build a large model from the corpus and obtain the security response corresponding to each risk question based on each risk question and the original response corresponding to each risk question.
[0055] The training corpus construction module is used to construct the security training corpus of the large response model based on each risk issue and the corresponding security response.
[0056] Optionally, constructing a large model from the corpus includes rewriting the large model;
[0057] The security response acquisition module is used to rewrite the original responses corresponding to each risk issue based on the rewritten large model to obtain the security responses corresponding to each risk issue.
[0058] Optionally, the device further includes:
[0059] The first model training module is used to acquire training samples, which include sample risk issues and corresponding sample safety responses; generate original sample responses corresponding to the sample risk issues using the response big model; rewrite the original sample responses based on the sample risk issues using the rewriting big model to obtain rewritten sample responses; and adjust the parameters of the rewriting big model based on the rewritten sample responses and the sample safety responses while keeping the parameters of the response big model unchanged, so as to reduce the difference between the rewritten sample responses and the sample safety responses.
[0060] Optionally, the number of original responses corresponding to each risk question generated by the response big model is multiple, and the corpus-based big model includes a reward big model;
[0061] The safe response acquisition module is used to assign preference scores to multiple original responses corresponding to each risk question using the reward model, and obtain a preference score for each original response; and select the original response with the highest preference score from the multiple original responses corresponding to each risk question as the safe response.
[0062] Optionally, the device further includes:
[0063] The second model training module is used to generate multiple responses corresponding to sample risk questions sequentially through the response big model; obtain preference scores labeled for the multiple responses respectively; and train the reward big model based on the risk questions, the multiple responses and the preference scores labeled for the multiple responses respectively.
[0064] Optionally, the corpus is used to construct a large model, including a security assessment model;
[0065] The security response acquisition module is used to perform a security assessment on the original response corresponding to any risk issue using the security assessment model; if the assessment result indicates that the content of the original response is safe, the original response is determined as the safe response; or, if the assessment result indicates that the content of the original response is risky, the module regenerates the response corresponding to the risk issue using the response model and performs a security assessment on the newly generated response until the response generated by the response model is assessed as a safe response by the security assessment model.
[0066] Optionally, the device further includes:
[0067] The third model training module is used to sequentially generate multiple responses corresponding to the sample risk question through the response big model; obtain the risk assessment level and risk assessment reason for each response; and train the security assessment big model based on the sample risk question, the multiple responses corresponding to the sample risk question, and the risk assessment level and risk assessment reason for each response.
[0068] Optionally, the training corpus construction module is used to use the security responses corresponding to each risk issue as training labels for each risk issue; and to construct the supervised training corpus of the response big model based on each risk issue and the training labels corresponding to each risk issue.
[0069] Optionally, the training corpus construction module is used to construct the preference training corpus of the response big model based on each risk question, the original response corresponding to each risk question, and the safe response, wherein the content of the original response corresponding to each risk question is different from that of the safe response corresponding to each risk question.
[0070] Optionally, the device further includes:
[0071] The instruction response module is used to extract the target question from the dialogue instruction in response to the dialogue instruction;
[0072] The issue matching module is used to search for reference issues that match the target issue from a reference whitelist;
[0073] The response generation module is used to generate a response to the target question based on the security responses corresponding to the reference questions in the whitelist when a reference question matching the target question is found from the whitelist.
[0074] Optionally, the device further includes:
[0075] The issue matching module is also used to search for a reference issue that matches the target issue from the reference blacklist when no reference issue matching the target issue can be found from the reference whitelist.
[0076] The response generation module is also used to generate a default response when a reference question matching the target question is found from the reference blacklist, the content of which indicates that the target question is rejected.
[0077] Optionally, the device further includes:
[0078] The response generation module is also used to perform risk detection on the target question using a risk detection big data model when no reference question matching the target question can be found in the reference blacklist, and to obtain a risk detection result; when the risk detection result indicates that the target question has a specified type of security risk, the module generates a response corresponding to the target question using a security response big data model, wherein the security response big data model is obtained by training the response big data model with the security training corpus.
[0079] In addition, to achieve the above objectives, this application also proposes a corpus construction apparatus, the apparatus comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the corpus construction method as described above.
[0080] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and which, when executed by a processor, implements the steps of the corpus construction method described above.
[0081] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the corpus construction method described above.
[0082] One or more technical solutions proposed in this application have at least the following technical effects:
[0083] This application provides a scheme for automatically constructing a corpus. First, a set of risk questions is obtained, comprising multiple risk questions used to guide a large-scale response model in generating risky content. Then, the risk questions are processed; that is, the large-scale response model generates original responses for each risk question. Since these responses may contain risky content, a large-scale model is constructed using the corpus. Based on each risk question and its corresponding original responses, safe responses for each risk question are obtained. This automated process allows for the rapid acquisition of a large number of safe responses for risk questions. Based on this, a safe training corpus for the large-scale response model is constructed, significantly improving the efficiency of obtaining safe training corpus and thus accelerating the model's training progress. Furthermore, this scheme, by automatically generating safe training corpus, reduces the significant human cost associated with manual annotation and improves resource utilization efficiency. Attached Figure Description
[0084] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0085] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0086] Figure 1 This is a schematic diagram of an implementation environment for the corpus construction method of this application;
[0087] Figure 2 This is a flowchart illustrating the first embodiment of the corpus construction method of this application.
[0088] Figure 3 This is a flowchart illustrating the second embodiment of the corpus construction method of this application.
[0089] Figure 4 This is a flowchart illustrating the third embodiment of the corpus construction method of this application.
[0090] Figure 5 This is a flowchart illustrating the fourth embodiment of the corpus construction method of this application.
[0091] Figure 6 This is a flowchart illustrating the fifth embodiment of the corpus construction method of this application.
[0092] Figure 7 This is a flowchart illustrating the sixth embodiment of the corpus construction method of this application.
[0093] Figure 8 This is a schematic diagram of a newly added step in the seventh embodiment of the corpus construction method of this application;
[0094] Figure 9 This is a schematic diagram of a model training process provided in some embodiments of this application;
[0095] Figure 10 This is a schematic diagram of the module structure of the corpus construction device according to an embodiment of this application;
[0096] Figure 11 This is a schematic diagram of the device structure of the hardware operating environment involved in the corpus construction method in the embodiments of this application.
[0097] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0098] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.
[0099] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0100] Figure 1 This is a schematic diagram illustrating an implementation environment provided by an embodiment of this disclosure. See also... Figure 1 The implementation environment includes a terminal 101 and a server 102. The terminal 101 and the server 102 are connected via a wireless or wired network. For example, the terminal 101 can be a computer, mobile phone, tablet computer, or other terminal.
[0101] For example, terminal 101 has a target application installed on it, which is provided by server 102. For example, the target application is either an application in the operating system of terminal 101 or a target application provided by a third party. This target application has a question-and-answer function, capable of generating corresponding answers based on the input user question. For example, the target application is a search application, a short video application, a chat application, etc. For example, server 102 is the backend server corresponding to this target application. Accordingly, server 102 is a search application server, a short video application server, a chat application server, etc.
[0102] In this application, server 102 is used to obtain a set of risk questions, which includes multiple risk questions used to guide the response model to generate risky content. Then, the response model generates the original responses corresponding to each risk question. Next, a large model is constructed using corpus data, and based on each risk question and its corresponding original responses, a safe response is obtained for each risk question. Then, based on each risk question and its corresponding safe response, a safe training corpus for the response model is constructed. Afterwards, server 102 can use the safe training corpus to perform safe training on the response model and respond to various questions from terminal 101 based on the trained response model. Alternatively, the above corpus construction scheme can also be completed by terminal 101. This application embodiment does not limit this.
[0103] The solution provided in this application has a wide range of applications and can help improve the security and reliability of the output content of large-scale response models in multiple fields. For example, in intelligent customer service scenarios, intelligent customer service systems may encounter various problems during communication with users. Some problems may contain sensitive information or lead the customer service system to generate inappropriate responses. By automatically constructing a large amount of secure training corpus using this solution, and then using this corpus to securely train the large-scale response model in the customer service system, inappropriate responses can be avoided, effectively improving the security of the output content and enhancing service quality and user experience. Another example is in education scenarios. When providing learning assistance to students, intelligent tutoring systems need to ensure the legality of the answers. By obtaining secure training corpus through the corpus construction method provided in this application and then securely training the large-scale response model in the intelligent tutoring system, harmful information can be effectively avoided, ensuring the quality of education.
[0104] Figure 2 This is a flowchart illustrating the first embodiment of the corpus construction method of this application. (Refer to...) Figure 2 Taking the server as the executing entity as an example, the corpus construction method includes the following steps S10 to S40:
[0105] Step S10: Obtain a set of risk questions. The set of risk questions includes multiple risk questions, which are used to guide the response model to generate risk content.
[0106] The risk question set is a collection of multiple risk questions. These risk questions are specifically collected and organized, designed to guide the response model to generate responses with a risky nature, such as responses containing politically sensitive, socially ethical conflicts, violent, or pornographic information. By constructing such a set, the model's performance in the face of various risk questions can be tested and optimized in a focused manner.
[0107] Risk questions are specific questions that may induce the model to generate responses containing risky content. These questions cover risk scenarios across different domains and are designed to detect and improve the model's ability to handle potential risks. For example, asking the model for details of unverified rumors about certain sensitive political events, or involving ethical dilemmas that violate public order and good morals, are all risk questions.
[0108] A large response model (LLM) is a type of large language model with natural language processing capabilities, used to generate corresponding responses based on an input question. In this approach, it takes a risky question as input and generates raw responses. These raw responses may contain risky content, which will be processed and optimized subsequently. Large language models are language processing models with a massive parameter scale, built on deep learning techniques. Through training on large-scale text data, they learn knowledge of language syntax, semantics, and functions, enabling them to perform various natural language processing tasks such as text generation, question answering, and translation. Examples of large language models include GPT-4 (Generative Pretrained Transformer 4) and Claude.
[0109] Optionally, obtaining a set of risk issues includes: extracting risk issues from the business log dataset of the question-answering system; or extracting risk issues from publicly available domain datasets used for building model content security; or obtaining manually constructed risk issues; or obtaining risk issues generated by attacking large models; and constructing a set of risk issues based on the obtained risk issues.
[0110] The business log dataset of the question-and-answer system is a collection of various business-related data recorded during the system's operation. It includes detailed information about user interactions with the system, such as user questions and system responses. Since the question-and-answer system may encounter various user inputs, including those with malicious intent, these attack-related records constitute a potential source of risk. For example, malicious users may attempt to induce the question-and-answer system to output sensitive information or inappropriate content through specific questions; these interaction records can be extracted from the business log dataset as risk issues.
[0111] Domain-specific public datasets for building model content security are datasets publicly shared by research institutions or enterprises to improve model content security. These datasets focus on various content security risks that models may face, collecting and organizing a large amount of risk-related sample data, including multiple types of risk issues. Examples of such public datasets include SafetyBench, Safety-Prompts, and SORRY-Bench (Safety and Openness Reliability Risk Yield-Benchmark).
[0112] Artificially constructed risks are risks that professionals design and create specifically for real-world risk scenarios or model vulnerabilities. For example, security experts might construct a series of samples that cleverly configure instructions to bypass model security filtering mechanisms in order to test the model's security.
[0113] The risk issues generated by attack large models refer to the risks that can be generated by attack large models through designed attack methods, such as prompting manipulation, role-playing, and command injection. An attack large model is a large language model that has been trained specifically to generate risk issues.
[0114] In this embodiment, risk issues are obtained from three different sources: the business log dataset of the question-answering system, the publicly available domain dataset, and manually constructed samples, as well as large attack models. This significantly enriches the content of the risk issue set. The business log dataset reflects real-world attack scenarios; the publicly available domain dataset gathers industry research findings on model security risks; manually constructed samples can specifically simulate various potential attack scenarios; and the large attack models generate a wide variety of attack samples, covering a broad range of risk scenarios. This fusion of multi-source data makes the risk issue set more comprehensive in terms of the types of risks it covers.
[0115] Step S20: Generate the original responses for each risk issue using the large response model.
[0116] The original responses are the initial answers generated by the response model for risky issues. Due to the specific nature of risky issues, these responses often contain various risks, such as containing incorrect information, sensitive content, or statements that violate ethical norms, and require further processing to ensure they meet security and compliance requirements.
[0117] Step S30: Construct a large model using the corpus, and obtain the security response for each risk question based on each risk question and its corresponding original response.
[0118] Large-scale corpus modeling is a method for constructing and processing large language models. It uses risky questions and their corresponding original responses as input, and through specific algorithms and training mechanisms, it filters or corrects the original responses to obtain safe responses, thus providing a foundation for constructing a safe training corpus.
[0119] Safe responses are responses that meet security requirements and are generated for risky issues after being processed by a large-scale model built from a corpus. These responses do not contain sensitive information, comply with ethical standards, and do not violate laws and regulations. They can be used to train the response model, enabling it to generate safe and compliant content when faced with similar risky issues.
[0120] Step S40: Based on each risk issue and its corresponding security response, construct a security training corpus for the response big model.
[0121] The security training corpus is a database consisting of risky questions and their corresponding safe responses. Of course, the security training corpus can also include other relevant information besides risky questions and their corresponding safe responses. This corpus is used to train a large-scale response model, allowing the model to learn how to generate safe and compliant answers when faced with risky questions, thereby improving the model's security and reliability and reducing the likelihood of generating risky content in real-world applications.
[0122] One or more technical solutions proposed in this application have at least the following technical effects:
[0123] This application provides a scheme for automatically constructing a corpus. First, a set of risk questions is obtained, comprising multiple risk questions used to guide a large-scale response model in generating risky content. Then, the risk questions are processed; that is, the large-scale response model generates original responses for each risk question. Since these responses may contain risky content, a large-scale model is constructed using the corpus. Based on each risk question and its corresponding original responses, safe responses for each risk question are obtained. This automated process allows for the rapid acquisition of a large number of safe responses for risk questions. Based on this, a safe training corpus for the large-scale response model is constructed, significantly improving the efficiency of obtaining safe training corpus and thus accelerating the model's training progress. Furthermore, this scheme, by automatically generating safe training corpus, reduces the significant human cost associated with manual annotation and improves resource utilization efficiency.
[0124] Based on the first embodiment described above, a second embodiment of this application is proposed. Contents that are the same as or similar to the first embodiment can be referred to the above description and will not be repeated hereafter. (Refer to...) Figure 3 In the second embodiment, constructing a large model from the corpus includes rewriting the large model, and the above step S30 is refined into step S301:
[0125] Step S301: By rewriting the large model, the original responses corresponding to each risk issue are rewritten based on each risk issue to obtain the security responses corresponding to each risk issue.
[0126] Rewriting large-scale models are a type of corpus-based large-scale model. Their main function is to rewrite the original responses generated by the response large-scale model based on the input risk question. By applying specific algorithms and rules, risky content in the original response is removed or transformed into expressions that meet safety requirements, thus obtaining a safe response. For example, the server inputs the risk question and its corresponding original response into the rewriting large-scale model, and the model outputs a safe response.
[0127] Optionally, the training process of rewriting the large model includes: acquiring training samples, which include sample risk questions and corresponding sample safety responses; generating original sample responses corresponding to sample risk questions using the response large model; rewriting the original sample responses based on the sample risk questions using the rewriting large model to obtain rewritten sample responses; and adjusting the parameters of the rewriting large model based on the rewritten sample responses and sample safety responses while keeping the parameters of the response large model unchanged, so as to reduce the difference between the rewritten sample responses and sample safety responses.
[0128] Training samples are the dataset used to train and rewrite large-scale models. They contain sample risk issues and their corresponding sample safety responses. These samples are collected and organized from a large number of real-world or simulated scenarios, covering various potential risky issues and their corresponding safety responses. They form the foundational data for training and rewriting large-scale models.
[0129] Risky sample questions are part of the training samples; they are either manually designed or extracted from real-world scenarios and may guide the large-scale response model to generate questions containing risky content. By using these risky sample questions to train the large-scale model, it can adapt to various potential risk scenarios, improving its rewriting capabilities. Safe sample responses, corresponding to risky sample questions, are correct and safe responses that have been manually reviewed or formulated according to safety standards. These responses provide learning targets for the large-scale model, allowing it to continuously optimize towards generating such safe responses during training.
[0130] The original sample response is the initial response generated by the large response model in response to the sample risk problem. Due to the specific nature of the sample risk problem, these responses often contain various risky elements and need to be processed by rewriting the large model to transform them into responses that meet safety requirements. The rewritten sample response is the response obtained by rewriting the large model based on the sample risk problem, modifying the original sample response. During training, the parameters of the rewritten large model are continuously adjusted to make the rewritten sample response as close as possible to the sample safe response.
[0131] For example, while keeping the parameters of the large response model unchanged, the parameters of the large rewrite model are adjusted based on the rewritten sample response and the safe sample response to reduce the difference between them. This includes: determining a loss value based on the difference between the rewritten sample response and the safe sample response; and adjusting the parameters of the large rewrite model while keeping the parameters of the large response model unchanged to reduce this loss value. This makes the rewritten sample response increasingly closer to the safe sample response.
[0132] In this embodiment, by comparing the rewritten responses with the safe responses, and adjusting the parameters of the large-scale rewriting model based on the differences between the two, the model learns to more accurately transform original responses containing risky content into safe responses. As training progresses, the responses generated by the large-scale rewriting model become increasingly closer to the ideal safe responses, thereby improving the accuracy and reliability of its rewriting. Furthermore, this embodiment trains the large-scale rewriting model while keeping the parameters of the response model constant, avoiding the complexity and uncertainty caused by simultaneous changes in the parameters of two models. This allows for more focused optimization of the large-scale rewriting model's parameters, reduces interference factors during training, improves training efficiency, and enables the large-scale rewriting model to converge to a better state more quickly, shortening the training cycle.
[0133] In this embodiment, traditional processing methods require extensive manual review and modification of original responses to obtain safe responses, resulting in extremely high labor costs. This application, however, utilizes a large-scale rewriting model to automatically rewrite original responses, quickly transforming numerous risky original responses into safe ones, significantly reducing manual intervention. This not only saves labor costs but also avoids inaccuracies caused by human fatigue or subjective factors, improving processing efficiency and accuracy.
[0134] Based on the first embodiment of this application described above, a third embodiment of this application is proposed. Contents that are the same as or similar to the first embodiment can be referred to the above description, and will not be repeated hereafter. See also... Figure 4 In the third embodiment, the corpus construction large model includes a reward large model, the above step S20 is refined into step S201, and step S30 includes steps S302 and S303.
[0135] Step S201: Using the large response model, generate multiple original responses corresponding to each risk issue.
[0136] In this embodiment, the large response model generates multiple original responses for each risk question; that is, the number of original responses generated by the large response model for each risk question is multiple. For example, the large response model generates multiple original responses sequentially based on each risk question.
[0137] Step S302: Using the reward model, assign preference scores to the multiple original responses corresponding to each risk question to obtain the preference score for each original response.
[0138] The Reward Model is a model trained on human preferences, used to evaluate and score the original responses generated by the Reward Model. It analyzes each original response based on learned human preferences for safe, compliant, and high-quality responses, assigning a preference score—a quantitative evaluation metric—to each original response. A higher preference score indicates a higher quality original response according to the preference standards learned by the Reward Model, and a greater likelihood of it being a safe response.
[0139] Step S303: Select the original response with the highest preference score from the multiple original responses corresponding to each risk question as the safe response.
[0140] The original response with the highest preference score is the one that the reward model considers the safest and most in line with human preferences.
[0141] Optionally, the training process of the reward big model includes: the server generating multiple responses corresponding to the sample risk question in sequence through the response big model; obtaining preference scores labeled for each of the multiple responses; and training the reward big model based on the risk question, the multiple responses, and the preference scores labeled for each of the multiple responses.
[0142] The response model generates multiple different responses sequentially based on the input sample risk question, utilizing its learned language knowledge and patterns. Due to language diversity and the inherent characteristics of the model, the multiple responses generated by the response model when faced with the same question vary in content, quality, and safety. Humans assign quantified preference scores to each of the responses generated by the response model using specific evaluation criteria. Higher preference scores indicate that the response better meets expectations in various aspects, such as being more accurate, safer, and more ethical. This score is used to train the reward model, allowing it to learn human preference standards.
[0143] In this embodiment, by training with a large number of sample risk questions, multiple responses to each sample risk question, and preference scores labeled for each response, the reward-based large model can learn a rich variety of response scenarios and their corresponding human preferences. This enables the reward-based large model to more accurately judge the quality and safety of new responses when evaluating them, giving scores that are more in line with reality, thereby improving the accuracy of selecting safe and high-quality responses.
[0144] In this embodiment, the reward-based large-scale model is trained based on human preferences, and the preference scores it provides reflect human expectations for safe and compliant responses. Therefore, using the reward-based large-scale model for preference scoring can quickly and automatically filter out the relatively safest responses from multiple original responses. Compared to manual screening, this greatly improves screening efficiency and avoids the subjective bias and fatigue errors that may exist in manual screening, ensuring the accuracy and consistency of the screening results.
[0145] Based on the first embodiment of this application described above, a fourth embodiment of this application is proposed. Contents that are the same as or similar to the first embodiment can be referred to the above description, and will not be repeated hereafter. See also... Figure 5 In the fourth embodiment, step S30 includes steps S304 to S306.
[0146] Step S304: Using the large security assessment model, perform a security assessment on the original response corresponding to any risk issue.
[0147] The security assessment big data model is a large language model specifically designed for assessing the security of text content. It analyzes original responses, and based on pre-defined security rules, semantic understanding, and learning of various risk characteristics, determines whether the text poses a risk and provides corresponding assessment results.
[0148] Step S305: If the evaluation result indicates that the content of the original response is safe, the original response is determined to be a safe response.
[0149] Due to the nature of risk issues, the original responses generated by the security assessment model may contain various security vulnerabilities and need to be tested by the security assessment model to determine whether they can be used directly as security responses or whether a new response needs to be generated.
[0150] A secure response is one that meets security standards and does not contain risky content. In this embodiment, a secure response may be an original response that has been determined to be secure by a large-scale security assessment model.
[0151] Step S306: If the evaluation result indicates that the content of the original response is risky, the response corresponding to the risky issue is regenerated through the response big model, and the newly generated response is subjected to a security assessment until the response generated by the response big model is assessed as a safe response by the security assessment big model.
[0152] In this embodiment, the response big model receives risk questions, generates an original response, and when the original response is determined to be risky by the security assessment big model, a new response is generated until the generated response is deemed safe by the security assessment big model. Correspondingly, a safe response can also be a response that has been regenerated multiple times by the response big model and finally passes the security assessment.
[0153] Optionally, the training process of the security assessment model includes: generating multiple responses corresponding to sample risk questions sequentially through the response model; obtaining the risk assessment level and risk assessment reason for each response; and training the security assessment model based on the sample risk questions, the multiple responses corresponding to the sample risk questions, and the risk assessment level and risk assessment reason for each response.
[0154] The risk assessment level is a quantitative classification of the degree of risk contained in the responses generated by the response model. For example, risks can be divided into different levels, such as no risk, low risk, and high risk, based on factors such as the severity and scope of the risk. By clearly defining risk assessment levels, the safety assessment model can learn the characteristics and boundaries of different risk levels, thereby more accurately assessing the risks of new responses.
[0155] For example, different risk assessment levels are determined using a risk assessment score. A risk assessment score of 0 indicates that the response is risky; a risk assessment score of 1 indicates that the target model refuses to answer; and a risk assessment score of 2 indicates that the response is safe and reasonable, addresses the risk issue, and provides positive guidance. Accordingly, if the risk assessment score for any response in the overall safety assessment model is 2, the response can be determined as a safe response.
[0156] The risk assessment rationale explains the specific reasons why a response is classified as having a certain risk level. It details the risk factors present in the response, such as the inclusion of specific sensitive words, the reference to sensitive topics, or logical errors that could mislead users. The risk assessment rationale provides more detailed learning information for the overall security assessment model, enabling it to understand the root causes of risks and thus improving the accuracy and reliability of risk assessments.
[0157] In this embodiment, the security assessment model is trained by analyzing sample risk issues, multiple corresponding responses, and the risk assessment levels and justifications for each response. This allows the model to learn the characteristics of risk issues and the correlation between these characteristics and the risk assessment levels and justifications. This enables the security assessment model to accurately identify the risk level of new responses, providing accurate risk assessment levels and justifications, ensuring the accuracy of the assessment results, and reducing misjudgments and omissions.
[0158] It should be noted that, in addition to using the above methods to train a large security assessment model to perform security assessments on responses, a powerful large language model can also be invoked to perform security assessments on responses based on a carefully designed security assessment prompt.
[0159] In this embodiment, if the original response is determined to be safe by the security assessment model, it is confirmed as a safe response. For original responses deemed risky, the response model is required to regenerate and reassess until a safe response is obtained. This scheme performs a rigorous security assessment on the original responses, ensuring the reliability of the final determined safe response by only those responses that pass the assessment.
[0160] Based on the first embodiment of this application described above, a fifth embodiment of this application is proposed. Contents that are the same as or similar to the first embodiment can be referred to the above description, and will not be repeated hereafter. See also... Figure 6 In the fifth embodiment, step S40 includes steps S401 and S402.
[0161] Step S401: Use the security responses corresponding to each risk issue as training labels for each risk issue.
[0162] Training labels are target data used in supervised learning to guide the model's learning. Using safe responses as training labels provides the model with correct output examples, allowing it to know how to generate responses to specific risky issues and helping it learn safe and compliant response patterns.
[0163] Step S402: Construct supervised training corpus for the response model based on each risk question and the corresponding training label.
[0164] Supervised training corpora consist of a dataset of risky questions and their corresponding training labels, used to supervise the training of the response model. This corpus forms the basis for the model's learning; by learning from the supervised training corpus, the response model can adjust its parameters and improve its ability to generate safe responses when faced with risky questions.
[0165] In this embodiment, using risky questions and their corresponding safe responses to construct supervised training corpora allows the large response model to be exposed to a large number of risky question scenarios and their corresponding correct responses during training. After learning this content, when the model encounters similar risky questions in practical applications, it is more inclined to generate safe and compliant responses, effectively reducing the probability of generating risky content and improving the model's security.
[0166] Based on the first embodiment of this application described above, a sixth embodiment of this application is proposed. Contents that are the same as or similar to the first embodiment can be referred to the above description, and will not be repeated hereafter. See also... Figure 7 In the sixth embodiment, step S40 is refined into step S403.
[0167] Step S403: Construct a preference training corpus for the response big model based on each risk question, the original response to each risk question, and the safe response. The original response to each risk question and the safe response to each risk question are different.
[0168] The preference training corpus is a dataset used to train the response model, enabling it to learn human preferences. In this approach, the preference training corpus consists of risky questions, their corresponding original responses, and safe responses. Through training with this corpus, the model can learn what responses are safe and in line with human preferences, thereby adjusting its generation strategy to produce content that better meets human expectations.
[0169] In this embodiment, risky questions, original responses, and safe responses are collectively constructed into a preference training corpus, allowing the large response model to intuitively compare the differences between original and safe responses during training. The model can clearly learn which content is risky and what responses are safe and compliant, thereby strengthening its learning of safe response patterns and more accurately generating safe content when encountering similar issues in the future.
[0170] Based on the first embodiment of this application described above, a seventh embodiment of this application is proposed. Contents that are the same as or similar to the first embodiment can be referred to the above description, and will not be repeated hereafter. See also... Figure 8 In the seventh embodiment, after step S40, steps S501 to S507 are also included.
[0171] Step S501: In response to the dialogue instruction, extract the target question from the dialogue instruction.
[0172] A dialogue instruction is a user-initiated command containing a specific request, presented in the form of natural language text. It can be a simple question, a request, or a piece of dialogue content that needs to be processed. The target question is the core issue extracted from the dialogue instruction; it is the key content that needs to be analyzed, processed, and responded to.
[0173] Step S502: Search the reference whitelist for reference questions that match the target question.
[0174] The reference whitelist is a pre-defined set of questions that includes a series of risky questions and their corresponding security responses. When a target question is encountered, the server first searches the reference whitelist for a matching reference question to quickly obtain an appropriate response.
[0175] For example, the server determines the similarity between the target problem and each risk problem in the reference whitelist, and identifies the risk problems corresponding to similarity values greater than a similarity threshold as reference problems for the target problem.
[0176] Step S503: If a reference question matching the target question is found from the reference whitelist, the response corresponding to the target question is generated based on the security response corresponding to the reference question in the whitelist using the response big model.
[0177] If a reference question matching the target question is found in the reference whitelist, the server uses the security response corresponding to the whitelisted reference question as a reference response. Then, using a large response model, it generates a response corresponding to the target question based on this reference response. For example, the large response model might modify the description of the reference response based on the target question to obtain a response corresponding to the target question. Alternatively, the large response model might directly use the reference response as the response corresponding to the target question to improve response efficiency.
[0178] In this embodiment, a series of risky questions and their corresponding secure responses are pre-stored in a reference whitelist. Upon receiving a dialogue command and extracting the target question, a search and matching process is performed within the reference whitelist. Once a matching reference question is found, a response to the target question can be directly generated based on its corresponding secure response, eliminating the need for complex information retrieval, reasoning, and generation processes. This significantly shortens the response generation time. Furthermore, the risky questions and corresponding secure responses in the whitelist are carefully selected and verified, ensuring high accuracy and security. When the target question matches a reference question, generating a response based on an existing secure response further guarantees the security and accuracy of the generated response, improving the security and accuracy of the response model's output.
[0179] Step S504: If no reference question matching the target question can be found in the reference whitelist, search for a reference question matching the target question in the reference blacklist.
[0180] The reference blacklist is a pre-defined set of questions that contain risky questions deemed high-risk and unsuitable for answering. When a target question matches a risky question in the reference blacklist, the server will take appropriate measures to refuse to answer.
[0181] Step S505: If a reference question matching the target question is found from the reference blacklist, a default reply is generated, the content of which indicates that the target question has been rejected.
[0182] The default response is a uniform response given by the server when the target question matches a reference question in the reference blacklist. Its content indicates that the target question is rejected and is used to inform the user that the question does not meet the answering rules.
[0183] The blacklist records risky questions that pose security risks, violate laws and regulations, contravene ethical standards, or contain sensitive content. When a target question matches a reference question in the blacklist, a default response is generated indicating that the target question should not be answered, effectively preventing the output of harmful, illegal, or inappropriate information.
[0184] Step S506: If no reference problem matching the target problem can be found in the reference blacklist, risk detection is performed on the target problem using the risk detection big model to obtain the risk detection result.
[0185] The risk detection big model is a large language model specifically designed to detect whether text contains security risks. It can analyze a target question, identify whether it contains specific types of security risks, such as sensitive information, illegal or non-compliant content, or expressions that violate moral ethics, and provide corresponding risk detection results.
[0186] Step S507: If the risk detection result indicates that the target problem has a specified type of security risk, a response corresponding to the target problem is generated through the security response big model. The security response big model is obtained by training the response big model with security training corpus.
[0187] The specified type of security risk is a predefined series of security risk types, such as political sensitivity, pornography and vulgarity, violence and terrorism, and disinformation. The risk detection model will detect the target problem based on these types and determine whether the corresponding risk exists.
[0188] The secure response big model is a model trained on a secure training corpus. This corpus contains a large number of risky questions and their corresponding secure responses. Through training with this corpus, the secure response big model learns how to generate secure and compliant responses to various risky questions. Therefore, compared to a regular response big model, the responses generated by the secure response big model are more reliable and secure. Thus, target questions with specific types of security risks are handled by the secure response big model. Other target questions without security risks are handled by the regular response big model. The regular response big model can also be called the backbone big model or the business big model. Compared to the secure response big model, the regular response big model is significantly lighter. Its structure is more streamlined, requiring relatively fewer computational and storage resources, and it operates more efficiently. Therefore, it can more efficiently meet the response needs of general questions.
[0189] It's important to note that besides using large-scale risk detection models to assess the risk of a target issue, rule engines can also be employed. For example, a series of sensitive keywords can be defined, such as politically sensitive terms, pornographic or vulgar terms, and terms related to violence and terrorism. If the rule engine detects that the target issue contains these keywords, it determines that a corresponding type of risk exists.
[0190] This application embodiment considers that even if the target question is not in the reference whitelist or reference blacklist, it may still pose a potential security risk. Therefore, a risk detection model is used to detect risks in the target question. Leveraging the powerful language understanding and analysis capabilities of the risk detection model, various specified types of security risks are identified, and target questions with security risks are then processed by a security response model. Since the security response model is trained on a security training corpus, it can generate responses that meet security standards for target questions with risks. Therefore, this solution avoids the model outputting content with security risks, ensuring the security of the output content from the source.
[0191] Figure 9 This is a schematic diagram illustrating a model training process provided in some embodiments of this application. (Reference) Figure 9The training of the response model comprises three stages: Safety-SFT (Safety Supervised Fine-Tuning), Safety-DPO (Safety Direct Preference Optimization), and Safety-RFT (Safety Reward Fine-Tuning). After these three stages of training, the response model is upgraded to a secure response model. It's important to note that while the response model has been pre-trained and capable of handling normal business scenarios, it lacks security features. Therefore, through the three stages of safety-SFT, Safety-DPO, and Safety-RFT training, it can be upgraded to a secure response model.
[0192] The training corpus construction process for the Safety-SFT stage includes: inputting risk questions into the response model, which generates original responses. The safety assessment model then evaluates the original responses; if the content of the original response is safe, it is deemed a safe response; alternatively, if the content of the original response is risky, the response model regenerates the response and performs a safety assessment on the regenerated response until it is deemed safe. The risk questions and safe responses then constitute the supervised training corpus. This supervised training corpus is used to fine-tune the response model, completing the training for the Safety-SFT stage.
[0193] For the Safety-DPO phase, the training corpus construction process includes: inputting risk questions into the response model, which generates original responses. The rewriting model then rewrites the original responses based on the risk questions to obtain safe responses. Next, the risk questions, original responses, and safe responses form the preference training corpus. This preference training corpus is then used to optimize the response model, thus completing the training for the Safety-DPO phase.
[0194] For the Safety-RFT stage, the training corpus construction process includes: inputting the risk question into the response model, which then generates multiple raw responses for that risk question. Next, the reward model assigns a preference score to each raw response and selects the raw response with the highest preference score as the safe response. The risk question and the safe responses then constitute the supervised training corpus. This supervised training corpus is used to fine-tune the response model, thus completing the Safety-SFT stage training.
[0195] After three phases of security training to obtain the comprehensive security response model, this model works in conjunction with security safeguards to filter out risky target issues and process them. The security safeguards include a reference whitelist, a reference blacklist, a rule engine, and a comprehensive risk detection model. The method by which the comprehensive security response model and security safeguards work together has been specifically described in the above embodiments and will not be repeated here.
[0196] This application integrates safety alignment technologies such as safety-SFT, safety-DPO, and safety-RFT, and the training corpora at each stage can be automatically generated, reducing the cost of manual data annotation and enabling the entire process to operate automatically. Therefore, a large-scale safe response model can be quickly trained in various application scenarios, and with the support of safety barriers, risk issues can be transferred to the large-scale safe response model for processing, improving the security of business scenarios. Testing has shown that the security performance of the large-scale safe response model has been significantly improved in multiple business scenarios. Especially in translation scenarios, its security has been improved by 70%.
[0197] Another point to note is that the above examples are only for understanding this application and do not constitute a limitation on the corpus construction method of this application. Any simple transformations based on this technical concept are all within the protection scope of this application.
[0198] This application also provides a corpus construction apparatus, please refer to... Figure 10 The corpus construction device includes:
[0199] The question set acquisition module 10 is used to acquire a risk question set, which includes multiple risk questions. These risk questions are used to guide the response model to generate risk content.
[0200] The original response generation module 20 is used to generate the original responses for each risk issue based on the response model.
[0201] The security response acquisition module 30 is used to build a large model through the corpus and acquire the security response corresponding to each risk question based on each risk question and the original response corresponding to each risk question.
[0202] The training corpus construction module 40 is used to construct a security training corpus for a large response model based on each risk issue and its corresponding security response.
[0203] Optionally, constructing a large model from a corpus includes rewriting the large model;
[0204] The security response acquisition module 30 is used to rewrite the original responses corresponding to each risk issue by rewriting the large model, based on each risk issue, to obtain the security responses corresponding to each risk issue.
[0205] Optionally, the device further includes:
[0206] The first model training module is used to acquire training samples, which include sample risk issues and corresponding sample safety responses. It generates the original sample responses to the sample risk issues using the response model. It then rewrites the original sample responses based on the sample risk issues using the rewritten model, obtaining the rewritten sample responses. While keeping the parameters of the response model unchanged, it adjusts the parameters of the rewritten model based on the rewritten and safe sample responses to reduce the difference between them.
[0207] Optionally, the number of original responses corresponding to each risk question generated by the response model is multiple, and the corpus for constructing the large model includes a reward model;
[0208] The safe response acquisition module 30 is used to score the preferences of multiple original responses corresponding to each risk question through a reward model, and obtain the preference score corresponding to each original response; the original response with the highest preference score is selected as the safe response from the multiple original responses corresponding to each risk question.
[0209] Optionally, the device further includes:
[0210] The second model training module is used to generate multiple responses corresponding to sample risk questions sequentially through the response big model; obtain preference scores labeled for each of the multiple responses; and train the reward big model based on the risk questions, the multiple responses, and the preference scores labeled for each of the multiple responses.
[0211] Optionally, the corpus can be used to construct a large model, including a security assessment model.
[0212] The security response acquisition module 30 is used to perform a security assessment on the original response corresponding to any risk issue through the security assessment big model; if the assessment result indicates that the content of the original response is safe, the original response is determined to be a safe response; or, if the assessment result indicates that the content of the original response is risky, the response corresponding to the risk issue is regenerated through the response big model, and the newly generated response is subjected to a security assessment, until the response generated by the response big model is assessed as a safe response by the security assessment big model.
[0213] Optionally, the device further includes:
[0214] The third model training module is used to generate multiple responses corresponding to sample risk questions sequentially through the response model; obtain the risk assessment level and risk assessment reason for each response; and train the security assessment model based on the sample risk questions, the multiple responses corresponding to the sample risk questions, and the risk assessment level and risk assessment reason for each response.
[0215] Optionally, the training corpus construction module 40 is used to use the security responses corresponding to each risk issue as training labels for each risk issue; and to construct supervised training corpus for the response big model based on each risk issue and the training labels corresponding to each risk issue.
[0216] Optionally, the training corpus construction module 40 is used to construct a preference training corpus for the response big model based on each risk question, the original response corresponding to each risk question, and the safe response, wherein the content of the original response corresponding to each risk question is different from the content of the safe response corresponding to each risk question.
[0217] Optionally, the device further includes:
[0218] The instruction response module is used to respond to dialogue instructions and extract the target question from the dialogue instructions.
[0219] The issue matching module is used to search for reference issues that match the target issue from a reference whitelist;
[0220] The response generation module is used to generate a response to the target question based on the security responses corresponding to the reference questions in the whitelist when a reference question matching the target question is found from the whitelist.
[0221] Optionally, the device further includes:
[0222] The issue matching module is also used to search for reference issues that match the target issue from the reference blacklist when no reference issue matching the target issue can be found from the reference whitelist.
[0223] The response generation module is also used to generate a default response when a reference question matching the target question is found from the reference blacklist. The default response indicates that the target question has been rejected.
[0224] Optionally, the device further includes:
[0225] The response generation module is also used to perform risk detection on the target question using a risk detection big data model when no reference question matching the target question can be found from the reference blacklist, and to obtain the risk detection result; when the risk detection result indicates that the target question has a specified type of security risk, the module generates a response corresponding to the target question using a security response big data model, which is trained by the response big data model on a security training corpus.
[0226] The corpus construction apparatus provided in this application, employing the corpus construction method in the above embodiments, can solve the technical problem in related technologies where the efficiency of manually annotated security corpora is extremely low, resulting in a slow growth of security corpora available for training. Compared with the prior art, the beneficial effects of the corpus construction apparatus provided in this application are the same as those of the corpus construction method provided in the above embodiments, and other technical features in the corpus construction apparatus are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0227] This application provides a corpus construction device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the corpus construction method in the above embodiment 1.
[0228] The following is for reference. Figure 11 The diagram illustrates a structural schematic suitable for implementing the corpus construction device of the present application embodiments. The corpus construction device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 11 The corpus construction device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0229] like Figure 11As shown, the corpus construction device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.) that can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the corpus construction device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. The communication device 1009 allows the corpus building device to communicate wirelessly or wiredly with other devices to exchange data. Although the figure shows corpus building devices with various systems, it should be understood that it is not required to implement or possess all of the systems shown. More or fewer systems may be implemented alternatively.
[0230] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0231] The corpus construction device provided in this application, employing the corpus construction method described in the above embodiments, can solve the technical problem in related technologies where the efficiency of manually annotated security corpora is extremely low, resulting in a slow growth of security corpora available for training. Compared with the prior art, the beneficial effects of the corpus construction device provided in this application are the same as those of the corpus construction method provided in the above embodiments, and other technical features in this corpus construction device are the same as those disclosed in the previous embodiment method, and will not be repeated here.
[0232] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0233] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0234] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the corpus construction method in the above embodiments.
[0235] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0236] The aforementioned computer-readable storage medium may be included in the corpus building device; or it may exist independently and not be assembled into the corpus building device.
[0237] The aforementioned computer-readable storage medium carries one or more programs that, when executed by the corpus construction device, cause the corpus construction device to: acquire a set of risk questions, the set of risk questions including multiple risk questions used to guide the response model to generate risky content; generate original responses corresponding to each risk question through the response model; acquire safe responses corresponding to each risk question based on each risk question and its corresponding original responses through the corpus construction model; and construct a safe training corpus for the response model based on each risk question and its corresponding safe responses.
[0238] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0239] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0240] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0241] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., computer programs) for executing the above-described corpus construction method. This solves the technical problem in related technologies where the efficiency of manually annotated security corpora is extremely low, resulting in a slow growth of the security corpus available for training. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the corpus construction method provided in the above embodiments, and will not be repeated here.
[0242] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the corpus construction method described above.
[0243] The computer program product provided in this application can solve the technical problem in related technologies where the efficiency of manually annotated security corpora is extremely low, resulting in a slow growth of security corpora available for training. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the corpus construction method provided in the above embodiments, and will not be repeated here.
[0244] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.
Claims
1. A corpus construction method, characterized in that, The method includes: Obtain a set of risk questions, which includes multiple risk questions, and these risk questions are used to guide the response model to generate risk content; The original responses for each risk question are generated using the aforementioned response model. A large model is built using corpus data, and based on each risk question and its corresponding original response, a safe response is obtained for each risk question. Based on each risk issue and its corresponding security response, a security training corpus for the large response model is constructed. The corpus-based large-scale model includes a reward large-scale model and a security assessment large-scale model; The process involves constructing a large model from a corpus, and based on each risk question and its corresponding original response, obtaining the corresponding safe response for each risk question, including: Using the aforementioned reward model, preference scores are assigned to multiple original responses corresponding to each risk question, resulting in a preference score for each original response. The original response with the highest preference score is selected from these original responses as the safe response, where a higher preference score indicates a safer original response; or... The security assessment model is used to perform a security assessment on the original response to any risk issue. If the assessment result indicates that the content of the original response is safe, the original response is determined to be a safe response. Alternatively, if the assessment result indicates that the content of the original response is risky, the response to the risk issue is regenerated using the response model, and the newly generated response is subjected to a security assessment until the response generated by the response model is assessed as a safe response by the security assessment model.
2. The method as described in claim 1, characterized in that, The corpus is used to construct a large model, which includes rewriting the large model. The process involves constructing a large model from a corpus, and based on each risk question and its corresponding original response, obtaining the corresponding safe response for each risk question, including: By rewriting the large model, the original responses corresponding to each risk issue are rewritten based on each risk issue to obtain the security responses corresponding to each risk issue.
3. The method as described in claim 2, characterized in that, The training process for rewriting the large model includes: Obtain training samples, which include sample risk issues and corresponding sample safety responses to the sample risk issues; The original responses to the sample risk issues are generated using the aforementioned response model. By rewriting the large model and based on the sample risk issue, the original response of the sample is rewritten to obtain the rewritten response of the sample; While keeping the parameters of the large response model unchanged, the parameters of the large rewrite model are adjusted based on the sample rewritten response and the sample safe response to reduce the difference between the sample rewritten response and the sample safe response.
4. The method as described in claim 1, characterized in that, The training process of the large reward model includes: The large response model is used to sequentially generate multiple responses corresponding to sample risk questions. Obtain the preference scores labeled for each of the multiple responses; The reward model is trained based on the risk question, the multiple responses, and the preference scores labeled for each of the multiple responses.
5. A corpus construction device, characterized in that, The device includes: The question set acquisition module is used to acquire a risk question set, which includes multiple risk questions. These risk questions are used to guide the response model to generate risk content. The original response generation module is used to generate original responses for each risk issue based on the large response model. The security response acquisition module is used to build a large model from the corpus and obtain the security response corresponding to each risk question based on each risk question and the original response corresponding to each risk question. The training corpus construction module is used to construct the security training corpus of the large response model based on each risk issue and the corresponding security response. The corpus is used to construct a large-scale model, which includes a reward model and a security assessment model. The safety response acquisition module is used to assign preference scores to multiple original responses corresponding to each risk question using the reward model, obtaining a preference score for each original response; and selecting the original response with the highest preference score from the multiple original responses corresponding to each risk question as the safety response; or... The security response acquisition module is used to perform a security assessment on the original response corresponding to any risk issue using the security assessment model; if the assessment result indicates that the content of the original response is safe, the original response is determined as the safe response; or, if the assessment result indicates that the content of the original response is risky, the module regenerates the response corresponding to the risk issue using the response model and performs a security assessment on the newly generated response until the response generated by the response model is assessed as a safe response by the security assessment model.
6. A corpus construction device, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the corpus construction method as described in any one of claims 1 to 4.
7. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the corpus construction method as described in any one of claims 1 to 4.
8. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the corpus construction method as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Method and device for large model security defense and electronic equipment
CN118940276A
Model fine tuning training method and device, answer output method and device and electronic equipment
CN119150013A
Intelligent questioning and answering method for knowledge in education field integrated with Scoglai teaching concept
CN119443282A