Reply method, system and equipment and storage medium

By obtaining and splicing the reference information of the language model for non-toxic and toxic instructions, the problem of inaccurate toxicity recognition of language models is solved, the balance of generation rate and leak rate is achieved, and the accuracy of the reply content is improved.

CN120216622APending Publication Date: 2025-06-27BEIJING ZITIAO NETWORK TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311774928.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-21
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the prior art, the language model is not accurate enough in terms of toxicity recognition, resulting in a low generation rate or a high leakage rate, and it is impossible to effectively balance the two.

Method used

By receiving the target instructions to be replied to by the language model, the reference information used to reply to non-toxic instructions and toxic instructions is obtained, and spliced ​​into the third reference information, and input the language model to generate the reply content.

Benefits of technology

By taking into account reference information of non-toxic and toxic instructions in a comprehensive way, the language model can balance the generation rate and leak rate, improving the accuracy and effectiveness of the reply content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216622A_ABST
    Figure CN120216622A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computers, and discloses a reply method, system and device and a storage medium, and the reply method comprises the steps: receiving a to-be-replied target instruction of a language model; obtaining first reference information used when the language model replies the non-toxic instruction, and obtaining second reference information used when the language model replies the toxic instruction; splicing the first reference information and the second reference information to obtain third reference information required for replying the target instruction; and inputting the target instruction and the third reference information into the language model, so that the language model generates reply content for the target instruction based on the third reference information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to a reply method, system, device, and storage medium. Background Art

[0002] A language model refers to a model that, by understanding the prompt information input by a user, uses the response information corresponding to the prompt information as the output content. Here, the output content of the language model is the answer content of the language model to the prompt information. For example, assuming that the prompt information input by the user is "Why is there climate change", the language model can use the relevant factors causing climate change as the answer.

[0003] Currently, in some scenarios, the language model needs to perform toxicity recognition on the prompt information input by the user. Specifically, toxic prompt information refers to prompt information that is not positive or does not allow a normal answer, such as prompt information related to illegal or unethical values. Non-toxic prompt information refers to prompt information that can be normally answered. For toxic prompt information, the language model can refuse to answer or guide the user in a positive direction. For example, assuming that the prompt information input by the user is "The benefits of XX", the language model can refuse to answer or use the disadvantages of XX as the answer. For non-toxic prompt information, the language model can answer normally.

[0004] However, in some technologies, the toxicity recognition of the language model for the prompt information is inaccurate, specifically manifested as too strong toxicity recognition ability, resulting in a low normal answer ability of the language model, that is, a low generation rate; or too weak toxicity recognition ability, resulting in too high a proportion of normal answers and toxic answers generated for toxic prompt information, that is, too high a leakage rate.

[0005] In view of this, there is an urgent need for a method that can balance the generation rate and the leakage rate. Summary of the Invention

[0006] In view of this, embodiments of the present disclosure provide a reply method, a reply system, an electronic device, and a computer-readable storage medium, which can balance the generation rate and the leakage rate of the language model.

[0007] On the one hand, the present disclosure provides a reply method, and the method includes:

[0008] Receiving a target instruction to be replied by the language model;

[0009] Obtaining first reference information used by the language model when replying to a non-toxic instruction, and obtaining second reference information used by the language model when replying to a toxic instruction;

[0010] Concatenating the first reference information and the second reference information to obtain third reference information required for replying to the target instruction;

[0011] Input the target instruction and the third reference information into the language model, so that the language model generates a response content for the target instruction based on the third reference information.

[0012] On the other hand, the present disclosure also provides a response system, which includes:

[0013] An instruction receiving module, configured to receive a target instruction to be replied by a language model;

[0014] A reference information obtaining module, configured to obtain first reference information used by the language model when replying to a non-toxic instruction, and obtain second reference information used by the language model when replying to a toxic instruction;

[0015] A reference information splicing module, configured to splice the first reference information and the second reference information to obtain third reference information required for replying to the target instruction;

[0016] An instruction replying module, configured to input the target instruction and the third reference information into the language model, so that the language model generates a response content for the target instruction based on the third reference information.

[0017] On the other hand, the present disclosure also provides a computer-readable storage medium, which is used to store a computer program. When the computer program is executed by a processor, the method described above is implemented.

[0018] On the other hand, the present disclosure also provides an electronic device, which includes a processor and a memory. The memory is used to store a computer program. When the computer program is executed by the processor, the method described above is implemented.

[0019] In the technical solutions of some embodiments of the present application, after receiving a target instruction to be replied by a language model, the first reference information used by the language model when replying to a non-toxic instruction is spliced with the second reference information used by the language model when replying to a toxic instruction to obtain third reference information. Since the third reference information includes both the first reference information and the second reference information, after inputting the third reference information and the target instruction into the language model, the language model can comprehensively consider from two aspects of non-toxic instructions and toxic instructions, so that the generated response content for the target instruction can achieve a balance between the generation rate and the omission rate. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The features and advantages of the present disclosure will be more clearly understood by referring to the accompanying drawings. The drawings are schematic and should not be construed as imposing any limitation on the present disclosure. In the drawings:

[0021] Figure 1 shows a schematic flowchart of a reply method provided by an embodiment of the present application;

[0022] Figure 2 shows a schematic diagram of an instruction receiving interface provided by an embodiment of the present application;

[0023] Figure 3 shows a schematic flowchart of a training data generation method provided by an embodiment of the present application;

[0024] Figure 4 shows a schematic structural diagram of an instruction category provided by an embodiment of the present application;

[0025] Figure 5 shows a schematic module diagram of a reply system provided by an embodiment of the present application;

[0026] Figure 6 shows a schematic diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0027] To make the purposes, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present disclosure without creative efforts shall fall within the protection scope of the present disclosure.

[0028] The present application provides a reply method, which can balance the generation rate and omission rate of a language model. The reply method can be applied to an electronic device. The electronic device includes, but is not limited to, a tablet computer, a desktop computer, a laptop computer, a server, etc. Please refer to Figure 1 , which is a schematic flowchart of a reply method provided by an embodiment of the present application. Figure 1 In

[0029] Step S11, receive a target instruction to be replied by the language model.

[0030] Specifically, instructions can refer to questions, requests or task statements expressed in natural language provided by the user. The language model understands the instructions and can output responses that match the instructions. For example, "Please write a science fiction novel" can be used as an instruction. Based on this instruction, the language model can output a science fiction novel as a response. For another example, "What is the result of 5 times 3" can be used as an instruction. Based on this instruction, the language model can output 15 as a response. For another example, "What are the effects of rising global temperatures" can be used as an instruction. Based on this instruction, the language model can output climate change, sea level rise, ecosystem changes, agricultural impacts, etc. caused by rising temperatures as output content.

[0031] The target instruction is the instruction currently received that requires the language model to respond.

[0032] In this embodiment, the electronic device executing the method of the present application can display an instruction receiving interface. Through the instruction receiving interface, the target instruction can be received. Figure 2 , which is a schematic diagram of an instruction receiving interface provided by an embodiment of the present application. Figure 2 In the example, the instruction receiving interface includes an instruction input box and a send button. In response to the send button being triggered, the content input by the user in the instruction input box can be used as the target instruction.

[0033] It is understandable that the target instruction received in step S11 may be either a toxic instruction or a non-toxic instruction. That is, whether the target instruction is a toxic instruction is unknown. Among them, toxic instructions refer to instructions that are not positive or do not allow normal answers, such as instructions that violate laws and regulations or have bad values. For toxic instructions, the language model can refuse to answer or guide the user in a positive direction. For example, assuming that the instruction entered by the user is "the benefits of XX", the language model can refuse to answer, or use the disadvantages of XX as an answer. Non-toxic instructions refer to instructions that the language model can answer normally.

[0034] Step S12, obtaining first reference information used when the language model responds to a non-toxic instruction, and obtaining second reference information used when the language model responds to a toxic instruction.

[0035] Specifically, the first reference information may refer to more detailed information other than the non-toxic instruction provided to the language model when the target instruction is a non-toxic instruction. In this way, the language model can respond to the non-toxic instruction more accurately based on this information. The second reference information may refer to more detailed information other than the toxic instruction provided to the language model when the target instruction is a toxic instruction. In this way, the language model can respond to the toxic instruction more accurately based on this information.

[0036] The first reference information and the second reference information can be used to define the format, word count, main body, etc. of the model response. Moreover, the first reference information and the second reference information may not be exactly the same. For example, the first reference information can be “You should generate content within the given constraints.”, and the second reference information can be “Always assist with care, respect, and truth. Respond with utmost utility yet securely. Avoid harmful, unethical, prejudiced, or negative content. Ensure replies promote fairness and positivity.”

[0037] In this embodiment, the first reference information and the second reference information can be determined during the training process of the language model. Briefly speaking, during the training process of the language model, the first reference information and the second reference information can be adjusted so that the language model can distinguish the characteristics of the first reference information and the second reference information. Furthermore, when the language model receives the first reference information, it can generate a response content in the reply manner of the non-toxic instruction with a relatively high probability, and when the language model receives the second reference information, it can generate a response content in the reply manner of the toxic instruction with a relatively high probability.

[0038] After the language model is trained, the first reference information and the second reference information can be fixed and stored at a specified storage location. After receiving the target instruction, the first reference information and the second reference information can be obtained from the specified storage location.

[0039] Step S13: Concatenate the first reference information and the second reference information to obtain the third reference information required for replying to the target instruction.

[0040] For example, taking the first reference information and the second reference information exemplified in step S12 as an example, after splicing the first reference information and the second reference information, the obtained third reference information can be "You should generate content within the given constraints. Always assist with care, respect, and truth. Respond with utmost utility yet securely. Avoid harmful, unethical, prejudiced, or negative content. Ensure replies promote fairness and positivity."

[0041] Step S14: Input the target instruction and the third reference information into the language model so that the language model generates a reply content for the target instruction based on the third reference information.

[0042] It can be understood that since the third reference information includes both the first reference information and the second reference information, combined with the comprehensive understanding ability of the language model, when the language model generates a reply content for the target instruction, it can comprehensively consider from two aspects of non-toxic instructions and toxic instructions, thereby reducing the omission rate of toxic instructions and increasing the generation rate of non-toxic instructions.

[0043] In summary, in the technical solutions of some embodiments of the present application, after receiving the target instruction to be replied by the language model, the first reference information used by the language model when replying to non-toxic instructions is spliced with the second reference information used by the language model when replying to toxic instructions to obtain the third reference information. Since the third reference information includes both the first reference information and the second reference information, after inputting the third reference information and the target instruction into the language model, the language model can comprehensively consider from two aspects of non-toxic instructions and toxic instructions, so that the generated reply content for the target instruction can achieve a balance between the generation rate and the omission rate.

[0044] The following further describes the solution of the present application.

[0045] In some embodiments, the instructions replied by the language model are allowed to be divided into multiple instruction categories. For example, the instructions can be divided into creation category, transformation category, knowledge reasoning category, summary extraction category, and other categories. Each instruction can belong to one of the instruction categories. The first reference information used by the language model when replying to non-toxic instructions under different instruction categories is different. For example, the first reference information used for non-toxic instructions under the creation category can be "You should generate imaginative and original content within the given constraints.", and the first reference information used for non-toxic instructions under the transformation category can be "You should convert content within the given constraints.". In view of this, obtaining the first reference information used by the language model when replying to non-toxic instructions described in step S12 above can include:

[0046] Determine the target instruction category to which the target instruction belongs, and obtain the first reference information used by the language model when replying to non-toxic instructions under the target instruction category.

[0047] In this embodiment, by classifying the instructions and setting different first reference information for instructions under different instruction categories, the language model can better distinguish non-toxic instructions of different categories, and then generate reply content that is more consistent with the non-toxic instructions, which can improve the accuracy of the reply content.

[0048] Furthermore, considering that for toxic instructions under different instruction categories, the language model can use the same or similar reply content for reply, so there is no need to set different second reference information for toxic instructions under different instruction categories. That is, toxic instructions under all instruction categories can have the same second reference information. In this way, the amount of data processing can be reduced.

[0049] The training process of the language model of the present application will be described below.

[0050] Those skilled in the art can understand that before model training, it is first necessary to collect training data and annotate the training data. In the present application, the training data can include toxic instruction-reply content pairs composed of toxic instructions and reply content for toxic instructions, and non-toxic instruction-reply content pairs composed of non-toxic instructions and reply content for non-toxic instructions. The annotation work of this kind of training data is relatively complex and has low efficiency. In view of this, the present application provides a training data generation method with high efficiency. Referring to Figure 3 , which is a schematic flowchart of the training data generation method provided by an embodiment of the present application.Figure 3 Among them, the training data generation method includes the following steps:

[0051] Step S31, obtain toxic seed instructions.

[0052] Specifically, the so-called toxic seed instructions refer to toxic instructions that can be used as seeds to generate more other instructions. In this embodiment, the toxic seed instructions can be used as seeds to generate more other sample toxic instructions and sample non-toxic instructions.

[0053] Furthermore, the toxic seed instructions can cover different instruction categories and toxicity categories. For example, assume that the instruction categories include creation, transformation, knowledge reasoning, summary extraction, and others, and the toxicity categories include illegal and unethical values. The toxic seed instructions can include:

[0054] Illegal and unethical value instructions under creation instructions;

[0055] Illegal and unethical value instructions under transformation instructions;

[0056] Illegal and unethical value instructions under knowledge reasoning instructions;

[0057] Illegal and unethical value instructions under summary extraction instructions;

[0058] Illegal and unethical value instructions under other instructions.

[0059] Step S32, input the toxic seed instructions into the trained first generation model, and the first generation model generates sample toxic instructions based on the toxic seed instructions.

[0060] Specifically, the number of sample toxic instructions under different toxicity categories can be determined according to the number of sample toxic instructions and sample non-toxic instructions required for each instruction category. Furthermore, the number of sample toxic instructions, the toxic seed instructions, and the toxicity category can be input into the trained first generation model, and the first generation model uses the toxic seed instructions as seeds to generate the required number of sample toxic instructions.

[0061] In this embodiment, when the first generation model generates sample toxic instructions, it can include the following steps:

[0062] 321) Generate candidate sample toxicity instructions, score the toxicity intensity of each candidate sample toxicity instruction, and label the toxicity category of each candidate sample toxicity instruction. Simply put, the first generation model can have a scoring function and a labeling function. For each generated candidate sample toxicity instruction, the first generation model can score each candidate sample toxicity instruction and label the toxicity category. Among them, the higher the score of the candidate sample toxicity instruction, the stronger the toxicity of the candidate sample toxicity instruction.

[0063] 322) After removing the candidate sample toxicity instructions with scores lower than the score threshold (i.e., not toxic enough) or incorrect toxicity categories, output the remaining candidate sample toxicity instructions as sample toxicity instructions.

[0064] In this way, through the first generation model, sample toxicity instructions can be generated.

[0065] In some other embodiments, the first generation model may not need to score or label the toxicity category of the candidate sample toxicity instructions, and directly output all the generated candidate sample toxicity instructions. Furthermore, through manual screening, the candidate sample toxicity instructions with insufficient toxicity or incorrect toxicity categories can be removed.

[0066] Of course, in some embodiments, a combination of the first generation model screening and manual screening can also be used to screen out sample toxicity instructions from the generated candidate sample toxicity instructions. Specifically, in these embodiments, the first generation model can have a scoring function and a toxicity category labeling function. Among the generated candidate sample toxicity instructions, if the score of a candidate sample toxicity instruction is lower than the first score threshold or the toxicity category is incorrect, the first generation model can directly remove the candidate sample toxicity instruction; if the score of a candidate sample toxicity instruction is higher than the second score threshold and the toxicity category is correct, the first generation model can output the candidate sample toxicity instruction as a sample toxicity instruction; if the score of a candidate sample toxicity instruction is between the first score threshold and the second score threshold, then the candidate sample toxicity instruction can be output as an instruction to be manually quality inspected to determine whether the candidate sample toxicity instruction can be used as a sample toxicity instruction through manual means. Among them, the above first score threshold is less than the second score threshold.

[0067] Step S33, based on the sample toxicity instructions, generate sample non-toxic instructions associated with the sample toxicity instructions.

[0068] Specifically, the sample non-toxic instructions can be obtained by expanding, modifying the sample toxicity instructions, etc.

[0069] Step S34, input the sample toxicity instructions and the sample non-toxic instructions into the trained second generation model, and the second generation model generates responses to the sample toxicity instructions and the sample non-toxic instructions.

[0070] In this embodiment, similar to the first generation model, when generating responses to sample non-toxic instructions and sample toxic instructions, the second generation model can also score the generated responses, and eliminate the responses with lower scores and the instructions corresponding to the responses. The relevant principle will not be elaborated here.

[0071] Thus, the training data of the language model can be obtained.

[0072] In some embodiments, the language model can be trained based on the following method:

[0073] Input the sample non-toxic instructions and the first sample reference information set for the sample non-toxic instructions into the language model, and adjust the parameters of the language model and the first sample reference information based on the first response content for the sample non-toxic instructions;

[0074] Input the sample toxic instructions and the second sample reference information set for the sample toxic instructions into the language model, and adjust the parameters of the language model and the second sample reference information based on the second response content for the sample toxic instructions.

[0075] Simply put, in the training stage of the language model, the language model can be trained based on the sample non-toxic instructions and toxic instructions respectively. In this way, the trained language model can better distinguish the instruction features of the sample non-toxic instructions and toxic instructions. In addition, during the training process of the language model, by adjusting the first sample reference information and the second sample reference information, the feature differences between these two sample reference information can be increased, which is beneficial for the language model to better learn the features of these two sample reference information.

[0076] Further, in some embodiments, the sample non-toxic instructions are allowed to be divided into multiple instruction categories, and different first sample reference information is allowed to be set for the sample non-toxic instructions under different instruction categories. Regarding the instruction categories, reference can be made to the relevant description in step S12, which will not be elaborated here.

[0077] The above training of the language model based on the sample non-toxic instructions can further include:

[0078] For any one of the above multiple instruction categories, input the sample non-toxic instructions under this instruction category and the first sample reference information set for the sample non-toxic instructions under this instruction category into the language model, and adjust the parameters of the language model and the first sample reference information set for the sample non-toxic instructions under this instruction category based on the first response content for the sample non-toxic instructions.

[0079] In this way, the trained language model can identify non-toxic instructions under different instruction categories and generate different response contents for non-toxic instructions under different instruction categories.

[0080] Furthermore, each instruction category can respectively include one or more sub-instruction categories, and different first sample reference information is allowed to be set for sample non-toxic instructions under different sub-instruction categories. For ease of understanding, refer to Figure 4 , which is a schematic structural diagram of an instruction category provided by an embodiment of the present application. Figure 4 In [the figure], the creation category is an instruction category. The creation category includes two sub-instruction categories: oral broadcast copywriting and novels. Different first sample reference information is allowed to be set for sample non-toxic instructions under the two sub-instruction categories of main body creation and novels.

[0081] The above training of the language model based on sample non-toxic instructions may include:

[0082] For any sub-instruction category under an instruction category, input the sample non-toxic instructions under the sub-instruction category and the first sample reference information set for the sample non-toxic instructions under the sub-instruction category into the language model, and adjust the parameters of the language model and the first sample reference information set for the sample non-toxic instructions under the sub-instruction category based on the first response content for the sample non-toxic instructions.

[0083] In this way, the trained language model can distinguish non-toxic instructions of different sub-categories under the same instruction category and generate different response contents.

[0084] Furthermore, in some embodiments, each instruction category and each sub-instruction category have their respective corresponding first sample reference information;

[0085] For any sub-instruction category under an instruction category, the first sample reference information for the sample non-toxic instructions under the sub-instruction category can be set based on the following method:

[0086] Concatenate the first sample reference information corresponding to the sub-instruction category and the first sample reference information of the instruction category to which the sub-instruction category belongs to obtain the first sample reference information set for the sample non-toxic instructions under the sub-instruction category.

[0087] For example, Figure 4Among them, assume that the first sample reference information for the creative category is "You should generate imaginative and original content within the given constraints.", and the first sample reference information for the voiceover script category is "Conforming to the genre and format of a voiceover script, the description is in first-person perspective and does not involve dialogues.". By concatenating the above two pieces of first sample reference information, we get "You should generate imaginative and original content within the given constraints. Conforming to the genre and format of a voiceover script, the description is in first-person perspective and does not involve dialogues.", which can be used as the first sample reference information for the sample non-toxic instructions under the voiceover script category.

[0088] The first sample reference information for the sample non-toxic instructions under the sub-instruction category obtained through this concatenation method can make the first sample reference information of different sub-categories under the same instruction category have both similarity and distinctiveness, facilitating the exploration of the language model's ability to integrate knowledge.

[0089] Furthermore, based on the first sample reference information corresponding to the instruction category, the language model trained according to the above method can already distinguish well between the non-toxic instructions of different sub-instruction categories under the same instruction category. In view of this, the first reference information used when the above-mentioned language model replies to the non-toxic instructions under the target instruction category may include:

[0090] For the non-toxic instructions under any sub-instruction category of the target instruction category, use the first reference information corresponding to the target instruction category as the first reference information used by the language model to reply to the non-toxic instructions under this sub-instruction category.

[0091] In this way, the amount of data processing can be reduced.

[0092] So far, the description of the voice enhancement method of this application is completed.

[0093] Corresponding to the voice enhancement method, the present application also provides a voice enhancement system. Please refer to Figure 5 , which is a schematic diagram of the module of the response system provided by an embodiment of the present application. Figure 5 In

[0094] The instruction receiving module is used to receive the target instruction to be replied by the language model;

[0095] The reference information acquisition module is used to acquire the first reference information used by the language model when replying to non-toxic instructions, and acquire the second reference information used by the language model when replying to toxic instructions;

[0096] The reference information splicing module is used to splice the first reference information and the second reference information to obtain the third reference information required for replying to the target instruction;

[0097] The instruction reply module is used to input the target instruction and the third reference information into the language model, so that the language model generates a reply content for the target instruction based on the third reference information.

[0098] Please refer to Figure 6 , which is a schematic diagram of the electronic device provided by an embodiment of the present application. The electronic device includes a processor and a memory. The memory is used to store a computer program. When the computer program is executed by the processor, the above method is implemented.

[0099] Among them, the processor can be a central processing unit (CPU). The processor can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. chips, or a combination of the above types of chips.

[0100] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the method in the embodiment of the present invention. The processor executes various functional applications and data processing of the processor by running the non-transitory software programs, instructions, and modules stored in the memory, that is, the method in the above method embodiment is implemented.

[0101] The memory may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created by the processor and the like. In addition, the memory may include a high-speed random access memory, and may also include a non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0102] An embodiment of the present application also provides a computer-readable storage medium for storing a computer program, which when executed by a processor, implements the above method.

[0103] Although the embodiments of the present disclosure have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A reply method, characterized in that, The method includes: Receiving a target instruction to be replied by a language model; Obtaining first reference information used by the language model when replying to a non-toxic instruction, and obtaining second reference information used by the language model when replying to a toxic instruction; Concatenating the first reference information and the second reference information to obtain third reference information required for replying to the target instruction; Inputting the target instruction and the third reference information into the language model, so that the language model generates a reply content for the target instruction based on the third reference information.

2. The method according to claim 1, wherein Instructions replied by the language model can be divided into multiple instruction categories, and the first reference information used by the language model when replying to non-toxic instructions under different instruction categories is different; The obtaining of the first reference information used by the language model when replying to a non-toxic instruction includes: Determining the target instruction category to which the target instruction belongs, and obtaining the first reference information used by the language model when replying to non-toxic instructions under the target instruction category.

3. The method according to claim 2, wherein Each of the instruction categories includes one or more sub-instruction categories, and there is respective corresponding first reference information for the instruction category and each sub-instruction category; The obtaining of the first reference information used by the language model when replying to non-toxic instructions under the target instruction category includes: For a non-toxic instruction under any sub-instruction category of the target instruction category, using the first reference information corresponding to the target instruction category as the first reference information used by the language model when replying to the non-toxic instruction under the sub-instruction category.

4. The method according to any one of claims 1, characterized in that, Before receiving the target instruction, the language model is trained based on the following method: Inputting a sample non-toxic instruction and first sample reference information set for the sample non-toxic instruction into the language model, and adjusting the parameters of the language model and the first sample reference information based on the first reply content for the sample non-toxic instruction; Inputting a sample toxic instruction and second sample reference information set for the sample toxic instruction into the language model, and adjusting the parameters of the language model and the second sample reference information based on the second reply content for the sample toxic instruction.

5. The method according to claim 4, wherein The sample non-toxic instructions can be divided into multiple instruction categories, and different first sample reference information can be set for sample non-toxic instructions under different instruction categories; Training the language model based on the sample non-toxic instructions includes: For any one of the multiple instruction categories, inputting the sample non-toxic instructions under the instruction category and the first sample reference information set for the sample non-toxic instructions under the instruction category into the language model, and adjusting the parameters of the language model and the first sample reference information set for the sample non-toxic instructions under the instruction category based on the first reply content for the sample non-toxic instructions.

6. The method according to claim 5, wherein Each of the instruction categories includes one or more sub-instruction categories, and different first sample reference information is allowed to be set for the sample non-toxic instructions under different sub-instruction categories; Training the language model based on the sample non-toxic instructions includes: For any sub-instruction category under an instruction category, input the sample non-toxic instructions under the sub-instruction category and the first sample reference information set for the sample non-toxic instructions under the sub-instruction category into the language model, and based on the first reply content for the sample non-toxic instructions, adjust the parameters of the language model and the first sample reference information set for the sample non-toxic instructions under the sub-instruction category.

7. The method according to claim 6, wherein Each instruction category and each sub-instruction category have their respective corresponding first sample reference information; For any sub-instruction category under an instruction category, the first sample reference information for the sample non-toxic instructions under the sub-instruction category is set based on the following method: Concatenate the first sample reference information corresponding to the sub-instruction category and the first sample reference information of the instruction category to which the sub-instruction category belongs to obtain the first sample reference information set for the sample non-toxic instructions under the sub-instruction category.

8. The method according to claim 6, wherein Before training the language model, the training data of the language model is constructed based on the following method: Obtain toxic seed instructions; Input the toxic seed instructions into a trained first generation model, and the first generation model generates sample toxic instructions based on the toxic seed instructions; Based on the sample toxic instructions, generate sample non-toxic instructions associated with the sample toxic instructions; Input the sample toxic instructions and the sample non-toxic instructions into a trained second generation model, and the second generation model generates replies for the sample toxic instructions and the sample non-toxic instructions.

9. A reply system, characterized in that, The system includes: An instruction receiving module for receiving a target instruction to be replied by a language model; A reference information obtaining module for obtaining the first reference information used by the language model when replying to non-toxic instructions, and obtaining the second reference information used by the language model when replying to toxic instructions; A reference information concatenating module for concatenating the first reference information and the second reference information to obtain the third reference information required for replying to the target instruction; An instruction replying module for inputting the target instruction and the third reference information into the language model, so that the language model generates a reply content for the target instruction based on the third reference information.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program, and when the computer program is executed by a processor, the method described in any one of claims 1 to 8 is implemented.

11. An electronic device, characterized in that, The electronic device includes a processor and a memory, the memory is used to store a computer program, and when the computer program is executed by the processor, the method described in any one of claims 1 to 8 is implemented.