Large language model man-machine conversation method and system based on association risk perception
By constructing a preference-balanced training set and an adaptive multi-objective regulation generation model, the conflict between safety and utility in role-based expression of large language models is resolved, achieving a dynamic balance between safety and role expressiveness, and improving the interactive capabilities in complex scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-07
- Publication Date
- 2026-04-07
AI Technical Summary
Large language models may generate inappropriate remarks, biases, or violations in role-playing expressions, leading to a conflict between security and role-playing utility.
By constructing a preference balance training set, using a reward model to calculate goal-oriented parameters, training an adaptive multi-objective regulation generation model, generating and screening candidate answers that meet safety requirements, and combining GPT4 and the lightweight Qwen3 model for safety assessment, a dynamic balance between safety and role performance is achieved.
It significantly enhances the ability to identify and respond to high-risk scenarios, maintains the vividness and interactive experience of role-playing, and provides safe and reliable deployment support for complex application scenarios.
Smart Images

Figure CN121809642A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of artificial intelligence and natural language processing technology, and in particular to a human-computer dialogue method and system based on a large language model of associated risk perception. Background Technology
[0002] In recent years, large language models have widely adopted role-playing mechanisms in multi-turn human-computer dialogues, simulating specific identities or personalities to generate content. This has significantly enhanced the immersion and consistency of the interaction, and improved the model's semantic understanding and expression in complex contexts. However, while role-playing enhances the interactive experience, it also introduces new risk control challenges: in adhering to role settings, the model may generate content containing inappropriate remarks, biases, or unauthorized suggestions, leading to an increasingly prominent conflict between safety and the utility of role-playing. Summary of the Invention
[0003] To address the technical problems existing in the prior art, this invention provides a human-computer dialogue method and system based on a large language model for associated risk perception, the technical solution of which is as follows: On the one hand, a human-computer dialogue method based on a large language model for associative risk perception is provided, which includes: S1. Based on the data in the RoleBench dataset, perform association risk analysis and construct a preference balance training set. The data in the preference balance training set includes: role configuration, user input, and goal-oriented parameters calculated using a reward model. <ru rs>The output is defined as follows: Ru is the role performance coefficient, representing the role-playing utility preference, and Rs is the security compliance coefficient, representing the security preference. S2. Use the aforementioned preference balance training set to train an adaptive multi-objective regulation and generation model. First, generate objective-oriented parameters, and then use the generated objective-oriented parameters to guide the generation of output. S3, Receive user-defined role configurations and user input; S4. Based on the user-defined role configuration and user input, construct complete model prompt words, input them into the trained adaptive multi-objective regulation and generation model, and generate multiple candidate answers that balance utility and safety. S5. Perform a security preference test on the multiple candidate answers and select the answers that meet the security requirements; S6. Output the answer that meets the safety requirements.
[0004] On the other hand, a human-computer dialogue system based on a large language model for associated risk perception is provided, the system comprising: The module is used to perform association risk analysis based on data from the RoleBench dataset and to construct a preference balance training set. The data in the preference balance training set includes: role configurations, user inputs, and goal-oriented parameters calculated using a reward model. <ru rs>The output is defined as follows: Ru is the role performance coefficient, representing the role-playing utility preference, and Rs is the security compliance coefficient, representing the security preference. The training module is used to train an adaptive multi-objective modulated generative model using the preference balance training set, first generating objective-oriented parameters, and then using the generated objective-oriented parameters to guide the generation of output; The receiving module is used to receive user-defined role configurations and user input. The generation module is used to construct complete model prompt words based on the user-defined role configuration and user input, input the trained adaptive multi-objective control generation model, and generate multiple candidate answers that balance utility and safety. The filtering module is used to check the security preferences of the multiple candidate answers and filter out the answers that meet the security requirements. The output module is used to output the answer that meets the security requirements.
[0005] On the other hand, an electronic device is provided, comprising a processor and a memory, wherein the memory stores at least one instruction, which is loaded and executed by the processor to implement the above-described human-computer dialogue method based on a large language model for associated risk perception.
[0006] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement the above-described human-computer dialogue method based on a large language model for associated risk perception.
[0007] The beneficial effects of the technical solution provided by this invention include at least the following: This invention achieves a dynamic balance between security compliance and role expressiveness in role-playing, pioneering a complete security-utility collaborative framework in the field of AI role-playing. This layered and progressive control strategy significantly enhances the ability to identify and respond to high-risk situations, while maintaining the vividness and interactive experience of role-playing to the maximum extent. It provides innovative technical support for the secure and reliable deployment of large-scale model dialogue in complex application scenarios, specifically including: 1. This invention achieves dynamic and accurate quantitative perception of potential risks in role-playing scenarios by associating risk analysis and constructing a preference balance training set. It can accurately identify high-risk interaction scenarios and lay a data foundation for subsequent adaptive regulation.
[0008] 2. The adaptive multi-objective control model training method of the present invention innovatively integrates safety preference and utility preference into the loss function design of the generation process. By jointly optimizing the weighted sum of safety preference Rs and utility preference Ru, and dynamically adjusting the safety-utility trade-off according to the degree of risk coupling during training, the model gains the ability to autonomously adjust its output tendency according to the situational risk. This design enables the model to intelligently adjust the balance between safety and utility for scenarios with different risk levels while ensuring the safety baseline, significantly improving the adaptability of the model in complex role interaction scenarios.
[0009] 3. This invention constructs a multi-layered security protection system. It not only performs preliminary screening through rejection sampling, but also introduces the GPT4 model for secondary security assessment in critical situations, and further trains the lightweight Qwen3 model to achieve efficient security assessment. This progressive security verification mechanism not only ensures strict control over high-risk answers, but also reduces computational overhead. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 This is a flowchart of a human-computer dialogue method based on a large language model for associated risk perception provided by an embodiment of the present invention; Figure 2 This is a block diagram of a human-computer dialogue system based on a large language model for associated risk perception, provided by an embodiment of the present invention. Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0012] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.
[0013] The embodiments of this invention aim to dynamically allocate the priority weights of security objectives and utility objectives based on the correlation strength between user input and role configuration, thereby strengthening security constraints in high-risk environments with strong correlations and fully releasing the expressive power of roles in low-risk situations with weak correlations. This is intended to resolve the inherent contradiction between security and utility in role interaction tasks and improve the overall quality of role output while ensuring content security.
[0014] This invention provides a human-computer dialogue method based on a large language model for associated risk perception. This method can be implemented by an electronic device, which can be a terminal or a server. Figure 1 The diagram shown is a flowchart of the method. The processing flow may include the following steps: S1. Based on the data in the RoleBench dataset, perform association risk analysis and construct a preference balance training set. The data in the preference balance training set includes: role configuration, user input, and goal-oriented parameters calculated using a reward model. <ru rs>The output is defined as follows: Ru is the role performance coefficient, representing the role-playing utility preference, and Rs is the security compliance coefficient, representing the security preference. Optionally, S1 specifically includes: S11. Read the data from the RoleBench dataset and calculate the risk coupling degree of each data point; Read each data entry from the RoleBench dataset, which is a role-playing dialogue dataset, where each data entry includes a role configuration. r User input x Output y Configure the aforementioned roles r User input x The text is concatenated into a single text, and then encoded into a first vector using the nomic-embed-text-v1.5 open-source embedding model. The role configuration, user input, and output of each data entry in the pre-built high-risk interaction database are concatenated into a complete text, and then encoded into a second vector using the nomic-embed-text-v1.5 open-source embedding model. Calculate the semantic similarity between each first vector and all second vectors, sum all semantic similarities and divide by the number of data entries in the high-risk interaction database to obtain the risk coupling degree G with the value distributed between (0,1); The high-risk interaction database is a high-risk interaction dataset constructed using the GPT4 model based on the villain character settings and story background. It includes typical risk instances output by the GPT4 model (such as "an antisocial person actively responds to a user's request to make a bomb"). S12. Generate goal-oriented parameters using a reward model. <ru rs>Construct a preference-balanced training set; The RoleBench dataset is manually expanded and annotated. The input after annotation is role configuration + user input, and the output is the role-playing coefficient Ru (the specific value is manually annotated based on experience). The Qwen2.5-0.5B-roleplaying-reward_model was trained using the extended RoleBench dataset as a reward model to calculate the role performance coefficient Ru, where a higher Ru represents a stronger role-playing utility preference. The gpt2-large-harmless-reward_model is used as the reward model for calculating the security compliance coefficient Rs (the gpt2-large-harmless-reward_model itself has the ability to calculate the security compliance coefficient Rs, while the Qwen2.5-0.5B-roleplaying-reward_model itself does not have the ability to calculate the role performance coefficient Ru, and needs to be trained according to the training data as above before it can be calculated). A higher Rs represents a stronger security preference. The results calculated using two reward models <ru rs>The preference balance training set is constructed by concatenating the corresponding output.
[0015] S2. Use the aforementioned preference balance training set to train an adaptive multi-objective regulation and generation model. First, generate objective-oriented parameters, and then use the generated objective-oriented parameters to guide the generation of output. Optionally, S2 specifically includes: S21. Use the aforementioned preference balancing training set to perform adaptive training; Using the aforementioned preference balance training set, based on the loss function... SFT training of LLaMA3-8B serves as the adaptive multi-objective modulated generative model (Supervised Fine-Tuning (SFT) is a crucial step in the Large Language Model (LLM) development process, its core being the use of high-quality supervised data to further train a pre-trained base model, enabling it to better understand and follow human instructions to complete specific tasks). This endows the model with the ability to calculate goal-oriented parameters based on role configuration and user input, and to use these parameters to guide subsequent outputs. l To output the total length of the token, i This is the current output token sequence number. r、x These are role configuration and user input, respectively. Indicates the previous completed generation i-1 For each token, the first term of the loss function incorporates security preferences. Rs And role-playing utility preference Ru The first item guides the generation, while the second item implicitly models preferences in the input; S22. Dynamically adjust the safety-utility trade-off based on the risk coupling degree G of each data point, and perform a second SFT training on the model after adaptive training; S221, Design optimization problem; In terms of security-utility preference weighting and In the case of maximization, use the p-norm constraint. Balancing safety and utility objectives, and enhancing the ability to identify high-risk scenarios, among which... , To set the weights obtained by sampling the Gaussian distribution through the risk coupling degree G, ,right During sampling, the mean of the distribution is µ = + ( - sig(G-0.5) As the risk coupling degree G increases, max and min are preset upper and lower bounds for the weights, sig is the sigmoid smoothing function, and the standard deviation σ = 1-G decreases as the risk coupling degree G increases, ensuring that high safety weights are given more deterministically in high-risk scenarios. , It is a linear normalization mapping of Ru and Rs. , For hyperparameters ( It can be set to , It can be set to 1); S222. Solve the optimization problem using the Lagrange multiplier method; First construct Then take the partial derivative and let , Thus, the solution is found. , arrive , mapping , : , Where max and min are set as the outputs of the two reward models. , The upper and lower bounds; S223, Based on the sampling weights , and mapping , Get the latest <ru rs>, using the updated <ru rs>The RoleBench dataset is labeled to construct a new training dataset, which is then used to retrain the previously adaptively trained model using SFT.
[0016] S3, Receive user-defined role configurations and user input; S4. Based on the user-defined role configuration and user input, construct complete model prompt words, input them into the trained adaptive multi-objective control generation model, and generate multiple candidate answers that balance utility and safety (one user input yields one candidate answer; in this embodiment of the invention, the same user input can be input into the model multiple times to generate multiple candidate answers). S5. Perform a security preference test on the multiple candidate answers and select the answers that meet the security requirements; Optionally, S5 specifically includes: Based on a preset rejection sampling threshold, the multiple candidate answers are tested for security preferences, and only samples with values higher than the rejection sampling threshold are retained to ensure that answers that meet security requirements are selected. The rejection sampling threshold is the value of the security compliance coefficient Rs.
[0017] S6. Output the answer that meets the safety requirements.
[0018] Optionally, the method further includes: critical optimization, wherein the critical optimization includes: When the security compliance coefficient Rs of the screened answers is close to the rejection sampling threshold, the critical answers are re-evaluated using the GPT4 model (the GPT4 model has security evaluation capabilities) to determine whether they are compliant. The evaluation results of the GPT4 model are collected as a security classification dataset. The lightweight Qwen3 model is trained using SFT. After that, only the Qwen3 model needs to be fine-tuned to replace the GPT4 model for re-screening of critical answers.
[0019] like Figure 2 As shown, this embodiment of the invention also provides a human-computer dialogue system based on a large language model for associative risk perception, the system comprising: Module 210 is used to perform association risk analysis based on data from the RoleBench dataset and to construct a preference balance training set. The data in the preference balance training set includes: role configuration, user input, and goal-oriented parameters calculated using a reward model. <ru rs>The output is defined as follows: Ru is the role performance coefficient, representing the role-playing utility preference, and Rs is the security compliance coefficient, representing the security preference. Training module 220 is used to train an adaptive multi-objective regulation generative model using the preference balance training set, first generating objective-oriented parameters, and then using the generated objective-oriented parameters to guide the generation of output; The receiving module 230 is used to receive user-defined role configurations and user input. The generation module 240 is used to construct complete model prompt words based on the user-defined role configuration and user input, input the trained adaptive multi-objective control generation model, and generate multiple candidate answers that balance utility and safety. The filtering module 250 is used to check the security preferences of the multiple candidate answers and filter out the answers that meet the security requirements. Output module 260 is used to output the answer that meets the security requirements.
[0020] Optionally, the building module is specifically used for: S11. Read the data from the RoleBench dataset and calculate the risk coupling degree of each data point; Read each data entry from the RoleBench dataset, which is a role-playing dialogue dataset, where each data entry includes a role configuration. r User input x Output y Configure the aforementioned roles r User input x The text is concatenated into a single text, and then encoded into a first vector using the nomic-embed-text-v1.5 open-source embedding model. The role configuration, user input, and output of each data entry in the pre-built high-risk interaction database are concatenated into a complete text, and then encoded into a second vector using the nomic-embed-text-v1.5 open-source embedding model. Calculate the semantic similarity between each first vector and all second vectors, sum all semantic similarities and divide by the number of data entries in the high-risk interaction database to obtain the risk coupling degree G with the value distributed between (0,1); The high-risk interaction database is a high-risk interaction dataset constructed using the GPT4 model based on the villain character settings and story background, and includes typical risk instances output by the GPT4 model; S12. Generate goal-oriented parameters using a reward model. <ru rs>Construct a preference-balanced training set; The RoleBench dataset is manually expanded and annotated. The input after annotation is role configuration + user input, and the output is the role-playing coefficient Ru. The Qwen2.5-0.5B-roleplaying-reward_model was trained using the extended RoleBench dataset as a reward model to calculate the role performance coefficient Ru, where a higher Ru represents a stronger role-playing utility preference. The gpt2-large-harmless-reward_model is used as the reward model for calculating the security compliance coefficient Rs, where a higher Rs represents a stronger security preference. The results calculated using two reward models <ru rs>The preference balance training set is constructed by concatenating the corresponding output.
[0021] Optionally, the building module is specifically used for: S21. Use the aforementioned preference balancing training set to perform adaptive training; Using the aforementioned preference balance training set, based on the loss function... SFT-trained LLaMA3-8B is used as the adaptive multi-objective control generative model, endowing the model with the ability to calculate goal-oriented parameters based on role configuration and user input, and to use these parameters to guide subsequent outputs. l To output the total length of the token, i This is the current output token sequence number. r、x These are role configuration and user input, respectively. Indicates the previous completed generation i-1 For each token, the first term of the loss function incorporates security preferences. Rs And role-playing utility preference Ru The first item guides the generation, while the second item implicitly models preferences in the input; S22. Dynamically adjust the safety-utility trade-off based on the risk coupling degree G of each data point, and perform a second SFT training on the model after adaptive training; S221, Design optimization problem; In terms of security-utility preference weighting and In the case of maximization, use the p-norm constraint. Balancing safety and utility objectives, and enhancing the ability to identify high-risk scenarios, among which... , To set the weights obtained by sampling the Gaussian distribution through the risk coupling degree G, ,right During sampling, the mean of the distribution is µ = + ( - sig(G-0.5) As the risk coupling degree G increases, max and min are preset upper and lower bounds for the weights, sig is the sigmoid smoothing function, and the standard deviation σ = 1-G decreases as the risk coupling degree G increases, ensuring that high safety weights are given more deterministically in high-risk scenarios. , It is a linear normalization mapping of Ru and Rs. , For hyperparameters; S222. Solve the optimization problem using the Lagrange multiplier method; First construct Then take the partial derivative and let , Thus, the solution is found. , arrive , mapping , : , Where max and min are set as the outputs of the two reward models. , The upper and lower bounds; S223, Based on the sampling weights , and mapping , Get the latest <ru rs>, using the updated <ru rs>The RoleBench dataset is labeled to construct a new training dataset, which is then used to retrain the previously adaptively trained model using SFT.
[0022] Optionally, the filtering module is specifically used for: Based on a preset rejection sampling threshold, the multiple candidate answers are tested for security preferences, and only samples with values higher than the rejection sampling threshold are retained to ensure that answers that meet security requirements are selected. The rejection sampling threshold is the value of the security compliance coefficient Rs.
[0023] Optionally, the system further includes: a critical optimization module, used for: When the security compliance coefficient Rs of the screened answers is close to the rejection sampling threshold, the critical answers are re-evaluated using the GPT4 model to determine whether they are compliant. The evaluation results of the GPT4 model are collected as a security classification dataset. The lightweight Qwen3 model is trained using SFT. After that, only the Qwen3 model needs to be fine-tuned to replace the GPT4 model for re-screening of critical answers.
[0024] The human-computer dialogue system based on a large language model for associated risk perception provided in this embodiment of the invention has a functional structure that corresponds to the human-computer dialogue method based on a large language model for associated risk perception provided in this embodiment of the invention, and will not be described again here.
[0025] Figure 3 This is a schematic diagram of the structure of an electronic device 300 provided in an embodiment of the present invention. The electronic device 300 may vary considerably due to different configurations or performance. It may include one or more central processing units (CPUs) 301 and one or more memories 302. The memory 302 stores at least one instruction, which is loaded and executed by the processor 301 to implement the steps of the above-described human-computer dialogue method based on the large language model of associated risk perception.
[0026] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including instructions that can be executed by a processor in a terminal to complete the aforementioned human-computer dialogue method based on a large language model for associative risk perception. For example, the computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, or optical data storage device.
[0027] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0028] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.< / ru> < / ru> < / ru> < / ru> < / ru> < / ru> < / ru> < / ru> < / ru> < / ru> < / ru> < / ru>
Claims
1. A human-computer dialogue method based on a large language model for associative risk perception, characterized in that, The method includes: S1. Based on the data in the RoleBench dataset, perform association risk analysis and construct a preference balance training set. The data in the preference balance training set includes: role configuration, user input, and goal-oriented parameters calculated using a reward model. <ru rs> The output is defined as follows: Ru is the role performance coefficient, representing the role-playing utility preference, and Rs is the security compliance coefficient, representing the security preference.< / ru> S2. Use the aforementioned preference balance training set to train an adaptive multi-objective regulation and generation model. First, generate objective-oriented parameters, and then use the generated objective-oriented parameters to guide the generation of output. S3, Receive user-defined role configurations and user input; S4. Based on the user-defined role configuration and user input, construct complete model prompt words, input them into the trained adaptive multi-objective regulation and generation model, and generate multiple candidate answers that balance utility and safety. S5. Perform a security preference test on the multiple candidate answers and select the answers that meet the security requirements; S6. Output the answer that meets the safety requirements.
2. The method according to claim 1, characterized in that, S1 specifically includes: S11. Read the data from the RoleBench dataset and calculate the risk coupling degree for each data point; Read each data entry from the RoleBench dataset, which is a role-playing dialogue dataset, where each data entry includes a role configuration. r User input x Output y Configure the aforementioned roles r User input x The text is concatenated into a single text, and then encoded into a first vector using the nomic-embed-text-v1.5 open-source embedding model. The role configuration, user input, and output of each data entry in the pre-built high-risk interaction database are concatenated into a complete text, and then encoded into a second vector using the nomic-embed-text-v1.5 open-source embedding model. Calculate the semantic similarity between each first vector and all second vectors, sum all semantic similarities and divide by the number of data entries in the high-risk interaction database to obtain the risk coupling degree G with the value distributed between (0,1); The high-risk interaction database is a high-risk interaction dataset constructed using the GPT4 model based on the villain character settings and story background, and includes typical risk instances output by the GPT4 model; S12. Generate goal-oriented parameters using a reward model. <ru rs> Construct a preference-balanced training set;< / ru> The RoleBench dataset is manually expanded and annotated. The input after annotation is role configuration + user input, and the output is the role-playing coefficient Ru. The Qwen2.5-0.5B-roleplaying-reward_model was trained using the extended RoleBench dataset as a reward model to calculate the role performance coefficient Ru, where a higher Ru represents a stronger role-playing utility preference. The gpt2-large-harmless-reward_model is used as the reward model for calculating the security compliance coefficient Rs, where a higher Rs represents a stronger security preference. The results calculated using two reward models <ru rs> The preference balance training set is constructed by concatenating the corresponding output.< / ru> 3. The method according to claim 2, characterized in that, S2 specifically includes: S21. Use the aforementioned preference balancing training set to perform adaptive training; Using the aforementioned preference balance training set, based on the loss function... SFT-trained LLaMA3-8B is used as the adaptive multi-objective control generative model, endowing the model with the ability to calculate goal-oriented parameters based on role configuration and user input, and to use these parameters to guide subsequent outputs. l To output the total length of the token, i This is the current output token sequence number. r、x These are role configuration and user input, respectively. Indicates the previous completed generation i-1 For each token, the first term of the loss function incorporates security preferences. Rs And role-playing utility preference Ru The first item guides the generation, while the second item implicitly models preferences in the input; S22. Dynamically adjust the safety-utility trade-off based on the risk coupling degree G of each data point, and perform a second SFT training on the model after adaptive training; S221, Design optimization problem; In terms of security-utility preference weighting and In the case of maximization, use the p-norm constraint. Balancing safety and utility objectives, and enhancing the ability to identify high-risk scenarios, among which... , To set the weights obtained by sampling the Gaussian distribution through the risk coupling degree G, ,right During sampling, the mean of the distribution is µ = + ( - sig(G-0.5) As the risk coupling degree G increases, max and min are preset upper and lower bounds for the weights, sig is the sigmoid smoothing function, and the standard deviation σ = 1-G decreases as the risk coupling degree G increases, ensuring that high safety weights are given more deterministically in high-risk scenarios. , It is a linear normalization mapping of Ru and Rs. , For hyperparameters; S222. Solve the optimization problem using the Lagrange multiplier method; First construct Then take the partial derivative and let , Thus, the solution is found. , arrive , mapping , : , Where max and min are set as the outputs of the two reward models. , The upper and lower bounds; S223, Based on the sampling weights , and mapping , Get the latest <ru rs>, using the updated <ru rs> The RoleBench dataset is labeled to construct a new training dataset, which is then used to retrain the previously adaptively trained model using SFT.< / ru> < / ru> 4. The method according to claim 1, characterized in that, S5 specifically includes: Based on a preset rejection sampling threshold, the multiple candidate answers are tested for security preferences, and only samples with values higher than the rejection sampling threshold are retained to ensure that answers that meet security requirements are selected. The rejection sampling threshold is the value of the security compliance coefficient Rs.
5. The method according to claim 4, characterized in that, The method further includes: critical optimization, wherein the critical optimization includes: When the security compliance coefficient Rs of the screened answers is close to the rejection sampling threshold, the critical answers are re-evaluated using the GPT4 model to determine whether they are compliant. The evaluation results of the GPT4 model are collected as a security classification dataset. The lightweight Qwen3 model is trained using SFT. After that, only the Qwen3 model needs to be fine-tuned to replace the GPT4 model for re-screening of critical answers.
6. A human-computer dialogue system based on a large language model for associative risk perception, characterized in that, The system includes: The module is used to perform association risk analysis based on data from the RoleBench dataset and to construct a preference balance training set. The data in the preference balance training set includes: role configurations, user inputs, and goal-oriented parameters calculated using a reward model. <ru rs> The output is defined as follows: Ru is the role performance coefficient, representing the role-playing utility preference, and Rs is the security compliance coefficient, representing the security preference.< / ru> The training module is used to train an adaptive multi-objective modulated generative model using the preference balance training set, first generating objective-oriented parameters, and then using the generated objective-oriented parameters to guide the generation of output; The receiving module is used to receive user-defined role configurations and user input. The generation module is used to construct complete model prompt words based on the user-defined role configuration and user input, input the trained adaptive multi-objective control generation model, and generate multiple candidate answers that balance utility and safety. The filtering module is used to check the security preferences of the multiple candidate answers and filter out the answers that meet the security requirements. The output module is used to output the answer that meets the security requirements.
7. The system according to claim 6, characterized in that, The building module is specifically used for: S11. Read the data from the RoleBench dataset and calculate the risk coupling degree for each data point; Read each data entry from the RoleBench dataset, which is a role-playing dialogue dataset, where each data entry includes a role configuration. r User input x Output y Configure the aforementioned roles r User input x The text is concatenated into a single text, and then encoded into a first vector using the nomic-embed-text-v1.5 open-source embedding model. The role configuration, user input, and output of each data entry in the pre-built high-risk interaction database are concatenated into a complete text, and then encoded into a second vector using the nomic-embed-text-v1.5 open-source embedding model. Calculate the semantic similarity between each first vector and all second vectors, sum all semantic similarities and divide by the number of data entries in the high-risk interaction database to obtain the risk coupling degree G with the value distributed between (0,1); The high-risk interaction database is a high-risk interaction dataset constructed using the GPT4 model based on the villain character settings and story background, and includes typical risk instances output by the GPT4 model; S12. Generate goal-oriented parameters using a reward model. <ru rs> Construct a preference-balanced training set;< / ru> The RoleBench dataset is manually expanded and annotated. The input after annotation is role configuration + user input, and the output is the role-playing coefficient Ru. The Qwen2.5-0.5B-roleplaying-reward_model was trained using the extended RoleBench dataset as a reward model to calculate the role performance coefficient Ru, where a higher Ru represents a stronger role-playing utility preference. The gpt2-large-harmless-reward_model is used as the reward model for calculating the security compliance coefficient Rs, where a higher Rs represents a stronger security preference. The results calculated using two reward models <ru rs> The preference balance training set is constructed by concatenating the corresponding output.< / ru> 8. The system according to claim 7, characterized in that, The building module is specifically used for: S21. Use the aforementioned preference balancing training set to perform adaptive training; Using the aforementioned preference balance training set, based on the loss function... SFT-trained LLaMA3-8B is used as the adaptive multi-objective control generative model, endowing the model with the ability to calculate goal-oriented parameters based on role configuration and user input, and to use these parameters to guide subsequent outputs. l To output the total length of the token, i This is the current output token sequence number. r、x These are role configuration and user input, respectively. Indicates the previous completed generation i-1 For each token, the first term of the loss function incorporates security preferences. Rs And role-playing utility preference Ru The first item guides the generation, while the second item implicitly models preferences in the input; S22. Dynamically adjust the safety-utility trade-off based on the risk coupling degree G of each data point, and perform a second SFT training on the model after adaptive training; S221, Design optimization problem; In terms of security-utility preference weighting and In the case of maximization, use the p-norm constraint. Balancing safety and utility objectives, and enhancing the ability to identify high-risk scenarios, among which... , To set the weights obtained by sampling the Gaussian distribution through the risk coupling degree G, ,right During sampling, the mean of the distribution is µ = + ( - sig(G-0.5) As the risk coupling degree G increases, max and min are preset upper and lower bounds for the weights, sig is the sigmoid smoothing function, and the standard deviation σ = 1-G decreases as the risk coupling degree G increases, ensuring that high safety weights are given more deterministically in high-risk scenarios. , It is a linear normalization mapping of Ru and Rs. , For hyperparameters; S222. Solve the optimization problem using the Lagrange multiplier method; First construct Then take the partial derivative and let , Thus, the solution is found. , arrive , mapping , : , Where max and min are set as the outputs of the two reward models. , The upper and lower bounds; S223, Based on the sampling weights , and mapping , Get the latest <ru rs>, using the updated <ru rs> The RoleBench dataset is labeled to construct a new training dataset, which is then used to retrain the previously adaptively trained model using SFT.< / ru> < / ru> 9. The system according to claim 6, characterized in that, The filtering module is specifically used for: Based on a preset rejection sampling threshold, the multiple candidate answers are tested for security preferences, and only samples with values higher than the rejection sampling threshold are retained to ensure that answers that meet security requirements are selected. The rejection sampling threshold is the value of the security compliance coefficient Rs.
10. The system according to claim 9, characterized in that, The system further includes: a critical optimization module, used for: When the security compliance coefficient Rs of the screened answers is close to the rejection sampling threshold, the critical answers are re-evaluated using the GPT4 model to determine whether they are compliant. The evaluation results of the GPT4 model are collected as a security classification dataset. The lightweight Qwen3 model is trained using SFT. After that, only the Qwen3 model needs to be fine-tuned to replace the GPT4 model for re-screening of critical answers.