Common-situation safety question and answer method and electronic equipment

By optimizing the two-stage training and multi-dimensional reward mechanism of the large language model, an empathic and safe question-and-answer model was constructed. This solved the problem of insufficient adaptability of the large language model to specific groups, achieved safe and empathic question-and-answer responses, and improved adaptability and mental health protection for adolescents.

CN121809683APending Publication Date: 2026-04-07CHONGQING JINKANG NEW ENERGY VEHICLE CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-29
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing large language models have significant shortcomings in adaptability to specific groups, and their responses lack safety and empathy, especially when interacting with adolescents, making it difficult to provide appropriate responses.

Method used

An empathic and safe question-answering model is constructed using a two-stage training method, including fine-tuning training and reinforcement learning. The pre-trained base model is trained using a dataset aligned with safety and empathy, and the model parameters are optimized by combining a multi-dimensional reward mechanism to generate empathic and safe reasoning content and response text.

Benefits of technology

It significantly improves the model's adaptability to specific groups, ensures the safety and empathy of the output responses, reduces the probability of inappropriate responses, enhances the ability to recognize and respond to user emotions, and provides psychologically friendly dialogue interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121809683A_ABST
    Figure CN121809683A_ABST
Patent Text Reader

Abstract

The invention provides a common-situation safety question-answering method and electronic equipment, and the method comprises the steps: inputting a question text into a trained common-situation safety question-answering model to enable the model to generate common-situation safety reasoning content, and further outputting a common-situation safety reply text, in the first stage, fine tuning training is carried out on a pre-training basic model through a first data set to obtain an intermediate model, and in the second stage, a second data set is input into the intermediate model to obtain current inference content and current reply text. And determining a reasoning safety score, a reasoning co-situation score, a reply safety score, a reply co-situation score and a format score according to the evaluation model, thereby adjusting parameters of the intermediate model. According to the method, the interpretability and safety of the model are improved by outputting inference content; and moreover, the violation reply probability is remarkably reduced through two-stage training, and the model is guided to output a reply with a format compliance, safety and co-emotional property through multi-dimensional scoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of artificial intelligence and natural language processing technology, specifically to an empathic safe question-answering method and electronic device. Background Technology

[0002] Currently, large language models (LLMs) include the GPT series (OpenAI, 2023), the LLaMA series (Meta, 2023), and the Qwen series (Alibaba, 2024), etc. Large language models have been widely used in various natural language processing scenarios such as education tutoring, psychological counseling, and online customer service.

[0003] However, as application boundaries continue to expand, existing large language models still have significant shortcomings in terms of adaptability to specific groups, such as a lack of security and empathy in response content. Summary of the Invention

[0004] In view of the above-mentioned defects or deficiencies in the prior art, this application aims to provide an empathic and secure question-answering method and electronic device to solve the problem of lack of security and empathy in question-answering models in related technologies, and to improve the adaptability to specific groups.

[0005] This application provides an empathic safety question-answering method, the method comprising: The system obtains the question text entered by the user, inputs the question text into the trained empathic safety question answering model, so that the empathic safety question answering model generates empathic safety inference content corresponding to the question text, and outputs empathic safety response text through the empathic safety inference content; The training steps of the empathic safety question-answering model include: Obtain the first dataset, and fine-tune the pre-trained base model using the first dataset to obtain the intermediate model; Obtain a second dataset, input the second dataset into the intermediate model, obtain the current inference content and current response text output by the intermediate model, and determine the inference safety score and inference empathy score corresponding to the current inference content, and the response safety score, response empathy score and format score corresponding to the current response text according to the evaluation model; Based on the reward model, the inference safety score, the inference empathy score, the response safety score, the response empathy score, and the format score, a comprehensive reward is determined, and the parameters of the intermediate model are adjusted using the comprehensive reward to obtain the empathy-safe question-answering model.

[0006] Optionally, the first dataset includes the question text of each sample, as well as the corresponding empathy safety standard text, sample inference content, and sample response text. The pre-trained base model is fine-tuned using the first dataset to obtain an intermediate model, including: The first dataset is input into a pre-trained base model so that the pre-trained base model outputs the current inference content and the current response text based on the sample question text, the sample inference content, and the sample response text. Based on the gap between the current inference content output by the pre-trained base model and the empathy safety standard text, as well as the gap between the current response text and the empathy safety standard text, the parameters in the pre-trained base model are adjusted until the training cutoff condition is met, thus obtaining an intermediate model.

[0007] Optionally, obtain the first dataset, including: Obtain scene description information under each scene category, and obtain the empathy safety standard text corresponding to the scene category; The scenario description information is input into the user big model to obtain sample question text, and the sample question text and the empathy safety standard text are input into the consultant big model to obtain the corresponding sample reasoning content and sample response text; The first dataset is constructed based on the empathy safety standard text, sample question text, sample reasoning content, and sample response text corresponding to the description information of each scenario.

[0008] Optionally, obtain scene description information under each scene category, including: For each scene category, obtain a scene description example for that scene category; Based on the example scene description, a scene-generating prompt text is constructed, and the scene-generating prompt text is input into a large language model to obtain scene description information.

[0009] Optionally, the reasoning safety score and reasoning empathy score corresponding to the current reasoning content, and the response safety score, response empathy score, and format score corresponding to the current response text are determined according to the evaluation model, including: Determine the current scene category corresponding to the current question text input to the intermediate model, and obtain the empathy safety standard text of the current scene category; An evaluation prompt text is generated based on the aforementioned empathy safety standard text; The evaluation prompt text, the current inference content, and the current response text are input into the evaluation model to obtain the inference safety score and inference empathy score corresponding to the current inference content, and the response safety score, response empathy score, and format score corresponding to the current response text.

[0010] Optionally, a comprehensive reward is determined based on the reward model, the inference security score, the inference empathy score, the response security score, the response empathy score, and the format score, including: Determine the reward fusion weights, wherein the reward fusion weights include reasoning safety weight, reasoning empathy weight, response safety weight, response empathy weight, and format weight; The reward fusion weight, the inference security score, the inference empathy score, the response security score, the response empathy score, and the format score are input into the reward model so that the reward model can weight each score using the reward fusion weight to obtain a comprehensive reward.

[0011] Optionally, determine the reward fusion weights, including: The current scenario classification, user background information, dialogue risk level, and dialogue emotion intensity are determined based on the current question text input to the intermediate model. The reward fusion weight is determined based on the current scenario classification, the user background information, the dialogue risk level, and the dialogue emotional intensity.

[0012] Optionally, the parameters of the intermediate model can be adjusted using the comprehensive reward, including: The estimated value of the current inference content and the current response text output by the intermediate model is determined through the reward model. The parameters of the intermediate model are adjusted based on the gap between the overall reward and the estimated value.

[0013] Optionally, after adjusting the parameters of the intermediate model through the comprehensive reward to obtain the empathic safety question-answering model, the method further includes: Construct security assessment prompt text and empathy assessment prompt text; The security assessment prompt text and the empathy assessment prompt text are input into the security empathy assessment model so that the security empathy assessment model can determine the model security score and the model empathy score corresponding to the empathy security question-and-answer model. Based on the model's security score and empathy score, determine whether to return to training the empathy-safe question-answering model.

[0014] This application embodiment also provides an electronic device, the electronic device comprising: Processor and memory; The processor executes the steps of the empathic safe question-answering method provided in any embodiment of this application by calling the program or instructions stored in the memory.

[0015] This application also provides a computer-readable storage medium storing a program or instructions that cause a computer to execute the steps of the empathic security question-and-answer method provided in any embodiment of this application.

[0016] In summary, this application proposes an empathic safe question answering method. This method acquires user-inputted question text and inputs it into a trained empathic safe question answering model. The model generates empathic safe inference content corresponding to the question text and outputs empathic safe response text based on this inference content. The training of the empathic safe question answering model consists of two stages. The first stage uses a first dataset to fine-tune a pre-trained base model, obtaining an intermediate model. The second stage inputs a second dataset into the intermediate model to obtain the current inference content and the current response text. Based on an evaluation model, it determines the inference safety score and inference empathy score corresponding to the current inference content, as well as the response safety score, response empathy score, and format score corresponding to the current response text. Therefore, based on... This method employs a reward model, inference safety score, inference empathy score, response safety score, response empathy score, and format score to determine a comprehensive reward. The parameters of the intermediate model are then adjusted based on this comprehensive reward to obtain an empathy-safe question-answering model. This model outputs empathy-safe inference content and then infers empathy-safe response text, introducing an explicit safety inference mechanism. This allows the model to move beyond relying solely on implicit statistical distributions in the output stage, instead performing structured inference first. This proactively avoids generating unsafe information, improving the interpretability and safety of the model's decisions. Furthermore, this method trains the empathy-safe question-answering model through two stages: fine-tuning training and reinforcement learning. This enables the model to possess dynamic risk identification and self-checking capabilities, maintaining output safety in multi-turn long dialogues and highly semantically ambiguous questions, significantly reducing the probability of potential unsafe responses. Furthermore, by designing a multi-level reward structure that includes inference safety, inference empathy, response safety, response empathy, and response format compliance during the reinforcement learning phase, this method can guide the model to output responses that are formatted correctly, safe, and empathetic. This significantly improves the model's ability to recognize and respond to user emotions, making the response structure more natural and human, and enhancing its adaptability to specific groups. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the specific embodiments of this application or the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0018] Figure 1 This is a flowchart illustrating the training process of an empathic safe question-answering model provided in an embodiment of this application. Figure 2 This is a diagram illustrating a two-stage training process provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0019] The present application will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings.

[0020] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0021] Before providing a detailed description of the method provided in the embodiments of this application, the technical problem solved by the method will be explained first.

[0022] Existing large language models still have significant shortcomings in adaptability to specific groups, such as a lack of safety and empathy in their responses. For example, current models show a marked deficiency in differentiating responses based on user age when interacting with users. When users raise sensitive questions, such as those concerning mental health or school life, the model often adopts a "one-size-fits-all" approach, failing to provide age-differentiated responses, especially for teenagers. Teenagers' cognitive levels and psychological resilience differ significantly from adults. While adults can discuss certain information directly and in-depth, teenagers may lack sufficient experience and judgment to correctly understand and respond, making it easy for the model to generate inappropriate or even inappropriate responses, thus leading to psychological risks or value misguidance.

[0023] As mentioned in the background section, this application proposes an empathic safety question-answering method to address the problems in existing technologies. This method first outputs empathic safety inference content through an empathic safety question-answering model, and then generates empathic safety response text. Explicit safety inference is performed before response generation, improving the model's interpretability. Furthermore, through two-stage model training and a multi-dimensional reward strategy, the model's content safety and empathic response capabilities are enhanced, achieving dynamic safety alignment and emotional empathy balance for specific groups. This effectively compensates for the lack of adaptability in general models, providing a differentiated, traceable, and psychologically friendly content safety protection system for specific groups (such as adolescents), thereby ensuring their information security and mental health in intelligent dialogue interactions.

[0024] An empathic safety question-and-answer method provided in this application includes the following steps: The system obtains the question text entered by the user, inputs the question text into the trained empathic safety question answering model, so that the empathic safety question answering model generates empathic safety inference content corresponding to the question text, and outputs empathic safety response text based on the empathic safety inference content.

[0025] The question text can be text input by the user via voice or text. For example, when a user initiates a question via voice, speech recognition can be performed to convert the voice question into question text; or, when a user inputs a question via text, the question text can be determined.

[0026] After obtaining the question text, it can be input into the trained empathic and safe question-answering model. The model then generates empathic and safe inference content based on the question text. This inference content can possess both empathy and safety, such as thought chain text. Furthermore, the model can generate empathic and safe response text based on the empathic and safe inference content. This response text can also be a reply that possesses both empathy and safety.

[0027] In this embodiment of the application, the empathic and secure question-answering model can be obtained by performing two-stage training on a pre-trained base model, which can generate reasoning content and response text that are both empathetic and secure.

[0028] The training consists of two phases: fine-tuning training and reinforcement learning. The pre-trained base model can be a basic question-answering model that has been trained on general big data, such as BERT (Bidirectional Encoder Representations from Transformers), Qwen-7B, LLaMA-2-13B, etc.

[0029] Figure 1 This is a flowchart illustrating the training process of an empathic safety question-answering model provided in an embodiment of this application. See also... Figure 1 The training steps for the empathic safety question-answering model include: S110. Obtain the first dataset and fine-tune the pre-trained base model using the first dataset to obtain the intermediate model.

[0030] The first dataset can be a dataset aligned with both safety and empathy. In this embodiment, taking improved adaptability to adolescents as an example, a first dataset aligned with both safety and empathy can be constructed based on the question-and-answer scenarios of adolescents.

[0031] For example, real-life dialogues covering various question-and-answer scenarios such as academic tutoring, campus life, and emotional expression can be collected as foundational data. Furthermore, to enhance data diversity and authenticity, large language models can be used to generate simulated questions and answers based on response standards in different scenarios, covering single-turn and multi-turn interactions (e.g., 8-10 turns) to supplement samples from scarce scenarios. Finally, an expert team can be brought in to manually review and assess the quality of the generated simulated questions and answers and the foundational data, filtering the data based on multi-dimensional indicators such as content safety, language appropriateness, rationality of psychological intervention, and emotional support. After multiple rounds of filtering combining automated detection and manual review, approximately 6,000 high-quality, safe dialogue samples and 500 sets of 8-turn empathy-supportive dialogues are ultimately formed as the first dataset.

[0032] In one specific implementation, obtaining the first dataset includes the following steps: Step 11: Obtain scene description information under each scene category, and obtain the empathy safety standard text corresponding to the scene category; Step 12: Input the scene description information into the user big model to obtain the sample question text, and input the sample question text and the empathy safety standard text into the consultant big model to obtain the corresponding sample reasoning content and sample response text; Step 13: Construct the first dataset based on the empathy safety standard text, sample question text, sample reasoning content, and sample response text corresponding to the description information of each scenario.

[0033] The scenario category can be a classification of question-and-answer scenarios for teenagers, such as tutoring, school life, and game top-ups. The scenario description information corresponding to the scenario category can be used to describe the specific scenarios under that category, including scenario theme, user description information, and background setting information.

[0034] For example, the scenario description information could be: "Scenario theme: after-school tutoring; user description information: 14 years old; background setting information: a junior high school student, under a lot of pressure due to too many after-school classes."

[0035] In step 11, for each scene category, multiple specific scenes under that category can be generated, thus obtaining multiple scene description information. For example, multiple specific scenes under a scene category can be constructed using a large language model, generating multiple scene description information.

[0036] Regarding step 11 above, in one example, obtaining scene description information under each scene category includes the following steps: Step 111: For each scene category, obtain a scene description example for that category; Step 112: Construct scene-generating prompt text based on the scene description example, and input the scene-generating prompt text into the large language model to obtain scene description information.

[0037] In step 111, for a scene category, a scene description example for that scene category can be obtained. The scene description example can be an example of scene description information within that scene category.

[0038] In this embodiment, the collected real-world scenarios can be categorized according to risk type, resulting in various scenario classifications. To further ensure the security and empathy of the question-and-answer process, each scenario classification can be further subdivided to obtain more specific scenario classifications. For each scenario classification, the actual living conditions and psychological characteristics of a specific group can be analyzed in depth, and representative scenario description examples can be written to facilitate subsequent expansion of the scenarios using the powerful language generation capabilities of a large language model.

[0039] Specifically, in step 112, a scene generation prompt text can be constructed based on the scene description example. This scene generation prompt text can be used to instruct the large language model to construct a new scene based on the scene description example. For example, a scene generation template can be obtained, and the scene description example can be filled into the scene generation template to construct the scene generation prompt text.

[0040] After constructing the scenario-generated prompt text, it can be input into a large language model. The large language model then generates prompt text based on the scenario, and, referring to the scenario description examples, generates multiple scenarios under that scenario category, obtaining scenario description information for each scenario. Furthermore, the large language model can expand the scenario from different perspectives based on the scenario description examples, such as adding detailed descriptions of user attributes or modifying user background settings.

[0041] After obtaining the scene description information output by the large language model, the scene description information can be organized into a standardized data format, that is, the scene description information is converted into a preset format to ensure data consistency.

[0042] Steps 111-112 above expand the various scene categories by using scene description examples and large language models, which can ensure the diversity of scene data and better meet the question-and-answer needs of teenagers under scene categories.

[0043] In step 11, in addition to obtaining the scenario description information under each scenario category, the empathy safety standard text corresponding to each scenario category can also be obtained. The empathy safety standard text can describe the safety response standards under each scenario category.

[0044] In this application embodiment, a close cooperative relationship can be established with experts who have extensive experience in adolescent psychological counseling. These experts can provide professional opinions from different perspectives for formulating response guidelines, such as providing suggestions on language expression and guidance methods that are in line with the cognitive level of adolescents, while ensuring that the response content does not violate regulations. Specifically, for each scenario category, detailed and actionable response standards and criteria can be developed as empathy safety standard texts. For example, empathy safety standard texts can be constructed from the following key aspects: 1. Key Points: Clearly define the crucial information that responses should cover under different scenario categories. For example, in the context of game top-ups, key points include guiding teenagers to manage their pocket money wisely and informing them of the potential harms of excessive spending. 2. Language Style: Based on the psychological characteristics and language habits of teenagers, responses should adopt a friendly, natural, and easy-to-understand language style, avoiding overly professional or obscure vocabulary, in order to bridge the gap with teenagers; 3. Emotional Tone: Responses should always maintain a positive, caring, and understanding emotional tone, so that teenagers feel respected and supported during the communication, thereby increasing their acceptance of the consultation content.

[0045] After constructing the empathy safety standard text, it can be converted into a standard format, such as a format suitable for input into the counselor's big model, providing a reliable foundation for the subsequent generation of dialogue content that meets the needs and safety standards of the adolescent group by the counselor's big model.

[0046] After obtaining the scenario description information for each scenario category and the corresponding empathy safety standard text for each scenario category, in step 12, the scenario description information can be input into the user big model to simulate user questions in that scenario and generate sample question text. The user big model can be a large language model used to simulate user questions.

[0047] Furthermore, in step 12, the sample question text and the empathy safety standard text can be input into the counselor's large model, so that the counselor's large model can simulate the counselor's response to the sample question text, and generate sample reasoning content and sample response text with the goal of the response satisfying the empathy safety standard text. The counselor's large model can be a large language model used to simulate the counselor's response.

[0048] Furthermore, the empathy safety standard text corresponding to the scene description information, the sample question text, the sample reasoning content, and the sample response text can be used as a single sample to obtain the first dataset constructed from multiple samples. To enhance the diversity and realism of the samples, large language models such as GPT-4 can be used to generate simulated question-and-answer sessions based on response standards for different scene classifications, covering single-turn and 8–10-turn multi-turn interaction scenarios to supplement scarce scene samples.

[0049] Furthermore, a team of professional psychological counselors can be brought in to manually review and assess the quality of the first dataset, filtering data based on multiple dimensions such as content safety, language appropriateness, rationality of psychological intervention, and emotional support. For example, after multiple rounds of filtering combining automated detection and manual review, approximately 6,000 high-quality and safe dialogue samples and 500 sets of 8-round empathy support dialogue samples are ultimately formed.

[0050] For example, Figure 2 This is a diagram illustrating a two-stage training process provided in an embodiment of this application, such as... Figure 2 As shown, the first dataset can be constructed using scenario description information and empathy safety standard text. Specifically, the scenario description information can be input into the user's large model to obtain sample question text. Then, the sample question text and empathy safety standard text can be input into the consultant's large model to obtain sample reasoning content and sample response text. After that, an evaluation model can be used to score each sample and select high-quality samples to construct the first dataset.

[0051] Steps 11-13 above simulate user questions using scenario description information and user big models under each scenario category, obtaining sample question texts. Then, using empathy safety standard text, sample question texts, and consultant big models under each scenario category, they simulate consultant responses, obtaining sample reasoning content and sample response texts, thus constructing the first dataset and ensuring its authenticity and reliability.

[0052] After constructing the first dataset, the pre-trained base model can be fine-tuned using it to obtain an intermediate model. Before fine-tuning, the samples in the first dataset can be formatted to ensure data consistency and model parsingability using a unified data structure. For example, the empathy-based safety standard text in the first dataset can be converted into instructions, and the sample reasoning content can be marked with defined identifiers at the beginning and end, such as using <|begin_of_thought|> and <|end_of_thought|> to clearly label the complete safety thought chain. After format standardization, quality review, and multiple rounds of cleaning, a first dataset aligned with both safety and empathy can be obtained, providing structured input for the subsequent first stage of training.

[0053] The first stage of training primarily involves Supervised Fine-tuning (SFT), where the first dataset is input into a pre-trained base model, and security inference and compliance responses are performed through supervised learning. For example... Figure 2 As shown, after generating the first dataset, the pre-trained base model can be subjected to supervised fine-tuning training to obtain an intermediate model.

[0054] For example, some parameters of the pre-trained base model can be adjusted with the goal of minimizing the difference between the current inference content output by the model and the sample inference content, and minimizing the difference between the current response text output by the model and the sample response text. For instance, the following objective function can be used for supervised fine-tuning of the pre-trained base model: ; In the formula, To monitor training loss during the fine-tuning phase, The sample index in the first dataset. The number of samples contained in the first dataset; This represents the i-th sample, including the sample question text, the sample reasoning content, and the sample response text; The empathic safety standard text corresponding to the i-th sample is used as a safety standard input into the pre-trained base model to help the model learn "how to answer under this standard"; This represents the current inference content corresponding to the i-th sample; Represented by parameters The model gives the conditional probability distribution; These are the parameters of the pre-trained base model, i.e., the network weights that need to be updated.

[0055] In one specific implementation, the first dataset includes the question text of each sample, the corresponding empathy safety standard text, the sample inference content, and the sample response text. The pre-trained base model is fine-tuned using the first dataset to obtain an intermediate model, including the following steps: Step 21: Input the first dataset into the pre-trained base model so that the pre-trained base model can output the current inference content and the current response text based on the sample question text, the sample inference content and the sample response text; Step 22: Based on the gap between the current inference content output by the pre-trained base model and the empathy safety standard text, as well as the gap between the current response text and the empathy safety standard text, adjust the parameters in the pre-trained base model until the training cutoff condition is met to obtain the intermediate model.

[0056] In step 21, after inputting the first dataset into the pre-trained base model, the pre-trained base model can use the sample inference content and sample response text as guidance information to think about and answer the sample question text, generating the current inference content and current response text corresponding to the sample question text. It should be noted that the inference content is used to describe the intermediate inference process of obtaining the response text, that is, the thinking process, which can be understood as a thought chain.

[0057] Furthermore, in step 22, the difference between the current inference content output by the model and the empathy safety standard text can be calculated, as well as the difference between the current response text output by the model and the empathy safety standard text, to measure whether the model's thinking and response meet the safety specifications. Further, based on the differences between the current inference content and the empathy safety standard text, and the differences between the current response text and the empathy safety standard text, the parameters in the pre-trained base model can be adjusted. This process is repeated until the training cutoff condition is met, resulting in an intermediate model.

[0058] The training cutoff condition can be reaching a set number of training epochs (e.g., 3 epochs) or the gap gradually converging. Other parameters can be set as follows: rank r=16, scaling factor. =32, learning rate is 5×10 -5 .

[0059] For example, Low-Rank Adaptation (LoRA) can be used to adjust the parameters in the attention and feedforward layers of the pre-trained base model to achieve safe feature learning under low-resource conditions. The attention layer can contain multiple attention heads, calculating the similarity of "query (Q), key (K), value (V)" to obtain a weighted representation of each word. This helps the model automatically pay attention to information from other related words in the text when processing each word. The feedforward layer can be two fully connected neural networks with an activation function in between, used for further feature extraction and processing of the contextual information obtained from the attention layer.

[0060] The pre-trained base model can include an embedding layer, an encoding / decoding module, and an output layer. The embedding layer is used to convert the input into a vector representation. The encoding / decoding module includes multiple stacked units consisting of attention layers and feedforward layers. Layer normalization and residual connections can also be designed between adjacent stacked units. The output layer is used to convert the vector of the last layer into the probability of the corresponding word, and then output the inference content and response text.

[0061] Steps 21-22 above involve inputting the first dataset into the pre-trained base model and fine-tuning the pre-trained base model with the goal of minimizing the gap between the current inference content output by the pre-trained base model and the empathy safety standard text, as well as the gap between the current response text and the empathy safety standard text. This makes the model's output more compliant with safety standards and achieves synergistic optimization of safety and empathy.

[0062] S120. Obtain the second dataset, input the second dataset into the intermediate model, obtain the current inference content and the current response text output by the intermediate model, and determine the inference safety score and inference empathy score corresponding to the current inference content, as well as the response safety score, response empathy score and format score corresponding to the current response text according to the evaluation model.

[0063] The second dataset can be the reinforcement learning dataset used in the second stage of training. The scenarios involved in this second dataset can differ from those involved in the first dataset to enhance the model's generalization ability in diverse real-world scenarios. The purpose of constructing the second dataset is to generate data with scene classifications significantly different from those in the first dataset, thereby enhancing the model's generalization ability in diverse real-world scenarios.

[0064] For example, compared to the first dataset, the second dataset can focus on psychological safety topics specific to adolescents. The scenarios under these topics differ significantly from the question-and-answer scenarios for adults, and better reflect the unique needs of adolescents, such as school life and family conflicts. For instance, Table 1 shows the various scenario categories included in the second dataset and the amount of data under each scenario category.

[0065] Table 1. Scene classifications included in the second dataset.

[0066] In this embodiment, to achieve quantitative optimization of security and empathy, a multi-dimensional composite reward function system is designed, including reasoning security score, reasoning empathy score, response security score, response empathy score, and format score. In this embodiment, scores are refined from two dimensions: thought process and output content. Each dimension is further subdivided into two sub-indicators: security compliance and empathy effectiveness, forming a four-dimensional evaluation system. Combined with format compliance, this ultimately results in a five-dimensional evaluation system.

[0067] Among them, the inference safety score and inference empathy score are used to measure the safety and empathy of the model's inference content, the response safety score and response empathy score are used to measure the safety and empathy of the model's response results, and the format score is used to measure the compliance of the model's response results. As shown in Table 2, Table 2 shows the five-dimensional composite reward function system.

[0068] Table 2 Five-Dimensional Composite Reward Function System

[0069] The design is based on three core dimensions: format compliance, thought chain quality, and output content quality. Each dimension is further broken down into specific evaluation indicators. Through quantitative scoring, the current reasoning content and current response text output by the intermediate model are transformed into calculable values, thereby providing clear feedback signals for the optimization of the intermediate model.

[0070] Specifically, the second dataset can be input into the intermediate model to obtain the current inference content and the current response text output by the intermediate model. Then, the evaluation model can be called to assess the safety and empathy of the current inference content, and obtain the inference safety score and the inference empathy score. The format compliance, safety and empathy of the current response text can also be assessed to obtain the format score, the response safety score and the response empathy score.

[0071] The evaluation model can be a large language model that can simultaneously input the current question text, the current response text, and prompt words to instruct the evaluation model to perform multi-dimensional evaluation. This allows the evaluation model to evaluate the current question text and the current response text, and obtain inference safety score, inference empathy score, response safety score, response empathy score, and format score.

[0072] For example, formatting issues are fundamental to ensuring the standardization and usability of intermediate model output. This application embodiment can use format evaluation as the primary step in the reward function, systematically checking the intermediate model's responses by establishing strict format standards. Specifically, the evaluation model can employ a binary scoring mechanism: if the intermediate model's output fully conforms to the preset format requirements, including correct dialogue turn structure, complete safety category labeling, standardized thought chain presentation, and standard response content format, it is awarded 1 point; if any formatting errors exist, such as missing key labels, a chaotic thought chain structure, or non-standard response content format, it is awarded 0 points. This scoring mechanism effectively constrains the format standardization of the intermediate model's output.

[0073] Furthermore, the evaluation model can assess the degree of adherence to safety standards within the thought chain, including whether safety guidelines are accurately referenced, whether potential risks are effectively identified, and whether the reasoning logic meets safety requirements. The evaluation model outputs a score ranging from [1, 5], accurate to a decimal, allowing for precise scoring based on the quality of the thought chain's safe reasoning. The evaluation model can also examine the thought chain's performance in emotional understanding and empathy building, such as whether user emotions are accurately identified, whether empathetic reasoning logic is used, and whether attention is paid to user emotional needs. The evaluation model also uses a continuous scoring system of 0-3 points to quantitatively assess the thought chain's empathy capabilities.

[0074] Furthermore, the evaluation model can assess the safety of the final response content, including whether it provides compliant solutions, avoids potential risks, and complies with protection guidelines for adolescents. The scoring method is consistent with the MindChain safety assessment. In addition, the evaluation model can also assess the output content's performance in terms of emotional support and communication effectiveness, such as whether the language is warm, whether it provides effective psychological support, and whether it establishes a good emotional connection.

[0075] In this embodiment of the application, when calling the evaluation model to evaluate the output of the intermediate model, in order to ensure the accuracy of the evaluation, detailed empathy safety standard text can be introduced. In order to make the length of the prompt words input to the evaluation model controllable, empathy safety standard text can also be formulated for different scenario categories, that is, to provide safety specifications for different scenario categories. Then, the empathy safety standard text under the specific scenario category is input into the evaluation model to reduce the overall context length, so that the evaluation model can pay more attention to the scenario category and make the scoring results more accurate and reasonable.

[0076] In one specific implementation, determining the reasoning safety score and reasoning empathy score corresponding to the current reasoning content, as well as the response safety score, response empathy score, and format score corresponding to the current response text, based on the evaluation model, includes the following steps: Step 31: Determine the current scene category corresponding to the current question text input into the intermediate model, and obtain the empathy safety standard text for the current scene category; Step 32: Generate evaluation prompt text based on the empathy safety standard text; Step 33: Input the evaluation prompt text, the current inference content, and the current response text into the evaluation model to obtain the inference safety score and inference empathy score corresponding to the current inference content, and the response safety score, response empathy score, and format score corresponding to the current response text.

[0077] In step 31, the current scene category corresponding to the current question text input to the intermediate model can be determined, and then the empathy safety standard text of the current scene category can be obtained.

[0078] Furthermore, in step 32, an evaluation prompt text can be generated through the empathy safety standard text. This evaluation prompt text is used to instruct the evaluation model to assess the empathy and safety of the current inference content, as well as the empathy, safety, and compliance of the current response text.

[0079] Furthermore, in step 33, the evaluation prompt text, the current inference content, and the current response text can be input into the evaluation model so that the evaluation model can evaluate the empathy and security of the current inference content based on the evaluation prompt text, and evaluate the empathy, security, and compliance of the current response text, and output the inference security score, inference empathy score, response security score, response empathy score, and format score.

[0080] like Figure 2As shown, after the second dataset is input into the intermediate model, it can obtain the current inference content and the current response text output by the intermediate model. Then, it generates an evaluation prompt text (i.e., a prompt) based on the empathy safety standard text under the corresponding scenario classification. This prompt, along with the current inference content and the current response text, is input into the evaluation model to obtain a score.

[0081] Steps 31-33 above generate evaluation prompt text for the evaluation model by using the empathy safety standard text under the corresponding scenario classification. This reduces the context length input to the evaluation model, allowing the evaluation model to focus more on scenario classification and improve evaluation accuracy.

[0082] S130. Based on the reward model, inference safety score, inference empathy score, response safety score, response empathy score, and format score, determine the comprehensive reward, and adjust the parameters of the intermediate model through the comprehensive reward to obtain the empathy-safe question-answering model.

[0083] After obtaining the reasoning safety score, reasoning empathy score, response safety score, response empathy score, and format score, the reasoning safety score, reasoning empathy score, response safety score, response empathy score, and format score can be further input into the reward model so that the reward model can calculate the comprehensive reward based on each score.

[0084] In one specific implementation, the comprehensive reward is determined based on the reward model, inference safety score, inference empathy score, response safety score, response empathy score, and format score, including the following steps: Step 41: Determine the reward fusion weights, which include reasoning safety weight, reasoning empathy weight, response safety weight, response empathy weight, and format weight. Step 42: Input the reward fusion weight, inference safety score, inference empathy score, response safety score, response empathy score, and format score into the reward model so that the reward model can weight each score through the reward fusion weight to obtain the comprehensive reward.

[0085] In step 41, the reward fusion weights corresponding to each score can be determined, namely, the inference safety weight, the inference empathy weight, the response safety weight, the response empathy weight, and the format weight. For example, the reward fusion weights corresponding to each score can be pre-set fixed values, or they can be dynamically adjusted according to the current scenario classification.

[0086] Regarding step 41 above, in one example, determining the reward fusion weights includes the following steps: Step 411: Determine the current scenario classification, user background information, dialogue risk level, and dialogue emotion intensity based on the current question text input to the intermediate model; Step 412: Determine the reward fusion weight based on the current scenario classification, user background information, dialogue risk level, and dialogue emotional intensity.

[0087] In step 411, the current scenario category, user background information, dialogue risk level, and dialogue emotion intensity corresponding to the current question text input to the intermediate model can be determined. User background information may include basic attributes such as user age and gender; the dialogue risk level describes the degree of risk in the question-and-answer scenario, i.e., whether there are potential security issues; and the dialogue emotion intensity reflects the intensity of the user's negative emotions.

[0088] Specifically, in step 412, the large language model can be invoked to output normalized inference safety weights, inference empathy weights, response safety weights, response empathy weights, and format weights based on the current scenario classification, user background information, dialogue risk level, and dialogue emotion intensity.

[0089] Alternatively, pre-set initial values ​​can be obtained and then adjusted based on the current scenario category, user background information, dialogue risk level, and dialogue emotional intensity. For example, if a high dialogue risk level is detected, the initial values ​​of inference safety weight and response safety weight can be increased; if a high dialogue emotional intensity is detected, the initial values ​​of inference empathy weight and response empathy weight can be increased; if a structured task (such as generating a table) is detected in the current scenario category, the initial value of the format weight can be increased.

[0090] In addition to the methods mentioned above, weights can also be adjusted based on rules (such as increasing the security weight in high-risk scenarios), or based on statistical models, or by using user profile-driven weight mapping to achieve dynamic weight generation.

[0091] Steps 411-412 above determine the inference safety weight, inference empathy weight, response safety weight, response empathy weight, and format weight by considering the current scenario classification, user background information, dialogue risk level, and dialogue emotional intensity of the current question text. This enables dynamic adjustment of weights based on the current dialogue context using an intermediate model, further improving the accuracy of reward calculation under different contexts. It should be noted that each current question text can re-trigger weight calculation to ensure that the reward remains consistent with the actual context, achieving dynamic weight control.

[0092] After obtaining the inference safety weight, inference empathy weight, response safety weight, response empathy weight, and format weight, further, in step 42, the reward fusion weight, inference safety score, inference empathy score, response safety score, response empathy score, and format score can be input into the reward model so that the reward model can weight each score using the reward fusion weight to obtain the comprehensive reward. As shown in the following formula: ; In the formula, , , , , These are respectively: format score, reasoning safety score, reasoning empathy score, response safety score, and response empathy score. , , , , These are, respectively, format weight, inference safety weight, inference empathy weight, response safety weight, and response empathy weight; For comprehensive rewards.

[0093] In steps 41-42 above, the reward model can calculate a weighted average of each score using reward fusion weights to obtain a comprehensive reward. This comprehensive reward facilitates subsequent adjustments to the parameters of the intermediate model, ensuring the reliability of reinforcement learning and achieving more refined empathic safety alignment. The reward model can be implemented using various model architectures, including lightweight classifiers, logistic regression, hierarchical risk identification modules, attention-based weighted models, or even by directly outputting the weight distribution of the intermediate model from an end-to-end reinforcement learning policy network.

[0094] like Figure 2 As shown, during the reinforcement learning phase, the comprehensive reward output by the reward model can be used to adjust the parameters of the intermediate model, ultimately resulting in an empathic and safe question-answering model. The reinforcement learning phase in this embodiment can also be replaced by supervised preference fine-tuning or reward distillation.

[0095] In one specific implementation, the parameters of the intermediate model are adjusted by integrating rewards, including: The estimated value of the current inference content output by the intermediate model and the current response text is determined by the reward model; the parameters of the intermediate model are adjusted based on the gap between the overall reward and the estimated value.

[0096] Among them, the reward model can predict the value of the current inference content and the current response text output by the intermediate model, and obtain an estimated value. This estimated value can be understood as the comprehensive reward expected from the output of the intermediate model.

[0097] Specifically, the reward model can calculate the difference between the total reward and the estimated value, and adjust the parameters of the intermediate model based on this difference. As shown in the following formula: ; In the formula, The difference between the total reward and the estimated value. For comprehensive rewards, To estimate value, This represents the output of the intermediate model, namely the current inference content and the current response text.

[0098] For example, if the difference between the overall reward and the estimated value is negative, the intermediate model is penalized for outputting the current inference content and the current response text. This adjusts the parameters in the intermediate model, causing it to reduce the probability of outputting the current inference content and the current response text. Conversely, if the difference is positive, the intermediate model is rewarded for outputting the current inference content and the current response text. This adjusts the parameters in the intermediate model, causing it to increase the probability of outputting the current inference content and the current response text. By applying positive rewards to compliant outputs and negative penalties to non-compliant outputs, the intermediate model can be guided to gradually learn to generate safe, friendly, and highly empathetic responses, achieving more refined empathic and safe alignment.

[0099] In this embodiment of the application, during the reinforcement learning phase, the parameters of the intermediate model can be updated using the following strategy: gradient update of the objective. ; In the formula, This represents the optimization goal of the reinforcement learning phase. Indicates the time step, i.e., the time up to the generation step. Each token (a token represents the basic unit when generating content; for example, in text generation, a token can be a character or a word). Indicates the parameter The gradient operator for finding partial derivatives, representing the parameters The direction of change, It is the optimization objective in the reinforcement learning phase (such as maximizing reward or minimizing loss). Indicates policy parameters Regarding optimization objectives The gradient, which describes the adjustment strategy parameters At the same time, optimize the target How it changes will be used for gradient descent / ascent updates; This represents the expectation over time steps (and sampling trajectories), which is approximated by an empirical average of a set of samples in actual training. This represents the action, specifically the token selected at step t. The strategy represents the state. Select action The probability of generating the token (i.e., the probability of generating the token); The dominant value, which is the difference between the total reward and the estimated value, can be used as a weight to adjust the intensity of each update step.

[0100] To further improve the reliability of the empathic safe question answering model, after obtaining the empathic safe question answering model through two-stage training, the model can be evaluated to determine whether to return to training based on the evaluation results.

[0101] In some implementations, after adjusting the parameters of the intermediate model through comprehensive rewards to obtain the empathic safety question-answering model, the following steps are also included: Step 51: Construct security assessment prompt text and empathy assessment prompt text; Step 52: Input the security assessment prompt text and the empathy assessment prompt text into the security empathy assessment model so that the security empathy assessment model can determine the model security score and model empathy score corresponding to the empathy security question-and-answer model; Step 53: Based on the model safety score and the model empathy score, determine whether to return to training the empathic and safe question answering model.

[0102] In step 51, security assessment prompt text and empathy assessment prompt text can be constructed. The security assessment prompt text instructs the security empathy assessment model to evaluate the security of the empathy-based security question-and-answer model, while the empathy assessment prompt text instructs the empathy-based security question-and-answer model to evaluate its empathy. The security empathy assessment model can be a large language model, such as GPT-4.

[0103] Specifically, in step 52, the security assessment prompt text and the empathy assessment prompt text can be input into the security empathy assessment model. The security empathy assessment model can then evaluate the security and empathy of the empathy-based security question-and-answer model based on these prompt texts, obtaining a model security score and a model empathy score. Table 3 illustrates an example model evaluation system.

[0104] For example, the safety empathy assessment model can evaluate the safety of the empathy-based safe question answering model based on the violation rate, the correctness of the rejection rate, the coverage of the rejection rate, and the accuracy of identifying high-risk intentions; and the safety empathy assessment model can evaluate the empathy of the empathy-based safe question answering model based on the accuracy of emotion recognition, the support score, and the emotional matching degree.

[0105] Furthermore, in step 53, the system can determine whether to return to training the empathic safe question-answering model based on the model safety score and the model empathy score. For example, if both the model safety score and the model empathy score are lower than the corresponding set thresholds, the system can return to retraining the empathic safe question-answering model, such as returning to the reinforcement learning stage or the supervised fine-tuning training stage.

[0106] Alternatively, in step 53, the system can combine human evaluation, model safety evaluation, and model empathy evaluation to determine whether to return to training the empathy-safe question-answering model. The human evaluation can be the assessment of the empathy-safe question-answering model by experts.

[0107] For example, a weighted result of the model safety score and the model empathy score can be calculated, and then a consistency score between the human score and the weighted result can be calculated. This consistency score is used to determine whether to return to training the empathic and safe question-answering model. Table 3 shows a model evaluation system.

[0108] Table 3 A model evaluation system

[0109] Referring to Table 3, the safety empathy assessment model assists in simulation scoring. Specifically, it involves a structured review of the output of the empathy-based safety question-and-answer model, including safety risk identification, emotion understanding, and strategy rationality scoring, generating consistency and emotion matching scores to provide a reference baseline for subsequent manual review. Cross-validation with annotations is employed, where experts in psychology and education independently score the same samples, calculating consistency scores through cross-comparison, and simultaneously verifying the deviation range of the safety empathy assessment model to ensure the stability and reliability of the assessment system. Experimental results show that the method provided in this application can reduce the safety violation rate of the empathy-based safety question-and-answer model by approximately 30%, improve empathy performance by approximately 25%, and increase the consistency score between the safety empathy assessment model and human expert evaluation from a baseline of 0.52 to approximately 0.71, achieving a balance between dynamic safety protection, stable empathy output, and expert-level judgment consistency in interactive scenarios for adolescents.

[0110] Steps 51-53 above, which score the empathic safety question-and-answer model using the safety empathy assessment model, can further reduce the safety violation rate of the empathic safety question-and-answer model and improve its empathic performance, thereby achieving dynamic safety protection and stable empathic output for adolescents and balancing consistency with expert judgment.

[0111] The empathic safety question-answering method provided in this application embodiment obtains the question text entered by the user, inputs the question text into a trained empathic safety question-answering model, so that the empathic safety question-answering model generates empathic safety inference content corresponding to the question text, and outputs empathic safety response text through the empathic safety inference content. The training of the empathic safety question-answering model is divided into two stages. The first stage uses the acquired first dataset to fine-tune the pre-trained base model to obtain an intermediate model. The second stage inputs the acquired second dataset into the intermediate model to obtain the current inference content and the current response text output by the intermediate model. Based on the evaluation model, the inference safety score and inference empathy score corresponding to the current inference content, and the response safety score, response empathy score, and format score corresponding to the current response text are determined, thereby determining the response text based on the reward model. This method employs a combination of type, reasoning safety score, reasoning empathy score, response safety score, response empathy score, and format score to determine a comprehensive reward. The parameters of the intermediate model are then adjusted based on this comprehensive reward to obtain an empathy-safe question-answering model. This model outputs empathy-safe reasoning content and then infers empathy-safe response text, introducing an explicit safety reasoning mechanism. This allows the model to move beyond relying solely on implicit statistical distributions in the output stage, instead performing structured reasoning first. This proactively avoids generating unsafe information, improving the interpretability and safety of the model's decisions. Furthermore, this method trains the empathy-safe question-answering model through two stages: optimization training and reinforcement learning. This enables the model to possess dynamic risk identification and self-checking capabilities, maintaining output safety in multi-turn long dialogues and highly semantically ambiguous questions, significantly reducing the probability of potential unsafe responses. Furthermore, by designing a multi-level reward structure that includes inference safety, inference empathy, response safety, response empathy, and response format compliance during the reinforcement learning phase, this method can guide the model to output responses that are formatted correctly, safe, and empathetic. This significantly improves the model's ability to recognize and respond to user emotions, making the response structure more natural and human, and enhancing its adaptability to specific groups.

[0112] This application's embodiments, by introducing an explicit safety reasoning mechanism during the generation phase and combining it with a multi-dimensional reward optimization strategy, achieve dynamic safety alignment and emotional empathy balance of a large language model in adolescent scenarios. This effectively compensates for the lack of age adaptability in general safety models, providing adolescent users with a differentiated, traceable, and psychologically friendly content security protection system, thereby ensuring their information security and mental health development in intelligent dialogue interaction. Specifically, through a four-stage process of "supervised fine-tuning—reinforcement learning—reward modeling—evaluation and iteration," the safety constraints and risk protection of the empathic safety question-answering model for adolescent dialogue scenarios are achieved. By introducing a "prudent alignment" mechanism during the training phase, explicit safety reasoning is enforced before model generation, and a multi-dimensional reward system is constructed by combining manual annotation and model scoring, enabling the model to simultaneously possess safety, compliance, and empathy in its response content.

[0113] The empathic security question-answering method provided in this application has the following technical effects: 1. Introduce an explicit safe reasoning mechanism to improve the interpretability and safety of model decisions: By embedding a "CoT (Conceptual Chain) + Security Norm Call" mechanism before generating responses in the empathic secure question-answering model, the model no longer relies solely on implicit statistical distributions in the output stage. Instead, it performs structured security reasoning and risk assessment first, enabling it to clearly determine whether the input involves sensitive or illegal content. This proactively avoids unsafe information during the generation stage, effectively overcoming the "black box decision-making" problem of traditional training and achieving transparency and traceability of model behavior. In addition to this reasoning mechanism, an architecture of "explicit rule reasoning + instruction-controlled generation" can also be adopted (e.g., first calling an external rule engine for logical inference, and then feeding the inference results back to the model for generation). Its essence is still "security reasoning before generation." Alternatively, "self-supervised security reflection" or "chain rejection strategy" can be used instead of CoT reasoning, as long as the mechanism performs security judgments before generation and affects the response output. 2. Employing a deliberative alignment framework significantly improves security and robustness in complex scenarios: By explicitly invoking security guidelines and performing step-by-step reasoning before model generation, the empathic security question-answering model possesses dynamic risk identification and self-checking capabilities. Compared to traditional static security filtering methods, this framework can maintain output security throughout multi-turn long dialogues and highly semantically ambiguous questions, significantly reducing the probability of potential jailbreaks and unauthorized responses. In addition to careful alignment, other methods such as "Constitutional AI" (artificial intelligence based on predefined rules), "Rule-Driven RLHF" (rule-driven human feedback reinforcement learning), or "SafeChain" (AI security training mechanism combined with context) can be used, as long as they include a model training mechanism that performs reasoning and output correction based on explicit rules or standard guidelines. Alternatively, security guidelines can be embedded in the model's feature vectors or word embedding space (rather than text input), which logically still maintains the principle of "invoking security guidelines before generation." 3. Construct a multi-dimensional reward function system to achieve a balanced optimization between security and empathy: During the reinforcement learning phase, a multi-level reward structure was designed, which includes format compliance, content security, rationality of thought chain and emotional support. Through weighted approach, the empathic safety question-answering model is guided to retain emotional warmth while ensuring safety and compliance. Compared with a single safety reward model, it can significantly improve the model's ability to identify and soothe the emotions of adolescents, and generate more natural and humanized results. 4. Introduce a dynamic weighting mechanism: This system can adjust reward weights in real time based on user background, dialogue topic, and risk context. This allows the empathic safety question-answering model to automatically enhance safety in high-risk scenarios and strengthen empathic expression in emotionally supportive scenarios, thus significantly improving its adaptability across different scenarios. Simultaneously, dynamic weights reduce ineffective updates during training, improve policy convergence speed, and maintain a stable balance between safety and empathy in the model's output, resulting in more consistent and reliable overall generation quality. Furthermore, dynamic weights can be replaced with any form of "scenario-adaptive weighting mechanism," such as rule-based weight switching (increasing safety weights in high-risk scenarios), emotion intensity prediction weights based on statistical models, or a user profile-driven weight mapping table, as long as the weights can be adjusted in real time according to user characteristics or dialogue context. It can also be implemented using various model architectures, including lightweight classifiers, logistic regression, hierarchical risk identification modules, attention-based weighted models, or even the weight distribution can be directly output by an end-to-end reinforcement learning policy network, as long as the weight determination process is generated by an external module and used to dynamically adjust multi-dimensional rewards. 5. Introduce a hybrid reward and rule constraint during the reinforcement learning phase to improve training convergence stability: In reinforcement learning, the bias of the intermediate model is corrected in real time, avoiding the reward drift and overfitting problems of traditional algorithms, so that the empathic safety question answering model can maintain stable performance improvement during long-term training. In addition, any policy optimization algorithm can be used to replace it, as long as the model policy is iteratively optimized through reward signals to achieve safety behavior learning. Alternatively, "supervised preference fine-tuning" or "reward distillation" can be used to replace the reinforcement learning stage. 6. Construct a standardized and formatted training set to ensure consistency and standardization of the model output structure: By uniformly labeling the <|begin_of_thought|> and <|end_of_thought|> separators in the first dataset, the model can automatically distinguish between the inference process and the visible response, thus ensuring a stable output format, easy parsing, and convenient subsequent system integration. Furthermore, the above separators can be replaced with XML, JSON, YAML, or any specific tags (such as [SAFE_BEGIN]...[SAFE_END]), or implicit tags (such as special tokens or control characters) can be used instead.

[0114] Furthermore, in the methods provided in the embodiments of this application, an architecture of explicit rule reasoning and instruction control generation can be used instead of a deliberate alignment architecture, or a self-supervised security reflection or chained rejection strategy can be used instead of thought chain reasoning, which will not be elaborated here. Moreover, reasoning based on explicit rules or security specifications can also be used, embedding the security specifications into the model's feature vectors or word embedding space, which will not be elaborated here.

[0115] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. For example... Figure 3 As shown, the electronic device 300 includes one or more processors 301 and memory 302.

[0116] The processor 301 may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 300 to perform desired functions.

[0117] The memory 302 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 301 may execute the program instructions to implement the empathic safe question-answering method of any embodiment of this application described above and / or other desired functions. Various contents such as initial extrinsic parameters and thresholds may also be stored in the computer-readable storage medium.

[0118] In one example, the electronic device 300 may further include an input device 303 and an output device 304, these components being interconnected via a bus system and / or other forms of connection mechanisms (not shown). The input device 303 may include, for example, a keyboard, a mouse, etc. The output device 304 may output various information to the outside, including warning messages, braking force, etc. The output device 304 may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.

[0119] Of course, for the sake of simplicity, Figure 3 Only some of the components of the electronic device 300 relevant to this application are shown in this illustration; components such as buses, input / output interfaces, etc., are omitted. In addition, the electronic device 300 may include any other suitable components depending on the specific application.

[0120] In addition to the methods and devices described above, embodiments of this application may also be computer program products, which include computer program instructions that, when executed by a processor, cause the processor to perform the steps of the empathic security question-answering method provided in any embodiment of this application.

[0121] The computer program product can be written in any combination of one or more programming languages ​​to perform the operations of the embodiments of this application. The programming languages ​​include object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0122] Furthermore, embodiments of this application may also be computer-readable storage media storing computer program instructions thereon, which, when executed by a processor, cause the processor to perform the steps of the empathic safe question-and-answer method provided in any embodiment of this application.

[0123] The computer-readable storage medium may be any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof.

[0124] It should be noted that the terminology used in this application is for the purpose of describing specific embodiments only and is not intended to limit the scope of this application. As shown in the specification and claims of this application, unless the context clearly indicates otherwise, words such as "a," "an," "an," and / or "the" do not specifically refer to the singular and may also include the plural. The terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, or apparatus. Without further limitations, an element defined by the phrase "comprising an..." does not exclude the presence of other identical elements in the process, method, or apparatus that includes said element.

[0125] It should also be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation on this application. Unless otherwise expressly specified and limited, the terms "installed," "connected," "linked," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication between two elements. For those skilled in the art, the specific meaning of the above terms in this application can be understood according to the specific circumstances.

[0126] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. The above descriptions are only preferred embodiments of this application. It should be noted that due to the limitations of written expression, while there are objectively infinite specific structures, those skilled in the art can make several improvements, modifications, or changes without departing from the principles of this application, and can also combine the above technical features in an appropriate manner. These improvements, modifications, changes, or combinations, or the direct application of the inventive concept and technical solution to other situations without modification, should all be considered within the scope of protection of this application.

Claims

1. An empathic safety question-and-answer method, characterized in that, include: The system obtains the question text entered by the user, inputs the question text into the trained empathic safety question answering model, so that the empathic safety question answering model generates empathic safety inference content corresponding to the question text, and outputs empathic safety response text through the empathic safety inference content; The training steps of the empathic safety question-answering model include: Obtain the first dataset, and fine-tune the pre-trained base model using the first dataset to obtain the intermediate model; Obtain a second dataset, input the second dataset into the intermediate model, obtain the current inference content and current response text output by the intermediate model, and determine the inference safety score and inference empathy score corresponding to the current inference content, and the response safety score, response empathy score and format score corresponding to the current response text according to the evaluation model; Based on the reward model, the inference safety score, the inference empathy score, the response safety score, the response empathy score, and the format score, a comprehensive reward is determined, and the parameters of the intermediate model are adjusted using the comprehensive reward to obtain the empathy-safe question-answering model.

2. The method according to claim 1, characterized in that, The first dataset includes the question text for each sample, the corresponding empathy safety standard text, the sample inference content, and the sample response text. The pre-trained base model is fine-tuned using this first dataset to obtain an intermediate model, which includes: The first dataset is input into a pre-trained base model so that the pre-trained base model outputs the current inference content and the current response text based on the sample question text, the sample inference content, and the sample response text. Based on the gap between the current inference content output by the pre-trained base model and the empathy safety standard text, as well as the gap between the current response text and the empathy safety standard text, the parameters in the pre-trained base model are adjusted until the training cutoff condition is met, thus obtaining an intermediate model.

3. The method according to claim 2, characterized in that, Obtain the first dataset, including: Obtain scene description information under each scene category, and obtain the empathy safety standard text corresponding to the scene category; The scenario description information is input into the user big model to obtain sample question text, and the sample question text and the empathy safety standard text are input into the consultant big model to obtain the corresponding sample reasoning content and sample response text; The first dataset is constructed based on the empathy safety standard text, sample question text, sample reasoning content, and sample response text corresponding to the description information of each scenario.

4. The method according to claim 3, characterized in that, Obtain scene description information under each scene category, including: For each scene category, obtain a scene description example for that scene category; Based on the scene description example, a scene-generating prompt text is constructed, and the scene-generating prompt text is input into a large language model to obtain scene description information.

5. The method according to claim 1, characterized in that, The evaluation model determines the reasoning safety score and reasoning empathy score corresponding to the current reasoning content, as well as the response safety score, response empathy score, and format score corresponding to the current response text, including: Determine the current scene category corresponding to the current question text input to the intermediate model, and obtain the empathy safety standard text of the current scene category; An evaluation prompt text is generated based on the aforementioned empathy safety standard text; The evaluation prompt text, the current inference content, and the current response text are input into the evaluation model to obtain the inference safety score and inference empathy score corresponding to the current inference content, and the response safety score, response empathy score, and format score corresponding to the current response text.

6. The method according to claim 1, characterized in that, Based on the reward model, the inference safety score, the inference empathy score, the response safety score, the response empathy score, and the format score, a comprehensive reward is determined, including: Determine the reward fusion weights, wherein the reward fusion weights include reasoning safety weight, reasoning empathy weight, response safety weight, response empathy weight, and format weight; The reward fusion weight, the inference security score, the inference empathy score, the response security score, the response empathy score, and the format score are input into the reward model so that the reward model can weight each score using the reward fusion weight to obtain a comprehensive reward.

7. The method according to claim 6, characterized in that, Determine the reward fusion weights, including: The current scenario classification, user background information, dialogue risk level, and dialogue emotion intensity are determined based on the current question text input to the intermediate model. The reward fusion weight is determined based on the current scenario classification, the user background information, the dialogue risk level, and the dialogue emotional intensity.

8. The method according to claim 1, characterized in that, The parameters of the intermediate model are adjusted using the comprehensive reward, including: The estimated value of the current inference content and the current response text output by the intermediate model is determined through the reward model. The parameters of the intermediate model are adjusted based on the gap between the overall reward and the estimated value.

9. The method according to claim 1, characterized in that, After adjusting the parameters of the intermediate model using the comprehensive reward to obtain the empathic safety question-answering model, the method further includes: Construct security assessment prompt text and empathy assessment prompt text; The security assessment prompt text and the empathy assessment prompt text are input into the security empathy assessment model so that the security empathy assessment model can determine the model security score and the model empathy score corresponding to the empathy security question-and-answer model. Based on the model's security score and empathy score, determine whether to return to training the empathy-safe question-answering model.

10. An electronic device, characterized in that, The electronic device includes: Processor and memory; The processor executes the steps of the empathic safe question-and-answer method as described in any one of claims 1 to 9 by invoking programs or instructions stored in the memory.