Persuasive purpose-oriented multi-round dialogue dynamic planning method, device and equipment and medium
By assessing the information sufficiency and security risks of dialogue messages and generating personalized responses according to preset priorities, the lack of goal orientation and security mechanisms in existing technologies is solved, thus achieving effective guidance and security of the dialogue system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-27
- Publication Date
- 2026-04-07
AI Technical Summary
The lack of clear goal orientation and structured intervention framework in existing technologies makes it difficult for dialogue systems to guide users to produce quantifiable behavioral changes, and the lack of security mechanisms makes it impossible to effectively evaluate the intervention effect.
By assessing the information sufficiency and security risks of dialogue messages, conditions are judged according to a preset priority order, and personalized responses are generated, including crisis management, exploratory responses, and persuasive responses. Pre-set strategies and resource databases are used to ensure the relevance and security of the responses.
It achieves goal-oriented guidance in the dialogue system, generates closely related personalized responses, ensures user safety, and improves the quantifiability and reliability of intervention effects.
Smart Images

Figure CN121809646A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device, and medium for multi-turn dialogue dynamic programming aimed at persuasion purposes. Background Technology
[0002] Multi-turn dialogue technology is evolving from simple question-and-answer sessions into intelligent agents capable of deep interaction and long-term companionship. Its core value lies in its ability to simulate human communication through continuous, context-sensitive exchanges, thereby building trust, guiding thought, and facilitating change. This technological characteristic elevates it beyond the scope of traditional tools, making it an ideal vehicle for digital interventions in complex areas such as mental health, behavioral change, and personal growth. A successful interventional dialogue system not only needs to understand language but also manage the dialogue process, maintain goal orientation, and remain sensitive to the user's emotional state and changes. This places extremely high demands on its architectural design and interaction strategies.
[0003] Currently, the mainstream intervention dialogue approach is end-to-end generative dialogue based on large language models. This approach leverages the powerful contextual understanding and natural language generation capabilities of large language models to simulate an empathetic and knowledgeable virtual partner. During interaction, the system typically does not pre-define a fixed dialogue flow or state, but rather dynamically generates a coherent, fluent, and supportive response based on each input from the user. Its core advantages lie in its high flexibility and naturalness of dialogue, enabling it to handle open-ended topics, provide immediate emotional comfort and broad advice, and attempt to alleviate users' negative emotions through high-quality "companionship" dialogue.
[0004] However, this intervention dialogue approach, which relies entirely on end-to-end generation, lacks goal orientation and measurable progress. The dialogue easily devolves into aimless chatter or emotional venting. While it may provide short-term comfort, it struggles to guide users towards concrete, observable behavioral changes, making the intervention's effectiveness difficult to quantify and evaluate. Furthermore, the uncontrollable and inconsistent application of its strategies means that while the model-generated content may appear reasonable, its underlying intervention logic remains a black box. It cannot guarantee that its recommendations consistently adhere to evidence-based clinical principles, potentially providing contradictory or inappropriate guidance in different rounds, lacking professionalism and reliability. Additionally, due to the ambiguity and passivity of safety boundaries, when faced with users' crisis signals or potential digital manipulation risks, the model may only generate formatted comforting statements, lacking a clear and enforceable crisis triage and protection mechanism, thus failing to fulfill the crucial responsibility of prioritizing safety. Summary of the Invention
[0005] This invention provides a method, apparatus, device, and medium for dynamic planning of multi-turn dialogues for persuasive purposes, which addresses the shortcomings of existing technologies that lack clear goal orientation and structured intervention frameworks, making it difficult for dialogue systems to guide users to produce quantifiable behavioral changes and resulting in the inability to assess intervention effects. It enables flexible switching of response decisions and ensures the generation of highly targeted and empathetic responses.
[0006] This invention provides a multi-turn dialogue dynamic programming method for persuasive purposes, comprising: assessing information sufficiency and security risks based on received dialogue messages to obtain predicted information sufficiency and predicted security risks; evaluating whether the judgment conditions of the corresponding priorities are met in sequence according to the predicted information sufficiency, predicted security risks, and the current exploration and persuasion rounds, in a preset priority order; determining the response type of the corresponding priority based on the satisfaction of the judgment conditions, and generating a corresponding personalized response based on the dialogue messages and context messages.
[0007] According to the present invention, a multi-turn dialogue dynamic programming method for persuasive purposes determines the corresponding priority response type based on the fulfillment of judgment conditions, and generates a corresponding personalized response based on the dialogue messages and context messages. The method includes: determining that the judgment condition is met when the predicted security risk reaches a preset security risk threshold, and determining the corresponding response type as a crisis management type; determining the security risk level based on the crisis management type, the dialogue messages, and context messages; and generating a crisis response by invoking security resources from a security resource database based on the security risk level and a corresponding preset crisis management strategy. The preset crisis management strategy is constructed prior to the intervention target based on the corresponding security risk level, and the security resource database is constructed prior to the availability of external support information and professional guidance content for reducing user psychological risk.
[0008] According to the present invention, a multi-turn dialogue dynamic programming method for persuasive purposes is provided. Based on satisfying judgment conditions, the method determines the response type of corresponding priority. Based on the dialogue messages and context messages, it generates a corresponding personalized response. Alternatively, based on a preset priority order, when all priority judgment conditions are not satisfied, the method further includes: determining that the satisfied judgment condition is that the number of persuasion turns is greater than or equal to the maximum number of persuasion turns or that the forced exploration flag is forced exploration; determining the corresponding response type as an exploration response type; based on the exploration response type, determining the dialogue context according to the dialogue messages and context messages; evaluating user behavior based on the dialogue messages and context messages, and making a fusion decision based on the dialogue context to generate an exploration response.
[0009] According to the present invention, a multi-turn dialogue dynamic programming method for persuasive purposes evaluates user behavior based on dialogue messages and contextual messages, including: using the antecedent-action-consequence-effect analysis method to capture triggering situations in the dialogue, locate user response behaviors, and analyze the consequences of the behaviors, generating an ABC chain; using the DARN-CAT classification method (desire, ability, reason, need, commitment, trigger, action) to identify user change signals and classify them according to change reasons and change language, obtaining change motivation; and using the decision balancing method to explore the advantages of change, disadvantages of change, advantages of maintaining the status quo, and disadvantages of maintaining the status quo when user hesitation language is detected, generating a decision balancing matrix to represent the overall decision-making process.
[0010] According to the present invention, a multi-turn dialogue dynamic programming method for persuasive purposes is provided. Based on satisfying judgment conditions, a response type with corresponding priority is determined. Based on dialogue messages and context messages, a corresponding personalized persuasive response is generated. The method further includes: determining a persuasive response type when the judgment condition is satisfied (e.g., the number of exploration rounds is greater than or equal to the maximum number of consecutive persuasion rounds), or when the judgment condition is satisfied (e.g., the predicted information sufficiency is greater than or equal to an information sufficiency threshold and the number of dialogue rounds is greater than or equal to the minimum total number of rounds required to trigger persuasion); based on the persuasive response type, determining whether the predicted information sufficiency assessed by the local route reaches a preset mode switching threshold; if the information sufficiency assessed by the local route reaches the preset mode switching threshold, determining the dialogue context and evaluating user behavior based on dialogue messages and context messages to generate an exploration response; if the predicted information sufficiency does not reach the preset mode switching threshold, identifying user intent based on dialogue messages and context messages; and generating a corresponding persuasive response based on user intent, dialogue messages, and context messages.
[0011] According to the present invention, a multi-turn dialogue dynamic programming method for persuasive purposes generates a corresponding personalized response based on user intent, dialogue messages, and contextual messages. The method includes: when the user intent is security concern, determining the security risk level based on the dialogue messages and contextual messages, and generating a crisis response by invoking security resources from a security resource database in conjunction with a corresponding preset crisis management strategy; wherein the preset crisis management strategy is constructed prior to intervention objectives based on the corresponding security risk level, and the security resource database is constructed prior to external support information and professional guidance content available for use in reducing user psychological risk; when the user intent is identification type, identifying the user state based on the dialogue messages and contextual messages, the user state including user behavior, motivation, and emotional state; generating an intervention strategy based on the identified user state, and executing the matching intervention strategy; wherein the intervention strategy is used to characterize the intervention direction and intervention objective configured to respond to the user's current needs.
[0012] According to the present invention, a multi-turn dialogue dynamic programming method for persuasive purposes generates and executes an intervention strategy based on the identified user state, including: determining the immediate response type and selecting components based on the identified user state; wherein, the components are pre-built and encapsulated functional modules of specific psychological techniques or communication logic based on different user states; extracting keywords or key phrases as evidence based on dialogue messages and context messages; retrieving the most matching script from a pre-set script library based on the selected components and evidence; wherein, the pre-set script library is constructed based on scripts configured with different components and evidence, and the scripts are used to represent pre-encapsulated intervention program templates for specific psychological goals; generating an intervention strategy based on the matching components, evidence, and the retrieved most matching script; and reading the user's historical interaction information from a pre-set memory bank and identifying the user's cognition based on the intervention strategy. Emotional characteristics; among them, user cognitive emotional characteristics are used to characterize the user's cognitive and emotional tendencies in the process of information processing and communication interaction; based on user cognitive emotional characteristics, the intervention strategy is adapted and adjusted, and the most matching communication tone is selected as a style tag and added to the intervention strategy to obtain a personalized intervention strategy; based on the personalized intervention strategy, combined with the immediate contextual characteristics extracted from the dialogue messages, contextual adaptation adjustment is performed to obtain a contextualized intervention strategy; among them, immediate contextual characteristics are used to characterize the specific emotional state, dialogue atmosphere or implicit potential needs shown by the user in the current dialogue round; based on the contextualized intervention strategy, corresponding skills are selected and combined from the dialogue micro-skill library, and transformed into natural, coherent and compliant language text to obtain the dialogue response text and return it; among them, the dialogue micro-skill library contains a variety of verified communication skills for building effective dialogues.
[0013] The present invention also provides a multi-turn dialogue dynamic planning device for persuasive purposes, comprising: an evaluation module, which evaluates information sufficiency and security risks based on received dialogue messages to obtain predicted information sufficiency and predicted security risks; a priority judgment module, which evaluates whether the judgment conditions of the corresponding priorities are met in sequence according to the predicted information sufficiency, predicted security risks, and the current exploration and persuasion rounds; and a response module, which determines the response type of the corresponding priority based on the satisfaction of the judgment conditions, and generates a corresponding personalized response based on the dialogue messages and context messages.
[0014] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements a multi-turn dialogue dynamic programming method for persuasive purposes as described above.
[0015] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a multi-turn dialogue dynamic programming method for persuasive purposes as described above.
[0016] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements a multi-turn dialogue dynamic programming method for persuasive purposes as described above.
[0017] The present invention provides a multi-turn dialogue dynamic planning method, apparatus, device, and medium for persuasive purposes. By assessing information sufficiency and security risks based on received dialogue messages, it proactively identifies potential risks during the dialogue and ensures that sufficient information is available to understand the user's current state and true needs before guiding user action. This makes subsequent guidance no longer blind but directional guidance based on a thorough understanding, avoiding hasty or erroneous responses when information is insufficient. Furthermore, it provides a clear decision-making mechanism by sequentially evaluating whether the corresponding priority judgment conditions are met according to a preset priority order. This allows the system to flexibly switch between different response decisions based on real-time evaluation results, avoiding decision confusion or logical conflicts. Based on meeting the judgment conditions, it determines the corresponding priority response type, thereby generating highly targeted and empathetic responses that make users feel understood and respected. It ensures that the generated responses are closely connected to the current dialogue context, making the dialogue smooth and natural. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0019] Figure 1 This is one of the flowcharts of the multi-turn dialogue dynamic programming method for persuasive purposes provided by the present invention; Figure 2 This is the second flowchart of the multi-turn dialogue dynamic programming method for persuasive purposes provided by the present invention; Figure 3 This is a schematic diagram of the structure of the multi-turn dialogue dynamic programming device for persuasive purposes provided by the present invention; Figure 4 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0021] Figure 1 This is a flowchart illustrating the multi-turn dialogue dynamic programming method for persuasive purposes provided by the present invention, as shown below. Figure 1 As shown, the method includes: S11, Based on the received dialogue messages, assess user tolerance, information sufficiency, and security risks to obtain predicted information sufficiency and predicted security risks; S12, based on the sufficiency of predicted information, predicted safety risks, and the current exploration and persuasion rounds, evaluate in sequence whether the judgment conditions of the corresponding priorities are met according to the preset priority order; S13. Based on the judgment conditions, determine the corresponding priority response type, and generate a corresponding personalized response according to the dialogue message and the context message.
[0022] It should be noted that the step numbers "S1N" in this manual do not represent the order of steps in a multi-turn dialogue dynamic programming method for persuasive purposes. The following will explain this in detail. Figure 2 This invention describes a multi-turn dialogue dynamic programming method for persuasive purposes.
[0023] Step S11: Based on the received dialogue messages, assess the information sufficiency and security risks to obtain the predicted information sufficiency and predicted security risks.
[0024] In this embodiment, the information sufficiency and security risks are assessed based on the received dialogue messages, including: performing Boolean detection on the received dialogue messages to obtain the predicted security risk SR, which is used to characterize whether there are specific self-harm or other-harm plans or behaviors that are seriously out of touch with reality in the dialogue messages; and retrieving the dialogue messages in the historical target rounds and the historical decisions and feedback corresponding to each dialogue message from the memory bank based on the received dialogue messages, and scoring them according to the preset scoring rules to obtain the predicted information sufficiency IS.
[0025] It should be added that the predicted security risks include a first identifier and a second identifier. The first identifier is used to indicate that there is a specific plan for self-harm or harm to others, or behavior that is seriously out of touch with reality in the dialogue message. When the first identifier appears, it is necessary to forcibly switch to the exploration mode. The second identifier is used to indicate that there is no specific plan for self-harm or harm to others, or behavior that is seriously out of touch with reality in the dialogue message. By predicting security risks, it is easy to directly determine the crisis management type when the first identifier is identified, and to call the crisis agent to execute the crisis management strategy. The specific processing can be referred to below, and will not be repeated here.
[0026] In addition, the specific historical decisions and feedback can be determined based on the information stored in the memory bank corresponding to the historical target rounds. For example, it can be ABC chains, background information, strategies that have been tried, user feedback, If-Then rules and delay rules, etc. ABC chains are used to represent trigger-behavior-short-term consequences-long-term costs. Background information includes relationship type, attachment tendency, etc., without further limitations here.
[0027] Accordingly, the preset scoring rules can be set according to the actual historical decisions, feedback and prior experience involved. For example, the more complete the ABC, the more specific the trigger, and the clearer the behavior cycle, the higher the IS. For another example, the more "to be supplemented / to be understood / to be formulated / to be evaluated" in the case text, the lower the IS. For yet another example, if it is just a fragmented feeling, without a behavior pattern or context, the IS will remain at a low level. No further restrictions are made here.
[0028] In an optional embodiment, after obtaining the sufficiency of the predicted information and the predicted security risk, the method further includes: identifying the dialogue message and determining whether it is a suggestion request; if it is determined to be a suggestion request, based on the second identifier of the predicted security risk and the sufficiency of the predicted information being greater than the first sufficiency threshold, determining it to be a persuasive response type, and executing the corresponding persuasive strategy, as can be referred to below, and will not be repeated here.
[0029] It is worth noting that before the t-th round of dialogue, the scheduler needs to maintain the state of the t-1 rounds of dialogue. The state specifically includes the number of consecutive exploration rounds (ER), consecutive persuasion rounds (PR), the number of conversation rounds (OR), user tolerance (UT), information sufficiency (IS), safety risk (SR), information sufficiency threshold (τ), the maximum allowed number of exploration rounds (ERMax), the minimum total number of rounds required to trigger persuasion (ORMin), the maximum allowed number of consecutive persuasion rounds (PRMax), the forced exploration flag (force), and a simple time index (t). ER can be determined based on the number of times the Explorer agent is invoked in a round of exploration, preventing the exploration phase from being extended indefinitely. PR can be determined based on the number of times the Persuader agent is invoked in the corresponding round of persuasion, preventing multiple consecutive rounds of suggestion output. OR is used to control the balance between "giving suggestions too early" and "delaying implementation for too long." UT ∈ [0, 1] is a long-term "comfort" estimate of the user's experience with structured guidance in the current session. IS ∈ [0, 1] 1] indicates whether the understanding of this case is sufficient to support a responsible recommendation; SR∈ {false, true} determines whether to short-circuit to the Crisis agent; τ controls "when to switch from exploration to recommendation"; force is pulled up by the Persuader's local router.
[0030] In an optional embodiment, the method further includes: assessing user tolerance based on the received dialogue messages and context messages to obtain a predicted user tolerance. It should be noted that the context messages can be selected from previous dialogue messages based on actual design requirements and prior experience as the context messages for the current dialogue.
[0031] Specifically, the predicted user tolerance is obtained by: analyzing emotional words, punctuation marks, and tone words in the dialogue messages to determine whether target emotional words exist, and determining whether the length of the user's reply message is less than a preset length threshold. The user then scores the message according to preset scoring rules and conducts a comprehensive evaluation to obtain the immediate tolerance. Based on the immediate tolerance and combined with the user tolerance from the previous round, the user tolerance at the current moment is updated to obtain the predicted user tolerance.
[0032] It should be noted that the target emotional language can be designed based on actual design needs, such as at least one of the following: imperative or irritating language, insulting or aggressive language, strings of emotionally suggestive punctuation marks, rejection language, and expressions of acceptance or cooperation. Strings of emotionally suggestive punctuation marks, such as question marks and exclamation marks, are not further limited here. Additionally, preset scoring rules can be designed according to actual design needs to convert the corresponding evaluation objects into scores, thereby facilitating the comprehensive evaluation to determine immediate tolerance. The comprehensive evaluation can be a combination of explicit linear evaluations, represented as follows: UTI t = clip(0.5-Short-Imper-Punc-Abuse-Refus+Accep,0,1) Among them, UTI t This indicates the immediate tolerance level for the current round. Short indicates the rating for short user responses, Imper indicates the rating for imperative tone, Punc indicates the rating for punctuation misuse, Abuse indicates the rating for content misuse, Rebus indicates the rating for refusing to answer, Accep indicates the rating for acceptability, and clip indicates the clipping function used to limit the final score to the range [0,1].
[0033] Furthermore, after obtaining the immediate tolerance level, it is smoothed using the user tolerance level from the previous round to update the current user tolerance level. When updating the current user tolerance level, an exponential moving average can be used, that is, assigning a preset weight to the immediate tolerance level and the user tolerance level from the previous round, and then summing them up. The corresponding weights can be set according to actual design needs and prior experience. For example, the immediate tolerance level is given a weight of 0.55 to ensure that it can adapt to the real-time changes in user emotions, and the user tolerance level from the previous round is given a weight of 0.45 to represent the system's trust in the previously accumulated tolerance levels. Thus, by weighting and summing, the change in the updated user tolerance level is relatively smooth. It is a long-range memory of "whether users are willing to be structured and guided" in recent rounds. It does not directly determine whether to choose Explore or Persuade, but rather adjusts the threshold later to make the same strategy perform differently for "more patient users" and "users who are easily pressured".
[0034] In an optional embodiment, after obtaining the predicted user tolerance, the method further includes: obtaining a threshold for the previous round of dialogue; wherein the threshold for the previous round of dialogue includes an information sufficiency threshold, a maximum number of consecutive persuasion rounds, a maximum number of persuasion rounds, a minimum total number of rounds required to trigger persuasion, and a forced exploration flag; and adaptively adjusting the threshold for the previous round of dialogue to obtain an updated threshold.
[0035] Specifically, the threshold of the previous round of dialogue is adaptively adjusted to obtain the updated threshold, including: preprocessing the threshold of the previous round of dialogue; adjusting the information sufficiency threshold, the maximum number of consecutive persuasion rounds, and the minimum total number of rounds required to trigger persuasion in the preprocessed threshold of the previous round of dialogue within a time window; coupling the adjusted information sufficiency threshold, the maximum number of consecutive persuasion rounds, and the maximum number of persuasion rounds based on the predicted user tolerance; and applying boundary constraints to all processed thresholds.
[0036] When the user actively accepts, it indicates that the current abstraction is usable, and the "information and number of rounds required before suggesting again" can be appropriately reduced to give the system more motivation. When the user explicitly refuses, it is more like a negative persuasion attribution: "This step is unwelcome, forcing the next round to return to exploration, and tightening the conditions for when to persuade again." Accordingly, the threshold of the previous round of dialogue is preprocessed, including: based on the effect of the previous round of persuasion, if the user actively accepts, the information sufficiency threshold is reduced by the first value, and the minimum total number of rounds required to trigger persuasion is updated using the difference between the minimum total number of rounds required to trigger persuasion and the second value, and the maximum value among the third value; if the user explicitly refuses, the forced exploration flag is updated to the first flag, the information sufficiency threshold is increased by the fourth value, the maximum consecutive persuasion rounds are updated using the sum of the maximum number of consecutive persuasion rounds and the fifth value, and the minimum number of consecutive persuasion rounds is updated using the sum of the minimum total number of rounds required to trigger persuasion and the seventh value, and the maximum value among the eighth value.
[0037] It should be added that the first to eighth values can be set according to actual design requirements and prior experience. For example, the first and fourth values can both be 0.05, the difference between the second, third, fifth, and seventh values can all be 1, the sixth value can be 6, and the eighth value can be 4. No further restrictions are imposed here. In addition, after preprocessing, the final persuasion result `last_persuaded_result` is reset to `none` to avoid the same round of feedback being reused.
[0038] In addition, when adjusting the time window for the information sufficiency threshold, the maximum number of consecutive persuasion rounds, and the minimum total number of rounds required to trigger persuasion in the preprocessed thresholds, a very coarse Just-in-Time Adaptive Interventions rhythm control is required. Specifically, in the early stage (t < 12), the preference is to "listen more first," that is, increase the suggestion threshold, increase the allowed number of exploration rounds, and increase the minimum number of rounds; in the later stage (t ≥ 16), the conversation is long enough, and to avoid "chatting without reaching a conclusion," the exploration space is appropriately compressed, and the system is encouraged to help users form a specific decision action, thereby avoiding two extremes: one is to give suggestions on the approach too early, and the other is to always stay at "just listening without helping to form a decision."
[0039] Specifically, in the early stages, the information sufficiency threshold can be increased by a first preset value, the maximum number of consecutive persuasion rounds can be increased by a second preset value, and the minimum total number of rounds required to trigger persuasion can be increased by a third preset value. The first, second, and third preset values can be set according to actual design needs or prior experience. For example, the first preset value can be 0.05, and the second and third preset values can be 1. No further limitations are made here.
[0040] Additionally, in the later stages, the information sufficiency threshold can be lowered by the fourth preset value. The maximum number of consecutive persuasion rounds is updated based on the difference between the maximum number of consecutive persuasion rounds and the fifth preset value, as well as the maximum value among the sixth preset value. The minimum number of rounds required to trigger persuasion is updated based on the difference between the minimum total number of rounds required to trigger persuasion and the seventh preset value, as well as the maximum value among the eighth preset value. The fourth, fifth, sixth, seventh, and eighth preset values can be set according to actual design requirements or prior experience. For example, the fourth preset value can be 0.05, the fifth, seventh, and eighth preset values can be 1, and the sixth preset value can be 2. No further limitations are imposed here.
[0041] Furthermore, when the adjusted information sufficiency threshold, the maximum number of consecutive persuasion rounds, and the maximum number of persuasion rounds are coupled and adjusted based on the predicted user tolerance, the lower the predicted user tolerance, the more easily users become agitated, while the higher the user tolerance, the more easily users accept structured coaching.
[0042] Specifically, when the predicted user tolerance is less than or equal to the first preset threshold, the maximum number of consecutive persuasion rounds is updated according to the first preset value, the maximum number of persuasion rounds is updated according to the second preset value, and the information sufficiency threshold is increased according to the third preset value; when the predicted user tolerance is greater than or equal to the second preset threshold, the maximum number of consecutive persuasion rounds is updated according to the fourth preset value, and the information sufficiency threshold is updated according to the fifth preset value; wherein, the third preset threshold is greater than the first preset threshold.
[0043] It should be added that the first, second, third, fourth, and fifth preset thresholds can be configured according to actual design needs or prior experience. For example, the first preset threshold can be 0.2, the second preset threshold can be 0.6, and the first to fifth preset values can be 2, 1, 0.05, 6, and 0.10 respectively. No further restrictions are imposed here. When the tolerance is low, the maximum number of consecutive explorations is limited to 2 rounds to avoid asking too many questions on the same topic for too long. The maximum number of consecutive persuasions is limited to 1 round to prevent multiple rounds of continuous "suggestion output". The information sufficiency threshold is also slightly lowered to more decisively conclude when the information is sufficient, reducing the user's decision-making burden. When the tolerance is high, the fourth preset value is set to 6 to give the Explorer a larger "exploration budget" to thoroughly structure the case and at the same time raise the information threshold to avoid giving suggestions too frequently when the understanding is not sufficient.
[0044] Finally, boundary constraints are applied to all processed thresholds, including: loading the safety boundaries of all thresholds and checking whether the corresponding processed threshold is within the corresponding safety boundary. If it exceeds the safety boundary, it is forcibly clamped to the nearest boundary value to facilitate researchers and safety audits. The table shows properties such as "the earliest number of rounds in which suggestions may appear", "the longest number of rounds of exploration", and "no 10 consecutive persuasions will occur".
[0045] Furthermore, the safety boundary for the information sufficiency threshold can be [0.45, 0.75], the safety boundary for the maximum number of consecutive persuasion rounds can be [2, 6], the safety boundary for the minimum total number of rounds required to trigger persuasion can be [1, 4], and the safety boundary for the maximum number of persuasion rounds can be [1, 3]. These can be set according to actual design requirements or prior experience, and are not further limited here.
[0046] Step S12: Based on the sufficiency of predicted information, predicted safety risks, and the current exploration and persuasion rounds, evaluate in sequence whether the judgment conditions of the corresponding priorities are met according to the preset priority order.
[0047] It should be added that the judgment conditions for each priority include at least one of the following: the predicted security risk reaches the preset security risk threshold, the number of persuasion rounds is greater than or equal to the maximum number of persuasion rounds or the request for forced exploration, the number of exploration rounds is greater than or equal to the maximum number of consecutive persuasion rounds, and the predicted information sufficiency is greater than or equal to the information sufficiency threshold and the number of conversation rounds is greater than or equal to the minimum total number of rounds required to trigger persuasion.
[0048] For example, refer to Figure 2 The preset priority order is from first to last, judging in the following order: the predicted security risk reaches the preset security risk threshold, the number of persuasion rounds is greater than or equal to the maximum number of persuasion rounds or a request for forced exploration is made, the number of exploration rounds is greater than or equal to the maximum number of consecutive persuasion rounds, and the predicted information sufficiency is greater than or equal to the information sufficiency threshold and the number of conversation rounds is greater than or equal to the minimum total number of rounds required to trigger persuasion. Accordingly, if the judgment condition of the corresponding priority is met, the corresponding response strategy is executed according to the corresponding response type; otherwise, the judgment condition of the next priority is evaluated, until the corresponding response strategy is executed or the judgment conditions of all priorities are evaluated. The specific process of executing the corresponding response strategy according to the corresponding response type can be referred to below, and will not be repeated here.
[0049] Step S13: Based on the judgment conditions, determine the response type with the corresponding priority, and generate the corresponding personalized response according to the dialogue message and the context message.
[0050] In this embodiment, based on the fulfillment of judgment conditions, the corresponding priority response type is determined. Based on the dialogue message and context message, a corresponding personalized response is generated. This includes: determining that the judgment condition is met when the predicted security risk reaches a preset security risk threshold, thus determining the corresponding response type as a crisis management type; based on the crisis management type, the security risk level is determined according to the dialogue message and context message; based on the security risk level and the corresponding preset crisis management strategy, security resources from the security resource database are invoked to generate a crisis response. The preset crisis management strategy is constructed based on the intervention target of the corresponding security risk level, and the security resource database is constructed based on available external support information and professional guidance content used to reduce user psychological risk. It should be noted that the scheduler determines the response type and invokes the corresponding intelligent agent. When the scheduler determines the response type to be a crisis management type, it invokes the corresponding crisis intelligent agent and switches to crisis management mode. The crisis intelligent agent executes the corresponding crisis management strategy to generate a crisis response, as described above, and will not be repeated here.
[0051] Specifically, based on the dialogue messages and contextual information, the security risk level is determined, including: based on the dialogue messages and contextual information, determining the user's timeliness, specificity, emotional state, and social support, and performing a weighted sum to determine the current crisis, and determining the corresponding security risk level based on a preset crisis level classification threshold; whereby, the timeliness of the expression includes immediacy and long-term nature, specificity is used to characterize the completeness of the user's plan, and social support is used to characterize whether the user is isolated and helpless or has the support of family and friends.
[0052] In addition, the preset crisis management strategies are configured in advance according to the corresponding security risk level. For example, the security risk level includes high-risk, medium-risk and low-risk levels. The crisis management strategies corresponding to the high-risk level can be to delay time, establish contact and provide resources. The crisis management strategies corresponding to the medium-risk level can be to listen and empathize, explore the problem and reduce pain. The crisis management strategies corresponding to the low-risk level can be to provide emotional support, normalize feelings and encourage seeking help. The specific strategies can be set according to actual design needs or prior experience, and no further restrictions are made here.
[0053] In addition, the security resource database includes geographic information, institutional information, and knowledge base information. Geographic information includes addresses and phone numbers of the nearest suicide intervention hotlines, hospital emergency rooms, and mental health assistance centers, which can be found based on the user's IP address or previously authorized location information. Institutional information includes nationwide mental health assistance hotlines and online psychological counseling platforms. Knowledge base information includes psychological tips on how to deal with extreme emotions and relaxation exercises, etc., without further limitations here.
[0054] In addition, in this embodiment, based on the satisfaction of the judgment conditions, the corresponding priority response type is determined, and a corresponding personalized response is generated according to the dialogue message and in combination with the context message. Alternatively, based on the preset priority order, if the judgment conditions of all priorities are not satisfied, the method further includes: determining that the satisfied judgment condition is that the number of persuasion rounds is greater than or equal to the maximum number of persuasion rounds or the forced exploration identifier is forced exploration, and determining the corresponding response type as an exploration response type; based on the exploration response type, the dialogue context is determined according to the dialogue message and in combination with the context message; based on the dialogue message and the context message, the user behavior is evaluated, and a fusion decision is made in combination with the dialogue context to generate an exploration response.
[0055] It should be noted that forced exploration is initiated by persuading the agent to actively send a request based on insufficient information or the need to return to exploration. This involves switching the forced exploration flag to the first flag indicating forced exploration. Additionally, when the scheduler determines the response type to be an exploration response, it invokes the corresponding exploration agent and switches to exploration mode. The exploration agent then executes the corresponding exploration response strategy to generate an exploration response, as described above, and will not be repeated here.
[0056] Furthermore, based on dialogue and contextual messages, user behavior is evaluated, including: using the antecedent-action-consequence (AAC) analysis method to capture triggering situations in the dialogue, locate user reaction behaviors, and analyze the consequences of these behaviors, generating an ABC chain; using the Desire, Capability, Reason, Need, Commitment, Trigger, Action (DARN-CAT) classification method to identify user change signals and categorize them by change reasons and change language, thus obtaining change motivations; and using the decision balancing method to explore the advantages of change, disadvantages of change, advantages of maintaining the status quo, and disadvantages of maintaining the status quo when user hesitation language is detected, generating a decision balancing matrix to represent the overall decision-making process.
[0057] Furthermore, based on the fulfillment of judgment conditions, the corresponding priority response type is determined. Based on the dialogue messages and contextual messages, a corresponding personalized persuasion response is generated. This includes: determining the corresponding response type as a persuasion response type when the judgment condition is met that the number of exploration rounds is greater than or equal to the maximum number of consecutive persuasion rounds, or when the judgment condition is met that the predicted information sufficiency is greater than or equal to the information sufficiency threshold and the number of conversation rounds is greater than or equal to the minimum total number of rounds required to trigger persuasion; based on the persuasion response type, determining whether the information sufficiency assessed by the local route reaches the preset mode switching threshold; if the information sufficiency assessed by the local route reaches the preset mode switching threshold, determining the dialogue context and evaluating user behavior based on the dialogue messages and contextual messages to generate an exploration response; if the predicted information sufficiency does not reach the preset mode switching threshold, identifying user intent based on the dialogue messages and contextual messages; and generating a corresponding persuasion response based on user intent, dialogue messages, and contextual messages.
[0058] It's important to note that by introducing exploration and persuasion rounds as criteria, the system prevents premature persuasion and avoids endless exploration that delays getting to the point. This ensures the intervention pace matches the user's psychological state, improving user experience and intervention effectiveness. Furthermore, when the scheduler determines the response type to be a persuasion response, it invokes the corresponding persuasion agent and switches to persuasion mode. When executing the corresponding persuasion response strategy, the persuasion agent's local routing first assesses the sufficiency of information in the current round. If information is insufficient, it sends a switch to exploration signal to the scheduler to switch to exploration mode.
[0059] It should be noted that local routing considers the dialogue type when assessing information sufficiency. Specifically: when the dialogue type is self-exploration or exploration required, the first information value is used as the assessed information sufficiency; when the dialogue type is a new event, the second information value is used as the assessed information sufficiency; for other cases, the predicted information sufficiency is used as the assessed information sufficiency. Additionally, the counter for consecutive persuasion rounds needs to be cleared.
[0060] Specifically, based on user intent, and combining dialogue messages and contextual messages, a corresponding persuasive response is generated. This includes: when the user intent is security concern, determining the security risk level based on dialogue messages and contextual messages, and, in conjunction with a corresponding preset crisis management strategy, invoking security resources from the security resource database to generate a crisis response; wherein, the preset crisis management strategy is constructed in advance based on the intervention objectives of the corresponding security risk level, and the security resource database is constructed in advance based on available external support information and professional guidance content used to reduce the user's psychological risk; when the user intent is identification type, identifying the user state based on dialogue messages and contextual messages, the user state includes user behavior, motivation, and emotional state; generating an intervention strategy based on the identified user state, and executing the matching intervention strategy; wherein, the intervention strategy is used to characterize the intervention direction and intervention objectives configured to respond to the user's current needs.
[0061] Further, based on the identified user state, an intervention strategy is generated and executed, including: determining the immediate response type and selecting components based on the identified user state; wherein, the components are pre-built and encapsulated functional modules of specific psychological techniques or communication logic based on different user states; extracting keywords or key phrases as evidence based on dialogue messages and context messages; retrieving the most matching script from a pre-set script library based on the selected components and evidence; wherein, the pre-set script library is constructed based on scripts configured with different components and evidence, and the scripts are used to represent pre-encapsulated intervention program templates for specific psychological goals; generating an intervention strategy based on the matching components, evidence, and the retrieved most matching script; and, based on the intervention strategy, reading the user's historical interaction information from a pre-set memory bank and identifying the user's cognitive and emotional characteristics; wherein, the user's cognitive... Emotional features are used to characterize users' cognitive and emotional tendencies during information processing and communication interactions. Based on these characteristics, intervention strategies are adaptively adjusted, and the most matching communication tone is selected as a style marker and added to the intervention strategy to obtain a personalized intervention strategy. Based on the personalized intervention strategy, contextualized intervention strategies are obtained by combining previously extracted immediate situational features from dialogue messages with contextualized intervention strategies. Immediate situational features characterize the specific emotional state, dialogue atmosphere, or implicit potential needs exhibited by the user in the current dialogue round. Based on the contextualized intervention strategy, corresponding skills are selected and combined from a dialogue micro-skills library and transformed into natural, coherent, and compliant language text to obtain and return the dialogue response text. The dialogue micro-skills library contains a variety of validated communication skills for building effective dialogues.
[0062] It should be added that, based on the personalized intervention strategy, and combined with the immediate contextual features extracted from the prior dialogue messages, context-adaptive adjustments are made to obtain a contextualized intervention strategy. This includes: based on the personalized intervention strategy and the immediate contextual features extracted from the prior dialogue messages, and based on a preset conflict judgment rule base, checking whether there is a conflict or inconsistency between the immediate contextual features and the personalized intervention strategy, and obtaining conflict detection results; wherein, the preset conflict judgment rule base is configured in advance according to actual conflicts, and the conflict detection results are used to characterize whether a conflict exists and the type of the corresponding conflict; performing pattern recognition or keyword matching on the personalized intervention strategy to obtain positive behavior information; determining the user state evolution trajectory based on the immediate contextual features and the immediate contextual features from previous rounds; and making a fusion decision based on the conflict detection results, positive behavior information, and user state evolution trajectory to determine the fine-tuning decision and fine-tune the personalized intervention strategy.
[0063] Furthermore, fine-tuning decisions include: if conflict is detected, adjusting the approach to address emotions first, then the issue, by adding a pre-emptive emotional response action to the personalized intervention strategy; if positive behavior is identified, adjusting the approach to seize opportunities and strengthen motivation, such as inserting an affirmation or summary step into the personalized strategy; if state evolution is detected, adjusting the approach to keep pace with the times and upgrade the strategy, such as switching from Guided Exploration to Execute Intent. This multi-layered architecture ensures that the intervention strategy is not a rigid script, but an intelligent process that can interact with the user in real time and evolve dynamically, thereby greatly improving the effectiveness of the intervention and the user experience.
[0064] In one optional embodiment, after generating the corresponding personalized response, the process includes: evaluating the response result to determine the response effect and improvement direction; updating each state; and storing the dialogue message and its corresponding ABC chain, DARN-CAT change signal, and the strategy executed by the corresponding response type in a memory bank.
[0065] Specifically, updating each state includes: when the mode is exploration, incrementing the number of consecutive exploration rounds and conversation rounds by 1, setting the number of consecutive persuasion rounds to 0, and adjusting the forced exploration flag to the second flag; when the mode is persuasion, setting the number of consecutive exploration rounds to 0, and incrementing the number of consecutive persuasion rounds and conversation rounds by 1; when the mode is crisis management, no update of the corresponding number of rounds is performed.
[0066] In summary, this embodiment of the invention assesses information sufficiency and security risks based on received dialogue messages, proactively identifies potential risks during the dialogue, and ensures that sufficient information is available to understand the user's current state and true needs before guiding user action. This makes subsequent guidance no longer blind but directional guidance based on a thorough understanding, avoiding hasty or erroneous responses when information is insufficient. Furthermore, it provides a clear decision-making mechanism by evaluating whether the corresponding priority judgment conditions are met according to a preset priority order. This allows the system to flexibly switch between different response decisions based on real-time evaluation results, avoiding decision confusion or logical conflicts. Based on meeting the judgment conditions, the corresponding priority response type is determined, thereby generating highly targeted and empathetic responses that make users feel understood and respected. This ensures that the generated responses are closely connected to the context of the current dialogue, making the dialogue smooth and natural.
[0067] The following describes the dynamic programming apparatus for multi-turn dialogue with persuasive purposes provided by the present invention. The dynamic programming apparatus for multi-turn dialogue with persuasive purposes described below can be referred to in correspondence with the dynamic programming method for multi-turn dialogue with persuasive purposes described above.
[0068] Figure 3 A schematic diagram of a multi-turn dialogue dynamic programming device for persuasive purposes is shown. The device includes: The evaluation module 31 evaluates the information sufficiency and security risks based on the received dialogue messages, and obtains the predicted information sufficiency and predicted security risks. The priority judgment module 32 evaluates whether the judgment conditions of the corresponding priority are met in turn, according to the sufficiency of the predicted information, the predicted security risks, and the current exploration round and persuasion round, in a preset priority order. The response module 33 determines the response type with the corresponding priority based on the judgment conditions, and generates a corresponding personalized response based on the dialogue message and the context message.
[0069] It should be noted that the specific principles of the embodiments of the present invention are the same as those of the method embodiments described above. For details, please refer to the method embodiments above. More detailed explanations will not be repeated here.
[0070] Figure 4 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 4 As shown, the electronic device may include a processor 410, a communications interface 420, a memory 430, and a communication bus 440. The processor 410, communications interface 420, and memory 430 communicate with each other via the communication bus 440. The processor 410 can call logical instructions in the memory 430 to execute a multi-turn dialogue dynamic programming method for persuasion purposes. This method includes: assessing user tolerance, information sufficiency, and security risk based on received dialogue messages to obtain predicted information sufficiency and predicted security risk; evaluating whether the judgment conditions for the corresponding priority are met sequentially according to a preset priority order based on the predicted information sufficiency, predicted security risk, and the current exploration and persuasion rounds; determining the response type for the corresponding priority based on the met judgment conditions; and generating a corresponding personalized response based on the dialogue messages and context messages.
[0071] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0072] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the multi-turn dialogue dynamic programming method for persuasion purposes provided by the above methods. The method includes: assessing user tolerance, information sufficiency, and security risks based on received dialogue messages to obtain predicted information sufficiency and predicted security risks; evaluating whether the judgment conditions of the corresponding priorities are met in sequence according to the predicted information sufficiency, predicted security risks, and the current exploration round and persuasion round, in a preset priority order; determining the response type of the corresponding priority based on the satisfaction of the judgment conditions; and generating a corresponding personalized response based on the dialogue messages and context messages.
[0073] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements a multi-turn dialogue dynamic programming method for persuasive purposes provided by the methods described above. The method includes: assessing user tolerance, information sufficiency, and security risk based on received dialogue messages to obtain predicted information sufficiency and predicted security risk; evaluating whether the judgment conditions of the corresponding priority are met in sequence according to the predicted information sufficiency, predicted security risk, and the current exploration round and persuasion round, in a preset priority order; determining the response type of the corresponding priority based on the satisfaction of the judgment conditions; and generating a corresponding personalized response based on the dialogue messages and in combination with context messages.
[0074] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0075] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0076] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A dynamic programming method for multi-turn dialogues aimed at persuasion, characterized in that, include: Based on the received dialogue messages, the information sufficiency and security risks are assessed to obtain the predicted information sufficiency and predicted security risks; Based on the sufficiency of the predicted information, the predicted security risks, and the current exploration and persuasion rounds, the judgment conditions of the corresponding priorities are evaluated in sequence according to the preset priority order. Based on the judgment conditions, the corresponding priority response type is determined, and a corresponding personalized response is generated according to the dialogue message and in combination with the context message.
2. The multi-turn dialogue dynamic programming method for persuasive purposes according to claim 1, characterized in that, Based on the judgment conditions, the corresponding priority response type is determined. According to the dialogue message and in conjunction with the context message, a corresponding personalized response is generated, including: When the judgment condition is met, i.e. the predicted security risk reaches the preset security risk threshold, the corresponding response type is determined to be the crisis management type. Based on the crisis management type, the security risk level is determined according to the dialogue messages and in conjunction with the context messages. Based on the security risk level and in conjunction with the corresponding preset crisis management strategy, security resources from the security resource database are invoked to generate a crisis response. The preset crisis management strategy is constructed in advance based on the intervention objectives of the corresponding security risk level, and the security resource database is constructed in advance based on available external support information and professional guidance content for reducing users' psychological risks.
3. The multi-turn dialogue dynamic programming method for persuasive purposes according to claim 1, characterized in that, Based on the fulfillment of judgment conditions, the corresponding priority response type is determined, and a corresponding personalized response is generated according to the dialogue message and in combination with the context message. Alternatively, if the judgment conditions for all priorities are not met based on a preset priority order, the process further includes: If the judgment condition is met, such as the number of persuasion rounds being greater than or equal to the maximum number of persuasion rounds or the forced exploration identifier being forced exploration, then the corresponding response type is determined to be an exploration response type. Based on the exploratory response type, the dialogue context is determined according to the dialogue messages and in conjunction with the context messages. Based on the dialogue messages and the context messages, the user behavior is evaluated, and a fusion decision is made in conjunction with the dialogue context to generate an exploratory response.
4. The multi-turn dialogue dynamic programming method for persuasive purposes according to claim 3, characterized in that, Based on the dialogue messages and the context messages, evaluate user behavior, including: Based on the dialogue messages and the context messages, using the cause-action-behavior-consequence-crit analysis method, the triggering context in the dialogue is captured, the user's reaction behavior is located, and the consequences of the behavior are analyzed to generate an ABC chain. Based on the dialogue messages and the context messages, the user's change signals are identified using the DARN-CAT classification method (desire, ability, reason, need, commitment, trigger, action) and categorized according to the reason for change and the language of change to obtain the change motivation. Based on the dialogue messages and the context messages, the decision balancing method is used to explore the advantages of change, disadvantages of change, advantages of maintaining the status quo, and disadvantages of maintaining the status quo when the user's hesitant language is detected, and to generate a decision balancing matrix to represent the overall decision.
5. The multi-turn dialogue dynamic programming method for persuasive purposes according to claim 1, characterized in that, Based on the fulfillment of judgment conditions, the corresponding priority response type is determined. Based on the dialogue message and in conjunction with context messages, a corresponding personalized persuasion response is generated, which also includes: When the judgment condition is satisfied that the number of exploration rounds is greater than or equal to the maximum number of consecutive persuasion rounds, or when the judgment condition is satisfied that the predicted information sufficiency is greater than or equal to the information sufficiency threshold and the number of conversation rounds is greater than or equal to the minimum total number of rounds required to trigger persuasion, the corresponding response type is determined to be the persuasion response type. Based on the persuasive response type, determine whether the information sufficiency of the local route assessment reaches the preset mode switching threshold. When the information sufficiency of the local route assessment reaches the preset mode switching threshold, determine the dialogue context and evaluate user behavior based on the dialogue message and context message to generate an exploratory response. When the sufficiency of the predicted information does not reach the preset mode switching threshold, the user intent is identified based on the dialogue message and the context message. Based on the user's intent, and combining the dialogue messages and the context messages, a corresponding persuasive response is generated.
6. The multi-turn dialogue dynamic programming method for persuasive purposes according to claim 5, characterized in that, Based on the user intent, and combining the dialogue messages and the context messages, a corresponding personalized response is generated, including: When the user's intent is security concern, the security risk level is determined based on the dialogue message and the context message. Then, in conjunction with the corresponding preset crisis management strategy, security resources in the security resource database are invoked to generate a crisis response. The preset crisis management strategy is constructed based on the intervention target corresponding to the security risk level, and the security resource database is constructed based on available external support information and professional guidance content for reducing the user's psychological risk. When the user intent is the identification type, the user state is identified based on the dialogue message and the context message. The user state includes user behavior, motivation and emotional state. Based on the identified user state, an intervention strategy is generated and the matching intervention strategy is executed; wherein the intervention strategy is used to characterize the intervention direction and intervention goal configured in response to the user's current needs.
7. The multi-turn dialogue dynamic programming method for persuasive purposes according to claim 6, characterized in that, Based on the identified user status, generate and execute intervention strategies, including: Based on the identified user state, determine the type of immediate response and select components; wherein, the components are pre-built and encapsulated functional modules based on different user states, using specific psychological techniques or communication logic. Based on the dialogue messages and the context messages, extract keywords or key phrases as evidence; Based on the selected components and the evidence, the most matching script is retrieved from a pre-set script library; wherein the pre-set script library is constructed based on scripts configured with different components and evidence, and the scripts are used to represent pre-packaged intervention templates for specific psychological goals. An intervention strategy is generated based on the matched components, the evidence, and the retrieved most matching script. Based on the intervention strategy, the user's historical interaction information is read from a preset memory bank, and the user's cognitive and emotional characteristics are identified; wherein, the user's cognitive and emotional characteristics are used to characterize the user's cognitive and emotional tendencies in the process of information processing and communication interaction. Based on the user's cognitive and emotional characteristics, the intervention strategy is adapted and adjusted. The most matching communication tone is selected as a style marker and added to the intervention strategy to obtain a personalized intervention strategy. Based on the personalized intervention strategy, and combined with the immediate contextual features extracted from the dialogue messages, contextual adaptation adjustments are made to obtain a contextualized intervention strategy; wherein, the immediate contextual features are used to characterize the specific emotional state, dialogue atmosphere, or implicit potential needs exhibited by the user in the current dialogue round. According to the contextualized intervention strategy, corresponding skills are selected and combined from the dialogue micro-skills library, and transformed into natural, coherent and compliant language text to obtain dialogue response text and return it; wherein, the dialogue micro-skills library contains a variety of verified communication skills for building effective dialogues.
8. A dynamic programming device for multi-turn dialogue aimed at persuasion, characterized in that, include: The evaluation module assesses information sufficiency and security risks based on the received dialogue messages, and obtains predicted information sufficiency and predicted security risks. The priority judgment module evaluates whether the judgment conditions of the corresponding priority are met in sequence according to the sufficiency of the predicted information, the predicted security risk, and the current exploration round and persuasion round, in a preset priority order. The response module determines the response type with corresponding priority based on the judgment conditions, and generates a corresponding personalized response based on the dialogue message and the context message.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the multi-turn dialogue dynamic programming method for persuasive purposes as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the multi-turn dialogue dynamic programming method for persuasive purposes as described in any one of claims 1 to 7.