A method and system for determining voice response strategies based on dynamic risk assessment
By employing dynamic risk assessment methods and large language models, fine-grained risk assessment and dynamic control of risks in voice dialogue systems have been achieved, resolving the issue of inconsistent strategy switching in existing technologies and improving system security and interactive experience.
Patent Information
- Application Number
- CN202511249251.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-03
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2045-09-03
AI Technical Summary
Existing voice dialogue systems lack a unified strategy switching mechanism when facing high risks and ambiguous intentions, resulting in a disconnect between security and interactive experience, and failing to smoothly transition between strategic rejection and gradual clarification within the same closed-loop process.
A dynamic risk assessment-based approach is adopted, which uses a pre-trained large language model to score the risk of user input and classifies it into high, medium and low risk levels. Strategic rejection, multi-round clarification or normal interaction processes are adopted respectively to achieve fine-grained assessment and dynamic control.
It improves the accuracy of risk identification and the consistency of interaction, reduces the false judgment rate, enhances the system's adaptability and user experience, and ensures flexible response in different risk scenarios.
Smart Images

Figure CN120748402B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of voice interaction technology, and in particular to a method and system for determining voice response strategies based on dynamic risk assessment. Background Technology
[0002] Currently, with the widespread application of large-scale pre-trained language models in voice dialogue systems, human-computer interaction is gradually shifting from rule-driven to an intelligent stage centered on deep learning. Neural network-based dialogue systems can learn natural language expressions and contextual relationships through massive amounts of data, achieving a smoother interactive experience. However, open-domain dialogue also brings security and compliance risks. When user input contains sensitive, illegal, or injection-attack intentions, the system must effectively identify and prevent the leakage of risky information or misleading answers while ensuring the quality of question-and-answer communication and user experience. This has become a fundamental problem that current voice interaction technology urgently needs to solve.
[0003] Existing voice dialogue security management often employs a dual-track triage mechanism: on the one hand, risk detection modules based on fixed thresholds or rules directly map high-risk or out-of-bounds requests to predefined rejection template responses; on the other hand, when the reliability of user intent recognition is insufficient, single-round or multi-round clarification follow-up questions are triggered to obtain supplementary information. Typical examples include: attack defense devices that use fixed templates to reject answers and truncate responses when a risk of sensitive information leakage is detected; template-based rejection is directly executed when the intent is identified as exceeding the business scope based on a dual-model judgment of function and intent; and multi-round question-and-answer systems that initiate multiple rounds of clarification when the intent is unclear, but do not conduct real-time risk reassessment. Although the above solutions improve security or intent recognition accuracy in their respective areas, they still have shortcomings such as fragmented branch decision-making, abrupt rejection responses, and a disconnect between the clarification link and risk assessment.
[0004] Therefore, although existing technologies can quickly intercept high-risk requests and gradually clarify ambiguous intentions, they lack a unified policy switching mechanism based on risk gradients. They cannot smoothly transition between strategic rejection and gradual clarification within a single framework, nor can they provide users with contextualized modification guidance or dynamically adjust the defense depth based on supplementary information. The core technical problem arising from this is: how to build a voice dialogue system that can perform fine-grained assessment and dynamic control of multi-level risk scenarios, and adaptively switch rejection and clarification policies within the same closed-loop process, so as to balance security compliance and a natural and smooth interactive experience. Summary of the Invention
[0005] This application provides a method and system for determining voice response strategies based on dynamic risk assessment. It constructs a system within a voice dialogue system that can perform fine-grained assessment and dynamic control of multi-level risk scenarios, and adaptively switch between rejection and clarification strategies within the same closed-loop process, thus balancing security and compliance with a natural and smooth interactive experience. This application provides the following technical solutions:
[0006] Firstly, this application provides a method for determining a voice response strategy based on dynamic risk assessment, the method comprising:
[0007] In response to a user-initiated voice interaction request, the system obtains user input and calls a pre-trained large language model to output a risk score.
[0008] The risk score is classified into levels based on a preset risk threshold, including high risk, medium risk and low risk.
[0009] If the risk level is high, immediately proceed with the strategic rejection response process;
[0010] If the risk level is medium risk, a multi-round clarification dialogue process will be initiated, and the risk score will be recalculated continuously based on the constructed clarification questions and the user's answers until the risk score is no longer in the medium risk range.
[0011] If the risk level is low, proceed directly to the normal interaction process.
[0012] In one specific implementation, obtaining user input and invoking a pre-trained large language model to output a risk score includes:
[0013] The speech recognition module transcribes the user's voice request into text, calls a pre-trained generative large language model to perform semantic analysis on the text input, extracts semantic features, contextual information and potential risk factors to comprehensively evaluate the semantic content of the user's request, and outputs a risk score S in real time through a built-in risk calculation mechanism.
[0014] In one specific implementation scheme, the risk score is classified into levels based on a preset risk threshold, and the risk levels include high risk, medium risk, and low risk.
[0015] Risk score S and pre-set dual thresholds and Perform the comparisons sequentially: when When, it is judged as high risk; when When, it is judged as medium risk; when At that time, it was determined to be low risk, among which .
[0016] In a specific feasible implementation, if the risk level is high risk, immediately initiating the strategic rejection response process includes:
[0017] Based on the text content input by the user and its contextual dialogue information, a pre-trained generative large language model is invoked to perform deep semantic parsing of the request. During the parsing process, specific feature information that triggers the high-risk judgment in the request is identified and extracted. Based on understanding the user's intent, the specific reasons why the request is judged as high-risk are clearly defined. The large language model generates a rejection response text based on the current context, and outputs the specific reasons for the rejection simultaneously.
[0018] In a specific feasible implementation, if the risk level is medium risk, a multi-round clarification dialogue process is initiated, continuously recalculating the risk score based on the constructed clarification questions and the user's answers, until the risk score is no longer in the medium risk range, including:
[0019] When the current risk score S is detected to be in the medium risk range, a vague and guiding clarification question is constructed based on the semantic feature analysis of the current input using a large language model. After the user provides supplementary information based on the clarification question, the pre-trained generative large language model is called again to perform semantic understanding and feature extraction on the new input, and the risk score S is updated and recalculated.
[0020] If the new risk score S rises to the high-risk range, the clarification process will be terminated immediately, and a strategic rejection process will be initiated, generating a rejection response that includes a risk description and recommendations.
[0021] If the new risk score S drops to the low-risk range, and the user's intent can be clearly identified based on the current semantic features, then the clarification process ends and the normal response process begins.
[0022] If the new risk score is still in the medium-risk range, the next round of clarification questions will be generated, the user will supplement the information again, the score will be reassessed, and the scoring recalculation process will be repeated until the score jumps out of the medium-risk range.
[0023] In a specific feasible implementation, if the risk level is low risk, directly entering the normal interaction process includes:
[0024] After outputting the risk score, if the score S is lower than the preset low-risk threshold... If the request is deemed a security request, the request content is directly input into the response generation module, which then calls the large language model to generate a standard response based on the current semantic content and context state.
[0025] In one specific implementation scheme, the step of classifying risk scores based on preset risk thresholds, where risk levels include high risk, medium risk, and low risk, further includes:
[0026] Let the two thresholds be the high-risk determination thresholds. Low-risk determination threshold At any given time t during a voice interaction round, the risk interval difference variable is set as follows:
[0027] ;
[0028] After each round of interaction The following updates will be made:
[0029] ;
[0030] in, , , , These are the adjustment parameters for nonlinear control, score fluctuation weighting, semantic fuzzy weighting, and symmetric offset feedback, respectively. This represents the standard deviation of the continuously output risk scores within the most recent voice interaction window. This represents the semantic ambiguity coefficient of the current voice input. Indicates the disturbance feedback term;
[0031] After each round of formula updates, a new risk buffer difference will be obtained. The dual thresholds are dynamically adjusted according to the following rules:
[0032] ;
[0033] .
[0034] Secondly, this application provides a system for determining voice response strategies based on dynamic risk assessment, employing the following technical solution:
[0035] A system for determining voice response strategies based on dynamic risk assessment includes:
[0036] The risk score output module is used to respond to user-initiated voice interaction requests, obtain user input, and call a pre-trained large language model to output a risk score.
[0037] The risk level classification module is used to classify the risk score into levels based on a preset risk threshold. The risk levels include high risk, medium risk, and low risk.
[0038] The high-risk determination module is used to immediately initiate a strategy-based rejection response process if the risk level is high.
[0039] The medium-risk determination module is used to enter a multi-round clarification dialogue process if the risk level is medium risk, and continuously recalculate the risk score based on the constructed clarification questions and the user's answers until the risk score is no longer in the medium-risk range.
[0040] The low-risk determination module is used to directly enter the normal interaction process if the risk level is low.
[0041] Thirdly, this application provides an electronic device, the device including a processor and a memory; the memory stores a program, the program being loaded and executed by the processor to implement a method for determining a voice response strategy based on dynamic risk assessment as described in the first aspect.
[0042] Fourthly, this application provides a computer-readable storage medium storing a program that, when executed by a processor, is used to implement a method for determining a voice response strategy based on dynamic risk assessment as described in the first aspect.
[0043] In summary, the beneficial effects of this application include at least the following:
[0044] (1) By using a pre-trained large language model to accurately extract the semantic features of user requests and deeply capture the potential risk features of the input text, the risk scoring is refined and highly accurate. Compared with traditional risk identification schemes based on rules or simple models, this scheme can improve the accuracy of risk identification to a certain extent, significantly reduce the probability of misjudgment and missed judgment, and effectively solve the problem of coarse risk classification in existing technologies.
[0045] (2) By designing a dual-threshold risk level assessment system, this system achieves fine-grained classification of user requests (high, medium, and low risk) and dynamically adjusts the processing strategy in real time. When the dual-threshold dynamic strategy is adopted, the false judgment rate of medium-risk requests can be improved to a certain extent, and the system can switch more flexibly between rejection and clarification, effectively solving the problems of interaction rigidity and security vulnerabilities caused by traditional single-path decision-making.
[0046] (3) In medium-risk scenarios, this system adopts a multi-round clarification dialogue and real-time risk reassessment linkage mechanism. By dynamically inquiring about the user's intent and recalculating the risk score in real time, it ensures that the risk can be accurately judged after obtaining more supplementary information. After adopting the multi-round clarification and real-time risk reassessment mechanism, the system's accuracy in security risk detection can be improved to a certain extent, which greatly improves the coherence and security of the interaction process and significantly solves the problem of the separation between the clarification link and the security assessment.
[0047] By introducing a multi-level risk classification mechanism and combining contextual semantic understanding with the risk scoring capabilities of a generative large language model, this system achieves refined risk identification, tiered response execution, and strategic interaction control of voice input content within a unified process framework. This effectively solves the technical problems of coarse interaction security control, single response strategies, and fragmented user experience in existing voice systems. The risk score S output by the large language model is used as a quantitative indicator, and a set of statically preset dual-risk judgment thresholds are set, allowing risk levels to be divided into high-risk, medium-risk, and low-risk levels, each corresponding to different processing strategy paths. Specifically, when the score falls into the high-risk range, the system immediately triggers a strategic rejection response process. Through model analysis of user intent and semantic risk factors, it outputs targeted, explanatory, and guiding rejection information to avoid invalid interactions or the generation of potentially inappropriate content. When the score is in the medium-risk range, the system initiates a multi-round clarification interaction process, guiding the user to gradually supplement or clarify semantic expressions while maintaining contextual continuity. The risk score is reassessed after each round of interaction until a clear response path is determined. When the score falls into the low-risk range, the system directly outputs regular voice response content, maintaining the naturalness and efficiency of the interaction. By employing the methods described above, the voice interaction system can flexibly switch response strategies when facing different risk scenarios. This ensures real-time blocking of high-risk content while preserving room for guiding and optimizing medium-risk semantics, thereby avoiding the risks of misjudgment and interaction interruption caused by a one-size-fits-all approach. In particular, during the handling of medium-risk situations, multi-round semantic clarification and large-model-assisted understanding enhance the system's ability to identify and resolve ambiguous expressions, leaps in semantics, or edge requests, effectively improving the system's fault tolerance and adaptability.
[0048] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, the preferred embodiments of this application are described in detail below with reference to the accompanying drawings. Attached Figure Description
[0049] Figure 1 This is a flowchart illustrating the method for determining a voice response strategy based on dynamic risk assessment in an embodiment of this application.
[0050] Figure 2 This is a schematic diagram of the overall process of determining the voice response strategy based on dynamic risk assessment in the embodiments of this application.
[0051] Figure 3 This is a structural block diagram of the system for determining a voice response strategy based on dynamic risk assessment in the embodiments of this application.
[0052] Figure 4 This is a block diagram of an electronic device that determines a voice response strategy based on dynamic risk assessment in an embodiment of this application. Detailed Implementation
[0053] The specific embodiments of this application will be described in further detail below with reference to the accompanying drawings and examples. The following examples are used to illustrate this application, but are not intended to limit the scope of this application.
[0054] Optionally, this application uses the method for determining a voice response strategy based on dynamic risk assessment provided in various embodiments as an example for application in an electronic device. The electronic device is a terminal or a server. The terminal can be a mobile phone, computer, tablet computer, etc. This embodiment does not limit the type of electronic device.
[0055] Reference Figure 1 This is a flowchart illustrating a method for determining a voice response strategy based on dynamic risk assessment, provided in one embodiment of this application. The method includes at least the following steps:
[0056] Step S101: In response to the user's voice interaction request, obtain the user's input and call the pre-trained large language model to output a risk score.
[0057] In step S101, in response to the user's interactive request initiated via voice, after receiving the request, a pre-trained generative large language model is invoked to perform semantic understanding and feature extraction on the input content, thereby completing the preliminary risk level assessment. The risk score S ranges from 0 to 1, with a higher value indicating a higher potential security risk of the request.
[0058] Specifically, the system first uses a speech recognition module to transcribe the user's voice request into text, ensuring the input is easy to process. Then, a pre-trained generative large language model is used to perform semantic analysis on the text input, extracting semantic features, contextual information, and potential risk factors. This model combines rich linguistic knowledge and risk identification rules to comprehensively evaluate the semantic content of the user's request. A built-in risk calculation mechanism outputs a risk score S in real time, with the score strictly limited to the range of 0 to 1, quantifying the security risk level of the input request.
[0059] It should be noted that the calculation process of this score is essentially within the scope of existing technology and has been widely used in current generative AI systems. Although the specific extraction methods of semantic features and the weight settings of sensitive factors may differ among systems, the overall process follows the consistent principles of semantic understanding, feature extraction, and risk-weighted scoring.
[0060] Optionally, this application selects the Dongfeng Big Model, developed by Speechocean, as the large language model. Based on the Transformer architecture, the Dongfeng Big Model, through pre-training on a large-scale corpus, possesses general semantic representation and natural language generation capabilities, and is a mature existing technology model in this field. In implementation, the Dongfeng Big Model performs semantic understanding on user input and outputs corresponding semantic features. Subsequently, the risk calculation mechanism set in the model receives the semantic features and calculates them, thereby realizing the function of outputting a risk score S.
[0061] Step S102: Classify the risk score into levels based on the preset risk threshold. The risk levels include high risk, medium risk and low risk.
[0062] In step S102, the core task is to apply the risk score S output by the model to a pre-set dual threshold. and Classify into levels, among which This maps the risk score into three levels: high risk, medium risk, and low risk, facilitating the refined allocation of subsequent interaction strategies.
[0063] Specifically, the risk score S is first compared sequentially with two thresholds: when When, it is judged as high risk; when When, it is judged as medium risk; when At that time, it is judged as low risk. In specific implementation, the score S and the threshold are... and The comparison and judgment are not performed in isolation, but rather make decisions based on a comprehensive set of factors, including the current interaction context, the semantic completeness of the user's request, and historical rating trends. This is particularly important when S is near a threshold boundary, allowing for a more nuanced assessment of the rating results. For example, if the rating is slightly higher than... However, if the user's request history shows that they frequently ask incomplete or vague questions, it's advisable to keep the request in the medium-risk range to trigger the necessary clarification process. Conversely, if the rating is slightly lower... However, user requests may contain strong contextual references and risk characteristics, which could trigger high-risk processes as appropriate. This dynamic fine-tuning mechanism does not change the threshold itself, but rather, in edge-scoring situations, it finely calibrates the classification results through joint judgment of multi-dimensional features.
[0064] Step S103: If the risk level is high, immediately proceed with the strategic rejection response process.
[0065] In step S103, when the risk score S reaches or exceeds the preset high-risk threshold... If a request is deemed high-risk, a strategic rejection response process is immediately initiated. At this stage, no further interactive clarification or risk reassessment is performed; instead, the strategic rejection response process corresponding to the high-risk request is immediately activated to achieve immediate blocking and appropriate guidance for such high-risk requests.
[0066] Specifically, the rejection response process includes the following steps: First, based on the user's input text content and its contextual dialogue information, a pre-trained generative large language model is invoked to perform deep semantic parsing of the request. During the parsing process, specific features in the request that may trigger a high-risk determination are identified and extracted, such as: whether it involves content that violates laws and regulations, whether it contains keywords in sensitive fields, whether there is an attempt to access the request without authorization, and whether there are misleading or offensive expressions. Through the above analysis, based on understanding the user's intent, the specific reasons why the request is determined to be high-risk are clearly defined.
[0067] Subsequently, based on the identification of risk factors, the large language model dynamically generates a context-adaptive rejection response text according to the current context. This text not only clearly informs the user that the request has been rejected by the system, but also simultaneously outputs the specific reasons for the rejection, such as: "Your request involves access to sensitive areas and cannot be processed at this time" or "This type of issue falls within the scope of system restrictions; it is recommended to adjust the expression." In addition, to improve the operability of the interaction and user cooperation, the rejection response also embeds specific and feasible modification suggestions, such as guiding the user to rephrase the request, avoid sensitive expressions, or use clearer language instructions.
[0068] The strategy-based rejection response generated by the large language model ensures that high-risk requests are blocked in a timely manner while providing users with valuable feedback. This reduces interaction security risks and avoids negative user experiences due to rejection. Furthermore, the non-template-based generation method of rejection responses enhances the diversity and contextual fit of the response content, significantly improving the system's flexibility and professionalism in handling complex semantic request scenarios.
[0069] Step S104: If the risk level is medium risk, proceed to a multi-round clarification dialogue process, continuously recalculating the risk score based on the constructed clarification questions and the user's answers, until the risk score is no longer in the medium risk range.
[0070] In step S104, when the risk score S is in the medium-risk range, a rejection response or direct response is not immediately generated. Instead, a processing flow centered on multiple rounds of clarification interaction and dynamic risk reassessment is initiated. This flow aims to gradually guide the user to clarify their intent for requests that are semantically ambiguous, incompletely expressed, or have potential sensitive factors but do not yet constitute a high risk, and to update the risk score and select decision branches in real time accordingly.
[0071] Specifically, when the current risk score S is detected to be in the medium-risk range, a vague but guiding clarification question is first constructed as a response to the current user request. The clarification question is generated based on the semantic feature analysis of the current input using a large language model. Its design aims to guide the user to provide more explicit and specific background information or behavioral intentions, thereby enabling the system to obtain more sufficient context for risk reassessment. After the user provides supplementary information based on the clarification question, the pre-trained generative large language model is invoked again to perform semantic understanding and feature extraction on the new input, updating and recalculating the risk score S. If the new risk score S rises to the high-risk range, the clarification process is immediately terminated, transitioning to a strategic rejection process, and a rejection response containing risk explanations and suggestions is generated. If the new risk score S falls to the low-risk range, and the user's intention can be clearly understood based on the current semantic features, the clarification process ends, and the normal response process begins. If the new risk score remains in the medium-risk range, the next round of clarification questions is generated, the user provides supplementary information again, and the score is reassessed. This recalculation process is repeated until the score falls out of the medium-risk range.
[0072] In implementation, this multi-round clarification and risk reassessment process possesses state preservation and contextual continuity capabilities. Each round of clarification is based on the previous semantic context and the latest input for synthetic modeling, ensuring that the model considers the continuity and evolutionary direction of the dialogue history in its judgments, avoiding semantic misjudgments or score fluctuations. Through the construction and execution of this multi-round clarification dialogue mechanism, on the one hand, the mechanism avoids a one-size-fits-all rejection judgment for unclear or slightly sensitive requests, technically reducing the probability of false rejection and ensuring the fault tolerance and flexibility of the voice interaction system; on the other hand, by guiding users to gradually improve their input and implementing real-time risk score updates, the system can dynamically adjust the risk judgment results, achieving refined and tiered risk handling responses; furthermore, this process also enhances user participation and feedback interpretability, allowing users to correct their requests based on their understanding of the issues, thereby enhancing the collaborative and adaptive capabilities of human-computer interaction.
[0073] Step S105: If the risk level is low, proceed directly to the normal interaction process.
[0074] In step S105, when the risk score S is in the low-risk range, it means that the current user request does not contain any identifiable security risks, sensitive information or compliance risks. Therefore, there is no need to implement any rejection or clarification strategy based on security controls. The user can directly enter the standard voice interaction process and output the normal response content.
[0075] Specifically, after the initial risk score is output, if the score S is lower than the system's preset low-risk threshold... If the request is deemed a security request, it is classified as such. At this point, subsequent clarification processes or risk reassessment mechanisms are not triggered. Instead, the request content is directly input into the response generation module, which uses a large language model to generate a standard response based on the current semantic content and context. The response content does not contain risk warnings, rejection guidance, or vague questions; it fully follows the user's intent to provide a service response, ensuring a natural and complete interactive experience. This step simplifies the processing path, accelerates the response pace, and improves service efficiency when the request risk is confirmed to be low, avoiding negative impacts on user experience due to excessive intervention. Simultaneously, clearly defined low-risk judgment boundaries ensure that the system only allows responses under risk-free conditions, constructing a secure and controllable closed loop for voice interaction.
[0076] Furthermore, as a preferred embodiment, in order to improve the accuracy of risk level judgment and the flexibility of strategy response in complex semantic interaction processes, this application further proposes a risk buffer update mechanism based on dynamic adjustment of threshold difference. This mechanism introduces a structurally innovative mathematical formula to achieve adaptive update of the risk division threshold interval, effectively enhancing the system's ability to respond to dynamic context changes.
[0077] Specifically, based on step S102, let the two thresholds on which the system risk level determination depends be the high-risk determination threshold. Low-risk determination threshold At any given time t during a voice interaction round, the risk interval difference variable is set as follows:
[0078] ;
[0079] This difference determines the width of the medium-risk interval, thus affecting the system's tolerance and handling strategy for boundary risk requests. To enable this risk buffer to dynamically adjust over time, the system performs adjustments after each round of interaction. The following updates will be made:
[0080] ;
[0081] in, , , , These are the adjustment parameters for nonlinear control, score fluctuation weight, semantic fuzzy weight, and symmetric offset feedback, which can be set according to risk control level, training data, or business strategy. This represents the standard deviation of risk scores output consecutively within the most recent voice interaction window, used to characterize the volatility of user intent during the current interaction. This parameter can be calculated by statistically analyzing recent scores using a sliding window. This represents the semantic ambiguity coefficient of the current speech input. This value is obtained by the large language model extracting features such as syntactic completeness, semantic certainty, and ambiguity of the user's text input, and then mapping them through a preset semantic ambiguity function. The range is [0,1], and the higher the value, the more ambiguous the expression. This represents the perturbation feedback item, calculated based on the degree of semantic intent shift between rounds (e.g., keyword shift, increased context span, etc.), which is used to compensate for scoring instability caused by input jumps. This represents the symmetry tension factor of the threshold structure. When the high and low risk thresholds are not symmetrically distributed around the center of the risk scoring interval, this term is introduced for fine-tuning to reconstruct the balance of the risk interval.
[0082] After each round of executing this update formula, a new risk buffer difference will be obtained. Then, the dual thresholds are dynamically adjusted according to the following rules:
[0083] ;
[0084] ;
[0085] In the design of the above formula, a dynamically adjustable risk buffer width variable is introduced and dynamically updated round by round through a formula that combines nonlinear mapping and multi-factor fusion. By introducing a difference variable and giving it the ability to change dynamically with interaction, the system can adjust the width of the medium-risk interval in real time based on the current semantic environment, user behavior volatility, and historical rating changes, thereby adapting to the risk perception sensitivity in different contexts: when the interaction is stable and the semantics are clear, the medium-risk interval is narrowed, and the system responds more decisively; when the interaction is unstable or the user's expression is ambiguous, the medium-risk interval is appropriately widened, and the system enters a multi-round clarification process to improve fault tolerance.
[0086] In traditional risk assessment and interaction systems, risk thresholds are often statically set, suitable for fixed scenarios. However, in real-world voice interaction applications, user behavior patterns, semantic features, and risk contexts are dynamically changing, making it difficult for fixed thresholds to consistently guarantee the accuracy and flexibility of the system's response. To address this issue, this solution introduces a dynamic dual-threshold update mechanism based on contextual state, interaction behavior, and system operation data. Different application scenarios have different risk tolerances; for example, medical voice assistants and general entertainment question-and-answer systems have significantly different risk judgment criteria. By dynamically updating the dual thresholds, the system can automatically adapt to current application needs based on historical interaction characteristics, enabling fine-grained strategy switching. Simultaneously, real-time updates to the risk thresholds help the system capture subtle adjustments in user behavior patterns and request content distribution during long-term operation, thus maintaining a high level of responsiveness to gray-zone risk scenarios and avoiding misjudgments (false rejections or false responses) caused by static classification.
[0087] In summary, combining Figure 2 This application introduces a multi-level risk classification mechanism, combining contextual semantic understanding with the risk scoring capabilities of a generative large language model. This enables refined risk identification, tiered response execution, and strategic interaction control of voice input content within a unified process framework, effectively solving technical problems in existing voice systems such as coarse interaction security control, simplistic response strategies, and fragmented user experience. The risk score S output by the large language model is used as a quantitative indicator, and a set of statically preset dual-risk judgment thresholds are set, allowing risk levels to be divided into high-risk, medium-risk, and low-risk levels, each corresponding to different processing strategy paths. Specifically, when the score falls into the high-risk range, the system immediately triggers a strategic rejection response process. Through model analysis of user intent and semantic risk factors, it outputs targeted, explanatory, and guiding rejection information to avoid invalid interactions or the generation of potentially inappropriate content. When the score is in the medium-risk range, the system initiates a multi-round clarification interaction process, guiding the user to gradually supplement or clarify semantic expressions while maintaining contextual continuity. The risk score is reassessed after each round of interaction until a clear response path is determined. When the score falls into the low-risk range, the system directly outputs regular voice response content, maintaining the naturalness and efficiency of the interaction. By employing the methods described above, the voice interaction system can flexibly switch response strategies when facing different risk scenarios. This ensures real-time blocking of high-risk content while preserving room for guiding and optimizing medium-risk semantics, thereby avoiding the risks of misjudgment and interaction interruption caused by a one-size-fits-all approach. In particular, during the handling of medium-risk situations, multi-round semantic clarification and large-model-assisted understanding enhance the system's ability to identify and resolve ambiguous expressions, leaps in semantics, or edge requests, effectively improving the system's fault tolerance and adaptability.
[0088] Figure 3 This is a structural block diagram of a system for determining a voice response strategy based on dynamic risk assessment, provided in one embodiment of this application. The system includes at least the following modules:
[0089] The risk score output module is used to respond to user-initiated voice interaction requests, obtain user input, and call a pre-trained large language model to output a risk score.
[0090] The risk level classification module is used to classify risk scores into levels based on preset risk thresholds. Risk levels include high risk, medium risk, and low risk.
[0091] The high-risk assessment module is used to immediately initiate a strategic rejection response process if the risk level is high.
[0092] The medium-risk determination module is used to initiate a multi-round clarification dialogue process if the risk level is medium. It continuously recalculates the risk score based on the constructed clarification questions and the user's answers until the risk score is no longer in the medium-risk range.
[0093] The low-risk assessment module is used to directly proceed to the normal interaction process if the risk level is low.
[0094] For relevant details, please refer to the above method implementation examples.
[0095] Figure 4 This is a block diagram of an electronic device provided in one embodiment of this application. The device includes at least a processor 401 and a memory 402.
[0096] Processor 401 may include one or more processing cores, such as a quad-core processor or an octa-core processor. Processor 401 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 401 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 401 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 401 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.
[0097] The memory 402 may include one or more computer-readable storage media, which may be non-transitory. The memory 402 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 402 are used to store at least one instruction, which is executed by the processor 401 to implement the method for determining a voice response strategy based on dynamic risk assessment provided in the method embodiments of this application.
[0098] In some embodiments, the electronic device may also optionally include: a peripheral device interface and at least one peripheral device. The processor 401, memory 402, and peripheral device interface can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface via a bus, signal line, or circuit board. Indicatively, peripheral devices include, but are not limited to: radio frequency circuits, touch displays, audio circuits, and power supplies.
[0099] Of course, electronic devices may also include fewer or more components, and this embodiment does not limit this.
[0100] Optionally, this application also provides a computer-readable storage medium storing a program that is loaded and executed by a processor to implement the method of determining a voice response strategy based on dynamic risk assessment in the above-described method embodiments.
[0101] Optionally, this application also provides a computer product including a computer-readable storage medium storing a program, which is loaded and executed by a processor to implement the method of determining a voice response strategy based on dynamic risk assessment as described in the above method embodiments.
[0102] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0103] The above embodiments merely illustrate several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. A method for determining a voice response strategy based on dynamic risk assessment, characterized in that, The method includes: In response to a user-initiated voice interaction request, the system obtains user input and calls a pre-trained large language model to output a risk score. The risk score is classified into levels based on a preset risk threshold, including high risk, medium risk and low risk. The risk score is classified into levels based on a preset risk threshold, including high risk, medium risk, and low risk. Let the two thresholds be the high-risk determination thresholds. Low-risk determination threshold At any given time t during a voice interaction round, the risk interval difference variable is set as follows: ; After each round of interaction The following updates will be made: ; in, , , , These are the adjustment parameters for nonlinear control, score fluctuation weighting, semantic fuzzy weighting, and symmetric offset feedback, respectively. This represents the standard deviation of the continuously output risk scores within the most recent voice interaction window. This represents the semantic ambiguity coefficient of the current voice input. Indicates the disturbance feedback term; After each round of formula updates, a new risk buffer difference will be obtained. The dual thresholds are dynamically adjusted according to the following rules: ; ; If the risk level is high, immediately proceed with the strategic rejection response process; If the risk level is medium risk, a multi-round clarification dialogue process will be initiated, and the risk score will be recalculated continuously based on the constructed clarification questions and the user's answers until the risk score is no longer in the medium risk range. If the risk level is low, proceed directly to the normal interaction process.
2. The method for determining a voice response strategy based on dynamic risk assessment according to claim 1, characterized in that, The process of obtaining user input and calling a pre-trained large language model to output a risk score includes: The speech recognition module transcribes the user's voice request into text, calls a pre-trained generative large language model to perform semantic analysis on the text input, extracts semantic features, contextual information and potential risk factors to comprehensively evaluate the semantic content of the user's request, and outputs a risk score S in real time through a built-in risk calculation mechanism.
3. The method for determining a voice response strategy based on dynamic risk assessment according to claim 2, characterized in that, The risk score is classified into levels based on a preset risk threshold, including high risk, medium risk, and low risk. Risk score S and pre-set dual thresholds and Perform the comparisons sequentially: when When, it is judged as high risk; when When, it is judged as medium risk; when At that time, it was determined to be low risk, among which .
4. The method for determining a voice response strategy based on dynamic risk assessment according to claim 1, characterized in that, If the risk level is high risk, the immediate initiation of the strategic rejection response process includes: Based on the text content input by the user and its contextual dialogue information, a pre-trained generative large language model is invoked to perform deep semantic parsing of the request. During the parsing process, specific feature information that triggers the high-risk judgment in the request is identified and extracted. Based on understanding the user's intent, the specific reasons why the request is judged as high-risk are clearly defined. The large language model generates a rejection response text based on the current context, and outputs the specific reasons for the rejection simultaneously.
5. The method for determining a voice response strategy based on dynamic risk assessment according to claim 3, characterized in that, If the risk level is medium risk, a multi-round clarification dialogue process will be initiated, continuously recalculating the risk score based on the constructed clarification questions and the user's answers, until the risk score is no longer in the medium risk range, including: When the current risk score S is detected to be in the medium risk range, a vague and guiding clarification question is constructed based on the semantic feature analysis of the current input using a large language model. After the user provides supplementary information based on the clarification question, the pre-trained generative large language model is called again to perform semantic understanding and feature extraction on the new input, and the risk score S is updated and recalculated. If the new risk score S rises to the high-risk range, the clarification process will be terminated immediately, and a strategic rejection process will be initiated, generating a rejection response that includes a risk description and recommendations. If the new risk score S drops to the low-risk range, and the user's intent can be clearly identified based on the current semantic features, then the clarification process ends and the normal response process begins. If the new risk score is still in the medium-risk range, the next round of clarification questions will be generated, the user will supplement the information again, the score will be reassessed, and the scoring recalculation process will be repeated until the score jumps out of the medium-risk range.
6. The method for determining a voice response strategy based on dynamic risk assessment according to claim 3, characterized in that, If the risk level is low, proceeding directly to the normal interaction process includes: After outputting the risk score, if the score S is lower than the preset low-risk threshold... If the request is deemed a security request, the request content is directly input into the response generation module, which then calls the large language model to generate a standard response based on the current semantic content and context state.
7. A system for determining a voice response strategy based on dynamic risk assessment, characterized in that, include: The risk score output module is used to respond to user-initiated voice interaction requests, obtain user input, and call a pre-trained large language model to output a risk score. The risk level classification module is used to classify the risk score into levels based on a preset risk threshold. The risk levels include high risk, medium risk, and low risk. The risk score is classified into levels based on a preset risk threshold, including high risk, medium risk, and low risk. Let the two thresholds be the high-risk determination thresholds. Low-risk determination threshold At any given time t during a voice interaction round, the risk interval difference variable is set as follows: ; After each round of interaction The following updates will be made: ; in, , , , These are the adjustment parameters for nonlinear control, score fluctuation weighting, semantic fuzzy weighting, and symmetric offset feedback, respectively. This represents the standard deviation of the continuously output risk scores within the most recent voice interaction window. This represents the semantic ambiguity coefficient of the current voice input. Indicates the disturbance feedback term; After each round of formula updates, a new risk buffer difference will be obtained. The dual thresholds are dynamically adjusted according to the following rules: ; ; The high-risk determination module is used to immediately initiate a strategy-based rejection response process if the risk level is high. The medium-risk determination module is used to enter a multi-round clarification dialogue process if the risk level is medium risk, and continuously recalculate the risk score based on the constructed clarification questions and the user's answers until the risk score is no longer in the medium-risk range. The low-risk determination module is used to directly enter the normal interaction process if the risk level is low.
8. An electronic device, characterized in that, The device includes a processor and a memory; the memory stores a program that is loaded and executed by the processor to implement a method for determining a voice response strategy based on dynamic risk assessment as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The storage medium stores a program that, when executed by a processor, is used to implement a method for determining a voice response strategy based on dynamic risk assessment as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Intelligent supervision method and system based on big data
CN119476946A
Industrial large model content security capability construction method, apparatus and device, and medium
CN120179786A