Voice instruction verification method and device, electronic equipment and storage medium

CN121641037BActive Publication Date: 2026-08-07CHINA UNICOM ONLINE INFORMATION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA UNICOM ONLINE INFORMATION TECHNOLOGY CO LTD
Filing Date
2025-12-08
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

当前语音认证方法依赖声纹特征比对与简单活体检测,未构建意图与行为模式的逻辑验证体系,难以抵御高保真语音重放、人工智能语音合成等新型攻击

Benefits of technology

[0051] According to a fourth aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are configured to cause the computer to perform the method described in the first aspect above.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121641037B_ABST
    Figure CN121641037B_ABST
Patent Text Reader

Abstract

The present disclosure provides a voice instruction verification method and device, electronic equipment and storage medium, and relates to the technical field of voice recognition, which analyzes user voice instructions to obtain instruction intent information and associated entity information and generates structured instructions, simultaneously extracts acoustic features of user voice instructions for emotion state analysis to generate emotion state indicators, can also query a preset knowledge graph to obtain user historical behavior patterns and social relationship data and collect real-time environmental context information, and then calculates a risk level based on these multi-dimensional information through a hierarchical risk assessment engine, and triggers corresponding security protection actions according to the risk level and starts a collaborative monitoring mechanism for a high risk level, so that the problem that the prior art only relies on acoustic features of voice signals themselves for verification, does not integrate and construct a multi-dimensional verification system, and lacks a collaborative monitoring mechanism requiring users to perform unnatural interaction actions can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of speech recognition technology, and in particular to a voice command verification method and apparatus, electronic device and storage medium. Background Technology

[0002] Voice recognition and security authentication technologies are the core support for modern intelligent interactive systems, and are widely used in highly sensitive scenarios such as financial transactions, government services, and smart homes. This technology constructs a multi-layered protection system for voice commands, covering the entire process from signal acquisition to intent parsing, through the collaborative operation of voiceprint recognition, liveness detection, and contextual verification. This includes key steps such as acoustic feature extraction, biometric comparison, and behavioral pattern analysis.

[0003] With the development of deepfake technology, existing solutions have systemic flaws in their signal-level defense capabilities, urgently requiring breakthroughs in the limitations of traditional physical feature verification frameworks. Current voice authentication methods rely on voiceprint feature comparison and simple liveness detection, failing to construct a logical verification system based on intent and behavioral patterns, making them vulnerable to new attacks such as high-fidelity voice playback and AI-generated speech. Furthermore, existing liveness detection schemes require users to perform unnatural interactive actions, inherently contradicting the convenience sought in voice interaction. More critically, related technologies lack a systematic integration of multi-dimensional contextual information such as user social relationships and geographical location, failing to identify abnormal associations through a three-dimensional verification mechanism of intent, context, and psychology. This not only creates serious security vulnerabilities but also leads to a dual risk of false positives and false negatives when the system responds to coercive scenarios. Summary of the Invention

[0004] This disclosure provides a voice command verification method, apparatus, electronic device, and storage medium. Its main objective is to at least partially address one of the technical problems in related technologies.

[0005] According to a first aspect of this disclosure, a voice command verification method is provided, comprising:

[0006] Parse user voice commands to obtain command intent information and associated entity information, and generate structured commands based on the command intent information and the entity information;

[0007] The acoustic features of the user's voice commands are extracted and emotional state analysis is performed to generate emotional state indicators.

[0008] The system queries a preset knowledge graph to obtain users' historical behavior patterns and social relationship data, collects real-time environmental context information, and queries and matches preset security guardrail policies. If the structured instruction violates the security guardrail policy, the authentication process is terminated and a preset blocking action is executed or a collaborative monitoring process is initiated.

[0009] If the safety barrier policy is not violated or does not exist, the risk level is calculated by a hierarchical risk assessment engine based on the structured instructions, the emotional state indicators, the historical behavior patterns, the social relationship data, and the real-time environmental context information.

[0010] The corresponding safety protection actions are triggered according to the risk level, wherein a collaborative monitoring mechanism is activated for high-risk levels.

[0011] Optionally, parsing user voice commands includes:

[0012] Based on natural language understanding, the user's voice commands are semantically classified to identify command intents that include at least financial transactions, device control, and account management types.

[0013] Extract entity objects corresponding to the intent of the instruction from the instruction text corresponding to the user's voice command, and perform category verification on the entity objects.

[0014] Optionally, the step of extracting the acoustic features of the user's voice commands and performing emotional state analysis to generate emotional state indicators includes:

[0015] Calculate the acoustic feature parameters of the speech signal corresponding to the user's voice command. The acoustic feature parameters include at least one of the following: fundamental frequency dynamic range, energy variation amplitude, harmonic noise ratio, and speech rate stability.

[0016] Based on the acoustic feature parameters, an emotional state index representing the user's psychological tension and coercion is generated through a nonlinear mapping model.

[0017] Optionally, the step of querying a preset knowledge graph to obtain user historical behavior patterns and social relationship data, collecting real-time environmental context information, and querying and matching preset security guardrail strategies includes:

[0018] Using user identifiers and entity objects as query keys, retrieve historical interaction statistics and social association strength between users and entity objects, and analyze and generate the user's historical behavior patterns and social relationship data;

[0019] The current geographic location coordinates, timestamp, and network connection type are obtained through the device positioning module, system clock, and network interface as the real-time environmental context information.

[0020] Using the target user's user identifier as an index, access the guardian role information associated with the target user and the safety fence policy pre-configured by the guardian in the preset knowledge graph.

[0021] Optionally, the calculation of risk level through the hierarchical risk assessment engine includes:

[0022] A rule engine is used to quickly filter and quantify the input structured instructions, emotional state indicators, historical behavior patterns, social relationship data, and real-time environmental context information to generate a preliminary risk value.

[0023] When the initial risk value falls within a preset uncertainty range, the semantic reasoning model is invoked to perform contextual analysis and output the corrected risk level.

[0024] Optional, also includes:

[0025] After the voice command is executed, the relation edge attributes and user behavior pattern statistics in the preset knowledge graph are updated based on the data from this interaction.

[0026] According to a second aspect of this disclosure, a voice command verification device is provided, comprising:

[0027] The parsing unit is used to parse user voice commands to obtain command intent information and associated entity information, and generate structured commands based on the command intent information and the entity information.

[0028] An extraction unit is used to extract the acoustic features of the user's voice commands and perform emotional state analysis to generate emotional state indicators.

[0029] The query unit is used to query a preset knowledge graph to obtain the user's historical behavior patterns and social relationship data, and collect real-time environmental context information. It also queries and matches a preset security guardrail policy. If the structured instruction violates the security guardrail policy, the authentication process is terminated and a preset blocking action is executed or a collaborative monitoring process is initiated.

[0030] The calculation unit is used to calculate the risk level based on the structured instructions, the emotional state indicators, the historical behavior patterns, the social relationship data, and the real-time environmental context information, through a hierarchical risk assessment engine, if the safety fence policy is not violated or does not exist.

[0031] The triggering unit is used to trigger corresponding safety protection actions according to the risk level, wherein a collaborative monitoring mechanism is preset to be activated at a high risk level.

[0032] Optionally, the parsing unit is also used for:

[0033] Based on natural language understanding, the user's voice commands are semantically classified to identify command intents that include at least financial transactions, device control, and account management types.

[0034] Extract entity objects corresponding to the intent of the instruction from the instruction text corresponding to the user's voice command, and perform category verification on the entity objects.

[0035] Optionally, the extraction unit is also used for:

[0036] Calculate the acoustic feature parameters of the speech signal corresponding to the user's voice command. The acoustic feature parameters include at least one of the following: fundamental frequency dynamic range, energy variation amplitude, harmonic noise ratio, and speech rate stability.

[0037] Based on the acoustic feature parameters, an emotional state index representing the user's psychological tension and coercion is generated through a nonlinear mapping model.

[0038] Optionally, the query unit is also used for:

[0039] Using user identifiers and entity objects as query keys, retrieve historical interaction statistics and social association strength between users and entity objects, and analyze and generate the user's historical behavior patterns and social relationship data;

[0040] The current geographic location coordinates, timestamp, and network connection type are obtained through the device positioning module, system clock, and network interface as the real-time environmental context information.

[0041] Using the target user's user identifier as an index, access the guardian role information associated with the target user and the safety fence policy pre-configured by the guardian in the preset knowledge graph.

[0042] Optionally, the computing unit is also used for:

[0043] A rule engine is used to quickly filter and quantify the input structured instructions, emotional state indicators, historical behavior patterns, social relationship data, and real-time environmental context information to generate a preliminary risk value.

[0044] When the initial risk value falls within a preset uncertainty range, the semantic reasoning model is invoked to perform contextual analysis and output the corrected risk level.

[0045] Optional, also includes:

[0046] The update unit is used to update the relation edge attributes and user behavior pattern statistics in the preset knowledge graph based on the interaction data after the voice command is executed.

[0047] According to a third aspect of this disclosure, an electronic device is provided, comprising:

[0048] At least one processor; and

[0049] A memory communicatively connected to the at least one processor; wherein,

[0050] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method described in the first aspect above.

[0051] According to a fourth aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are configured to cause the computer to perform the method described in the first aspect above.

[0052] According to a fifth aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the method described in the first aspect above.

[0053] The voice command verification method, apparatus, electronic device, and storage medium disclosed herein parse user voice commands to obtain command intent information and associated entity information, and generate structured commands. Simultaneously, it extracts the acoustic features of user voice commands for emotional state analysis to generate emotional state indicators. It queries a preset knowledge graph to obtain user historical behavior patterns and social relationship data, and collects real-time environmental context information. It also queries and matches preset security barrier policies. If the structured command violates the security barrier policy, the authentication process is terminated and a preset blocking action is executed or a collaborative monitoring process is initiated. Furthermore, based on this multi-dimensional information, a layered risk assessment is performed. The engine calculates the risk level and triggers corresponding security protection actions based on the risk level. It also presets a collaborative monitoring mechanism to activate for high-risk levels. Therefore, it can solve the problems in existing technologies that rely solely on the acoustic characteristics of the voice signal itself for verification, fail to integrate command intent, user emotional state, historical behavior patterns, social relationships, and real-time environmental context to build a multi-dimensional verification system, lack a collaborative monitoring mechanism, and require users to perform unnatural interactive actions for liveness detection. This achieves the technical effect of effectively resisting forgery attacks at the signal level, ensuring the convenience of voice interaction, forming a collaborative security protection network at the family level, and realizing the unified technical effect of voice authentication security, convenience, and collaborative protection.

[0054] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0055] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0056] Figure 1 A flowchart illustrating a voice command verification method provided in an embodiment of this disclosure;

[0057] Figure 2This is a schematic diagram of the structure of a voice command verification device provided in an embodiment of the present disclosure;

[0058] Figure 3 A schematic block diagram of an example electronic device provided for embodiments of this disclosure. Detailed Implementation

[0059] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0060] The following description, with reference to the accompanying drawings, outlines a voice command verification method, apparatus, electronic device, and storage medium according to embodiments of the present disclosure.

[0061] Figure 1 This is a flowchart illustrating a voice command verification method provided in an embodiment of the present disclosure.

[0062] like Figure 1 As shown, the method includes the following steps:

[0063] Step 101: Parse the user's voice command to obtain command intent information and associated entity information, and generate a structured command based on the command intent information and the entity information.

[0064] In the embodiments of this disclosure, the voice authentication process first performs semantic-level parsing on the acquired user voice commands. The core objective of this parsing is to extract the core command intent information carried by the command, as well as the entity information directly associated with the command intent information. The command intent information clarifies the core purpose of the user initiating the voice command, while the entity information consists of key parameter-like information required to support the realization of that core purpose. After obtaining the command intent information and associated entity information, a structured command with a standardized format is constructed based on their correspondence. This allows the key content of the command to be presented in a form that can be directly recognized, read, and invoked by subsequent processing stages, providing a unified and standardized input basis for subsequent voice authentication-related stages such as risk assessment. As one implementation method, natural language understanding technology can be used to parse the voice commands. For example, the command intent can be determined through intent recognition, entity information can be obtained through entity extraction, and then the two can be integrated into a structured command according to preset rules.

[0065] By semantically parsing and structuring user voice commands, accurate and standardized core command information is provided for subsequent multi-dimensional evaluation of voice authentication. This avoids subsequent processing deviations caused by unstructured command information and ensures the efficient recall of key command content, laying the foundation for improving the accuracy and processing efficiency of the entire voice authentication process.

[0066] Step 102: Extract the acoustic features of the user's voice commands and perform emotional state analysis to generate emotional state indicators.

[0067] In the embodiments of this disclosure, during the voice authentication process, for the acquired user voice commands, relevant features characterizing the user's emotional state are further extracted from their acoustic dimensions. Based on these acoustic features, the user's emotional state when issuing the voice command is analyzed, ultimately generating an emotional state index that can quantitatively reflect the user's real-time psychological state. This index can serve as a key reference for determining whether the user is in a normal psychological state in the subsequent risk assessment stage, adding a psychological verification dimension to voice authentication. As one implementation method, the underlying acoustic features of the voice commands (such as fundamental frequency variation, energy jitter, harmonic noise ratio, and speech rate) can be calculated and modeled. By quantitatively analyzing these features, it can be determined whether the user is in an abnormal physiological or psychological state such as tension or coercion, thereby generating a corresponding emotional state score as an emotional state index.

[0068] By introducing emotional state analysis, voice authentication is supplemented with psychological verification criteria, breaking through the limitations of traditional voice authentication which relies solely on the physical characteristics of sound signals. It can effectively capture the emotional changes of users caused by abnormal situations (such as coercion), providing support for the subsequent accurate identification of abnormal voice commands, thereby enhancing the voice authentication system's ability to defend against complex security risks.

[0069] Step 103: Query the preset knowledge graph to obtain the user's historical behavior patterns and social relationship data, and collect real-time environmental context information. Query and match the preset security guardrail policy. If the structured instruction violates the security guardrail policy, the authentication process is terminated and a preset blocking action is executed or a collaborative monitoring process is initiated.

[0070] In the embodiments of this disclosure, during the context information construction stage of voice authentication, historical behavior pattern data and social relationship data related to the current user are extracted by accessing a preset knowledge graph. The knowledge graph, as a carrier storing the user's past interaction patterns and interpersonal relationship information, provides historical behavior pattern data reflecting the user's long-term operational habits, and social relationship data reflecting the association attributes between the user and the object related to the command. Both provide historical reference for judging the rationality of the command. Simultaneously, real-time environmental context information directly related to the current voice command triggering scenario is collected. This information reflects the scenario characteristics of the current interaction and, together with the aforementioned historical data, constitutes a user interaction context covering both historical and real-time dimensions, providing comprehensive contextual support for subsequent risk assessment. As one implementation method, the user identifier and command-related entity identifier in the structured command can be used as indexes to query the knowledge graph to obtain data such as the user's social relationship with related entities, historical interaction frequency, and monetary habits. Simultaneously, real-time environmental context information such as the current geographical location, time, and network connection method is collected through the device interface.

[0071] By integrating historical contextual data and real-time environmental information from the knowledge graph, the system overcomes the shortcomings of traditional voice authentication, which relies solely on voice signals and lacks multi-dimensional contextual associations. This allows subsequent risk assessments to combine user history and current scenario for comprehensive judgment, avoiding misjudgments due to limited information and providing crucial data support for improving the accuracy and comprehensiveness of voice authentication. After completing the preset knowledge graph query and real-time contextual information collection, and before entering the hierarchical risk assessment engine, the system performs a high-priority pre-check of security guardrail policies. This check aims to perform rapid and mandatory compliance verification of current structured commands based on semantic-level security rules pre-set by the guardian, enabling immediate interception of extremely high-risk or obviously abnormal commands and forming a proactive defense barrier. Specifically, the system uses the current user's unique identifier as an index to access the guardian role information associated with the user and the set of security guardrail policies pre-configured by the guardian, stored in the preset knowledge graph. The security barrier policy is a constraint rule based on semantic logic. Typical forms include, but are not limited to: absolute limits on single or daily transaction amounts; lists of specific entities prohibited from transactions or operations (such as unfamiliar receiving accounts or high-risk merchants); specific time periods or geographical locations where commands are prohibited; and whitelists of trusted operations that allow seamless access. The system quickly matches and logically judges the key elements in the structured command (such as command intent, entity objects, and transaction amount) with all relevant security barrier policies retrieved. If the matching result indicates that the current command violates any security barrier policy (e.g., the transfer amount exceeds a preset limit, or the recipient is on a prohibited list), the entire voice authentication process is immediately terminated. The system will not proceed to subsequent tiered risk assessment but will directly execute the preset blocking action bound to the violation policy. This could include returning a clear rejection prompt to the user, simply logging without performing any operation, or, more importantly, automatically triggering a collaborative monitoring process, instantly sending a warning notification containing details of the current command and the reason for the violation to the terminal device of the preset guardian, requesting remote intervention and decision-making by the guardian. If the matching result indicates that the current instruction does not violate any security fence policies, or the current user has not configured any fence policies, the authentication process continues and proceeds to the subsequent tiered risk assessment steps.

[0072] By introducing a pre-emptive check using a safety barrier strategy, a rapid decision-making threshold defined by user trust relationships (guardians) is established before complex, multi-dimensional risk assessments. This allows for near-zero-latency blocking of known high-risk patterns, significantly improving the system's response speed and certainty in defending against pre-defined risk scenarios. This not only strengthens protection for groups such as the elderly and children, achieving the "proactive defense through pre-configured strategies" emphasized in the technical solution, but also ensures that in most compliant and normal interactions, this check step is seamless and silent, without affecting the user experience, thus achieving a better balance between security and convenience.

[0073] Step 104: If the safety barrier policy is not violated or does not exist, the risk level is calculated by a hierarchical risk assessment engine based on the structured instructions, the emotional state indicators, the historical behavior patterns, the social relationship data, and the real-time environmental context information.

[0074] In the embodiments of this disclosure, during the risk assessment stage of voice authentication, the previously acquired structured instructions, emotional state indicators, user historical behavior patterns, social relationship data, and real-time environmental context information are first integrated. This multi-dimensional information is used as a unified input and imported into a hierarchical risk assessment engine for comprehensive calculation to determine the risk level. The hierarchical risk assessment engine achieves accurate risk judgment through a phased assessment logic. First, based on preset basic assessment rules, a preliminary quantitative analysis of the multi-dimensional information is performed to filter out clearly defined low-risk and high-risk situations. For risk situations that fall within a fuzzy range after the preliminary assessment, further refined judgment is achieved through deeper reasoning analysis, ultimately outputting a unique risk level. This ensures that the assessment covers both efficient judgment in basic scenarios and accurate recognition in complex scenarios. As one implementation method, a preliminary risk score can be obtained by scoring the multi-dimensional information according to preset rules using a fast rule filter. If the preliminary risk score falls within a preset fuzzy range, a large language model is invoked for deep reasoning to generate a final risk score, thereby classifying the risk level.

[0075] By integrating multi-dimensional information and employing a hierarchical evaluation logic, the system avoids the judgment bias caused by relying on single pieces of information in traditional risk assessments. This ensures the efficiency of assessments in basic scenarios while improving the accuracy of risk assessments in complex scenarios. It provides a reliable risk basis for subsequent targeted triggering of security protection actions, effectively enhancing the comprehensiveness and robustness of voice authentication risk assessment.

[0076] Step 105: Trigger corresponding safety protection actions according to the risk level, wherein a collaborative monitoring mechanism is activated for a high-risk level.

[0077] In the embodiments of this disclosure, during the security response phase of voice authentication, based on the risk level calculated by the hierarchical risk assessment engine, protective measures adapted to the risk level are matched and triggered from a preset set of security protection actions. This achieves differentiated security handling of voice commands with different risk levels, ensuring both security and interaction efficiency. Specifically, for preset high-risk levels, a collaborative monitoring mechanism is activated. This mechanism, by associating with a preset monitoring subject, allows the monitoring subject to participate in the security decision-making process for high-risk commands, forming a collaborative security protection closed loop between the user and the monitoring subject, avoiding the limitations of single-dimensional security judgment. As one implementation method, it can be preset that low-risk levels correspond to seamless permission for command execution, medium-risk levels require local secondary biometric authentication, and high-risk levels directly refuse execution. If a monitoring scenario is involved under high-risk levels, a risk warning or collaborative authorization request can be sent to the monitoring subject via network communication, allowing the monitoring subject to remotely confirm before executing the final action.

[0078] By matching risk levels with security protection actions, the accuracy of voice authentication security responses is achieved, avoiding excessive protection from affecting the interactive experience. At the same time, the activation of the collaborative monitoring mechanism under high-risk levels fills the gap in traditional technology's lack of collaborative security protection at the family level, effectively improving the protection capabilities for users with weaker safety awareness, such as the elderly and children, and further strengthening the security coverage and humanistic care attributes of the voice authentication system.

[0079] The voice command verification method disclosed herein parses user voice commands to obtain command intent information and associated entity information, and generates structured commands. Simultaneously, it extracts the acoustic features of user voice commands for emotional state analysis to generate emotional state indicators. It queries a preset knowledge graph to obtain user historical behavior patterns and social relationship data, and collects real-time environmental context information. It also queries and matches preset security guardrail policies. If the structured command violates the security guardrail policy, the authentication process is terminated and a preset blocking action is executed or a collaborative monitoring process is initiated. Based on this multi-dimensional information, a hierarchical risk assessment engine calculates the risk level, triggers corresponding security protection actions according to the risk level, and a preset collaborative monitoring mechanism is initiated for high-risk levels. Therefore, it can solve the problems of existing technologies that rely solely on the acoustic features of the voice signal itself for verification, fail to integrate command intent, user emotional state, historical behavior patterns, social relationships, and real-time environmental context to construct a multi-dimensional verification system, lack a collaborative monitoring mechanism, and require users to perform unnatural interactive actions for liveness detection. This achieves the technical effect of effectively resisting signal-level forgery attacks, ensuring the convenience of voice interaction, forming a family-level collaborative security protection network, and realizing the unified technical effect of voice authentication security, convenience, and collaborative protection.

[0080] As a specific implementation of this disclosure, based on the basic solution, the parsing of user voice commands is further defined as follows: semantically classifying the user voice commands based on natural language understanding to identify command intents that include at least financial transactions, device control, and account management types; extracting entity objects corresponding to the command intents from the command text corresponding to the user voice commands, and performing category verification on the entity objects.

[0081] Specifically, when parsing user voice commands, the user's voice command is first converted into corresponding command text using speech-to-text technology. Then, the natural language understanding module integrated into the smart device performs semantic classification on the command text. This natural language understanding module is equipped with a specially trained intent recognition model. The training data contains a large number of text samples labeled with command intent types, covering core scenarios such as financial transactions (e.g., transfers, credit card repayments, fund purchases), device control (e.g., smart home light switches, air conditioner temperature adjustments, appliance mode switching), and account management (e.g., account password resets, login authorization settings, account binding / unbinding). The model analyzes the semantic relationships, key feature words, and sentence structure logic in the command text to output the command intent type that matches the current voice command, ensuring that the recognition result includes at least the above three types of command intents. After determining the command intent, the natural language understanding module further... The entity extraction function is activated to extract entity objects corresponding to the current intent from the instruction text. For example, under the intent of financial transaction, the payee's account information and transaction amount are extracted; under the intent of device control, the target device name and control parameter threshold are extracted; and under the intent of account management, the account identifier to be operated and the operation permission level are extracted. Then, the preset entity category verification rules are called to verify the category of the extracted entity objects. For example, it verifies whether the transaction amount is valid data that meets the numerical format requirements, whether the target device name exists in the system's preset list of controllable devices, and whether the account identifier matches the preset account coding standard. Only when the entity object passes the category verification will it enter the structured instruction generation process.

[0082] By accurately classifying semantics to clarify the type of instruction intent, a basis is provided for formulating differentiated processing strategies for different risk levels (such as high-risk financial transactions and low-risk equipment control). At the same time, the category verification of entity objects can filter invalid or abnormal entity data, avoiding deviations in the subsequent authentication process due to incorrect entity information. This effectively improves the accuracy and reliability of the instruction parsing process and lays the foundation for the accuracy of subsequent risk assessment.

[0083] As a specific implementation of this disclosure, based on the basic scheme, the extraction of acoustic features of the user's voice command and the analysis of emotional state to generate an emotional state index are further defined, including: calculating acoustic feature parameters of the voice signal corresponding to the user's voice command, wherein the acoustic feature parameters include at least one of fundamental frequency dynamic range, energy variation amplitude, harmonic noise ratio, and speech rate stability; and generating the emotional state index characterizing the user's psychological tension and stress state through a nonlinear mapping model based on the acoustic feature parameters.

[0084] Specifically, when extracting acoustic features from user voice commands and performing emotional state analysis, the original speech signal corresponding to the user's voice command is first preprocessed. Noise interference is reduced through pre-emphasis, framing, and windowing operations. Then, signal processing algorithms such as short-time Fourier transform and linear predictive coding are used to calculate acoustic feature parameters. Specifically, when calculating the fundamental frequency dynamic range, the fundamental frequency value of each frame of the speech signal is obtained through a fundamental frequency extraction algorithm (such as the YIN algorithm). The maximum and minimum values ​​of the fundamental frequency across all frames of the speech signal are statistically analyzed, and the difference between the two values ​​represents the fundamental frequency dynamic range. When calculating the energy variation amplitude, the short-time energy of each frame of the speech signal is calculated, and the dispersion of the short-time energy across all frames is calculated using variance or range formulas to obtain the energy variation amplitude. When calculating the harmonic noise ratio, the harmonic components and noise components in the speech signal are separated, and the harmonic noise ratio is determined by the ratio of their energy. When calculating speech rate stability, the speech signal is first processed by speech endpoint detection and word segmentation. The time interval between adjacent words is statistically analyzed, and the standard deviation of the interval time is calculated. The smaller the standard deviation, the higher the speech rate stability. At least one of the above acoustic feature parameters is selected as the basis for analysis. Subsequently, the calculated acoustic feature parameters are input into a pre-trained nonlinear mapping model. This model is built based on deep learning frameworks (such as LSTM and CNN). The training data is a set of speech samples labeled with the user's psychological tension (such as calm, mild tension, and severe tension) and coercion state (such as no coercion and coercion). The model learns the nonlinear relationship between acoustic feature parameters and psychological state and outputs an emotional state index quantified from 0 to 100. The higher the score, the higher the user's psychological tension or the greater the possibility of a coercive state. This score is used as the emotional state index.

[0085] By selecting acoustic feature parameters that are strongly correlated with users' emotional states, the basic data for emotion analysis is ensured to be targeted and effective. The nonlinear mapping model can accurately fit the correlation between acoustic features and complex psychological states, avoiding the problem that linear models cannot capture subtle differences in emotional changes. The resulting emotion state index can accurately represent users' real-time psychological states, providing a reliable basis for identifying abnormal scenarios such as coercion in subsequent risk assessments, and effectively improving the accuracy of the emotion state analysis process.

[0086] As a specific implementation of this disclosure, based on the basic solution, the method further defines the query of the preset knowledge graph to obtain user historical behavior patterns and social relationship data, and to collect real-time environmental context information, and to query and match preset safety barrier policies. This includes: using user identifiers and entity objects as query keys to retrieve the historical interaction statistical characteristics and social association strength between the user and the entity objects, and analyzing and generating the user historical behavior patterns and social relationship data; obtaining the current geographical location coordinates, timestamp, and network connection type as the real-time environmental context information through the device positioning module, system clock, and network interface; and using the target user's user identifier as an index to access the guardian role information associated with the target user in the preset knowledge graph and the safety barrier policies pre-configured by the guardian.

[0087] Specifically, the system integrates three key data preparation steps—knowledge graph query, real-time context collection, and security guardrail strategy matching—into an efficient and collaborative contextual information construction and security screening operation. First, the system uses the user identifier (e.g., user ID) obtained after parsing the current voice command and the entity object (e.g., receiving account, target device) as the joint query key to concurrently access the preset knowledge graph. Based on this key, the graph query engine retrieves and extracts historical interaction statistical features between nodes and edges in the graph, such as the cumulative number of transactions, average transaction amount, and transaction time distribution patterns in financial transaction scenarios, as well as the quantitative value of social association strength calculated based on communication records, kinship annotations, or shared social circles. The system then calls the built-in behavioral pattern analysis model to summarize and generate structured, quantifiable user historical behavioral pattern data based on these statistical features (e.g., "preferring small-amount, high-frequency transfers," "rarely operating devices at night"); simultaneously, based on the strength of social association and the relationship types explicitly stated in the graph (e.g., "relatives," "colleagues," "friends"), corresponding social relationship data is generated. Secondly, to obtain real-time environmental context information, the system concurrently calls the underlying hardware and system interfaces of the smart device: it obtains formatted current geographic coordinates through a device positioning module integrating GPS, BeiDou, or base station positioning; it reads a high-precision system clock to generate a precise timestamp containing year, month, day, hour, minute, and second; and it scans currently active network interfaces to identify and record the specific type of network connection (e.g., "Home Wi-Fi SSID: Home_Network", "Operator: China Mobile 5G"). Finally, specifically for matching security fence policies, the system uses the user identifier of the aforementioned target user as an independent index to perform another targeted query on the preset knowledge graph. This query focuses on the subgraph structure representing the "guardianship" relationship in the graph, locating all guardian nodes connected to the user node through relationship edges such as "hasGuardian" (has a guardian), and then extracting the complete set of security fence policies associated with these guardian nodes, stored in the form of attribute or policy nodes. This set of policies consists of semantic-level rules pre-configured by the guardian based on their understanding of the user and their intention to protect them. Examples include "prohibit transfers to accounts in list {A, B, C}", "single transaction amount must not exceed X yuan", and "sensitive operations are only allowed near the geographical location 'home address'". After acquiring all the above data, the system outputs the user's historical behavior patterns, social relationship data, real-time environmental context information, and the security guardrail policy set to the subsequent processing module.

[0088] By integrating and concurrently processing multi-dimensional contextual information queries with security policy retrieval, this implementation significantly improves the system's data preparation efficiency in the early stages of the authentication process and reduces latency that may result from serial queries. More importantly, it ensures that historical patterns, social contexts, real-time scenarios, and manually preset security boundaries relied upon for risk assessment are considered collaboratively within the same logical framework. This design not only provides comprehensive and consistent input for the subsequent layered risk assessment engine but also enables security guardrail policies to accurately match and judge based on the richest contextual information (including newly acquired historical and real-time data), thereby intercepting high-risk instructions that clearly violate user habits, social common sense, or the will of guardians at the first moment, achieving both immediacy and foresight in security protection.

[0089] As a specific implementation of this disclosure, based on the basic solution, the calculation of risk level through the hierarchical risk assessment engine is further defined as follows: using a rule engine to quickly filter and quantify the input structured instructions, emotional state indicators, historical behavior patterns, social relationship data, and real-time environmental context information to generate a preliminary risk value; when the preliminary risk value falls into a preset uncertainty range, a semantic reasoning model is invoked to perform contextual association analysis and output a corrected risk level.

[0090] Specifically, when calculating the risk level using the hierarchical risk assessment engine, the rule engine module is first activated. This module contains multiple sets of quantitative scoring rules designed for voice authentication scenarios. For example, the rules include: "If the entity object in the structured command does not appear in the user's historical behavior pattern (i.e., the first interaction), the risk score increases by a preset value; if the network connection type in the real-time environmental context information is a trusted network labeled with a knowledge graph (such as home default Wi-Fi), the risk score decreases by a preset value; if the emotional state indicator indicates that the user has a tendency to be tense or coerced, the risk score increases by a preset value; if social relationship data shows that the entity object and the user have a high degree of correlation (such as relatives or frequently contacted persons), the risk score decreases by a preset value," etc. The rule engine matches the input structured command, emotional state indicator, historical behavior pattern, social relationship data, and real-time environmental context information with the above preset rules one by one and accumulates the scores according to the weights corresponding to the rules, finally generating a preliminary risk value ranging from 0 to 100 points. Next, the system determines whether the preliminary risk value falls within a preset uncertainty range (e.g., 30-90 points). If it does not fall within this range, the level corresponding to the preliminary risk value is directly determined as the risk level. If it falls within this range, a pre-trained semantic reasoning model (e.g., a large language model) is invoked. This model receives the above multi-dimensional information as input and, through analyzing the logical relationships between the information (e.g., the comprehensive reasoning that "the transaction amount is higher than the historical average, but the geographical location is the user's permanent residence, and the recipient is a highly related relative with a normal emotional state"), adjusts and optimizes the preliminary risk value and outputs the corrected risk level. During the correction process, it also refers to historical similar situation judgment cases stored in the knowledge graph to improve the reliability of the reasoning.

[0091] By using a rule engine to quickly filter and quantify multi-dimensional information, the efficiency of risk assessment is ensured, and response delays caused by complex calculations are avoided. At the same time, semantic reasoning models are used to handle the uncertain range of the initial risk value, and the risk level is corrected through deep contextual association analysis. This effectively solves the problem that single rule judgment is difficult to deal with complex scenarios, significantly improves the accuracy of risk level calculation, and provides a more reliable basis for the reasonable triggering of subsequent security protection actions.

[0092] As a specific implementation of this disclosure, based on the basic solution, the embodiment of this disclosure further includes: after the voice command is executed, updating the relation edge attributes and user behavior pattern statistics in the preset knowledge graph based on the interaction data.

[0093] Specifically, after the voice command is executed according to the corresponding security protection action, the system immediately initiates the update process of the preset knowledge graph. First, key interaction data is extracted from the entire process data of this voice authentication interaction. This data includes the structured command generated this time (including command intent and entity object information), emotional state indicators, real-time environmental context information (such as geographical location and network connection type at the time of interaction), and the timestamp of command execution completion. Subsequently, the system locates the relationship edges between user nodes and entity objects involved in this interaction within the pre-defined knowledge graph (e.g., the "transfer to" relationship edge in a financial transaction scenario, or the "control" relationship edge in a device control scenario). It then writes the extracted interaction data, such as timestamps, emotional state indicators, and geographical locations, into the knowledge graph as new or updated attributes of these relationship edges, overwriting or supplementing existing historical attribute records. Simultaneously, the system calls the built-in statistical analysis module to update user behavior pattern statistics based on the interaction data. For example, under the financial transaction intent, it recalculates the cumulative interaction frequency and the latest average transaction amount between the user and the entity object, or updates the interaction frequency statistics for specific time periods (e.g., weekday evenings) and specific geographical locations (e.g., home areas). Under the device control intent, it updates the user's monthly number of device operations and the statistical percentage of commonly used operation parameters. The updated statistics are then synchronously written into the attribute fields associated with the user behavior pattern in the knowledge graph, completing the entire knowledge graph update operation.

[0094] By updating the relational edge attributes and user behavior pattern statistics of the knowledge graph in a timely manner after each voice command is executed, the knowledge graph can continuously reflect the latest user interaction habits and relationships, avoiding deviations in subsequent risk assessments due to data lag in the knowledge graph. At the same time, the updated user behavior pattern statistics can provide a more accurate reference standard for identifying abnormal interactions (such as transactions deviating from the historical average amount or device operation during infrequent periods), further improving the accuracy and reliability of subsequent voice authentication risk assessments.

[0095] It should be noted that the embodiments of this disclosure may include multiple steps. For ease of description, these steps are numbered, but these numbers are not a limitation on the execution time slots or execution order between the steps; these steps can be implemented in any order, and the embodiments of this disclosure do not limit this.

[0096] Corresponding to the aforementioned voice command verification method, this disclosure also proposes a voice command verification device. Since the device embodiments of this disclosure correspond to the aforementioned method embodiments, details not disclosed in the device embodiments can be referred to the aforementioned method embodiments, and will not be repeated here.

[0097] Figure 2This is a schematic diagram of the structure of a voice command verification device provided in an embodiment of the present disclosure, as shown below. Figure 2 As shown, it includes:

[0098] The parsing unit 21 is used to parse the user's voice command to obtain command intent information and associated entity information, and generate a structured command based on the command intent information and the entity information.

[0099] Extraction unit 22 is used to extract the acoustic features of the user's voice command and perform emotional state analysis to generate emotional state indicators;

[0100] The query unit 23 is used to query a preset knowledge graph to obtain the user's historical behavior patterns and social relationship data and collect real-time environmental context information, and query to match the preset security guardrail policy. If the structured instruction violates the security guardrail policy, the authentication process is terminated and a preset blocking action is executed or a collaborative monitoring process is initiated.

[0101] The calculation unit 24 is used to calculate the risk level based on the structured instructions, the emotional state indicators, the historical behavior patterns, the social relationship data, and the real-time environmental context information, through a hierarchical risk assessment engine, if the safety fence policy is not violated or does not exist.

[0102] Triggering unit 25 is used to trigger corresponding safety protection actions according to the risk level, wherein a collaborative monitoring mechanism is preset to be activated at a high risk level.

[0103] The voice command verification device disclosed herein parses user voice commands to obtain command intent information and associated entity information, and generates structured commands. Simultaneously, it extracts the acoustic features of user voice commands for emotional state analysis to generate emotional state indicators. It queries a preset knowledge graph to obtain user historical behavior patterns and social relationship data, and collects real-time environmental context information. It also queries and matches preset security guardrail policies. If the structured command violates the security guardrail policy, the authentication process is terminated and a preset blocking action is executed or a collaborative monitoring process is initiated. Based on this multi-dimensional information, a hierarchical risk assessment engine calculates the risk level, triggers corresponding security protection actions according to the risk level, and initiates a collaborative monitoring mechanism for preset high-risk levels. Therefore, it can solve the problems of existing technologies that rely solely on the acoustic features of the voice signal itself for verification, fail to integrate command intent, user emotional state, historical behavior patterns, social relationships, and real-time environmental context to construct a multi-dimensional verification system, lack a collaborative monitoring mechanism, and require users to perform unnatural interactive actions for liveness detection. This achieves the technical effect of effectively resisting signal-level forgery attacks, ensuring the convenience of voice interaction, forming a collaborative security protection network at the family level, and realizing the unified technical effect of voice authentication security, convenience, and collaborative protection.

[0104] Furthermore, in one possible implementation of this embodiment, the parsing unit 21 is also used for:

[0105] Based on natural language understanding, the user's voice commands are semantically classified to identify command intents that include at least financial transactions, device control, and account management types.

[0106] Extract entity objects corresponding to the intent of the instruction from the instruction text corresponding to the user's voice command, and perform category verification on the entity objects.

[0107] Furthermore, in one possible implementation of this embodiment, the extraction unit 22 is also used for:

[0108] Calculate the acoustic feature parameters of the speech signal corresponding to the user's voice command. The acoustic feature parameters include at least one of the following: fundamental frequency dynamic range, energy variation amplitude, harmonic noise ratio, and speech rate stability.

[0109] Based on the acoustic feature parameters, an emotional state index representing the user's psychological tension and coercion is generated through a nonlinear mapping model.

[0110] Furthermore, in one possible implementation of this embodiment, the query unit 23 is also used for:

[0111] Using user identifiers and entity objects as query keys, retrieve historical interaction statistics and social association strength between users and entity objects, and analyze and generate the user's historical behavior patterns and social relationship data;

[0112] The current geographic location coordinates, timestamp, and network connection type are obtained through the device positioning module, system clock, and network interface as the real-time environmental context information.

[0113] Using the target user's user identifier as an index, access the guardian role information associated with the target user and the safety fence policy pre-configured by the guardian in the preset knowledge graph.

[0114] Furthermore, in one possible implementation of this embodiment, the computing unit 24 is also used for:

[0115] A rule engine is used to quickly filter and quantify the input structured instructions, emotional state indicators, historical behavior patterns, social relationship data, and real-time environmental context information to generate a preliminary risk value.

[0116] When the initial risk value falls within a preset uncertainty range, the semantic reasoning model is invoked to perform contextual analysis and output the corrected risk level.

[0117] Furthermore, in one possible implementation of this embodiment, such as Figure 2 As shown, it also includes:

[0118] The update unit 26 is used to update the relation edge attributes and user behavior pattern statistics in the preset knowledge graph based on the interaction data after the voice command is executed.

[0119] It should be noted that the foregoing explanation of the method embodiments also applies to the apparatus of this embodiment, and the principle is the same, so it is not limited in this embodiment.

[0120] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0121] Figure 3 A schematic block diagram of an example electronic device 300 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0122] like Figure 3 As shown, the electronic device 300 includes a computing unit 301, which can perform various appropriate actions and processes based on a computer program stored in ROM (Read-Only Memory) 302 or a computer program loaded from storage unit 308 into RAM (Random Access Memory) 303. The RAM 303 may also store various programs and data required for the operation of the electronic device 300. The computing unit 301, ROM 302, and RAM 303 are interconnected via a bus 304. An I / O (Input / Output) interface 305 is also connected to the bus 304.

[0123] Multiple components in electronic device 300 are connected to I / O interface 305, including: input unit 306, such as keyboard, mouse, etc.; output unit 307, such as various types of displays, speakers, etc.; storage unit 308, such as disk, optical disk, etc.; and communication unit 309, such as network card, modem, wireless transceiver, etc. Communication unit 309 allows electronic device 300 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0124] The computing unit 301 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 301 include, but are not limited to, CPUs (Central Processing Units), GPUs (Graphics Processing Units), various special-purpose AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, DSPs (Digital Signal Processors), and any suitable processor, controller, microcontroller, etc. The computing unit 301 performs the various methods and processes described above, such as the voice command verification method. For example, in some embodiments, the voice command verification method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 308. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 300 via ROM 302 and / or communication unit 309. When the computer program is loaded into RAM 303 and executed by the computing unit 301, one or more steps of the methods described above may be performed. Alternatively, in other embodiments, the computing unit 301 may be configured to perform the aforementioned voice command verification method by any other suitable means (e.g., by means of firmware).

[0125] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application-Specific Standard Products), SOCs (System-on-Chips), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0126] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0127] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, RAM, ROM, EPROM (Electrically Programmable Read-Only Memory) or flash memory, optical fiber, CD-ROM (Compact Disc Read-Only Memory), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0128] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0129] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include LANs (Local Area Networks), WANs (Wide Area Networks), the Internet, and blockchain networks.

[0130] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. A server can be a cloud server, also known as a cloud computing server or cloud host, a hosting product within the cloud computing service system that addresses the shortcomings of traditional physical hosts and VPS (Virtual Private Server) services, such as high management difficulty and weak business scalability. Servers can also be servers for distributed systems or servers incorporating blockchain technology.

[0131] It's important to note that artificial intelligence (AI) is the study of enabling computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily include computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graph technologies.

[0132] The various numerical designations such as "first," "second," etc., used in this disclosure are merely for ease of description and are not intended to limit the scope of the embodiments of this disclosure, nor do they indicate a sequential order.

[0133] At least one of the features described in this disclosure can also be described as one or more, and multiple features can be two, three, four or more, and this disclosure does not impose any limitations. In the embodiments of this disclosure, for a technical feature, the technical features in that technical feature are distinguished by "first", "second", "third", "A", "B", "C" and "D", etc., and there is no sequential order or size order among the technical features described by "first", "second", "third", "A", "B", "C" and "D".

[0134] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0135] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A voice command verification method, characterized in that, include: Parse user voice commands to obtain command intent information and associated entity information, and generate structured commands based on the command intent information and the entity information; The acoustic features of the user's voice commands are extracted and emotional state analysis is performed to generate emotional state indicators. The system queries a preset knowledge graph to obtain users' historical behavior patterns and social relationship data, collects real-time environmental context information, and queries and matches preset security guardrail policies. If the structured instruction violates the security guardrail policy, the authentication process is terminated and a preset blocking action is executed or a collaborative monitoring process is initiated. The process of querying a preset knowledge graph to obtain user historical behavior patterns and social relationship data, collecting real-time environmental context information, and querying and matching preset safety barrier strategies includes: using user identifiers and entity objects as query keys to retrieve historical interaction statistical characteristics and social association strength between users and entity objects, and analyzing and generating the user historical behavior patterns and social relationship data; obtaining the current geographical location coordinates, timestamp, and network connection type as the real-time environmental context information through the device positioning module, system clock, and network interface; and using the target user's user identifier as an index to access the guardian role information associated with the target user in the preset knowledge graph and the safety barrier strategies pre-configured by the guardian. If the safety barrier policy is not violated or does not exist, the risk level is calculated using a hierarchical risk assessment engine based on the structured instructions, the emotional state indicators, the historical behavior patterns, the social relationship data, and the real-time environmental context information. The calculation of the risk level using the hierarchical risk assessment engine includes: using a rule engine to quickly filter and quantify the input structured instructions, emotional state indicators, historical behavior patterns, social relationship data, and real-time environmental context information to generate a preliminary risk value; when the preliminary risk value falls within a preset uncertainty range, a semantic reasoning model is invoked to perform contextual analysis and output a corrected risk level. The corresponding safety protection actions are triggered according to the risk level, wherein a collaborative monitoring mechanism is activated for high-risk levels.

2. The method according to claim 1, characterized in that, The parsing of user voice commands includes: Based on natural language understanding, the user's voice commands are semantically classified to identify command intents that include at least financial transactions, device control, and account management types. Extract entity objects corresponding to the intent of the instruction from the instruction text corresponding to the user's voice command, and perform category verification on the entity objects.

3. The method according to claim 1, characterized in that, The step of extracting the acoustic features of the user's voice commands and performing emotional state analysis to generate emotional state indicators includes: Calculate the acoustic feature parameters of the speech signal corresponding to the user's voice command. The acoustic feature parameters include at least one of the following: fundamental frequency dynamic range, energy variation amplitude, harmonic noise ratio, and speech rate stability. Based on the acoustic feature parameters, an emotional state index representing the user's psychological tension and coercion is generated through a nonlinear mapping model.

4. The method according to claim 1, characterized in that, Also includes: After the voice command is executed, the relation edge attributes and user behavior pattern statistics in the preset knowledge graph are updated based on the data from this interaction.

5. A voice command verification device, characterized in that, include: The parsing unit is used to parse user voice commands to obtain command intent information and associated entity information, and generate structured commands based on the command intent information and the entity information. An extraction unit is used to extract the acoustic features of the user's voice commands and perform emotional state analysis to generate emotional state indicators. The query unit is used to query a preset knowledge graph to obtain the user's historical behavior patterns and social relationship data, and collect real-time environmental context information. It also queries and matches a preset security guardrail policy. If the structured instruction violates the security guardrail policy, the authentication process is terminated and a preset blocking action is executed or a collaborative monitoring process is initiated. The query unit is also used to: use the user identifier and entity object as query keys to retrieve the historical interaction statistical characteristics and social association strength between the user and the entity object, and analyze and generate the user's historical behavior patterns and social relationship data. The device positioning module, system clock, and network interface are used to obtain the current geographical location coordinates, timestamp, and network connection type as the real-time environmental context information; the user identifier of the target user is used as an index to access the guardian role information associated with the target user and the safety fence policy pre-configured by the guardian in the preset knowledge graph; The calculation unit is configured to calculate the risk level based on the structured instructions, the emotional state indicators, the historical behavior patterns, the social relationship data, and the real-time environmental context information, using a hierarchical risk assessment engine, if the safety barrier policy is not violated or does not exist. The calculation unit is also configured to: use a rule engine to quickly filter and quantify the input structured instructions, the emotional state indicators, the historical behavior patterns, the social relationship data, and the real-time environmental context information to generate a preliminary risk value. When the initial risk value falls within a preset uncertainty range, the semantic reasoning model is invoked to perform contextual analysis and output the corrected risk level. The triggering unit is used to trigger corresponding safety protection actions according to the risk level, wherein a collaborative monitoring mechanism is preset to be activated at a high risk level.

6. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-4.

7. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-4.

8. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method according to any one of claims 1-4.

Citation Information

Patent Citations

  • Payment verification method, mobile terminal, and computer-readable storage medium

    CN109064182A

  • Speech recognition authentication method and system based on multi-modal features and dynamic evaluation

    CN120748413A