Intelligent voice interaction and emotion analysis system
By building an emotion-driven adaptive service architecture and combining multimodal emotion recognition and privacy protection, we have solved the shortcomings of traditional voice interaction systems in emotion perception and privacy security, achieved high-precision emotion recognition and data security, and adapted to the dynamic decision-making needs of financial scenarios.
Patent Information
- Application Number
- CN202510681065.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-09-26
AI Technical Summary
Traditional voice interaction systems have defects in emotional perception and are unable to cope with the semantically ambiguous expressions and high-risk transaction scenarios of elderly users. In addition, their privacy protection mechanisms are weak, which easily leads to the risk of identity theft.
A lightweight NLP model based on the Transformer architecture is used to parse voice commands. This is combined with a multimodal emotion recognition module and a dynamic service strategy engine, integrated with voiceprint spectrum analysis and risk-aware decision trees, and deployed with privacy protection modules and edge computing optimization to build an emotion-driven adaptive service architecture.
It achieves high-precision emotional state perception, enhances the system's defense capabilities against disguised emotional attacks, supports exclusive services for elderly users, ensures data security and continuous service availability, and adapts to the dynamic decision-making needs of financial scenarios.
Smart Images

Figure CN120708593A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of self-service terminals, and in particular to an intelligent voice interaction and emotion analysis system. Background Art
[0002] Traditional voice interaction systems rely on keyword matching and fixed speech templates, and have defects such as lack of emotional perception and insufficient ability to understand complex business. They are unable to cope with the dynamic decision-making needs of elderly users with ambiguous semantic expressions or high-risk transaction scenarios; existing technologies lack the ability to integrate multimodal emotional features, and the privacy protection mechanism is weak, which can easily lead to the risk of identity impersonation due to the leakage of voiceprint data. Summary of the Invention
[0003] The purpose of the present invention is to provide an intelligent voice interaction and emotion analysis system to solve the above-mentioned problems in the prior art, thereby solving all or one of the above-mentioned problems in the prior art.
[0004] In order to solve the above technical problems, the specific technical solutions of the present invention are as follows: The present invention provides an intelligent voice interaction and emotion analysis system, comprising: The voice command parsing engine, the multimodal emotion recognition module and the dynamic service strategy engine are used to realize user intention recognition and emotional state perception through the fusion analysis of natural language processing and voiceprint emotion features, and adapt to the adaptive service response in financial scenarios.
[0005] As an improved solution, the voice command parsing engine is deployed with a lightweight NLP model based on the Transformer architecture to support fuzzy semantic error correction and multi-round dialogue context management, and is compatible with dialect pronunciation features and financial professional terminology libraries.
[0006] As an improved solution, the multimodal emotion recognition module integrates voiceprint spectrum analysis, speech rate feature extraction and semantic emotion mapping algorithms, which are used to construct user emotion portraits through a cross-modal attention mechanism, and the recognition accuracy reaches industry-leading levels.
[0007] As an improved solution, a risk-aware decision tree is provided in the dynamic service strategy engine, which is used to automatically trigger a manual review process and a secondary verification mechanism when it is detected that the user's emotional volatility exceeds a threshold or involves a high-risk transaction.
[0008] As an improved solution, the intelligent voice interaction and sentiment analysis system also deploys a user portrait knowledge base, which is used to integrate historical interaction records and biometric data through a federated learning framework to support incremental optimization of the exclusive speech library for the elderly user group.
[0009] As an improved solution, the intelligent voice interaction and emotion analysis system also deploys a multi-channel feedback enhancement engine, which is used to dynamically adjust the interaction response speed and information presentation density by combining the voice response intonation characteristics and touch operation timing data.
[0010] As an improved solution, the intelligent voice interaction and sentiment analysis system is also deployed with a privacy protection enhancement module, which adopts the federated voiceprint embedding technology and the differential privacy noise injection mechanism to ensure the irreversibility of the emotional feature data after desensitization processing.
[0011] As an improved solution, the intelligent voice interaction and sentiment analysis system is also deployed with an edge computing optimization engine, which is used to achieve real-time analysis of voice commands in offline environments through model lightweight distillation technology, and support continuous service availability in network interruption scenarios.
[0012] As an improved solution, the intelligent voice interaction and emotion analysis system is also deployed with a cross-modal adversarial verification module, which is used to simulate abnormal emotional characteristics through generative adversarial networks and enhance the system's defense capabilities against disguised emotional attacks.
[0013] As an improved solution, the intelligent voice interaction and sentiment analysis system is also deployed with a service strategy evolution engine, which is used to continuously optimize the emotional response strategy library based on reinforcement learning algorithms to adapt to financial product iterations and regulatory policy changes.
[0014] The beneficial effects of the technical solution of the present invention are: by constructing an emotion-driven adaptive service architecture, the present invention innovatively integrates voiceprint embedding technology, federated learning and adversarial verification mechanism, thereby breaking through the technical bottlenecks of traditional voice systems in terms of emotion perception accuracy, service strategy flexibility and data security. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0016] Figure 1 It is a schematic diagram of the architecture of the intelligent voice interaction and sentiment analysis system described in an embodiment of the present invention. DETAILED DESCRIPTION
[0017] The preferred embodiments of the present invention are described in detail below with reference to the accompanying drawings so that the advantages and features of the present invention can be more easily understood by those skilled in the art, thereby making a clearer and more precise definition of the protection scope of the present invention.
[0018] In the description of the present invention, it should be noted that the embodiments described in the present invention are only part of the embodiments of the present invention, rather than all of the embodiments; based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative work are within the scope of protection of the present invention.
[0019] The terms "first," "second," and the like in the specification and claims herein and in the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate so that the embodiments of the present invention described herein can be implemented in orders other than those illustrated or described herein. In addition, the terms "including" and "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, apparatus, product, or device comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such process, method, product, or device.
[0020] Embodiment, this embodiment provides an intelligent voice interaction and emotion analysis system, such as Figure 1 Shown, including: The voice command parsing engine, the multimodal emotion recognition module and the dynamic service strategy engine are used to realize user intention recognition and emotional state perception through the fusion analysis of natural language processing and voiceprint emotion features, and adapt to the adaptive service response in financial scenarios.
[0021] As an improved solution, the voice command parsing engine is deployed with a lightweight NLP model based on the Transformer architecture to support fuzzy semantic error correction and multi-round dialogue context management, and is compatible with dialect pronunciation features and financial professional terminology libraries.
[0022] As an improved solution, the multimodal emotion recognition module integrates voiceprint spectrum analysis, speech rate feature extraction and semantic emotion mapping algorithms, which are used to construct user emotion portraits through a cross-modal attention mechanism, and the recognition accuracy reaches industry-leading levels.
[0023] As an improved solution, a risk-aware decision tree is provided in the dynamic service strategy engine, which is used to automatically trigger a manual review process and a secondary verification mechanism when it is detected that the user's emotional volatility exceeds a threshold or involves a high-risk transaction.
[0024] As an improved solution, the intelligent voice interaction and sentiment analysis system also deploys a user portrait knowledge base, which is used to integrate historical interaction records and biometric data through a federated learning framework to support incremental optimization of the exclusive speech library for the elderly user group.
[0025] As an improved solution, the intelligent voice interaction and emotion analysis system also deploys a multi-channel feedback enhancement engine, which is used to dynamically adjust the interaction response speed and information presentation density by combining the voice response intonation characteristics and touch operation timing data.
[0026] As an improved solution, the intelligent voice interaction and sentiment analysis system is also deployed with a privacy protection enhancement module, which adopts the federated voiceprint embedding technology and the differential privacy noise injection mechanism to ensure the irreversibility of the emotional feature data after desensitization processing.
[0027] As an improved solution, the intelligent voice interaction and sentiment analysis system is also deployed with an edge computing optimization engine, which is used to achieve real-time analysis of voice commands in offline environments through model lightweight distillation technology, and support continuous service availability in network interruption scenarios.
[0028] As an improved solution, the intelligent voice interaction and emotion analysis system is also deployed with a cross-modal adversarial verification module, which is used to simulate abnormal emotional characteristics through generative adversarial networks and enhance the system's defense capabilities against disguised emotional attacks.
[0029] As an improved solution, the intelligent voice interaction and sentiment analysis system is also deployed with a service strategy evolution engine, which is used to continuously optimize the emotional response strategy library based on reinforcement learning algorithms to adapt to financial product iterations and regulatory policy changes.
[0030] It should be noted that the above examples are only for explaining the present invention and are not intended to limit the scope of protection of the present invention.
[0031] It should also be understood that in the embodiments herein, the term "and / or" merely describes an association between associated objects, indicating that three possible relationships exist. For example, "A and / or B" could represent: A alone, A and B simultaneously, or B alone. Furthermore, the character " / " in this document generally indicates an "or" relationship between the associated objects.
[0032] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the composition and steps of each example according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this document.
[0033] In the several embodiments provided herein, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices, or units, or can be an electrical, mechanical, or other form of connection.
[0034] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the embodiments herein.
[0035] In addition, the functional units in the various embodiments herein may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0036] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this article is essentially or the part that contributes to the existing technology, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of this article. The aforementioned storage medium includes: various media that can store program code, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0037] The above descriptions are merely embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention's description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. An intelligent voice interaction and sentiment analysis system, characterized in that: include: The voice command parsing engine, the multimodal emotion recognition module and the dynamic service strategy engine are used to realize user intention recognition and emotional state perception through the fusion analysis of natural language processing and voiceprint emotion features, and adapt to the adaptive service response in financial scenarios.
2. The intelligent voice interaction and sentiment analysis system according to claim 1, characterized in that: The voice command parsing engine is equipped with a lightweight NLP model based on the Transformer architecture to support fuzzy semantic error correction and multi-round dialogue context management, and is compatible with dialect pronunciation features and a financial terminology library.
3. The intelligent voice interaction and emotion analysis system according to claim 1, characterized in that: The multimodal emotion recognition module integrates voiceprint spectrum analysis, speech rate feature extraction and semantic emotion mapping algorithms, which are used to construct user emotion portraits through a cross-modal attention mechanism, and the recognition accuracy reaches industry-leading levels.
4. The intelligent voice interaction and emotion analysis system according to claim 1, characterized in that: The dynamic service strategy engine is provided with a risk-aware decision tree, which is used to automatically trigger a manual review process and a secondary verification mechanism when it is detected that the user's emotional volatility exceeds a threshold or involves a high-risk transaction.
5. The intelligent voice interaction and emotion analysis system according to claim 1, characterized in that: The intelligent voice interaction and sentiment analysis system also deploys a user portrait knowledge base, which is used to integrate historical interaction records and biometric data through a federated learning framework to support incremental optimization of the exclusive speech library for the elderly user group.
6. The intelligent voice interaction and emotion analysis system according to claim 1, characterized in that: The intelligent voice interaction and emotion analysis system is also deployed with a multi-channel feedback enhancement engine, which is used to combine the voice response intonation characteristics and touch operation timing data to dynamically adjust the interaction response speed and information presentation density.
7. The intelligent voice interaction and sentiment analysis system according to claim 1, characterized in that: The intelligent voice interaction and sentiment analysis system is also equipped with a privacy protection enhancement module, which uses federated voiceprint embedding technology and differential privacy noise injection mechanism to ensure the irreversibility of emotional feature data after desensitization processing.
8. The intelligent voice interaction and sentiment analysis system according to claim 1, characterized in that: The intelligent voice interaction and sentiment analysis system is also deployed with an edge computing optimization engine, which is used to achieve real-time analysis of voice commands in offline environments through model lightweight distillation technology, and support continuous service availability in network interruption scenarios.
9. The intelligent voice interaction and emotion analysis system according to claim 1, characterized in that: The intelligent voice interaction and emotion analysis system is also equipped with a cross-modal adversarial verification module, which is used to simulate abnormal emotional characteristics through generative adversarial networks and enhance the system's defense capabilities against disguised emotional attacks.
10. The intelligent voice interaction and emotion analysis system according to claim 1, characterized in that: The intelligent voice interaction and sentiment analysis system is also deployed with a service strategy evolution engine, which is used to continuously optimize the emotional response strategy library based on reinforcement learning algorithms to adapt to financial product iterations and regulatory policy changes.