Distributed voice interaction module system and event coordination processing method
By integrating voice interaction units with a central semantic coordination unit through a distributed voice interaction unit system, the problems of cross-module semantic integration and privacy protection in existing voice systems are solved. This achieves an efficient and personalized voice interaction experience with reduced latency, making it suitable for scenarios such as smart security, education and companionship, and emotional toys.
Patent Information
- Application Number
- CN202511178135.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-21
- Publication Date
- 2025-11-14
AI Technical Summary
Most existing voice interaction devices are single-point architectures, lacking cross-module semantic integration capabilities, resulting in fragmented voice interaction experiences. It is difficult to achieve coherent and contextual semantic reasoning and response. Furthermore, existing systems lack semantic event collaborative processing mechanisms, which cannot effectively integrate voice commands and sensing data from multiple modules, making them prone to misjudgments or delays. Moreover, reliance on cloud services poses privacy risks and application limitations.
A distributed voice interaction unit system is provided, comprising a voice interaction unit and a central semantic coordination unit. It has voice wake-up, recognition, semantic understanding and voice response functions, supports semantic event collaborative architecture, integrates voice events and sensing data from multiple units, includes a context construction module and a voice personality proxy mechanism to achieve personalized interaction and context continuation, and has data synchronization and semantic priority ranking logic. It also supports multi-model orchestration and a privacy-preserving semantic summarization layer.
It improves the overall interactive quality, immediacy, and scalability of voice systems deployed in multiple modules, achieves high availability and scalability of cross-scenario voice interaction, reduces response latency, enhances privacy protection, and supports dynamic adjustment of personalized voice personality models and cross-module collaborative decision-making.
Smart Images

Figure CN120954404A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of speech recognition, natural language processing, and human-computer interaction, specifically to a distributed speech interaction unit system and its collaborative processing method for semantic events. This system combines edge speech recognition, adaptive semantic understanding, modular deployment, and a cross-device semantic event coordination mechanism, making it suitable for improving the continuity of semantic understanding, the personalization of responses, and the scalability of the system in multi-terminal speech processing scenarios. Background Technology
[0002] With the rapid development of voice artificial intelligence (AI) and natural language processing (NLP) technologies, voice interfaces have been widely used in smart homes, educational toys, medical companionship, and security monitoring. However, most existing voice interaction devices adopt a single-point architecture, which can only process voice input from a single module and lacks the ability to integrate semantics across modules. This results in a fragmented voice interaction experience, making it difficult to achieve coherent and contextual semantic reasoning and response.
[0003] Furthermore, while distributed voice module deployments offer flexibility, most existing systems lack semantic event collaboration mechanisms. This hinders the effective integration of voice commands and sensor data from multiple modules, and also lacks event arbitration and prioritization logic. This is particularly problematic in smart security, multi-person environments, or public spaces, easily leading to misjudgments or delays. On the other hand, voice devices generally rely on cloud services for semantic parsing, which limits their application in latency- and privacy-sensitive scenarios (such as access control, children's rooms, and medical spaces). Regarding personalized interaction and contextual memory, current voice systems are generally limited by static databases and fixed response sequences. They cannot dynamically adjust based on user tone, historical interactions, and preferences, nor can they accumulate and evolve individual voice personality models, reducing the adaptability and long-term value of the interaction.
[0004] Although context engineering has been applied to large language models (LLMs) to improve contextual reasoning performance, existing designs are mostly device-oriented, lacking cross-module semantic event streaming integration and context memory proxy architecture, as well as flexible collaboration mechanisms and data security layer support between various speech units.
[0005] Therefore, this invention proposes a scalable distributed voice unit interaction system, the core of which includes a voice interaction unit and a central semantic coordination unit, and has the following characteristics: it supports the deployment of voice units in multiple terminal devices, and has voice wake-up, recognition, semantic understanding and voice response functions:
[0006] (a) Provides a semantic event collaboration framework that integrates multi-unit speech events and sensing data.
[0007] (b) It contains a context construction module and a voice personality proxy mechanism to support personalized interaction and context continuation.
[0008] (c) Each voice interaction unit has data synchronization and semantic priority sorting logic, which can be used for real-time decision-making and event fusion.
[0009] (d) A semantic summarization layer that can flexibly bridge multiple language models and provides privacy protection.
[0010] This invention can be applied to various scenarios such as smart security, educational companionship, emotional toys, and long-term care interaction, improving the overall interactive quality, immediacy, and scalability of voice systems under multi-module deployment. Summary of the Invention
[0011] To achieve the above objectives, the present invention provides the following technical solution:
[0012] A distributed voice interaction unit system, comprising:
[0013] (i) At least one voice interaction unit, which can be modularly deployed in one or more terminal devices, each of the voice interaction units comprising:
[0014] (a) Local speech processing module for voice wake-up, speech recognition and semantic parsing;
[0015] (b) Semantic event generation module, used to convert speech input into semantic events with structured fields.
[0016] Semantic Event Object;
[0017] (c) Short-term memory module, used to store recent conversation context and preferences;
[0018] (ii) Central Semantic Event Coordination Unit (AIOrchestration Layer), used to receive signals from various voice interactions.
[0019] The semantic events of a unit, whose functions include:
[0020] (a) Multi-LLM Orchestration Engine, based on latency tolerance,
[0021] Privacy levels, role authorization, regional regulations, and service quality indicators are used to dynamically select cloud or edge language models for inference.
[0022] (b) The Persona Proxy and Memory Synchronization Module triggers the synchronization of short-term and long-term memory and adjusts the voice personality parameters based on the frequency of semantic events, emotional relevance, or explicit user instructions.
[0023] (c) A memory synchronization and version control mechanism is used to maintain consistency of short-term and long-term memory data across multiple voice interaction units;
[0024] (d) Response intent generation module, used to output response text, specified voice units,
[0025] The structured response intent of the speech style parameters and memory update prompts is transmitted back to the corresponding speech interaction unit.
[0026] Preferably, the voice interaction unit has independent local voice interaction capability, and when it detects an event that requires collaborative decision-making, it transmits semantic event objects to the central semantic event coordination unit (AIOrchestration Layer). The semantic event objects include statement intent, emotion index, context association marker and module identification code.
[0027] Preferably, the central semantic event coordination unit (AIOrchestration Layer) further includes an event orchestrator for establishing an event sequence management module based on semantic events to maintain contextual consistency and prioritization across event processing.
[0028] Preferably, the semantic event object includes at least: event type, sentence summary, sentiment index, confidence level, source module identifier, timestamp, context marker, and privacy level marker. These fields serve as the basis for event priority ranking, time sequence alignment, and multi-model orchestration decisions. The semantic event buffer and priority processing module further sorts and arbitrates events based on their urgency, context relevance, source credibility, confidence level, and timestamp differences. In the case of redundant events from multiple sources, it performs deduplication and time sequence alignment according to the confidence level threshold and source priority rules.
[0029] Preferably, the Safe Abstraction Layer includes a semantic masking module, which performs the following steps based on the privacy level label of a semantic event before it is sent to the semantic reasoning module:
[0030] (a) Identify and tag personally identifiable information contained in the event field, such as voiceprint metrics, mood index, geolocation, or identity code;
[0031] (b) Automatically select the corresponding masking method for different privacy levels, including: field deletion, obfuscation, category generalization or dynamic replacement value;
[0032] (c) Generate a semantic event object summary that includes a masking strategy summary and a security level label to facilitate risk adjustment in subsequent model selection.
[0033] Preferably, the Multi-LLM Orchestration Engine includes:
[0034] (a) Policy Selector, used to select and score models based on at least the following factors: event delay tolerance, semantic complexity, data privacy level, role authorization permissions, local regulations and system service quality indicators;
[0035] (b) Model Router: Based on the output of the policy selector, the semantic events are directed to an applicable language model, including but not limited to: edge deployment model, local lightweight model, and cloud-based large language model, and supports dynamic hot switching and context maintenance between heterogeneous models.
[0036] Preferably, the central semantic event coordination unit (AIOrchestration Layer) supports cloud deployment, edge computing configuration, or a hybrid architecture thereof, and can dynamically switch its operating location according to latency requirements and service quality indicators (SLO).
[0037] Preferably, the Persona Proxy includes a set of personalized voice style parameters for the user, including but not limited to speech rate, tone, emotional expression, vocal lexicon preferences, and interactive role settings. The Persona Proxy module includes personalized updates and generation based on the user's voice interaction history, and can dynamically adjust its voice response style based on context evaluation and user interaction history, and record the source module and context conditions of each update to support version tracking, rollback, and weight adjustment strategies.
[0038] Preferably, the memory synchronization module adopts a hybrid event-driven and time-window strategy, and uses a synchronization index including module identification code, context frame identification code, version hash and timestamp to maintain the consistency of short-term and long-term memory among multiple voice interaction units. The memory synchronization module adopts a hybrid event-driven and time-window strategy, and includes a short-term memory module and a long-term memory module, which are used to save the voice interaction process and semantic preference information, respectively. The memory module has a semantic data persistence strategy.
[0039] Preferably, the structured data output by the response intent generation module includes at least the response text, intent type, specified output voice interaction unit, voice style parameters and memory update prompts, and the central semantic event coordination unit generates a decision audit log for each reasoning and decision, recording decision parameters, model selection and masking strategy.
[0040] Preferably, the central semantic event coordination unit supports cross-modal event fusion, including semantically integrating events generated by external modules (e.g., sensors, image systems, third-party data sources) with voice events to improve contextual understanding and decision-making accuracy.
[0041] Preferably, the voice interaction unit and the central semantic event coordination unit (AIOrchestration Layer) can exchange data through any or more of the following communication methods, including but not limited to SubG, Zigbee, Thread, Wi-Fi, and BLE.
[0042] A method for coordinated processing of distributed voice interaction and semantic events includes the following steps:
[0043] (a) The voice interaction unit receives voice input and performs local speech recognition (ASR) and natural language understanding (NLU) to generate a semantic event object;
[0044] (b) Store the semantic event object into the event buffer queue according to its timestamp, context marker and confidence level;
[0045] (c) Perform event sorting, conflict arbitration, redundancy removal and time alignment on the semantic event objects;
[0046] (d) The semantic event object is de-identified through a Safe Abstraction Layer and a privacy level label is added;
[0047] (e) Based on the latency tolerance, privacy level and user role authorization of the event, dynamically select cloud or edge models for semantic reasoning through the Multi-LLM Orchestration Engine;
[0048] (f) Generate a response intent based on the reasoning result, wherein the response intent includes response text, target voice interaction unit identification code and memory update prompt;
[0049] (g) Update the short-term memory and long-term memory modules according to the stated response intent, and adjust the corresponding Persona Proxy parameters.
[0050] Preferably, the generation of the semantic event object includes the following steps:
[0051] (a) Convert the received voice signal into text data;
[0052] (b) Perform semantic parsing on the text data to identify user intent;
[0053] (c) Extracting emotional features, user state, and contextual markers from the semantic parsing process;
[0054] Add a version identifier or hash value to the semantic event object to track data consistency and integrity during cross-module processing.
[0055] Preferably, the Safe Abstraction Layer includes the following processing steps:
[0056] (a) Identify the sensitive fields contained in each semantic event object based on the privacy level label attached to it;
[0057] (b) Perform de-identification processing on the identified sensitive fields, including field deletion, data obfuscation, or label replacement;
[0058] (c) While retaining contextual data and intent parameters that can be used for decision reasoning, ensure that semantic events do not contain identifiable personal privacy information before being transmitted to the semantic reasoning module.
[0059] Preferably, the multi-model orchestration engine includes a policy selector and a model router. The policy selector scores and selects models based on event priority and latency tolerance. The model router guides cloud models, edge models, or hybrid deployment models according to the policy to perform semantic reasoning tasks.
[0060] Preferably, the memory synchronization includes triggering a memory update operation between the short-term memory module and the long-term memory module based on the event frequency, emotional intensity, or explicit instructions from the user, and ensuring that the memory state between different voice interaction units remains consistent through a synchronization consistency index.
[0061] Preferably, the cross-modal event fusion includes: receiving state input from the sensor module, image events generated by the image processing module, and semantic events transmitted by the voice interaction unit; and converging the multi-source events into a unified context framework for integrated processing, so that the central semantic event coordination unit can perform semantic reasoning and response intent generation.
[0062] (a) Fragmented user experience: Most existing voice systems are single-point devices and cannot connect semantics and context.
[0063] (b) Lack of semantic coordination mechanism: Distributed modules cannot coordinate with each other to process voice input, resulting in inconsistent decision-making.
[0064] (c) Model-dependent centralized architecture: Current speech inference relies heavily on cloud-based models, which have high latency and privacy risks.
[0065] (d) Unable to evolve with user growth: lacks memory mechanism and personalized agent, and has a rigid interaction style.
[0066] (e) Lack of context construction process: lack of context engineering to guide large language models to make effective reasoning.
[0067] (f) Lack of flexibility in model binding: Current systems are mostly single LLM bindings, which are difficult to adapt across products and regions.
[0068] Based on the above problems, the present invention proposes the following technical solutions. This system can be flexibly deployed on different terminal devices (such as security gateways, lamps, cameras, smart toys, voice boxes, etc.). Each voice interaction unit has basic voice processing capabilities and can perform semantic integration, context construction and model reasoning through the central semantic event coordination unit, so as to achieve high availability and scalability of cross-scene voice interaction.
[0069] The system of this invention includes the following units and mechanisms:
[0070] (a) Voice Interaction Unit:
[0071] Each module includes audio receiving and playback units, local voice wake-up, automatic speech recognition (ASR), natural language understanding (NLU), and text-to-speech (TTS) capabilities, enabling it to independently process voice commands and generate corresponding responses. The module also features event reporting functionality, transmitting semantic events back to the central semantic event coordination unit, and can store short-term memory data such as user tone preferences, interaction styles, and task status.
[0072] (b) Central Semantic Event Coordination Unit (AIOrchestration Layer):
[0073] This unit is used to receive and process semantic events returned from multiple speech modules. It possesses semantic integration, priority ranking, event arbitration, task decision-making, and cross-module semantic synchronization capabilities, supporting flexible operation in both single-module and multi-module configurations. Internally, it includes a context construction module that dynamically constructs context based on historical semantic events, the user's voice persona proxy, and the current situation to support language model inference.
[0074] (c) Language Model Adaptation and Orchestration Engine (Multi-LLM Orchestration Engine):
[0075] It supports the integration and routing of multiple large language models (such as GPT, Claude, Gemini, etc.). It can automatically select the most suitable model for processing based on product attributes, event attributes, user identity, region, latency tolerance, and data sensitivity. Furthermore, it dynamically adjusts the model strategy based on product characteristics, region, user preferences, and interaction history.
[0076] (d) Personality Memory Module and Persona Proxy:
[0077] The module can create an individual voice personality for each user, including tone, word style, emotional response and learning progress, and can be updated synchronously by the central semantic event coordination unit through the OTA mechanism to achieve consistency of personalized memory and continuous learning across modules.
[0078] (e) Safe Abstraction Layer:
[0079] Converting voice into an abstract semantic event format before data transmission avoids transmitting the original audio, enhances privacy protection, and also helps with cross-module event standardization and semantic processing.
[0080] Module flexible deployment and application scenario scalability
[0081] The system can be applied to various scenarios such as security, education, trendy toys, long-term care, and emotional support. It provides asynchronous voice interaction, cross-device semantic decision-making, and multi-sensor fusion, and has high scalability and platform application potential.
[0082] The "voice interaction unit" described in this invention refers to a functional module capable of receiving, processing, and responding to voice signals. It can be implemented as a hardware module, a software logic element, or a voice processing subsystem embedded in a terminal device. This definition aims to encompass various design and deployment forms and is not limited to a specific physical configuration. Through the above design, this invention can significantly improve the semantic accuracy, real-time response capability, and personalized interaction quality of voice interaction systems, forming a platform-centric scalable voice AI solution and providing a technical architecture with practical deployment value and patent protection potential. This invention, through semantic event objectification, hierarchical coordinated reasoning, and multi-model dynamic adaptation, can effectively improve the voice system's ability to handle multi-module semantic consistency, reduce response latency, enhance personalized interaction accuracy, and control voice privacy levels in distributed deployments. Attached Figure Description
[0083] Figure 1The architecture diagram of the distributed voice interaction system described in this invention is as follows. The system includes one or more voice interaction units (100), a central semantic event coordination unit (200) (AIOrchestrationLayer), and at least one large language model (600) (LLM) server. The semantic event transmission, language model query, and response processing flow between the units are described.
[0084] Figure 2 The internal module composition diagram of the voice interaction unit (100) described in the invention includes an audio input / output module (110), a wake word recognition module (120), a local voice processing module (130) (including ASR / NLU / TTS), a short-term memory module (150), and a wireless communication module (160), which are used to perform local voice event processing and transmit semantic events back to the central semantic event coordination unit (200);
[0085] Figure 3 This is a schematic diagram of a voice interaction process in one embodiment of the present invention. The voice interaction unit (100) receives a voice input unit (110), analyzes the semantics through a local voice processing module (130), generates a semantic event (300), and transmits it to a central semantic event coordination unit (200). The event then sequentially passes through a semantic event cache (305), a context construction module (210), a multi-model orchestration engine (280), and a personalized voice response generation module, generating a voice response and outputting it to the audio output unit (110). During the response process, external language model services can also be invoked, and user memory updates can be performed.
[0086] Figure 4 This is a schematic diagram of the voice interaction system structure of the present invention, showing that the system can operate in a single-module or multi-module collaborative mode. Each voice interaction unit (100) has basic voice processing and short-term memory functions, and can interact with the central semantic event coordination unit (200) through wireless communication to realize semantic event processing and multi-module collaboration;
[0087] Figure 5A : Flowchart for the voice interaction module to process voice input and generate semantic events (300);
[0088] Figure 5B : A schematic diagram of the data interaction process for context construction (210) of the central semantic coordination unit;
[0089] Figure 5C : A schematic diagram of the data flow for semantic inference (320) processing and speech response generation;
[0090] Figure 5D This is a schematic diagram illustrating the process of triggering memory update (390) and module synchronization based on the prompts in the response intent after a semantic response is generated.
[0091] Figure 5E This is a schematic diagram of semantic event coordination processing under a multi-voice module collaborative architecture in one embodiment of the present invention. It shows that after multiple voice interaction units generate a semantic event (300), they construct the context (210) and process the semantic processing module (240) through the central semantic event coordination unit (200), and drive the voice output module (350) to execute personalized voice feedback according to the response intent (330). This process also illustrates the calling and memory push mechanism of the voice personality agent (230).
[0092] Figure 5F This is a flowchart for semantic event priority processing. The voice interaction system receives semantic events (S110), extracts key features (S110), and then performs preliminary sorting (S120), redundant event arbitration (S130), temporal alignment (S140), conditional scoring (S150), and priority level allocation (S160) in sequence. Finally, it determines whether to output based on the conditions (S170), and those that meet the conditions are sent to the coordination layer for processing (S180).
[0093] Figure 6A : This is a "Privacy Image and Model Routing Architecture Diagram" in one embodiment of the present invention, showing that after the semantic event object is masked by the security abstraction layer (260) according to the privacy level (660), an appropriate inference strategy (640) is selected according to the determined privacy level, including edge big model (641), cloud big model (642) or hybrid inference mode (643) to support semantic processing decisions based on privacy sensitivity.
[0094] Figure 6B This is a flowchart of the "Privacy Image and Model Routing Strategy Processing" in one embodiment of the present invention. It describes that after a semantic event object enters the privacy level image module (S210), a mask rule is applied through the security abstraction layer (S220). If the level is less than or equal to Partial, it will further enter the model routing strategy (S240) to select an appropriate inference model; otherwise, it will refuse to be passed to the model inference path to protect privacy.
[0095] Figure 7AThis describes the voice module's voice personality parameter synchronization process. When a change occurs in voice style parameters (such as speech rate and tone) at the module level (S710), it reports to the central coordination unit (S720), which then performs a version comparison (S730). If there is no difference, the process ends (S740). If there is a difference, an update event is generated (S750) and pushed to other voice modules (S760). Each module updates its voice personality agent accordingly (S770), completing the synchronization (S780).
[0096] Figure 7B This is a schematic diagram of an embodiment of the present invention, illustrating that when the voice interaction units (100A-100C) hold different versions of voice personality agents (711A-711C), the central semantic event coordination unit produces a consistent version of the voice personality agent (711D) through the difference merging module (720) and the conflict resolution module (730), and synchronizes it back to each voice module to maintain the consistency of voice style. Detailed Implementation
[0097] To more clearly illustrate the technical content and implementation methods of the present invention, a specific embodiment of the present invention will be described in detail below with reference to the accompanying drawings. Those skilled in the art can easily understand other advantages and effects of the present invention from the attributes disclosed in this specification. See also Figures 1 to 7B This invention proposes a voice interaction system with distributed semantic coordination and personalized voice interaction functions, which includes the following main components.
[0098] like Figure 1 As shown, the present invention provides a distributed voice interaction system, the system architecture of which includes:
[0099] 1. Voice Interaction Unit (100)
[0100] Each voice interaction unit (100) can be deployed as a detachable module or embedded subsystem in one or more terminal devices (e.g., smart gateways, light switches, smart lighting fixtures, toys, cameras, or standalone voice boxes), and includes:
[0101] (a) Audio Input / Output Unit (110): Receives voice signals from the user.
[0102] And play a voice response.
[0103] (b) Wake Word Detection (120): Listen for and identify wake words to trigger subsequent speech recognition.
[0104] (c) Local processing engine (130) (LightASR / NLU / TTS): Performs speech-to-text (ASR), semantic understanding (NLU) and speech synthesis (TTS) on the module side, with basic semantic processing and response capabilities.
[0105] (d) Short-term memory (150): Stores temporary user preferences, tone, context, and task status.
[0106] (e) Wireless Communication Module (160) (RF Communication): Communicates with the central semantic event coordination unit through protocols such as WIFI, BLE, Sub-GHz, Zigbee, Z-Wave, and Thread.
[0107] Each voice interaction unit (100) can independently process voice input, generate semantic events (300), and upload the events wirelessly to the central semantic event coordination unit (200), which then constructs and integrates the context to enhance the cross-device voice interaction experience.
[0108] 2. Central Semantic Event Coordination Unit (200) (AIOrchestration Layer)
[0109] This is the core processing unit of the system of this invention, used to integrate and analyze the semantic events returned by each speech module. It includes:
[0110] (a) Context Construction Module (210): Constructs contextual information based on event history, task status, and personal memories.
[0111] (b) Event Prioritization (220): Sorting and arbitrating events based on time, urgency, or source.
[0112] (c) Persona Proxy (230): Each user has an independent persona configuration to generate personalized voice responses.
[0113] (d) Semantic Processing Module (240) is used to integrate contextual information and semantic events to perform semantic reasoning and response generation.
[0114] (e) Long-term Memory Store (250): Stores historical interaction data, user preferences and language styles for model reasoning and personality development.
[0115] (f) Safe Abstraction Layer (260): Abstracts voice events to avoid transmitting raw audio and improve privacy protection.
[0116] (g) User Account Management (270): Multi-user identification and authorization management module, which supports multi-user login, identification and separate storage of individual data.
[0117] (h) Multi-LLM Orchestration Engine (280): Automatically selects and connects to the appropriate language model (LLM) based on the user context, task nature, and latency requirements.
[0118] (i) Response Generation Module (290): Generates personalized natural language speech responses based on model output and contextual information.
[0119] 3. Language Model Group (600) (LLMs)
[0120] The system can bridge and support multiple LLMs (e.g., (610)GPT, (620)Claude, (630)Gemini, (640)Ernie), and can dynamically select the model for querying and responding based on regional regulations, cybersecurity investigations, or mission characteristics. Figure 2 As shown, in one embodiment of the present invention, the voice interaction unit (100) can be independently deployed in one or more terminal devices in a modular form, and has autonomous voice processing capabilities. Its internal module architecture includes the following functional modules:
[0121] 1. Audio Input / Output Unit (110)
[0122] This unit includes a microphone and a speaker, responsible for receiving external voice input and playing voice output, respectively. The audio signal, after being filtered and gain-controlled in the pre-amplifier stage, is then transmitted to the voice wake-up module for preliminary analysis.
[0123] 2. Wake Word Detection (120)
[0124] This unit continuously monitors ambient sounds and uses a lightweight model to identify pre-defined wake words (such as "Hey Buddy") or abnormal noises (such as alarms, screams, crying, or broken glass). When recognition is successful, it triggers subsequent voice processing and enters an active listening state. This wake word detection mechanism not only supports the monitoring of traditional command words and abnormal events, but also enables more natural and proactive voice interaction through semantic reasoning models, improving user experience and device intelligence.
[0125] 3. Local speech processing module (130) (Light ASR / NLU / TTS)
[0126] This module integrates the following sub-modules to complete the speech understanding process.
[0127] (a) ASR (Automatic Speech Recognition) module: Converts speech signals into text.
[0128] (b) NLU (Natural Language Understanding) module: parses the semantics and intent in a statement.
[0129] (c) TTS (Text-to-Speech) module: Converts system responses into speech for playback to the user. Audio input signals are processed via a 110→120→130 processing chain, sequentially completing ASR→NLU→response generation→TTS output. This module can independently handle common voice events, such as: light control, voice dialogue, emotion response, or learning game controls.
[0130] 4. Short-term memory (150)
[0131] Short-term memory can provide a reference for identifying the current context, and can play an immediate role in continuous dialogue, multi-turn interaction or personalized learning tasks. This unit temporarily stores the user's recent interaction data, such as: recent semantic instructions, response status, tone preference, emotional state, learning status, etc., to improve subsequent contextual understanding and interaction fluency.
[0132] 5. Wireless Communication Module (160) (RF Communication)
[0133] Supports protocols such as WIFI, BLE, Sub-GHz, Zigbee, Z-Wave, and Thread, for use in:
[0134] (a) Transmit the semantic event (300) to the central semantic event coordination unit (200).
[0135] (b) Receive the response returned by the central semantic event coordination unit and perform voice broadcasting or device control.
[0136] like Figure 3 As shown, an embodiment of the present invention illustrates the semantic event processing flow from the voice interaction unit (100) to the central semantic event coordination unit (200).
[0137] When a user issues a voice command, the voice signal is first received by the audio input unit (110) inside the voice interaction unit (100), and then processed by the local speech processing module (130) (ASR / NLU) for speech-to-text conversion and semantic parsing. After parsing, it is converted into a structured semantic event (300), which includes semantic type, voice summary, timestamp, and source identification, and is transmitted wirelessly to the central semantic event coordination unit (200). In the central coordination unit, the semantic event (300) is first received by the semantic event receiving and caching module (305) and temporarily stored in the event queue. This module can sort and initially filter events based on timestamps.
[0138] The continuation processing is performed by the context construction module (210), which builds a contextual framework based on the current user's context information and takes into account existing memory data (from memory modules or history not shown in the figure) to construct semantic coherence.
[0139] Once the context is constructed, the data flow enters the multi-model orchestration engine (280) (Semantic Inference & LLMsModelSelection). This module has a built-in multi-model collaboration mechanism that, based on the context complexity and event type, selects whether to call an external large language model (LLMs 600), such as GPT, Claude, or Gemini, to complete the inference computation. For simple commands, the built-in local model can handle them directly to reduce latency.
[0140] After reasoning is completed, the system enters the response generation module (290) to generate the semantically corresponding response intent and transmits it back to the voice interaction unit (100), where the voice output unit (110) performs speech synthesis and playback. Finally, the response content and processing process are synchronously recorded in the response logging and memory update module (390) to continuously update personalized data and enhance future context processing capabilities.
[0141] Figure 4This is a schematic diagram of the operational architecture of a voice interaction system according to one embodiment of the present invention, which simultaneously supports application scenarios of a single-unit voice device and a multi-unit cooperative voice device. In the single-module operation scenario, the voice interaction unit (100) internally includes: a wake-word detection module (120), a local voice processing module (130), a short-term memory module (150), and a wireless communication module (160). This unit can independently process voice input, perform real-time voice recognition and semantic analysis, and generate voice output content through text-to-speech (TTS) to respond to user commands. The processed semantic event can be transmitted to the central semantic event coordination unit (200) through the wireless module (160) to trigger subsequent semantic reasoning and voice response procedures.
[0142] The central semantic event coordination unit (200) comprises multiple semantic processing modules, including: a context construction module (210), an event prioritization module (220), a persona proxy module (230), a semantic processing module (240), a long-term memory module (250), a safe abstraction layer (260), a user account management module (270), a multi-LLM orchestration engine (280), and a response generation module (290). This architecture supports diverse semantic processing tasks such as contextual reasoning, persona synchronization, response optimization, and memory updating.
[0143] In the multi-module collaborative operation mode, multiple voice interaction units (100) (such as voice microphones, lighting modules, home control units, etc.) are distributed in the user's field. Each module has voice processing and communication functions, can independently receive voice input, and transmit semantic events back to the central unit (200). At this time, the central unit can integrate and allocate events according to the multi-module collaborative task memory block, including: an event priority sorting module (410), a context sharing module (420), a task allocation module (430), and a response task coordination module (440). This process can realize semantic synchronization, task relay, and collaborative response between multiple modules. For example, when the bedroom voice module detects an emergency command, the living room module can assist in making a sound or notifying the security platform.
[0144] Through Figure 4 As shown in the structure, the present invention can realize a flexible deployment architecture of single module and multi-mode collaboration, support functions such as cross-device voice event processing, context sharing, semantic processing and memory management, thereby improving the scalability and real-time response capability of the voice interaction system in smart home, toy, care device and security scenarios.
[0145] like Figure 5A As shown, when the user's voice input is detected by the voice interaction unit (100), the system will initiate the voice processing flow through the wake word detection module (120) and perform sentence recognition and semantic parsing through the local voice processing module (130). The parsing result will be encapsulated into a semantic event object (300), which includes fields such as sentence summary, event type, voice unit source identification, and timestamp. The aforementioned semantic event (300) will be transmitted to the central semantic event coordination unit (200) through the wireless communication module (160) to enter the subsequent semantic processing flow.
[0146] like Figure 5BAs shown, after a semantic event (300) is transmitted to the central semantic event coordination unit (200), the context construction module (210) first performs event parsing. This module is responsible for analyzing the type, statement summary, and source identifier of the semantic event, and calling the long-term memory store (250) and the persona proxy module (230) to retrieve contextual information and personalized parameters related to the user. Through the fusion of the above information, the context construction module (210) establishes a context frame (310), which includes contextual elements such as the user's historical interaction records, preference settings, tone style, and task progress, providing a foundation for subsequent semantic reasoning and response generation.
[0147] In one embodiment of the present invention, the memory and proxy mechanism of the voice interaction system is designed as follows:
[0148] (a) Short-Term Memory (150): Located within the voice interaction unit (100), this module temporarily stores recent user voice input, interactive content, and contextual fragments. It supports fast lookup and conditional synchronization mechanisms, and can report and synchronize necessary data to the long-term memory module (250) when a specific event is triggered or a time interval is reached.
[0149] (b) Long-Term Memory Store (250): Located in the central semantic event coordination unit (200), this store is used to store information such as long-term user preferences, interaction history, semantic tags, and contextual evolution records across devices and time. This module supports shared access and consistency maintenance among multiple speech modules and can accept synchronization requests from the short-term memory module.
[0150] (c) Persona Proxy Module (230): Based on the user identification code returned by the voice interaction unit, this module maintains the user's personalized voice parameters, such as speech rate, tone, intonation, and emotional expression tendencies. This module can also dynamically adjust based on the contextual framework (310) returned by the context construction module, enabling subsequent voice...
[0151] The synthesized output maintains the style consistency expected by the user.
[0152] The three modules mentioned above can be synchronized and dynamically adjusted through semantic event triggering conditions and system strategies. For example, the short-term memory module (150) can trigger synchronization with the long-term memory module (250) when it detects frequently repeated phrases or task changes, while the voice personality agent module (230) can automatically correct voice parameters based on user feedback or task completion rate to improve the consistency of personalization and interactive experience. The above modules also participate in the subsequent memory update process (such as...). Figure 5D As shown in the figure, individual updates and strategy synchronization can be performed based on the content of the voice response.
[0153] like Figure 5C As shown, the semantic events (300) generated by the voice interaction unit (100) are processed by the central semantic event coordination unit. They first enter the context construction module (210) to form a context frame (310), which integrates information such as the semantic event content, recent interaction background, user preferences, and task context. This context frame (310) serves as the input for the semantic inference module (320). The semantic inference module selects a model and performs semantic parsing based on the context and the user's voice personality characteristics to generate a response intent (330). This response intent includes...
[0154] (a) Intent Type;
[0155] (b) Response Text;
[0156] (c) Target operation action;
[0157] (d) Specify the target module;
[0158] (e) Voice style and emotion parameter settings (Persona Parameters);
[0159] (f) Memory Update Instruction;
[0160] (g) Voice unit identification code and timestamp, etc.
[0161] After semantic reasoning is completed, the system transmits the response intent (330) to the speech synthesis module (340) (TTSModule), produces speech output content (350), and plays it to complete the speech response process.
[0162] like Figure 5D As shown, when the response intent includes a prompt for a memory update to be synchronized, the system will enter the memory update process (360). This process can automatically execute one or more of the following modules' data write and update actions based on the content:
[0163] (a) Long-term Memory Store (250): Used to store user-related preference settings, semantic mirroring, interaction history or task completion records, and can support sharing among multiple voice modules.
[0164] (b) Short-term memory module (150): mainly deployed on the local end of the voice interaction unit, with fast access capability, used to record recent interactive statements, task context and temporary state.
[0165] (c) Voice Personality Proxy Module (230): Stores the user's voice based on their identification information.
[0166] Parameters such as speech pattern, speech rate, emotional expression, and response preferences serve as the basis for generating personalized voice output. In multi-module voice interaction application scenarios, the central semantic event coordination unit (200) can also automatically push the above data to other voice modules according to the synchronization strategy, ensuring contextual consistency and speech style uniformity, improving cross-module user experience, and the above modules also participate in the context construction stage (such as... Figure 5B As shown in the figure, it serves as an important basis for semantic event understanding and response preparation.
[0167] like Figure 5E As shown, in a multi-voice module collaborative deployment scenario, each voice interaction unit (100) can locally detect voice wake-up keywords (120) and perform voice recognition and semantic parsing (130), and independently generate semantic events (300) to report to the central semantic event coordination unit (200). The central semantic event coordination unit (200) includes a context construction module (210) and a semantic processing module (240), which can integrate the context, interaction history and module status based on the received semantic events (300) to perform semantic parsing and task judgment.
[0168] When the semantic processing module (240) determines that a task requires collaboration across voice modules, the central semantic event coordination unit (200) will automatically assign the main processing module and one or more auxiliary modules to handle the task together. It will then retrieve the user's voice style parameters through the voice personality agent (230) to generate a structured response intent (330), which includes response text, target action, output module number, memory update prompts, and other information. The response intent (330) will be transmitted to the designated module's voice output module (350) to execute voice playback. As needed, relevant memory or contextual parameters will be pushed back to the voice personality agent module (230) or other voice interaction units (100) via the "Memory Pushing" mechanism to achieve contextual consistency and personalized response experience across multiple modules. For details of the "Memory Pushing" synchronization process, please refer to [link to relevant documentation]. Figure 7A The example shown illustrates a memory synchronization update implementation.
[0169] like Figure 5F As shown, the semantic events generated by the voice interaction unit (100), as a sorting module within the central semantic event coordination unit (200), will execute the following priority processing flow to ensure consistency and response efficiency in multi-module event processing:
[0170] (a) Receiving semantic events (S100): The system receives semantic event objects (300) reported by the voice interaction module.
[0171] It includes core fields such as event type, timestamp, confidence level, and context label.
[0172] (b) Extracting key features (S110): The system parses the event field and extracts information such as event type, emotional intensity, contextual markers, source credibility and confidence level, which are used as the basis for subsequent ranking and scoring.
[0173] (c) Preliminary Priority Ranking (S120): Based on the characteristics of the event category (e.g., alarms take precedence over queries) and the source module (e.g., the main module takes precedence over the slave module), a preliminary hierarchical classification and ranking is performed to establish a rough priority queue.
[0174] (d) Redundant event arbitration (S130): When multiple voice modules report similar events at the same time (such as multiple modules detecting the same alarm sound), the system selects and retains the most reliable version based on parameters such as confidence level, time difference, and source module, and marks the others as redundant.
[0175] (e) Time alignment (S140): The system refers to the event timestamp and adjusts the order of multiple events to ensure that the processing order error caused by upload delay will not affect the system's judgment.
[0176] (f) Event scoring mechanism (S150): Quantify and score key features such as confidence level, emotional intensity, and contextual relevance.
[0177] Used for priority ranking.
[0178] (g) Priority level labeling (S160): The system labels events as high priority, medium priority or low priority based on the scoring results to facilitate subsequent scheduling and model routing decisions.
[0179] (h) Determine output conditions (S170): The system determines the output based on set conditions (such as scoring threshold, event status, system load).
[0180] Determine whether to send the event to the semantic processing module for further processing.
[0181] (i) Output to the coordination layer (S180): Events that meet the conditions will be transmitted to the central semantic event coordination unit (200).
[0182] To construct context and generate responses.
[0183] This processing mechanism can significantly reduce the system burden caused by recurring events, improve the accuracy of event reasoning, and enhance the stability and real-time response capability of the voice interaction system in multi-user, multi-device scenarios.
[0184] like Figure 6A As shown, semantic events are first assessed for their privacy level by the privacy mapping module (610), and then enter the safe abstraction layer (260) for masking and abstraction processing. Based on the event's privacy level and semantic characteristics, the event is then processed by the model routing strategy (640), a submodule of the multi-model orchestration engine (280), which is used to select specific inference paths from different paths. The system will select the appropriate inference path from the following modules:
[0185] (a) Edge Large Model Module (641) (Edge LLM);
[0186] (b) Cloud Large Model Module (642) (Cloud LLM);
[0187] (c) Hybrid Inference Engine (643)
[0188] This design balances privacy protection and inference performance, and through the overall scheduling mechanism of the multi-model orchestration engine (280), it supports the initial parsing of lightweight models in local modules to improve computational efficiency.
[0189] like Figure 6B As shown in the flowchart, this diagram illustrates the privacy handling mechanism for semantic events within the central semantic coordination layer. Before an event enters the semantic processing module, the system completes privacy protection according to the following steps:
[0190] (a) Privacy Level Mirroring (S210): Classified by event source, speaker characteristics, and semantic type, categorized as P0
[0191] Levels: (Public), P1 (Partially Masked), P2 (Sensitive), P3 (Private).
[0192] (b) Security abstraction processing (S220): Perform field masking processing, covering data such as user name, location, and mood.
[0193] (c) Determine whether inference is allowed (S230): The subsequent model inference process can only proceed if the event privacy level is lower than or equal to P1.
[0194] (d) Submitting the model routing strategy (S240): Qualified events are submitted to the model routing module for inference and assignment, corresponding to... Figure 6A Module selection strategy.
[0195] This process ensures that high-privacy-level events do not enter the cloud or edge inference modules without authorization and enhances data security.
[0196] like Figure 7A As shown, when any voice interaction module detects a change in local personality parameters (such as speech rate, tone, etc.) (S710), it automatically triggers a reporting mechanism to send the change information back to the central semantic event coordination unit (200) (S720). After receiving the updated parameters, the central semantic event coordination unit (200) performs a version comparison process (S730) to determine whether the reported parameters are consistent with the standard personality parameters stored in the central database.
[0197] If the comparison results are identical, the process ends and synchronization is deemed unnecessary (S740). If a difference exists, the central unit generates an update event (750) and broadcasts it to other voice interaction modules via a synchronization mechanism (S760). Each receiving module automatically updates its local voice personality agent based on the event content (S770) to achieve synchronization consistency of voice style parameters among multiple modules. The process ultimately ends at step (S780).
[0198] Figure 7BThis is a schematic diagram of an embodiment of the present invention, illustrating how the system performs a voice personality synchronization procedure when the personality proxy data (711A-711C) held by the voice interaction units (100A-100C) are inconsistent. First, each voice interaction unit reports its local voice personality proxy version (711A-711C). The central semantic event coordination unit (200) compares and aggregates the content through the difference merging module (720), and in the conflict resolution module (730), based on priority logic and authorization settings, produces a consistent voice personality proxy version (740) (labeled 711D). This version is the standard voice style data after conflict resolution and will be synchronously transmitted back to each voice interaction unit to maintain the consistency of personality style in the overall voice interaction system. It balances individual module learning with collective style unification, supports dynamic merging, conflict arbitration, and OTA synchronization, and is suitable for interactive systems requiring cross-module voice style consistency.
[0199] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A distributed voice interaction unit system, characterized in that, Include: (a) At least one voice interaction unit, modularly deployable in one or more terminal devices, each voice interaction unit having local voice processing capabilities, including at least: i. Semantic event generation module, used to convert local speech processing results into semantic event objects with structured fields. These fields include at least event type, sentence summary, sentiment index, confidence level, source identifier, timestamp, context marker, and privacy level marker. This field structure is suitable for the consistent representation of speech events and cross-modal events. ii. Short-term memory module, used to store recent conversation context and preferences; iii. A communication module for transmitting semantic event objects to the central semantic event coordination unit; (b) A central semantic event coordination unit (AIOrchestration Layer) for receiving semantic events from various speech interaction units, including at least: i. Safe Abstraction Layer, which automatically performs masking processing such as field deletion, obfuscation, category generalization or dynamic replacement value based on the privacy level tag in the semantic event object, and generates an event summary with masking policy summary and security level label; ii. Multi-LLM Orchestration Engine: Based on latency tolerance, semantic complexity, privacy level, user role authorization, regional regulations and service quality indicators, it dynamically selects edge-deployed models, local lightweight models or cloud-based large language models for inference, and supports real-time hot switching between heterogeneous models and continuous maintenance of context. iii. The Persona Proxy and Memory Synchronization Module triggers the synchronization of short-term and long-term memory based on the frequency of semantic events, emotional relevance, or explicit user commands, and dynamically adjusts the voice personality parameters based on historical interaction data. iv. Memory synchronization and version control mechanisms, including Diff-Merge, conflict arbitration, and version governance strategies, are used to maintain consistency of short-term and long-term memory data across multiple voice interaction units. It also supports over-the-air (OTA) updates, broadcast synchronization, and rollback operations; v. The response intent generation module is used to output a structured response intent containing response text, a specified voice interaction unit identification code, voice style parameters and memory update prompts, and to send the response intent back to the corresponding voice interaction unit.
2. The distributed voice interaction unit system according to claim 1, characterized in that: The semantic event object includes at least: event type, statement summary, sentiment index, confidence level, source module identifier, timestamp, context marker, and privacy level marker; Furthermore, the Central Semantic Event Coordination Unit (AIOrchestration Layer) has an event prioritization module, which is used for: (a) Based on event urgency, context relevance, source credibility, confidence level and timestamp differences, receive multi-source semantic event objects are sorted and prioritized; (b) When a multi-source redundant event is detected, deduplication and time alignment are performed based on the confidence threshold and the source priority rule; (c) The semantic event objects after sorting and arbitration are processed by determining the processing node based on the current latency tolerance and quality of service (SLO) index, and the running location can be dynamically switched between cloud deployment, edge computing configuration, or a hybrid architecture; wherein, the sorting and arbitration are based on semantic hierarchical features and cross-source consistency scores, the features including event urgency, context relevance, user authorization, privacy level labels, and cross-modal consistency index; the quality of the original audio signal is not a necessary condition; (d) Maintain consistency between the context frame and the event handling state when switching processing nodes.
3. The distributed voice interaction unit system according to claim 1, characterized in that: The Safe Abstraction Layer includes a semantic masking module, which performs the following steps based on the privacy level label of a semantic event before it is sent to the semantic reasoning module: (a) Based on the privacy level marker attached to the semantic event object, automatically identify the sensitive information contained in its fields, including but not limited to personally identifiable information, voiceprint indicators, emotion index, geographical location, identity code or device serial number; (b) Apply corresponding de-identification strategies for different privacy levels, including: field deletion, numerical obfuscation, category generalization, or context-based dynamic replacement value generation. (c) After the masking process is completed, an event summary containing a summary of the masking strategy adopted and a security level label is generated, and sufficient contextual information to support context construction and semantic reasoning is retained. (d) The security level label and the masking strategy summary are incorporated into the event summary and sent back, so that the Multi-LLM Orchestration Engine can directly call them as one of the constraint factors in model selection and running node switching.
4. A distributed voice interaction unit system according to claim 1, characterized in that: The Multi-LLM Orchestration Engine includes: (a) Policy Selector, used to select and score models based on at least the following factors: event delay tolerance, semantic complexity, data privacy level, role authorization permissions, local regulations and system service quality indicators; (b) Model Router, based on the output of the policy selector, directs semantic events to an applicable language model, including but not limited to: edge deployment model, local lightweight model, and cloud-based large language model; (c) Supports instant hot switching between heterogeneous language models, and performs serialization and recovery of the intermediate context frame during the switching process; prioritizes compliant running nodes when there are conflicts in privacy levels or regional regulations; (d) Parallel queries and weighted integration of results can be performed across multiple language models to improve the accuracy of inference.
5. A method for distributed voice interaction and semantic event coordination processing in the system as described in claims 1-4, characterized in that, Includes the following steps: (a) The voice interaction unit receives voice input and performs speech recognition (ASR) and natural language understanding (NLU) processing locally to generate a structured semantic event object containing event type, sentence summary, sentiment index, confidence level, source identification code, timestamp, context label and privacy level label. (b) Store the semantic event object into the event buffer queue according to its timestamp, context marker and confidence level; (c) Perform event sorting, conflict arbitration, redundancy removal and time alignment on the semantic event objects; (d) The semantic event object is de-identified through a Safe Abstraction Layer and a privacy level label is added; (e) Based on the latency tolerance, privacy level and user role authorization of the event, dynamically select cloud or edge models for semantic reasoning through the Multi-LLM Orchestration Engine, and maintain the consistency of the contextual framework. (f) Generate a response intent based on the reasoning result, wherein the response intent includes response text, target voice interaction unit identification code and memory update prompt; (g) The response intent is sent back to the corresponding voice interaction unit, and the synchronization operation of short-term memory and long-term memory is triggered according to the memory update prompt. During the synchronization process, the parameters of the voice personality proxy are updated to maintain the consistency of voice style and interactive experience across modules.
6. The method for coordinated processing of distributed voice interaction and semantic events in a distributed voice interaction module system according to claim 5, characterized in that, The generation of the semantic event object includes the following steps: (a) Convert the received voice signal into text data. This conversion can be performed in the local voice processing module of the voice interaction unit to support offline or low-latency operation mode. (b) Perform semantic parsing on the text data to identify user intent, task type and context-related parameters; (c) Extract emotional features, user status and context markers from the semantic parsing process, and generate privacy level markers based on the current dialogue context; (d) Generate a version identifier code and / or hash value for the semantic event object to track the data consistency and integrity during cross-module transmission and processing, and perform version comparison and integrity verification when the central semantic event coordination unit or other voice interaction unit receives the event. (e) When a version conflict or data anomaly is detected, an event recovery or retransmission mechanism is triggered to ensure that semantic events remain consistent and traceable across multiple modules.
7. The method for coordinated processing of distributed voice interaction and semantic events in a distributed voice interaction module system according to claim 5, characterized in that, The Safe Abstraction Layer includes the following processing steps: (a) Based on the privacy level marker attached to the semantic event object, automatically identify the sensitive information contained in its fields, including but not limited to personally identifiable information, voiceprint indicators, emotion index, geolocation, identity code or device serial number; (b) Apply corresponding de-identification strategies for different privacy levels, including field deletion, numerical obfuscation, category generalization, or context-adaptive dynamic replacement value generation. (c) While maintaining the integrity of contextual information and intent parameters related to semantic reasoning, generate a processed semantic event object and a corresponding event summary. The event summary includes a summary of the masking strategy adopted and a security level label. The event summary is provided to the Multi-LLM Orchestration Engine as one of the constraint factors for model selection and running node switching.
8. A distributed voice interaction unit system according to claim 1, characterized in that: The central semantic event coordination unit supports cross-modal event fusion processing, including: (a) Receive heterogeneous event data generated by external modules, including but not limited to sensor events, image recognition events and third-party system data; (b) Convert the heterogeneous events into semantic event objects with structured fields, the fields having the same field structure as the semantic event objects generated by the voice interaction unit, to support unified processing; (c) Based on the contextual framework, perform semantic processing and contextual reasoning to generate response intent or trigger multi-module collaborative processing, and calculate cross-source consistency indicators for the fused event set, including but not limited to weighted confidence, time-series alignment score or source reputation. d) Based on the fused contextual framework and reasoning results, trigger the coordinated response of one or more voice interaction units, including but not limited to voice response, visual warning, device control or cross-system command issuance.