Auxiliary communication system based on anthropomorphic voice substitution mechanism and interaction method thereof
The auxiliary communication system, which uses an anthropomorphic voice narration mechanism, solves the problem of low user willingness to communicate in existing technologies. It enables proactive response stimulation and data analysis in non-real-time scenarios and is applicable to multiple scenarios such as family education and social guidance.
Patent Information
- Application Number
- CN202510909002.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2025-10-24
AI Technical Summary
Existing voice interaction technologies lack flexible guidance mechanisms and personalized emotional connection processes, making it difficult to stimulate users' willingness to express themselves, especially in non-real-time communication scenarios where there is a lack of effective auxiliary communication systems.
It adopts a human-like voice narration mechanism and uses a distributed architecture-based auxiliary communication system to build a closed loop of voice interaction by leveraging edge computing and cloud analysis. This includes human-like voice style synthesis, behavior recognition, and data feedback mechanisms, supporting asynchronous task playback and user response collection, reducing user resistance, and increasing the initiation rate and participation in communication.
It enables users to proactively respond in non-real-time communication scenarios, reduces communication pressure, increases the communication initiation rate, provides quantifiable behavioral data analysis, and supports multi-scenario deployment and personalized intervention.
Smart Images

Figure CN120833784A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of voice interaction, human-computer emotional interaction, artificial intelligence voice synthesis and auxiliary communication system, in particular to an auxiliary communication system based on a personified role voice reporting mechanism and an interaction method thereof, which is suitable for user groups with weak communication willingness or insensitive response to traditional human voice guidance, and is widely used in family education, social guidance, emotional accompaniment, learning training and other scenes. BACKGROUND
[0002] In real life, part of the user groups, especially children, emotionally sensitive groups or people with low communication willingness, often show slow response, no response or communication avoidance when facing direct language guidance (such as oral instructions from parents or teachers) due to psychological resistance, attention shift or social avoidance.
[0003] The existing voice interaction technology mostly constructs human-computer communication process in the way of "instruction recognition-voice broadcast", lacks flexible guidance mechanism and personalized emotional connection process, and is difficult to stimulate the user's initiative to express willingness. In addition, traditional voice assistants or learning aids usually use single neutral voice output, lack of role characteristics, tone changes and context adaptation, and are difficult to build a real and credible communication atmosphere.
[0004] Research and practice show that part of the users have natural trust and interaction desire for personified roles (such as dolls, animation characters, virtual friends). When the information is expressed by a user-familiar and trusted role in a specific voice style, the user is more likely to receive the content and produce a response behavior. Therefore, if the speech of the intervenor or the guide is reported by a "third-party role" in voice, and is output by an intelligent device with appropriate tone and rhythm, it will help to alleviate the communication pressure and stimulate the user's participation in communication behavior.
[0005] At present, there is still a lack of a complete auxiliary communication system that integrates role voice synthesis, behavior response recognition, task scheduling and data feedback closed loop, especially in non-real-time and non-face-to-face communication guidance scenarios, there is a lack of effective technical path to support the user's initiative to express and continue interaction. SUMMARY
[0006] The present application provides an auxiliary communication system based on a personified voice reporting mechanism and an asynchronous interaction method thereof, aiming to solve the problems of emotional acceptance difficulty, low initiative response and high pressure in traditional human-computer or human-to-human communication. It is especially suitable for children, adolescents and other groups of people who have emotional connection with specific roles. By "replacing expression" of the speech intention of the intervenor by a role trusted or loved by the user, the user's response willingness and participation enthusiasm are improved.
[0007] I. Technical implementation goal
[0008] The core technical objectives of the present application include; A non-human intermediary voice interaction mechanism is constructed, and the language content of the intervenor is conveyed through the mode of "role narration", so as to reduce the communication resistance; Voice guidance behavior is embedded in the natural context of the user, and the communication behavior is triggered by taking the user familiar object (such as a doll) as a carrier, so as to improve the communication initiation rate; Asynchronous voice task playing and user response collection are realized, so as to avoid the pressure on the user for instant interaction; A structured response data collection and behavior recording mechanism is provided, so as to form behavior data that can be quantified and tracked for intervention effect; Remote intervention, one-to-many task management and personalized voice style configuration are supported, so as to facilitate multi-scene deployment and continuous intervention.
[0009] II. System composition architecture
[0010] The present system adopts a distributed architecture and is mainly composed of the following three subsystems:
[0011] 1. User-side voice interaction device (user side): installed in dolls, toys, figurines, cloth devices or other interactive structures, integrated with edge computing units and voice processing modules, and having the following functions: Voice collection and instruction execution: built-in high-sensitivity microphone for collecting patient voice; Local voice synthesis: supporting synthesis of specified role tone through pre-installed TTS model; Voice playing module: loudspeaker outputs voice content with emotional intonation to guide the user to imitate; Edge recognition: local preliminary recognition of response content (such as keywords, speech rate, volume); Safe caching and return: record the collected data and support offline caching, and connect to upload after connection.
[0012] 2. Intervention-side interaction platform (intervention communication side): doctors, therapists and parents configure content through mobile devices or web platforms, including: Task editing and dispatching module: configure voice content, task type (prompt, question, motivation), time scheduling strategy; Style synthesis setting module: select role tone (such as cartoon characters, family simulation, etc.), speech rate, intonation; User management module: supports one-to-many user group management, and configures patient portraits and preference settings; Behavior record browsing module: browse historical interaction data and logs according to time axis or task classification; AI suggestion engine (optional): recommend next task template or adjust content based on user performance.
[0013] 3. Cloud control and analysis platform (service background): Provide data hub and model calculation service: Multi-device scheduling management: Register the voice device, bind the ID, and realize unified control; Task synchronization and delivery: Translate the intervention end task into a standard message and distribute it to the voice device according to the set time; Data reception and cleaning: Receive user voice data and behavior logs, and perform preprocessing and labeling; Evaluation and analysis module: Analyze user response based on algorithm model (such as response delay, keyword coverage, and speech speed deviation); Log archiving and trend modeling: Generate individual intervention trajectory graph, voice training growth curve, and reaction emotion trend; Data interface module: Support API to third-party rehabilitation platform to realize data synchronization and intervention collaboration.
[0014] Three, key function modules and their coordination
[0015] The system of the present application constructs a complete closed loop around "task delivery - voice narration playback - user response recognition - behavior data collection - feedback analysis and strategy update", mainly including the following five types of function modules:
[0016] 1. Synthetic module of human-like voice style This module relies on a pre-built multi-role voice synthesis library and TTS model to support the synthesis of intervention end input content into specific voice style role output, such as cartoon characters, relatives simulation, virtual animals, etc. Voice style can load emotion adjustment parameters such as "happy", "comfort", "encouragement", "calm", etc. to enhance the voice appeal and user trust, and improve the narration effect. The module supports individual setting and role binding, and has the ability of timbre caching and compression transmission, which is suitable for low bandwidth environment.
[0017] 2. Local speech recognition module (edge ASR) User response behavior is preliminarily identified by the device end, and based on a lightweight speech recognition model (such as keyword matching, small sample adaptation or micro RNN structure) to determine whether the user has made an effective voice response, identify its keyword hit, speech speed, voice energy and other parameters. It can realize local cache recognition result, reduce the dependence on network stability, and support offline retransmission mechanism and edge data preprocessing strategy.
[0018] 3. Timed and event triggered scheduling module Task execution supports multiple scheduling mechanisms, including: Time table trigger: such as "play good morning narration every day at 8:00"; Event trigger: such as "play reward voice after completing the last task"; Intelligent adaptation: adjust the task playing interval according to the response frequency; In-module task queue and conflict detection mechanism to ensure that high-priority tasks are not covered, and failed tasks can be set to retry mechanism and record the failure reason.
[0019] 4. Data recording and behavior log module The system conducts structured data collection for each interaction behavior, including task ID, voice playing time, user response time, whether to respond, voice feature indicators, environmental noise level, and generates standardized logs. Log data is encrypted by default and cached locally, and the upload period is set (such as every 2 hours / daily / after task completion) to ensure data security and resume transmission.
[0020] 5. Cloud intervention feedback and AI assisted decision module The cloud platform aggregates device reporting data, automatically forms daily / weekly interaction reports, and provides user behavior trend charts such as "response time shortening curve", "keyword completion rate heat map", "speech speed stability line chart", etc. to help interveners identify progress changes. In addition, the system has an AI suggestion engine (such as rule-based recommendation or small model clustering recommendation) that can dynamically adjust recommended task templates, tone style and task difficulty based on user historical behavior data, achieving intelligent auxiliary optimization of the intervention process.
[0021] Four, the essential difference from traditional interaction scheme
[0022] 1. Emotional intermediary transformation of interaction path: Instead of "intervener direct speech" form of voice output, it constructs a "person-anthropomorphic character-person" voice interaction intermediary bridge in the form of "user familiar role voice statement intervention content", which greatly reduces user resistance.
[0023] 2. The interaction mode is changed from synchronous to asynchronous: The system supports asynchronous playback of tasks, delayed response of users, breakpoint resume and scheduling management, avoiding the pressure on users caused by "immediate response", and encouraging natural rhythm exchange generation.
[0024] 3. Voice content has the ability of role feeling and emotional expression: Unlike "system voice" or "human voice repetition" in traditional systems, the invention emphasizes the "affinity expression" and "tone personality customization" of voice, enhancing the emotional link with users.
[0025] 4. Build a structured data collection and closed-loop feedback system: From task creation to response analysis, to personalized task optimization, it forms a complete data closed-loop path, adapting to long-term observation and intervention optimization.
[0026] 5. Portable, scalable, and platformable deployment: The device structure has modular characteristics and can be deployed in dolls, figurines, child intelligent devices, education terminals, and other carriers. The system supports multi-terminal, multi-user synchronous collaboration and adapts to the deployment needs of multiple environments such as families, special education schools, rehabilitation platforms, and community service stations.
[0027] Five, data processing and feedback closed-loop path
[0028] The system forms a complete closed-loop path around voice task interaction: 1. Task definition and role setting: The interveners input language content, select anthropomorphic role voice, set tone style and trigger rules (timing / event) through the control terminal; 2. Voice package generation and task distribution: The cloud performs TTS synthesis and parameter packaging on the input content, generates a task voice package with role timbre and emotion parameters, and distributes it to the designated user device; 3. Device playback and user reception: The user-side device automatically plays at the set time or event trigger after receiving the task, guiding the user to naturally generate response behavior; 4. User response collection and edge recognition: The device collects response voice through the microphone and uses the edge ASR model to identify whether the response, keyword matching degree, speech speed change, sound energy, and other features; 5. Data return and behavior modeling: The data is uploaded to the cloud through the communication module, unified filing, and forms user behavior timeline, response feature indicators, and interaction completion rate; 6. Intervention feedback and strategy adjustment: The interveners can view the behavior report through the control platform, or automatically recommend and adjust the content, role, or rhythm by the system, and enter the next round of intervention. BRIEF DESCRIPTION OF DRAWINGS
[0029] Figure 1 : System overall architecture diagram, showing the data flow and task flow between the three subsystems (user end, intervention end, cloud end) of the invention; Figure 2 : Voice interaction device structure diagram, showing the main hardware modules and their functions inside the device; DETAILED DESCRIPTION
[0030] In order to make the purpose, technical scheme and beneficial effects of the present application more clear and explicit, the present application is further described in detail below in combination with the drawings and examples. The present application is not limited to the following specific examples, and any equivalent replacement or improvement within the spirit and principle of the present application shall be included in the protection scope of the present application.
[0031] Example one: daily language guidance training of family
[0032] A 5-year-old child Xiaochen with mild language expression disorder has poor response to the mother's voice stimulation, but has obvious emotional attachment to his familiar "dinosaur doll". The mother inputs "Xiaobao, it's time to eat now ~ come on" through the mobile phone App. The system synthesizes it into the personification voice of the dinosaur role and sends it to the doll built-in device to play. Xiaochen gradually shows response behavior after hearing the familiar role voice. The system records the response duration and keyword feedback for parents to evaluate the interaction effect.
[0033] Example two: multi-person collaborative theme course guidance scene
[0034] Special education teachers set up "fruit recognition" tasks and configure multiple voice roles (such as "Banana Brother" and "Apple Sister") to pronounce "I am a banana" and "Who am I" in turn, encouraging students to imitate. Each student is equipped with an independent voice doll device, and the system records the number of imitations and keyword recognition matching rate of each student to form a data report for teachers to evaluate.
[0035] Example three: remote emotion guidance of psychological therapist
[0036] The therapist sets up daily voice tasks for children with social disorders, and the system plays through the patient's favorite doll role: "Today we are going to try to say hello to others, okay?" The patient completes the task in a stress-free environment, and the system records the feedback for the therapist to review and adjust.
[0037] Example four: speech therapist teaching intervention scene
[0038] In the language expression training course, the teacher sets up daily practice tasks through the system, such as: "Please say 'apple' again." "Great, let's say 'banana' together." Multiple voice devices play corresponding role voices in order to students, prompting and encouraging them. After the student completes the pronunciation, the system identifies the response keywords and speech speed characteristics and provides feedback through encouraging voice (such as "Great, do it again!"), improving interaction participation. Data is recorded in real time and uploaded to the platform for teaching analysis.
[0039] Example five: multi-sensory auxiliary teaching of special education teacher
[0040] In special education schools, teachers use multiple voice dolls to play instructions for daily tasks at the same time: "Today we put on our shoes and go to the playground." Students follow the familiar voices to complete the tasks, and the system automatically counts the frequency of following and the length of reaction time to provide evaluation reports for teachers.
Claims
1. An assisted communication system based on an anthropomorphic voice proxy mechanism, characterized in that, Comprise: A voice interaction device, with voice output module, voice style synthesis module, voice input module, edge recognition module, communication module and local cache module, embedded in a carrier familiar to the user; An intervention platform for interveners to input text or voice information, set tone style, emotion parameters and play strategy; A cloud control platform for generating proxy voice tasks, scheduling content, receiving response data and conducting behavior analysis; The voice interaction device receives proxy voice tasks and plays them, guiding the user to respond; The edge recognition module identifies user response information and uploads it through the communication module; The intervention platform adjusts the task content according to the user data to form a continuous interaction loop.
2. The system of claim 1, wherein, The voice interaction device is embedded in a figurative structure, including but not limited to dolls, figurines, cloth toys, voice terminals or other structures.
3. The system of claim 1, wherein, The voice style synthesis module supports role sound library calling, including but not limited to cartoon characters, relative simulation sound, animal imitation sound, etc.
4. The system of claim 1, wherein, The intervention platform provides task timing scheduling function, which can set task playing time, repetition period and priority.
5. The system of claim 1, wherein, The cloud control platform supports multi-user management, which can batch issue tasks to multiple voice interaction devices.
6. An interactive method based on the system of claim 1, characterized in that: Comprise the following steps: Interveners input text or voice information through the platform; The system converts the content into anthropomorphic voice form and generates tasks; The user terminal receives the task and plays it at the set time; Users respond with voice or behavior; The edge recognition module analyzes the response information; Return to the cloud for analysis and intervention feedback update.
7. The method of claim 6, wherein, The playing process is asynchronous playing, supporting timing trigger, semantic trigger or user behavior trigger.
8. The method of claim 6, wherein, The recognition method of user response includes keyword recognition, response delay measurement, etc.