An Adaptive Low-Intrusion Voice Companion Method and System Based on Human Multi-Dimensional State Perception

CN122575350APending Publication Date: 2026-08-14GUANGDONG JIAZHICHUANG TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-30
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0003]本发明的目的在于克服现有技术的不足,提供一种基于人类多维度状态感知的自适应低打扰语音陪伴方法、系统、终端设备及计算机可读存储介质,构建多维度感知、极致低打扰控制、同源真人语音交互与离线隐私保护相结合的人机交互方案,解决现有技术中感知片面、打扰严重、隐私泄露风险高、交互自然度低的技术问题

Benefits of technology

1. 构建全场景多维度用户状态感知体系,通过融合场景环境、情绪生理、行为节律、终端使用状态与实时交互意图,全面提升智能交互的用户理解能力与人性化交互水平;

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

This invention discloses an adaptive, low-intrusion voice companionship method, system, terminal device, and computer-readable storage medium based on multi-dimensional human state perception. It achieves voice companionship, emotional soothing, and reminders in an extremely low-intrusion manner—without pop-ups, vibration, interruptions, or touching on taboos—by perceiving the user's environment, emotional physiology, behavioral rhythms, terminal usage status, and real-time interaction intentions, combined with the user's personalized characteristics and preset taboo boundaries. Utilizing a multi-emotional, homogeneous voice library recorded by the same human, it provides voice companionship, emotional comfort, and reminders in an extremely low-intrusion manner, without pop-ups, vibration, interruptions, or touching on taboos. This invention supports local offline priority operation, offers strong privacy and security, low hardware adaptation costs, and broad commercial application scenarios, and can be widely used in various smart terminals such as mobile phones, in-vehicle systems, and smart homes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence, human-computer interaction, edge intelligence and voice interaction, and in particular to an adaptive low-intrusion voice companionship method, system, terminal device and computer-readable storage medium based on human multi-dimensional state perception. Background Technology

[0002] Existing voice interaction systems generally suffer from several problems, including: limited perception dimensions (capable of recognizing only user-initiated voice commands and failing to perceive user context, emotional state, and behavioral rhythms); mechanical interaction logic lacking emotional warmth; high intrusiveness (frequent pop-ups, vibrations, and notification sounds disrupting users' daily lives); insufficient privacy and security (large amounts of user voice data are uploaded to the cloud, posing a risk of leakage); and over-reliance on cloud computing power, rendering them unusable without a network connection. These issues make it difficult to achieve human-centered, seamless, comfortable, secure, and controllable personalized companionship interaction, and fail to meet users' diverse, personalized, and low-intrusion needs. Summary of the Invention

[0003] The purpose of this invention is to overcome the shortcomings of the prior art and provide an adaptive low-intrusion voice companionship method, system, terminal device and computer-readable storage medium based on human multi-dimensional state perception. It constructs a human-computer interaction solution that combines multi-dimensional perception, extremely low-intrusion control, homogeneous real human voice interaction and offline privacy protection, and solves the technical problems of one-sided perception, serious disturbance, high risk of privacy leakage and low naturalness of interaction in the prior art.

[0004] The core technical solution of this invention is as follows: By comprehensively recognizing the user's real-time status through multi-dimensional perception, it breaks through the limitations of traditional single-command recognition. By establishing clear and actionable rules for minimizing disruption to users, unnecessary interference can be reduced. By using a database of real human voices from the same source, we can enhance the naturalness, realism, and emotional warmth of voice interaction. By prioritizing local offline operation, we can improve privacy, security, and terminal stability.

[0005] Beneficial effects 1. Construct a multi-dimensional user status perception system across all scenarios, and comprehensively improve the user understanding ability and humanized interaction level of intelligent interaction by integrating scenario environment, emotional physiology, behavioral rhythm, terminal usage status and real-time interaction intent; 2. Establish standardized and implementable rules for minimally disruptive interaction, fundamentally addressing the industry pain point of existing voice systems excessively disturbing users, and significantly improving user comfort and long-term acceptance. 3. By using a homogeneous real human voice library, the awkwardness caused by splicing multiple voice sources is avoided, effectively improving the naturalness, realism and emotional companionship of voice interaction; 4. Implements a local offline-first operation architecture, with core data processing completed on the terminal first, reducing cloud data transmission, significantly reducing the risk of user privacy leakage, and enhancing data security; 5. Based on a universal accelerometer and gyroscope sensor for terminals, it is compatible with most mainstream smart terminals, requires no additional hardware modification, has low hardware cost, and has a wide range of compatibility. 6. Supports continuous iteration and optimization based on user feedback to improve interaction fit and personalization capabilities; 7. It can be widely used in multiple fields such as consumer electronics, smart cars, smart homes, and wearable devices, with a wealth of commercial application scenarios. Detailed Implementation

[0006] The present invention will be further described in detail below with reference to specific embodiments. The following embodiments are for illustrative purposes only and are not intended to limit the scope of protection of the present invention.

[0007] The method and system described in this invention can be applied to various smart terminals with accelerometer sensors, gyroscope sensors and voice playback functions, such as smartphones, tablets, smartwatches, smart cars, smart speakers, smart TVs, and smart home terminals. All core data processing flows are executed locally on the terminal first, without the need to rely on cloud servers.

[0008] Example 1: Silent Adaptation in Rest Scenarios When a user places their smartphone face down on a table, the built-in accelerometer and gyroscope collect linear acceleration and angular velocity data to determine that the device is in a horizontal, screen-down position. Combining this with the current system time and the user's historical circadian rhythm data, the system determines that the user is in a resting state. The system automatically triggers an ultra-quiet, low-interference mode, disabling unnecessary voice announcements and retaining only gentle voice prompts for preset emergency matters. Throughout this process, there are no pop-ups, vibrations, or notification sounds, ensuring no disturbance to the user's rest.

[0009] Example 2: User Emotional Soothing Adaptation The terminal collects the user's voice acoustic features through the microphone, extracts parameters such as voice frequency, loudness, and rhythm, and analyzes to determine whether the user is in a low mood. It retrieves user preference data stored in the personalized and taboo memory modules to determine the appropriate soothing tone. It calls the corresponding soothing voice material from the same source real human voice library and uses the low-interference control module to deliver a voice reassurance broadcast with a gentle volume, slow speed, and low frequency, without actively interrupting the user's current behavior or triggering any visual or vibration prompts.

[0010] Example 3: Low-Distraction Reminder for Driving Scenarios When applied to in-vehicle intelligent terminals, the system collects vehicle driving status data through accelerometers and gyroscopes, and combines this data with the driver's operating rhythm and voice acoustic characteristics to determine if the driver is in a normal driving state. If the system detects signs of driver fatigue, it provides a safe rest reminder in a very low-frequency, serious, and gentle tone, without any unnecessary announcements, multimedia interference, or bright light prompts, thus avoiding distraction and strictly adhering to the rule of minimal disturbance in driving scenarios.

[0011] Example 4: User Taboo Avoidance Execution Users can set up their own "Do Not Disturb" periods by interacting with the terminal to refuse certain types of notifications. The system records this information in the personalization and taboo memory module and stores it as the user's taboo threshold data. In subsequent operation, the system automatically avoids the corresponding notifications, remains silent throughout the "Do Not Disturb" period, and does not trigger related voice services, eliminating the need to repeatedly confirm the user's preferences.

[0012] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. An adaptive, low-intrusion voice companionship method based on human multi-dimensional state perception, characterized in that, Includes the following steps: The user's status is perceived from multiple dimensions through the terminal, including at least the scene environment status, emotional and physiological status, behavioral rhythm status, terminal usage status, and real-time interaction intent. A comprehensive decision is made based on the user's multi-dimensional status, personalized characteristics, and preset taboos and bottom lines. The system uses a library of real human voices with multiple emotions and tones, recorded by the same person, for voice output. Perform voice companionship, emotional soothing, task reminders, or silent adaptation according to the preset rules for minimal disruption; The rules for minimal disruption include: no pop-up notifications, no unnecessary device vibrations, no notification sounds, no background music, no flashing lights, no interruptions to the user's focused behavior, and no violation of the user's preset taboos.

2. The method according to claim 1, characterized in that, The multi-dimensional state perception includes: The terminal's operating status is used to identify the scene and environment. Collect speech acoustic features to identify user emotions and physiological states; Collect acceleration, angular velocity, touch interaction behavior, and circadian rhythm data to identify user behavior posture, terminal usage status, and real-time interaction intent; A personalized user feature model is built based on historical interaction data, user feedback, preference settings, and taboo information.

3. The method according to claim 1, characterized in that, The core perception, comprehensive decision-making, and data storage processes all support local offline priority operation. They do not collect complete raw voice data from users. They only perform necessary semantic recognition when users actively wake up and issue clear voice commands, and do not upload users' core personal privacy data to cloud servers.

4. The method according to claim 1, characterized in that, The terminal adaptively adjusts the tone, speed, volume, frequency, duration and triggering timing of the voice output according to the user's real-time status, and dynamically adapts and adjusts based on user feedback to achieve personalized interaction optimization.

5. The method according to claim 1, characterized in that, When the system detects that the user is focused on working, studying, resting, sleeping, driving, or using the public space, it automatically enters an extremely quiet and low-interference mode, only performing gentle voice output when necessary and appropriate for the scenario.

6. The method according to claim 1, characterized in that, The terminal stores and strictly adheres to user-preset taboo boundaries, which include offensive tone, disliked language, time periods for refusing to be disturbed, and personal emotional interaction boundaries.

7. The method according to claim 1, characterized in that, The terminal's posture and handheld status are obtained solely through data calculation from the accelerometer and gyroscope sensors, without relying on cameras, infrared sensors, or other visual sensors, and the entire posture calculation process is completed locally on the terminal.

8. The method according to claim 1, characterized in that, The source-based real-person voice library consists of voice clips with different emotions and tones recorded by the same real person.

9. An adaptive, low-intrusion voice companion system based on human multi-dimensional state perception, characterized in that, include: The multi-dimensional state perception module is used to acquire user scene environment, emotional physiology, behavioral rhythm, terminal usage status and real-time interaction intent; The personalization and taboo memory module is used to store user preferences, historical interaction data, and preset taboo boundaries; The integrated decision-making module is used to generate interactive decision-making instructions based on multi-dimensional status data and user personalized characteristics. The same-source real-person voice library module is used to store voice materials with multiple emotions and tones recorded by the same real person; The low-interference control module is used to perform voice companionship, emotional soothing, reminders, or silent control according to preset rules for the lowest possible level of disturbance. The dual-mode security engine is used to implement a dual-mode operating architecture that prioritizes local offline operation and allows cloud access as an option, while also protecting user privacy data.

10. The system according to claim 9, characterized in that, The multi-dimensional state perception module includes at least one of the following: scene perception unit, voice emotion recognition unit, behavior posture perception unit, and circadian rhythm analysis unit.

11. The system according to claim 9, characterized in that, The system is applied to smartphones, smartwatches, tablets, smart cars, smart speakers, smart home terminals, or in-vehicle smart terminals.

12. The system according to claim 9, characterized in that, The voice companionship, emotional soothing, and reminder services can be performed by virtual characters, AI interactive roles, or digital humans. Their interaction logic, emotional response, voice expression, and triggering strategies strictly adhere to the user's multi-dimensional state, personalized characteristics, taboo bottom lines, and the rule of minimal disturbance.

13. A terminal device, characterized in that, It includes a processor and a memory, the memory storing a computer program, which, when executed by the processor, implements the adaptive low-interference voice companion method according to any one of claims 1 to 8.

14. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the adaptive low-interference voice companion method according to any one of claims 1 to 8.