Special AI accompanying housekeeper system and method for elder based on AI large model and voiceprint recognition model
Through AI-driven voiceprint recognition and multi-agent collaboration systems, the problem of traditional security and intelligent voice assistants being unable to combine security and emotional interaction has been solved, enabling comprehensive services of 24/7 security monitoring and multi-scenario companionship.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGDONG ZHUOSHANG NETWORK TECH CO LTD
- Filing Date
- 2026-02-24
- Publication Date
- 2026-05-15
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies cannot effectively combine security monitoring with emotional companionship. Traditional security systems infringe on privacy and lack emotional interaction, while smart voice assistants cannot accurately identify identities or understand complex life scenarios, resulting in the failure to comprehensively address the safety and emotional needs of elderly people living alone.
Employing an AI-driven voiceprint recognition and multi-agent collaboration system, it provides 24/7 security monitoring and multi-scenario companionship services through environmental sound analysis, identity recognition, emotion understanding, and large language model processing.
It achieves 24/7 seamless security monitoring, accurately identifies the elderly's identity and emotions, and provides personalized multi-scenario services to meet the elderly's needs for daily living assistance and emotional comfort.
Smart Images

Figure CN122045371A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of intelligent elderly care technology, artificial intelligence and voice processing technology, specifically to an AI companionship management system and method for the elderly that integrates sound scene recognition, voiceprint recognition and large language model technology. Background Technology
[0002] Elderly people living alone face the dual challenges of inadequate security monitoring and lack of emotional companionship. Existing technological solutions are clearly fragmented: security monitoring systems (such as cameras and infrared sensors) only focus on security incidents, easily infringing on privacy and failing to provide emotional interaction; traditional intelligent voice assistants, while capable of simple conversations, lack accurate user identification, the ability to understand complex life scenarios (such as unusual sounds), and the capacity for deep emotional cognition and personalized companionship. Therefore, there is an urgent need for a comprehensive solution that integrates proactive security warnings, identity recognition, and emotionally-driven intelligent interaction. Summary of the Invention
[0003] This invention aims to overcome the shortcomings of existing technologies and provide an AI butler system and method that can ensure the safety of the elderly around the clock and provide human-like, multi-scenario companionship. The technical solution proposed in this invention can understand the environmental sound context, accurately identify the identity and emotions of the elderly, and provide comprehensive services ranging from daily living assistance to emotional comfort through multi-agent collaboration driven by a large language model. To achieve the above objectives, the present invention adopts the following technical solution: The core method flow of this invention is as follows: Figure 1 As shown, it includes seven main steps: Step 1: Audio Data Acquisition. The system continuously collects raw audio streams from the elderly person's living space using hardware devices such as microphones. Step 2: Preliminary sound classification. Using the trained sound classification model, the audio is analyzed in real time to initially distinguish it into ambient sounds (such as footsteps and collision sounds) and human voices. Step 3: Ambient Sound Early Warning Processing. For the categorized ambient sounds, the system quickly compares them with a pre-built abnormal warning sound tag library (containing dangerous sound samples such as falling, breaking glass, etc.). Once a match is successful, the system immediately generates an early warning message and notifies preset emergency contacts via SMS, APP push, etc., to achieve proactive safety monitoring. Step 4: Voice Identification. For the categorized voices, the system activates a high-precision voiceprint recognition model to extract the speaker's voiceprint features and compares them with the voiceprint feature database of registered family members. This step accurately determines whether the speaker is a registered user or a visitor. If the speaker is a registered user, the system can accurately identify their specific identity (e.g., father, mother), laying the foundation for subsequent personalized services. Step 5: Personality Profiling and Intelligent Interaction Processing. Once the speaker is identified as a registered senior citizen, the system uses Automatic Speech Recognition (ASR) technology to convert the senior citizen's voice query or dialogue into text. This text is uploaded to the large language model processing unit. After speech-to-text conversion, based on the personality model algorithm, multimodal analysis is performed on the text content (vocabulary, sentence structure) and voice characteristics (tone, speech rate) to perceive the user's emotions (such as happiness, sadness, anxiety) and speaking style in real time, and dynamically update the user's personality profile (such as "prefers concise responses" and "responds with a calm and gentle demeanor"). Step 6: Intent Understanding and Response Generation. The large language model performs deep semantic understanding and intent recognition on the text. Based on the identified different intents (such as checking the weather, expressing emotions, reminiscing about the past, seeking help, etc.), the large language model integrates the current intent, the real-time updated dynamic personality profile, and historical interaction records stored in the database (M5), and can automatically continue past chat topics, calling the corresponding functional agents for collaborative processing. For example... Figure 3 As shown, these intelligent agents cover communication scenarios (such as initiating video calls), emotional cognition (such as emotion recognition and reassurance), emotional communication scenarios (such as chatting and reading news), and daily life scenarios (such as reminding someone to take medication and controlling home appliances). After processing, each intelligent agent returns the results to the large language model, which integrates and generates coherent, natural, and empathetic response text. Step 7: Speech Synthesis and Output. The system uses high-quality text-to-speech (TTS) technology to convert the generated response text into natural and friendly speech. This speech signal is sent back to the original hardware device (such as a smart speaker) for playback, completing a full interactive loop. Accordingly, the present invention also provides a system for implementing the above method, the architecture of which is as follows: Figure 2 As shown, the system mainly consists of an environmental sensor module (M1), a voiceprint recognition module (M2), an AI large-scale model processing module (M3), a user profile and historical interaction module (M4), and a response execution module (M5), connected in sequence. These five modules work together to realize the entire process from environmental perception to intelligent decision-making and then to multimodal response. (iv) Description of the attached drawings
[0004] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Figure 1 The following is a detailed flowchart of the AI companionship management method for the elderly based on a large AI model and a voiceprint recognition model provided by this invention. Figure 2 The present invention provides an architecture diagram of an AI companionship and management system and method for the elderly based on a large AI model and a voiceprint recognition model. Figure 3 The multifunctional application scenario mind map provided by the system of this invention demonstrates its application in different dimensions such as communication, social interaction, emotion, spirit, and life. Figure 4 : A schematic diagram of the application scenarios provided by the system of this invention. (V) Detailed Implementation Methods
[0005] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. The described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the present invention. See Figure 1 , Figure 2 , Figure 3 and Figure 4 When implementing the system of this invention, the hardware devices are deployed in areas where the elderly frequently move around, such as the living room and bedroom. The environmental sensor module (M1) is responsible for audio acquisition. The acquired data first enters the voiceprint recognition module (M2). The sound classification unit (M201) of this module determines whether the current sound is ambient noise or someone speaking. Assuming a loud "bang" is detected, the warning matching unit (M202) will immediately compare it with the "fall sound" sample library. If the features are highly similar, the warning process is triggered instantly, and the communication unit of the response execution module (M5) sends an alarm to the child's mobile phone: "Possible fall, please check immediately." If the system detects that an elderly person is saying, "I'm feeling a bit down today, I want to hear something cheerful," the voice classification unit (M201) identifies it as a human voice, and the voiceprint feature extraction and matching unit (M203) confirms it is Grandma Zhang herself. The speech is then sent to the AI large-scale model processing module (M3) for text recognition. The large language model processing unit understands that this is an expression of emotion and a search for comfort (corresponding to...). Figure 3(This relates to the scenarios of "emotional cognition" and "spiritual communication"). Therefore, it simultaneously coordinates the emotional communication agent (M302) (providing empathetic responses) and the knowledge-based question-answering agent (M303) (retrieving funny jokes or opera excerpts). The large language model merges the outputs of both, generating the following text: "Grandma Zhang, hearing you say that makes me anxious too. Let me tell you something interesting, or how about we listen to your favorite Huangmei Opera, 'The Female Imperial Envoy'?" This text is sent to the response execution module (M5), where the speech synthesis unit converts it into a concerned tone of voice, which is then played through a speaker. The system may record this emotional distress event. The system utilizes a personality profile and topic continuation function to create a "caring for children and enjoying life" personality profile (M4) for Grandma Zhang. When she inquires about her daughter's meals, the system recognizes her identity and voice, analyzes her concerned emotions, reinforces the profile's features, and searches historical interaction records (M4) to find the topic of "daughter learning to ride a bicycle." The system generates a response: "Grandma, don't worry, she should have eaten. You said last time she fell while learning to ride a bike but didn't cry, she's so strong." This response, played through voice synthesis, demonstrates personalized care. For everyday needs, such as when an elderly person says, "Please turn on the living room light for me," the system, after recognizing the person's identity and intent, will send a light-on command to the smart lighting fixture through the smart home control interface (M501) of the response execution module (M5). The implementation cases demonstrate how this invention seamlessly integrates security alerts, identity verification, intent understanding, and multi-scenario services. This invention is not just a tool, but a smart, secure, and empathetic AI-powered solution for accompanying the elderly. In summary, the above description is merely a preferred embodiment of the invention and does not constitute a limitation thereof. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this invention should be included within the protection scope of this invention.
Claims
1. The AI companionship management system and method for the elderly based on AI large-scale models and voiceprint recognition models, characterized in that, In step 3, the abnormal warning sound tag library contains samples of falling sounds, breaking sounds, and abnormal cries for help.
2. The AI companionship management system and method for the elderly based on AI large-scale models and voiceprint recognition models is characterized in that, In step 4, if the voice is identified as that of an unfamiliar person, a safety alert record is generated, and a decision is made based on preset rules as to whether to notify the guardian. If the voice is that of a registered user, it is matched using voiceprint features. The voiceprint feature database contains voiceprint templates of at least different family members such as the father and mother, enabling the system to accurately identify and distinguish specific individuals.
3. The AI companionship management system and method for the elderly based on AI large-scale models and voiceprint recognition models are characterized in that, In step 5, if the user is confirmed to be a registered user, the voice is parsed into text using a voice recognition model and uploaded to a large language model for processing to generate response text. This process combines the user's historical interaction preferences and real-time voice emotion to generate highly personalized response content.
4. The AI companionship management system and method for the elderly based on AI large-scale models and voiceprint recognition models is characterized in that, In step 6, the large language model can call or manage external smart devices based on the intent recognition results, and perform control operations including health monitoring devices, security monitoring devices, and smart home devices.
5. The AI companionship management system and method for the elderly based on AI large-scale models and voiceprint recognition models is characterized in that, In step 6, the personality big model algorithm quantifies and assesses emotional state and speaking style by analyzing word choice, sentence structure, and changes in tone, speed, and energy in the speech signal.
6. The AI companionship management system and method for the elderly based on AI large-scale models and voiceprint recognition models is characterized in that, In step 6, the analysis of the elderly’s historical interaction data, health records and real-time voice emotions based on the large language model and personality model generates personalized response text. It will actively query and associate the topic content of the user’s most recent chat, thereby achieving the coherence of the conversation and the automatic continuation of the topic.
7. The AI companionship management system and method for the elderly based on AI large-scale models and voiceprint recognition models is characterized in that, include: The environmental sensor module (M1) is used to acquire audio data from hardware devices; The voiceprint recognition module (M2) is connected to the environmental sensor module (M1) and is used to perform sound classification, environmental sound warning matching and human voiceprint recognition. The AI large model processing module (M3) is connected to the voiceprint recognition module (M2), and includes a voice recognition unit, a large language model processing unit and multiple functional intelligent agents, used to convert speech into text, understand intent and generate response text; The user profile and historical interaction module (M4) is connected to the AI big model processing module (M3) and is used to store and maintain the personality profile data of registered users (including personality preferences, emotional patterns, and preference tags) and historical interaction records (including dialogue topics, time, and device control records) to provide data support for personalized responses to the big language model. The response execution module (M5), connected to the user profile and history interaction module (M4), includes a speech synthesis unit and a communication control unit, used to convert text responses into audio and output them, or to send warnings and control commands.
8. The AI companionship management system and method for the elderly based on AI large-scale models and voiceprint recognition models, characterized in that, The voiceprint recognition module (M2) includes: The sound classification unit (M201) is used to distinguish between ambient sound and human voice; The early warning matching unit (M202) is used to compare ambient sound with the early warning tag library; The voiceprint feature extraction and matching unit (M203) is used to extract, identify, and compare human voiceprint features.
9. The AI companionship management system and method for the elderly based on AI large-scale models and voiceprint recognition models, characterized in that, The functional agents in the AI large model processing module (M3) include at least: daily reminder agent (M301), emotional communication agent (M302), knowledge question answering agent (M303), and communication assistance agent (M304).
10. The AI companionship management system and method for the elderly based on AI large-scale models and voiceprint recognition models, characterized in that, The response execution module (M5) is also connected to a smart home control interface (M501) for controlling external devices according to the instructions of the AI large model processing module (M3).