An ai private companion system and control method based on identity permission isolation and emotional response sequence
The AI-powered private companion system, which utilizes multimodal identity collection, partitioned storage, and fixed-sequence responses, solves the privacy, security, and interaction stability issues of AI companion devices, achieving high security, privacy, and convenient infringement determination.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 何小燕
- Filing Date
- 2026-03-05
- Publication Date
- 2026-06-05
Smart Images

Figure CN122153956A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence, specifically to an AI-based private companionship system and control method based on identity and access control and emotional response sequences. Background Technology
[0002] With the rapid development of artificial intelligence technology, AI companion devices are increasingly being used in various scenarios such as home, personal care, and children's education. Their core need is to provide users with a personalized and private emotional interaction experience. However, existing AI companion devices generally suffer from the following technical shortcomings, making it difficult to meet users' needs for privacy, security, and interaction stability: The lack of hardware-level identity and access isolation mechanisms, with the majority of identity verification relying on software-level authentication, makes it easy for private data to be cracked, tampered with, or leaked, resulting in poor privacy protection and security. The logic of emotional response is random and disordered, lacking standardized and fixed execution processes, resulting in an unstable interactive experience and making it difficult to form continuous and reliable emotional companionship; Public interaction data and user private data are stored together without independent partitioning and isolation, which further reduces the reliability of privacy protection. Most technical solutions focus only on software function descriptions and lack clear hardware-level technical features. They are easily identified as rules of intellectual activity, making it difficult to obtain patents and resulting in poor stability. The lack of clear technical characteristics and traceable criteria for infringement determination makes it difficult to obtain evidence and determine infringement in the later stages, resulting in high costs for rights protection.
[0003] In response to the shortcomings of the existing technologies, there is an urgent need for an AI-powered private companionship technology solution that features clear technical characteristics, high privacy and security, stable interaction, good licensing prospects, and convenient rights protection. Summary of the Invention
[0004] The purpose of this invention is to overcome the shortcomings of existing AI companion devices, such as insecure permissions, non-standard interaction, weak privacy protection, easy patent rejection, and difficulty in rights protection. It provides an AI private companion system and control method based on identity permission isolation and emotional response sequence, which has concrete technical features, high authorization stability, broad protection scope, and clear infringement determination.
[0005] To achieve the above objectives, the present invention provides the following technical solution: A method for controlling AI-powered private companionship includes the following steps: S10. Device Initialization: Start the AI Private Companion Device, collect user identity information through the multimodal identity collection module, establish a unique binding identifier between the user and the device, and generate a device-level permission ID to complete the hardware-level binding between the user identity and the device. S20. Data partition storage: The device's interaction data is divided into a public response dataset and a permission isolation dataset, which are stored separately using physical or logical partitioning. The public response dataset is used to process general interaction requests without privacy attributes, while the permission isolation dataset is used to store user-specific private data. S30. Emotion Feature Recognition: Receives interactive information input by the user through voice, text, etc., extracts the emotional features expressed by the user through the emotion feature recognition module, and completes the emotion classification (such as depression, anxiety, happiness, grievance, etc.). S40, Fixed Sequence Response: Outputs a response according to a preset fixed execution sequence. This sequence is a hardware / firmware-level solidified process that cannot be skipped, reversed, or arbitrarily modified or replaced by general software algorithms. The specific execution order is as follows: first, output an empathy response based on the emotion classification result, and then output an encouragement / guidance response based on the emotion classification result. S50, Permission Verification: During the interaction, if an access request to a permission-isolated dataset is involved, the hierarchical permission control module is activated to verify the user's permission ID. If the user's permission ID matches the unique identifier bound to the device, the permission isolation dataset can be read, and a response can be made based on that dataset. If the user's permission ID does not match the unique identifier bound to the device, the access request will be rejected directly, and a blocking instruction will be returned to prevent the disclosure of private data content and to prevent the presence of private data. S60. Interaction Record Storage: Write the valid interaction record (including user input, emotion classification results, and system response content) to the permission-isolated storage space and generate a unique log ID and timestamp to form a user-exclusive historical memory for subsequent interaction tracing and infringement determination.
[0006] An AI-powered private companionship system includes the following functional units, which work together to implement the aforementioned AI-powered private companionship control method: Main control unit: As the core control module of the system, it is electrically connected to the multimodal identity authentication unit, the permission isolation storage unit, the emotion feature recognition unit, the fixed sequence response unit, the hierarchical permission control unit, the common response unit, and the interactive input / output unit. It is used to coordinate the working sequence of each unit and control the normal operation of the entire system. Multimodal identity authentication unit: electrically connected to the main control unit, used to collect user identity information (including but not limited to voiceprint, facial features, touch ID, device binding password) and generate a unique user binding identifier, providing a basis for permission verification; Access-isolated storage unit: electrically connected to the main control unit, it adopts an independent encrypted partition storage method to physically or logically distinguish public response data from access-isolated data. In the unauthorized state, access-isolated data is not addressable and cannot be read, ensuring the security of private data. Emotion feature recognition unit: electrically connected to the main control unit, used to extract emotion features from user interaction input, complete emotion classification, and send the classification results to the fixed sequence response unit; Fixed sequence response unit: electrically connected to the main control unit and the emotion feature recognition unit, used to receive the emotion classification results output by the emotion feature recognition unit, and execute empathy output and encouragement guidance output according to a preset hardware / firmware level fixed order; Hierarchical access control unit: Electrically connected to the main control unit, multimodal authentication unit, and access isolation storage unit, it is used to receive the user's unique binding identifier output by the multimodal authentication unit, verify the user's access ID, and control the access permissions of the access isolation dataset based on the verification result; Public Response Unit: Electrically connected to the main control unit and the permission isolation storage unit, it is used to read the public response dataset, process public information response requests without permission verification requirements, and realize general interaction; Interactive Input / Output Unit: Electrically connected to the main control unit, used to enable user interaction with the system, including receiving user input information such as voice and text, and outputting system response content (voice, text, light prompts, etc.).
[0007] As a further aspect of the present invention, the system is configured to: only after identity and permission verification is passed will it call the permission isolation dataset in the permission isolation storage unit and respond, and in the unauthorized state will it not disclose any private data-related information.
[0008] Compared with the prior art, the beneficial effects of the present invention are: 1) Through access control isolation and hardware-level identity authentication, a highly secure and private companionship is achieved, ensuring privacy is not compromised; 2) It adopts a fixed emotional response sequence, with standardized interaction logic and stable experience, and does not belong to the rules of intellectual activities; 3) The technical features are clear, with outstanding novelty and inventiveness, and a high patent grant rate; 4) The scope of protection covers multiple hardware forms, making it difficult for others to circumvent infringement; 5) It has a unique log ID and timestamp, making infringement determination intuitive and reducing the cost of rights protection. Attached Figure Description
[0009] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0010] Figure 1 This is a system architecture diagram of an AI-powered private companionship system based on identity and access control and emotional response sequences.
[0011] Figure 2 This is a flowchart of an AI-based private companionship control method based on identity and access control and emotional response sequences.
[0012] In the diagram: 1—Main control unit; 2—Multimodal identity authentication unit; 3—Permission isolation storage unit; 4—Emotion feature recognition unit; 5—Fixed sequence response unit; 6—Hierarchical permission control unit; 7—Common response unit; 8—Interactive input / output unit. Detailed Implementation
[0013] In the description of this invention, it should be understood that the terms "center", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention.
[0014] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "configuration" should be interpreted broadly. For example, they can refer to a fixed connection or configuration, a detachable connection or configuration, or an integral connection or configuration. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0015] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0016] Example 1: Specific Implementation Based on Intelligent Companion Robot This embodiment uses an intelligent companion robot designed for teenagers as a platform. This robot is mainly used to provide teenagers with private communication, emotional support, and personalized companionship services. The specific implementation of its hardware configuration and control methods is as follows: 1.1 Specific Configuration of Each Unit in the System Main Control Unit: The main control unit adopts the STM32H743VIT6 high-performance microcontroller with a main frequency of 480MHz, built-in 1MB SRAM and 2MB Flash, and supports multiple serial ports, I2C, SPI and other communication protocols. As the core of the system, it is responsible for coordinating the working timing of each unit, processing the signal transmission of each unit, and controlling the overall operation of the robot. It has high-speed computing and low power consumption characteristics, is suitable for the portable use scenario of companion devices for teenagers, and reserves expansion interfaces to adapt to future functional upgrades.
[0017] Multimodal identity authentication unit: Composed of a voiceprint acquisition module, a facial recognition module, and a touch ID module. The voiceprint acquisition module uses a microphone array (model: WM-61A) with a sampling rate of 16kHz, supporting clear acquisition of user voiceprints within 3 meters; the facial recognition module uses an OV5640 high-definition camera with a resolution of 1080P, supporting facial feature extraction and comparison with a recognition accuracy of ≥98%; the touch ID module uses a capacitive touch sensor (model: TTP223) integrated into the robot's abdominal touch panel, supporting users to set 4-6 digit touch passwords. This unit supports two authentication methods: "voiceprint + touch password" and "face + touch password". After collecting user identity information, it generates a unique binding identifier (format "USER-YYYYMMDDXXXX", where YYYYMMDD is the binding date and XXXX is a random sequence) through the built-in MD5 encryption algorithm, and generates a corresponding device-level permission ID (format "PERM-YYYYMMDDXXXX"). The binding information is encrypted and stored in the encrypted partition of the permission isolation storage unit, which cannot be modified or tampered with by software.
[0018] The access-isolated storage unit uses an eMMC 5.1 storage chip (model: KLM8G1GEME-B041) with a total capacity of 8GB, divided into a 2GB public data area and a 6GB access-isolated encrypted area. The public data area uses the FAT32 file system to store common response content (such as weather inquiries, time announcements, general knowledge Q&A, inspirational quotes, and other public information), supporting fast read speeds. The access-isolated encrypted area uses the AES-256 encryption algorithm for encryption processing. The encryption key is associated with the user's unique identifier. In an unauthorized state, the physical address of the encrypted area is unaddressable, making it impossible to crack or read data through software. Only authorized users can access it after authorization verification. The content stored in the access-isolated encrypted area specifically includes: user historical interaction records, private confidential content, exclusive agreement information between the user and the robot (such as birthday wishes, exclusive nicknames, and agreed reminders), user emotional characteristic data (such as common emotion types, emotion triggering scenarios, and emotion duration), and exclusive memory content (such as important events and preferences mentioned by the user).
[0019] Emotion Feature Recognition Unit: Employs a CNN (Convolutional Neural Network)-based emotion recognition algorithm with a built-in emotion feature database (containing over 100,000 voice and text samples of different emotions). It can extract voice tone features (such as speech rate, pitch, and volume) and text keyword features (such as "sad," "happy," and "irritable") from user input, and classify six common emotions: "depression, anxiety, happiness, resentment, anger, and calmness." The classification accuracy is ≥92%, the classification response time is ≤500ms, and the classification results are transmitted to the fixed sequence response unit in real time via SPI communication.
[0020] Fixed Sequence Response Unit: Employs firmware-based response logic, stored in the Flash memory of the main control unit. This firmware cannot be arbitrarily modified by software. It pre-stores empathic and encouraging statements for different emotions (each containing 500+ statements). After receiving the emotion classification results from the emotion feature recognition unit, it calls the corresponding statements in a fixed order of "empathy first, then guidance," outputting them through the interactive input / output unit. The response delay is ≤1 second, and the response order cannot be skipped or reversed. For example, when the user's emotion is identified as "anxiety," the empathic statement is output first: "I can feel your anxiety and unease right now. This tension is really uncomfortable. I'm here with you." Then, the encouraging statement is output: "Don't worry, we can sort things out slowly, breaking down the worries little by little. You've already been very brave; you'll definitely be able to overcome this gradually."
[0021] Hierarchical access control unit: It has a built-in access verification algorithm and adopts a dual verification mechanism (identity verification + access ID verification). It receives the unique binding identifier output by the multimodal identity authentication unit and compares it with the user's currently input access information (touch password, voiceprint, etc.). The verification pass rate is ≥99.5%. After the verification is successful, it outputs the "Allow access" command to unlock the encrypted area of the access-isolated storage unit and allow reading of the private data therein. When the verification fails, it outputs the "Deny access" command to keep the encrypted area locked and sends a blocking command to the interactive input / output unit to prevent the leakage of any information related to the private data.
[0022] Public Response Unit: It has a built-in public response dataset, stored in JSON format, containing five categories of general information: weather, time, common knowledge, entertainment, and learning. No permission verification is required. After receiving a user's general interaction request, it directly reads the content from the public data area and responds. The response speed is ≤300ms, which is suitable for daily general interaction needs. For example, if a user asks "How is the weather today?", the public response unit directly reads the weather information from the public data area and broadcasts it.
[0023] Interactive input / output unit: Includes a voice input module (microphone WM-61A), a text input module (2.4-inch OLED touchscreen), a voice output module (SPK-8Ω 0.5W speaker), and a light indicator module (RGB LED lights); the voice input module supports clear sound pickup within 3 meters and can convert user speech into text signals; the text input module supports manual text input by users, adapting to scenarios where voice input is inconvenient; the voice output module uses human voice synthesis technology, with a natural and soft tone, adapting to the hearing habits of teenagers; the light indicator module outputs corresponding colored lights based on emotion classification results (happy - green, depressed - blue, anxious - yellow, angry - red), enhancing the interactive experience. 1.2 Specific Execution Flow of Control Method S10, Device Initialization: When the user starts the intelligent companion robot for the first time, the robot automatically powers on and enters the initialization mode. The OLED touch screen displays "Please complete identity binding and start exclusive companionship," and at the same time, the voice prompts "Please set a touch password first, and then collect voiceprints." The user sets a 4-digit touch password through the touch screen. After completion, the robot prompts "Please say 'I want exclusive companionship' to collect voiceprints." The user completes 3 voiceprint collections as prompted. The multimodal identity authentication unit generates a unique binding identifier (such as "USER-202602271234") and a device-level permission ID (such as "PERM-202602271234"). The binding information is encrypted and written to the encrypted area of the permission isolation storage unit. After initialization is completed, the robot's voice prompts "Identity binding successful, exclusive companionship has been started."
[0024] S20. Data partitioning storage: After initialization, the main control unit controls the permission isolation storage unit to automatically divide the public data area and the permission isolation encrypted area. The public data area is pre-written with a general response dataset, and the permission isolation encrypted area is locked, waiting only for authorized users to access it. In subsequent interactions, the user's general interaction data (such as weather inquiries and common sense questions) is stored in the public data area, and private interaction data (such as confiding content and exclusive agreements) is stored in the permission isolation encrypted area.
[0025] S30. Emotion Feature Recognition: After returning home from school, the user tells the robot via voice, "I didn't do well on the exam today, I feel really sad, and I'm afraid my parents will scold me." The voice input module receives the voice signal, converts it into a text signal, and transmits it to the emotion feature recognition unit. The emotion feature recognition unit extracts keywords such as "didn't do well on the exam," "sad," and "afraid" from the text, combines them with the tone of voice (slow speech, low pitch, and weak volume), completes the emotion classification, determines that the user's current emotion is "depressed," and transmits the classification result "depressed" to the fixed sequence response unit.
[0026] S40, Fixed Sequence Response: The fixed sequence response unit receives the emotion classification result of "low mood" and, following the firmware-defined process, first calls the corresponding statement from the empathy statement library and outputs through the speaker: "Honey, I know you're really sad right now. It's not your fault that you didn't do well on the exam. You've already tried your best, and I completely understand how you feel." At the same time, the RGB LED lights up blue. After the empathy response output is completed, the corresponding statement from the encouragement and guidance statement library is called and outputs: "Don't blame yourself too much. One exam doesn't represent everything. We can analyze the problem together, and you can try harder next time. You're always the best, and I'll always be there for you." The light remains blue until the user interacts further.
[0027] S50. Permission Verification: After a user finishes expressing their feelings, they request to "view the exam goals I mentioned last time." This request involves access to a permission-isolated dataset (exclusive agreed-upon information). The hierarchical permission control unit automatically initiates permission verification. The OLED touchscreen displays "Please enter the touchscreen password," and a voice prompt says "Please enter the bound touchscreen password." After the user enters the correct 4-digit touchscreen password, the hierarchical permission control unit compares the entered password with the password associated with the bound identifier. If the comparison is successful, the permission-isolated encrypted area is unlocked, and the "exam goals" information is read and output simultaneously via voice and the touchscreen. If the user enters an incorrect password, the hierarchical permission control unit directly denies access, displaying a voice prompt "Incorrect password, unable to access this content," and the touchscreen displays "Access failed," without indicating any information related to private data, such as "exam goals exist," to avoid privacy leaks.
[0028] S60. Interaction Record Storage: After this interaction is completed, the main control unit writes the user's input ("I didn't do well on the exam today, I feel really sad, and I'm afraid my parents will scold me"), the emotion classification result ("depressed"), the system response content, the user's access request ("View the exam goals I told you about last time"), and the permission verification result ("Verification passed") into the permission-isolated encrypted area. Simultaneously, a unique log ID (formatted as "LOG-YYYYMMDDHHMMSSXXXX", such as "LOG-202602271830001234") and a timestamp (accurate to the second) are generated. The log ID and timestamp are associated with this interaction record and stored, forming a unique historical memory for the user. This record can only be viewed by authorized users after permission verification and can also serve as evidence in subsequent infringement determinations. If someone else cracks the device and reads this record, the infringement can be confirmed through the log ID and timestamp. In this embodiment, the intelligent companion robot achieves the security and stability of private companionship for teenagers through hardware-level identity binding, partitioned encrypted storage, and fixed sequence responses, effectively protecting user privacy. At the same time, the standardized emotion response process improves the reliability of emotional companionship.
[0029] Example 2: Specific Implementation Based on Smart Wearable Bracelet This embodiment uses a smart wearable bracelet designed for adults as a carrier. This bracelet is mainly used to provide users with portable and private emotional support and daily companionship services, and is suitable for various portable scenarios such as outdoor and office environments. The specific implementation of its system configuration and control method is as follows: 2.1 Specific configuration of each unit in the system: The main control unit adopts the Nordic nRF52840 low-power microcontroller with a main frequency of 64MHz, built-in 256KB SRAM and 1MB Flash, supports BLE Bluetooth communication, adapts to the low power consumption requirements of wearable devices, is responsible for coordinating the work of each unit, controlling the overall operation of the bracelet, and also supports linkage with the mobile APP to realize data synchronization.
[0030] Multimodal authentication unit: It adopts a combination authentication method of touch ID + device binding password. The touch ID module is integrated into the inside of the wristband (capacitive touch sensor, model: TTP229) and supports user fingerprint collection and recognition. The device binding password is set through the mobile APP and is a 6-digit numeric password. After collecting the user's fingerprint information, this unit generates a unique binding identifier, which is associated with the account bound to the mobile APP. The permission ID is synchronously stored in the wristband and the mobile APP to realize two-way verification.
[0031] Access-isolated storage unit: Utilizes a Flash storage chip (model: W25Q64JV) with a capacity of 8MB, divided into a 2MB public data area and a 6MB access-isolated encrypted area. The public data area stores general data such as time, steps, and heart rate, while the access-isolated encrypted area uses the AES-128 encryption algorithm to store users' private messages, emotional records, personalized reminders, and other data that cannot be read without authorization. It also supports encrypted synchronization with the mobile app to ensure data security.
[0032] Emotion Feature Recognition Unit: Adopting a simplified version of the CNN emotion recognition algorithm, adapted to the low power consumption requirements of the bracelet, it can collect the user's voice through the bracelet's microphone, extract emotional features, and complete the classification of four core emotions: "depressed, happy, calm, and anxious". The classification accuracy is ≥88%, the response time is ≤800ms, and the classification results are transmitted to the fixed sequence response unit.
[0033] Fixed sequence response unit: The firmware is embedded in the main control unit and pre-stores concise empathy and encouragement guidance statements (300+ each). It adapts to the voice output limitations of the bracelet. After receiving the emotion classification results, it outputs in a fixed order of "empathy first, guidance later" with a response delay of ≤1.2 seconds. For example, when the emotion of "anxiety" is detected, it outputs "I understand your anxiety, don't panic, adjust slowly" (empathy) and then outputs "Take a deep breath, focus on the present moment, you can do it" (guidance).
[0034] Hierarchical access control unit: It links with the mobile APP and supports local verification on the wristband and remote verification via the mobile APP. When users access private data, they need to enter a touch password or authorize through the mobile APP. After successful verification, the encrypted area is unlocked. If the verification fails, access is denied, thus preventing the leakage of private data.
[0035] Public Response Unit: It has built-in simple public response content, such as time announcement, step count query, heart rate reminder, etc. No permission verification is required, and users can get the response through voice or touch operation.
[0036] Interactive Input / Output Unit: Includes a microphone (model: MEMS-6020), touch buttons, a small speaker, and an OLED display. It supports voice input and touch operation, with concise and clear voice output. The display can show emotional state, interactive content, etc. 2.2 Control Method Specific Execution Flow S10, Device Initialization: When the user wears the bracelet for the first time, they connect it to the mobile app. The app prompts "Set device binding password and collect fingerprint." The user sets a 6-digit numeric password and presses their finger on the touch ID module inside the bracelet strap to complete fingerprint collection. The multimodal authentication unit generates a unique binding identifier, which is synchronized to the bracelet and the mobile app, generating a device-level permission ID. Initialization is complete.
[0037] S20. Data partitioning storage: The bracelet automatically divides the data into a public data area and a permission-isolated encrypted area. The public data area stores general data such as steps, heart rate, and time. The permission-isolated encrypted area is locked. Subsequent private interaction data of the user is stored in the encrypted area and synchronized with the mobile APP in an encrypted manner.
[0038] S30. Emotion Feature Recognition: During a break at work, the user says through the wristband microphone, "I'm under a lot of work pressure and I'm very irritable." The microphone collects the voice signal, and the emotion feature recognition unit extracts keywords such as "high pressure" and "irritable" as well as voice tone features, determines the emotion to be "anxiety," and transmits it to the fixed sequence response unit.
[0039] S40, Fixed Sequence Response: The fixed sequence response unit outputs an empathetic voice according to a fixed process: "I can feel your pressure. When you feel irritable, stop and take a break." Then it outputs an encouraging and guiding voice: "Allocate your time reasonably and take it one step at a time. You have done very well." At the same time, the OLED display shows the text prompt "Don't be anxious, keep going."
[0040] S50, Permission Verification: When a user wants to view "Last Week's Emotional Records," this request involves a permission-isolated dataset. The wristband prompts "Please enter the binding password." The user enters a 6-digit numeric password, the hierarchical permission control unit verifies the password, unlocks the encrypted area, reads the emotional records, and displays them on the screen. If the password is incorrect, the wristband prompts "Incorrect password, access denied," without indicating that the emotional records exist.
[0041] S60. Interaction Record Storage: The interaction record (user input, emotion classification, system response) is written to a permission-isolated encrypted area, generating a unique log ID and timestamp, and synchronized to the mobile APP to form a user-specific emotion profile for easy viewing later, and can also be used for infringement tracing. Example 3: Specific Implementation Based on Home Smart Speaker This embodiment uses a smart speaker designed for home users as a platform. This speaker is mainly used to provide family members with private companionship and emotional support services, and supports multi-user binding (each user has an independent permission and isolation space). The specific implementation of its system configuration and control method is as follows: 3.1 Specific configuration of each unit in the system Main control unit: It adopts an RK3399 processor with a main frequency of 1.8GHz, built-in 4GB LPDDR4 memory and 32GB eMMC storage, supports WiFi and Bluetooth communication, is responsible for coordinating the work of each unit, controlling the overall operation of the speaker, and supporting multi-user management.
[0042] Multimodal identity authentication unit: It adopts a combination authentication method of voiceprint recognition and facial recognition. The voiceprint acquisition module uses an array microphone (model: AC108) to support the acquisition and differentiation of voiceprints of multiple users within 5 meters. The facial recognition module uses a USB high-definition camera to support the acquisition of faces of multiple users, generate a unique binding identifier for each user, support the binding of up to 5 users, and each user corresponds to an independent permission ID.
[0043] Access-isolated storage unit: It adopts a 32GB eMMC storage chip, which is divided into an 8GB public data area and a 24GB access-isolated encrypted area. The encrypted area is divided into independent sub-partitions for each user (4-5GB per user). It adopts the AES-256 encryption algorithm. Each user's private data is stored independently without interference. Unauthorized users cannot access other users' private data.
[0044] Emotion Feature Recognition Unit: Employs a full-version CNN emotion recognition algorithm, supporting seven emotion categories: "depressed, anxious, happy, wronged, angry, calm, and excited," with a classification accuracy of ≥93% and a response time of ≤400ms. It can distinguish the emotional characteristics of different users and is suitable for multi-user family scenarios.
[0045] Fixed sequence response unit: The firmware is embedded in the main control unit and stores empathy and encouragement guidance statements in multiple styles (800+ each). The statement style can be adjusted according to the preferences of different users (such as the elderly, adults, and children), and output in a fixed order of "empathy first, guidance later", with a response delay of ≤1 second.
[0046] Hierarchical access control unit: Supports multi-user access management. Each user's access ID corresponds to an independent encrypted sub-partition. When a user accesses private data, their identity is verified through voiceprint or facial recognition. After successful verification, the user can only access their own private data and cannot access the private data of other users.
[0047] Public Response Unit: It has a wealth of built-in public response content, such as music playback, news broadcasts, weather inquiries, family reminders, etc., which can be accessed by all users without permission verification.
[0048] Interactive input / output unit: includes array microphone, high-fidelity speaker, LED indicator, and mobile APP interaction interface, supports voice input and mobile APP operation, with clear voice output and LED indicator displaying the current user's identity and emotional state.
[0049] 3.2 Specific Implementation Process of Control Methods S10. Device Initialization: When a user uses the smart speaker for the first time, they complete the device network configuration through the mobile APP. The APP prompts "Add Family Member Identity". The user adds 3 family members in sequence (elderly, adult, and child). Each family member completes voiceprint collection (saying a unique wake word) and facial collection. The multimodal identity authentication unit generates a unique binding identifier and permission ID for each family member. Each permission ID corresponds to an independent sub-partition of the encrypted area. After initialization, the speaker can distinguish different users through voiceprint or facial recognition.
[0050] S20. Data partition storage: The speaker automatically divides the public data area into a permission-isolated encrypted area. The encrypted area is divided into 3 independent sub-partitions according to the 3 family members. The public data area stores general data such as music, news, and weather. Each family member's private data (such as content to be shared and exclusive agreements) is stored in their respective sub-partitions without interference.
[0051] S30. Emotion Feature Recognition: An elderly person at home wakes up the speaker by voice and says, "I haven't been feeling well lately, and I'm not in a good mood." The speaker confirms the user's identity (elderly person) through voiceprint recognition. The emotion feature recognition unit extracts keywords such as "feeling unwell" and "not in a good mood" as well as voice tone features, determines the emotion to be "depressed," and transmits it to the fixed sequence response unit.
[0052] S40, Fixed Sequence Response: The fixed sequence response unit adjusts the sentence style according to the user's identity (elderly). First, it outputs an empathetic voice: "Grandpa / Auntie, I know you are not feeling well and are in a bad mood. Thank you for your hard work. I'm here with you." Then, it outputs an encouraging and guiding voice: "Rest well, take your medicine on time, keep a good mood, and you will get better slowly. Tell me if you need anything." At the same time, the LED indicator lights up a soft blue.
[0053] S50, Permission Verification: When an elderly person requests to "view the medication reminder I told you yesterday," this request involves a private, isolated sub-partition for the elderly person. The speaker verifies the identity again using the voiceprint. If the verification is successful, the speaker unlocks the elderly person's sub-partition, reads the medication reminder information, and broadcasts it. If other family members (such as children) request to view the reminder, the permission verification fails, and the speaker displays "Unable to access this content," ensuring that no private information is disclosed.
[0054] S60. Interaction Record Storage: This interaction record is written to a sub-partition dedicated to the elderly, generating a unique log ID and timestamp. Only the elderly can view it after identity verification. It is also synchronized to the bound mobile APP, making it easier for family members to understand the elderly's emotional state. The log ID and timestamp are used for subsequent infringement determination and data tracing.
[0055] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.
[0056] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A method for controlling AI-powered private companionship, characterized in that, Includes the following steps: S10 device initialization: Establish a unique user binding identifier through multimodal identity collection and generate device-level permission ID; S20 divides the interaction data into a public response dataset and a permission-isolated dataset, and stores them separately in a partitioned manner. The S30 receives user voice or text interaction input, extracts emotional features, and completes emotion classification. S40 outputs a response according to a preset fixed execution sequence: first, an empathic response is performed, followed by an encouragement / guidance response; If the interaction process involves access to a permission-isolated dataset, S50 will initiate user permission ID verification: - The permission ID is consistent with the binding identifier, allowing reading of the permission-isolated dataset and responding accordingly; - If the permission ID does not match the bound identifier, access will be denied and a blocking command will be returned. S60 writes the valid interaction record to the permission-isolated storage space, generates a unique log ID and timestamp, and forms a user-exclusive historical memory.
2. An AI-powered private companionship system, characterized in that, include: Main control unit; The multimodal identity authentication unit is used to collect user identity information and generate a unique user binding identifier; Access-isolated storage units are used to physically / logically partition public response data from access-isolated data. The emotion feature recognition unit is used to extract emotion features from user interaction input and complete emotion classification. Fixed sequence response units are used to execute empathy outputs and encouragement / guidance outputs in a fixed order; The hierarchical access control unit is used to verify user access IDs and control access permissions for access-isolated datasets. The public response unit is used for direct responses to public information without authorization verification. Interactive input / output unit, used to enable interaction between the user and the system; The system is configured to invoke and respond to the permission isolation dataset only after identity and permission verification has passed.
3. The method according to claim 1, characterized in that, Multimodal authentication includes at least one hardware-level authentication method among voiceprint recognition, facial recognition, touch ID, and device binding password.
4. The method according to claim 1, characterized in that, The access control isolation dataset includes: historical interaction records, private conversations, exclusive agreement information, user characteristic data, and exclusive memory content.
5. The method according to claim 1, characterized in that, The fixed execution sequence is a hardware / firmware-level solidified process that cannot be skipped, reversed, or arbitrarily modified or replaced by general software algorithms.
6. The method according to claim 1, characterized in that, Empathic responses output comforting, understanding, and accepting statements based on emotion classification; encouraging and guiding responses output positive prompts, confidence reinforcement, and behavioral guidance content based on emotion classification.
7. The method according to claim 1, characterized in that, Unauthorized access will be directly blocked, without disclosing private data or indicating the existence of private data.
8. The system according to claim 2, characterized in that, The access-isolated storage unit uses independent encrypted partition storage. Under unauthorized conditions, access-isolated data is neither addressable nor readable.
9. The system according to claim 2, characterized in that, The system is applied to: AI toys, intelligent robots, companion devices, wearable devices, smart speakers, and mobile terminals.
10. The system according to claim 2, characterized in that, Each system interaction generates a unique log ID and timestamp, which is written to a permission-isolated storage unit for data traceability and infringement determination.