AI-based nursing bed and digital human interaction method and system thereof

By using an AI-based nursing bed system combined with digital human interaction technology, the problem of the traditional nursing bed's monotonous interaction methods has been solved, enabling personalized and emotional interaction, and improving the patient's user experience and nursing efficiency.

CN121661252APending Publication Date: 2026-03-13GUANGZHOU LIJIE MEDICAL EQUIP TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-09
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Existing nursing beds lack convenient and emotionally engaging interaction methods between patients and the bed, failing to meet patients' needs for emotional support and information access. Furthermore, their responses are relatively simplistic and cannot be intelligently controlled based on patients' individual needs.

Method used

The AI-based nursing bed receives digital family member summoning commands from the user via an interactive trigger module. It then uses a digital human interaction terminal to retrieve the corresponding virtual avatar model of the digital family member, enabling interaction with the user and control of the nursing bed itself. This includes a display screen, an audio acquisition module, an audio output module, and a nursing bed control module, generating a realistic virtual avatar for natural interaction with the user.

Benefits of technology

It enhances the intelligence level of nursing beds, provides a personalized, convenient and humanized nursing experience, alleviates patients' loneliness, enhances users' confidence in rehabilitation, and improves the efficiency of nursing work, and has broad application prospects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121661252A_ABST
    Figure CN121661252A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent nursing, and particularly discloses an AI-based nursing bed and a digital human interaction method and system thereof. The interaction triggering module is used for receiving a digital relative calling instruction input by a user; the digital person interaction terminal is arranged on the bedside side of the nursing bed body and used for calling a corresponding digital relative virtual image model stored in a digital person model library based on the digital relative calling instruction and interacting with a user or controlling the nursing bed body based on the corresponding digital relative virtual image model; comprising a display screen, an audio acquisition module, an audio output module and a nursing bed body control module, the use convenience and the humanization degree of the nursing bed are improved, and the experience feeling and the comfort degree of the user in the nursing process are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent nursing technology, and in particular to an AI-based nursing bed and its digital human interaction method and system. Background Technology

[0002] In the field of healthcare, nursing beds are crucial equipment for patients during treatment and rehabilitation. Improving their functionality and patient experience has a profound impact on patients' physical and mental health and the recovery process. With the increasing aging population and rising demands for higher quality healthcare, the functions of traditional nursing beds are no longer sufficient to meet the increasingly diverse needs of patients. The rapid development of artificial intelligence (AI) technology has brought new opportunities to the healthcare field.

[0003] However, existing nursing beds lack convenient and emotionally engaging interaction methods for patients, failing to meet their needs for emotional support and information access. When patients are in hospitals or nursing homes, they often experience loneliness due to separation from loved ones, and traditional nursing beds struggle to provide the same interactive experience. Furthermore, existing nursing beds offer limited responses to patient commands, lacking the ability to interact with vivid and natural digital avatars based on individual patient needs, and cannot intelligently control the bed itself. This reduces patient satisfaction and comfort during use, failing to fully meet the current demand for intelligent and humanized services in the medical and nursing field.

[0004] Therefore, this invention proposes an AI-based nursing bed and its digital human interaction method and system. Summary of the Invention

[0005] This invention provides an AI-based nursing bed and its digital human interaction method and system, with the nursing bed itself providing the basic support. An interaction trigger module receives user-input commands to summon a digital family member, providing an entry point for activating special interactive functions. The digital human interaction terminal, located at the bedside, includes a display screen, an audio acquisition module, an audio output module, and a nursing bed control module. Based on commands, it can retrieve the corresponding virtual image model of a digital family member from a digital human model library, enabling interaction with the user and control of the nursing bed. This not only provides emotional companionship and alleviates loneliness but also allows users to conveniently control the nursing bed by interacting with a familiar digital family member, improving the ease of use and humanization of the nursing bed, and enhancing the user's experience and comfort during care.

[0006] This invention provides an AI-based nursing bed, comprising: The nursing bed itself; The interactive trigger module is used to receive digital family member summoning commands input by the user; The digital human interaction terminal is located at the head of the nursing bed and is used to retrieve the corresponding digital family virtual image model stored in the digital human model library based on the digital family summoning command. It can then interact with the user or control the nursing bed based on the corresponding digital family virtual image model. The terminal includes a display screen, an audio acquisition module, an audio output module, and a nursing bed control module.

[0007] This invention provides an AI-based digital human interaction method for a nursing bed, comprising: A digital virtual avatar model of a relative is generated based on pre-inputted description data of the relative, and the digital virtual avatar model of the relative is stored in a digital human model library; Parse the target digital relative role identifier in the user's input digital relative summoning command, and retrieve the target digital relative virtual image model corresponding to the target digital relative role identifier from the digital human model library; Analyze users' emotional state and summoning intentions based on real-time collected user monitoring audio data; Based on the behavioral and behavioral feature library of the target digital relative virtual avatar model, the user's emotional state and summoning intention, and the historical interaction records between the user and the corresponding target digital relative virtual avatar model, the target digital relative virtual avatar is dynamically rendered to obtain the current digital human response animation. The current digital human response animation is displayed on the screen. At the same time, when the summoning intent includes a control command for the nursing bed itself, the corresponding nursing bed itself is controlled based on the corresponding control command.

[0008] Preferably, a digital virtual avatar model of a relative is generated based on pre-inputted description data of the relative, and the digital virtual avatar model of the relative is stored in a digital human model library, including: Extract the family's voice features, facial features, language habits, and behavioral habits from the pre-input description of the family member. Digital virtual image models of relatives are generated based on the voice features, facial morphology features, language habit features, and behavioral habit features of relatives. Store the digital avatars of your loved ones in the digital human model library.

[0009] Preferably, the process involves parsing the target digital relative role identifier in the user-input digital relative summoning command and retrieving the corresponding virtual image model of the target digital relative from the digital human model library, including: Match the user's input digital family member summoning command with a list of corresponding command-style character identifiers to identify the target digital family member character identifier in the user's input digital family member summoning command; Retrieve the virtual image model of the target digital relative corresponding to the target digital relative role identifier from the digital human model library.

[0010] Preferably, the analysis of users' emotional state and summoning intentions based on real-time collected user monitoring audio data includes: Analyze users' emotional state based on real-time collected user monitoring audio data; User behavior data is analyzed based on real-time collected user monitoring audio data. The summoning intent is determined based on the user's emotional state, verbal and behavioral data, and summoning commands.

[0011] Preferably, based on the behavioral and behavioral feature library of the target digital relative virtual avatar model, the user's emotional state and summoning intention, and the historical interaction records between the user and the corresponding target digital relative virtual avatar model, the target digital relative virtual avatar is dynamically rendered to obtain the current digital human's response animation, including: The summoning response result is determined based on the historical interaction records between the user and the corresponding target digital relative virtual avatar model and the summoning intent. Multiple response mapping matrices were analyzed based on the historical interaction records between users and the corresponding target digital family virtual avatar models. Based on the behavioral and speech habit feature library of the target digital family virtual avatar model and the user's emotional state analysis, a standard response format for the summoning response result is determined. Generate the optimal response form for the summon response results based on all response mapping matrices and the standard response form of the summon response results; Based on the summoning response results and the corresponding optimal response form, the target digital family virtual image is dynamically rendered to obtain the current digital human's response animation.

[0012] Preferably, multiple response mapping matrices are analyzed based on historical interaction records between the user and the corresponding target digital relative virtual avatar model, including: Extract the user's summoning intent, prior emotional state, corresponding response form, and user satisfaction in each historical interaction record between the user and the corresponding target digital family virtual image model. All historical interaction processes with the same summoning intention and corresponding response form are summarized as the first historical interaction process set, and all historical interaction processes with the same prior emotional state and corresponding response form are summarized as the second historical interaction process set. Based on the summoning intent, prior emotional state, corresponding response form and user satisfaction of all historical interaction processes in each first historical interaction process set, the first mapping coefficient between each prior emotional state and corresponding response form under the corresponding summoning intent is analyzed. Based on the summoning intent, prior emotional state, corresponding response form and user satisfaction of all historical interaction processes in each second historical interaction process set, the second mapping coefficient between each summoning intent and corresponding response form under the corresponding prior emotional state is analyzed. The first and second mapping coefficients between each pre-existing emotional state and each response form under each summoning intention are used to determine the final mapping coefficients between each pre-existing emotional state and each response form under each summoning intention; The final mapping coefficients between each pre-existing emotional state and the same response form under each summoning intention are represented by a matrix to obtain the response mapping matrix for each response form.

[0013] Preferably, the standard response format for determining the summoning response result based on the behavioral habit feature library of the target digital relative virtual avatar model and the user's emotional state analysis includes: In the behavioral and speech habit feature database of the target digital family virtual avatar model, multiple behavioral and speech response features of the target digital family virtual avatar under the user's emotional state were retrieved; Based on the summoning intent, the standard response format of the summoning response result is determined by retrieving multiple verbal and behavioral response features of the target digital relative's virtual image in the user's emotional state.

[0014] Preferably, the optimal response form for generating the summon response result based on the standard response form of all response mapping matrices and summon response results includes: The maximum value of the user's emotional state and summoning intention in all response mapping matrices of all response forms is taken as the first target matrix element. The response form corresponding to the response mapping matrix to which the first target matrix element belongs is taken as the candidate response form. The matching degree between the candidate response form and the standard response form is calculated as the first matching degree. The corresponding matrix elements of the user's emotional state and summoning intent in the response mapping matrix in the standard response form are used as the second target matrix elements, and the matching degree of the first target matrix elements and the second target matrix elements is determined as the second matching degree. The co-selection reliability of the candidate response form and the standard response form is calculated based on the first and second matching degrees between the candidate response form and the standard response form. When the overall reliability is not less than the reliability threshold, the optimal response form of the summon response result is obtained by summing the form performance values ​​of the largest reliability among the form performance values ​​of the same dimension in the candidate response form and the standard response form. When the overall reliability is less than the reliability threshold, the response form with the highest overall reliability among the candidate response form and the standard response form is taken as the best response form for the call response result.

[0015] This invention provides an AI-based digital human interaction system for a nursing bed, comprising: The digital human pre-built module is used to generate a digital virtual image model of a relative based on pre-inputted description data of the relative, and to store the digital virtual image model of the relative in the digital human model library; The virtual model summoning module is used to parse the target digital relative role identifier in the digital relative summoning command input by the user, and retrieve the virtual image model of the target digital relative corresponding to the target digital relative role identifier from the digital human model library; The emotion and intent analysis module is used to analyze the user's emotional state and calling intent based on real-time collected user monitoring audio data. The virtual dynamic rendering module is used to dynamically render the target digital relative's virtual image based on the target digital relative's speech and behavior feature library, the user's emotional state and summoning intention, and the historical interaction records between the user and the corresponding target digital relative's virtual image model, to obtain the current digital human's response animation. The interactive output control module is used to display the current digital human response animation on the screen. At the same time, when the summoning intent includes a control command for the nursing bed itself, the corresponding nursing bed itself is controlled based on the corresponding control command.

[0016] The beneficial effects of this invention compared to existing technologies are as follows: Integrating AI technology into nursing beds not only enhances the intelligence level of the beds but also creates a more personalized, convenient, and humanized nursing experience for patients. Digital human interaction technology, as an important branch of AI, provides a new direction for expanding the functionality of nursing beds by creating realistic virtual avatars that interact naturally with users. It can build a closer and more efficient communication bridge between patients and caregivers. Patients can interact with digital avatars to obtain health guidance, entertainment, and other services, alleviating psychological stress caused by illness or hospitalization and enhancing their confidence in recovery. For caregivers, this system can assist in the communication and guidance of some basic nursing tasks, improving the efficiency of nursing work. Furthermore, this technology is expected to deeply integrate with fields such as telemedicine and smart homes in the future, building a more complete smart healthcare ecosystem with broad application prospects and development potential.

[0017] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in this application.

[0018] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0019] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a schematic diagram of the internal components and modules of the AI-based nursing bed in an embodiment of the present invention; Figure 2 This is a schematic diagram of a digital human interaction method for an AI-based nursing bed in an embodiment of the present invention; Figure 3 This is a schematic diagram of the AI-based digital human interaction system for a nursing bed in an embodiment of the present invention. Detailed Implementation

[0020] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0021] like Figure 1 As shown, the present invention provides an implementation of an AI-based nursing bed, comprising: The nursing bed itself; The interactive trigger module is used to receive digital family member summoning commands input by the user; The digital human interaction terminal is located at the head of the nursing bed and is used to retrieve the corresponding digital family virtual image model stored in the digital human model library based on the digital family summoning command. It can then interact with the user or control the nursing bed based on the corresponding digital family virtual image model. The terminal includes a display screen, an audio acquisition module, an audio output module, and a nursing bed control module.

[0022] In this embodiment, the nursing bed body is the basic physical support part of the entire nursing bed equipment, providing patients with basic lying space and serving as the basic platform for patient care and interaction with digital humans.

[0023] In this embodiment, the digital family member summoning command is an instruction sent by the user to the nursing bed system to evoke a specific digital family member's virtual image. For example, if a user wants to summon their deceased grandfather for companionship and conversation, they would input a digital family member summoning command containing the character identifier "grandfather" through the interaction trigger module.

[0024] In this embodiment, the digital human model library is a database used to store virtual avatar models of digital relatives.

[0025] In this embodiment, the digital family virtual avatar model is a model generated using these features that can mimic the speech and behavior of a family member. For example, the generated virtual avatar model of a father can communicate with the user in the father's tone and language habits, and try to restore the image of a father as much as possible in terms of appearance and behavior.

[0026] In this embodiment, when the system retrieves the corresponding virtual avatar model of a digital relative from the digital human model library based on the digital relative summoning command, it interacts with the user based on this model. On one hand, the system displays a digital human response animation on the screen, and the audio output module plays voices that match the characteristics of the virtual avatar, chatting with the user to relieve boredom, sharing life experiences, understanding the user's emotions, and mimicking the tone of a relative to communicate, thus providing emotional companionship. On the other hand, if the user's summoning intention includes control commands for the nursing bed itself, such as adjusting the angle or height of the nursing bed, the system will operate the nursing bed itself according to the corresponding commands through the nursing bed control module, improving the convenience and user-friendliness of the nursing bed.

[0027] In this embodiment, a display screen is positioned at the head of the nursing bed to show the current digital human response animation dynamically rendered based on the corresponding digital family virtual avatar model. The audio acquisition module is responsible for collecting the user's voice in real time, providing audio data for the system to analyze the user's emotional state, speech and behavior data, and determine the summoning intent. The audio output module plays the voice information generated by the system based on the digital family virtual avatar model, realizing voice interaction with the user. The nursing bed control module controls various functions of the nursing bed based on the nursing bed control commands contained in the user's summoning intent, such as adjusting the angle and height of the nursing bed. These modules work together to realize the interaction between the nursing bed and the user based on the digital family virtual avatar model and the control of the nursing bed.

[0028] like Figure 2 As shown, this invention provides an implementation method for a digital human interaction method for an AI-based nursing bed, comprising: A digital virtual avatar model of a relative is generated based on pre-inputted description data of the relative, and the digital virtual avatar model of the relative is stored in a digital human model library; Parse the target digital relative role identifier in the user's input digital relative summoning command, and retrieve the target digital relative virtual image model corresponding to the target digital relative role identifier from the digital human model library; Analyze users' emotional state and summoning intentions based on real-time collected user monitoring audio data; Based on the behavioral and behavioral feature library of the target digital relative virtual avatar model, the user's emotional state and summoning intention, and the historical interaction records between the user and the corresponding target digital relative virtual avatar model, the target digital relative virtual avatar is dynamically rendered to obtain the current digital human response animation. The current digital human response animation is displayed on the screen. At the same time, when the summoning intent includes a control command for the nursing bed itself, the corresponding nursing bed itself is controlled based on the corresponding control command.

[0029] In this embodiment, the pre-input description data of the relatives serves as the foundational information for creating a digital virtual avatar model of the relatives. It covers various characteristics of the relatives, including their voice features, such as timbre and tone of voice; their facial features, such as the shape and outline of their facial features; their language habits, such as commonly used words and catchphrases; and their behavioral habits, such as habitual actions and postures.

[0030] In this embodiment, a digital virtual avatar model of a relative is generated based on pre-inputted description data of the relative. According to the pre-input description data, the system utilizes language synthesis techniques and related algorithms, such as acoustic model training algorithms and language model algorithms; 3D modeling techniques and related algorithms, such as geometric modeling algorithms and texture mapping algorithms; and animation generation techniques and related algorithms, such as keyframe animation algorithms and physics-based animation algorithms, to generate the digital virtual avatar model. First, the system extracts the relative's voice features, facial morphology features, language habit features, and behavioral habit features from the description data. Then, it integrates these features using professional modeling and rendering techniques to construct a digital virtual avatar model.

[0031] In this embodiment, the target digital relative role identifier is an identifier used in the user's input digital relative summoning command to explicitly specify the specific digital relative to be summoned. It can be a term of address for the relative, such as "Dad" or "Grandma," or other specific identifiers that can uniquely identify a particular digital relative.

[0032] In this embodiment, the user monitors audio data: In this embodiment, the user monitors audio data by collecting the sound data emitted by the user in real time through the audio acquisition module.

[0033] In this embodiment, the user's emotional state and summoning intention refer to the user's current emotional state, such as happiness, sadness, anger, or boredom, determined by analyzing the user's monitored audio data. Summoning intention refers to the purpose the user wants to achieve by inputting a digital relative summoning command. This may include chatting with the digital relative, obtaining information, seeking companionship, or controlling the nursing bed itself, such as adjusting its height or angle.

[0034] In this embodiment, the speech and behavior feature library of the target digital relative virtual avatar model is a database that stores the speech and behavior features of the corresponding digital relative virtual avatar in various situations. These features are generated based on pre-input relative description data, including language expression patterns, commonly used vocabulary, tone characteristics, and corresponding behavioral habits under different emotional states.

[0035] In this embodiment, the historical interaction record between the user and the corresponding target digital relative virtual image model records relevant information about each interaction between the user and the specific digital relative virtual image model, including the summoning intention during each interaction, the user's pre-interaction emotional state, the corresponding response form given by the system, and the user's satisfaction with the response.

[0036] In this embodiment, the current digital human response animation is generated by the system dynamically rendering the target digital relative virtual avatar based on its behavioral and speech habit feature library, the user's emotional state and summoning intent, and historical interaction records between the user and the corresponding target digital relative virtual avatar model. This animation is displayed to the user on a screen, accompanied by voice playback from the audio output module, vividly presenting the interaction process between the digital relative and the user. The actions and expressions of the digital relative virtual avatar in the animation are matched to the current interaction context. For example, when the user is sad, the digital relative virtual avatar may make comforting expressions and actions, making the interaction more realistic and natural, and enhancing the user's experience.

[0037] In this embodiment, when the summoning intent includes a control command for the nursing bed body, the corresponding nursing bed body is controlled based on the corresponding control command: if the system determines through analysis that the user's summoning intent includes a control command for the nursing bed body, such as adjusting the height, angle, or backrest tilt of the nursing bed, the system will transmit these control commands to the nursing bed body control module.

[0038] To extract various features from pre-input descriptions of relatives to generate digital virtual avatar models and store them in a digital human model library for subsequent personalized interaction, this paper proposes generating digital virtual avatar models based on pre-input relative description data and storing them in a digital human model library, including: Extract the family's voice features, facial features, language habits, and behavioral habits from the pre-input description of the family member. Digital virtual image models of relatives are generated based on the voice features, facial morphology features, language habit features, and behavioral habit features of relatives. Store the digital avatars of your loved ones in the digital human model library.

[0039] In this embodiment, the voice characteristics of relatives refer to the unique vocal features presented when relatives speak. This includes the timbre of a relative's voice, such as some people's voices being deep and others being clear; intonation, such as the way they speak, whether it is calm or passionate; speech rate, that is, how fast they speak; and pronunciation habits, such as certain dialect pronunciations, the way certain words are pronounced, etc.

[0040] In this embodiment, the facial morphological characteristics of a loved one are a description of the facial appearance. This includes the shape of the facial features, such as whether the eyes are large or small, double or single eyelids; the height and width of the nose; the size and shape of the mouth and lips. It also includes the facial contour, such as whether it is a round, oval, or square face, as well as some special facial markings, such as moles or dimples.

[0041] In this embodiment, the language habits of relatives reflect their habitual ways of expressing themselves. This includes commonly used vocabulary, such as catchphrases and words frequently used in specific situations; sentence structure, whether they prefer short or long sentences and the characteristics of their sentence organization; and language style, whether they are humorous, serious, or gentle and friendly.

[0042] In this embodiment, the behavioral habits of relatives refer to the habitual patterns exhibited by relatives in their daily actions. This may include physical movements, such as habitually crossing their arms over their chest or having a unique walking posture; facial expressions, such as frequently smiling or frowning; and behaviors in specific scenarios, such as liking to eat snacks while watching TV or resting their chin on their hand while thinking.

[0043] In this embodiment, a digital virtual image model of a relative is generated based on the relative's voice features, facial morphology features, language habit features, and behavioral habit features: First, based on the facial features of the loved one, a virtual facial model that resembles the appearance of the loved one is constructed using 3D modeling technology, with detailed depiction of facial features, facial contours, and special markings.

[0044] Then, based on the voice characteristics of relatives, speech synthesis technology is used to simulate a voice with similar timbre, tone, speech rate and pronunciation habits, so that the virtual avatar can communicate with the user in a voice like that of a relative.

[0045] Furthermore, by combining the language habits and characteristics of relatives, we set the language expression mode of the virtual character in different situations, including commonly used vocabulary, sentence structure and language style, to ensure that the words spoken by the virtual character conform to the language habits of relatives.

[0046] Then, based on the behavioral habits and characteristics of family members, corresponding actions and expressions are added to the virtual avatar, allowing it to perform body movements, facial expressions, and specific scene behaviors just like a real family member during interaction. By organically combining these features, a digital virtual avatar model of a family member is generated that highly replicates the characteristics of a family member in terms of appearance, voice, and behavior, providing users with a more realistic and considerate companionship interaction experience.

[0047] To accurately retrieve digital family member images by matching user input commands with a list of character identifiers and retrieving the corresponding virtual avatar model from a digital human model library, this paper proposes parsing the target digital family member character identifier in the user's input command and retrieving the corresponding virtual avatar model from the digital human model library. The process includes: Match the user's input digital family member summoning command with a list of corresponding command-style character identifiers to identify the target digital family member character identifier in the user's input digital family member summoning command; Retrieve the virtual image model of the target digital relative corresponding to the target digital relative role identifier from the digital human model library.

[0048] In this embodiment, the instruction format refers to the specific format or standard that a user should follow when inputting a digital command to summon a loved one. Examples include voice input, keypad input, and gesture input.

[0049] In this embodiment, the list of character identifiers in the form of instructions is a list containing various possible target digital family member character identifiers under a single instruction format. Each character identifier in the list corresponds to a specific digital family member virtual avatar model or family member character.

[0050] In this embodiment, the user-input digital family member summoning command is matched with a list of corresponding command-format role identifiers to identify the target digital family member role identifier in the user-input digital family member summoning command: the user-input command content is compared with each item in the list of corresponding command-format role identifiers to find the matching role identifier.

[0051] To determine call-to-action intent by analyzing real-time user monitoring audio data, including emotional state and behavioral data, and to enable the system to more accurately understand user needs, this paper proposes an analysis of user emotional state and call-to-action intent based on real-time user monitoring audio data, including: Analyze users' emotional state based on real-time collected user monitoring audio data; User behavior data is analyzed based on real-time collected user monitoring audio data. The summoning intent is determined based on the user's emotional state, verbal and behavioral data, and summoning commands.

[0052] In this embodiment, the system analyzes the user's emotional state based on real-time collected user monitoring audio data: The system processes and interprets the real-time collected user monitoring audio data to determine the user's current emotional state. Specifically, it considers multiple dimensions of the audio, such as tone of voice. A high-pitched and fast-paced tone may indicate excitement or agitation, while a low-pitched and slow-paced tone may suggest depression or frustration. Volume is also a factor; loud talking may reflect anger or excitement, while soft talking may suggest calmness or shyness. Furthermore, details such as interjections and sighs in the audio should not be overlooked; for example, frequent sighs may indicate helplessness or sorrow. By comprehensively analyzing these audio features, the system can accurately identify different emotional states such as happiness, sadness, anger, and anxiety. For example, if a user speaks loudly and quickly, accompanied by cheerful interjections, the system can analyze that the user is in a happy emotional state.

[0053] In this embodiment, user speech and behavior data is analyzed based on real-time collected user monitoring audio data: the speech content in the audio is converted into text form to obtain the specific words spoken by the user. Then, natural language processing technology performs in-depth analysis on these texts, such as word segmentation, breaking sentences down into meaningful lexical units; part-of-speech tagging, determining the part of speech of each word; and named entity recognition, identifying entity information such as names of people, places, and organizations in the text. Through these processes, the system can clearly understand the content expressed by the user, that is, the user's speech and behavior data.

[0054] In this embodiment, the summoning intent is determined based on the user's emotional state, verbal and behavioral data, and summoning command. For example, if the user's emotional state is happy, their verbal and behavioral data is "I had a good experience today and I want to share it," and their summoning command is to summon a digital friend, then the system can determine that the user's summoning intent is to share the good thing that happened to them with their digital friend. Similarly, if the user is anxious, their verbal and behavioral data mentions "feeling unwell," and they simultaneously issue a summoning command to call for caregivers, the system can determine that the user's summoning intent is to seek help from caregivers regarding their physical discomfort. By integrating this information, the system can more accurately grasp the user's needs.

[0055] To obtain the current digital human's response animation through multi-step dynamic rendering by combining the behavioral and behavioral feature database of the target digital relative virtual avatar, the user's emotional state and summoning intent, and the historical interaction records between the user and the corresponding target digital relative virtual avatar model, a method is proposed to dynamically render the target digital relative virtual avatar to obtain the current digital human's response animation, including: The summoning response result is determined based on the historical interaction records between the user and the corresponding target digital relative virtual avatar model and the summoning intent. Multiple response mapping matrices were analyzed based on the historical interaction records between users and the corresponding target digital family virtual avatar models. Based on the behavioral and speech habit feature library of the target digital family virtual avatar model and the user's emotional state analysis, a standard response format for the summoning response result is determined. Generate the optimal response form for the summon response results based on all response mapping matrices and the standard response form of the summon response results; Based on the summoning response results and the corresponding optimal response form, the target digital family virtual image is dynamically rendered to obtain the current digital human's response animation.

[0056] In this embodiment, the summoning response is determined based on the user's historical interaction records and summoning intent with the corresponding target digital family member virtual avatar model. The system analyzes these records and combines them with the current summoning intent to infer a suitable summoning response. For example, if a user frequently summons their digital mother when feeling lonely, and historical interaction records show that the digital mother provides comfort and shares interesting stories after the user expresses loneliness, resulting in high user satisfaction, then when the user summons their digital mother again due to loneliness, based on historical interaction records and the current summoning intent of "seeking companionship and comfort," the system's determined summoning response might be the digital mother offering comforting words and sharing an interesting story. The system uses historical interaction records as a reference to ensure that the summoning response better aligns with the user's past preferences and needs.

[0057] In this embodiment, the summoning response result is the response content determined by the system based on the historical interaction records between the user and the corresponding target digital relative virtual avatar model and the current summoning intention. It can be a voice reply from the digital relative virtual avatar, a specific facial expression, or a combination of interactive behaviors.

[0058] To extract various information from the historical interaction records between users and target digital family virtual avatar models, a multi-step analysis is conducted to determine the final mapping coefficients between the pre-summoning emotional state and the response form under each summoning intention, and this is then matrix-represented to obtain the response mapping matrix for generating the optimal response form. This paper proposes to analyze multiple response mapping matrices based on the historical interaction records between users and corresponding target digital family virtual avatar models, including: Extract the user's summoning intent, prior emotional state, corresponding response form, and user satisfaction in each historical interaction record between the user and the corresponding target digital family virtual image model. All historical interaction processes with the same summoning intention and corresponding response form are summarized as the first historical interaction process set, and all historical interaction processes with the same prior emotional state and corresponding response form are summarized as the second historical interaction process set. Based on the summoning intent, prior emotional state, corresponding response form and user satisfaction of all historical interaction processes in each first historical interaction process set, the first mapping coefficient between each prior emotional state and corresponding response form under the corresponding summoning intent is analyzed. Based on the summoning intent, prior emotional state, corresponding response form and user satisfaction of all historical interaction processes in each second historical interaction process set, the second mapping coefficient between each summoning intent and corresponding response form under the corresponding prior emotional state is analyzed. The first and second mapping coefficients between each pre-existing emotional state and each response form under each summoning intention are used to determine the final mapping coefficients between each pre-existing emotional state and each response form under each summoning intention; The final mapping coefficients between each pre-existing emotional state and the same response form under each summoning intention are represented by a matrix to obtain the response mapping matrix for each response form.

[0059] In this embodiment, the user's summoning intent, prior emotional state, corresponding response form, and user satisfaction in each historical interaction process are as follows: The summoning intent is the purpose that the user wants to achieve by initiating this interaction, such as wanting to chat to relieve boredom, obtain information, or control the nursing bed.

[0060] Pre-interaction emotional state refers to the user's emotional state before initiating interaction, such as happiness, sadness, or anxiety.

[0061] The corresponding response format is the way the system responds to the user's call based on the target digital family member virtual avatar model, such as the content of the voice reply, the display of animated actions, etc.

[0062] User satisfaction reflects the degree of satisfaction a user has with the response to an interaction. It can usually be expressed numerically or on a scale of 1 to 5, with 5 being very satisfied.

[0063] In this embodiment, based on the summoning intent, prior emotional state, corresponding response form, and user satisfaction of all historical interaction processes in each first historical interaction process set, the first mapping coefficient between each prior emotional state and corresponding response form under the corresponding summoning intent is analyzed: Step 1: The system first comprehensively collects all historical interaction data between the user and the corresponding target digital family member virtual avatar model. This data records in detail the summoning intention, prior emotional state, corresponding response form, and user satisfaction for each interaction.

[0064] For example, it records many similar interaction records such as "User A, in a state of boredom, summoned a story with the intention of hearing it, and the digital relative responded by telling a folk tale, and the user gave a satisfaction score of 4 out of 5" and "User B, in a state of curiosity, summoned a story with the intention of hearing it, and the digital relative responded by telling a folk tale, and the user gave a satisfaction score of 3".

[0065] Step 2: Based on the summoning intent and the corresponding response, classify all historical interaction processes. Group historical interaction processes with the same summoning intent and the same response into one set, namely the first historical interaction process set.

[0066] For example, select all interactions where the intention to be summoned is "want to hear a story" and the response is "tell a folk tale," and form a first-historical interaction set. This set contains the interaction records of different users in various pre-emotional states regarding the combination of "want to hear a story - tell a folk tale."

[0067] Step 3: In each first historical interaction process set, for the specific call intention of "wanting to hear a story", begin to analyze the degree of correlation between different pre-emotional states (such as boredom, curiosity, excitement, etc.) and the response form of "tell a folk tale".

[0068] Specifically, this includes: first, counting the number of interactions under each pre-emotional state. For example, in the interaction set of "Want to hear a story - tell a folk tale", there were 20 interactions under the emotion of boredom, 15 interactions under the emotion of curiosity, and so on.

[0069] Next, the total user satisfaction score for each pre-emotional state is calculated. For example, the total satisfaction score for 20 interactions under the boredom state is 80 points (the satisfaction scores for each interaction are added together), and the total satisfaction score for 15 interactions under the curiosity state is 45 points.

[0070] Step 4: Divide the total user satisfaction for each pre-emotional state by the number of interactions in that pre-emotional state to obtain the average satisfaction for each pre-emotional state.

[0071] For boredom, the average satisfaction score is 80 points ÷ 20 times = 4 points; for curiosity, the average satisfaction score is 45 points ÷ 15 times = 3 points.

[0072] Step 5: The average satisfaction value can be used as a reflection of the degree of correlation between the corresponding pre-emotional state and the response form of "telling a folk tale", which is the first mapping coefficient.

[0073] The calculations above show that the first mapping coefficient for boredom is 4 points, and the first mapping coefficient for curiosity is 3 points.

[0074] In this embodiment, based on the summoning intent, prior emotional state, corresponding response form, and user satisfaction of all historical interaction processes in each second historical interaction process set, the second mapping coefficient between each summoning intent and corresponding response form under the corresponding prior emotional state is analyzed using the same logical method as the first mapping coefficient determined in the previous step.

[0075] In this embodiment, the first and second mapping coefficients between each pre-existing emotional state and each response form under each summoning intention are used to determine the final mapping coefficient between each summoning intention and each pre-existing emotional state and each response form. For example, a weighted average is used, assigning different weights to the first and second mapping coefficients based on their importance in determining the final relationship, and then calculating the final mapping coefficient. For instance, for the combination of the summoning intention of "wanting to hear a story," the pre-existing emotional state of "boredom," and the response form of "telling a folk tale," if the weight of the first mapping coefficient is set to 0.6 and the weight of the second mapping coefficient is set to 0.4, and the first mapping coefficient is known to be 0.8 and the second mapping coefficient to be 0.7, then the final mapping coefficient = 0.6 × 0.8 + 0.4 × 0.7 = 0.76.

[0076] In this embodiment, the final mapping coefficients between each pre-existing emotional state and the same response form under each summoning intent are represented in a matrix form to obtain the response mapping matrix for each response form: For each response type, the final mapping coefficients corresponding to different calling intentions and pre-existing emotional states are organized into a matrix, resulting in a response mapping matrix for each response type. The rows and columns of the matrix can correspond to different calling intentions and pre-existing emotional states, respectively.

[0077] To determine a standard response format for a summoning response that better reflects the characteristics of the virtual avatar, based on the behavioral and behavioral feature database of the target digital relative's virtual avatar model and the user's emotional state, combined with the summoning intent, this paper proposes a standard response format for summoning responses based on the behavioral and behavioral feature database of the target digital relative's virtual avatar model and the user's emotional state analysis. This includes: In the behavioral and speech habit feature database of the target digital family virtual avatar model, multiple behavioral and speech response features of the target digital family virtual avatar under the user's emotional state were retrieved; Based on the summoning intent, the standard response format of the summoning response result is determined by retrieving multiple verbal and behavioral response features of the target digital relative's virtual image in the user's emotional state.

[0078] In this embodiment, multiple behavioral and verbal response features of the target digital family virtual avatar model under the user's emotional state are retrieved from the behavioral and verbal habit feature library of the target digital family virtual avatar model. When the system obtains the user's emotional state, it searches the behavioral and verbal habit feature library of the target digital family virtual avatar model. For example, if the user is in a happy emotional state, the system will find various behavioral and verbal behaviors that the target digital family virtual avatar might exhibit when facing a happy user. These behaviors may include specific verbal expressions, such as words of congratulations or praise; and corresponding actions or expressions, such as smiling or clapping. These are the multiple behavioral and verbal response features of the target digital family virtual avatar under the user's happy emotional state.

[0079] In this embodiment, multiple verbal and behavioral response features refer to a series of verbal and behavioral response characteristics that the target digital family virtual image model may produce in a specific emotional state.

[0080] In this embodiment, the standard response form of the summoning response result is determined by retrieving multiple verbal and behavioral response features of the target digital relative virtual avatar under the user's emotional state based on the summoning intent: After obtaining multiple verbal and behavioral response features of the target digital relative virtual avatar under the user's current emotional state, the system further filters and determines the response based on the user's summoning intent. For example, if the user is in a sad mood and the summoning intent is to seek comfort, the system will select content that matches the intention to seek comfort from the multiple verbal and behavioral response features under the sad mood. This might involve selecting words and tones that emphasize comfort and encouragement, as well as corresponding comforting actions, such as a gentle pat on the shoulder animation, etc., combining these to form a complete response pattern, which is the standard response form of the summoning response result.

[0081] To improve the reliability and rationality of the response by determining the target matrix elements in the response mapping matrix, calculating the matching degree and co-selecting reliability, and determining the optimal response form of the call response result based on the reliability, this paper proposes to generate the optimal response form of the call response result based on the standard response form of all response mapping matrices and call response results, including: The maximum value of the user's emotional state and summoning intention in all response mapping matrices of all response forms is taken as the first target matrix element. The response form corresponding to the response mapping matrix to which the first target matrix element belongs is taken as the candidate response form. The matching degree between the candidate response form and the standard response form is calculated as the first matching degree. The corresponding matrix elements of the user's emotional state and summoning intent in the response mapping matrix in the standard response form are used as the second target matrix elements, and the matching degree of the first target matrix elements and the second target matrix elements is determined as the second matching degree. The co-selection reliability of the candidate response form and the standard response form is calculated based on the first and second matching degrees between the candidate response form and the standard response form. When the overall reliability is not less than the reliability threshold, the optimal response form of the summon response result is obtained by summing the form performance values ​​of the largest reliability among the form performance values ​​of the same dimension in the candidate response form and the standard response form. When the overall reliability is less than the reliability threshold, the response form with the highest overall reliability among the candidate response form and the standard response form is taken as the best response form for the call response result.

[0082] In this embodiment, the matching degree between the candidate response form and the standard response form is calculated. This matching is considered from multiple dimensions, such as language expression, comparing whether the vocabulary, sentence structure, and semantics used in the two are similar; and in terms of actions and expressions, observing whether the type and amplitude of actions and changes in facial expressions match. By averaging or weighting these dimensions, a numerical value is obtained to represent the matching degree.

[0083] In this embodiment, the matching degree of the first target matrix element and the second target matrix element is determined: for example, the difference between the two values ​​is calculated, and then the matching degree is determined by comparing the difference with a certain benchmark value, where the benchmark value is, for example, equal to the first target matrix element or the second target matrix element.

[0084] In this embodiment, the co-selection reliability of the candidate response form and the standard response form is calculated based on the first matching degree and the second matching degree between the candidate response form and the standard response form: Co-selection reliability is used to comprehensively evaluate the reliability of both candidate and standard response forms being selected as the final response. A weighted average approach may be used, assigning different weights to the first and second matching degrees, and then calculating the weighted sum as the co-selection reliability. For example, suppose the weight of the first matching degree is 0.6 and the weight of the second matching degree is 0.4.

[0085] In this embodiment, the reliability threshold is a pre-set standard value used to determine whether the candidate response form and the standard response form are reliable enough to determine whether they can be combined to generate the optimal response form. For example, a value of 0.75 is used.

[0086] In this embodiment, the optimal response form is obtained by summing the form performance values ​​with the highest reliability among the form performance values ​​of the same dimension in the candidate response form and the standard response form: These two response formats may have their own performance values ​​on different dimensions, such as the verbal expression dimension and the action / facial expression dimension. For each dimension, the system compares the reliability of the candidate response format and the standard response format's performance values ​​on that dimension, and selects the performance value with the highest reliability. For example, on the verbal expression dimension, if a certain expression in the candidate response format has a reliability of 85% and the corresponding expression in the standard response format has a reliability of 80%, then the expression in the candidate response format is selected for that dimension; on the action / facial expression dimension, if a certain action in the candidate response format has a reliability of 78% and the corresponding action in the standard response format has a reliability of 82%, then the action in the standard response format is selected for that dimension. The sum of the performance values ​​with the highest reliability selected on each dimension constitutes the optimal response format for the call-to-response result.

[0087] In this embodiment, the method for determining the overall reliability of the candidate response form and the standard response form, and the reliability of the formal performance values ​​of each dimension in the candidate response form and the standard response form, includes: Step 1: Preprocess historical data to complete the dimensional division and range labeling of multidimensional vectors: The dimensions of a user's pre-emotional state include: emotion type, emotion intensity, and emotion duration; and the range of values ​​for each dimension, for example, dividing emotion intensity into three ranges: low (0-3 points), medium (4-7 points), and high (8-10 points). The dimensions of summoning intent include: intent type, urgency, and clarity; and the range of single-dimensional values, for example, dividing intent clarity into three ranges: vague (0-4 points), medium (5-7 points), and clear (8-10 points). The response format includes: speech rate, facial expression intensity, and movement frequency, as well as range indicators for single-dimensional values. For example, speech rate can be divided into three ranges: slow (0-3), medium (4-7), and fast (8-10). The remaining dimensions (such as emotion type, action frequency, etc.) are each divided into three ranges according to the same logic, and each range is identified by a unique ID, such as low=1, medium=2, high=3.

[0088] Scene partition combination identifier: Scene partitions are identified by a triplet consisting of “emotion intensity range ID + intent clarity range ID + response dimension range ID”, for example, “low emotion intensity + clear intent + medium speech rate” corresponds to the combination (1,3,2).

[0089] Step 2: Calculate mapping coefficients by combining different scenarios to construct a knowledge graph. Samples were extracted from historical interaction records, including the pre-interaction emotional state, call-to-action intention, response form, and user satisfaction (0-10 points) for each interaction.

[0090] The samples were categorized by scenario-based partitioning, and the frequency of occurrence of different response dimension ranges and corresponding total satisfaction were counted for each combination.

[0091] The matching relationship between each scene partition combination and the response dimension range is calculated using the following formula: Mapping coefficient = (total satisfaction of a certain response dimension range under this combination ÷ total occurrence of this range) ÷ 10. The result is normalized to 0-1. The higher the value, the more suitable the response dimension is in this scenario.

[0092] Knowledge graph construction: Nodes include scene partition combination nodes, such as triplet identifiers and response dimension range nodes, such as medium speech speed; the edge weights between nodes are set as mapping coefficients between the corresponding scene partitions and response dimension ranges.

[0093] Step 3: Locate the starting point of the graph based on the current interaction scenario, crawl the path, and accumulate weights: Starting with the scene partitioning of the current interaction determined based on real-time emotional state and summoning intent, crawl all reachable response dimension range nodes in the graph; Single-dimensional path weight = the product of the mapping coefficients of all edges on the path, which serves as the basis for the adaptability of that response dimension.

[0094] Step 4: Calculate the single-dimensional reliability and overall reliability of the candidate response form and the standard response form respectively: For each dimension of both candidate and standard response formats, the cumulative weight of its corresponding path in the knowledge graph is used as the reliability of that dimension. Taking speech rate as an example: First, determine the scene partition combination of the current interaction, such as high emotional intensity + clear intent, and use this as the starting node of the knowledge graph; Find all paths from the starting node to nodes in each range of the speech rate dimension, such as slow speech rate, medium speech rate, and fast speech rate, in the graph. For each path, the mapping coefficients of the edges traversed are multiplied sequentially to obtain the weight of a single path; The path weights of all nodes reaching the same speech rate range are added together, and the sum is the cumulative weight of that dimension in the current scenario, which is the reliability of that dimension.

[0095] For example, there are two paths from high emotional intensity + clear intention to slow speech rate. The product of the mapping coefficients of path 1 is 0.6 (edge ​​1 coefficient 0.6), and the product of the mapping coefficients of path 2 is 0.3 (edge ​​1 coefficient 0.5 × edge 2 coefficient 0.6). Therefore, the cumulative weight of slow speech rate is 0.6 + 0.3 = 0.9, that is, the reliability of this dimension is 0.9.

[0096] Weights are assigned based on the importance of each dimension (speech rate 40%, facial expression intensity 30%, and action frequency 30%), and the reliability of each dimension is weighted and summed: Overall reliability = (speech rate reliability × 40%) + (facial expression intensity reliability × 30%) + (action frequency reliability × 30%).

[0097] like Figure 3 As shown, the present invention provides an AI-based digital human interaction system for a nursing bed, comprising: The digital human pre-built module is used to generate a digital virtual image model of a relative based on pre-inputted description data of the relative, and to store the digital virtual image model of the relative in the digital human model library; The virtual model summoning module is used to parse the target digital relative role identifier in the digital relative summoning command input by the user, and retrieve the virtual image model of the target digital relative corresponding to the target digital relative role identifier from the digital human model library; The emotion and intent analysis module is used to analyze the user's emotional state and calling intent based on real-time collected user monitoring audio data. The virtual dynamic rendering module is used to dynamically render the target digital relative's virtual image based on the target digital relative's speech and behavior feature library, the user's emotional state and summoning intention, and the historical interaction records between the user and the corresponding target digital relative's virtual image model, to obtain the current digital human's response animation. The interactive output control module is used to display the current digital human response animation on the screen. At the same time, when the summoning intent includes a control command for the nursing bed itself, the corresponding nursing bed itself is controlled based on the corresponding control command.

[0098] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of this invention and its equivalents, this invention also intends to include these modifications and variations.

Claims

1. An AI-based nursing bed, characterized in that, include: The nursing bed itself; The interactive trigger module is used to receive digital family member summoning commands input by the user; The digital human interaction terminal is located at the head of the nursing bed and is used to retrieve the corresponding digital family virtual image model stored in the digital human model library based on the digital family summoning command. It can then interact with the user or control the nursing bed based on the corresponding digital family virtual image model. The terminal includes a display screen, an audio acquisition module, an audio output module, and a nursing bed control module.

2. A digital human interaction method for an AI-based nursing bed, characterized in that: include: A digital virtual avatar model of a relative is generated based on pre-inputted description data of the relative, and the digital virtual avatar model of the relative is stored in a digital human model library; Parse the target digital relative role identifier in the user's input digital relative summoning command, and retrieve the target digital relative virtual image model corresponding to the target digital relative role identifier from the digital human model library; Analyze users' emotional state and summoning intentions based on real-time collected user monitoring audio data; Based on the behavioral and behavioral feature library of the target digital relative virtual avatar model, the user's emotional state and summoning intention, and the historical interaction records between the user and the corresponding target digital relative virtual avatar model, the target digital relative virtual avatar is dynamically rendered to obtain the current digital human response animation. The current digital human response animation is displayed on the screen. At the same time, when the summoning intent includes a control command for the nursing bed itself, the corresponding nursing bed itself is controlled based on the corresponding control command.

3. The AI-based digital human interaction method for a nursing bed according to claim 2, characterized in that, Digital virtual avatar models of relatives are generated based on pre-inputted descriptions of the relatives, and these models are stored in a digital human model library, including: Extract the family's voice features, facial features, language habits, and behavioral habits from the pre-input description of the family member. Digital virtual image models of relatives are generated based on the voice features, facial morphology features, language habit features, and behavioral habit features of relatives. Store the digital avatars of your loved ones in the digital human model library.

4. The AI-based digital human interaction method for a nursing bed according to claim 2, characterized in that, The system parses the target digital relative role identifier in the user-input digital relative summoning command and retrieves the corresponding virtual image model of the target digital relative from the digital human model library, including: Match the user's input digital family member summoning command with a list of corresponding command-style character identifiers to identify the target digital family member character identifier in the user's input digital family member summoning command; Retrieve the virtual image model of the target digital relative corresponding to the target digital relative role identifier from the digital human model library.

5. The AI-based digital human interaction method for a nursing bed according to claim 2, characterized in that, Based on real-time collected user monitoring audio data analysis, the system identifies users' emotional states and calling intentions, including: Analyze users' emotional state based on real-time collected user monitoring audio data; User behavior data is analyzed based on real-time collected user monitoring audio data. The summoning intent is determined based on the user's emotional state, verbal and behavioral data, and summoning commands.

6. The AI-based digital human interaction method for a nursing bed according to claim 2, characterized in that, Based on the behavioral and behavioral feature database of the target digital relative virtual avatar model, the user's emotional state and summoning intent, and the historical interaction records between the user and the corresponding target digital relative virtual avatar model, the target digital relative virtual avatar is dynamically rendered to obtain the current digital avatar's response animation, including: The summoning response result is determined based on the historical interaction records between the user and the corresponding target digital relative virtual avatar model and the summoning intent. Multiple response mapping matrices were analyzed based on the historical interaction records between users and the corresponding target digital family virtual avatar models. Based on the behavioral and speech habit feature library of the target digital family virtual avatar model and the user's emotional state analysis, a standard response format for the summoning response result is determined. Generate the optimal response form for the summon response results based on all response mapping matrices and the standard response form of the summon response results; Based on the summoning response results and the corresponding optimal response form, the target digital family virtual image is dynamically rendered to obtain the current digital human's response animation.

7. The AI-based digital human interaction method for a nursing bed according to claim 6, characterized in that, Based on the analysis of historical interaction records between users and corresponding target digital family virtual avatar models, multiple response mapping matrices were derived, including: Extract the user's summoning intent, prior emotional state, corresponding response form, and user satisfaction in each historical interaction record between the user and the corresponding target digital family virtual image model. All historical interaction processes with the same summoning intention and corresponding response form are summarized as the first historical interaction process set, and all historical interaction processes with the same prior emotional state and corresponding response form are summarized as the second historical interaction process set. Based on the summoning intent, prior emotional state, corresponding response form and user satisfaction of all historical interaction processes in each first historical interaction process set, the first mapping coefficient between each prior emotional state and corresponding response form under the corresponding summoning intent is analyzed. Based on the summoning intent, prior emotional state, corresponding response form and user satisfaction of all historical interaction processes in each second historical interaction process set, the second mapping coefficient between each summoning intent and corresponding response form under the corresponding prior emotional state is analyzed. The first and second mapping coefficients between each pre-existing emotional state and each response form under each summoning intention are used to determine the final mapping coefficients between each pre-existing emotional state and each response form under each summoning intention; The final mapping coefficients between each pre-existing emotional state and the same response form under each summoning intention are represented by a matrix to obtain the response mapping matrix for each response form.

8. The AI-based digital human interaction method for a nursing bed according to claim 6, characterized in that, Based on the behavioral and behavioral feature database of the target digital family virtual avatar model and the user's emotional state analysis, the standard response format for the summoning response result is derived, including: In the behavioral and speech habit feature database of the target digital family virtual avatar model, multiple behavioral and speech response features of the target digital family virtual avatar under the user's emotional state were retrieved; Based on the summoning intent, the standard response format of the summoning response result is determined by retrieving multiple verbal and behavioral response features of the target digital relative's virtual image in the user's emotional state.

9. The AI-based digital human interaction method for a nursing bed according to claim 6, characterized in that, The optimal response form for the summon response results is generated based on all response mapping matrices and the standard response form of the summon response results, including: The maximum value of the user's emotional state and summoning intention in all response mapping matrices of all response forms is taken as the first target matrix element. The response form corresponding to the response mapping matrix to which the first target matrix element belongs is taken as the candidate response form. The matching degree between the candidate response form and the standard response form is calculated as the first matching degree. The corresponding matrix elements of the user's emotional state and summoning intent in the response mapping matrix in the standard response form are used as the second target matrix elements, and the matching degree of the first target matrix elements and the second target matrix elements is determined as the second matching degree. The co-selection reliability of the candidate response form and the standard response form is calculated based on the first and second matching degrees between the candidate response form and the standard response form. When the overall reliability is not less than the reliability threshold, the optimal response form of the summon response result is obtained by summing the form performance values ​​of the largest reliability among the form performance values ​​of the same dimension in the candidate response form and the standard response form. When the overall reliability is less than the reliability threshold, the response form with the highest overall reliability among the candidate response form and the standard response form is taken as the best response form for the call response result.

10. A digital human interaction system for an AI-based nursing bed, characterized in that: include: The digital human pre-built module is used to generate a digital virtual image model of a relative based on pre-inputted description data of the relative, and to store the digital virtual image model of the relative in the digital human model library; The virtual model summoning module is used to parse the target digital relative role identifier in the digital relative summoning command input by the user, and retrieve the virtual image model of the target digital relative corresponding to the target digital relative role identifier from the digital human model library; The emotion and intent analysis module is used to analyze the user's emotional state and calling intent based on real-time collected user monitoring audio data. The virtual dynamic rendering module is used to dynamically render the target digital relative's virtual image based on the target digital relative's speech and behavior feature library, the user's emotional state and summoning intention, and the historical interaction records between the user and the corresponding target digital relative's virtual image model, to obtain the current digital human's response animation. The interactive output control module is used to display the current digital human response animation on the screen. At the same time, when the summoning intent includes a control command for the nursing bed itself, the corresponding nursing bed itself is controlled based on the corresponding control command.