Intelligent accompanying method and emotion monitoring device
By using an emotion monitoring device to detect and provide real-time feedback on changes in patients' voice, facial expressions, and body posture, and by using a semantic analysis model to judge emotions and conduct voice interaction, the problem of existing smart beds failing to monitor psychological changes has been solved, thereby improving the quality of care and patients' living experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- DONGGUAN DERUCCI BEDDING CO LTD
- Filing Date
- 2024-12-31
- Publication Date
- 2026-04-21
AI Technical Summary
Existing smart bed technology mainly focuses on the patient's physical needs, failing to effectively monitor and respond to the patient's psychological changes, neglecting the psychological problems that arise during long-term bed rest, resulting in the inability to detect and alleviate patients' anxiety and depression in a timely manner.
The system uses an emotion monitoring device to detect the user's emotional state in real time, acquires voice features, facial features, and body posture changes, uses a semantic analysis model to judge the emotion, and adjusts the user's emotion through voice interaction to output personalized audio responses.
It enables real-time monitoring and feedback of patients' emotions, regulates patients' emotions through voice interaction, improves the quality of care and life experience, and alleviates psychological problems in a timely manner.
Smart Images

Figure CN119770823B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of smart bed technology, specifically to a smart care method and an emotion monitoring device. Background Technology
[0002] In modern nursing, bedridden patients, due to prolonged confinement to bed, not only face physical discomfort but also often experience psychological problems such as loneliness, anxiety, and depression. These issues can exacerbate their condition and affect the recovery process. Therefore, designing a smart bed to alleviate user discomfort is an important direction in the current development of smart bed technology.
[0003] Current smart beds typically focus primarily on the patient's physical care, such as mattress pressure adjustment, body rotation, and vital sign monitoring. While these functions can effectively alleviate physical discomfort, they do not pay sufficient attention to the patient's mental health. They cannot monitor and provide real-time feedback on the patient's psychological changes, nor do they offer effective emotional support mechanisms, thus neglecting the psychological problems that arise during prolonged bed rest. Summary of the Invention
[0004] This application discloses an intelligent companionship method and an emotion monitoring device, which detects and provides feedback on the user's emotional state in real time, so as to regulate the user's emotions through voice interaction, improve the quality of care for the user, and thus enhance the user's life experience.
[0005] The first aspect of this application discloses an intelligent companionship method applied to an emotion-monitoring intelligent bed, wherein the intelligent bed is equipped with an emotion monitoring device, and the method includes:
[0006] Acquire user's emotional data, including voice features, facial expression features, and body posture change features; judge different emotional data to obtain different emotional states; obtain the user's target emotional state based on different emotional states and preset weight information for different emotional states, including normal state and abnormal state.
[0007] If the target's emotional state is abnormal, issue the first feedback signal;
[0008] Based on the first feedback signal, the semantic analysis model guides the user to conduct voice interaction and outputs an audio response to the user's voice information.
[0009] As an optional implementation, in the first aspect of this embodiment, different emotional data are judged to obtain different emotional states, including:
[0010] Collect users' voice information and obtain voice features based on the voice information. Voice features include Mel frequency cepstral coefficients, pitch, energy, zero crossover rate, and speech rate.
[0011] Voice features are input into an emotion classification model to obtain emotion categories, and the first emotional state is obtained based on the emotion category.
[0012] As an optional implementation, in the first aspect of this embodiment, collecting the user's voice information and obtaining voice features based on the voice information includes:
[0013] Convert voice information into digital voice information;
[0014] Based on the digital voice information, obtain the noise-removed digital voice information;
[0015] Obtain first digital information whose frequency is greater than a threshold frequency from the noise-removed digital speech information; add a preset frequency to the frequency of the first digital information to obtain second digital information.
[0016] Speech features are obtained based on the second digital information and the third digital information whose frequencies are less than a threshold frequency in the digital speech information after noise removal.
[0017] As an optional implementation, in the first aspect of this embodiment, the smart bed is further equipped with a camera device to judge different emotional data and obtain different emotional states, and also includes:
[0018] Capture user facial video using a camera device;
[0019] Identify key facial regions from facial video;
[0020] Obtain facial key points in key facial regions, and obtain expression features based on facial key points. Expression features include motion information of facial key points.
[0021] Motion information is input into the emotion classification model to obtain the emotion category, and the second emotional state is obtained based on the emotion category.
[0022] As an optional implementation, in the first aspect of this embodiment, different emotional data are judged to obtain different emotional states, which further includes:
[0023] Collect pressure data on the smart bed, and obtain the first pressure data after noise removal based on the pressure data;
[0024] The value of the first pressure data is compressed to a preset range to obtain the second pressure data;
[0025] Convert the second pressure data into time series data;
[0026] Body posture change characteristics are obtained from time-series data, including pressure distribution information, pressure change information, and pressure change frequency information.
[0027] The body posture change features are input into the emotion classification model to obtain the emotion category, and the third emotional state is obtained based on the emotion category.
[0028] As an optional implementation, in the first aspect of this embodiment, the semantic model includes a first semantic model, which guides the user to perform voice interaction based on a first feedback signal through a semantic analysis model, and outputs an audio response corresponding to the user's voice information, including:
[0029] In response to the first feedback signal, the system outputs first voice information based on the first semantic model. The first voice information is used to inquire about the user's current mood.
[0030] Upon receiving the second voice information from the user, the second voice information is analyzed according to the first semantic model to obtain the semantic content of the second voice information;
[0031] The third voice information is generated based on the semantic content and output. The third voice information is used to provide feedback to the user's second voice information.
[0032] As an optional implementation, in a first aspect of this embodiment, after emitting first speech information according to a first semantic model in response to a first feedback signal, the method further includes:
[0033] If no second voice message is received from the user within the first preset time period, a fourth voice message is issued to guide the user in communication.
[0034] As an optional implementation, in a first aspect of this embodiment, the semantic model includes a second semantic model, and after emitting first speech information according to the first semantic model in response to a first feedback signal, the method further includes:
[0035] Based on the speech features, the speech information density within a second preset duration is obtained, and it is determined whether the speech information density is greater than a set threshold. The second preset duration is greater than the first preset duration.
[0036] If the voice information density is greater than a set threshold, a second feedback signal is sent.
[0037] In response to the second feedback signal, the system interacts with the user via voice through the second semantic model and outputs an audio response corresponding to the user's voice information. The second semantic model is different from the first semantic model.
[0038] As an optional implementation, in the first aspect of this embodiment, the method further includes:
[0039] When the voice information density is less than a set threshold, a third feedback signal is sent.
[0040] In response to the third feedback signal, a fifth voice message is issued; the fifth voice message is used to remind the user to select the target working mode of the emotion monitoring device, which is meditation music playback, short story playback or soothing music playback.
[0041] The second aspect of this application discloses an emotion monitoring device, including an emotion perception module, an emotion judgment module, and an intelligent voice interaction system. The emotion judgment module is connected to the emotion perception module and also connected to the intelligent voice interaction system.
[0042] The emotion perception module is used to acquire the user's emotion data and send the emotion data to the emotion judgment module. The emotion data includes voice features, facial expression features, and body posture change features.
[0043] The emotion judgment module is used to judge different emotion data separately to obtain different emotion states; and,
[0044] This is used to obtain the user's target emotional state based on different emotional states and preset weight information for those emotional states. Emotional states include normal and abnormal states; and...
[0045] Used to issue the first feedback signal when the target's emotional state is abnormal;
[0046] The intelligent voice interaction system is used to guide users to conduct voice interaction based on the first feedback signal through a semantic analysis model, and output audio responses to the corresponding user's voice information.
[0047] Compared with related technologies, the embodiments of this application have at least the following beneficial effects:
[0048] The intelligent companionship method disclosed in this application is applied to an emotion monitoring smart bed. The smart bed is equipped with an emotion monitoring device. The method includes: acquiring the user's emotion data, including voice features, facial expression features, and body posture change features; judging different emotion data to obtain different emotion states; obtaining the user's target emotion state based on different emotion states and preset weight information for different emotion states; issuing a first feedback signal when the target emotion state is abnormal; and guiding the user to perform voice interaction based on the first feedback signal through a semantic analysis model, and outputting an audio response corresponding to the user's voice information. This method can detect and provide feedback on the patient's emotional state in real time, regulating the patient's emotions through voice dialogue and user emotions through voice interaction, improving the quality of care for the user, and thus enhancing the user's life experience. Attached Figure Description
[0049] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0050] Figure 1 A flowchart illustrating an intelligent caregiving method provided in this application embodiment;
[0051] Figure 2 A flowchart illustrating an intelligent companionship method for obtaining emotional states, provided as an embodiment of this application;
[0052] Figure 3 A flowchart illustrating another intelligent companionship method for obtaining emotional states provided in an embodiment of this application;
[0053] Figure 4 A flowchart illustrating another intelligent companionship method for obtaining emotional states provided in an embodiment of this application;
[0054] Figure 5 A flowchart illustrating another intelligent companionship method for obtaining emotional states provided in an embodiment of this application;
[0055] Figure 6 A flowchart illustrating another intelligent care method provided in this application embodiment;
[0056] Figure 7 A flowchart illustrating another intelligent care method provided in this application embodiment;
[0057] Figure 8 A flowchart illustrating another intelligent care method provided in this application embodiment;
[0058] Figure 9 This is a schematic diagram of an emotion detection device provided in an embodiment of this application. Detailed Implementation
[0059] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0060] It should be noted that the terms "first, second, third" used in the embodiments of this application are used to distinguish similar or different objects and do not represent a specific order of objects. It can be understood that "first, second, third" can be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0061] It should be noted that the terms "comprising" and "having," and any variations thereof, in the embodiments and accompanying drawings of this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the steps or units listed, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices.
[0062] In modern nursing care, patients who are bedridden for extended periods not only face physical discomforts such as bedsores and muscle atrophy due to lack of movement and positional changes, but also often experience psychological problems such as loneliness, anxiety, and depression caused by prolonged confinement to bed. Studies have shown that patients who are bedridden lack interaction with the outside world and emotional support, making them prone to negative emotions. If these emotions are not addressed and intervened in a timely manner, they may worsen and ultimately affect their physical health.
[0063] Current smart bed technologies primarily focus on the patient's physical needs, such as automatically adjusting mattress pressure, assisting with positional changes, and monitoring vital signs to reduce discomfort and prevent common complications. While these technologies offer a more comfortable care experience to some extent, most smart beds currently fail to effectively address the patient's psychological state. This often leads to caregivers missing the optimal window for psychological intervention when patients experience low mood or anxiety.
[0064] This application discloses an intelligent companionship method and an emotion monitoring device, which detects and provides feedback on the user's emotional state in real time, so as to regulate the user's emotions through voice interaction, improve the quality of care for the user, and thus improve the user's life experience. The following are detailed descriptions.
[0065] The intelligent companionship method disclosed in this application can detect and provide feedback on the user's emotional state in real time, and is applicable to various scenarios, such as hospitals, clinics, nursing homes, home care, schools, and enterprises.
[0066] In hospital and clinic settings, the intelligent companionship method and emotion monitoring device disclosed in this application are applicable to intensive care units, psychiatric wards, and post-operative recovery facilities. In nursing homes and home care settings, elderly people, due to declining physiological functions, are prone to loneliness, anxiety, depression, and other psychological problems, especially those who are bedridden or live alone. The intelligent companionship method and emotion monitoring device disclosed in this application can monitor the elderly's emotional state in real time and provide timely feedback to soothe their emotions and alleviate their psychological problems. In school environments, especially for middle school and university students, academic pressure, social problems, and confusion about self-identity often lead to emotional issues such as anxiety, depression, and excessive stress. The intelligent companionship method and emotion monitoring device disclosed in this application can promptly detect students' emotional states and provide timely feedback to soothe their emotions. Similarly, in corporate environments, especially high-pressure, high-intensity work environments, emotional fluctuations can lead to decreased productivity or employee turnover. The intelligent companionship method and emotion monitoring device disclosed in this application can provide timely psychological counseling to employees, helping them maintain mental health and thereby improving overall work efficiency.
[0067] Please see Figure 1 , Figure 1 A flowchart of an intelligent companionship method provided in this application embodiment, applied to an emotion monitoring smart bed, the smart bed being equipped with an emotion monitoring device, the flowchart including at least the following steps S101-S102:
[0068] Step S101: Obtain the user's emotional data, which includes voice features, facial expression features, and body posture change features;
[0069] Optionally, emotional data may also include physiological data and sleep data, without specific limitations.
[0070] The physiological data can include heart rate, skin conductance, respiratory rate, blood pressure, body temperature, and electroencephalogram (EEG) data, without specific limitations. The sleep data can include sleep duration, sleep cycles, sleep interruptions, and sleep quality, without specific limitations.
[0071] In this step, by setting up an emotion monitoring device on the smart bed, different types of emotion data of the user can be obtained in real time. This allows for a comprehensive assessment of the user's emotional state from multiple dimensions, which can greatly improve the accuracy and effectiveness of emotion monitoring.
[0072] Step S102: Judge different emotional data separately to obtain different emotional states;
[0073] In this step, different types of data are judged independently to identify different emotional states. This process can distinguish the changing patterns of different emotional data, helping to comprehensively capture subtle changes in users' emotions and effectively avoiding emotion recognition errors caused by inaccurate judgment of a single data dimension.
[0074] Step S103: Obtain the user's target emotional state based on different emotional states and their preset weights.
[0075] Emotional states include normal states and abnormal states;
[0076] The preset weighting information was obtained by engineers through numerous experiments. It primarily assigns varying degrees of importance to different types of emotional data in determining the accuracy of user emotion assessment. By allocating these preset weights to different types of emotional data, more accurate emotional state monitoring results can be obtained. For example, in some cases, voice data may be more sensitive in identifying anxiety, while body language data may be more effective in identifying anger. In this way, the system can selectively adjust the weights of different emotional data, thereby improving the overall accuracy of emotion monitoring and making the assessment of emotional state more consistent with the user's actual emotional expression.
[0077] Step S104: If the target's emotional state is abnormal, issue the first feedback signal;
[0078] Optionally, the first feedback signal may include an emotion regulation instruction issued by the system to the emotion monitoring device module, and the emotion monitoring device shall proceed to the next step according to the emotion regulation instruction. The first feedback signal may also include a simple reminder, prompt or emotion regulation suggestion voice signal, with the purpose of making the user aware of the existence of the emotional problem and stimulating him to make adjustment or further interaction.
[0079] In this step, when the emotion monitoring device detects an abnormal emotional state in the target, it can promptly issue an initial feedback signal. This feedback signal serves as a warning, alerting the user that their emotional state may be problematic, thus providing conditions for subsequent intervention and support. For example, the monitoring device identifies abnormal emotional fluctuations in the user and promptly provides feedback to the user, helping them to recognize their current emotional issues and providing an early warning for emotional adjustment.
[0080] Step S105: Based on the first feedback signal, guide the user to perform voice interaction through the semantic analysis model and output the corresponding audio response to the user's voice information.
[0081] In this step, when the emotion monitoring device detects an anomaly, it can guide the user to engage in voice interaction based on the initial feedback signal and output a corresponding audio response through a semantic analysis model. This feedback method can provide a more personalized and emotional response based on the user's actual emotional needs. For example, when the system detects that the user is in an anxious state, it can provide suggestions for relieving emotions through voice interaction, or help the user alleviate emotional stress through a relaxed and pleasant tone. This voice interaction method not only enhances the system's user-friendliness but also effectively improves the user's emotional relief experience.
[0082] The emotion monitoring device can acquire emotion data including voice features; it can be understood that the smart bed is equipped with a microphone sensor, and the emotion monitoring device collects the user's voice information through the built-in microphone sensor, and then obtains voice features through the voice information.
[0083] The emotion data acquired by the emotion monitoring device also includes facial expression features. It can be understood that the smart bed has a built-in camera, and the emotion monitoring device uses the camera to collect facial videos of the user, and then obtains facial expression features based on the facial videos.
[0084] The emotion data acquired by the emotion monitoring device also includes body posture characteristics. It can be understood that the smart mattress is equipped with motion sensors, and the emotion monitoring device obtains the user's body posture characteristics by acquiring pressure data on the smart mattress.
[0085] Optionally, the emotion monitoring device can also acquire physiological characteristics. It is understood that the smart bed can be equipped with various sensors, and the emotion monitoring device acquires physiological characteristics through these sensors. For example, it can acquire heart rate data by embedding a photoplethysmography (PPG) sensor in the smart bed's mattress; acquire skin conductance data by equipping the smart bed or wearable device with a skin conductance sensor; acquire respiratory rate data by embedding a pressure sensor or airflow sensor in the smart bed's mattress or pillow; acquire pressure data by setting a vibration sensor or airbag sensor in the mattress; acquire body temperature data by setting an infrared sensor or contact sensor in the smart bed's mattress or pillow; and acquire brainwave data by integrating an electroencephalogram (EEG) sensor in the smart bed's mattress or pillow.
[0086] Optionally, the emotion monitoring device can also acquire sleep characteristics, which can be obtained by monitoring the user's sleep behavior and physiological characteristics. It is understood that the smart bed can be equipped with multiple sensors, and the emotion monitoring device uses these sensors and sleep monitoring algorithms to acquire sleep characteristics. For example, by setting up accelerometers and heart rate sensors on the smart bed and combining them with sleep monitoring algorithms, sleep time, sleep cycles, and sleep interruptions can be acquired; by setting up multiple sensors on the smart bed, physiological characteristics can be acquired, and sleep quality can be obtained based on these physiological characteristics and sleep monitoring algorithms.
[0087] For example, firstly, emotion data collection is the foundation of the entire process. The system collects users' emotion data in real time through various sensors, cameras, microphones, and other devices. Emotion data can include three dimensions: voice features, facial expression features, and body language features. Voice features include tone of voice, speech rate, volume, and changes in tone, which can reflect the user's emotional fluctuations; facial expression features reflect subtle changes in emotion through changes in facial expressions, such as anger, joy, and anxiety; body language features reflect the user's emotional state through body posture, movements, and body language, for example, tension may manifest as tense shoulders and restless hands and feet.
[0088] Secondly, after collecting this multi-dimensional emotional data, the emotional state is determined. The emotion monitoring device analyzes and integrates voice features, facial expression features, and body posture changes using a pre-set emotion recognition model to determine the emotional state represented by each feature. For example, if a user speaks faster and in a higher tone, it may indicate that the user is tense or anxious; if the user's facial expression shows a frown or downturned corners of the mouth, it may indicate that the user is depressed; if the user's body posture is curled up or avoidance, it may indicate some kind of unease or fear. Through the analysis of these features, the emotion monitoring device comprehensively judges the user's emotional state, which is divided into two categories: "normal state" and "abnormal state." By combining the pre-set weight information corresponding to each emotional state, the system obtains the user's target emotional state. When the system determines that the user's target emotional state is abnormal, the system will issue a first feedback signal. The purpose of this signal is to remind the user that their current emotional state is unstable or unhealthy and requires guidance or intervention.
[0089] Finally, based on the initial feedback signal, the system further guides voice interaction. Through a semantic analysis model, the system deeply understands the user's voice input, generates corresponding emotional guidance content, and sends this content as an audio response for personalized voice communication with the user. This audio response typically uses speech synthesis technology to convert the feedback content into a natural and easily understood speech form. The audio response may include warm comfort, suggestions for emotion regulation techniques, or even some encouraging positive language. Therefore, this companionship method, through multi-dimensional emotional data collection, accurate emotion judgment, intelligent feedback mechanisms, and voice interaction guidance, forms a closed-loop emotion regulation system that can provide users with precise emotional support in real-time, dynamic emotional changes, helping them regulate negative emotions, improve emotional stability, and promote emotional health.
[0090] Optionally, semantic analysis models can help machines understand contextual relationships, implicit meanings, and various semantic levels in language. There are many types of semantic analysis models, which can be selected according to actual needs. For example, there are rule-based semantic analysis models, vector space-based semantic analysis models, context-based pre-trained semantic analysis models, graph-based semantic analysis models, deep learning-based semantic analysis models, and multimodal semantic analysis models, etc., without specific limitations here.
[0091] Rule-based semantic analysis models typically use manually defined rules and dictionaries for semantic analysis. Examples include word sense disambiguation and syntactic-semantic combination, but no specific restrictions are imposed here.
[0092] In a semantic analysis model based on vector space, the meaning of words and sentences is represented by vectors. This can usually be modeled using methods such as the bag-of-words model or TF-IDF, without making any specific restrictions here.
[0093] Context-based pre-trained language models are those that are pre-trained on large-scale corpora and can capture the contextual information of words, sentences, and paragraphs well. Examples include Bidirectional Encoder Representations from Transformers (BERT), Generative Pretrained Transformer (GPT), and Text-to-Text Transfer Transformer (T5), etc. No specific restrictions are made here.
[0094] Graph-based semantic analysis models typically use graph neural networks for analysis, such as knowledge graphs, and no specific restrictions are imposed here.
[0095] There are many semantic analysis models based on deep learning, and you can choose according to your actual needs, such as recurrent neural networks, long short-term memory, and gated recurrent units, etc., without making specific restrictions here.
[0096] Multimodal semantic analysis models integrate information from different modalities for joint analysis, such as Contrastive Language-Image Pretraining (CLIP) and Dall·E. No specific restrictions are imposed here, and the appropriate model can be selected according to actual needs.
[0097] In one embodiment, when the emotion monitoring device detects that the user's target state is abnormal, the emotion monitoring device can also send information about the user's abnormal emotion to the terminal device of the user's relatives or caregivers, so as to remind the user's relatives or caregivers to pay attention to the user's emotional state in a timely manner and take corresponding measures to alleviate the user's emotions.
[0098] In one embodiment, the emotion monitoring device also uses password login, encryption protocols during data transmission, and secure data storage to protect user privacy. For example, family members or caregivers need to log in with a password to authenticate with the emotion monitoring device before they can access the user's emotional state data. Furthermore, if family members or caregivers want to access the user's emotional data, the emotion monitoring device encrypts the data using TLS or SSL encryption protocols before transmission to ensure that data transmission between the emotion monitoring device and the family member or caregiver is not obtained or tampered with by others. Additionally, data stored on the emotion monitoring device uses storage encryption algorithms, such as AES, to ensure that even in the event of a data breach, the stored emotional data cannot be read by third parties, thereby protecting user privacy.
[0099] In some embodiments, the method in step S102, which involves judging different emotional data to obtain different emotional states, can be as follows:
[0100] Collect users' voice information and obtain voice features based on the voice information. Voice features include Mel frequency cepstral coefficients, pitch, energy, zero crossover rate, and speech rate.
[0101] Voice features are input into an emotion classification model to obtain emotion categories, and the first emotional state is obtained based on the emotion category.
[0102] For a clearer understanding of the above methods, please refer to [link / reference]. Figure 2 , Figure 2 A flowchart of an intelligent companionship method for obtaining emotional states provided in an embodiment of this application is included, which at least includes the following steps S201-S202:
[0103] Step S201: Collect the user's voice information and obtain voice features based on the voice information.
[0104] Among them, speech features include Mel frequency cepstral coefficients, pitch, energy, zero crossover rate, and speech rate;
[0105] Mel-frequency cepstral coefficients are one of the most commonly used features in speech signal processing to extract the audio characteristics of speech. They simulate the perceptual characteristics of human hearing, converting the spectrum to a Mel scale, which is more consistent with the human ear's sensitivity to different frequencies. They can capture the speaker's speech features (such as timbre, pitch, and voice quality). In emotion recognition, because emotions affect the pitch and timbre of speech (e.g., anger, happiness, or tension), they can be used to distinguish different emotional states.
[0106] Pitch refers to the highness or lowness of a sound, usually determined by frequency. When a person speaks, the frequency of vibration of the vocal cords determines the pitch. Changes in pitch can reflect emotional fluctuations. For example, when angry or anxious, a person's pitch tends to rise, while in a relaxed or contemplative state, the pitch may be lower. In emotion recognition, changes in pitch can help the system determine whether a user is in an excited or tense state.
[0107] Energy refers to the intensity or amplitude of a speech signal over a given period of time. It reflects the volume of the speaker's voice and typically varies depending on their emotions. For example, when angry, people's voices are usually more intense and have higher energy; while when sad or depressed, the energy is lower. Changes in energy can also help distinguish the intensity and type of emotion.
[0108] The zero-crossing rate refers to the number of times a speech signal waveform crosses zero, that is, the number of times the signal switches from the positive half-cycle to the negative half-cycle or from the negative half-cycle to the positive half-cycle. It is often used to describe the roughness of the signal, that is, the clarity of speech. A higher zero-crossing rate may indicate that the sound has more high-frequency components, such as a tense, excited, or angry tone; a lower zero-crossing rate may be associated with a calm and relaxed state.
[0109] Speech rate refers to the number of syllables or words spoken per unit of time. The speed of speech is directly related to the speaker's psychological and emotional state. Generally, people speak faster when they are tense, anxious, or excited, and slower when they are calm or sad. Changes in speech rate are an important signal of emotion regulation, and are particularly effective in identifying emotions such as anxiety, anger, and excitement.
[0110] Step S202: Input the speech features into the emotion classification model to obtain the emotion category, and obtain the first emotional state based on the emotion category.
[0111] The sentiment classification model can be a traditional machine learning model, which can be selected according to actual needs, such as Support Vector Machine (SVM), Decision Trees, Random Forest, K-Nearest Neighbor (KNN), Naive Bayes, etc., without specific restrictions.
[0112] The available emotion categories include pleasure, sadness, anger, fear, surprise, disgust, anxiety, confidence, depression, excitement, frustration, and calmness, without specific restrictions.
[0113] The primary emotional state is determined as normal or abnormal based on the positive and negative emotions represented by these emotion categories. For example, if the acquired emotion category is positive (pleasure), the primary emotional state is normal. If the acquired emotion category is negative (anxiety), the primary emotional state is abnormal.
[0114] In some embodiments, the acquired raw speech signal can also be input into a deep learning model to obtain the emotion category. The deep model can be selected according to actual needs. For example, Convolutional Neural Networks (CNN), Recurrent Neural Networks (RNN), Deep Neural Networks (DNN), and Transformer models with self-attention mechanisms can be selected according to actual needs.
[0115] In one embodiment, the method in step S201: collecting the user's voice information and obtaining voice features based on the voice information, can be:
[0116] Convert voice information into digital voice information;
[0117] Based on the digital voice information, obtain the noise-removed digital voice information;
[0118] Obtain first digital information whose frequency is greater than a threshold frequency from the noise-removed digital speech information; add a preset frequency to the first digital information frequency to obtain second digital information.
[0119] Speech features are obtained based on the second digital information and the third digital information whose frequencies are less than a threshold frequency in the digital speech information after noise removal.
[0120] For a clearer understanding of the above methods, please refer to [link / reference]. Figure 3 , Figure 3 A flowchart of another intelligent companionship method for obtaining emotional states provided in an embodiment of this application is included, which at least includes the following steps S301-S305:
[0121] Step S301: Collect the user's voice information and convert the voice information into digital voice information;
[0122] Step S302: Obtain the noise-removed digital speech information based on the digital speech information;
[0123] Step S303: Obtain first digital information with a frequency greater than a threshold frequency in the digital speech information after noise removal processing; add a preset frequency to the first speech signal frequency to obtain second digital information.
[0124] Step S304: Obtain speech features based on the second digital information and the third digital information whose frequencies are less than the threshold frequency in the digital speech information after noise removal.
[0125] Step S305: Input the speech features into the emotion classification model to obtain the emotion category, and obtain the first emotional state based on the emotion category.
[0126] In this method, clearer digital speech information is obtained by removing noise from the digital speech information, and clearer digital speech information can be obtained by adding high-frequency features to the digital speech information. That is, the first digital information with a frequency greater than a threshold frequency is obtained from the digital speech information, and a preset frequency is added to the first speech signal frequency to obtain the second digital information. This is because the mid-to-high frequency components in the speech information usually contain more detailed information, such as timbre and tone changes, thereby improving the intelligibility and clarity of the speech, so as to facilitate the subsequent extraction of speech features.
[0127] In some embodiments, the smart bed is also equipped with a camera to collect facial video information of the user. Regarding the method in step S102: judging different emotional data to obtain different emotional states, it can be:
[0128] Capture user facial video using a camera device;
[0129] Identify key facial regions from facial video;
[0130] Obtain facial key points in key facial regions, and obtain facial expression features based on facial key points. Facial expression features include motion information of facial key points.
[0131] Motion information is input into the emotion classification model to obtain the emotion category, and the second emotional state is obtained based on the emotion category.
[0132] For a clearer understanding of the above methods, please refer to [link / reference]. Figure 4 , Figure 4 A flowchart illustrating another intelligent companionship method for obtaining emotional states provided in this application embodiment, the flowchart including at least the following steps S401-S404:
[0133] Step S401: Capture the user's facial video using a camera device;
[0134] Step S402: Obtain key facial regions based on the facial video;
[0135] Step S403: Obtain facial key points in key facial regions, and obtain facial expression features based on facial key points;
[0136] Among them, facial expression features include motion information of key facial points;
[0137] Step S404: Input the motion information into the emotion classification model to obtain the emotion category, and obtain the second emotional state based on the emotion category.
[0138] This method first requires capturing facial video of the user using a camera. This video is the foundation of sentiment analysis, providing dynamic changes in the user's facial expressions. The facial video capture typically consists of multiple consecutive frames, which capture the evolution of the user's facial expressions over time. Next, key facial regions, including the eyes, mouth, eyebrows, and nose, are obtained from the facial video. Computer vision algorithms, such as Haar Cascade, MTCNN, Dlib, and MediaPipe, can be used to locate and define these key regions. Then, facial key points are obtained, such as the center of the pupil, the position of the upper and lower lips, and the starting point of the eyebrows. These key points are typically obtained using facial calibration models, such as Dlib's 68 facial calibration points or MediaPipe's 468 calibration points. Finally, the changes in the positions of these key points are tracked and input into the sentiment classification model to analyze changes in facial expressions. For example, an upward turn of the corners of the mouth may indicate a smile, while a downward turn of the eyebrows may indicate anger or confusion.
[0139] In some embodiments, the method in step S102, which involves judging different emotional data to obtain different emotional states, can also be:
[0140] Collect pressure data on the smart bed, and obtain the first pressure data after noise removal based on the pressure data;
[0141] The value of the first pressure data is compressed to a preset range to obtain the second pressure data;
[0142] Convert the second pressure data into time series data;
[0143] Body posture change characteristics are obtained from time-series data, including pressure distribution information, pressure change information, and pressure change frequency information.
[0144] The body posture change features are input into the emotion classification model to obtain the emotion category, and the third emotional state is obtained based on the emotion category.
[0145] For a clearer understanding of the above methods, please refer to [link / reference]. Figure 5 , Figure 5 A flowchart of another intelligent companionship method for obtaining emotional states provided in an embodiment of this application, the flowchart including at least the following steps S501-S505:
[0146] Step S501: Collect pressure data on the smart bed, and obtain the first pressure data after noise removal based on the pressure data;
[0147] Step S502: Compress the value of the first pressure data to a preset range to obtain the second pressure data;
[0148] Step S503: Convert the second pressure data into time series data;
[0149] Step S504: Obtain postural change characteristics based on time-series data;
[0150] Among them, the characteristics of body posture change include pressure distribution information, pressure change information, and pressure change frequency information;
[0151] Step S505: Input the body posture change features into the emotion classification model to obtain the emotion category, and obtain the third emotional state based on the emotion category.
[0152] In this method, the smart bed is equipped with multiple pressure sensors that can be placed within the mattress to detect the user's weight and position. These pressure sensors can capture the user's movements, positional changes, and pressure distribution at different locations in real time. The pressure data output by these sensors is typically pressure values over a specific time interval, representing the pressure intensity at different locations. After obtaining the pressure data from the sensors, to make the data more accurate and meaningful, filtering algorithms (such as low-pass filtering, Kalman filtering, etc.) are usually used to remove noise caused by environmental factors, equipment errors, or the sensor's own technical limitations, thus obtaining the noise-removed first pressure data. Next, to facilitate subsequent processing and analysis, the first pressure data needs to be compressed or normalized to a preset range to obtain the second pressure data. This preset range is usually a standard range (such as 0 to 1, or a specific pressure intensity interval), allowing comparisons on the same scale and avoiding data inconsistencies caused by differences in the sensitivity of different sensors. Finally, the second pressure data is converted into time-series data, taking into account the time sequence and variation patterns of the data, making it more meaningful for understanding the user's state. Finally, based on time-series data, information on stress distribution, stress variation, and stress frequency changes are obtained. This data on postural changes is then input into an emotion classification model to determine the emotion category, thereby identifying the third emotional state. For example, based on stress frequency variation information, it can be determined whether a user frequently adjusts their posture over a period of time. If frequent posture changes are detected within a certain period, it indicates that the user is experiencing discomfort or anxiety.
[0153] In some embodiments, the above-described steps for obtaining the first, second, and third emotional states can be simultaneously monitored and obtained in the emotion monitoring device. Alternatively, two of any three emotional states can be monitored, or one of any three emotional states can be monitored. Users can choose according to their actual needs. It is understood that the emotion monitoring device consumes more energy when monitoring multiple emotional states simultaneously, and is more energy-efficient when monitoring fewer types of emotional states.
[0154] In some embodiments, a fourth emotional state may be obtained based on physiological data, and a fifth emotional state may be obtained based on sleep data; no specific limitations are imposed here.
[0155] In some embodiments, the semantic model includes a first semantic model, which, in the method of step S105, guides the user to perform voice interaction based on the first feedback signal through a semantic analysis model and outputs an audio response corresponding to the user's voice information, can be:
[0156] In response to the first feedback signal, the system outputs first voice information based on the first semantic model. The first voice information is used to inquire about the user's current mood.
[0157] Upon receiving the second voice information from the user, the second voice information is analyzed according to the first semantic model to obtain the semantic content of the second voice information;
[0158] The third voice information is generated based on the semantic content and output. The third voice information is used to provide feedback to the user's second voice information.
[0159] For a clearer understanding of the above methods, please refer to [link / reference]. Figure 6 , Figure 6 A flowchart of another intelligent care method provided in the embodiments of this application, the flowchart including at least the following steps S601-S607:
[0160] Step S601: Obtain the user's emotional data, which includes voice features, facial expression features, and body posture change features;
[0161] Optionally, emotional data may also include physiological data and sleep data, without specific limitations.
[0162] The physiological data can include heart rate, skin conductance, respiratory rate, blood pressure, body temperature, and electroencephalogram (EEG) data, without specific limitations. The sleep data can include sleep duration, sleep cycles, sleep interruptions, and sleep quality, without specific limitations.
[0163] Step S602: Judge different emotional data separately to obtain different emotional states;
[0164] Step S603: Obtain the user's target emotional state based on different emotional states and their preset weights.
[0165] Emotional states include normal states and abnormal states;
[0166] The preset weight information was obtained by technicians through multiple experiments. It mainly refers to the allocation of different emotional data to the accuracy of user emotion judgment. When the preset weight information is allocated to different types of emotional data, more accurate emotional state monitoring results can be obtained.
[0167] Step S604: If the target's emotional state is abnormal, issue the first feedback signal;
[0168] Optionally, the first feedback signal may include an emotion regulation instruction issued by the system to the emotion monitoring device module, and the emotion monitoring device shall proceed to the next step according to the emotion regulation instruction;
[0169] Optionally, the first feedback signal may also include a simple reminder, prompt, or voice signal suggesting emotion regulation, with the aim of making the user aware of the existence of emotional problems and prompting them to regulate or further interact.
[0170] Step S605: In response to the first feedback signal, output the first voice information according to the first semantic model. The first voice information is used to ask the user about their current mood.
[0171] Step S606: Upon receiving the second voice information from the user, analyze the second voice information according to the first semantic model to obtain the semantic content of the second voice information;
[0172] Step S607: Generate corresponding third voice information based on semantic content, and output the third voice information. The third voice information is used to provide feedback to the user's second voice information.
[0173] In some embodiments, after step S605, i.e. after emitting first speech information according to the first semantic model in response to the first feedback signal, the method further includes:
[0174] If no second voice message is received from the user within the first preset time period, a fourth voice message is issued to guide the user in communication.
[0175] For a clearer understanding of the above methods, please refer to [link / reference]. Figure 7 , Figure 7 A flowchart of another intelligent care method provided in the embodiments of this application, the flowchart including at least the following steps S701-S706:
[0176] Step S701: Obtain the user's emotional data, which includes voice features, facial expression features, and body posture change features;
[0177] Optionally, emotional data may also include physiological data and sleep data, without specific limitations.
[0178] The physiological data can include heart rate, skin conductance, respiratory rate, blood pressure, body temperature, and electroencephalogram (EEG) data, without specific limitations. The sleep data can include sleep duration, sleep cycles, sleep interruptions, and sleep quality, without specific limitations.
[0179] Step S702: Judge different emotional data separately to obtain different emotional states;
[0180] Step S703: Obtain the user's target emotional state based on different emotional states and preset weight information for different emotional states;
[0181] Emotional states include normal states and abnormal states;
[0182] The preset weight information was obtained by technicians through multiple experiments. It mainly refers to the allocation of different emotional data to the accuracy of user emotion judgment. When the preset weight information is allocated to different types of emotional data, more accurate emotional state monitoring results can be obtained.
[0183] Step S704: If the target's emotional state is abnormal, issue the first feedback signal;
[0184] Optionally, the first feedback signal may include an emotion regulation instruction issued by the system to the emotion monitoring device module, and the emotion monitoring device shall proceed to the next step according to the emotion regulation instruction;
[0185] Optionally, the first feedback signal may also include a simple reminder, prompt, or voice signal suggesting emotion regulation, with the aim of making the user aware of the existence of emotional problems and prompting them to regulate or further interact.
[0186] Step S705: In response to the first feedback signal, output the first voice information according to the first semantic model. The first voice information is used to ask the user about their current mood.
[0187] Step S706: If no second voice message is received from the user within the first preset waiting time, a fourth voice message is sent to guide the user's communication.
[0188] Optionally, the fourth voice message can be gentle and caring, such as: "Hello, I can sense that you seem a little unhappy. If you'd like, I'm always here to listen. Would you like to talk to me now?"
[0189] Alternatively, the fourth voice message can be encouraging, such as: "I understand that speaking isn't always easy, but I'm here to listen. If you have anything to say, or just want some quiet time, that's perfectly fine."
[0190] Alternatively, the fourth voice message can also be emotionally reassuring, such as, “If you are feeling uncomfortable right now, you can relax and take a break. If you need any help, or just need someone to keep you company, I’m here for you.”
[0191] For example, if the emotion detection device detects an abnormal emotion in a user, the device will send a first voice message based on a speech model to ask the user about their current mood. The purpose is to determine the user's emotional perception and reaction to external information. If the user does not react to external information within a first preset time, or is in a state of not wanting to speak, it indicates that the user's emotions are abnormal. Therefore, a fourth voice message needs to be sent to guide the user to communicate. This fourth voice message conveys care and understanding, avoids putting pressure on the user, and gives them enough space to choose whether to communicate further, respecting the user's emotional needs and boundaries.
[0192] In some embodiments, the semantic model includes a second semantic model, and after step S605, i.e. after emitting first speech information according to the first semantic model in response to the first feedback signal, the method further includes:
[0193] Based on the speech features, the speech information density within a second preset duration is obtained, and it is determined whether the speech information density is greater than a set threshold. The second preset duration is greater than the first preset duration.
[0194] If the voice information density is greater than a set threshold, a second feedback signal is sent.
[0195] In response to the second feedback information, the system interacts with the user via voice through the second semantic model and outputs an audio response corresponding to the user's voice information. The second semantic model is different from the first semantic model.
[0196] In some embodiments, after obtaining the voice information density within a second preset duration based on voice features and determining whether the voice information density is greater than a set threshold, a third feedback signal is sent if the voice information density is less than the set threshold.
[0197] In response to the third feedback signal, a fifth voice message is issued; the fifth voice message is used to remind the user to select the target working mode of the emotion monitoring device, which is meditation music playback, short story playback or soothing music playback.
[0198] For a clearer understanding of the above methods, please refer to [link / reference]. Figure 8 , Figure 8 A flowchart of another intelligent care method provided in the embodiments of this application, the flowchart including at least the following steps S801-S808:
[0199] Step S801: Obtain the user's emotional data, which includes voice features, facial expression features, and body posture change features;
[0200] Optionally, emotional data may also include physiological data and sleep data, without specific limitations.
[0201] The physiological data can include heart rate, skin conductance, respiratory rate, blood pressure, body temperature, and electroencephalogram (EEG) data, without specific limitations. The sleep data can include sleep duration, sleep cycles, sleep interruptions, and sleep quality, without specific limitations.
[0202] Step S802: Judge different emotional data separately to obtain different emotional states;
[0203] Step S803: Obtain the user's target emotional state based on different emotional states and their preset weights;
[0204] Emotional states include normal states and abnormal states;
[0205] The preset weight information was obtained by technicians through multiple experiments. It mainly refers to the allocation of different emotional data to the accuracy of user emotion judgment. When the preset weight information is allocated to different types of emotional data, more accurate emotional state monitoring results can be obtained.
[0206] Step S804: If the target's emotional state is abnormal, issue the first feedback signal;
[0207] Optionally, the first feedback signal may include an emotion regulation instruction issued by the system to the emotion monitoring device module, and the emotion monitoring device shall proceed to the next step according to the emotion regulation instruction;
[0208] Optionally, the first feedback signal may also include a simple reminder, prompt, or voice signal suggesting emotion regulation, with the aim of making the user aware of the existence of emotional problems and prompting them to regulate or further interact.
[0209] Step S805: In response to the first feedback signal, output the first voice information according to the first semantic model. The first voice information is used to ask the user about their current mood.
[0210] Step S806: Obtain the speech information density within the second preset duration based on the speech features, and determine whether the speech information density is greater than a set threshold.
[0211] Step S807: When the voice information density is greater than a set threshold, a second feedback signal is sent. In response to the second feedback information, the user is interacted with via voice through the second semantic model, and an audio response corresponding to the user's voice information is output.
[0212] Step S808: When the voice information density is less than a set threshold, a third feedback signal is sent, and in response to the third feedback signal, a fifth voice message is sent.
[0213] The fifth voice message is used to remind the user to select the target working mode of the emotion monitoring device, which is meditation music playback, short story playback, or soothing music playback.
[0214] Optionally, the target working mode can also be broadcasting hilarious short clips, fun quizzes, etc., without specific restrictions.
[0215] For example, after the emotion detection device determines that the user's target emotional state is abnormal, and the emotion detection device emits first voice information according to the first semantic model, it waits for a second preset time period, obtains the density of the user's voice information received within that second preset time period, and determines whether the obtained voice information density is greater than a set threshold, and whether the second preset time period is greater than the first preset time period. This is because the strength of the user's desire to confide can be determined by judging the density of voice information received by the emotion detection device after a relatively long time interval.
[0216] When the density of voice information exceeds a set threshold, it can be determined that the user has a strong desire to express themselves. Based on this strong desire, the emotion detection device can invoke a second semantic model to communicate with the user. The content logic generated by this second semantic model differs from that of the first semantic model. The first semantic model focuses on guiding the user's communication, while the second semantic model emphasizes engaging in more fun and relaxed communication. Therefore, this second semantic model can possess characteristics such as humor and emotional perception, diversity, randomness, creativity and interactivity, and high fluency.
[0217] When the density of voice information is less than a set threshold, it can be determined that the user has little desire to confide, i.e., is not in a state of wanting to talk. Therefore, the emotion detection device sends a third feedback signal. Based on the third feedback signal, the emotion detection device sends a fifth message, prompting the user to select the target working mode of the emotion detection device. After the user selects the target working mode, if the selected target working mode is meditation music playback, the emotion detection device can play meditation music to help the user relax, improve concentration, improve mood, promote sleep, and enhance mental health. Alternatively, if the selected target working mode is short story playback, the emotion detection device can bring the user a pleasant emotional experience through this light and entertaining method, promoting mental health, language development, emotional resonance, creativity, and cognitive abilities in many ways. In addition, users can also unconsciously gain a lot of mental growth and emotional comfort. Alternatively, if the selected target working mode is soothing music playback, the emotion detection device will also play soothing music to help the user reduce stress and alleviate emotions.
[0218] This method uses different emotion-soothing modes based on the density of voice information to provide users with more emotionally relevant feedback, helping them to better regulate their emotions and psychological state, thereby achieving a more effective emotion-soothing effect.
[0219] In some embodiments, the emotion monitoring device also includes a network search function, allowing users to select the meditation music track, story title, or soothing music title to be played according to their preferences. For example, a user can verbally express the meditation music track they wish to hear to the emotion monitoring device. After receiving the voice information, the emotion monitoring device uses the network search function to search for the meditation music track by name, and plays the meditation music after it is found.
[0220] In some embodiments, the emotion monitoring device can monitor the user's emotions in real time during the process of soothing the user's emotions, and provide real-time feedback to correct the user's emotional state. If the user's emotional state is detected to be normal for a period of time, the emotion monitoring device can verbally remind the user whether they want to take a break. After receiving feedback from the user, the device can continue to soothe the user's emotions based on the user's feedback. For example, if the user chooses not to take a break, the emotion monitoring device will continue to soothe the user's emotions. If the user chooses to take a break, the emotion monitoring device will stop working.
[0221] This application also discloses an emotion detection device; please refer to [link / reference needed]. Figure 9 , Figure 9This is a schematic diagram of an emotion detection device disclosed in an embodiment of this application, including an emotion perception module 91, an emotion judgment module 92, and an intelligent voice interaction system 93. The emotion judgment module 92 is connected to the emotion perception module 91 and also connected to the intelligent voice interaction system 93.
[0222] The emotion perception module 91 is used to acquire the user's emotion data and send the emotion data to the emotion judgment module. The emotion data includes voice features, facial expression features and body posture change features.
[0223] The emotion judgment module 92 is used to judge different emotion data separately to obtain different emotion states; and,
[0224] This is used to obtain the user's target emotional state based on different emotional states and preset weight information for those emotional states. Emotional states include normal and abnormal states; and...
[0225] Used to issue the first feedback signal when the target's emotional state is abnormal;
[0226] The intelligent voice interaction system 93 is used to guide the user to conduct voice interaction through a semantic analysis model based on the first feedback signal, and output the corresponding audio response to the user's voice information.
[0227] Based on the above-described intelligent companionship method, this application provides an electronic device, including: a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the computer program, it implements the visual system positioning and detection algorithm described above.
[0228] Based on the above-described intelligent companionship method, this application embodiment also provides a computer program product, including a computer program that, when executed by a processor, implements the visual system positioning and detection algorithm described above.
[0229] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a non-volatile computer-readable storage medium. When executed, the program can include the processes of the embodiments of the methods described above. The storage medium can be a magnetic disk, optical disk, ROM, etc.
[0230] Any references to memory, storage, databases, or other media used herein may include non-volatile and / or volatile memory. Suitable non-volatile memory may include ROM, Programmable ROM (PROM), Erasable PROM (EPROM), Electrically Erasable PROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM), which is used as an external cache. By way of illustration and not limitation, RAM may take many forms, such as Static RAM (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchlink DRAM (SLDRAM), Rambus DRAM (RDRAM), and Direct Rambus DRAM (DRDRAM).
[0231] It should be understood that the phrase "one embodiment" or "an embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Those skilled in the art should also recognize that the embodiments described in the specification are optional embodiments, and the actions and modules involved are not necessarily essential to this application.
[0232] In the various embodiments of this application, it should be understood that the sequence number of each process does not necessarily imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0233] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; they can be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0234] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0235] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that there can be three kinds of relationships. For example, object A and / or object B can represent three situations: object A exists alone, object A and object B exist simultaneously, and object B exists alone.
[0236] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0237] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.
[0238] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.
[0239] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method or device embodiments.
[0240] The intelligent companionship method disclosed in the embodiments of this application has been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only for the purpose of helping to understand the method and its core ideas. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A smart bed, characterized in that, The smart bed is equipped with an emotion monitoring device and a camera device. The emotion monitoring device is configured to, Acquire user's emotional data, which includes voice features, facial expression features, and body posture change features; Different emotional data are judged to obtain different emotional states. The user's target emotional state is obtained based on the different emotional states and the preset weight information of the different emotional states. The emotional states include normal state and abnormal state. If the target's emotional state is abnormal, a first feedback signal is issued; Based on the first feedback signal, the user is guided to perform voice interaction through a semantic analysis model, and an audio response corresponding to the user's voice information is output. The step of judging different emotional data to obtain different emotional states includes: Collect the user's voice information, obtain the voice features based on the voice information; input the voice features into an emotion classification model to obtain an emotion category, and obtain a first emotional state based on the emotion category; The camera device captures a video of the user's face; the facial key regions are obtained from the video; the facial key points of the facial key regions are obtained; the facial expression features are obtained from the facial key points; the facial expression features are input into the emotion classification model to obtain the emotion category; and the second emotional state is obtained from the emotion category. The pressure data on the smart bed is collected, and first pressure data after noise removal is obtained based on the pressure data; the value of the first pressure data is compressed to a preset range to obtain second pressure data; the second pressure data is converted into time series data; the body posture change features are obtained based on the time series data, the body posture change features are input into the emotion classification model to obtain the emotion category, and a third emotional state is obtained based on the emotion category.
2. The smart bed according to claim 1, characterized in that, The speech features include Mel frequency cepstral coefficients, pitch, energy, zero crossover rate, and speech rate.
3. The smart bed according to claim 1, characterized in that, The emotion monitoring device is configured to collect the user's voice information and obtain voice features based on the voice information, including: Convert the voice information into digital voice information; Based on the digital voice information, obtain the noise-removed digital voice information; Obtain first digital information whose frequency is greater than a threshold frequency from the noise-removed digital speech information; add a preset frequency to the frequency of the first digital information to obtain second digital information. The speech features are obtained based on the second digital information and the third digital information in the noise-removed digital speech information where the frequency is less than the threshold frequency.
4. The smart bed according to claim 1, characterized in that, The facial expression features include motion information of the key facial points.
5. The smart bed according to claim 1, characterized in that, The body posture change characteristics include pressure distribution information, pressure change information, and pressure change frequency information.
6. The smart bed according to claim 1, characterized in that, The semantic analysis model includes a first semantic model, and the emotion monitoring device is configured to guide the user to perform voice interaction based on the first feedback signal through the semantic analysis model, and output an audio response corresponding to the user's voice information, including: In response to the first feedback signal, the system outputs first voice information based on the first semantic model, the first voice information being used to inquire about the user's current mood. Upon receiving the second voice information from the user, the second voice information is analyzed according to the first semantic model to obtain the semantic content of the second voice information; Based on the semantic content, corresponding third voice information is generated and output. The third voice information is used to provide feedback to the user on the second voice information.
7. The smart bed according to claim 6, characterized in that, The emotion detection device is further configured to, after issuing first speech information based on the first semantic model in response to the first feedback signal... If no second voice message is received from the user within a first preset time period, a fourth voice message is issued to guide the user in communication.
8. The smart bed according to claim 7, characterized in that, The semantic analysis model includes a second semantic model, and the emotion detection device is further configured to, after emitting first speech information according to the first semantic model in response to the first feedback signal... Based on the speech features, obtain the speech information density within a second preset duration, and determine whether the speech information density is greater than a set threshold, wherein the second preset duration is greater than the first preset duration. If the voice information density is greater than the set threshold, a second feedback signal is sent. In response to the second feedback signal, the system interacts with the user via voice through the second semantic model and outputs an audio response corresponding to the user's voice information. The second semantic model is different from the first semantic model.
9. The smart bed according to claim 8, characterized in that, The emotion detection device is also configured to: If the voice information density is less than the set threshold, a third feedback signal is sent. In response to the third feedback signal, a fifth voice message is issued; the fifth voice message is used to remind the user to select the target working mode of the emotion monitoring device, the target working mode being meditation music playback, short story playback, or soothing music playback.
Citation Information
Patent Citations
Man-machine interactive method and apparatus used for intelligent robot
CN106843458A
Dynamic interaction method, server, electronic equipment and storage medium
CN112148850A