Health interaction method and device based on emotion recognition, equipment and storage medium
By collecting and analyzing users' multimodal data, utilizing pre-trained emotion recognition models, adjusting interaction strategies, and generating personalized voice response information, the problem of low interactive experience in intelligent voice assistants is solved, and the user experience and emotional support capabilities are improved.
Patent Information
- Application Number
- CN202511371924.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-24
- Publication Date
- 2026-02-06
AI Technical Summary
Existing intelligent voice assistants offer a low level of user experience, lack personalization and emotional support capabilities, and are unable to meet the needs of ordinary users.
By collecting multimodal data from users, including voice signals, user images, and physiological parameters, feature extraction is performed and the data is input into a pre-trained emotion recognition model to determine the user's emotion tags. Based on the emotion tags, the interaction strategy is adjusted to generate personalized voice response information.
It enhances the personalization and emotional support capabilities of user interaction, improving the user experience, especially for the elderly and patients with chronic diseases, by providing more caring health explanations and making the test data easy to understand.
Smart Images

Figure CN121478112A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent interaction, in particular to a health interaction method and device based on emotion recognition, equipment and storage medium. BACKGROUND
[0002] With the rapid development of artificial intelligence technology, intelligent voice assistants are increasingly widely used in daily life. However, the existing intelligent voice assistants only directly output the problems queried by users, and ordinary users lack professional knowledge and often have difficulty in understanding the specific meanings. Moreover, the interaction mode is less interesting, it is difficult to interact with users individually, the emotional care ability is weak, and the user experience is low. SUMMARY
[0003] The embodiments of the present application provide a health interaction method and device based on emotion recognition, equipment and storage medium, to at least solve the technical problem of low user experience of voice interaction mode in related technologies.
[0004] According to an aspect of an embodiment of the present application, a health interaction method based on emotion recognition is provided, comprising:
[0005] Collecting multi-modal data of a user, the multi-modal data comprising voice signals, user images and physiological parameters;
[0006] Performing feature extraction on the multi-modal data to obtain multi-modal features;
[0007] Inputting the multi-modal features into a pre-trained emotion recognition model to obtain an emotion probability distribution of the user, and determining an emotion label of the user based on the emotion probability distribution;
[0008] Determining an interaction strategy based on the emotion label, generating a health explanation text corresponding to the voice signals of the user, and obtaining voice reply information based on the interaction strategy and the health explanation text.
[0009] In one embodiment, determining an interaction strategy based on the emotion label, generating a health explanation text corresponding to the voice signals of the user, and obtaining voice reply information based on the interaction strategy and the health explanation text, comprises:
[0010] Adjusting the interaction strategy based on the emotion label, the interaction strategy comprising at least one or more of tone, speed, and rhetoric;
[0011] Generating a health explanation text corresponding to the voice signals of the user based on the physiological parameters;
[0012] Converting the health explanation text into reply voice, and post-processing the reply voice based on the interaction strategy to generate the voice reply information.
[0013] In an embodiment, the health interpretation text corresponding to the user voice signal is generated based on the physiological parameter, comprising:
[0014] A detection result is determined according to the physiological parameter;
[0015] In the case of an abnormal detection result, personalized life data of the user is obtained, and an abnormal reason and a health suggestion are analyzed based on the personalized life data;
[0016] A preset extensible template library is used to convert the detection result, the abnormal reason and the health suggestion into an initial health interpretation text;
[0017] The initial health interpretation text is input into a large language model to obtain the health interpretation text.
[0018] In an embodiment, the personalized life data of the user is obtained, and the abnormal reason and the health suggestion are analyzed based on the personalized life data, comprising:
[0019] Behavior logs of the user in a preset period are collected, and the behavior logs include diet logs, exercise logs and emotion logs;
[0020] Behavior feature extraction and cluster analysis are performed based on the behavior logs to obtain a behavior pattern of the user;
[0021] The abnormal reason and the health suggestion are determined based on the behavior pattern.
[0022] In an embodiment, feature extraction is performed on the multi-modal data to obtain multi-modal features, comprising:
[0023] Voice activity detection and microphone array beamforming preprocessing are performed on the voice signal to obtain a preprocessed voice signal, and acoustic features are extracted from the preprocessed voice signal;
[0024] Face detection is performed on the user image to obtain facial image features;
[0025] The physiological parameter is filtered to obtain physiological parameter features;
[0026] The multi-modal features are obtained based on the acoustic features, the facial image features and the physiological parameter features;
[0027] The multi-modal features are time-aligned to obtain time-aligned multi-modal data.
[0028] In an embodiment, the multi-modal features are input into a pre-trained emotion recognition model to obtain an emotion probability distribution of the user, and an emotion label of the user is determined based on the emotion probability distribution, comprising:
[0029] obtaining a context vector based on user historical interaction data;
[0030] concatenating the multi-modal feature and the context vector into a long vector to obtain a fusion vector in each time window;
[0031] inputting the fusion vector into a pre-trained emotion recognition model to obtain an emotion probability distribution of the user;
[0032] determining the emotion label based on the emotion with the highest probability.
[0033] In an embodiment, before inputting the multi-modal feature into the pre-trained emotion recognition model, further comprising:
[0034] constructing the emotion recognition model, wherein the emotion recognition model adopts a multi-layer Transformer network structure, and each layer contains a multi-head self-attention mechanism and a feedforward neural network;
[0035] constructing a training data set based on the fusion vectors of multiple time windows and corresponding emotion labels;
[0036] training the emotion recognition model using the training data set, a cross-entropy loss function, and a modality consistency regularization method.
[0037] According to another aspect of the embodiments of the present application, a health interaction device based on emotion recognition is provided, comprising:
[0038] a collection module configured to collect multi-modal data of a user, wherein the multi-modal data includes voice signals, user images, and physiological parameters;
[0039] a feature extraction module configured to perform feature extraction on the multi-modal data to obtain multi-modal features;
[0040] an emotion recognition module configured to input the multi-modal features into a pre-trained emotion recognition model to obtain an emotion probability distribution of the user, and determine an emotion label of the user based on the emotion probability distribution;
[0041] an interaction module configured to determine an interaction strategy based on the emotion label, generate a health explanation text corresponding to the voice signal of the user, and obtain voice reply information based on the interaction strategy and the health explanation text.
[0042] According to another aspect of the embodiments of the present application, an electronic device is also provided, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the health interaction method based on emotion recognition by using the computer program.
[0043] According to a further aspect of the embodiment of the present application, a computer readable storage medium is also provided, which stores a computer program. The computer program is configured to execute the health interaction method based on emotion recognition when running.
[0044] The technical solution provided by the embodiment of the present application can include the following beneficial effects:
[0045] The present application collects multi-modal data of the user, including voice signals, user images and physiological parameters. After feature extraction, the data is input into a pre-trained emotion recognition model to obtain the emotion probability distribution of the user, and the emotion label of the user is determined accordingly. Based on the emotion label, the system can dynamically adjust the interaction strategy, and finally form the voice reply information. This multi-modal fusion method not only improves the accuracy of emotion recognition, but also dynamically adjusts the communication strategy according to the emotional state of the user, provides a more caring and personalized interactive experience, and is especially suitable for groups such as the elderly, patients with chronic diseases and other groups that need emotional care. In addition, the present application can return a colloquial health explanation text, making the detection data easy to understand and improving the user experience.
[0046] Further, feedback of the user on the health suggestions can be collected, transcribed into text and subjected to sentiment analysis and topic modeling, so as to continuously optimize the interaction strategy and the generation quality of the health explanation text, and discover the user demand. BRIEF DESCRIPTION OF DRAWINGS
[0047] The accompanying drawings, which are included to provide a further understanding of the present application, constitute a part of the present application and illustrate the illustrative embodiments of the present application and their description serve to explain the present application, and do not constitute an improper limitation on the present application. In the drawings:
[0048] Figure 1 is a flowchart of a health interaction method based on emotion recognition according to an embodiment of the present application;
[0049] Figure 2 is a flowchart of another health interaction method based on emotion recognition according to an embodiment of the present application;
[0050] Figure 3 is a schematic diagram of a health interaction device based on emotion recognition according to an embodiment of the present application;
[0051] Figure 4 is a structural schematic diagram of an optional electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0052] In order for those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should be within the scope of protection of the present application.
[0053] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0054] The emotion recognition-based health interaction method of the embodiments of the present application will be described in detail below in conjunction with the drawings. As shown in the figure, the method mainly includes the following steps: Figure 1 The method mainly includes the following steps:
[0055] S101 collects multi-modal data of the user, and the multi-modal data includes voice signals, user images and physiological parameters.
[0056] In an embodiment, when the user interacts with the terminal, the multi-modal acquisition unit synchronously acquires voice, facial video and physiological signals.
[0057] Specifically, the voice signals are acquired through a microphone array set by the terminal, the user images are acquired through a camera, and the physiological signals of the user are acquired through a data acquisition interface module, which integrates multiple communication protocols such as Bluetooth Low Energy (BLE), Wi-Fi and USB, to realize connection with different types of health detection devices. For example, heart rate, body temperature, blood pressure and other information.
[0058] S102 extracts features from the multi-modal data to obtain multi-modal features.
[0059] In an embodiment of the present application, the multi-modal data is subjected to feature extraction to obtain multi-modal features, which includes voice activity detection and microphone array beamforming preprocessing of the voice signals, and extracting acoustic features from the preprocessed voice signals.
[0060] Specifically, the collected speech signal is pre-processed, specifically including voice activity detection (VAD) and microphone array beamforming. Voice activity detection is used to identify speech segments and non-speech segments, remove silence and background noise, and improve the quality of the speech signal. Microphone array beamforming enhances the speech signal in the target direction by adjusting the weight of the microphone array, and suppresses noise and interference in other directions. After these preprocessing steps, the resulting speech signal is clearer and more stable. Subsequently, acoustic features such as fundamental frequency (F0), intensity and voiceprint are extracted from the pre-processed speech signal, which will be used for subsequent emotion recognition and speech analysis.
[0061] Further, the user image is subjected to face detection to obtain facial image features.
[0062] Specifically, MTCNN (Multi-Task Cascaded Convolutional Neural Networks) is used to process each frame of video data. MTCNN first detects the face region, and then aligns the key points of the face, such as eyes, nose and mouth, to ensure the consistency and standardization of the face image. In this way, high-quality facial image features can be effectively extracted, providing accurate data basis for subsequent emotion recognition and facial expression analysis.
[0063] Further, the physiological parameters are filtered to obtain physiological parameter features. Based on the acoustic features, facial image features and physiological parameter features, multi-modal features are obtained.
[0064] In an embodiment of the present application, the timestamps of the speech signal, user image and physiological parameters are also extracted respectively; the timestamps of the user image and physiological parameters are aligned with the timestamps of the audio frames of the speech signal; the multi-modal data after timestamp alignment is subjected to interpolation processing to obtain time series aligned multi-modal data.
[0065] Specifically, in order to realize the synchronous processing of multi-modal data, the timestamps of the speech signal, user image and physiological parameters are first extracted respectively. These timestamps are used to mark the specific time points of each modal data. Then, taking the audio frames of the speech signal as the reference, the timestamps of the user image and physiological parameters are aligned with the timestamps of the audio frames. This step ensures the consistency of data in different modalities in time. Finally, the multi-modal data after timestamp alignment is subjected to interpolation processing to fill in the small differences in time, thereby obtaining time series aligned multi-modal data. This process enables speech, image and physiological signals to be analyzed on the same time sequence.
[0066] S103 inputs the multi-modal features into the pre-trained emotion recognition model to obtain an emotion probability distribution of the user, and determines an emotion label of the user based on the emotion probability distribution.
[0067] In an embodiment of the present application, a context vector is obtained based on user historical interaction data; in each time window, the multi-modal features and the context vector are spliced into a long vector to obtain a fusion vector; the fusion vector is input into the pre-trained emotion recognition model to obtain an emotion probability distribution of the user; and an emotion label is determined based on the emotion with the highest probability.
[0068] Specifically, the user's recent round of interaction text or voice transcription content is encoded into a context vector and multi-modal features using a BERT, RoBERTa or other pre-trained language model encoder to improve the accuracy of emotion recognition and context awareness.
[0069] In each time window, the multi-modal features and the context vector are spliced into a long vector to obtain a fusion vector.
[0070] In each time window, the acoustic features of the voice signal, the facial image features of the user image, and the physiological features of the physiological parameters are spliced to obtain a window-level vector. This process integrates features of different modalities together, providing a comprehensive feature representation for subsequent emotion recognition and health monitoring.
[0071] Optionally, the multi-modal features and the context vector can also be fused to obtain a fusion vector.
[0072] Further, the fusion vector is input into the pre-trained emotion recognition model to obtain an emotion probability distribution of the user.
[0073] Specifically, the input features of the model include the following categories:
[0074] Acoustic features: including fundamental frequency and intensity, these features can reflect the acoustic characteristics of the voice, and are helpful for emotion recognition. AU vector: facial action unit vector, with a dimension of 17. These features can capture the changes of facial expressions, and are an important basis for emotion recognition. HRV indicator: heart rate variability (HRV) indicator, such as RMSSD (Root Mean Square of Successive Differences), used to reflect the changes of physiological signals, which is helpful for identifying emotional state. The context vector can also be input, such as the summary vector of the last 5 rounds of conversation, which is used to provide the context information of the conversation to help the model better understand the background of the current emotion.
[0075] The emotion recognition model adopts a multi-layer Transformer model architecture, for example, a 4-layer Transformer model, each layer containing a multi-head self-attention mechanism and a feedforward neural network. The multi-head self-attention mechanism can capture features in different subspaces, improving the model's ability to fuse multi-modal data. The feedforward neural network performs a nonlinear transformation on the output of the self-attention mechanism, further extracting features. A residual connection and layer normalization are added after each sublayer to improve the training stability and performance of the model.
[0076] The model outputs an emotion probability distribution. In one implementation, the model outputs the probability distribution of 8 emotions, for example: anxiety: 0.8, calm: 0.2, anger: 0.1, excitement: 0.05, sadness: 0.03, happiness: 0.02, surprise: 0.01, disgust: 0.01. The model also outputs a confidence value representing the credibility of the emotion recognition result.
[0077] The emotion label is determined based on the emotion with the highest probability. In one implementation scenario, the probability value of anxiety is the highest, and the emotion label of the user is determined as anxiety.
[0078] In one embodiment of the present application, before inputting the multi-modal features into the pre-trained emotion recognition model, it further includes: constructing an emotion recognition model, the emotion recognition model adopts a multi-layer Transformer network structure, each layer containing a multi-head self-attention mechanism and a feedforward neural network; based on the fusion vectors of multiple time windows and the corresponding emotion labels, constructing a training data set; training the emotion recognition model using the training data set, a cross-entropy loss function, and a modal consistency regularization method.
[0079] S104 determines an interaction strategy based on the emotion label, generates a health explanation text corresponding to the user's voice signal, and obtains voice reply information based on the interaction strategy and the health explanation text.
[0080] In one embodiment of the present application, the interaction strategy is first regulated based on the emotion label, and the interaction strategy at least includes one or more of tone, speech rate, and rhetoric.
[0081] According to the emotional state of the user, various interaction strategies can be adjusted to achieve a more personalized, considerate and effective interaction experience. For example, according to the emotion, the speech rate is adjusted, anxiety or tension: the speech rate is appropriately slowed down to give the user more time to understand and react. Calm or relaxed: the speech rate remains normal or slightly faster to maintain the smoothness of the conversation. Anger: the speech rate remains stable to avoid speeding up or slowing down to avoid exacerbating the emotion. Excitement or happiness: the speech rate can be appropriately accelerated to increase the vitality of the interaction.
[0082] The tone can also be adjusted according to the emotion, such as soft and soothing for anxiety or tension, using a gentle tone. For anger, the tone remains steady, avoiding being too high or too low to escalate the emotion. For excitement or happiness, the tone can be appropriately raised to increase the vitality of the interaction. For sadness or depression, the tone is low and gentle, expressing sympathy and support.
[0083] The tone can also be adjusted according to the emotion, such as soft and soothing for anxiety or tension, using a gentle tone. For anger, the tone remains steady, avoiding being too high or too low to escalate the emotion. For excitement or happiness, the tone can be appropriately raised to increase the vitality of the interaction. For sadness or depression, the tone is low and gentle, expressing sympathy and support.
[0084] In the case of anxiety or tension, more explanations and support are provided to guide the user step by step. For example, explain the operation process in steps and provide detailed guidance. For anger, remain calm and provide solutions to avoid further irritation. For example, provide multiple solutions for the user to choose from and express the willingness to help. For excitement or happiness, increase the interest and participation of the interaction, such as asking relevant questions or suggestions and encouraging the user to share more. For sadness or depression, provide emotional support, listen to the user's feelings and express sympathy and understanding. For example, provide comforting words and encourage the user to express their emotions.
[0085] Further, based on the physiological parameters, health interpretation text corresponding to the user's voice signal is generated.
[0086] The detection result is determined according to the physiological parameters. The normal, high and low numerical interval of the health index is obtained from the authoritative medical guidelines, and the real-time collected physiological parameters are matched with these intervals, and the normal, high or low state of the numerical value is judged and the result is output. For example, the medical guidelines stipulate that the normal interval of blood pressure is 90-139 mmHg for systolic pressure and 60-89 mmHg for diastolic pressure, the high interval is 140-159 mmHg for systolic pressure and 90-99 mmHg for diastolic pressure, and the low interval is less than 90 mmHg for systolic pressure and less than 60 mmHg for diastolic pressure. If the systolic pressure is detected to be 145 mmHg and the diastolic pressure is 92 mmHg, the system determines that the blood pressure data is in the high state after matching.
[0087] In the case of abnormal detection result, the user's personalized life data is obtained, and the abnormal reason and health suggestion are analyzed based on the personalized life data.
[0088] Further, the identity of the user is determined based on the voiceprint features of the user, or the identity of the user is determined based on the face image of the user.
[0089] Collecting behavior logs of a user for a preset period, the behavior logs including diet logs, exercise logs, and mood logs; performing behavior feature extraction and clustering analysis based on the behavior logs to obtain a behavior pattern of the user; determining an abnormal reason and a health suggestion based on the behavior pattern.
[0090] Specifically, to help users better manage their health, the system first collects the user's diet, exercise, and mood records. The diet records include the types, quantities, and times of meals each day; the exercise records cover the types, durations, intensities, and times of exercise; and the mood records include the user's emotional state. These data are preprocessed, including data cleaning, standardization, and feature extraction, to ensure the accuracy and consistency of the data. Subsequently, the system uses clustering analysis to classify the user's behavior patterns and lifestyle habits. By analyzing the clustering results, potential behavior patterns are identified, such as high-calorie diet, lack of exercise, and frequent anxiety, etc. Finally, the system analyzes the abnormal reasons based on these potential behavior patterns, for example, if the user measures blood glucose within 30 minutes after a meal and the value increases, the system will combine the measurement time and diet logs to output an explanation "may be related to the meal just eaten". And provide personalized health suggestions for the user, such as adjusting the dietary structure, developing an exercise plan, and providing emotional management methods, to help users improve their lifestyle habits and prevent health problems.
[0091] Further, a preset extensible template library is used to convert the detection results, abnormal reasons, and health suggestions into initial health explanation texts.
[0092] The natural language generation module converts the detection results and abnormal reasons into initial health explanation texts through a preset extensible template library. The corresponding template is selected and filled with specific data to generate easy-to-understand health explanations.
[0093] Further, the initial health explanation texts are input into a large language model to obtain health explanation texts. Further optimization generates more natural and accurate health explanation texts to improve user understanding and acceptance.
[0094] Specifically, the generated initial texts are input into a large language model, combined with the user's historical health data, lifestyle habits, and contextual conversations, to generate more natural and personalized health explanation texts.
[0095] The system can return colloquial health interpretation texts, making detection data easy to understand and improving user experience. Feedback texts or voices from users on health suggestions are collected, transcribed into texts, and subjected to sentiment analysis and topic modeling. Through sentiment analysis, the user's satisfaction with health suggestions is understood, and topic modeling extracts the core content and focus of user feedback. Based on user feedback, the large language model is trained and optimized to improve the quality and personalization of generated texts. The interaction strategy and the quality of health interpretation text generation are continuously optimized to discover user needs.
[0096] The health interpretation text is converted into a reply voice, and the reply voice is post-processed based on the interaction strategy to generate voice reply information.
[0097] In the embodiments of the present application, the system first converts the generated health interpretation text into a voice form through text-to-speech (TTS) technology to form a preliminary reply voice. Subsequently, according to the previously determined interaction strategy, the reply voice is post-processed, which may include adjusting the speech rate of the voice to match the emotional state of the user, changing the intonation to convey appropriate emotional color, or optimizing the volume of the voice to ensure clarity and comfort. Or add some comforting and encouraging rhetoric. Through these post-processing steps, the system can generate more personalized and emotional voice reply information, thereby improving the user's interaction experience and making them feel understood and cared for.
[0098] In one embodiment, it also includes physiological parameters recorded by a sliding window, and detects whether the user has a trend anomaly based on the physiological parameter data in the sliding window.
[0099] Specifically, by maintaining a sliding window to record the user's daily health detection data in the last 14 days or the last 100 measurement data. First, calculate the Z-score based on the data in the sliding window. If the Z value is greater than 3, it is judged that the data is abnormal. The exponential weighted moving average (EWMA) method can also be used to detect the trend of the data. If the EWMA value exceeds the set threshold for K consecutive times, a trend anomaly alarm is triggered, so that the user's trend health anomaly can be discovered and warned in time.
[0100] In the presence of abnormal conditions, the corresponding warning level is determined based on the correspondence between the abnormal value and the preset warning level. The warning level includes emergency, high risk, and attention. If it is an emergency state, the instantaneous value exceeds the critical threshold, the system immediately prompts the user to seek emergency medical treatment through voice, and at the same time pushes alarm information to family members or emergency telephone, ensuring that the user can quickly obtain professional medical assistance.
[0101] In order to facilitate understanding of the method of the embodiments of the present application, the following will be described in conjunction with the accompanying Figure 2 further description.
[0102] AsFigure 2 As shown, first, the multi-modal data of the user is collected, including speech signals, user images, and physiological parameters. Then, feature extraction and cross-modal alignment are performed on these data to obtain multi-modal features. These features are then input into a pre-trained emotion recognition model to obtain the emotion probability distribution of the user, and based on this, the emotion label is determined. Based on the emotion label, the system generates a spoken health explanation text corresponding to the user's voice instruction, and determines the interaction strategy, including tone, speed, and rhetoric, etc. Subsequently, the system converts the text into a voice reply and performs post-processing to generate the final voice reply information. In addition, the system also continuously monitors the user's physiological parameters, and if an abnormal trend is found, active intervention is performed.
[0103] The present application proposes an innovative health management voice interaction method, aiming to realize accurate identity and emotion recognition by integrating multi-modal data such as voice, image, and physiological parameters. The system first collects multi-modal data of the user, then performs feature extraction and cross-modal alignment to obtain unified multi-modal feature representation. These features are then input into a pre-trained emotion recognition model to determine the user's emotional state. Based on the identified emotion label, the system dynamically adjusts the interaction strategy, including tone, speed, and rhetoric, to provide personalized health interaction experience and enhance user stickiness. In addition, the system also has the function of active health management, by continuously monitoring the user's physiological parameters and analyzing their trends, combined with emotion anomaly detection, to realize early intervention and promote the user's health management.
[0104] According to another aspect of the embodiments of the present application, a health interaction device based on emotion recognition for implementing the above-mentioned health interaction method based on emotion recognition is also provided. As shown, Figure 3 The device comprises:
[0105] The acquisition module 301 is configured to acquire multi-modal data of the user, and the multi-modal data comprises speech signals, user images, and physiological parameters.
[0106] The feature extraction module 302 is configured to perform feature extraction on the multi-modal data to obtain multi-modal features.
[0107] The emotion recognition module 303 is configured to input the multi-modal features into a pre-trained emotion recognition model to obtain the emotion probability distribution of the user, and determine the emotion label of the user based on the emotion probability distribution.
[0108] The interaction module 304 is configured to determine the interaction strategy based on the emotion label, generate a health explanation text corresponding to the user's speech signals, and obtain voice reply information based on the interaction strategy and the health explanation text.
[0109] It should be noted that the health interaction device based on emotion recognition provided by the above-mentioned embodiments is only exemplified by the division of the above-mentioned functional modules when executing the health interaction method based on emotion recognition. In actual application, the above-mentioned functions can be completed by different functional modules according to the needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the health interaction device based on emotion recognition and the health interaction method based on emotion recognition provided by the above-mentioned embodiments belong to the same concept, which embodies the realization process details of the method embodiments, which will not be repeated here.
[0110] According to another aspect of the embodiments of the present application, an electronic device corresponding to the health interaction method based on emotion recognition provided by the above-mentioned embodiments is also provided to execute the above-mentioned health interaction method based on emotion recognition.
[0111] Please refer to Figure 4 , which shows a schematic diagram of an electronic device provided by some embodiments of the present application. As Figure 4 shown, the electronic device includes a processor 400, a memory 401, a bus 402 and a communication interface 403, the processor 400, the communication interface 403 and the memory 401 are connected through the bus 402; the memory 401 stores a computer program executable on the processor 400, and the processor 400 executes the computer program to execute the health interaction method based on emotion recognition provided by any of the preceding embodiments of the present application.
[0112] Among them, the memory 401 can contain a high-speed random access memory (RAM: Random Access Memory), and can also include a non-volatile memory, such as at least one disk memory. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 403 (which can be wired or wireless), and the Internet, wide area network, local network, metropolitan area network, etc. can be used.
[0113] The bus 402 can be an ISA bus, a PCI bus, or an EISA bus, etc. The bus can be divided into address bus, data bus, control bus, etc. Among them, the memory 401 is used to store programs, and the processor 400 executes the programs after receiving the execution instructions. The health interaction method based on emotion recognition disclosed in any of the preceding embodiments of the present application can be applied to the processor 400 or realized by the processor 400.
[0114] The processor 400 can be an integrated circuit chip with a processing capability of signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit of hardware in the processor 400 or the instruction in the form of software. The processor 400 described above can be a general processor, including a central processing unit (CPU), a network processor (NP), etc.; can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a ready programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. Each method, step and logic block diagram disclosed in the embodiments of the present application can be implemented or executed. The general processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as a hardware code processor for execution, or a combination of hardware and software modules in the code processor for execution. The software module can be located in a random access memory, a flash memory, a read only memory, a programmable read only memory or an electrically erasable programmable memory, a register, etc. The storage medium in the art. The storage medium is located in the memory 401, and the processor 400 reads the information in the memory 401, and combines the hardware to complete the steps of the above method.
[0115] The electronic device provided by the embodiments of the present application and the health interaction method based on emotion recognition provided by the embodiments of the present application have the same inventive concept, and have the same beneficial effects as the method adopted, run or implemented by the application program stored therein.
[0116] According to another aspect of the embodiments of the present application, a computer readable storage medium corresponding to the health interaction method based on emotion recognition provided by the preceding embodiments is also provided, and the computer readable storage medium has a computer program (i.e. program product) stored thereon. When the computer program is run by a processor, the health interaction method based on emotion recognition provided by any of the preceding embodiments is executed.
[0117] It should be noted that examples of the computer readable storage medium can also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory or other optical, magnetic storage medium, which will not be described one by one here.
[0118] The computer readable storage medium provided by the above embodiments of the present application and the health interaction method based on emotion recognition provided by the embodiments of the present application have the same inventive concept, and have the same beneficial effects as the method adopted, run or implemented by the application program stored therein.
[0119] The technical features of the above embodiments can be combined in any manner. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described, however, as long as the combinations of the technical features do not contradict each other, they should be considered to be within the scope of the present disclosure.
[0120] The above embodiments only express several implementation manners of the present application, and the description is relatively specific and detailed, but it should not be understood as a limitation on the patent scope of the present application. It should be pointed out that, for ordinary skilled persons in the art, several modifications and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.
Claims
1. A health interaction method based on emotion recognition, characterized in that, include: Collect multimodal data from users, including speech signals, user images, and physiological parameters; Feature extraction is performed on the multimodal data to obtain multimodal features; The multimodal features are input into a pre-trained emotion recognition model to obtain the user's emotion probability distribution, and the user's emotion label is determined based on the emotion probability distribution. Based on the emotion tags, an interaction strategy is determined, and a health explanation text corresponding to the user's voice signal is generated. Based on the interaction strategy and the health explanation text, voice response information is obtained.
2. The method according to claim 1, characterized in that, Based on the emotion tags, an interaction strategy is determined, and a health explanation text corresponding to the user's voice signal is generated. Based on the interaction strategy and the health explanation text, voice response information is obtained, including: The interaction strategy is adjusted based on the emotion tag, and the interaction strategy includes at least one or more of tone, speaking speed, and speech. Generate a health explanation text corresponding to the user's voice signal based on the physiological parameters; The health explanation text is converted into a response voice, and the response voice is post-processed based on the interaction strategy to generate the voice response information.
3. The method according to claim 2, characterized in that, Based on the physiological parameters, a health explanation text corresponding to the user's voice signal is generated, including: The test results are determined based on the physiological parameters. In the event of abnormal test results, the user's personalized lifestyle data is obtained, and the cause of the abnormality and health recommendations are analyzed based on the personalized lifestyle data. A pre-defined, extensible template library is used to transform the detection results, causes of abnormalities, and health recommendations into initial health explanation text; The initial health explanation text is input into the large language model to obtain the health explanation text.
4. The method according to claim 3, characterized in that, Acquire users' personalized lifestyle data, and analyze the causes of anomalies and provide health recommendations based on the personalized lifestyle data, including: Collect user behavior logs for a preset time period, including diet logs, exercise logs, and emotion logs; Based on the behavior logs, behavioral features are extracted and cluster analysis is performed to obtain the user's behavior patterns; The causes of the abnormalities and health recommendations are determined based on the behavioral patterns described.
5. The method according to claim 1, characterized in that, Feature extraction is performed on the multimodal data to obtain multimodal features, including: The speech signal is preprocessed by speech activity detection and microphone array beamforming to obtain a preprocessed speech signal, and acoustic features are extracted from the preprocessed speech signal. Face detection is performed on the user image to obtain facial image features; The physiological parameters are filtered to obtain physiological parameter characteristics; The multimodal features are obtained based on the acoustic features, facial image features, and physiological parameter features. The multimodal features are time-aligned to obtain time-aligned multimodal data.
6. The method according to claim 1, characterized in that, The multimodal features are input into a pre-trained emotion recognition model to obtain the user's emotion probability distribution. Based on the emotion probability distribution, the user's emotion label is determined, including: Context vectors are obtained based on users' historical interaction data; Within each time window, the multimodal features and the context vector are concatenated into a long vector to obtain the fused vector; The fused vector is input into a pre-trained emotion recognition model to obtain the user's emotion probability distribution; The emotion label is determined based on the emotion with the highest probability.
7. The method according to claim 1, characterized in that, Before inputting the multimodal features into the pre-trained emotion recognition model, the following steps are also included: The emotion recognition model is constructed, which adopts a multi-layer Transformer network structure, with each layer containing a multi-head self-attention mechanism and a feedforward neural network. A training dataset is constructed based on the fusion vectors from multiple time windows and their corresponding sentiment labels; The emotion recognition model is trained using the aforementioned training dataset, cross-entropy loss function, and modality consistency regularization.
8. A health interaction device based on emotion recognition, characterized in that, include: The acquisition module is used to acquire the user's multimodal data, which includes voice signals, user images, and physiological parameters. The feature extraction module is used to extract features from the multimodal data to obtain multimodal features; The emotion recognition module is used to input the multimodal features into a pre-trained emotion recognition model to obtain the user's emotion probability distribution, and to determine the user's emotion label based on the emotion probability distribution. The interaction module is used to determine the interaction strategy based on the emotion tag, generate health explanation text corresponding to the user's voice signal, and obtain voice response information based on the interaction strategy and the health explanation text.
9. An electronic device, characterized in that, It includes a processor and a memory storing program instructions, the processor being configured to perform the emotion-based health interaction method as described in any one of claims 1 to 7 when executing the program instructions.
10. A computer-readable medium, characterized in that, It stores computer-readable instructions that are executed by a processor to implement a health interaction method based on emotion recognition as described in any one of claims 1 to 7.