Intelligent accompanying large model-based elderly user emotional state interaction method and system

CN120780265AInactive Publication Date: 2025-10-14HEALTH HOPE (BEIJING) TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510963365.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-14
Publication Date
2025-10-14
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies make it difficult to monitor the emotional state of elderly users in real time and accurately, especially through multimodal data fusion, and are unable to provide timely feedback to guardians and personalized interactions when emotions are abnormal, resulting in inefficient emotional care.

Method used

Using a large model of intelligent companionship, multimodal data (physiological data and voice data) of elderly users is collected in real time. A standardized time series data set is generated through multimodal emotion fusion technology, abnormal emotions are determined, and prompt information is sent to guardians for personalized voice interaction.

Benefits of technology

It realizes real-time and accurate monitoring of the emotional state of elderly users, timely discovers potential emotional problems, improves the timeliness and effectiveness of emotional care, and enhances the mental health and user experience of elderly users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120780265A_ABST
    Figure CN120780265A_ABST
Patent Text Reader

Abstract

The invention provides an elderly user emotional state interaction method and system based on an intelligent accompanying large model. The senior user emotional state interaction method comprises the steps that daily data information of senior users is collected in real time, data processing is conducted on the daily data information, and a standardized time sequence data set corresponding to the daily data information is generated; fusing the daily data information of the elderly user by using multi-modal emotion fusion to obtain fused emotion parameters, and judging whether the elderly user has emotion abnormity or not by using the emotion parameters of the elderly user; and when the senior user has the abnormal emotion, sending prompt information to a guardian, and performing voice interaction with the senior user according to the emotion state of the senior user. The system comprises modules corresponding to the steps of the method.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application provides an elderly user emotional state interaction method and system based on an intelligent companion large model, and belongs to the technical field of emotional data processing. BACKGROUND

[0002] The physical and mental health problems of the elderly have attracted widespread attention, and emotional state care is particularly important.

[0003] At present, there are many deficiencies in the emotional state monitoring and interaction of elderly users. On the one hand, traditional emotional monitoring methods mostly rely on artificial periodic inquiry or simple questionnaire survey. This method is not only inefficient, but also difficult to capture the emotional changes of the elderly in real time and accurately. Artificial inquiry may be limited by time interval and cannot timely discover the emotional problems of the elderly in specific situations; questionnaire survey has the problems of low cooperation degree of the elderly and careless filling, resulting in poor data accuracy. On the other hand, although some existing intelligent devices have certain monitoring functions, they often only focus on physiological indicator monitoring, such as heart rate, blood pressure, etc., and lack sufficient attention and effective means for emotional state monitoring. Even if some devices involve emotional monitoring, they mostly rely on single modal data, such as text analysis or voice tone analysis to determine emotions, which is difficult to comprehensively and accurately understand the complex emotional state of the elderly. Because the emotional expression of the elderly may be implicit, multi-modal fusion (such as combining voice, expression, behavior, etc.) can more accurately understand their true emotions. In addition, when the emotional state of the elderly user is abnormal, the existing interaction and feedback mechanism is not perfect. When the emotional state of the elderly is detected to be abnormal, it is often difficult to timely and effectively convey key information to the guardian, and it is also difficult to provide personalized and targeted voice interaction according to the specific emotional state of the elderly, which cannot truly meet the emotional care and psychological comfort needs of the elderly.

[0004] Therefore, there is an urgent need for an elderly user emotional state interaction method based on an intelligent companion large model, which can collect daily data information of the elderly user in real time and comprehensively, accurately determine the emotional state through multi-modal emotion fusion, timely feedback to the guardian and effectively perform voice interaction when the emotional state is abnormal, so as to improve the quality and efficiency of emotional care for the elderly user. SUMMARY

[0005] The application provides an elderly user emotional state interaction method and system based on an intelligent companion large model, which solves the technical problems in the prior art, and the technical solutions adopted are as follows: The elderly user emotional state interaction method based on the intelligent companion large model comprises: Real-time collection of daily data information of the elderly user, and data processing of the daily data information, to generate a standardized time series data set corresponding to the daily data information; Fusion of the daily data information of the elderly user by multi-modal emotion fusion to obtain a fused emotion parameter; Determination of whether the elderly user has emotional abnormalities by using the emotion parameter of the elderly user; When the elderly user has emotional abnormalities, send a prompt message to the guardian, and perform voice interaction with the elderly user according to the emotional state of the elderly user.

[0006] Further, real-time collection of daily data information of the elderly user, and data processing of the daily data information, to generate a standardized time series data set corresponding to the daily data information, comprising: Collecting physiological data of the elderly user by using a wearable device; wherein the physiological data includes heart rate data, step frequency data and body temperature data; Collecting voice data of the elderly user by using a voice collection device; Recognizing the voice content of the voice data of the elderly user to obtain dialogue text data and audio waveform data corresponding to the voice content of the elderly user; Retrieving physiological data, dialogue text data and audio waveform data of the elderly user to form multi-source data; Timestamp alignment of the multi-source data to establish a unified time axis t; Normalizing each type of data information contained in the multi-source data, and generating a standardized time series data set by using the normalized multi-source data.

[0007] Further, fusion of the daily data information of the elderly user by multi-modal emotion fusion to obtain a fused emotion parameter, comprising: Obtaining a voice emotion parameter by using voice data in the daily data information; Obtaining a physiological emotion parameter by using physiological data in the daily data information; Fusing the voice emotion parameter and the physiological emotion parameter to obtain a fused emotion parameter.

[0008] Further, obtaining a voice emotion parameter by using voice data in the daily data information, comprising: Retrieving dialogue text data and audio waveform data in the voice data from the standardized time series data set; The audio amplitude parameter in the audio waveform data; Comparing the audio amplitude parameter with a preset amplitude parameter threshold; retrieve a time period when the audio amplitude parameter exceeds the amplitude parameter threshold as a calibration time period; extract conversation text data corresponding to the calibration time period; use the trained recurrent neural network model to obtain an emotion category from the conversation text data, and compare the emotion category with an emotion category table stored in a database to obtain an emotion arousal coefficient corresponding to the emotion category; extract heart rate data and step frequency data corresponding to the calibration time period as calibration heart rate data and calibration step frequency data; use the calibration heart rate data and the calibration step frequency data in combination with the emotion arousal coefficient to obtain a speech emotion parameter.

[0009] Further, using the calibration heart rate data and the calibration step frequency data in combination with the emotion arousal coefficient to obtain a speech emotion parameter includes: using the calibration heart rate data and the calibration step frequency data to obtain a calibration coefficient corresponding to each calibration time period; using the calibration coefficient corresponding to each calibration time period in combination with the emotion arousal coefficient corresponding to the emotion category appearing in each calibration time period to obtain a speech emotion parameter.

[0010] Further, using the physiological data in the daily data information to obtain a physiological emotion parameter includes: retrieve heart rate data, step frequency data, and body temperature data from the physiological data in the standardized time series data set; use the heart rate data, step frequency data, and body temperature data in the physiological data to obtain a heart rate data standard deviation, a step frequency data standard deviation, and a body temperature data standard deviation; use the heart rate data standard deviation, the step frequency data standard deviation, and the body temperature data standard deviation in combination with the corresponding heart rate data average value, step frequency data average value, and body temperature data average value to obtain a first heart rate data coefficient, a first step frequency data coefficient, and a first body temperature data coefficient corresponding to the heart rate data, the step frequency data, and the body temperature data; retrieve heart rate data, step frequency data, and body temperature data corresponding to all calibration time periods from the standardized time series data set to obtain a heart rate data standard deviation, a step frequency data standard deviation, and a body temperature data standard deviation corresponding to all calibration time periods as a whole; use the heart rate data standard deviation, the step frequency data standard deviation, and the body temperature data standard deviation corresponding to all calibration time periods as a whole in combination with the minimum value of the heart rate data, the minimum value of the step frequency data, and the minimum value of the body temperature data corresponding to all calibration time periods as a whole to obtain a second heart rate data coefficient, a second step frequency data coefficient, and a second body temperature data coefficient corresponding to the heart rate data, the step frequency data, and the body temperature data; The physiological emotional parameter is obtained by using the first heart rate data coefficient, the first step frequency data coefficient and the first body temperature data coefficient in combination with a second heart rate data coefficient, a second step frequency data coefficient and a second body temperature data coefficient.

[0011] Further, the voice emotional parameter and the physiological emotional parameter are fused to obtain a fused emotional parameter, including: The voice emotional parameter and the physiological emotional parameter are called. The voice emotional parameter and the physiological emotional parameter are dynamically weighted to obtain a dynamic weight value corresponding to the voice emotional parameter and the physiological emotional parameter. The voice emotional parameter and the physiological emotional parameter are fused by using the dynamic weight value corresponding to the voice emotional parameter and the physiological emotional parameter.

[0012] Further, the voice emotional parameter and the physiological emotional parameter are dynamically weighted to obtain a dynamic weight value corresponding to the voice emotional parameter and the physiological emotional parameter, including: The historical reliability coefficients of the wearable device and the voice collection device are extracted. The historical reliability coefficients of the wearable device and the voice collection device are used. The emotional excitement coefficient corresponding to each calibration time period is called. The standard deviation of the emotional excitement coefficient corresponding to all calibration time periods is obtained by using the emotional excitement coefficient corresponding to the calibration time period. The dynamic weight value corresponding to the voice emotional parameter and the physiological emotional parameter is obtained by using the standard deviation of the emotional excitement coefficient corresponding to the calibration time period in combination with the historical reliability coefficients of the wearable device and the voice collection device.

[0013] Further, whether the elderly user has emotional abnormalities is determined by using the emotional parameter of the elderly user, including: The emotional parameter of the elderly user is compared with a preset emotional parameter threshold. When the emotional parameter of the elderly user exceeds the preset emotional parameter threshold, it is determined that the elderly user has emotional abnormalities.

[0014] The elderly user emotional state interaction system based on the intelligent accompanying large model, the elderly user emotional state interaction system includes: The data acquisition module is used for real-time acquisition of daily data information of the elderly user, and the daily data information is processed to generate a standardized time series data set corresponding to the daily data information. The emotional abnormality determination module is used for fusing the daily data information of the elderly user by using multi-modal emotional fusion to obtain a fused emotional parameter. an emotion abnormality determination module configured to determine whether the elderly user has an emotion abnormality based on the emotion parameter of the elderly user; a voice interaction module configured to send a prompt message to a guardian when the elderly user has an emotion abnormality, and to perform voice interaction with the elderly user according to the emotion state of the elderly user.

[0015] The present application has the following advantages: The emotion state interaction method and system for elderly users based on the intelligent companion large model provided by the present application adopts a multi-modal data fusion manner, overcomes the limitations of single-modal monitoring, comprehensively analyzes the emotions of elderly users from multiple dimensions, can more accurately capture subtle and complex emotional changes, improves the accuracy of emotion monitoring, and timely discovers potential emotional problems of the elderly. Real-time data acquisition and processing can obtain the emotion state information of the elderly user at the first time, quickly respond to emotional abnormality, avoid the deterioration of emotional problems due to monitoring lag, and strive for valuable time for timely intervention and help. The prompt message is sent to the guardian in a timely manner, so that the guardian can remotely and timely understand the emotional state of the elderly user, facilitate the guardian to make corresponding decisions and actions, enhance the timeliness and effectiveness of emotional care for the elderly, and provide more comprehensive care for the elderly. According to the specific emotion state of the elderly user, personalized voice interaction is performed, the communication is more in line with the current mental state of the elderly, more easily causes emotional resonance, gives the elderly more considerate emotional comfort, improves the experience and satisfaction of the elderly user using the intelligent companion product, to a certain extent, relieves the loneliness of the elderly, and promotes the mental health of the elderly. BRIEF DESCRIPTION OF DRAWINGS

[0016] Figure 1 a flowchart of the method of the present application; Figure 2 a system block diagram of the system of the present application. DETAILED DESCRIPTION

[0017] The preferred embodiments of the present application will be described below with reference to the accompanying drawings, and it should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application, and are not used to limit the present application.

[0018] The present application provides an emotion state interaction method for elderly users based on an intelligent companion large model, as shown in Figure 1 The emotion state interaction method for elderly users based on the intelligent companion large model includes the following steps: S1, real-time collection of daily data information of an elderly user, and data processing of the daily data information to generate a standardized time series data set corresponding to the daily data information; S2, fusion of the daily data information of the elderly user by using a multi-modal emotion fusion to obtain a fused emotion parameter; S3, determining whether the elderly user has emotional abnormalities based on the emotional parameters of the elderly user; S4, when the elderly user has emotional abnormalities, sending a prompt message to the guardian, and performing voice interaction with the elderly user according to the emotional state of the elderly user.

[0019] The working principle of the above technical solution is that the daily data information of the elderly user is collected in real time by various sensors (such as a microphone to collect voice data, a camera to collect facial expression data, a motion sensor to collect behavior data, etc.). Then, the collected raw data is processed to remove noise and invalid data, and arranged in chronological order to generate a standardized time series dataset, providing a standardized and ordered data basis for subsequent analysis. Using multi-modal emotion fusion technology, the voice, expression, behavior, and other different modal data in the standardized time series dataset are integrated to obtain emotional parameters. According to the obtained emotional parameters, a comparison analysis is performed with the pre-set normal emotional parameter range or threshold value. If the emotional parameters exceed the normal range, such as excessive negativity, anxiety, agitation, and other abnormal conditions, it is determined that the elderly user has emotional abnormalities. These thresholds and ranges are determined based on statistical analysis of a large amount of emotional data of elderly users and professional knowledge in the field of psychology. Once it is determined that the elderly user has emotional abnormalities, the system immediately sends a prompt message to the guardian, which can include a description of the current emotional state of the elderly user, related factors that may cause emotional abnormalities, etc. At the same time, according to the specific emotional state of the elderly user, the intelligent companion model calls the corresponding voice interaction strategy and speech template to communicate with the elderly user. For example, for a depressed elderly person, comforting, encouraging, and positive guiding words are given; for an anxious elderly person, content to calm the emotions and answer doubts is provided to alleviate their negative emotions.

[0020] The effect of the above technical solution is that the multi-modal data fusion method overcomes the limitations of single modal monitoring, comprehensively analyzes the emotions of the elderly user from multiple dimensions, can more accurately capture subtle and complex emotional changes, improves the accuracy of emotional monitoring, and timely discovers potential emotional problems of the elderly. Real-time data collection and processing can obtain the emotional state information of the elderly user at the first time, quickly respond to emotional abnormality, avoid the deterioration of emotional problems due to monitoring lag, and strive for valuable time for timely intervention and assistance. Sending prompt messages to guardians in a timely manner enables guardians to remotely and timely understand the emotional state of the elderly user, facilitating guardians to make appropriate decisions and actions, enhancing the timeliness and effectiveness of emotional care for the elderly, and providing more comprehensive care for the elderly. According to the specific emotional state of the elderly user, personalized voice interaction is performed, making the communication more in line with the current mental state of the elderly, more easily causing emotional resonance, giving the elderly more considerate emotional comfort, improving the experience and satisfaction of the elderly user using the intelligent companion product, to some extent, alleviating the sense of loneliness of the elderly, and promoting their mental health.

[0021] In one embodiment of the present application, daily data information of an elderly user is collected in real time, and the daily data information is processed to generate a standardized time series data set corresponding to the daily data information, comprising: S101, collecting physiological data of an elderly user by using a wearable device; wherein the physiological data includes heart rate data, step frequency data and body temperature data; S102, collecting voice data of the elderly user by using a voice collection device; S103, identifying the voice content of the voice data of the elderly user to obtain conversation text data and audio waveform data corresponding to the voice content of the elderly user; S104, calling the physiological data, conversation text data and audio waveform data of the elderly user to form multi-source data; S105, time stamp alignment of the multi-source data to establish a unified time axis t; S106, normalizing each type of data information contained in the multi-source data, and generating a standardized time series data set by using the normalized multi-source data.

[0022] The structure of the standardized time series data set is as follows: Wherein, X(t) represents the standardized time series data set; H represents the normalized heart rate data; S represents the normalized step frequency data; T represents the normalized body temperature data; A represents the normalized audio waveform data; and D represents the conversation text data.

[0023] The working principle of the above technical solution is as follows: the physiological condition of the elderly user is monitored in real time by wearable devices (such as smart bands, smart watches, etc.), and heart rate data (reflecting the frequency of heartbeats, which can reflect the stress state of the body, etc.), step frequency data (which can assist in determining the amount of exercise and physical activity) and body temperature data (which reflect the basic physiological function state of the body) are continuously collected. These devices use built-in sensors (such as heart rate sensors, accelerometers, temperature sensors, etc.) to collect data. Voice acquisition devices (such as microphones, smart speakers, etc.) are used to record the voice information of the elderly user. These devices convert sound signals into electrical signals to achieve preliminary acquisition of voice data. Speech recognition technology is used on the collected voice data. On the one hand, the voice content is converted into dialogue text data, which facilitates understanding of the information expressed by the elderly user from the semantic level; on the other hand, audio waveform data is obtained, which is used to analyze the acoustic characteristics of the voice, such as tone, speed, volume, etc. These characteristics can reflect the emotional state of the elderly user when speaking. The physiological data, dialogue text data and audio waveform data of the elderly user collected are combined together to form multi-source data. These different types of data describe the state of the elderly user from multiple dimensions, providing comprehensive information for subsequent comprehensive analysis. Each data point in the multi-source data is labeled with a timestamp, and then different types of data are aligned on the time axis according to the timestamp to establish a unified time axis. This ensures that data from different sources has consistency in the time dimension, facilitating subsequent comprehensive analysis based on time series, such as analyzing the correlation between physiological state and voice expression at a certain specific time. For each type of data in the multi-source data, a normalization method (such as min-max normalization) is used to map the data to a specific interval (such as [0, 1]). Through normalization processing, the differences in dimension and value range of different types of data are eliminated, making different data comparable. Finally, the normalized multi-source data is arranged in chronological order to form a standardized time series data set, providing a standardized data basis for subsequent applications such as sentiment analysis and health monitoring based on this data set.

[0024] The effect of the above technical solution is that data is collected from two important dimensions of physiology and voice, covering physiological indicators such as heart rate, step frequency, body temperature, and information such as voice content and acoustic characteristics, comprehensively reflecting the physical and language expression conditions of the elderly user, and providing rich data support for in-depth understanding of the state of the elderly user. The uniform time axis is established through the time stamp alignment, which guarantees the synchronization and correlation of different types of data in time, so that the subsequent analysis based on time series is more accurate and reliable, and the change rule and mutual relationship of different dimension data at the same time or different time can be mined. The normalization processing eliminates the dimension and value range difference of the data, so that all kinds of data are in the same comparable scale, which is convenient for using unified algorithm and model to process and analyze the data, improves the efficiency and accuracy of data analysis, and is beneficial to subsequent more accurate emotional state judgment, health risk assessment and other applications using standardized time series data set. The generated standardized time series data set structure is clear and standard, which provides a high-quality data basis for various elderly user care applications (such as emotion analysis application, health monitoring and early warning application) developed based on the data set, facilitates the calling and implementation of subsequent algorithms and models, and improves the performance and reliability of related applications.

[0025] In an embodiment of the present application, the daily data information of the elderly user is fused by using multi-modal emotion fusion to obtain the fused emotional parameters, including: S201, obtaining voice emotional parameters from voice data in the daily data information; S202, obtaining physiological emotional parameters from physiological data in the daily data information; S203, fusing the voice emotional parameters and the physiological emotional parameters to obtain the fused emotional parameters.

[0026] The working principle of the above technical solution is that for the voice data in the daily data information, first, the acoustic features and semantic content of the voice are analyzed from two aspects. In terms of acoustic features, the pitch, speed, volume, fundamental frequency and other information of the voice are extracted, for example, a rising tone and fast speed may reflect an excited emotion, while a slow and deep tone may be related to a low emotion; through audio signal processing technology, these acoustic features are converted into quantifiable parameters. In terms of semantic content, natural language processing technology is used to analyze the emotional content of the dialogue text data, identify the emotional tendency expressed in the text, such as positive, negative, neutral, etc., and convert the semantic emotional information into corresponding parameters. Finally, the acoustic feature parameters and semantic content parameters are comprehensively analyzed, and a specific algorithm (such as weighted summation, machine learning model, etc.) is used to calculate the voice emotional parameters that can fully reflect the emotions expressed by the voice.

[0027] For physiological data (heart rate data, step frequency data, and body temperature data) in daily data information, the relationship between physiological indicators and emotional state is analyzed based on the correlation knowledge of physiology and psychology. For example, sudden acceleration of heart rate may be related to emotions such as tension, anxiety, etc.; a significant decrease in step frequency may indicate a depressed mood and lack of energy; and slight changes in body temperature may also reflect the body's response to stress emotions to some extent. By establishing a mapping model of physiological indicators and emotional state (such as a regression model, a neural network model, etc. trained based on a large amount of sample data), physiological data is converted into physiological emotional parameters that can represent emotional state.

[0028] The acquired voice emotional parameters and physiological emotional parameters are integrated using multi-modal fusion technology. Based on the complementarity and correlation between different modal data, multi-modal fusion technology uses weighted fusion, model fusion (such as deep neural network fusion model), etc. to comprehensively calculate voice emotional parameters and physiological emotional parameters. By setting appropriate weights or training the fusion model, the two parameters complement each other, eliminating the limitations of single modal data, so as to obtain the fused emotional parameters, which more comprehensively and accurately reflect the real emotional state of the elderly user.

[0029] The above technical solutions have the following effects: combining voice and physiological data for emotion analysis makes up for the shortcomings of single modal data in emotion recognition. Voice data reflects emotions from the aspects of language expression and sound characteristics, and physiological data provides emotional clues from the perspective of physical physiological response. The two modal data complement and supplement each other, which can more accurately capture the emotional changes of the elderly user, improve the accuracy of emotion recognition, and reduce misjudgment and omission. Voice emotional parameters and physiological emotional parameters quantify emotions from different angles, and the fused emotional parameters cover information in multiple dimensions such as language expression, sound characteristics, and physical physiological response, which can more comprehensively and stereoscopically reflect the emotional state of the elderly user, avoiding the bias in understanding emotions caused by the one-sidedness of a single information source. In practical applications, single modal data may be affected by environmental factors (such as voice data affected by noise interference, physiological data affected by device errors) or individual differences (such as unclear voice expression of the elderly, different physiological indicator base values), resulting in inaccurate data. Multi-modal emotion fusion integrates multiple modal data, and when one modal data is abnormal, other modal data can provide complementary information, reducing the dependence on single data and enhancing the robustness and stability of the emotion recognition system. Accurate and comprehensive emotional parameters provide a reliable basis for subsequent emotional care for the elderly. When the system judges that the emotional state of the elderly user is abnormal based on the fused emotional parameters, it can more accurately understand the essence of the emotional problem, thereby providing more targeted voice interaction and emotional support, improving the quality of emotional care for the elderly, and enhancing the experience and satisfaction of the elderly in using the intelligent companion system.

[0030] In one embodiment of the present application, voice emotion parameters are obtained from voice data in the daily data information, including: S2011, dialogue text data and audio waveform data in the voice data are called from the standardized time-series data set; S2012, audio amplitude parameters in the audio waveform data are obtained; S2013, the audio amplitude parameters are compared with preset amplitude parameter thresholds; S2014, a time period in which the audio amplitude parameters exceed the amplitude parameter thresholds is called as a calibration time period; S2015, dialogue text data corresponding to the calibration time period is extracted; S2016, emotion categories are obtained from the dialogue text data by using a trained recurrent neural network model, and the emotion categories are compared with emotion category tables stored in a database to obtain emotion arousal coefficients corresponding to the emotion categories; the emotion arousal coefficients are set according to experience and actual application, the value range of the emotion arousal coefficients is (0, 1], and the larger the value is, the more intense the emotion is; S2017, heart rate data and step frequency data corresponding to the calibration time period are extracted as calibration heart rate data and calibration step frequency data; S2018, voice emotion parameters are obtained from the calibration heart rate data and the calibration step frequency data in combination with the emotion arousal coefficients.

[0031] The working principle of the technical solution is as follows: dialogue text data and audio waveform data in the speech data are obtained from the standardized time series data set. The dialogue text data contains the specific content of the speech of the elderly user, and the audio waveform data records the acoustic characteristics of the speech. The audio amplitude parameters in the audio waveform data are extracted, and the amplitude size can reflect the information such as the volume of the speech. The amplitude parameters are compared with the preset amplitude parameter threshold value, and the time period in which the audio amplitude parameter exceeds the threshold value, i.e. the calibration time period, is found. The amplitude exceeding the threshold value often means that the volume of the speech is larger and the emotion may be more excited in these time periods. The dialogue text data corresponding to the calibration time period is extracted, and the trained recurrent neural network model is used for alignment and processing. The recurrent neural network model has a memory function and can process sequence data. Through the analysis of the text semantics, the emotion categories contained therein, such as anger, joy, sadness, etc. are identified. The identified emotion categories are compared with the emotion category table stored in the database to obtain the emotion excitement coefficient corresponding to the emotion category. The emotion excitement coefficient is set according to experience and actual application situation, and is used to quantify the intensity of emotion. The heart rate data and step frequency data corresponding to the calibration time period are extracted as calibration heart rate data and calibration step frequency data. Heart rate and step frequency can reflect the stress state and emotional state of the body to a certain extent, and the heart rate and step frequency often increase when the emotion is excited. The calibration heart rate data, calibration step frequency data and emotion excitement coefficient are combined, and the emotional information in the speech text and the changes of physiological indicators are comprehensively considered, and the speech emotion parameter is calculated through a specific algorithm (such as weighted summation, etc.). The parameter can more comprehensively reflect the emotional state of the elderly user in speech communication.

[0032] The above technical scheme has the effects that: the acoustic characteristics (audio amplitude) of the voice, the semantic information (dialog text emotion analysis), and the physiological indicators (heart rate and step frequency) are comprehensively considered to analyze the emotions of the elderly user in multiple dimensions. Compared with single voice analysis or physiological indicator monitoring, this comprehensive method can more accurately capture the emotional changes of the elderly user, reduce misjudgments caused by single factor fluctuations, and improve the accuracy of emotion recognition. By setting the amplitude parameter threshold, focusing on the calibration time period with large audio amplitude, and analyzing these possible emotional moments, the analysis efficiency is improved, and the emotional fluctuations of the elderly user are more accurately focused on, so that potential emotional problems can be found in time. The emotion excitement coefficient is introduced to convert different emotion categories into specific numerical values, which facilitates the quantification and comparison of the intensity of emotions. This helps to more intuitively understand the emotional state of the elderly user and provides a more explicit basis for subsequent emotional care and intervention. The voice emotion parameters are calculated in combination with the physiological indicators (heart rate and step frequency) of the elderly user, considering the influence of individual differences on emotional expression. Different elderly people may have different physiological indicator changes when they are emotionally excited, and by combining individual physiological data, more personalized emotion evaluation can be achieved to provide services that better meet the needs of the elderly user.

[0033] In an embodiment of the present application, the voice emotion parameters are obtained by using the calibration heart rate data and calibration step frequency data in combination with the emotion excitement coefficient, which includes: Step 1: Obtain the calibration coefficient corresponding to each calibration time period by using the calibration heart rate data and calibration step frequency data. The calibration coefficient corresponding to each calibration time period is obtained by the following formula: Wherein, B represents the calibration coefficient; P x represents the calibration heart rate data corresponding to each calibration time period; P xc represents a preset heart rate reference value; f represents the calibration step frequency data corresponding to each calibration time period; f c represents a preset step frequency reference value; represents the multiple relationship of the current calibration heart rate data relative to the preset heart rate reference value, reflecting the degree of deviation of the heart rate from the normal reference level in the time period. The larger the value, the higher the current heart rate relative to the reference value. The step frequency reflects the proportional relationship of the step frequency relative to the reference step frequency, and the exponential function exp(−x) (here x= ) increases with ​decreases with the increase of the step frequency. It reflects the influence of the step frequency on the result, the closer the part of the result is to 1, the more the step frequency change contributes to the emotion-related physiological state, and the step frequency acceleration may reflect the body movement change of the user due to the emotion, and the step frequency is converted into a coefficient part that can be analyzed cooperatively with the heart rate through this calculation. The product of the relative value of the heart rate and the correlation value of the step frequency is squared. The squaring operation limits the value range of the result on the one hand, and comprehensively considers the heart rate and the step frequency factors on the other hand, so that the calibration coefficient more reasonably reflects the joint action of the two, and obtains the final calibration coefficient B. The heart rate normalized result and the step frequency mapping result are nonlinearly integrated to obtain a calibration coefficient B corresponding to the emotion physiological basis under the joint action of the heart rate and the step frequency in the calibration time period, so that the influences of the heart rate and the step frequency are reasonably weighted and fused for subsequent emotion parameter calculation.

[0034] Step 2, using the calibration coefficient corresponding to each calibration time period and the emotion arousal coefficient corresponding to the emotion category appearing in each calibration time period to obtain the speech emotion parameter.

[0035] The speech emotion parameter is obtained by the following formula: Wherein, Y represents the speech emotion parameter; n represents the number of calibration time periods corresponding to; B i represents the calibration coefficient corresponding to the i-th calibration time period; Q i represents the maximum value of the emotion arousal coefficient corresponding to all emotion types appearing in the i-th calibration time period; Q bi represents the standard deviation of the emotion arousal coefficient corresponding to all emotion types appearing in the i-th calibration time period. The whole is used to adjust the size of the denominator to balance the influence of the discreteness of the emotion arousal degree on the result. The physiological indicators (heart rate, step frequency) and the emotion arousal degree are comprehensively considered, and the multiplication of the two is the cooperative calculation of the physiology and the emotion intensity, which reflects the correlation between the physiological state (heart rate, step frequency reflection) and the strongest emotion intensity in the calibration time period, and combines the physiological basis and the emotion intensity, such as the multiplication of the physiological abnormality (heart rate, step frequency change) and the most prominent emotion intensity, which highlights the influence degree of the emotion and the physiology cooperation in the time period. Through the operation of the numerator and the denominator, the calibration coefficient, the maximum emotion arousal coefficient, and the discreteness of the emotion arousal coefficient are combined to comprehensively consider the influence of multiple factors on the speech emotion parameter.

[0036] The working principle of the above technical solution is that: the core of this step is to measure the deviation of the physiological state of the elderly user in the time period relative to the normal state by comparing the calibration heart rate data and the calibration step frequency data in each calibration time period with the preset reference value, and then obtaining the calibration coefficient. This step comprehensively considers the calibration coefficient of each calibration time period and the emotion excitement coefficient corresponding to the emotion type occurring in the time period, thereby obtaining the voice emotion parameter.

[0037] The effect of the above technical solution is: this technical solution combines physiological indicators (heart rate and step frequency) and emotion excitement coefficients corresponding to emotion categories to calculate voice emotion parameters, which can more comprehensively and accurately reflect the emotional state of elderly users compared to single-dimensional evaluation methods. The change of physiological indicators is the objective manifestation of emotional excitement, and the emotion excitement coefficient quantifies the intensity of emotion from the semantic level, and the combination of the two makes the emotion evaluation more accurate. The introduction of the standard deviation of the emotion excitement coefficient considers the stability of emotion in each calibration time period. In some cases, even if the maximum value of the emotion excitement coefficient is high, if the standard deviation is also large, it means that the emotional fluctuation is intense, and there may be emotional instability. By comprehensively considering the maximum value and the standard deviation, the emotional changes can be captured more carefully, and the accuracy of emotion evaluation can be improved. The actual calibration heart rate data and calibration step frequency data of each elderly user are used, and are compared with the preset reference value, considering the differences in physiological characteristics between individuals. The change amplitude of heart rate and step frequency may be different when different elderly users are emotionally excited, and this method can evaluate emotions according to the actual situation of individuals to realize personalized emotion monitoring. The obtained voice emotion parameter is a specific numerical value, which is convenient for quantitative analysis and comparison of the emotional state of elderly users. Caregivers or systems can discover emotional abnormalities of elderly users in time according to the parameter, and take appropriate measures such as providing psychological comfort or adjusting the accompanying strategy, improving the practicality of emotion monitoring and intervention. In obtaining the calibration coefficient, two physiological indicators, heart rate and step frequency, are considered comprehensively, and the two are fused through a specific functional relationship, which is more comprehensive than a single indicator in reflecting physiological state. In obtaining the voice emotion parameter, not only the calibration coefficient is combined, but also the maximum and dispersion degree of the emotion excitement coefficient are considered, and multi-dimensional information fusion makes the result more accurately describe the emotion in the voice. The processing method of each factor in the above two formulas of this embodiment avoids that a single factor fluctuation has too great an impact on the result, and enhances the stability of the system under different emotional fluctuations. The parameters (such as heart rate reference value, step frequency reference value, etc.) in the formula can be adjusted according to different application scenarios or individual differences, and can adapt to changes in various emotion types and different emotion excitement degrees, so that the system can better obtain the calibration coefficient and voice emotion parameter in different users and different situations, improving the adaptability.

[0038] Meanwhile, the existing multi-dimensional data means often simply and juxtapose physiological data (heart rate, step frequency) and emotional data (text, voice) for analysis, and only make "data superposition". The above technical solution of the embodiment constructs a dynamic coupling relationship through a formula, and in the calibration coefficient formula, the heart rate and the step frequency are not independently involved in the calculation, but are associated through the ratio with the reference value and the exponential function, to simulate the dynamic response of the physiological indicators with the emotional fluctuation (such as the nonlinear correlation between the heart rate acceleration and the step frequency change rate and the emotional excitement degree). In the voice emotional parameter formula, the physiological calibration coefficient and the emotional excitement coefficient are further integrated, so that the "physiological response" and the "emotional intensity / fluctuation" form a closed loop association, which is not a simple addition of physiological data + emotional data, but a dynamic interaction of physiological change driving emotional calculation and emotional characteristics feeding back physiological weight, which is more in line with the linkage mechanism of physiology and psychology when the actual emotion is generated. The existing multi-dimensional data-driven interaction often stays in a one-way process of triggering a fixed response through emotional recognition. The above technical solution of the embodiment triggers the precise early warning through the voice emotional parameter (Y) calculated by the formula, which is not a fuzzy emotional label, but a quantitative emotional abnormality degree value. The system can push the warning level degree to the guardian more accurately based on the numerical interval of Y (such as threshold determination "intervention is needed"), to avoid non-discriminatory prompts. Meanwhile, the calculation process of the calibration coefficient and the emotional excitement coefficient realizes the disassembly of the "emotional abnormality cause" (physiological fluctuation, emotional intensity / stability).

[0039] In one embodiment of the present application, physiological emotional parameters are obtained from physiological data in the daily data information, including: S2021, heart rate data, step frequency data and body temperature data in the physiological data are called from the standardized time series data set; S2022, heart rate data standard deviation, step frequency data standard deviation and body temperature data standard deviation are obtained from the heart rate data, the step frequency data and the body temperature data in the physiological data; S2023, first heart rate data coefficient, first step frequency data coefficient and first body temperature data coefficient corresponding to the heart rate data, the step frequency data and the body temperature data are obtained by ratio processing of the heart rate data standard deviation, the step frequency data standard deviation and the body temperature data standard deviation and the corresponding heart rate data average value, step frequency data average value and body temperature data average value; S2024, heart rate data standard deviation, step frequency data standard deviation and body temperature data standard deviation corresponding to all calibration time periods are obtained from the heart rate data, the step frequency data and the body temperature data corresponding to all calibration time periods in the standardized time series data set; S2025, utilize the standard deviation of heart rate data, the standard deviation of step frequency data and the standard deviation of body temperature data corresponding to all calibration time periods as a whole to perform ratio processing with the minimum value of heart rate data, the minimum value of step frequency data and the minimum value of body temperature data corresponding to all calibration time periods as a whole, to obtain second heart rate data coefficient, second step frequency data coefficient and second body temperature data coefficient corresponding to the heart rate data, the step frequency data and the body temperature data; S2026, utilize the first heart rate data coefficient, the first step frequency data coefficient and the first body temperature data coefficient to combine the second heart rate data coefficient, the second step frequency data coefficient and the second body temperature data coefficient to obtain the physiological emotional parameter.

[0040] Wherein, the physiological emotional parameter is obtained by the following formula: Wherein, S represents the physiological emotional parameter; P x01 and P x02 respectively represent the first heart rate data coefficient and the second heart rate data coefficient; f 01 and f 02 respectively represent the first step frequency data coefficient and the second step frequency data coefficient; T 01 and T 02 respectively represent the first body temperature data coefficient and the second body temperature data coefficient. First, multiply the calculation results of heart rate, step frequency and body temperature data respectively, which is to comprehensively consider the common influence of the change characteristics of the three physiological data on the physiological emotional parameter. The cubic root operation is to adjust the numerical range of the product result, on the other hand, to maintain the relative balance of the three physiological data in the calculation, to avoid the change of a certain data to excessively dominate the final result, so as to obtain the physiological emotional parameter S which comprehensively reflects the change of physiological data.

[0041] The working principle of the technical solution is as follows: the heart rate, step frequency and body temperature data in the physiological data are obtained from the standardized time series data set, which are stored in the data set after pre-collection and processing, and provide a basis for subsequent analysis. First, the standard deviation of the heart rate, step frequency and body temperature data is calculated, which measures the dispersion degree of the data and reflects the fluctuation of the data. Then, the standard deviation of each data is processed by ratio with the corresponding average value to obtain the first heart rate data coefficient, the first step frequency data coefficient and the first body temperature data coefficient. These coefficients can reflect the dispersion degree of the data relative to the average value, and are used to preliminarily describe the change characteristics of the physiological data. For all calibration time periods (determined in the previous voice data analysis), the standard deviation of the overall corresponding heart rate, step frequency and body temperature data is calculated again. The standard deviation is processed by ratio with the minimum value of the overall corresponding heart rate, step frequency and body temperature data in the calibration time period to obtain the second heart rate data coefficient, the second step frequency data coefficient and the second body temperature data coefficient. This group of coefficients further describes the change of the physiological data in a specific time period from another angle, i.e., the dispersion degree relative to the minimum value. The two groups of coefficients obtained above are combined to calculate the physiological emotion parameter through a specific formula. This parameter integrates the change characteristics of the physiological data in the whole and specific time period to reflect the emotional state of the elderly user, because the change of the physiological data is often related to the emotional state.

[0042] The above technical solution has the following effects: through multi-dimensional analysis and calculation of physiological data, the physiological data is converted into a quantitative physiological emotion parameter, so that the emotional state of the elderly user can be presented in the form of a numerical value, facilitating subsequent analysis and judgment, and providing an objective and measurable index for the determination of emotional abnormalities. Not only the overall change of heart rate, step frequency and body temperature data is considered (through the first group of coefficients), but also the calibration time period that may be related to emotion is analyzed separately (through the second group of coefficients), and the physiological emotion parameter is obtained by comprehensively considering the two aspects, which more comprehensively captures the emotional information contained in the physiological data, and improves the accuracy and reliability of the emotional state evaluation of the elderly user. The accurate physiological emotion parameter can help the system to more accurately judge the emotional state of the elderly user, and when combined with other information such as voice emotion parameter, it can more comprehensively understand the emotional state of the elderly user, thereby providing a strong basis for subsequent decisions such as whether to send a prompt information to the guardian and how to interact with the elderly user through voice, which helps to discover the emotional problems of the elderly user in time and take corresponding measures. At the same time, the data is retrieved from the standardized time series data set, which ensures the consistency and standardization of the data. The standardized data can reduce the interference caused by data format, unit, dimension and other problems, so that the calculation result is more stable and reliable. In different time points or different individuals, the data can be processed under the unified standard, avoiding the result fluctuation caused by data difference. By calculating the first heart rate data coefficient, the first step frequency data coefficient, the first body temperature data coefficient, the second heart rate data coefficient, the second step frequency data coefficient and the second body temperature data coefficient, and combining them to calculate the physiological emotion parameter, the influence of different data characteristics and different time period data can be balanced. This comprehensive calculation method can reduce the influence of individual abnormal data on the final result, so that the physiological emotion parameter is more stable when facing data fluctuation.

[0043] Meanwhile, the scheme comprehensively considers heart rate, step frequency, and body temperature, three kinds of physiological data, each of which reflects the physiological state of the human body from different angles. The change of heart rate can directly reflect the activity intensity of the heart, the step frequency can reflect the movement state and rhythm of the body, and the body temperature is related to the physiological processes such as metabolism and endocrine of the body. By analyzing these three kinds of data at the same time, the changes of the physiological state of the human body can be more comprehensively captured, so as to more accurately infer the emotional state. For example, when emotional excitement, the heart rate may accelerate, the step frequency may become faster, and the body temperature may also slightly rise, and the comprehensive analysis of these data can more accurately judge the degree of emotional excitement. When calculating the first heart rate data coefficient, the first step frequency data coefficient, and the first body temperature data coefficient, the standard deviation and the average value of each physiological data are processed by ratio, which not only considers the overall level (average value) of the data, but also considers the dispersion degree (standard deviation) of the data. The standard deviation reflects the fluctuation of the data, and by combining it with the average value, the characteristics of the physiological data can be more accurately described. In this way, the relationship between physiological data and emotions can be more accurately captured, and the accuracy of physiological emotional parameter acquisition can be improved. In the scheme, not only the correlation coefficient of daily physiological data (first coefficient) is calculated, but also the data of the calibration time period is analyzed to obtain the second heart rate data coefficient, the second step frequency data coefficient, and the second body temperature data coefficient. The calibration time period may be a time period when some specific situations or events occur, and by analyzing the data of these time periods, the change characteristics of physiological data under specific situations can be more accurately captured. For example, when facing a stressful event, physiological data may change significantly in the calibration time period, and by calculating the coefficient of the calibration time period, the influence of such changes on emotions can be more accurately reflected. Combining the first coefficient and the second coefficient to obtain the physiological emotional parameter can comprehensively consider the physiological data characteristics in daily life and specific situations, further improving the accuracy of the physiological emotional parameter. This way can avoid the problem of incomplete information caused by only considering the data of a single time period, making the obtained physiological emotional parameter more accurately reflect the true emotional state of the individual. The data is retrieved from the standardized time series data set, which is usually recorded and organized in chronological order. By directly obtaining data from such a data set, the latest physiological data information can be quickly obtained, so as to timely reflect the current physiological state and emotional state of the individual. For example, in the application of real-time monitoring of emotional state, the latest physiological data can be obtained in time and the physiological emotional parameter can be calculated, which is of great significance for timely discovering emotional changes and taking corresponding measures. The standardized data set can also ensure the consistency and standardization of the data, facilitating rapid data processing and analysis. After obtaining the data, it can be directly calculated according to the predetermined algorithm, without complex data preprocessing and conversion, thereby improving the timeliness of physiological emotional parameter acquisition.By calculating the correlation coefficients of heart rate, step frequency and body temperature data of different users, the individual physiological characteristics of each user can be reflected. The physiological data of different users may differ in average value, standard deviation and variation law, etc. By calculating the individualized coefficients, the unique physiological characteristics of each user can be better adapted. Considering the data of the calibration period can better adapt to the physiological changes of the user in different situations. Different users may have different physiological responses when facing the same situation. By analyzing the data of the calibration period, the physiological change characteristics of each user in a specific situation can be more accurately captured, thereby improving the adaptability between the physiological emotional parameters and the user. The scheme can dynamically adjust the physiological emotional parameters according to the real-time changes of the user's physiological data. Over time and with changes in the user's state, physiological data will change continuously. By calculating and updating the correlation coefficients in real time, the physiological emotional parameters can timely reflect these changes. For example, when the user changes from a quiet state to a motion state, the physiological data such as heart rate, step frequency and body temperature will change significantly. The scheme can timely capture these changes and adjust the physiological emotional parameters, thereby better adapting to the dynamic changes of the user. This dynamic adaptation capability enables the scheme to provide accurate physiological emotional parameters in different user states and situations, improving the adaptability with the user. Whether in daily activities or in special situations, the user can rely on the scheme to obtain physiological emotional parameters that match their current state.

[0044] On the other hand, the different statistical feature changes of heart rate, step frequency and body temperature are comprehensively considered. By calculating the difference values of the two groups of coefficients of each physiological data respectively and combining them, the physiological state change information is comprehensively captured. Compared with single physiological data or single statistical dimension, the correlation between physiology and emotion can be more comprehensively reflected, and the comprehensiveness and integrity of the physiological emotional parameters are improved. The exponential function is used as the denominator to normalize the changes of each physiological data, avoiding the extreme influence of a large or small coefficient value on the final result. The cubic root operation also balances the role of the three physiological data, so that the result will not deviate too much due to the strong fluctuations of a certain physiological data, enhancing the stability of the physiological emotional parameter calculation result. Detailed consideration of the differences of each physiological data in different statistical dimensions can accurately capture the subtle changes of physiological data, which are closely related to emotional state. By reasonably calculating and combining these changes, the physiological emotional parameters obtained can more accurately reflect the true emotional state, improving the accuracy in emotional evaluation.

[0045] In an embodiment of the present application, the voice emotional parameters and the physiological emotional parameters are fused to obtain fused emotional parameters, comprising: S2031, retrieve voice emotional parameters and physiological emotional parameters; S2032, dynamically setting weights for the voice emotion parameter and the physiological emotion parameter to obtain dynamic weight values corresponding to the voice emotion parameter and the physiological emotion parameter; S2033, combining the voice emotion parameter and the physiological emotion parameter by using the dynamic weight values corresponding to the voice emotion parameter and the physiological emotion parameter to obtain a fused emotion parameter.

[0046] The fused emotion parameter is obtained by the following formula: wherein, K represents the fused emotion parameter; w 01 and w 02 respectively represent the dynamic weight values corresponding to the voice emotion parameter and the physiological emotion parameter; S represents the physiological emotion parameter; and Y represents the voice emotion parameter.

[0047] The working principle of the above technical solution is as follows: the voice emotion parameter Y and the physiological emotion parameter S are obtained from the results obtained by previous processing. The voice emotion parameter is obtained by analyzing voice data (such as dialogue text, audio waveform, etc.), which reflects the emotional information contained in the voice expression of the elderly user; the physiological emotion parameter is calculated based on physiological data (such as heart rate, step frequency, body temperature, etc.), which reflects the emotional state reflected by the physiological state of the elderly user. The dynamic weight values corresponding to the voice emotion parameter and the physiological emotion parameter are obtained by dynamically setting weights for the voice emotion parameter and the physiological emotion parameter. The setting of the dynamic weight considers various factors, such as the historical reliability coefficient of the device, the standard deviation of the emotional excitement coefficient, and the standard deviation of the two emotion parameters in the historical data. Through the comprehensive consideration of these factors, appropriate weights can be allocated to the voice emotion parameter and the physiological emotion parameter according to different situations to adapt to the importance of the two kinds of data to the final emotion judgment in different scenarios. The dynamic weight values are combined with the voice emotion parameter and the physiological emotion parameter to calculate the fused emotion parameter by the above formula. The formula weights the voice emotion parameter and the physiological emotion parameter according to their corresponding dynamic weights, thereby obtaining an emotion parameter that comprehensively considers the voice and physiological information, and more comprehensively reflects the real emotional state of the elderly user.

[0048] The effect of the above technical solution is that the information of two dimensions of voice and physiology is comprehensively used for emotion evaluation, overcoming the limitation of single modal data in emotion recognition. The voice data can reflect the emotion at the language expression level, and the physiological data can reflect the emotion at the body physiological reaction level. The combination of the two can confirm and supplement each other, more accurately capture the emotional changes of the elderly users, reduce the misjudgment and omission, and improve the accuracy of emotion evaluation. The setting of dynamic weight enables the system to flexibly adjust the importance of the voice emotion parameter and the physiological emotion parameter according to the actual situation. For example, when the voice collection device has high reliability and the emotion excitement coefficient fluctuates greatly, a higher weight can be allocated to the voice emotion parameter; and when the physiological data can better reflect the emotional changes, the weight of the physiological emotion parameter will be increased accordingly. This dynamic adjustment mechanism enables the system to better adapt to different scenes and data quality, and improves the adaptability and reliability of emotion evaluation. The fused emotion parameter integrates the information of voice and physiology, providing more comprehensive and rich emotion information for subsequent emotion analysis and processing. The caregivers or intelligent companion systems can better understand the emotional state of the elderly users based on the comprehensive parameter, so as to provide more targeted emotional care and support. By fusing multi-modal data and adopting dynamic weight setting, the dependence of the system on a single data source is reduced. When one kind of data source is abnormal (such as voice collection device failure, physiological data fluctuation anomaly), the information of the other kind of data source can still ensure the accuracy of emotion evaluation to a certain extent, enhancing the robustness and stability of the system.

[0049] In an embodiment of the present application, dynamic weight setting is performed on the voice emotion parameter and the physiological emotion parameter, and dynamic weight values corresponding to the voice emotion parameter and the physiological emotion parameter are obtained, including: Step 1, extract the historical reliability coefficient of the wearable device and the voice collection device; Wherein, the historical reliability coefficient is obtained by the following formula: Wherein, g represents the historical reliability coefficient; M s represents the number of correct data detection times of the wearable device and the voice collection device in a unit period; M z represents the total number of data detection times of the wearable device and the voice collection device in a unit period; represents the proportion of correct data detected by the device in the total detection data in a unit period, which directly reflects the correctness of the device detection data. The higher the ratio, the better the accuracy of the device detection data in the period. The correct data detection proportion is multiplied by , in order to map the proportion value to a new interval for subsequent processing by a sine function Here it acts as a scaling factor, which The value range (0 to 1) is mapped to 0 to , prepare for the input of the sine function. By transforming the sine function, the change of the reliability coefficient can be made relatively smooth, avoiding sudden changes. Using the sine function to calculate the historical reliability coefficient makes the change of the coefficient smooth. The introduction of a sine function as a reliability coefficient prevents drastic fluctuations in the reliability coefficient caused by small changes in the correct data detection ratio. Calculated using a sine function, the historical reliability coefficient changes relatively smoothly, facilitating a more stable system assessment of device reliability. Limiting the value range of the historical reliability coefficient to 0–1 normalizes device reliability. This ensures comparability across devices. Regardless of the total amount of detection data and the amount of correct detection data, this unified coefficient can be used to measure the reliability level. This facilitates the appropriate consideration of device reliability in subsequent system decisions, such as dynamic weighting. By considering the ratio of correct data detections to total data detections within a unit cycle, the device's detection accuracy over a given period of time can be accurately reflected. This ratio is a key indicator of device reliability. The historical reliability coefficient calculated based on this ratio provides the system with more accurate device reliability information, helping the system more appropriately assign weights when processing voice emotion parameters and physiological emotion parameters, thereby improving the accuracy of overall emotion assessment.

[0050] Step 2: Using historical reliability coefficients of the wearable device and the voice acquisition device; Step 3: retrieve the emotional excitement coefficient corresponding to each calibration time period; Step 4: using the emotional excitement coefficient corresponding to the calibration time period to obtain the standard deviation of the emotional excitement coefficient corresponding to all calibration time periods; Step 5: Use the standard deviation of the emotional excitement coefficient corresponding to the calibration time period in combination with the historical reliability coefficient of the wearable device and the voice acquisition device to obtain the dynamic weight values ​​corresponding to the voice emotion parameters and the physiological emotion parameters.

[0051] The dynamic weight values ​​corresponding to the voice emotion parameters and physiological emotion parameters are obtained by the following formula: Among them, w 01 and w 02 Respectively represent the dynamic weight values ​​corresponding to the voice emotion parameters and physiological emotion parameters; S b represents the standard deviation of the physiological and emotional parameters obtained from the historical data; Yb denotes the standard deviation of the speech emotion parameters obtained in the historical data; g 01 and g 02 denote the historical reliability coefficients of the wearable device and the speech collection device, respectively. Q b denotes the standard deviation of the emotion arousal coefficient corresponding to the calibration time period; wherein, and By numerator denominator operation, the influence of data dispersion degree (reflected by the sine function part) and emotion arousal coefficient (reflected by the denominator) is combined to obtain a proportional value that comprehensively reflects multiple factors, and then multiplied by the device historical reliability coefficient g 01 , g 02 , the final dynamic weight value corresponding to the speech emotion parameters and the physiological emotion parameters is obtained. The formula comprehensively considers the device historical reliability, the dispersion degree of physiological and speech emotion parameters, and the emotion arousal coefficient. Multiple factors jointly determine the dynamic weight value, which can more comprehensively evaluate the importance of speech emotion parameters and physiological emotion parameters under different conditions compared to considering only a single factor, making the weight distribution more reasonable. The use of nonlinear functions such as sine function and exponential function makes the calculation process of dynamic weight value relatively smooth. Avoiding the sudden change of weight caused by the slight change of a factor, enhancing the system stability. At the same time, the weight can be dynamically adjusted according to the change of device reliability and the dispersion degree of emotion parameters. Different devices have different reliabilities at different times, and the emotion parameter fluctuation is also different. This formula can flexibly adapt to these changes, reasonably distribute the weight, and ensure accurate evaluation of speech and physiological emotion parameters under various conditions, improving the adaptability of the system.

[0052] The working principle of the above technical solution is as follows: First, extract the historical reliability coefficients of the wearable device and the speech collection device. This coefficient reflects the reliability of the device in the historical data collection process. The higher the coefficient, the more accurate and reliable the data collected by the device. Retrieve the emotion arousal coefficients corresponding to each calibration time period. These coefficients are obtained in the previous speech data processing and reflect the intensity of emotion in each calibration time period. Calculate the standard deviation of the emotion arousal coefficients corresponding to all calibration time periods. This standard deviation measures the fluctuation of the emotion arousal coefficients between different calibration time periods. A larger standard deviation means that the emotion fluctuates more, and more attention may be needed to the accuracy and reliability of the data. Calculate the dynamic weight value corresponding to the speech emotion parameters and the physiological emotion parameters by combining the standard deviation of the physiological emotion parameters in the historical data, the standard deviation of the speech emotion parameters, and the historical reliability coefficients of the wearable device and the speech collection device. Here, the fluctuation of the two emotion parameters in the historical data and the reliability of the device are considered, and the weight is dynamically allocated according to these factors.

[0053] The technical scheme has the following effects: by considering the historical reliability coefficient of the device, the quality of data from different sources can be quantitatively evaluated. For data collected by devices with high reliability, a higher weight is given, so that accurate and reliable data can be more fully utilized in emotion evaluation, reducing the deviation of emotion evaluation caused by device errors, and improving the accuracy of overall emotion evaluation. The introduction of the standard deviation of the emotion excitement coefficient and the standard deviation of the emotion parameter in the historical data enables the weight to be dynamically adjusted according to the fluctuation of the emotion. When the emotion fluctuates greatly, the system can more flexibly allocate the weight to better capture the change in emotion and improve the sensitivity and response capability to abnormal emotional situations. The technical scheme comprehensively considers the voice emotion parameter and the physiological emotion parameter, and through dynamic weight setting, it can reasonably integrate the two kinds of data according to different situations, fully utilize the advantages of multi-source data, more comprehensively and accurately reflect the emotional state of the elderly user, and avoid the limitations of a single data source. Dynamic weight setting enables the system to adaptively adjust the weight according to the device reliability and the emotion fluctuation, reduces the influence of device failure, abnormal data and other factors on the emotion evaluation result, enhances the robustness and stability of the system, and ensures that accurate emotion evaluation can be provided in various situations. On the other hand, considering the historical reliability coefficient of the wearable device and the voice collection device, the data credibility can be evaluated according to the past performance of the device. The data collected by the device with high reliability has a more reasonable proportion in the weight calculation, reducing the parameter deviation caused by device errors, so that the comprehensive result of the voice emotion parameter and the physiological emotion parameter is closer to the real emotional state. At the same time, combined with the standard deviation of the emotion excitement coefficient, the degree of emotion fluctuation can be measured to further accurately depict the change in emotion and improve the accuracy of the final emotion evaluation. By adjusting the weight through the historical reliability coefficient, the influence of temporary failure or abnormal data of a device on the result is avoided. When the device reliability is low, the corresponding parameter weight is reduced, so that the overall emotion evaluation result is more stable. By using the standard deviation of the emotion excitement coefficient, the influence of emotion fluctuation on the weight can be smoothed to prevent abnormal weight calculation caused by short-term and violent emotion fluctuation, and the stability of the system is maintained. Different devices have different reliabilities, and the reliability of the same device may also change in different scenes. By dynamically adjusting the weight based on the historical reliability coefficient, the performance change of the device can be adapted. Combined with the standard deviation of the emotion excitement coefficient, the weight can be reasonably adjusted in different emotional scenes, such as emotional excitement, so that the system can work effectively in various devices and emotional situations.

[0054] In an embodiment of the present application, whether the elderly user has an emotional abnormality is determined by using the emotion parameters of the elderly user, comprising: S301, comparing the emotion parameters of the elderly user with a preset emotion parameter threshold; S302, when the emotion parameters of the elderly user exceed the preset emotion parameter threshold, it is determined that the elderly user has an emotional abnormality.

[0055] The working principle of the above technical solution is to compare the obtained emotional parameters of the elderly user (calculated through a series of previous steps, including voice emotional parameters and physiological emotional parameters after processing and fusion) with pre-set emotional parameter thresholds. These pre-set emotional parameter thresholds are determined based on a large amount of emotional data of elderly users and related psychological research and practical experience, representing the parameter range under normal emotional state. If the emotional parameters of the elderly user exceed the pre-set emotional parameter thresholds, it means that the emotional state of the elderly user has exceeded the normal range, and the system determines that the elderly user has emotional abnormalities accordingly.

[0056] The effect of the above technical solution is that by setting clear thresholds and comparing, it can quickly and intuitively determine whether the elderly user has emotional abnormalities. Once the emotional parameters exceed the thresholds, an alarm can be sent out or appropriate measures can be taken in time, which helps to discover potential emotional problems of the elderly user in time and avoid further deterioration of the problem. This judgment method based on data and pre-set thresholds provides a relatively objective judgment standard, reducing the interference of human factors. Compared with simply relying on subjective observation or experience judgment, it is more accurate and reliable, and can provide strong support for subsequent intervention and care. The pre-set emotional parameter thresholds can be personalized according to individual differences, health conditions, and living habits of different elderly users. For example, for elderly users with certain specific diseases or special psychological conditions, thresholds that are more in line with their actual situation can be set to improve the accuracy and pertinence of emotional abnormality determination. Clear emotional abnormality determination results can provide decision-making basis for intelligent companion systems or caregivers. For example, when the system determines that the elderly user has emotional abnormalities, it can automatically trigger the function of sending prompt information to the caregiver, or adjust the voice interaction strategy with the elderly user to provide more targeted emotional support and care.

[0057] The elderly user emotional state interaction system based on the intelligent companion large model, as shown in Figure 2 includes: A data acquisition module for real-time acquisition of daily data information of the elderly user, and data processing of the daily data information to generate standardized time series data sets corresponding to the daily data information; An emotional abnormality determination module for fusing the daily data information of the elderly user using multi-modal emotional fusion to obtain fused emotional parameters; An emotional abnormality determination module for determining whether the elderly user has emotional abnormalities using the emotional parameters of the elderly user; A voice interaction module for sending prompt information to a caregiver when the elderly user has emotional abnormalities, and performing voice interaction with the elderly user according to the emotional state of the elderly user.

[0058] The working principle of the above technical solution is: real-time collection of daily data information of the elderly user through various sensors (such as microphone to collect voice data, camera to collect facial expression data, motion sensor to collect behavior data, etc.). Then process the collected raw data, remove noise and invalid data, and arrange and organize them in chronological order to generate a standardized time series dataset, providing a standardized and ordered data basis for subsequent analysis. Use multi-modal emotion fusion technology to integrate the data of different modalities such as voice, expression, and behavior in the standardized time series dataset to obtain emotion parameters. According to the obtained emotion parameters, compare and analyze with the pre-set normal emotion parameter range or threshold value. If the emotion parameter exceeds the normal range, such as excessive negative, anxiety, agitation, etc. that do not conform to the normal situation, it is determined that the elderly user has emotional abnormalities. These thresholds and ranges are determined based on statistical analysis of a large number of emotional data of elderly users and professional knowledge in the field of psychology. Once it is determined that the elderly user has emotional abnormalities, the system immediately sends a prompt message to the guardian, which can include the current emotional state description of the elderly user, related factors that may cause emotional abnormalities, etc. At the same time, according to the specific emotional state of the elderly user, the intelligent companion big model calls the corresponding voice interaction strategy and speech template to communicate with the elderly user. For example, for a depressed elderly person, give comforting, encouraging, and positive guidance words; for an anxious elderly person, provide content to alleviate their negative emotions and answer their doubts.

[0059] The effect of the above technical solution is: using multi-modal data fusion, overcoming the limitations of single modal monitoring, comprehensively analyzing the emotions of the elderly user from multiple dimensions, can more accurately capture subtle and complex emotional changes, improve the accuracy of emotional monitoring, and timely discover potential emotional problems of the elderly. Real-time data collection and processing can obtain the emotional state information of the elderly user at the first time, quickly respond to emotional abnormality, avoid the deterioration of emotional problems due to monitoring lag, and strive for valuable time for timely intervention and help. Send prompt information to the guardian in a timely manner, so that the guardian can remotely and timely understand the emotional state of the elderly user, facilitate the guardian to make corresponding decisions and actions, enhance the timeliness and effectiveness of emotional care for the elderly, and provide more comprehensive care for the elderly. According to the specific emotional state of the elderly user, personalized voice interaction is carried out, making the communication more in line with the current state of mind of the elderly, more easily causing emotional resonance, giving the elderly more thoughtful emotional comfort, improving the experience and satisfaction of the elderly user using the intelligent companion product, to a certain extent, alleviating the sense of loneliness of the elderly, and promoting their mental health.

[0060] Obviously, many modifications and variations of the present application are possible in light of the above teachings. It is, therefore, to be understood that within the scope of the appended claims and their equivalents, the application can be practiced otherwise than as specifically described.

Claims

1. A method for interacting with the emotional state of elderly users based on a large intelligent companionship model, characterized in that: The elderly user emotional state interaction method includes: Collecting daily data information of elderly users in real time, and performing data processing on the daily data information to generate a standardized time series data set corresponding to the daily data information; Multimodal emotion fusion is used to fuse the daily data information of the elderly user to obtain fused emotion parameters; wherein the multimodal emotions include voice emotion parameters and physiological emotion parameters, and the voice emotion parameters are determined by combining the calibration coefficient corresponding to each calibration time period with the emotion excitement coefficient corresponding to the emotion category appearing in each calibration time period; the physiological emotion parameters are set by combining the first heart rate data coefficient, the first cadence data coefficient, and the first body temperature data coefficient corresponding to the standardized time series data set with the second heart rate data coefficient, the second cadence data coefficient, and the second body temperature data coefficient corresponding to the calibration time period; Determining whether the elderly user has abnormal emotions by using the emotional parameters of the elderly user; When the elderly user has abnormal emotions, a prompt message is sent to the guardian, and voice interaction is performed with the elderly user according to the emotional state of the elderly user.

2. The method for interacting with the emotional state of elderly users according to claim 1, characterized in that: The daily data information of elderly users is collected in real time, and the daily data information is processed to generate a standardized time series data set corresponding to the daily data information, including: Using wearable devices to collect physiological data of elderly users; wherein the physiological data includes heart rate data, cadence data and body temperature data; Use voice collection equipment to collect voice data of elderly users; Recognizing the speech content of the elderly user's speech data, and obtaining conversation text data and audio waveform data corresponding to the elderly user's speech content; Retrieving the elderly user's physiological data, conversation text data, and audio waveform data to form multi-source data; Performing time stamp alignment on the multi-source data to establish a unified time axis t; Each type of data information contained in the multi-source data is normalized, and a standardized time series data set is generated using the normalized multi-source data.

3. The method for interacting with the emotional state of elderly users according to claim 1, characterized in that: The daily data information of the elderly user is fused using multimodal emotion fusion to obtain fused emotion parameters, including: Acquiring voice emotion parameters using voice data in the daily data information; Obtaining physiological emotion parameters using the physiological data in the daily data information; The voice emotion parameter and the physiological emotion parameter are fused to obtain a fused emotion parameter.

4. The method for interacting with the emotional state of elderly users according to claim 3, characterized in that: Acquiring voice emotion parameters using the voice data in the daily data information includes: Retrieving the dialogue text data and audio waveform data in the speech data from the standardized time series data set; The audio amplitude parameter in the audio waveform data; Comparing the audio amplitude parameter with a preset amplitude parameter threshold; Retrieving a time period during which the audio amplitude parameter exceeds the amplitude parameter threshold as a calibration time period; Extracting the conversation text data corresponding to the marked time period; Use the trained recurrent neural network model to obtain emotion categories from the conversation text data, and compare the emotion categories with the emotion category table stored in the database to obtain the emotional excitement coefficient corresponding to the emotion category; Extracting the heart rate data and cadence data corresponding to the calibration time period as the calibration heart rate data and calibration cadence data; The calibrated heart rate data and the calibrated step frequency data are combined with the emotional excitement coefficient to obtain the voice emotion parameter.

5. The method for interacting with the emotional state of elderly users according to claim 4, characterized in that: The speech emotion parameter is obtained by combining the calibrated heart rate data and the calibrated step frequency data with the emotion agitation coefficient, including: Obtaining a calibration coefficient corresponding to each calibration time period using the calibrated heart rate data and the calibrated cadence data; The speech emotion parameters are obtained by using the calibration coefficient corresponding to each calibration time period and the emotion excitement coefficient corresponding to the emotion category appearing in each calibration time period.

6. The method for interacting with the emotional state of elderly users according to claim 3, characterized in that: Obtaining physiological and emotional parameters using the physiological data in the daily data information includes: Retrieving heart rate data, cadence data, and body temperature data in the physiological data from the standardized time series data set; Using the heart rate data, cadence data, and body temperature data in the physiological data, obtain a standard deviation of the heart rate data, a standard deviation of the cadence data, and a standard deviation of the body temperature data; Performing ratio processing on the standard deviation of the heart rate data, the standard deviation of the cadence data, and the standard deviation of the body temperature data with their corresponding average values ​​of the heart rate data, the average values ​​of the cadence data, and the average values ​​of the body temperature data to obtain a first heart rate data coefficient, a first cadence data coefficient, and a first body temperature data coefficient corresponding to the heart rate data, the cadence data, and the body temperature data; Retrieving the heart rate data, cadence data, and body temperature data corresponding to all calibration time periods from the standardized time series data set to obtain the heart rate data standard deviation, cadence data standard deviation, and body temperature data standard deviation corresponding to all calibration time periods as a whole; Perform ratio processing on the standard deviation of the heart rate data, the standard deviation of the cadence data, and the standard deviation of the body temperature data corresponding to all the calibration time periods as a whole, and the minimum value of the heart rate data, the minimum value of the cadence data, and the minimum value of the body temperature data corresponding to the calibration time period as a whole, to obtain the second heart rate data coefficient, the second cadence data coefficient, and the second body temperature data coefficient corresponding to the heart rate data, the cadence data, and the body temperature data; The physiological emotion parameter is obtained by using the first heart rate data coefficient, the first cadence data coefficient and the first body temperature data coefficient in combination with the second heart rate data coefficient, the second cadence data coefficient and the second body temperature data coefficient.

7. The method for interacting with the emotional state of elderly users according to claim 3, characterized in that: Fusing the voice emotion parameter and the physiological emotion parameter to obtain the fused emotion parameter includes: Retrieving voice emotion parameters and physiological emotion parameters; Dynamically weighting the voice emotion parameters and the physiological emotion parameters to obtain dynamic weight values ​​corresponding to the voice emotion parameters and the physiological emotion parameters; The dynamic weight values ​​corresponding to the voice emotion parameters and the physiological emotion parameters are combined with the voice emotion parameters and the physiological emotion parameters to obtain the fused emotion parameters.

8. The method for interacting with the emotional state of elderly users according to claim 7, characterized in that: Dynamically weighting the voice emotion parameters and the physiological emotion parameters to obtain dynamic weight values ​​corresponding to the voice emotion parameters and the physiological emotion parameters includes: Extract historical reliability coefficients of wearable devices and voice acquisition devices; Utilizing historical reliability coefficients of the wearable device and the voice acquisition device; Retrieve the emotional excitement coefficient corresponding to each calibration time period; Obtaining the standard deviation of the emotional excitement coefficients corresponding to all calibration time periods using the emotional excitement coefficient corresponding to the calibration time period; The standard deviation of the emotional excitement coefficient corresponding to the calibration time period is combined with the historical reliability coefficient of the wearable device and the voice acquisition device to obtain the dynamic weight values ​​corresponding to the voice emotion parameter and the physiological emotion parameter.

9. The method for interacting with the emotional state of elderly users according to claim 1, characterized in that: Determining whether the elderly user has abnormal emotions by using the elderly user's emotional parameters includes: comparing the emotional parameter of the elderly user with a preset emotional parameter threshold; When the emotion parameter of the elderly user exceeds a preset emotion parameter threshold, it is determined that the elderly user has abnormal emotions.

10. An elderly user emotional state interaction system based on an intelligent companionship model, characterized by: The elderly user emotional state interaction system includes: A data collection module is used to collect daily data information of elderly users in real time, and perform data processing on the daily data information to generate a standardized time series data set corresponding to the daily data information; An emotional anomaly determination module is configured to fuse the daily data information of the elderly user using multimodal emotion fusion to obtain fused emotional parameters; wherein the multimodal emotions include voice emotional parameters and physiological emotional parameters, and the voice emotional parameters are determined by combining the calibration coefficient corresponding to each calibration time period with the emotional excitement coefficient corresponding to the emotion category appearing in each calibration time period; the physiological emotional parameters are set by combining the first heart rate data coefficient, the first cadence data coefficient, and the first body temperature data coefficient corresponding to the standardized time series data set with the second heart rate data coefficient, the second cadence data coefficient, and the second body temperature data coefficient corresponding to the calibration time period; An emotional abnormality determination module, configured to determine whether the elderly user has emotional abnormality by using the emotional parameters of the elderly user; The voice interaction module is used to send a prompt message to the guardian when the elderly user has abnormal emotions, and to perform voice interaction with the elderly user according to the emotional state of the elderly user.

Citation Information

Cited By

  • AI question and answer expert model construction method and system based on psychological counseling and medium

    CN121171262A

  • Intelligent elder accompanying robot and method based on emotion recognition

    CN121468611A