Old person emotional accompanying and health management large model of voice synthesis of mixed children
By constructing a large-scale model for emotional companionship and health management for the elderly, and utilizing multimodal physiological data acquisition and voice emotion control parameter adjustment, the problem of independent physiological health monitoring and voice synthesis was solved, enabling the synchronous emotional expression of the children's digital clone voice and the elderly's health status.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- XIAN CHENHUANTI DATA CO LTD
- Filing Date
- 2026-04-16
- Publication Date
- 2026-05-29
Smart Images

Figure CN122116869A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent elderly care and emotional companionship technology, and more specifically, to a large-scale model for emotional companionship and health management of the elderly that integrates the voice synthesis of children. Background Technology
[0002] With the accelerating aging of the population, the demand for daily emotional companionship and health management for elderly people living alone or in empty nests continues to grow. Currently, technological solutions in this field mainly follow three paths: First, a companion interaction system combining voice cloning and large language models, which replicates voiceprint features by collecting voice samples from children and uses large language models to generate dialogue content to simulate daily communication with family members; second, a multimodal emotion recognition interaction system, which judges the elderly person's emotional state through facial expression analysis and voice emotion recognition, and adjusts the response content and tone accordingly; and third, a multi-parameter physiological monitoring and health early warning system, which continuously collects, analyzes trends in, and alerts to abnormal physiological indicators such as blood pressure, heart rate, and blood sugar.
[0003] In the existing solutions described above, physiological health monitoring and emotional expression in speech are handled by separate subsystems. Health monitoring data is only used to generate numerical reports and trigger early warning notifications, while the emotional tone of the synthesized speech is determined based on a preset fixed style or semantic analysis results from the dialogue text. As a result, when the elderly person's real-time physiological indicators fluctuate or become abnormal, the emotional expression characteristics of the children's voices output by the system do not change accordingly in terms of speech rate, tone, pauses, etc. For example, an elderly person with high blood pressure or a fast heart rate may still receive a child's voice with a brisk pace and rising tone, resulting in a mismatch between the emotional expression in the speech and the elderly person's physical condition. In addition, although some solutions can recognize and mirror the emotions expressed by the elderly person, their emotional regulation relies on the emotions already expressed by the elderly person. They cannot proactively perceive and pre-understand the physical condition of the elderly person before they speak, thus making it difficult to reflect continuous attention and synchronous empathy for the elderly person's health status at the level of emotional expression. Therefore, a large-scale model for emotional companionship and health management for the elderly, which integrates the voice synthesis of children, is proposed to address the above problems. Summary of the Invention
[0004] To overcome the aforementioned deficiencies of the prior art, embodiments of the present invention provide a large-scale model for elderly emotional care and health management that integrates children's voice synthesis. The problem to be solved is that physiological health monitoring data and emotional expression of voice synthesis are independent of each other, and the emotional tone of the children's digital avatar voice cannot be synchronously correlated with the elderly's real-time health status, resulting in the problem that the voice emotional output does not match the elderly's physical condition.
[0005] To achieve the above objectives, this invention provides a large-scale model for elderly emotional care and health management that integrates children's voice synthesis.
[0006] This system continuously collects multimodal physiological data from the elderly, transforms the deviation of physiological characteristic parameters into an objective description of the elderly's physical condition, and then uses a lightweight language model to generate a text describing the care and feelings from the perspective of the children. It also analyzes the emotional tone and level of concern from the text and finally maps the above information into voice emotion control parameters to regulate the children's voice synthesis, so that the output of the children's digital avatar voice is synchronized with the elderly's real-time health status in terms of emotional expression.
[0007] The system includes a physiological state monitoring module, a deviation quantization module, a semantic conversion module, an emotion analysis module, a parameter mapping module, and a child speech synthesis module.
[0008] The physiological state monitoring module continuously collects multimodal physiological data from the elderly and extracts physiological characteristic parameters, including blood pressure, heart rate, and pulse wave transit time. Through continuous collection of multimodal physiological data, objective quantitative evidence reflecting the elderly's real-time physical condition is obtained, providing a data foundation for subsequent health status assessment and voice emotion regulation.
[0009] Furthermore, the multimodal physiological data collected by the physiological state monitoring module also includes one or more of blood glucose data, heart rate variability data, or environmental sensor data, wherein the environmental sensor data includes one or more of ambient temperature, ambient humidity, or ambient light intensity.
[0010] By expanding the dimensions of physiological and environmental data collection, the information content of subsequent physiological deviation vectors is enriched, enabling health status assessment to integrate more influencing factors and improving the comprehensiveness of the representation of the elderly's physical condition.
[0011] The deviation quantification module calculates the degree of deviation by comparing the current collected values of each physiological characteristic parameter with the individualized baseline value of the elderly person using a deviation calculation formula, and then combines the various deviation degrees in sequence into a physiological deviation vector. Through the quantitative calculation of the degree of deviation, physiological indicators of different dimensions are uniformly transformed into comparable deviation measures, and the overall fluctuation of multiple physiological indicators of the elderly person relative to their normal state is structurally represented in vector form.
[0012] Furthermore, the deviation calculation formula executed by the deviation quantization module is as follows: Degree of deviation ; in, These are the current collected values of the physiological characteristic parameters. This is a personalized baseline value for this physiological characteristic parameter. This is the threshold for the normal fluctuation range of this physiological characteristic parameter. This calculation formula normalizes the absolute deviation between the current collected value and the personalized baseline value using the normal fluctuation range threshold as a reference, so that the degree of deviation of different physiological indicators has a unified dimension and comparability.
[0013] The semantic transformation module, based on the degree of deviation of each parameter in the physiological deviation vector, maps indicators whose deviation exceeds the normal range into natural language description fragments through semantic mapping rules and concatenates them into objective situation description text. The objective situation description text is then filled into a preset child care perspective prompt template and input into a lightweight language model deployed on the edge computing unit to obtain the care feeling description text output by the lightweight language model.
[0014] This module transforms numerical physiological deviation information into emotionally charged natural language descriptions, simulating the feelings and tone of children after observing their parents' physical condition, thus providing a semantic basis for voice emotion regulation.
[0015] Furthermore, the semantic mapping rules in the semantic transformation module include the correspondence between physiological indicators, deviation direction, deviation degree intervals, and natural language description fragments. This correspondence discretizes continuous deviation degree values into multiple intervals, each interval corresponding to a descriptive text that conforms to everyday expression habits, achieving a regularized and interpretable transformation from numerical values to semantics.
[0016] Furthermore, the pre-defined child-care perspective prompt template includes prompt text that constrains the lightweight language model to describe its inner feelings in the first person, as if it were written by a child. This prompt template, through role setting and output requirements, constrains the language model's generative behavior, ensuring that its output description of care and feelings focuses on inner feelings and tone, providing semantically clear input for subsequent sentiment analysis.
[0017] The emotion analysis module performs emotional semantic analysis on the text describing the feelings of care, identifies core words indicating emotional state, categorizes the emotional tone into one of worry, concern, or reassurance, matches the intensity modifiers of the core words with a preset intensity level mapping table to determine the degree of concern score, and determines the speech rate tendency based on the emotional tone and the degree of concern score. This module extracts quantifiable emotional features from the text describing the feelings of care, establishes a mapping from natural language emotion to numerical emotion parameters, and provides clear quantitative input for subsequent speech control.
[0018] Furthermore, the sentiment analysis module compares the core vocabulary with a preset sentiment tone lexicon to categorize sentiment tone, and compares the intensity modifiers with a preset intensity level mapping table to determine the degree of concern score. Through a rule-based approach of lexicon matching and mapping table lookup, sentiment category and intensity information are stably extracted from natural language text, ensuring the certainty and consistency of sentiment analysis results.
[0019] The parameter mapping module generates speech emotion control parameters according to the mapping relationship between the emotional tone category and the concern level score. The speech emotion control parameters include speech rate factor and fundamental frequency offset. This module converts the qualitative category and quantitative score obtained from emotion analysis into acoustic control parameters that can be directly executed by the speech synthesis engine, establishing a parameterized mapping channel from emotional state to speech expression features.
[0020] Furthermore, the mapping relationship executed by the parameter mapping module is as follows: When the emotional tone is worried, the fundamental frequency offset is negative and the speech rate factor is inversely proportional to the concern score, which makes the tone of voice tend to be low and the speech rate slows down as the level of worry increases, simulating the natural tone change of children when they are worried about their parents' health. When the emotional tone is concerned, the fundamental frequency offset is zero or negative and the absolute value is less than the absolute value of the fundamental frequency offset when the tone is worried. The speech rate factor is taken within the normal speech rate range, so that the voice is slightly lower while maintaining a stable tone, reflecting appropriate attention and reminder. When the emotional tone is reassuring, the fundamental frequency offset is zero or positive, and the speech rate factor is within or above the normal speech rate range, making the tone of voice natural or slightly rising, conveying a sense of peace and relaxation.
[0021] Furthermore, the voice emotion control parameters also include pause insertion markers and energy envelope adjustment parameters. When the concern score exceeds a threshold, the parameter mapping module inserts a pause marker at the phrase boundary that is 200 to 400 milliseconds longer than the regular pause duration, and increases the attenuation coefficient of the voice loudness envelope. By adding pauses at phrase boundaries and softening the loudness envelope, the voice rhythm becomes more stable and the tone more concerned, further enhancing the match between the voice emotion expression and the elderly person's health status.
[0022] Upon receiving a dialogue command triggered by the elderly person or a system-timed greeting command, the child's voice synthesis module obtains the child's response text generated by a large language model. This response text, along with a pre-set child's voiceprint embedding vector and the latest voice emotion control parameters, is input into the voice synthesis engine to drive the engine to output the child's digital avatar voice. This module integrates health-state-driven emotion control parameters with the dialogue content and the child's vocal characteristics, ensuring that the output child's digital avatar voice reflects emotional expressions adapted to the elderly person's real-time physical condition in terms of speech rate, intonation, and other acoustic dimensions.
[0023] Furthermore, it also includes a child voiceprint registration module, which pre-collects child speech samples and extracts voiceprint features to obtain the child voiceprint embedding vector. During the synthesis process, the child speech synthesis module adjusts the phoneme duration prediction parameters according to the speech rate factor and adjusts the fundamental frequency curve generation parameters according to the fundamental frequency offset. The child voiceprint embedding vector provides a basis for timbre consistency, and by scaling the phoneme duration using the speech rate factor and shifting the fundamental frequency curve using the fundamental frequency offset, independent control of the synthesized speech rate and intonation is achieved without altering the timbre characteristics.
[0024] Furthermore, a forced modulation module is also included. When the deviation of any key physiological indicator in the physiological deviation vector exceeds a safety threshold, the forced modulation module blocks the output of the semantic conversion module and the emotion analysis module, and provides a forced modulation parameter set to the child's speech synthesis module. This forced modulation parameter set includes a forced speech rate factor below the normal speech rate range and a negative forced fundamental frequency offset. This module provides a forced emotion modulation path independent of the conventional semantic analysis process when the elderly person's physical condition shows significant abnormalities, ensuring that the child's speech output in such cases consistently conveys a speech expression with clear signs of concern.
[0025] The technical effects and advantages of this invention are as follows: The technical solution of this invention constructs a multi-level processing flow from multimodal physiological data acquisition to the generation of voice emotion control parameters. It uses the elderly person's real-time health status as a pre-driving signal for voice emotion expression, ensuring that the emotional characteristics of the children's digital avatar's voice are synchronized with the elderly person's physical condition. Specific technical effects and advantages are described below.
[0026] First, the deviation quantification module quantifies the degree of deviation of each physiological characteristic parameter from the personalized baseline value and combines them into a physiological deviation vector in a fixed order. This vector reflects the fluctuation of multiple physiological indicators of the elderly in a structured form in real time, providing a unified and quantified health status input for subsequent semantic transformation.
[0027] Second, the semantic transformation module converts indicators that deviate beyond the normal range into natural language descriptive fragments using preset semantic mapping rules, and then concatenates them into an objective situation description text. Based on this, the objective situation description text is filled into a children's care perspective prompt template, driving a lightweight language model deployed on the edge computing unit to generate a care feeling description text. This care feeling description text simulates the feelings and tone that children experience after perceiving their elderly parent's current physical condition, providing a semantic basis for voice emotion regulation.
[0028] Third, the emotion analysis module identifies core vocabulary related to emotional states from the text describing feelings of care, categorizing the emotional tone as worry-based, concerned, or reassured, and determining the level of concern score by matching intensity modifiers with a preset mapping table. The parameter mapping module generates speech emotion control parameters, including speech rate factor, fundamental frequency offset, pause insertion markers, and energy envelope adjustment parameters, based on the emotional tone category and the level of concern score. The generation of these parameters is correlated with the elderly person's current physiological deviation, providing a traceable and quantitative basis for speech emotion regulation.
[0029] Fourth, after a dialogue is triggered, the child's voice synthesis module inputs the response text generated by the large language model, the pre-set child's voiceprint embedding vector, and the voice emotion control parameters driven by physiological state into the voice synthesis engine. The voice synthesis engine adjusts the phoneme duration prediction parameters based on the speech rate factor, adjusts the fundamental frequency curve generation parameters based on the fundamental frequency offset, and modulates the rhythm and loudness envelope of the synthesized speech based on the pause insertion markers and energy envelope adjustment parameters. The resulting digital avatar of the child's voice has speech rate, intonation, pauses, and other emotional expression characteristics that are adapted to the elderly person's current physical condition.
[0030] Through the aforementioned technical solution, when the elderly person's various physiological indicators are within the normal fluctuation range, the children's voice tone is natural and stable; when abnormal deviations such as high blood pressure or fast heart rate are detected, the children's speech speed slows down, the fundamental frequency drops, and caring pauses are inserted. This mechanism does not rely on recognizing the elderly person's proactive expression of emotions, but is based on continuous perception and quantitative analysis of physiological state, which helps improve the realism and empathic consistency of the digital avatar in emotional companionship scenarios. Attached Figure Description
[0031] Figure 1 This is a system module framework diagram of the present invention. Detailed Implementation
[0032] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0033] Example 1 As attached Figure 1 The model shown integrates voice synthesis for elderly emotional support and health management. This system continuously collects multimodal physiological data from the elderly, and its specific implementation is as follows: The system hardware includes multi-parameter physiological monitoring equipment deployed in elderly people's residences, edge computing units, voice interaction terminals, and cloud-based large language model servers connected via communication networks.
[0034] The multi-parameter physiological monitoring device integrates a blood pressure sensor, a heart rate sensor, a pulse sensor, a blood glucose monitoring module, and an environmental sensor.
[0035] The edge computing unit uses an embedded processor with neural network inference capabilities, runs the Linux operating system, and deploys the deviation quantization module, semantic conversion module, sentiment analysis module, parameter mapping module, forced control module, and the speech synthesis engine part of the child speech synthesis module.
[0036] The voice interaction terminal includes a microphone array and a speaker. The microphone array supports wake word detection and voice activity detection and is connected to the edge computing unit via an I2S interface. The speaker receives and plays audio data via the I2S interface.
[0037] As a specific hardware implementation, the blood pressure sensor adopts an oscillometric cuff-type blood pressure monitor module, which automatically inflates and measures every 30 minutes. Each measurement outputs systolic and diastolic blood pressure and communicates with the edge computing unit through a UART serial interface.
[0038] The heart rate sensor uses a photoplethysmography (PPG) sensor to continuously acquire pulse wave signals at a sampling frequency of 100Hz and outputs digital heart rate values via an I2C bus.
[0039] The pulse sensor includes a single-lead ECG electrode and a transmissive photoelectric sensor, which respectively collect ECG signals and fingertip pulse wave signals. The sampling frequency is 256Hz, and the raw waveform data is output through the SPI interface.
[0040] The blood glucose monitoring module uses a continuous blood glucose monitoring sensor. The probe is implanted in the subcutaneous tissue and measures the glucose concentration of the interstitial fluid every 5 minutes and converts it into a blood glucose value. The data is then sent to the edge computing unit via Bluetooth Low Energy protocol.
[0041] The environmental sensors include a DS18B20 digital temperature sensor, a DHT22 capacitive humidity sensor, and a BH1750 photodiode illuminance sensor, which output ambient temperature, ambient humidity, and ambient illuminance via a single bus, an I2C bus, and an I2C bus, respectively.
[0042] The edge computing unit can be an NVIDIA Jetson Orin Nano embedded processor, and the voice interaction terminal can be a ReSpeaker circular 4-microphone array and a full-band speaker unit.
[0043] The following provides a detailed explanation of the specific implementation methods of each module.
[0044] The physiological state monitoring module is as follows: The physiological state monitoring module operates in a multi-threaded manner, with each sensor corresponding to an independent data acquisition thread, and the threads synchronously accessing the shared data buffer through a mutex lock.
[0045] First, the blood pressure sensor thread triggers a measurement every 30 minutes. After the measurement is completed, the systolic blood pressure value, diastolic blood pressure value and timestamp are encapsulated into a data frame and sent to the edge computing unit through the UART interface.
[0046] Secondly, the heart rate sensor thread reads heart rate values from the I2C bus at a frequency of 100Hz, calculates the average heart rate value every 2 seconds, and encapsulates the average heart rate value with a timestamp and writes it into a circular buffer.
[0047] Secondly, the pulse sensor thread synchronously acquires ECG signals and fingertip pulse wave signals. For each cardiac cycle, it calculates the time difference between the peak of the ECG R wave and the peak of the main pulse wave. The average value of five consecutive cardiac cycles is taken as the pulse wave transmission time, which is then encapsulated with a timestamp and written into a circular buffer.
[0048] Then, the blood glucose monitoring module thread reads the blood glucose value and corresponding timestamp from the Bluetooth receive buffer and writes them directly into the circular buffer.
[0049] Finally, the environmental sensor thread reads the temperature, humidity, and illuminance values every 10 seconds, encapsulates them with a timestamp, and writes them into a circular buffer.
[0050] The circular buffer uses fixed-size circular storage units, each storing a set of timestamp-aligned physiological characteristic parameters. When new data is written and overwrites old data, the oldest record is automatically discarded.
[0051] The deviation quantization module is as follows: The deviation quantization module reads the latest set of physiological characteristic parameters from the circular buffer every 30 seconds, performs deviation calculation, and the calculation results are used by the semantic conversion module and the forced regulation module.
[0052] The personalized baseline values were determined as follows: during the system initialization phase, various physiological characteristic parameters of the elderly were continuously collected for 72 hours in a resting state. The resting state was determined by a combination of heart rate and accelerometer data, requiring the heart rate to be within the resting heart rate range and the combined amplitude of the three-axis acceleration to be less than 0.1g.
[0053] For the physiological characteristic parameter sequences collected in the resting state, after removing outliers that exceed 3 times the standard deviation, the arithmetic mean of the remaining collected values is taken as the personalized baseline value of the parameter and stored in the local non-volatile storage of the edge computing unit.
[0054] For blood pressure values, personalized baseline values are established for the morning peak period and the non-morning peak period. The morning peak period is defined as 6:00 to 10:00 a.m. and the non-morning peak period is the rest of the time.
[0055] The deviation quantization module performs the following deviation calculation formula for each physiological characteristic parameter: Degree of deviation = in, These are the current collected values of the physiological characteristic parameters. This is a personalized baseline value for this physiological characteristic parameter. This is the threshold value for the normal fluctuation range of this physiological characteristic parameter.
[0056] Normal fluctuation range threshold Based on the elderly person's historical data, the following was obtained: the physiological characteristic parameter sequence of the resting state during the initial stage was taken, and its standard deviation was calculated. , will 3 As The values of different physiological characteristic parameters. The values are calculated and stored separately.
[0057] For systolic blood pressure, The value range is limited to 10 mmHg to 20 mmHg; Regarding heart rate, The value range is limited to 8 times per minute to 15 times per minute; Regarding pulse propagation time The value is limited to 10% to 20% of its baseline value.
[0058] The deviation quantification module arranges the calculated deviations in a preset fixed order and combines them into a physiological deviation vector. The preset fixed order is: systolic blood pressure deviation, heart rate deviation, pulse wave transit time deviation, blood glucose deviation, and heart rate variability deviation.
[0059] For physiological characteristic parameters for which no valid data was obtained, the corresponding deviation component is set to zero. The physiological deviation vector is stored in shared memory as a binary data structure for subsequent modules to read.
[0060] The semantic transformation module is as follows: The semantic transformation module reads the latest physiological deviation vector from shared memory, performs semantic transformation, and sends the generated care and feeling description text to the sentiment analysis module through a message queue.
[0061] Semantic mapping rules are stored in JSON format in the local storage of the edge computing unit. The rule file defines the mapping relationship between physiological indicator types, deviation directions, deviation degree ranges, and natural language description fragments.
[0062] For systolic blood pressure, when the deviation is greater than or equal to 1.0 and less than 1.2, Greater than At that time, it was mapped to the descriptive fragment "blood pressure is slightly high"; When the deviation is greater than or equal to 1.2 and Greater than At that time, it was mapped to the descriptive fragment "blood pressure was much higher than usual"; When the deviation is greater than or equal to 1.0 and Less than At that time, it was mapped to the descriptive fragment "blood pressure is lower than usual".
[0063] For heart rate indicators, when the deviation is greater than or equal to 1.0 and X is greater than B, it is mapped to the descriptive fragment "heart rate is faster than usual"; When the deviation is greater than or equal to 1.0 and X is less than B, it is mapped to the description segment "heartbeat is slower than usual".
[0064] For blood glucose indicators, when the blood glucose value is lower than the preset lower limit of the normal range, it is mapped to the description fragment "blood glucose is a little low", and when the blood glucose value is higher than the preset upper limit of the normal range, it is mapped to the description fragment "blood glucose is high".
[0065] All mapping rules are loaded into a hash table structure in memory during system initialization for fast retrieval.
[0066] The semantic transformation module iterates through each component of the physiological deviation vector. For components with a deviation degree greater than or equal to 1.0, it retrieves the corresponding natural language description fragment from the hash table based on the physiological indicator type, deviation direction, and deviation degree interval. The deviation direction is determined by comparing the magnitudes of X and B. When X is greater than B, the deviation direction is positive; when X is less than B, the deviation direction is negative.
[0067] For components with a deviation of less than 1.0, no descriptive fragment is generated.
[0068] When the deviation of all components is less than 1.0, a preset normal state description fragment, "All indicators are quite normal," is generated. All retrieved description fragments are then concatenated according to the component order in the physiological deviation vector, with fragments separated by commas, to form an objective state description text.
[0069] The pre-defined child-care perspective prompt template is a fixed string constant containing the following content: "Suppose you are the child of an elderly person. The following is an observation of your parents' current physical condition: {objective condition description text}. As a child who cares about your parents' health, facing these circumstances..." The semantic conversion module performs a string replacement operation, replacing the placeholder "{objective condition description text}" with the actual generated objective condition description text, forming the complete prompt text.
[0070] The lightweight language model is deployed in the ONNXRuntime inference framework on the edge computing unit.
[0071] The model takes a text sequence as input, with a maximum input length of 512 words. The model performs autoregressive inference on the input prompt text, generating an output sequence word by word, with a maximum output length of 128 words. After inference, the generated word sequence is decoded into text, serving as a description of the feelings of care and concern.
[0072] The sentiment analysis module is as follows: The sentiment analysis module receives text describing feelings of care from the message queue, performs sentiment semantic analysis, and writes the analysis results to shared memory for the parameter mapping module to read.
[0073] The preset emotional tone vocabulary is stored in the form of a UTF-8 encoded text file.
[0074] The vocabulary for words related to worry includes: worry, tense, nervous, afraid, unable to let go, anxious, uneasy, apprehensive, heartbroken, and concerned. The vocabulary for words expressing concern includes: to remind, to pay attention, to ask, to get more rest, to remind, to be mindful, to instruct, to take care of, and to pay attention to. The reassuring word library includes: okay, quite stable, no problem, reassured, at ease, steady, stable, normal, and secure. The words in each word library are loaded into three independent hash sets during system initialization.
[0075] The sentiment analysis module calls the jieba Chinese word segmentation library to segment the text describing the caring feeling, obtaining a word sequence. Each word in the word sequence is matched one by one with the three hash sets: If it matches successfully with the worried word library, the sentiment tone is set to worried; If it does not match the worried word library but matches successfully with the concerned word library, it is set to concerned; If it only matches successfully with the reassuring word library, it is set to reassuring.
[0076] If it matches multiple word libraries simultaneously, the final sentiment tone category is determined according to the priority rule that the worried type takes precedence over the concerned type, and the concerned type takes precedence over the reassuring type.
[0077] The preset intensity level mapping table is stored in memory in the form of key-value pairs. The specific corresponding relationship is: the concern degree scores corresponding to the degree adverbs "a bit", "slightly", and "somewhat" are 0.3; When there is no modifier or the modifier is "average" or "ordinary" are 0.5; When the degree adverbs are "very", "extremely", "especially", "extremely", "extremely", "feeling tight in the heart" are 0.8.
[0078] The sentiment analysis module searches forward starting from the position of the matched core word in the word sequence, and takes the first degree adverb found as the intensity modifier. Search in the mapping table with this intensity modifier as the key to obtain the concern degree score . If there is no degree adverb before the core word, the default concern degree score of 0.5 is assigned.
[0079] The speech rate tendency is determined according to the sentiment tone and the concern degree score as follows: When the sentiment tone is worried and is greater than or equal to 0.5, the speech rate tendency is set to slow; When the sentiment tone is concerned, the speech rate tendency is set to normal; When the sentiment tone is reassuring, the speech rate tendency is set to normal or slightly faster.
[0080] The sentiment analysis module encapsulates the sentiment tone category, the concern degree score and the speech rate tendency into a data structure and writes it to the specified area of the shared memory.
[0081] The parameter mapping module is as follows: The parameter mapping module reads the emotion parsing results from the shared memory, generates voice emotion control parameters, and writes the parameters into the shared memory for use by the child speech synthesis module.
[0082] The parameter mapping module performs the following mapping based on the sentiment tone category.
[0083] When the emotional tone is one of worry, the fundamental frequency offset Take negative values, and the absolute value varies with the score of the level of concern. The increase is gradual and stepwise: like If less than 0.7, then Take -2 semitones, if If greater than or equal to 0.7, then Take -3 semitones; speech rate factor according to Calculation, where The preset adjustment coefficient is set to 0.3. The value range is from 0.7 to 1.0. The higher the The smaller.
[0084] When the emotional tone is caring, the fundamental frequency offset Take -1 semitone, speech rate factor Take 0.95.
[0085] When the emotional tone is reassured, the fundamental frequency offset Take 0 semitones or +1 semitone, speech rate factor Choose 1.0 or 1.05.
[0086] Voice emotion control parameters also include pause insertion markers and energy envelope adjustment parameters. When the concern level score... When the threshold of 0.5 is exceeded, the parameter mapping module sets the pause insertion flag to valid and calculates the pause increment. , The value is in the range of 200 milliseconds to 400 milliseconds. They are positively correlated; at the same time, the attenuation coefficient in the energy envelope adjustment parameter is increased. Set to the range of 1.3 to 1.5 with Values that show a positive correlation.
[0087] The parameter mapping module writes the generated speech emotion control parameters into shared memory in the form of a structure, which contains a speech rate factor. Fundamental frequency offset , pause insertion mark, pause increment and attenuation factor .
[0088] The child's speech synthesis module is as follows: After receiving a dialogue trigger command, the child's voice synthesis module executes the voice synthesis process and finally outputs the voice of the child's digital clone with health perception and empathy.
[0089] The dialogue triggering commands include commands triggered by the elderly and system-timed greeting commands. The elderly-triggered dialogue commands are generated by the wake-word detection module, which uses the Snowboy hot word detection engine to continuously detect the audio stream collected by the microphone array in real time. When the detection confidence exceeds 0.8, a trigger signal is sent to the children's speech synthesis module. System-timed greeting commands are generated by cron jobs according to a daily schedule of 8:00, 12:00, and 18:00.
[0090] Upon receiving the trigger signal, the child's speech synthesis module sends a dialogue generation request to the cloud-based large language model service. The request content is sent via HTTPS POST and includes the current dialogue context and a summary of the elderly person's health status. The cloud-based large language model, fine-tuned using dialogue data from the elderly care field, generates a response text in the child's voice upon receiving the request and returns it to the edge computing unit.
[0091] The speech synthesis engine consists of a cascaded VITS acoustic model and a HiFi-GAN vocoder, both deployed on edge computing units. The VITS acoustic model receives the response text, child voiceprint embedding vectors, and speech emotion control parameters as input, and outputs an 80-dimensional Mel-spectrum feature sequence; the HiFi-GAN vocoder receives the Mel-spectrum feature sequence and outputs a monophonic speech waveform with a sampling rate of 22.05kHz.
[0092] During the synthesis process, the speech synthesis engine performs the following adjustments.
[0093] First, the duration predictor in the VITS acoustic model outputs a sequence of duration frames for each phoneme. . will sequence Each element is multiplied by the speech rate factor Rounding down yields the adjusted sequence of continuous frames. Subsequent upsampling operations are based on Alternative conduct.
[0094] Secondly, fundamental frequency offset Using semitones as units, through conversion relationships Convert to frequency ratio During the waveform synthesis stage, the generated fundamental frequency curve is multiplied point by point. Perform frequency scaling.
[0095] Furthermore, when the pause insertion flag is valid, the speech synthesis engine inserts a silence segment at the phrase boundary position corresponding to the punctuation mark in the response text. The duration of the silence segment is equal to the increase in pause size. .
[0096] Finally, when the energy envelope adjustment parameter is valid, the speech synthesis engine multiplies the energy values of each frequency band in the Mel spectrum by a factor that linearly decays from 1 to [the value]. The coefficient sequence is such that the loudness at the end of the phrase is multiplied by the attenuation factor. Accelerates decay.
[0097] The synthesized speech waveform is output to the speaker for playback via the ALSA audio driver.
[0098] The mandatory control module is as follows: The forced regulation module runs in an independent thread and checks the physiological deviation vector in the shared memory every 30 seconds, providing the system with the ability to forcibly intervene in the voice emotion expression under health risk conditions.
[0099] The forced control module maintains a global flag. The default value is the first logical state. When a deviation of systolic blood pressure or heart rate greater than or equal to 1.5 is detected, it will... Set to the second logical state. The semantic conversion module and the sentiment analysis module check before the start of each processing cycle. State: If If the second logical state is reached, the current processing is skipped, and no data is output to subsequent modules.
[0100] when When set to the second logical state, the forced modulation module writes the forced modulation parameter set into the shared memory of the child speech synthesis module. The forced modulation parameter set contains the forced speech rate factor. Take 0.75, forced fundamental frequency offset Take -3 semitones, enable forced pause insertion mark, and increase forced pause amount. Take 300 milliseconds.
[0101] When both the deviation in systolic blood pressure and the deviation in heart rate decrease to below 1.5, Restore to the first logical state and resume normal data flow.
[0102] The specific details of the children's voiceprint registration module are as follows: During the system initialization phase, the child voiceprint registration module guides the child to register their voiceprint through a voice interaction terminal, thus establishing the vocal foundation for the child's digital avatar. The system plays a preset registration prompt tone, and the child reads the preset registration text aloud as prompted. The microphone array captures the read-out speech at a 16kHz sampling rate and 16-bit quantization precision, saving it as a mono PCM format file.
[0103] The voiceprint extraction model employs the ECAPA-TDNN architecture. The input is an 80-dimensional Mel-frequency cepstral coefficient feature sequence extracted from a PCM file, with a frame length of 25 milliseconds and a frame shift of 10 milliseconds. The model outputs a 192-dimensional voiceprint embedding vector, which, after L2 normalization, is stored as a binary floating-point array in the local path of the edge computing unit. The child speech synthesis module reads the voiceprint embedding vector from this path during each synthesis.
[0104] As an alternative hardware implementation, blood pressure values can be continuously estimated using an optical sensor array on a wrist-worn wearable device. The device employs photoplethysmography combined with pulse wave transit time (PRT) for cuffless continuous blood pressure estimation, sampling at a frequency of 100Hz, and outputs the estimated blood pressure value every 10 seconds via Bluetooth. Heart rate values are obtained by converting the RR interval from the device's built-in ECG sensor. Pulse wave transit time is calculated by simultaneously acquiring ECG signals and radial artery pulse wave signals.
[0105] As another implementation method for semantic mapping rules, these rules can be stored in XML format. The rule file contains multiple rule nodes, each defining physiological indicator attributes, deviation direction attributes, deviation degree range attributes, and descriptive text sub-nodes. The semantic transformation module uses the libxml2 library to parse the rule file and construct an in-memory rule tree structure. When traversing the physiological deviation vector components, matching nodes are retrieved in the rule tree based on the component attributes to extract the descriptive text. The lightweight language model is deployed within the inference framework of the edge computing unit.
[0106] As another approach to sentiment analysis, the sentiment analysis module can use the THULAC Chinese word segmentation tool to segment the text describing feelings of care and concern. This tool is based on a conditional random field model. The content of the sentiment tone lexicon and intensity level mapping table is the same as in the aforementioned method. The segmented sequence is then matched with the lexicon, and the matching logic is also the same.
[0107] As another way to implement parameter mapping, when the emotional tone is worried, the speech rate factor... The calculation can be performed using a nonlinear mapping function. ,in The preset curvature coefficient is set to 0.4, so that... Follow The magnitude increases and then gradually decreases. The mapping relationship under other emotional tones is the same as described above.
[0108] As another approach to voiceprint extraction, the voiceprint extraction model can employ the ResNet-34 architecture, using the AAM-Softmax loss function during training. The model input is a 64-dimensional Mel-frequency cepstral coefficient feature, and the output is a 128-dimensional voiceprint embedding vector. The extracted vector is stored in the same way as described above.
[0109] Through the detailed description of the above specific embodiments, those skilled in the art can clearly understand the specific implementation methods of each module of the present invention, the input and output interfaces, and the data flow relationship between modules, and select appropriate hardware configurations and algorithm parameters according to actual application scenarios to implement the present invention.
[0110] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A large-scale model for elderly emotional companionship and health management that integrates children's voice synthesis, characterized in that, include: The physiological state monitoring module continuously collects multimodal physiological data of the elderly and extracts physiological characteristic parameters, including blood pressure, heart rate and pulse wave transit time. The deviation quantification module calculates the degree of deviation by comparing the current collected values of each physiological characteristic parameter with the individualized baseline value of the elderly person according to the deviation calculation formula, and combines the degree of deviation of each item in order into a physiological deviation vector. The semantic transformation module, based on the degree of deviation of each parameter in the physiological deviation vector, maps the indicators whose deviation exceeds the normal range into natural language description fragments and concatenates them into objective situation description text through semantic mapping rules. The objective situation description text is then filled into a preset child care perspective prompt template and input into a lightweight language model deployed on the edge computing unit to obtain the care feeling description text output by the lightweight language model. The sentiment analysis module performs sentiment semantic analysis on the text describing the feelings of care, identifies core words that indicate emotional state, classifies the emotional tone into one of worry, concern or reassurance, matches the intensity modifiers of the core words with a preset intensity level mapping table to determine the degree of concern score, and determines the speech rate tendency based on the emotional tone and the degree of concern score. The parameter mapping module generates voice emotion control parameters according to the mapping relationship between the emotional tone category and the concern level score. The voice emotion control parameters include speech rate factor and fundamental frequency offset. The children's voice synthesis module, upon receiving a dialogue command triggered by the elderly or a system-timed greeting command, obtains the children's response text generated by the large language model, inputs the response text, the preset children's voiceprint embedding vector, and the latest voice emotion control parameters into the voice synthesis engine, and drives the voice synthesis engine to output the children's digital clone voice.
2. The large-scale model for elderly emotional companionship and health management integrating children's voice synthesis as described in claim 1, characterized in that, The multimodal physiological data collected by the physiological state monitoring module also includes one or more of blood glucose data, heart rate variability data, or environmental sensor data, wherein the environmental sensor data includes one or more of ambient temperature, ambient humidity, or ambient light intensity.
3. The large-scale model for elderly emotional companionship and health management integrating children's voice synthesis as described in claim 1, characterized in that, The deviation calculation formula executed by the deviation quantization module is: Degree of deviation = ; Where X is the current collected value of the physiological characteristic parameter, and B is the personalized baseline value of that physiological characteristic parameter. This is the threshold value for the normal fluctuation range of this physiological characteristic parameter.
4. The large-scale model for elderly emotional companionship and health management integrating children's voice synthesis as described in claim 1, characterized in that, The semantic mapping rules in the semantic transformation module include the correspondence between physiological indicators, deviation direction, deviation degree range and natural language description fragments.
5. The large-scale model for elderly emotional companionship and health management integrating children's voice synthesis as described in claim 1, characterized in that, The preset child care perspective prompt template includes prompt text that constrains the lightweight language model to describe its inner feelings in the first person as a child.
6. The large-scale model for elderly emotional companionship and health management integrating children's voice synthesis as described in claim 1, characterized in that, The sentiment analysis module compares the core vocabulary with a preset sentiment tone lexicon to classify sentiment tone, and compares the intensity modifiers with a preset intensity level mapping table to determine the degree of concern score.
7. The large-scale model for elderly emotional companionship and health management integrating children's voice synthesis as described in claim 6, characterized in that, The mapping relationship executed by the parameter mapping module is as follows: When the emotional tone is one of worry, the fundamental frequency offset takes a negative value and the speech rate factor is inversely proportional to the score of concern. When the emotional tone is concerned, the fundamental frequency offset is zero or negative and its absolute value is less than the absolute value of the fundamental frequency offset when the tone is worried, and the speech rate factor is within the normal speech rate range. When the emotional tone is reassuring, the fundamental frequency offset is zero or a positive value, and the speech rate factor is within the normal speech rate range or higher.
8. The large-scale model for elderly emotional companionship and health management integrating children's voice synthesis as described in any one of claims 1 to 5, characterized in that, The voice emotion control parameters also include pause insertion markers and energy envelope adjustment parameters; when the concern score exceeds the threshold, the parameter mapping module inserts a pause marker at the phrase boundary that is 200 to 400 milliseconds longer than the normal pause duration, and increases the attenuation coefficient of the voice loudness envelope.
9. The large-scale model for elderly emotional companionship and health management integrating children's voice synthesis as described in any one of claims 1 to 5, characterized in that, It also includes a forced regulation module, which blocks the output of the semantic conversion module and the emotion analysis module when the deviation of any key physiological indicator in the physiological deviation vector exceeds the safety threshold, and provides a forced regulation parameter set to the child speech synthesis module. The forced regulation parameter set includes a forced speech rate factor below the normal speech rate range and a negative forced fundamental frequency offset.
10. The large-scale model for elderly emotional companionship and health management integrating children's voice synthesis as described in any one of claims 1 to 5, characterized in that, It also includes a child voiceprint registration module, which pre-collects child voice samples and extracts voiceprint features to obtain the child voiceprint embedding vector; the child voice synthesis module adjusts the phoneme duration prediction parameters according to the speech rate factor and adjusts the fundamental frequency curve generation parameters according to the fundamental frequency offset during the synthesis process.