Intelligent Health Robot Integrating Intelligent Conversation and Health Intervention

Through the intelligent health robot system, voice collection, compensation and multimodal recognition technology are used to solve the problem of vague accent and expression of the elderly, precise intervention and emotional support for elderly health management are achieved, and the adaptability and accuracy of the health robot are improved.

CN120183386BActive Publication Date: 2025-07-22NANJING MEDICAL UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510661847.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-07-22
Estimated Expiration
2045-05-22

AI Technical Summary

Technical Problem

The prior art is difficult to accurately capture the problems of accents, expressions, and missing words in the elderly, resulting in the inability to effectively evaluate the emotional and psychological state of the elderly and affect the intervention of health.

Method used

The intelligent health robot combining intelligent dialogue and health intervention extracts formant characteristics and pause discrimination through the voice acquisition unit, uses the voice compensation unit to correct missing words and tone, and the multimodal recognition unit outputs voice text and demand intention space. The rehabilitation guidance module generates a rehabilitation strategy intervention plan to realize intelligent dialogue and health intervention.

Benefits of technology

It significantly improves the accuracy and reliability of elderly health management, overcomes language interaction disorders, realizes in-depth analysis of complex physiological and psychological states, provides personalized rehabilitation plans and emotional support, and enhances the adaptability of healthy robots to elderly language disorder groups.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120183386B_ABST
    Figure CN120183386B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of intelligent voice, and particularly relates to an intelligent health robot that combines intelligent conversation and health intervention, including a speech recognition module and a rehabilitation guidance module. The speech recognition module includes a speech acquisition unit, a speech compensation unit, and a multi-modal recognition unit. First, the speech acquisition unit collects the user's speech, compensates for the formant features and pause discrimination through a discriminant filter to obtain the speech feature space to be recognized. Then, the speech compensation unit corrects the speech omission words and tones to obtain the corrected speech feature space to be recognized. The multi-modal recognition unit then outputs the speech text and the demand intention space, and combines with a preset evaluation interval to obtain the intention emotion state space. Finally, the rehabilitation guidance module generates a rehabilitation strategy intervention plan according to the demand intention space and the emotion state space, calls the reinforcement inference strategy library, and outputs it in real-time speech, realizing intelligent conversation and health intervention.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of intelligent voice, and particularly relates to an intelligent health robot that combines intelligent dialogue and health intervention. Background Art

[0002] Intelligent health robots have great potential in assisting the health management of the elderly due to the convenience of voice interaction. However, there are many problems in practical applications. The elderly have strong accents, ambiguous expressions, and are prone to omitting words. Traditional speech recognition and sentiment analysis technologies are difficult to accurately capture key information, resulting in the inability to effectively evaluate the emotional and psychological state of the elderly, and further affecting the intervention of the elderly's health status. Existing technologies, such as the patent with the authorization announcement number CN114141366B, propose a stroke rehabilitation evaluation and auxiliary analysis method based on speech multi-task learning. By constructing a multi-task learning model including a bottom-layer Mel spectrogram deep residual network, a long short-term memory network, and a top-layer fully connected neural network, and combining the mean square error loss function and the cross-entropy loss function weighted superposition method, the speech function impairment of stroke is evaluated. However, this method mainly focuses on the rehabilitation evaluation of specific diseases and does not optimize for the problems commonly existing among the elderly, such as accents, ambiguous expressions, and word omission. It has limitations when actually used for the daily health management of the elderly and extensive speech interaction scenarios. Summary of the Invention

[0003] Aiming at the deficiencies of the existing technology, the present invention proposes an intelligent health robot that combines intelligent dialogue and health intervention, including a speech recognition module and a rehabilitation guidance module. The speech recognition module includes a speech acquisition unit, a speech compensation unit, and a multi-modal recognition unit. First, the speech acquisition unit collects the user's speech, extracts the compensated formant features and pause discrimination through a discriminant filter to obtain the speech feature space to be recognized. Then, the speech compensation unit corrects the speech word omission and tone to obtain the corrected speech feature space to be recognized. The multi-modal recognition unit outputs the speech text and the demand intention space accordingly, and combines the preset evaluation interval to obtain the intention emotion state space. Finally, the rehabilitation guidance module generates a rehabilitation strategy intervention plan according to the demand intention space and the emotion state space, and outputs it in real-time voice, realizing intelligent dialogue and health intervention.

[0004] To achieve the above object, the present invention provides the following technical solutions:

[0005] An intelligent health robot that combines intelligent dialogue and health intervention, including: a speech recognition module, a rehabilitation guidance module; the speech recognition module includes a speech acquisition unit, a speech compensation unit, and a multi-modal recognition unit;

[0006] The user voice information is collected by the voice collection unit, and for the collected voice information, through a preset discriminant filter, the formant feature compensation and pause discrimination of the voice high-frequency attenuation region are performed to obtain the voice feature space to be recognized;

[0007] The voice feature space to be recognized is input into the voice compensation unit, and through a preset word meaning compensation space and tone calibration space, voice missing words and voice part and tone correction are performed to obtain the corrected voice feature space to be recognized;

[0008] Based on the corrected voice feature space to be recognized, the voice text and demand intention space are obtained through the multi-modal recognition unit;

[0009] At the same time, based on the voice text, the corrected voice feature space to be recognized, and the voice feature space to be recognized, combined with a preset voice emotion evaluation interval, the intention emotion state space is obtained;

[0010] Based on the demand intention space and the intention emotion state space, through the reinforcement inference strategy library preset by the rehabilitation guidance module, a rehabilitation strategy intervention plan is obtained for real-time voice output.

[0011] Specifically, the acquisition process of the voice feature space to be recognized includes:

[0012] According to the collected voice information, through the time delay difference estimation algorithm combined with the multi-path discrimination mechanism and a preset abnormal source discrimination threshold, the abnormal discrimination of the collected voice source is performed, and according to the discrimination result, combined with the RANSAC algorithm, the sound source determined to be an abnormal source is eliminated to obtain the target voice state set to be recognized and the corresponding collection distance;

[0013] Based on the target voice state set to be recognized combined with the corresponding collection distance, through the sound energy attenuation principle and the proportion of the sound waves in the direction of the non-abnormal sound source determined, the sound intensity reflection enhancement of the target voice state set to be recognized collected is performed to obtain the enhanced target voice state set to be recognized;

[0014] Based on the enhanced target voice state set to be recognized, through a preset sound frequency - age - pathological feature correlation matrix combined with the corresponding user health parameters, the formant feature of the voice high-frequency attenuation region corresponding to the age and health state in the enhanced target voice state set to be recognized is enhanced and compensated twice through a hierarchical filter;

[0015] At the same time, through a voice pause discrimination model combined with the elderly voice health index, the speech rate fluctuation curve and pause distribution of different age groups, the abnormal pause discrimination of the target voice state set to be recognized after the second enhancement is performed, and the position and duration corresponding to the abnormal pause are marked in the voice frequency band corresponding to the target voice state set to be recognized after the second enhancement to obtain the voice feature space to be recognized.

[0016] Specifically, the construction process of the semantic compensation space and the tone calibration space includes:

[0017] Obtain the standard dialect texts corresponding to different regions, the speech dialect texts recognized by different age groups, and the corresponding Mandarin expressions, and perform corresponding relationship annotation to obtain compensation word order pairs;

[0018] Construct a hierarchical dialect text parsing space according to the compensation word order pairs in combination with the Penn Chinese Treebank;

[0019] The hierarchical dialect text parsing space includes the probability missing character segments corresponding between the standard dialect texts of different regions and the speech dialect texts recognized by different age groups and the corresponding missing expression positions, as well as the probability missing character segments corresponding between the standard dialect texts and the Mandarin expressions and the corresponding missing expression position information.

[0020] Specifically, the construction process of the semantic compensation space and the tone calibration space also includes:

[0021] Based on the missing expression positions and the corresponding probability missing character segments corresponding between the speech dialect texts recognized by different age groups in each region of the hierarchical dialect text parsing space, use the CRF model trained by missing word annotation to judge the part-of-speech of the missing character segments and obtain the part-of-speech discrimination result of the missing character segments;

[0022] According to the judgment result and the probability corresponding to the missing character segment, combine the marked abnormal pause positions and durations through the fine-tuned GraphSAGE model, and perform random walks on the standard dialect text library constructed by the standard dialect text to obtain the hierarchically compensated missing character segments corresponding to the speech text and the corresponding compensation accuracy rate;

[0023] The hierarchically compensated missing character segments include grammatical quantifier compensation, real-word semantic compensation, and emotional expression vocabulary replacement compensation;

[0024] Based on the speech dialect texts recognized by different age groups in each region, the corresponding hierarchically compensated missing character segments, and the corresponding compensation accuracy rate, construct the semantic compensation space.

[0025] Specifically, the construction process of the semantic compensation space and the tone calibration space also includes:

[0026] Obtain the dialect pronunciation rule libraries corresponding to different dialect regions, and establish corresponding initial consonant conversion rules, final conversion rules, and tone mapping rules based on the dialect pronunciation rule libraries and the standard pronunciation rule libraries;

[0027] Establish a triple conversion mapping space between Mandarin pronunciation and the dialect prosody and tones of different age groups and health states in the corresponding regions based on the initial consonant conversion rules, final conversion rules, and tone mapping rules;

[0028] Based on a triple conversion mapping space combined with a reinforcement learning algorithm, by adjusting the speech rate, reference frequency of the input dialect speech and the corresponding Mandarin speech, and the reverberation quantity of different regional dialects, simulation reinforcement training is carried out to obtain a triple conversion mapping space that meets the preset conversion accuracy threshold;

[0029] Based on the dialect speech, the corresponding Mandarin speech and the triple conversion mapping space that meets the preset conversion accuracy threshold, a tone calibration space is obtained.

[0030] Specifically, the process of obtaining the corrected speech feature space to be recognized includes:

[0031] According to the speech feature space to be recognized, through a speech slicing algorithm, according to the positions corresponding to the marked abnormal pauses, speech slicing is carried out and connection annotation is carried out on the slice positions to obtain a speech feature slice space to be recognized;

[0032] According to the speech feature slice space to be recognized, through a semantic compensation space, missing character segments are discriminated between the corresponding speech features internally and between adjacent speech feature segments to be recognized with connection annotation, and the part-of-speech and semantic coordinate positions of the corresponding character segments are obtained;

[0033] According to the part-of-speech and semantic coordinate positions of the corresponding character segments, through the semantic compensation space, corresponding compensation for grammar quantifiers, real-word semantics and emotional expression vocabulary is carried out on the corresponding speech feature slices to be recognized, and the corresponding compensation accuracy is marked at the corresponding compensation position points to obtain a speech feature segment space to be recognized with semantic compensation.

[0034] Specifically, the process of obtaining the corrected speech feature space to be recognized further includes:

[0035] The speech feature segment space to be recognized with semantic compensation is input into the tone calibration space, combined with the region, corresponding age group and health status to which the corresponding speech feature segment belongs, and the corresponding initial consonant conversion rule, final conversion rule and tone mapping rule are called to perform standardized mapping correction on the initial consonants, finals and tones of the characters corresponding to each speech feature segment to be recognized with semantic compensation, and a speech feature space to be recognized with tone calibration is obtained;

[0036] Based on a simulation algorithm combined with a reinforcement learning algorithm, and using the compensation accuracy difference corresponding to semantic compensation and the tone conversion accuracy difference as a training reward function, cyclic simulation training is carried out on the semantic compensation and tone calibration processes to obtain a corrected speech feature space to be recognized.

[0037] Specifically, the process of obtaining the intention emotion state space includes:

[0038] Based on the corrected speech feature space to be recognized, the context-related features of the speech are extracted through the Wav2Vec 2.0 model, and frame-level alignment is performed on the extracted speech segments;

[0039] The aligned speech segments are decoded for acoustic feature sequences, corrected for spelling, and punctuation is predicted through a context-aware speech recognition model to obtain the text information space of the speech conversion;

[0040] Based on the text information space of the speech conversion, intent classification is performed through a Chinese pre-trained intent decomposition model combined with a medical knowledge graph configured with temporal reasoning, and entity extraction is performed on the corresponding key intent content words and emotion expression vocabulary to obtain the intent classification result and the intent keyword space.

[0041] Specifically, the process of obtaining the intent emotion state space further includes:

[0042] Based on the intent classification result, the intent keyword space, the corrected speech feature space to be recognized, and the pitch perturbation, amplitude perturbation, and harmonic-to-noise ratio in the speech feature space to be recognized, the intonation contour slope and energy envelope variance obtained through the wavelet algorithm are used to construct an emotion analysis feature space;

[0043] Based on the pre-trained multi-modal emotion recognition model combined with the preset emotion recognition types and the preset health level-disease type-emotion type association matrix, a multi-modal emotion intent recognition model is constructed;

[0044] The emotion analysis feature space is input into the multi-modal emotion intent recognition model to obtain the intent emotion type and level, and based on the intent emotion type and level, combined with the change trend of the emotion state constructed from the corresponding user historical emotion data, potential psychological problem features are obtained;

[0045] Based on the intent emotion type and level and the corresponding potential psychological problem features, the intent emotion state space is obtained.

[0046] Specifically, the construction process of the voice frequency-age-pathological feature association matrix includes:

[0047] Based on the dialect types corresponding to different regions, the speech information of 500 users of different ages, health status levels, and disease types in the corresponding regions is obtained, and the voice frequency and fundamental frequency fluctuation characteristics of the corresponding users are obtained through a voice preprocessing algorithm;

[0048] Based on the voice frequency, fundamental frequency fluctuation characteristics, age, health status level, and disease type information of the corresponding users, the correlation degrees between any two of the voice frequency, age, health status level, and disease type are obtained through a correlation algorithm;

[0049] Based on the correlation degrees between voice frequency, age, health status level, and disease type pairwise, a voice frequency - age - pathological feature correlation matrix is obtained.

[0050] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0051] In view of the deficiencies of the prior art, the present invention constructs a multi - level linked intelligent voice processing and health intervention system, effectively overcoming the language interaction barriers in the health management of the elderly population. In particular, based on the high - frequency attenuation formant feature extraction and dynamic pause discrimination technology, the acoustic characteristics of pathological speech are accurately captured, significantly improving the feature representation ability under the condition of fuzzy pronunciation; combined with the context - aware semantic compensation and tone calibration mechanism, the system reconstructs the incomplete semantics and corrects the dialect deviation to ensure the integrity and accuracy of key health information; through the multi - modal emotion intention recognition model that fuses speech text, acoustic features, and historical interaction data, it breaks through the limitations of traditional single - modal analysis and realizes the in - depth analysis of complex physiological and psychological states; relying on the dynamic decision - making architecture of the enhanced reasoning strategy library, it maps the accurately identified health needs and emotional states to personalized rehabilitation plans, forming a closed - loop management from voice interaction to health intervention. This process greatly enhances the adaptability of health robots to the elderly with language disorders and provides a reliable guarantee for timely and accurate emotional support and health monitoring. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 It is a module diagram of an intelligent health robot combining intelligent dialogue and health intervention according to an embodiment of the present invention;

[0053] Figure 2 It is a working flow chart of an intelligent health robot combining intelligent dialogue and health intervention according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] Embodiment

[0055] Please refer to Figure 1 and Figure 2 An embodiment provided by the present invention: An intelligent health robot combining intelligent dialogue and health intervention includes: a speech recognition module, a rehabilitation guidance module; the speech recognition module includes a speech acquisition unit, a speech compensation unit, and a multi - modal recognition unit;

[0056] The speech recognition module is used for the acquisition and compensated recognition of speech information to obtain a speech text and a demand intention space;

[0057] The speech acquisition unit is used for the acquisition of speech information, and performs formant feature extraction and pause discrimination on the high - frequency attenuation region of the speech through a preset discriminant filter to obtain a speech feature space to be recognized;

[0058] The voice compensation unit is used to perform voice missing word and voice part and tone correction on the voice feature space to be recognized through a preset word meaning compensation space and tone calibration space, so as to obtain a corrected voice feature space to be recognized;

[0059] The multi-modal recognition unit is used to obtain a voice text and a demand intention space according to the corrected voice feature space to be recognized through a configured multi-modal semantic-emotional recognition model, and obtain an intention emotional state space according to the voice text, the corrected voice feature space to be recognized, and the voice feature space to be recognized;

[0060] The rehabilitation guidance module is used to obtain a rehabilitation strategy intervention plan according to the intention emotional state space and the demand intention space combined with a reinforcement inference strategy library configured with a reinforcement self-learning model, and perform real-time voice output.

[0061] Further, the specific working processes corresponding to the voice collection unit, the voice compensation unit, the multi-modal recognition unit, and the rehabilitation guidance module in this embodiment include:

[0062] The user voice information is collected through the voice collection unit, and the collected voice information is subjected to formant feature compensation and pause discrimination in the voice high-frequency attenuation region through a preset discriminant filter to obtain a voice feature space to be recognized;

[0063] Further, the acquisition process of the voice feature space to be recognized in this embodiment includes:

[0064] According to the collected voice information, the voice source of the collected voice is abnormally discriminated through a time delay difference estimation algorithm combined with a multi-path discrimination mechanism and a preset abnormal source discrimination threshold, and the sound source determined to be an abnormal source is removed according to the discrimination result combined with the RANSAC algorithm, so as to obtain a target voice state set to be recognized and the corresponding collection distance;

[0065] Based on the target voice state set to be recognized combined with the corresponding collection distance, the sound intensity reflection enhancement of the collected target voice state set is performed through the sound energy attenuation principle and the proportion of the sound wave in the direction of the non-abnormal sound source determined, so as to obtain an enhanced target voice state set;

[0066] Further, the exemplary process corresponding to the enhanced target voice state set in this embodiment includes:

[0067] Instance scenario description:

[0068] Scenario: An elderly user is talking to a smart health robot in the kitchen, with the background noise of a range hood and the sound of a TV.

[0069] Goal: Extract the user's voice from the mixed sound sources and enhance its clarity.

[0070] Specific implementation steps:

[0071] (1)Voice source discrimination:

[0072] The calculation process of time delay difference estimation (TDOA) includes:

[0073] Collect voice signals using a 4-channel circular microphone array;

[0074] Utilize the configured TDOA to calculate the time difference for each sound source to reach different microphones;

[0075] The multi-path discrimination process includes:

[0076] Based on the time delay difference estimation result, combined with a preset threshold (such as 0.5ms), discriminate the abnormal source;

[0077] Example: The time delay difference of the range hood noise > 1ms, determined as an abnormal source;

[0078] The process of removing abnormal sound sources by the RANSAC algorithm includes:

[0079] Use the RANSAC algorithm to fit the normal sound source direction and remove the abnormal points deviating from the fitting curve;

[0080] Example: The user's voice direction is 45°, remove the abnormal sound sources (such as the TV sound) deviating from ±10°;

[0081] (2)Sound intensity reflection enhancement:

[0082] The process of calculating the sound energy attenuation includes:

[0083] According to the sound energy attenuation formula constructed based on the spherical wave propagation principle, calculate the attenuation coefficient;

[0084] Example: At a distance of 1.2 meters, the attenuation coefficient = 1.44;

[0085] The process of enhancing non-abnormal sound sources includes:

[0086] Perform weighted enhancement on the sound waves in the direction of the non-abnormal sound source;

[0087] Example: Enhance the voice signal in the 45° direction and suppress the noise in other directions;

[0088] The process of intensity reflection includes:

[0089] Perform gain compensation on the voice signal according to the attenuation coefficient;

[0090] Example: Gain the user's voice signal by 1.44 times.

[0091] Based on the enhanced set of speech states to be recognized for the target, combined with the corresponding user health parameters through a preset voice frequency - age - pathological feature correlation matrix, a hierarchical filter is used to perform secondary enhancement compensation on the formant features in the voice high - frequency attenuation region corresponding to the age and health status in the enhanced set of speech states to be recognized for the target;

[0092] Furthermore, the construction process of the voice frequency - age - pathological feature correlation matrix in this embodiment includes:

[0093] Based on the dialect types corresponding to different regions, obtain the voice information of 500 users corresponding to different age groups, health status levels, and disease types within the corresponding regions, and obtain the voice frequency and fundamental frequency fluctuation characteristics of the corresponding users through a voice pre - processing algorithm;

[0094] Based on the voice frequency, fundamental frequency fluctuation characteristics, age, health status level, and disease type information of the corresponding users, through a correlation algorithm, obtain the correlation degrees between any two of the voice frequency, age, health status level, and disease type;

[0095] Based on the correlation degrees between any two of the voice frequency, age, health status level, and disease type, obtain the voice frequency - age - pathological feature correlation matrix.

[0096] Furthermore, the specific implementation process of constructing the voice frequency - age - pathological feature correlation matrix in this embodiment includes:

[0097] (1) Data collection:

[0098] Collect the voice data of 500 users in different dialect areas (such as the Wu dialect area and the Cantonese dialect area);

[0099] Multi - dimensional annotation: Annotate the age, health status (such as diabetes, Parkinson's disease), and disease stage (such as Hoehn - Yahr stage) of the users;

[0100] Example:

[0101] Collect the voice data of 500 diabetes patients over 70 years old in the Wu dialect area, and annotate the blood glucose control level (such as HbA1c value);

[0102] (2) Feature extraction:

[0103] Voice pre - processing:

[0104] Use a pre - processing algorithm to extract features such as the fundamental frequency (F0), formants (F1 - F3), and harmonic - to - noise ratio (HNR);

[0105] Example: The 4 - 6Hz tremor feature of Parkinson's patients;

[0106] Pathological feature extraction:

[0107] Extract acoustic features related to diseases (such as the vocal tremor frequency in Parkinson's disease and the voice energy attenuation in diabetes).

[0108] Example: The voice energy of diabetic patients significantly attenuates above 3000Hz.

[0109] (3)Correlation analysis:

[0110] Pearson correlation coefficient calculation:

[0111] Calculate the correlation between voice frequency and age, health status;

[0112] Example: The correlation coefficient between the fundamental frequency fluctuation of users over 70 years old and the degree of anxiety = 0.82.

[0113] Construct a voice frequency - age - pathological feature correlation matrix to store the correlation between each pair. Exemplarily, the following is a specific example of diabetes and Parkinson's disease in the voice frequency - age - pathological feature correlation matrix, specifically:

[0114] {

[0115] "Over 70 years old": {

[0116] "Diabetes": {"Fundamental frequency fluctuation": 0.75, "Formant dispersion": 0.68},

[0117] "Parkinson's disease": {"Tremor frequency": 0.82, "Harmonic - to - noise ratio": 0.70}

[0118] }

[0119] };

[0120] (4)Results:

[0121] Voice frequency - age - pathological feature correlation matrix: Used to guide the design of hierarchical filters.

[0122] Exemplarily, the hierarchical filter configuration process in this embodiment includes:

[0123] (1)Age stratification:

[0124] For users over 55 years old, set high - density Gammatone filters in the 200 - 400Hz frequency band.

[0125] Example: One filter every 25Hz to enhance pathological features (such as tremor signals).

[0126] (2)Pathological stratification:

[0127] For Parkinson's patients, increase the filter density in the 4 - 6Hz tremor frequency band.

[0128] Example: Set the filter density three times the normal value in the 4 - 6 Hz frequency band.

[0129] (3) Secondary enhancement compensation:

[0130] Energy compensation:

[0131] Perform energy compensation on the target frequency band to improve speech clarity.

[0132] Example: Enhance the energy of the 200 - 400 Hz frequency band by 3 dB.

[0133] (4) Dynamic adjustment:

[0134] Dynamically adjust the filter parameters according to the user's health status.

[0135] Example: When it is detected that the user's blood sugar control is poor, enhance the energy compensation in the frequency band above 3000 Hz.

[0136] Meanwhile, through a speech pause discrimination model that combines the elderly speech health index, the speech rate fluctuation curve and pause distribution of different age groups, perform abnormal pause discrimination on the target set of speech states to be recognized after secondary enhancement, and mark the positions and durations corresponding to the abnormal pauses in the speech frequency band corresponding to the target set of speech states to be recognized after secondary enhancement, to obtain the speech feature space to be recognized.

[0137] Furthermore, the acquisition process corresponding to the abnormal pause discrimination and the elderly speech health index in this embodiment includes:

[0138] Exemplarily, the example process of pause detection in this embodiment includes:

[0139] First, it is necessary to obtain the speech samples of the elderly through professional speech acquisition equipment. During the acquisition process, ensure that the environment is quiet and reduce external noise interference to obtain high - quality speech data;

[0140] Feature extraction: Pre - process the collected speech samples, remove interference signals such as noise, and then extract features such as Jitter, Shimmer, and pause duration;

[0141] Furthermore, in this embodiment, Jitter is analyzed by analyzing the pitch period of the speech signal, calculating the difference between adjacent pitch periods, and then statistically calculating statistics such as the average value and standard deviation of these differences as the quantization index of Jitter;

[0142] For the calculation of Shimmer, it is necessary to analyze the amplitude change of the speech signal; first extract the amplitude of the speech signal within each pitch period, and then calculate the relative change amount of the adjacent period amplitudes as the quantization index of Shimmer;

[0143] Furthermore, in this embodiment, Jitter, which is "pitch period jitter" in Chinese, refers to the degree of difference between adjacent pitch periods in a speech signal; it reflects the stability of vocal cord vibration. The larger the pitch period jitter, the more unstable the vocal cord vibration, and the worse the speech quality may be. For example, in some speech disorders caused by laryngeal diseases or neurological problems, the Jitter value will increase significantly.

[0144] Shimmer, which is "amplitude perturbation" in Chinese, represents the variation of the amplitudes of speech signals in adjacent pitch periods. It measures the stability of the sound intensity generated during vocal cord vibration. The higher the Shimmer value, the greater the amplitude fluctuation of the speech signal, which may also reflect abnormalities in speech production. For example, under the influence of some respiratory function disorders or vocal cord lesions, the Shimmer value will deviate from the normal range.

[0145] LSTM model prediction: Use the LSTM model to predict the pause type (natural / pathological).

[0146] Example: A 1.8-second pause is detected and determined to be a pathological pause.

[0147] Exemplarily, the process of marking abnormal pause examples in this embodiment includes:

[0148] Position and duration marking:

[0149] Mark the position and duration of abnormal pauses in the speech frequency band. Specifically, by analyzing the time axis of the speech signal and combining the start and end times of the pause, determine the position of the abnormal pause on the time axis. For example, after detection, it is found that the abnormal pause starts at 2.5 seconds and ends at 4.3 seconds.

[0150] Example: Mark the time axis from 2.5 to 4.3 seconds as an abnormal pause.

[0151] Comprehensively calculate the extracted Jitter, Shimmer, and pause duration to obtain the elderly speech health index.

[0152] Input the to-be-recognized speech feature space into the speech compensation unit, and through the preset semantic compensation space and tone calibration space, perform speech missing word and speech voice part and tone correction to obtain the corrected to-be-recognized speech feature space;

[0153] Furthermore, the construction process of the semantic compensation space and tone calibration space in this embodiment includes:

[0154] Obtain the standard dialect texts corresponding to different regions, the speech dialect texts recognized by different age groups, and the corresponding Mandarin expression texts, and perform corresponding relationship annotation to obtain compensation word order pairs;

[0155] Construct a hierarchical dialect text parsing space according to the compensatory word order in combination with the Penn Chinese Treebank;

[0156] The hierarchical dialect text parsing space includes the probability missing character segments corresponding between the standard dialect texts in different regions and the spoken dialect texts recognized by different age groups, as well as the corresponding missing expression positions, and the probability missing character segments corresponding between the standard dialect text and the Mandarin expression text, as well as the corresponding missing expression position information;

[0157] Based on the missing expression positions corresponding between the spoken dialect texts recognized by different age groups in each region of the hierarchical dialect text parsing space and the corresponding probability missing character segments, use the CRF model trained by missing word annotation to judge the part-of-speech of the missing character segments, and obtain the part-of-speech discrimination results of the missing character segments;

[0158] According to the judgment result and the probability corresponding to the missing character segment, combine the marked abnormal pause positions and durations through the fine-tuned GraphSAGE model, and perform random walks on the standard dialect text library constructed by the standard dialect text to obtain the hierarchically compensated missing character segments corresponding to the speech text and the corresponding compensation accuracy rate;

[0159] The hierarchically compensated missing character segments include grammatical quantifier compensation, real-word semantic compensation, and emotional expression vocabulary replacement compensation;

[0160] Furthermore, in this embodiment, a three-level compensation mechanism is constructed according to the above-mentioned grammatical quantifier compensation, real-word semantic compensation, and emotional expression vocabulary replacement compensation, as shown in Table 1 specifically;

[0161] Table 1 Three-level compensation mechanism

[0162]

[0163] As shown in Table 1 above, it includes compensation levels, trigger conditions, candidate word sources, and examples, where L1, L2, and L3 respectively correspond to the missing compensation for grammatical structures, the missing of entity associations, and the replacement compensation for fuzzy emotional expressions.

[0164] Furthermore, the construction process of the high-frequency missing word library in this embodiment includes:

[0165] Collect 1000 hours of elderly health dialogue voices of different age groups, health levels, and regions, and perform common missing word annotations (such as quantifiers, conjunctions). At the same time, analyze the annotated voices to obtain the missing word patterns of patients of different age groups, health levels, and regions and the frequencies of the corresponding words being omitted or left out. Mark the words with an omission or omission frequency greater than 5 times as high-frequency missing words, and obtain the high-frequency missing word library based on the obtained high-frequency missing words.

[0166] Further, the process of obtaining the co-occurrence thesaurus + medical knowledge graph in Table 1 of this embodiment is specifically as follows:

[0167] Extract the association relationships of symptoms - drugs - examination items from databases such as PubMed, interface with the electronic medical record system of community hospitals, extract patients' chief complaints and diagnosis results, and collect the disease history and medication records in users' health archives to obtain medical graph information;

[0168] Based on the medical graph information, use the BiLSTM-CRF model combined with the N-gram model to perform medical entity extraction and entity co-occurrence frequency statistics; for example, the original text: Recently, high blood pressure and dizziness; extraction result: blood pressure (symptom), dizziness (symptom); extraction result: blood pressure - dizziness (frequency = 0.75);

[0169] Fuse the co-occurrence relationships, co-occurrence frequency statistics with medical literature data, and construct a medical knowledge graph through graph algorithms.

[0170] Construct a semantic compensation space based on the speech dialect texts identified for different age groups in each region, the corresponding hierarchical compensated missing character segments, and the corresponding compensation accuracy rates;

[0171] Obtain the dialect pronunciation rule libraries corresponding to different dialect regions, and establish the corresponding initial consonant conversion rules, final consonant conversion rules, and tone mapping rules based on the dialect pronunciation rule libraries and the standard pronunciation rule libraries;

[0172] Further, the initial consonant conversion rules, final consonant conversion rules, and tone mapping rules in this embodiment are specifically shown in Table 2:

[0173] Table 2 Pronunciation Rule Knowledge Base

[0174]

[0175] Further, as shown in Table 2 above, in this embodiment, only the pronunciation rule knowledge bases corresponding to the language systems in the Wu dialect, Cantonese, and Sichuan-Chongqing regions are constructed here. In the actual application process, according to the specific region where the corresponding user is located and the differences between the specific dialect phonology and tones used and Mandarin, those skilled in the art can construct and add them to the above pronunciation rule knowledge bases according to the specific actual situation.

[0176] Establish a triple conversion mapping space between Mandarin pronunciation and the dialect phonology and tones of different age groups and physical health states in the corresponding regions based on the initial consonant conversion rules, final consonant conversion rules, and tone mapping rules;

[0177] Based on the triple transformation mapping space combined with the reinforcement learning algorithm, through adjusting the speech rate, reference frequency, and reverberation quantity of different regional dialects of the input dialect speech and the corresponding Mandarin speech, perform simulated reinforcement training to obtain a triple transformation mapping space that meets the preset transformation accuracy threshold;

[0178] Based on the dialect speech, the corresponding Mandarin speech, and the triple transformation mapping space that meets the preset transformation accuracy threshold, obtain the tone calibration space.

[0179] Exemplarily, the specific implementation process of constructing the semantic compensation space in this embodiment includes:

[0180] Principle:

[0181] Based on the hierarchical dialect text parsing space, use the CRF model and the GraphSAGE model to predict the missing character segments.

[0182] Implementation process:

[0183] Prediction of missing character segments:

[0184] Use the CRF model to predict the part-of-speech of the missing characters (such as quantifiers, entity nouns).

[0185] Example: Predict that the missing character in "buy __ apples" is "piece".

[0186] Semantic reasoning:

[0187] Use the GraphSAGE model to perform random walks on the medical knowledge graph to predict the missing entities.

[0188] Example: Predict that the missing entity in "recently __ not feeling well" is "blood pressure".

[0189] Result:

[0190] Semantic compensation space: Contains compensation rules for grammatical quantifiers, real-word semantics, and emotional expression vocabulary.

[0191] Exemplarily, the process of constructing the tone calibration space in this embodiment includes:

[0192] Principle:

[0193] Based on the dialect pronunciation rule library, establish conversion rules for initials, finals, and tones.

[0194] Implementation process:

[0195] Rule mapping:

[0196] Construct initial conversion rules (such as "zh" → "z"), final conversion rules (such as "ian" → "i").

[0197] Example: Convert "zhī" of users in the Wu dialect area to "zī".

[0198] Reinforcement training:

[0199] Use the reinforcement learning algorithm to adjust the speech rate, fundamental frequency, and reverberation quantity, and optimize the conversion accuracy rate.

[0200] Tone calibration space: Includes the standardized mapping rules of initials, finals, and tones.

[0201] Furthermore, the acquisition process of the corrected speech feature space to be recognized in this embodiment includes:

[0202] According to the speech feature space to be recognized, use the speech slicing algorithm to perform speech slicing and connect and label the slice positions according to the positions corresponding to the marked abnormal pauses, and obtain the speech feature slice space to be recognized;

[0203] According to the speech feature slice space to be recognized, use the semantic compensation space to discriminate the missing character segments between the corresponding speech features internally and between adjacent speech feature segments to be recognized with connection labels, and obtain the part-of-speech and semantic coordinate positions of the corresponding character segments;

[0204] According to the part-of-speech and semantic coordinate positions of the corresponding character segments, use the semantic compensation space to perform corresponding compensation for the grammatical quantifiers, real-word semantics, and emotional expression vocabulary of the corresponding speech feature slices to be recognized, and mark the corresponding compensation accuracy rate at the corresponding compensation position points, and obtain the speech feature segment space to be recognized with semantic compensation;

[0205] Input the speech feature segment space to be recognized with semantic compensation into the tone calibration space, combine the region to which the corresponding speech feature segment to be recognized belongs with the corresponding age group and health status, and call the corresponding initial conversion rules, final conversion rules, and tone mapping rules to perform standardized mapping correction on the initials, finals, and tones of the characters corresponding to each speech feature segment to be recognized with semantic compensation, and obtain the speech feature space to be recognized after tone calibration;

[0206] Based on the simulation algorithm combined with the reinforcement learning algorithm, and using the difference in compensation accuracy rate corresponding to semantic compensation and the difference in tone conversion accuracy rate as the training reward function, perform cyclic simulation training on the semantic compensation process and the tone calibration process, and obtain the corrected speech feature space to be recognized.

[0207] Based on the corrected speech feature space to be recognized, obtain the speech text and demand intention space through the multi-modal recognition unit;

[0208] At the same time, based on the speech text, the corrected speech feature space to be recognized, and the speech feature space to be recognized, combined with the preset speech emotion evaluation interval, obtain the intention emotion state space;

[0209] Furthermore, the process of obtaining the intended emotional state space in this embodiment includes:

[0210] Based on the corrected speech feature space to be recognized, extract the context-related features of the speech through the Wav2Vec 2.0 model, and perform frame-level alignment on the extracted speech segments;

[0211] Perform acoustic feature sequence decoding, spelling correction, and punctuation prediction on the aligned speech segments through a context-aware speech recognition model to obtain the text information space of the speech conversion;

[0212] Based on the text information space of the speech conversion, perform intent classification through a Chinese pre-trained intent decomposition model combined with a medical knowledge graph configured with temporal reasoning, and extract entities for the corresponding key intent real words and emotional expression words to obtain the intent classification result and the intent keyword space.

[0213] Specifically, the process of obtaining the intended emotional state space further includes:

[0214] Based on the intent classification result, the intent keyword space, the corrected speech feature space to be recognized, and the pitch perturbation, amplitude perturbation, and harmonic-to-noise ratio in the speech feature space to be recognized, construct an emotional analysis feature space through the wavelet algorithm to obtain the intonation contour slope and energy envelope variance;

[0215] Based on the pre-trained multi-modal emotion recognition model, combined with the preset emotion recognition type and the preset health level-disease type-emotion type association matrix, construct a multi-modal emotion intent recognition model;

[0216] Input the emotional analysis feature space into the multi-modal emotion intent recognition model to obtain the intended emotion type and level, and based on the intended emotion type and level, combined with the long-term change trend of the emotional state constructed from the corresponding user historical emotional data, obtain the potential psychological problem features;

[0217] Based on the intended emotion type and level and the corresponding potential psychological problem features, obtain the intended emotional state space.

[0218] Furthermore, the health level stratification in this embodiment is specifically:

[0219] Divide the health level according to the user's health record (such as healthy, sub-healthy, diseased).

[0220] Furthermore, the disease type adaptation in this embodiment:

[0221] Set specific emotional assessment intervals for different disease types (such as Parkinson's disease, diabetes).

[0222] Exemplarily, in this embodiment, the acquisition process of the above-mentioned intended emotional state space is illustrated by the following example:

[0223] Implementation case:

[0224] Input:

[0225] Corrected speech feature space: "My blood sugar has been high these days and I feel dizzy."

[0226] User profile: 70 years old, from the Wu dialect area, diabetes, health level: disease.

[0227] Processing process:

[0228] Speech text recognition: Generate the text "My blood sugar has been high these days and I feel dizzy."

[0229] Requirement intention analysis:

[0230] Intention classification: Health consultation.

[0231] Entity extraction: Symptom = "dizziness", Time = "these days".

[0232] Requirement reasoning: {"Measure blood sugar", "Consult the cause of dizziness", "Adjust medication"}.

[0233] Emotional state classification:

[0234] Acoustic feature analysis: Jitter = 0.08, Shimmer = 0.12.

[0235] Text emotional analysis: Negative score = 0.85.

[0236] Multimodal fusion: Generate emotional feature vectors.

[0237] Emotional classification: Anxiety = 0.75, Depression = 0.15.

[0238] Emotional polarity judgment: Negative (anxiety, depression).

[0239] Output:

[0240] Speech text and requirement intention space: {"Measure blood sugar", "Consult the cause of dizziness", "Adjust medication"}.

[0241] Intended emotional state space: {"Anxiety": 0.75, "Depression": 0.15, "Calm": 0.10}.

[0242] Based on the said requirement intention space and intended emotional state space, through the reinforcement reasoning strategy library preset by the rehabilitation guidance module, a rehabilitation strategy intervention plan is obtained for real-time voice output.

[0243] Furthermore, the steps for constructing the enhanced reasoning strategy library in this embodiment include:

[0244] Step 1: Knowledge collection and organization, specifically including:

[0245] Data sources:

[0246] Medical literature: Collect a large number of authoritative literature related to rehabilitation medicine, including rehabilitation treatment guidelines for various diseases, clinical research results, etc.;

[0247] Expert experience: Obtain the rehabilitation treatment strategies and experience accumulated by experts in long-term clinical practice. For example, obtain the ideas and methods of senior rehabilitation physicians for formulating personalized rehabilitation plans for patients of different ages and different severities of illness;

[0248] Clinical case data: Organize the actual case data of the hospital's rehabilitation department, including patients' basic information, disease diagnosis, rehabilitation process records, rehabilitation effect evaluation, etc. For example, collect the rehabilitation treatment data of 1000 knee injury patients in a hospital within one year and analyze the relationship between different treatment plans and rehabilitation effects.

[0249] Step 2: Knowledge representation and encoding, specifically including:

[0250] Variables:

[0251] K: Represents the knowledge set, and the collected knowledge is represented and encoded in a structured manner so that the computer can understand and process it.

[0252] Knowledge units: Decompose knowledge into small knowledge units. For example, the rehabilitation knowledge unit for hypertensive patients can include diet control knowledge (such as daily salt intake not exceeding 6 grams), exercise advice knowledge (such as 150 minutes of moderate-intensity aerobic exercise per week), drug treatment knowledge (types and dosages of antihypertensive drugs taken on time), etc.

[0253] Encoding methods:

[0254] Production rules: Encode knowledge in the form of production rules, that is, "if (condition), then (action)". For example, "if a patient has diabetes and poor blood sugar control, then it is recommended to adjust the diet structure, increase dietary fiber intake, and reduce the intake of high-sugar foods".

[0255] Semantic network: Construct a semantic network to represent the relationships between knowledge, such as the associations between diseases and symptoms, rehabilitation measures and diseases, and rehabilitation effects and rehabilitation measures. For example, connect "stroke" with "limb hemiplegia", "rehabilitation training (physical therapy, occupational therapy)", etc. through semantic relationships.

[0256] Step 3: Inference engine construction, specifically including:

[0257] Variable:

[0258] I: An inference engine that retrieves and infers a suitable rehabilitation strategy intervention plan from a reinforcement inference strategy library according to the input demand intention space and intention emotion state space.

[0259] Input information: Includes the user's demand intention (such as "I want to relieve low back pain"), emotion state (such as "anxiety"), and basic health information (such as age, gender, medical history, etc.).

[0260] Inference algorithm:

[0261] Forward inference: Starting from the information input by the user, according to the rules and knowledge in the reinforcement inference strategy library, gradually deduce possible rehabilitation strategies. For example, if the user inputs "I am a 60-year-old female with knee pain and the pain has recently worsened", the inference engine, based on the rehabilitation knowledge about knee pain in the knowledge base and combined with the patient's age and gender information, deduces possible rehabilitation strategies such as "It is recommended to apply hot compress to relieve pain and perform knee joint rehabilitation training 3 times a week, including knee joint flexion and extension exercises and quadriceps femoris exercises".

[0262] Backward inference: First assume a rehabilitation goal, and then search in the knowledge base for the conditions and strategies that can achieve this goal. For example, assume the goal is "to improve the patient's self-care ability", and the inference engine searches in the knowledge base for rehabilitation measures for different diseases and patient conditions to determine how to achieve this goal.

[0263] Step 4: Strategy evaluation and update, specifically including:

[0264] Evaluation indicators:

[0265] Rehabilitation effect: By tracking and evaluating the patients who actually apply the rehabilitation strategy intervention plan, observe their rehabilitation effects, such as the degree of relief of disease symptoms, the recovery of physical functions, etc. For example, for fracture patients, evaluate indicators such as fracture healing time and the degree of recovery of limb movement range.

[0266] User satisfaction: Collect the feedback and satisfaction evaluation of patients on the rehabilitation strategy, and understand the acceptance degree and subjective feelings of patients towards the rehabilitation plan. For example, through questionnaires or patient interviews, ask patients about their satisfaction with aspects such as the intensity of rehabilitation training and the rationality of diet suggestions.

[0267] Update mechanism:

[0268] According to the evaluation results, update and optimize the strategies in the enhanced reasoning strategy library. If it is found that a certain rehabilitation strategy has poor effects in actual application or low patient satisfaction, analyze the reasons and adjust the corresponding knowledge and rules. For example, if it is found that a certain type of rehabilitation training program is likely to cause secondary injuries to patients in a specific age group, modify the applicable conditions of the program in the knowledge base or adjust the training intensity and methods. At the same time, continuously incorporate new medical research results, expert experience, and clinical case data into the enhanced reasoning strategy library to maintain the timeliness and accuracy of its knowledge.

[0269] Exemplarily, the generation process of the rehabilitation strategy intervention plan in this embodiment includes:

[0270] Principle:

[0271] Generate personalized rehabilitation strategies according to the demand intention space and the intention emotion state space.

[0272] Implementation process:

[0273] Strategy matching:

[0274] Match rehabilitation strategies according to the demand intention (such as "measuring blood sugar") and the emotion state (such as "anxiety").

[0275] Example: Generate a strategy of "measuring blood sugar immediately and playing soothing music".

[0276] Dynamic adjustment:

[0277] Dynamically adjust the strategy parameters according to user feedback and historical data.

[0278] Example: If the user feedback is "the music is ineffective", then adjust it to "deep breathing training".

[0279] Result:

[0280] Rehabilitation strategy intervention plan: {"measuring blood sugar", "playing soothing music", "deep breathing training"}.

[0281] Through multimodal speech processing and intelligent compensation technology, this embodiment significantly improves the accuracy and reliability of elderly health management, providing a more intelligent and personalized health management solution for the elderly population. In particular, for the unique problems of mixed accents, vague expressions, and missing words among the elderly, the system uses high-frequency attenuation formant feature extraction and dynamic semantic compensation mechanisms, combined with a dual-channel analysis model of acoustic features and text emotions, to effectively overcome the emotional misjudgment caused by interrupted pronunciation and achieve a closed-loop management from voice interaction to health intervention. Specifically, the voice acquisition unit accurately captures the target voice and suppresses environmental noise through a multi-microphone beamforming array and the RANSAC algorithm. Combining the principle of sound energy attenuation and hierarchical filters, it enhances the pathological voice features (such as the 4-6Hz tremor signal of Parkinson's patients), significantly improving the voice clarity and feature representation ability. For example, in a kitchen environment, the system can effectively eliminate the noise of the range hood, accurately locate the direction of the user's voice, and compensate for the voice intensity attenuation caused by distance through the reverse enhancement mapping of sound energy attenuation, increasing the signal-to-noise ratio of the voice signal by more than 10dB. The voice compensation unit realizes a three-level compensation mechanism of grammar quantifier compensation, real-word semantic compensation, and emotional expression vocabulary replacement compensation by constructing a high-frequency missing word library, a co-occurrence word library + medical knowledge graph, and an emotion dictionary + historical dialogue analysis, ensuring the integrity and accuracy of key health information. For example, when the user's expression is vague (such as "I've been... that... always thirsty..."), the system can automatically correct the voice semantics and tone or complete the missing words according to the context and medical knowledge graph (such as "My blood sugar has been high these days"), deleting pauses, useless words, and correcting tone errors (such as correcting "thirsty" to "high"), and inserting degree adverbs in combination with the emotion dictionary (such as "I'm

especially

[0282] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation modes, which are merely illustrative rather than restrictive. Under the guidance of the present invention, ordinary technicians in the field may also change, modify, replace and modify the above-mentioned embodiments without departing from the scope of protection of the purpose of the present invention and the claims, and all of these are within the protection of the present invention.

[0283] If the disclosed technical solution involves personal information, the product using the disclosed technical solution has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing the personal information. If the disclosed technical solution involves sensitive personal information, the product using the disclosed technical solution has obtained the individual's separate consent before processing the sensitive personal information, and at the same time meets the "explicit consent" requirement. For example, on personal information collection devices such as cameras, clear and prominent signs are set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that he or she agrees to the collection of his or her personal information; or on the device that processes personal information, when the personal information processing rules are notified by obvious signs / information, the individual's authorization is obtained through pop-up information or by asking the individual to upload his or her personal information; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.

Claims

1. An intelligent health robot integrating intelligent conversation and health intervention, characterized in that, It includes a speech recognition module and a rehabilitation guidance module; the speech recognition module includes a speech acquisition unit, a speech compensation unit, and a multi-modal recognition unit; The user's speech information is collected through the speech acquisition unit, and for the collected speech information, through a preset discriminant filter, the formant feature compensation and pause discrimination of the speech high-frequency attenuation region are performed to obtain a speech feature space to be recognized; The speech feature space to be recognized is input into the speech compensation unit, and through a preset word meaning compensation space and tone calibration space, speech missing words and speech voice part and tone correction are performed to obtain a corrected speech feature space to be recognized; Based on the corrected speech feature space to be recognized, the speech text and demand intention space are obtained through the multi-modal recognition unit; At the same time, based on the speech text, the corrected speech feature space to be recognized, and the speech feature space to be recognized, combined with a preset speech emotion evaluation interval, an intention emotion state space is obtained; Based on the demand intention space and the intention emotion state space, through the enhanced inference strategy library preset by the rehabilitation guidance module, a rehabilitation strategy intervention plan is obtained for real-time speech output.

2. The intelligent health robot integrating intelligent conversation and health intervention according to claim 1, characterized in that, The acquisition process of the speech feature space to be recognized includes: According to the collected speech information, through the time delay difference estimation algorithm combined with the multi-path discrimination mechanism and a preset abnormal source discrimination threshold, the abnormal discrimination of the collected speech source is performed, and according to the discrimination result, combined with the RANSAC algorithm, the sound source determined to be an abnormal source is removed to obtain a target speech state set to be recognized and the corresponding acquisition distance; Based on the target speech state set to be recognized combined with the corresponding acquisition distance, through the sound energy attenuation principle and the proportion of sound waves in the direction of the non-abnormal sound source determined, the sound intensity reflection enhancement of the target speech state set to be recognized is performed to obtain an enhanced target speech state set to be recognized; Based on the enhanced target speech state set to be recognized, through a preset sound frequency-age-pathological feature correlation matrix combined with the corresponding user health parameters, the formant feature of the speech high-frequency attenuation region corresponding to the age and health status in the enhanced target speech state set is secondarily enhanced and compensated through a hierarchical filter; At the same time, through a speech pause discrimination model that combines the elderly speech health index, the speech speed fluctuation curve and pause distribution of different age groups, the abnormal pause discrimination of the secondarily enhanced target speech state set is performed, and the position and duration corresponding to the abnormal pause are marked in the speech frequency band corresponding to the secondarily enhanced target speech state set to obtain a speech feature space to be recognized.

3. The intelligent health robot integrating intelligent dialogue and health intervention according to claim 2, wherein The construction process of the word meaning compensation space and the tone calibration space includes: Obtain the standard dialect texts corresponding to different regions, the speech dialect texts recognized by different age groups, and the corresponding Mandarin expression texts, and perform corresponding relationship annotation to obtain compensation word order pairs; According to the compensation word order pairs, combined with the Penn Chinese Tree Bank, construct a hierarchical dialect text parsing space; The hierarchical dialect text parsing space includes the probability missing character segments and corresponding missing expression positions corresponding to the standard dialect texts in different regions and the speech dialect texts recognized by different age groups, as well as the probability missing character segments and corresponding missing expression position information corresponding to the standard dialect texts and the Mandarin expression texts.

4. The intelligent health robot integrating intelligent dialogue and health intervention according to claim 3, characterized in that, The construction process of the semantic compensation space and the tone calibration space further includes: Based on the missing expression positions and corresponding probability missing character segments corresponding to the speech dialect texts recognized by different age groups in each region of the hierarchical dialect text parsing space, a CRF model trained by missing word annotation is used to judge the part-of-speech of the missing character segments, and the part-of-speech discrimination result of the missing character segments is obtained; According to the judgment result and the probability corresponding to the missing character segments, a fine-tuned GraphSAGE model is combined with the marked abnormal pause positions and durations to perform random walks on the standard dialect text library constructed by the standard dialect texts, and the hierarchically compensated missing character segments corresponding to the speech texts and the corresponding compensation accuracy rates are obtained; The hierarchically compensated missing character segments include grammar quantifier compensation, real-word semantic compensation, and emotional expression vocabulary replacement compensation; Based on the speech dialect texts recognized by different age groups in each region, the corresponding hierarchically compensated missing character segments and the corresponding compensation accuracy rates, a semantic compensation space is constructed.

5. The intelligent health robot integrating intelligent conversation and health intervention according to claim 4, characterized in that, The construction process of the semantic compensation space and the tone calibration space further includes: Obtain the dialect pronunciation rule libraries corresponding to different dialect regions, and establish corresponding initial consonant conversion rules, final conversion rules, and tone mapping rules based on the dialect pronunciation rule libraries and the standard pronunciation rule libraries; Based on the initial consonant conversion rules, final conversion rules, and tone mapping rules, establish a triple conversion mapping space between Mandarin pronunciation and the dialect prosody and tones of different age groups and health states in the corresponding regions; Based on the triple conversion mapping space and combined with the reinforcement learning algorithm, by adjusting the speech rates, reference frequencies of the input dialect speech and the corresponding Mandarin speech, and the reverberation amounts of different regional dialects, perform simulated reinforcement training to obtain a triple conversion mapping space that meets the preset conversion accuracy threshold; Based on the dialect speech, the corresponding Mandarin speech, and the triple conversion mapping space that meets the preset conversion accuracy threshold, obtain the tone calibration space.

6. The intelligent health robot integrating intelligent dialogue and health intervention according to claim 5, characterized in that, The acquisition process of the corrected speech feature space to be recognized includes: According to the speech feature space to be recognized, use the speech slicing algorithm to perform speech slicing according to the positions corresponding to the marked abnormal pauses and connect and label the slice positions to obtain the speech feature slice space to be recognized; According to the speech feature slice space to be recognized, use the semantic compensation space to judge the missing character segments between the corresponding speech features internally and between adjacent speech feature segments to be recognized with connection labels, and obtain the part-of-speech and semantic coordinate positions of the corresponding character segments; According to the part-of-speech and semantic coordinate positions of the corresponding character segments, use the semantic compensation space to perform corresponding compensations for grammar quantifiers, real-word semantics, and emotional expression vocabulary on the corresponding speech feature slices to be recognized, and mark the corresponding compensation accuracy rates at the corresponding compensation position points to obtain the speech feature segment space to be recognized with semantic compensation.

7. The intelligent health robot integrating intelligent conversation and health intervention according to claim 6, characterized in that, The process of obtaining the corrected speech feature space to be recognized further includes: Input the speech feature segment space with semantic compensation to the tone calibration space. Combining the region to which the corresponding speech feature segment to be recognized belongs, the corresponding age group, and the health status, call the corresponding initial consonant conversion rule, final consonant conversion rule, and tone mapping rule to perform standardized mapping correction on the initial consonants, final consonants, and tones of the characters corresponding to each speech feature segment with semantic compensation, and obtain the speech feature space to be recognized after tone calibration; Based on the simulation algorithm combined with the reinforcement learning algorithm, and using the compensation accuracy difference corresponding to semantic compensation and the tone conversion accuracy difference as the training reward function, perform cyclic simulation training on the semantic compensation and tone calibration processes to obtain the corrected speech feature space to be recognized.

8. The intelligent health robot integrating intelligent dialogue and health intervention according to claim 7, characterized in that, The process of obtaining the intended emotional state space includes: Based on the corrected speech feature space to be recognized, extract the context-related features of the speech through the Wav2Vec 2.0 model, and perform frame-level alignment on the extracted speech segments; Perform acoustic feature sequence decoding, spelling correction, and punctuation prediction on the aligned speech segments through a context-aware speech recognition model to obtain the text information space of speech conversion; Based on the text information space of speech conversion, perform intent classification through a pre-trained Chinese intent decomposition model combined with a medical knowledge graph configured with temporal reasoning, and extract entities for the corresponding key intent content words and emotion expression words to obtain the intent classification result and the intent keyword space.

9. The intelligent health robot integrating intelligent dialogue and health intervention according to claim 8, characterized in that, The process of obtaining the intended emotional state space further includes: Based on the intent classification result, the intent keyword space, the corrected speech feature space to be recognized, and the fundamental frequency perturbation, amplitude perturbation, and harmonic-to-noise ratio in the speech feature space to be recognized, construct an emotional analysis feature space through the wavelet algorithm to obtain the intonation contour slope and energy envelope variance; Based on the pre-trained multi-modal emotion recognition model combined with the preset emotion recognition type and the preset health level-disease type-emotion type association matrix, construct a multi-modal emotion intent recognition model; Input the emotional analysis feature space into the multi-modal emotion intent recognition model to obtain the intended emotion type and level, and based on the intended emotion type and level, combine the change trend of the emotional state constructed from the corresponding user's historical emotion data to obtain the potential psychological problem features; Based on the intended emotion type, level, and the corresponding potential psychological problem features, obtain the intended emotional state space.

10. The intelligent health robot integrating intelligent dialogue and health intervention according to claim 9, characterized in that, The process of constructing the sound frequency-age-pathology feature association matrix includes: Based on the dialect types corresponding to different regions, obtain the speech information of 500 users corresponding to different age groups, health status levels, and disease types within the corresponding regions, and obtain the sound frequency and fundamental frequency fluctuation characteristics of the corresponding users through the sound preprocessing algorithm; Based on the sound frequency, fundamental frequency fluctuation characteristics, age, health status level, and disease type information of the corresponding users, obtain the correlation degrees between any two of the sound frequency, age, health status level, and disease type through the correlation algorithm; Based on the correlation degrees between pairwise combinations of sound frequency, age, health status level, and disease type, a sound frequency-age-pathological feature correlation matrix is obtained.

Citation Information

Patent Citations

  • Auxiliary analysis method for stroke rehabilitation assessment based on speech multi-task learning

    CN114141366B

  • Intelligent voice dialogue system and method based on AI scene

    CN118782036A

  • System and method for steering care plan actions by detecting tone, emotion, and / or health outcome

    WO2021071971A1