An assistive turning device for elderly care and its voice control method
By acquiring the breathy, weak, and dialect feature vectors of elderly people, combining them with the noise of the nursing environment, and integrating semantic state and device state logic, the problem of misoperation in elderly speech recognition in nursing scenarios has been solved, achieving higher recognition accuracy and device control stability.
Patent Information
- Application Number
- CN202511300698.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-12
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-09-12
AI Technical Summary
Existing speech recognition technology is difficult to adapt to the voice characteristics of the elderly, especially in care settings where it is easily affected by noise, resulting in low recognition accuracy and the risk of misoperation.
By acquiring exclusive acoustic feature vectors of elderly people’s breathy voice, weak voice and dialect characteristics, and combining them with changes in the noise of the nursing environment to determine the robustness weight of the voice signal, and by integrating semantic state features and device state logic, fault-tolerant voice command execution can be achieved.
It improves the accuracy and stability of voice control for the elderly, avoids misoperation, and enhances the intelligence and safety of nursing equipment.
Smart Images

Figure CN120808768B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of voice control technology, and more specifically, to an assistive turning device for elderly care and its voice control method. Background Technology
[0002] Voice control refers to the method by which users control the device to perform actions by giving voice commands (such as "turn to the left", "stop", "increase angle"). The voice recognition method converts natural language into control signals that the device can execute. Assisted turning devices refer to electric or mechanical structures installed on hospital beds or nursing beds to help the elderly automatically turn to the left or right and adjust their lying angle, thereby reducing the risk of bedsores and reducing the workload of caregivers.
[0003] Voice control methods for assisted turning devices in elderly care refer to the use of voice recognition (converting speech into machine-readable commands) to control the device. This allows elderly people with limited mobility to control the device simply by speaking, without needing to manually press buttons, touch screens, or rely on others for assistance. It's a user-friendly feature designed to enhance the elderly's self-care ability and simplify the operation process. It enables the elderly to control the device independently when conscious, improving the care experience and making the care process more intelligent and humane. It is particularly suitable for elderly people with limited mobility or difficulty expressing themselves verbally, allowing them to control the device through the most natural way of "speaking," without complex physical movements. However, in existing technologies, voice recognition... While primarily designed for adults, existing speech recognition models are not optimized for the vocal characteristics of the elderly. They struggle to adapt to speech distortions and unclear pronunciation caused by physiological features such as missing teeth, respiratory illnesses, slow speech, and heavy accents. Furthermore, they lack the ability to adapt to common elderly expressions, weak intonation, or breathy pronunciation. Additionally, the nursing environment often involves sudden noise disturbances (such as monitoring alarms, conversations among caregivers, and bed movement noise), easily leading to speech recognition errors or missed judgments. This results in low accuracy in recognizing elderly speech and poses a risk of misoperation. Therefore, how to achieve fault-tolerant voice command execution in elderly care scenarios by fusing speech recognition confidence with device status logic, thereby reducing the misrecognition rate when elderly people use voice-controlled turning devices, has become a challenge for the industry. Summary of the Invention
[0004] This application provides an assistive turning device for elderly care and its voice control method, which can realize fault-tolerant voice command execution in elderly care scenarios based on the fusion control of voice recognition confidence and device status logic.
[0005] In a first aspect, this application provides a voice control method for an assistive turning device for elderly care, the method comprising the following steps:
[0006] Obtain a unique acoustic feature vector that includes the breathy voice characteristics, weak voice characteristics, and dialect characteristics of the elderly;
[0007] The robustness weights of speech signals under different nursing action scenarios are determined by combining the specific acoustic feature vectors with sudden noise changes in elderly people under different nursing environments.
[0008] Based on the nursing voice commands of the elderly in the current nursing action scenario and the exclusive acoustic feature vector, the semantic state features of the elderly in the current nursing action scenario are determined. Based on the semantic state features and the robustness weight of the voice signals in different nursing action scenarios, the effective value of the nursing voice commands of the elderly in the current nursing action scenario is determined.
[0009] The nursing voice commands of the elderly in the current nursing action scenario are classified into action categories to obtain the corresponding turning action category of the elderly in the current nursing action scenario, and then the action logic relationship between the turning action category and the current turning action state of the auxiliary turning device is determined.
[0010] The credibility of the elderly person's voice control command execution in the current nursing action scenario is determined by the recognition of the valid value and the association of the action logic, and then the turning action control command of the assistive turning device is generated based on the credibility of the command execution.
[0011] In some embodiments, determining the robustness weights of speech signals under different nursing action scenarios by combining the dedicated acoustic feature vector with sudden noise changes in elderly individuals under different nursing environments specifically includes:
[0012] The noise interference experienced by elderly individuals in different care environments was determined by sudden changes in noise levels.
[0013] The elderly’s voice signals collected under different nursing action scenarios are compared with the exclusive acoustic feature vector to obtain the acoustic feature matching error of the elderly’s voice signals under the noise interference.
[0014] The robustness weights of speech signals under different nursing action scenarios are determined by the acoustic feature matching error.
[0015] In some embodiments, combining the elderly person's voice commands in the current care action scenario with the specific acoustic feature vector includes:
[0016] Acquire the nursing voice commands of the elderly in the current nursing action scenario;
[0017] The nursing voice commands are enhanced based on the specific acoustic feature vector;
[0018] The semantic state features of elderly people in the current nursing action scenario are determined by the nursing voice commands with enhanced understanding.
[0019] In some embodiments, determining the effective value of the nursing voice command for the elderly in the current nursing action scenario based on the semantic state features and the robustness weights of the voice signals under different nursing action scenarios specifically includes:
[0020] The semantic state features are combined with the robustness weights of the speech signals under different nursing action scenarios for evaluation to obtain the collaborative score between the semantic state features and the robustness weights.
[0021] The effective value of the recognition of nursing voice commands for the elderly in the current nursing action scenario is determined based on the collaborative score.
[0022] In some embodiments, the nursing voice commands given by the elderly person in the current nursing action scenario are classified into action categories to obtain the corresponding turning action category for the elderly person in the current nursing action scenario. Specifically, these categories include:
[0023] The nursing voice commands given by the elderly in the current nursing action scenario are converted into text to obtain the written expression of the elderly's voice;
[0024] Based on the textual expression of the elderly person's speech, keywords are extracted and semantics are matched for the nursing voice instructions to obtain the core action words of the elderly person in the current nursing action scenario;
[0025] By identifying the meaning of the elderly person's actions in the current nursing scenario using the core action words, the corresponding turning-over action category for the elderly person in the current nursing scenario can be obtained.
[0026] In some embodiments, determining the action logic association between the turning-over action category and the current turning-over action state of the assisted turning-over device specifically includes:
[0027] Read the current turning motion status of the assisted turning device;
[0028] Determine the association between the current rolling-over action state and the rolling-over action category;
[0029] The associated state establishes a logical association between the rolling-over action category and the current rolling-over action state.
[0030] In some embodiments, determining the credibility of voice control command execution by the elderly person in the current care action scenario through the identification of valid values and the association of action logic specifically includes:
[0031] The executable strength of voice control for the elderly in the current nursing action scenario is determined by the action logic association.
[0032] The reliability of the elderly person's voice control command execution in the current nursing action scenario is determined based on the executable strength and the recognized valid value.
[0033] In some embodiments, the execution of confidence-based control instructions for generating turning-over movements of the assisted turning-over device based on the instructions specifically includes:
[0034] Preset trusted execution threshold;
[0035] When the credibility of the instruction execution is greater than or equal to the credibility execution threshold, the nursing voice instruction of the elderly in the current nursing action scenario is determined to be executable, and then the turning action control instruction of the assistive turning device is generated.
[0036] When the credibility of the instruction execution is less than the credibility execution threshold, the nursing voice instruction for the elderly in the current nursing action scenario is marked as unexecutable, and the elderly are prompted to reconfirm through voice feedback.
[0037] In some embodiments, the sudden noise change in the nursing environment refers to a background noise source that suddenly appears during the nursing process and has unstable spectral characteristics.
[0038] Secondly, this application provides an assistive turning device for elderly care, including a voice control unit, the voice control unit comprising:
[0039] The acquisition module is used to acquire exclusive acoustic feature vectors containing the breathy voice characteristics, weak voice characteristics, and dialect characteristics of the elderly;
[0040] The processing module is used to determine the robustness weights of the speech signal under different nursing action scenarios by combining the dedicated acoustic feature vector with the sudden noise changes of the elderly in different nursing environments.
[0041] The processing module is further configured to determine the semantic state features of the elderly in the current nursing action scenario based on the nursing voice commands of the elderly in the current nursing action scenario and the exclusive acoustic feature vector, and to determine the effective recognition value of the nursing voice commands of the elderly in the current nursing action scenario based on the semantic state features and the robustness weight of the voice signals in different nursing action scenarios.
[0042] The processing module is also used to determine the action category of the nursing voice commands of the elderly in the current nursing action scenario, obtain the turning action category of the elderly in the current nursing action scenario, and then determine the action logic association between the turning action category and the current turning action state of the auxiliary turning device.
[0043] The execution module is used to determine the credibility of the elderly person's voice control command execution in the current nursing action scenario by the recognition valid value and the action logic association, and then generate the turning action control command of the assisted turning device based on the command execution credibility.
[0044] The technical solutions provided by the embodiments disclosed in this application have the following beneficial effects:
[0045] In this application, a unique acoustic feature vector containing the breathy, weak, and dialectal characteristics of the elderly is obtained. The robustness weights of the speech signals in different nursing action scenarios are determined by combining the unique acoustic feature vector with sudden noise changes in different nursing environments. The semantic state features of the elderly in the current nursing action scenario are determined by combining the nursing speech commands of the elderly with the unique acoustic feature vector. Based on the semantic state features and the robustness weights of the speech signals in different nursing action scenarios, the effective recognition value of the nursing speech commands of the elderly in the current nursing action scenario is determined. The nursing speech commands of the elderly in the current nursing action scenario are classified into action categories to obtain the corresponding turning action category, and then the action logic association between the turning action category and the current turning action state of the assisted turning device is determined. The credibility of the voice control command execution of the elderly in the current nursing action scenario is determined by the effective recognition value and the action logic association, and then the turning action control command of the assisted turning device is generated based on the command execution credibility.
[0046] Therefore, in this application, firstly, the effective value of the nursing voice command for the elderly in the current nursing action scenario is determined based on the semantic state features and the robustness weight of the voice signal under different nursing action scenarios. This effectively quantifies the recognizability of the current voice command in complex environments, avoiding erroneous control due to insufficient confidence or semantic ambiguity, thus significantly enhancing the controllability and fault tolerance of the elderly's voice recognition process. This is the basic guarantee for subsequent voice command execution logic judgment. Secondly, the action logic association between the turning action category and the current turning action state of the assisted turning device is determined. By judging whether the elderly's voice request has execution significance (such as avoiding repeated left turning, directional conflict, etc.) through action logic, it can effectively compensate for misjudgments caused by auditory interference or unclear expression during the voice recognition process, avoiding the direct transmission of voice recognition errors to the assisted turning device. This reflects the constraining integration of the voice control system with the device's behavioral logic, so as to achieve fault-tolerant control and avoidance of misoperation. Then, through the effective recognition value and the action logic... By establishing a logical association to determine the credibility of voice control commands executed by the elderly in the current nursing scenario, this approach effectively avoids relying solely on recognition validity to determine whether to execute commands, thus improving the intelligence and stability of the voice control strategy. For example, even when the voice expression is somewhat ambiguous but the command is highly consistent with the current device status, it can still be judged as "credible." This decision-making mechanism based on command execution credibility significantly improves the actual fault-tolerant control capability of the elderly's voice interaction process under multi-source interference in the nursing setting. Finally, based on the command execution credibility, the turning action control commands for the assisted turning device are generated. This effectively blocks false triggers when faced with high-risk voice inputs such as weak voices, wheezing, or inverted speech patterns from the elderly, preventing the assisted turning device's actions from becoming disconnected from the elderly user's intentions and causing accidental injury. This further improves the fault-tolerant control closed loop, ensuring that the intelligent response process of the assisted turning device is more robust, safe, and humane. In summary, this solution can achieve fault-tolerant voice command execution in elderly care scenarios based on the fusion control of voice recognition confidence and device status logic. Attached Figure Description
[0047] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0048] Figure 1 This is an exemplary flowchart of a voice control method for an assistive turning device for elderly care, according to some embodiments of this application;
[0049] Figure 2This is an exemplary flowchart illustrating the determination of semantic state features according to some embodiments of this application;
[0050] Figure 3 This is an exemplary flowchart illustrating the action category determination according to some embodiments of this application;
[0051] Figure 4 This is a schematic diagram of the structure of a voice control unit according to some embodiments of this application;
[0052] Figure 5 This is a schematic diagram of the structure of a computer device that implements a voice control method for an assistive turning device for elderly care, according to some embodiments of this application. Detailed Implementation
[0053] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0054] refer to Figure 1 The figure is an exemplary flowchart of a voice control method for an assistive turning device for elderly care, according to some embodiments of this application. The method mainly includes the following steps:
[0055] In step 101, a unique acoustic feature vector containing the breathy characteristics, weak voice characteristics, and dialect characteristics of the elderly is obtained.
[0056] In practical implementation, obtaining a unique acoustic feature vector containing the breathy, weak, and dialectal characteristics of the elderly can be achieved as follows: First, a voice acquisition device equipped with an array microphone can be used to acquire the natural speech commands of the elderly during care through multiple channels. The acquired raw speech data of the elderly is then processed using the Voice Activity Detection (VAD) method to remove silence segments and calibrate speech segments. Next, multi-channel speech enhancement techniques, such as a fusion method based on beamforming and spectral subtraction, are applied to the speech segments to suppress background noise and highlight the main speech signal of the elderly. Subsequently, the Short-Time Fourier Transform (STFT) spectrum is extracted from the enhanced speech signal, and the Mel-Frequency Cepstral Coefficients are further calculated. Acoustic features such as coefficients (MFCC), formant distribution, and speech energy envelope are identified. A deep neural network model with an attention mechanism (such as BiLSTM-Attention) is constructed to identify aerophonic characteristics (i.e., low-energy regions with a high proportion of expiratory components due to incomplete vocal cord closure), weak voice characteristics (i.e., low overall speech intensity and narrowed dynamic range), and dialect characteristics (i.e., language region features such as phoneme substitution, intonation shift, and speech rate variations at the speech unit level). Finally, the three types of labeled features—aerophonic, weak, and dialect—are fused with their corresponding acoustic features in a multi-dimensional manner to construct a multi-dimensional vector representing the individual speech characteristics of the elderly. This vector serves as the elderly-specific acoustic feature vector for subsequent semantic state modeling and recognition reliability assessment. Other methods can also be used in other embodiments, and are not limited here.
[0057] It should be noted that the aerophonic characteristics in this application refer to the phenomenon of significant exhalation noise and high-frequency energy diffusion in the speech spectrum, which reflects the weakening of the elderly's voice; the weak voice characteristics in this application refer to the characteristic of low overall sound pressure level in speech, which reflects the weak tone of the elderly suffering from illness; the dialectal characteristics in this application refer to the sound change or word change speech features generated in a specific regional context, which are interfering with the standard Mandarin model; in addition, the exclusive acoustic feature vector in this application is a multi-dimensional numerical representation used to characterize the individual pronunciation pattern, speech signal structure and change characteristics of the elderly. As the key input data for the personalized speech recognition ability of the elderly model, the exclusive acoustic feature vector can enhance the adaptability and accuracy of subsequent semantic analysis and action command judgment under the diverse speech input of the elderly.
[0058] In step 102, the robustness weights of the speech signal under different nursing action scenarios are determined by combining the dedicated acoustic feature vector with the sudden noise changes of the elderly in different nursing environments.
[0059] In some embodiments, determining the robustness weights of speech signals under different nursing action scenarios by combining the specific acoustic feature vector with sudden noise changes in elderly people under different nursing environments can be achieved through the following steps:
[0060] The noise interference experienced by elderly individuals in different care environments was determined by sudden changes in noise levels.
[0061] The elderly’s voice signals collected under different nursing action scenarios are compared with the exclusive acoustic feature vector to obtain the acoustic feature matching error of the elderly’s voice signals under the noise interference.
[0062] The robustness weights of speech signals under different nursing action scenarios are determined by the acoustic feature matching error.
[0063] In practice, determining noise interference in different nursing environments for elderly individuals can be achieved through the following methods: Sound environment sensing technology can be used to monitor potential sudden noise sources during elderly care in real time, such as the operation of instruments by caregivers, bed movement, and the operation of oxygen concentrators. This can be done by setting up multiple microphone arrays for omnidirectional sound sampling, and using a combination of short-time energy detection and spectral entropy analysis. Entropy is used to detect sudden noise features in the acquired audio to identify the noise intensity and frequency range of sudden noise changes, which serves as noise interference for the elderly in different care environments. The voice signals of the elderly collected under different care scenarios are compared with the dedicated acoustic feature vector. The acoustic feature matching error of the elderly's voice signal under the noise interference can be achieved by: collecting the actual nursing instruction voices issued by the elderly in different care scenarios (such as getting up, turning over, lying flat), and performing frame-by-frame matching and comparison with the dedicated acoustic feature vector. Then, dynamic time warping algorithms and spectral difference measurement methods (such as logarithmic spectral distance) combined with the identified noise intensity and frequency range in the noise interference can be used to evaluate the acoustic feature matching error between the elderly's voice signal and the dedicated acoustic feature vector. The robustness weights of speech signals under different nursing action scenarios can be determined by the acoustic feature matching error as follows: the acoustic feature matching error under different noise interference can be used as an input index through a multi-channel robustness evaluation function. Combined with the energy retention and spectral stability of the speech signal, a robustness score of the speech signal under different nursing action scenarios is output. Then, the score is normalized to generate a value between 0 and 1 as the robustness weight of the speech signal under different nursing action scenarios. This weight is used to measure the stability and recognizability of the speech signal under sudden noise interference in different nursing action scenarios. It can be used as an adaptive compensation factor for subsequent nursing scenarios to improve the stability and execution reliability of elderly people's voice command recognition under high noise conditions. Other methods can also be used to determine the robustness weights in other embodiments, which are not limited here.
[0064] It should be noted that, in this application, sudden noise changes in the nursing environment refer to background noise sources that suddenly appear and have unstable spectral characteristics during the nursing process; nursing action scenarios in this application represent the operational states at different stages of elderly care, such as before turning over, during turning over, and during nursing intervention; noise interference in this application refers to non-target sound components in the elderly care environment that interfere with speech recognition, which are used to identify the external pressure conditions faced by the elderly in speech command recognition; acoustic feature matching error in this application represents the degree of deviation between the elderly's speech and standard features in the time or frequency domain, which is used to assess the impact of nursing environment noise on the stability of the elderly's speech expression; robustness weight in this application represents the stability score of elderly speech command recognition under different nursing action scenarios, which aims to improve the accuracy and reliability of subsequent elderly speech command recognition stages in non-ideal environments.
[0065] In step 103, the semantic state features of the elderly in the current nursing action scenario are determined by combining the nursing voice commands of the elderly in the current nursing action scenario with the dedicated acoustic feature vector. The effective recognition value of the nursing voice commands of the elderly in the current nursing action scenario is determined based on the semantic state features and the robustness weight of the voice signals in different nursing action scenarios.
[0066] In some embodiments, reference Figure 2 As shown, this figure is an exemplary flowchart for determining semantic state features in some embodiments of this application. In this embodiment, the semantic state features of the elderly person in the current nursing action scenario can be determined by the following steps based on the nursing voice commands of the elderly person in the current nursing action scenario combined with the specific acoustic feature vector:
[0067] First, in step 1031, the nursing voice commands of the elderly in the current nursing action scenario are obtained;
[0068] Secondly, in step 1032, the nursing voice command is enhanced based on the specific acoustic feature vector;
[0069] Finally, in step 1033, the semantic state features of the elderly person in the current nursing action scenario are determined by the enhanced understanding of the nursing voice instructions.
[0070] In specific implementation, acquiring nursing voice commands from the elderly in the current nursing action scenario can be achieved in the following way: acquiring nursing voice commands actively issued by the elderly during the current nursing action (such as turning over or assisting in lifting), such as "help me turn over" or "move a little to the left," through voice acquisition devices within the nursing scenario; enhancing the understanding of the nursing voice commands based on the specific acoustic feature vector can be achieved in the following way: dividing the nursing voice commands into short-time frames, extracting their basic acoustic parameters (such as MFCC, zero-crossing rate, energy spectrum, etc.) and frame-level temporal information of the elderly nursing voice sequence, and then calling pre-trained acoustic models (such as Conformer models or Factorized Time-Delay Neural Networks (TDNN-F structures)) and language models (such as Bidirectional Encoder Representations from Transformers (BERT) or Long Short-Term Memory networks). A joint decoder (Memory, LSTM) is used to fuse the contextual semantics and speech sequence signal of the nursing voice instruction to obtain a preliminary semantic vector. This preliminary semantic vector is then cross-aligned and attention-weighted modeled with the elderly person's specific acoustic feature vector (including their breathy frequency band, weak sound intensity, and dialect pronunciation patterns). An acoustic consistency attention mechanism is used to enhance the understanding of blurred breathy boundaries, unstable intonation, or dialectal expressions, thereby correcting the nursing voice instruction by reducing deviations caused by speech rate and unclear pronunciation. This strengthens the key semantic information in the elderly person's speech content, thus enhancing the understanding of the nursing voice instruction. The attention-weighted modeling refers to using an attention mechanism to strengthen the mapping between key speech frames and feature vectors, improving the semantic parsing's ability to understand complex speech instructions. The response sensitivity of non-speech features; the semantic state features of the elderly in the current nursing action scenario can be determined by the understanding-enhanced nursing voice instructions in the following way: the semantic orientation and state sentiment of the nursing voice instructions in the current nursing action scenario can be extracted by combining the semantic prior knowledge of the understanding-enhanced nursing voice instructions and the current nursing action scenario labels (such as "turning over before sleep" or "sudden call") through graph convolutional neural networks or semantic scene inference networks. For example, the intention category, context sentiment, semantic coherence, tone intensity, expression confidence, etc. of the nursing voice instructions for the elderly. The final extracted results are used as the semantic state features of the elderly in the current nursing action scenario. Other methods can also be used in other embodiments, which are not limited here.
[0071] It should be noted that the nursing voice commands in the current nursing action scenario in this application represent the voice content of the elderly actively expressing their needs during the nursing process under the current nursing operation; the understanding enhancement in this application refers to the process of adjusting and enhancing the acoustic and semantic content of the elderly's current nursing voice commands by combining personalized acoustic features, so as to improve the recognizability and semantic clarity of the nursing voice commands. The semantic state features in this application are used to characterize the voice intention that the elderly want to express in the current nursing action scenario, which is an important basis for subsequent nursing action logic judgment and execution control.
[0072] In some embodiments, determining the effective value of nursing voice commands for the elderly in the current nursing action scenario based on the semantic state features and the robustness weights of voice signals under different nursing action scenarios can be achieved through the following steps:
[0073] The semantic state features are combined with the robustness weights of the speech signals under different nursing action scenarios for evaluation to obtain the collaborative score between the semantic state features and the robustness weights.
[0074] The effective value of the recognition of nursing voice commands for the elderly in the current nursing action scenario is determined based on the collaborative score.
[0075] In specific implementation, the semantic state features are combined with the robustness weights of the speech signals under different nursing action scenarios for evaluation. The collaborative score between the semantic state features and the robustness weights can be achieved in the following way: the semantic state features and the robustness weights of the speech signals under the current corresponding nursing action scenario can be input into the fusion evaluation model. This model can use a multi-layer perceptron (MLP). A Perceptron (MLP) structure or a weighted scoring model based on an attention mechanism (such as a transformer encoder) is used to combine the semantic information strength and anti-interference ability of the input. In the combined evaluation stage, the model comprehensively scores the similarity, complementarity, and dynamic cooperation relationship between semantic state features and robustness weights based on a pre-trained speech semantic alignment function to calculate the collaborative score between semantic state features and robustness weights. The effective value of the nursing voice command recognition for the elderly in the current nursing action scenario can be determined based on the collaborative score in the following way: the collaborative score can be input into a linear mapping layer or a sigmoid normalization function to map the collaborative score to the interval [0, 1], and the final value is used as the effective value of the nursing voice command recognition for the elderly in the current nursing action scenario. The higher the value, the more reliable the nursing voice command is in the current nursing action scenario and the more suitable it is as the execution input for voice turning control. Other methods can also be used in other embodiments, which are not limited here.
[0076] It should be noted that the collaborative score in this application reflects the degree to which elderly care voice commands possess both semantic clarity and anti-interference stability in the current care action scenario; the combined evaluation in this application refers to the process of jointly processing semantic state features and robustness weights, and quantifying their joint reliability through a fusion evaluation model, which can be used to comprehensively judge the credibility of elderly care voice commands; the valid recognition value in this application reflects the reliability of the accurate recognition and semantic judgment of elderly care voice commands in the current care action scenario, and it is a key threshold judgment basis before executing voice control logic.
[0077] In step 104, the nursing voice commands of the elderly in the current nursing action scenario are classified into action categories to obtain the corresponding turning action category of the elderly in the current nursing action scenario, and then the action logic association between the turning action category and the current turning action state of the auxiliary turning device is determined.
[0078] In some embodiments, reference Figure 3 As shown, this figure is an exemplary flowchart of action category determination in some embodiments of this application. The action category determination of the elderly person's voice commands in the current nursing action scenario, to obtain the corresponding turning-over action category for the elderly person in the current nursing action scenario, can be achieved through the following steps:
[0079] The nursing voice commands given by the elderly in the current nursing action scenario are converted into text to obtain the written expression of the elderly's voice;
[0080] Based on the textual expression of the elderly person's speech, keywords are extracted and semantics are matched for the nursing voice instructions to obtain the core action words of the elderly person in the current nursing action scenario;
[0081] By identifying the meaning of the elderly person's actions in the current nursing scenario using the core action words, the corresponding turning-over action category for the elderly person in the current nursing scenario can be obtained.
[0082] In specific implementation, the text conversion of nursing voice commands from the elderly in the current nursing action scenario can be achieved in the following way: A speech recognition model (such as a Connectionist Temporal Classification (CTC) or Transformer-based acoustic-language integrated model) can be used to transcribe the nursing voice commands from the elderly in the current nursing action scenario, converting them into standardized written expressions. Based on the written expression of the elderly's voice, keyword extraction and semantic matching can be performed on the nursing voice commands to obtain the core action words of the elderly in the current nursing action scenario. This can be achieved in the following way: Based on a pre-constructed action semantic mapping vocabulary, keyword extraction and semantic matching can be performed on the transcribed written expression of the elderly's voice to extract core action words with instructive meaning, such as "turn over," "turn to the left," and "turn to the right." The core action words, such as "turn back," "return to center," "turn a little more," and "turn 30 degrees to the left," are used to determine the meaning of the elderly person's actions in the current nursing scenario. This can be achieved by using a trained action classification model, such as a BERT or BiLSTM-based multi-classifier, and inputting the core action words into the model to determine the meaning of the elderly person's actions in the current nursing scenario and identify the corresponding turning action category, such as "turning over to the left side," "turning over to the right side," and "returning to supine position." Other methods can also be used in other embodiments, which are not limited here.
[0083] It should be noted that the action category determination in this application refers to the process of mapping the identified elderly person's voice information to specific action types according to a preset classification standard; the text expression in this application refers to the result of transcribing the elderly person's nursing voice instructions into natural language text through voice recognition technology, which is a readable sentence or phrase; the core action words in this application refer to the key verbs or verb phrases extracted from the transcribed text, which represent the core control intent in the elderly person's nursing voice instructions; the turning action category in this application refers to the standardized turning action corresponding to the nursing voice instruction in the current nursing action scenario, which serves as the final execution target of the control instructions for the assistive turning device; in addition, this step improves the accuracy of voice control in real nursing scenarios and the error tolerance of instruction recognition through semantic alignment and context awareness.
[0084] In some embodiments, determining the action logic association between the turning-over action category and the current turning-over action state of the assisted turning-over device can be achieved through the following steps:
[0085] Read the current turning motion status of the assisted turning device;
[0086] Determine the association between the current rolling-over action state and the rolling-over action category;
[0087] The associated state establishes a logical association between the rolling-over action category and the current rolling-over action state.
[0088] In specific implementation, reading the current turning motion status of the assisted turning device can be achieved in the following way: The current turning motion status of the assisted turning device can be obtained in real time through a built-in motion status reading module (such as a position sensor, angle encoder, or actuator feedback system). This includes the device's current turning direction (e.g., left lateral, right lateral, supine) and execution progress (e.g., completed, in progress, standby), etc., thus reading the current turning motion status of the assisted turning device. The association between the current turning motion status and the turning motion category can be determined using... The following approach is used: A state-action mapping model bound to the motion structure of the assisted turning device can be utilized. This model, based on a nursing rule knowledge base or an action switching flowchart developed by experts, determines whether the target turning action (i.e., the turning action category corresponding to the elderly person in the current nursing action scenario) conflicts with the current device state (e.g., the user issues a "turn left" command when the elderly person is already in a left lateral decubitus position). The resulting judgment (consistent, conflicting, continuous, or repetitive) is used as the association state between the two to determine if a reasonable logical relationship exists between the turning action category and the current turning action state. The action logic association between the turning action category and the current turning action state can be constructed using the following method: If the association state indicates that the target turning action and the current device state are logically continuous, the action logic association between the turning action category and the current turning action state is considered valid, indicating a high degree of feasibility and rationality in the turning action transition (i.e., a high association weight). If there is a conflict or redundancy, it can be adjusted to a feasible transitional action through an action scheduling logic module (e.g., state machine judgment or rule engine). The system uses actions such as "first return to position, then roll to the right" or "currently rolled 30 degrees to the right, needs to roll 15 degrees more to the right" to construct a logical relationship diagram between the rolling action category and the current rolling action state. This diagram serves as the logical association between the two actions, and a weight is assigned to the association (this weight represents the feasibility and rationality of the rolling action transition, with a value between 0 and 1). This forms execution prior information that can be referenced by subsequent control modules, thereby improving the rationality and continuity of the executed actions and effectively preventing the execution of repeated, conflicting, or invalid control commands. Other methods can also be used in other embodiments, which are not limited here.
[0089] It should be noted that the current turning-over action state in this application refers to the specific physical state description of the actual execution position and movement stage of the assisted turning-over device; the associated state in this application represents the logical relationship type between the target turning-over action and the current turning-over state of the assisted turning-over device, such as the judgment results of consistency, conflict, continuity, repetition, etc.; the action logic association in this application refers to the execution dependency logic between the target turning-over action in the elderly care voice command and the current turning-over state of the assisted turning-over device. It is the intermediate bridge between the elderly's voice recognition action result and the control device's execution action, used to ensure that the action connection during the operation of the assisted turning-over device is reasonable, safe, and efficient.
[0090] In step 105, the credibility of the elderly person's voice control command execution in the current nursing action scenario is determined by the identification valid value and the action logic association, and then the turning action control command of the assisted turning device is generated based on the command execution credibility.
[0091] In some embodiments, determining the reliability of voice control command execution by the elderly in the current care action scenario through the identification of valid values and the association of action logic can be achieved by the following steps:
[0092] The executable strength of voice control for the elderly in the current nursing action scenario is determined by the action logic association.
[0093] The reliability of the elderly person's voice control command execution in the current nursing action scenario is determined based on the executable strength and the recognized valid value.
[0094] In specific implementation, determining the executable strength of voice control by the elderly in the current nursing action scenario through the action logic association can be achieved in the following way: the executable strength of voice control by the elderly in the current nursing action scenario can be determined based on the action logic association. For example, if the assisted turning device is currently "turning to the left" and the elderly person's nursing voice command in the current nursing action scenario is "continue turning to the left", then the executable strength can be set between 0.8 and 1 (indicating that the action logic is reasonable). If the nursing voice command is "turn to the right", then the executable strength can be set between 0.4 and 0.5 (indicating that the action logic has some conflict). If the elderly person's nursing voice command is "turn to the right" but the assisted turning device is currently at the end of the right turn, then the executable strength is 0 (indicating that the action logic conflicts). The executable strength is between 0 and 1. Based on the executable strength and the recognized valid value, the elderly person's voice control is determined. The credibility of voice control command execution by the elderly in the current nursing action scenario can be achieved in the following way: the executable strength and the recognized effective value can be weighted and fused through linear combination, fuzzy inference, or confidence fusion function (such as product fusion or softmax normalization, attention weight-based fusion network). The weighting coefficient can be set according to historical experience data (with priority given to action safety). Finally, a confidence index is obtained to characterize whether the nursing voice command of the elderly in the current nursing action scenario is worth executing by the device. It is ensured that the fusion output result is between 0 and 1 and has discriminative power. The output confidence index is then used as the credibility of voice control command execution by the elderly in the current nursing action scenario, and is used in the subsequent voice control process to determine whether to adopt the current nursing voice command of the elderly. Other methods can also be used to determine this in other embodiments, which are not limited here.
[0095] It should be noted that the executable strength in this application represents a quantitative value of the feasibility of executing nursing voice commands in the current turning state of the device. It is used to constrain whether the voice recognition result can be reasonably executed in the current context, so as to reflect the state perception capability of the assistive turning device. The command execution credibility in this application refers to a comprehensive scoring index of whether the elderly person's current voice control command is worth executing. It is used to reflect the acceptability of the elderly person's current voice command and the coordination between the state of the assistive turning device. As the basis for the final judgment on whether the nursing voice command is executed, it can improve the reliability and accurate response capability of the elderly person's voice control.
[0096] In some embodiments, generating turnover motion control instructions for the assisted turnover device based on the confidence level of the instructions can be achieved through the following steps:
[0097] Preset trusted execution threshold;
[0098] When the credibility of the instruction execution is greater than or equal to the credibility execution threshold, the nursing voice instruction of the elderly in the current nursing action scenario is determined to be executable, and then the turning action control instruction of the assistive turning device is generated.
[0099] When the credibility of the instruction execution is less than the credibility execution threshold, the nursing voice instruction for the elderly in the current nursing action scenario is marked as unexecutable, and the elderly are prompted to reconfirm through voice feedback.
[0100] In specific implementation, the preset reliable execution threshold can be achieved in the following way: it can be trained based on a large number of real nursing scenario samples, and the distribution of the success rate of control command execution under different command execution reliability can be statistically analyzed, thereby presetting a reasonable reliable execution threshold (such as 0.7). In other embodiments, it can also be set in other ways, which are not specifically limited here. The generation of the turning action control command of the assisted turning device can be achieved in the following way: when the command execution reliability is greater than or equal to the reliable execution threshold, the nursing voice command of the elderly in the current nursing action scenario is determined to be executable, and the corresponding turning action category is used as the action type index. The preset action execution command is matched with the current turning action state of the assisted turning device and the device action parameters (such as movement angle, speed, duration) to generate a specific turning action control command. This includes controlling the target turning action (e.g., "turning to the left"), execution parameters (e.g., "45 degrees, 30-second slow speed"), and control timing to drive the turning action to complete a precise movement. When the confidence level of the instruction execution is less than the confidence execution threshold, the nursing voice instruction for the elderly in the current nursing action scenario is marked as unexecutable, and the elderly are prompted to reconfirm via voice feedback. This can be achieved in the following way: when the confidence level of the instruction execution is less than the confidence execution threshold, the nursing voice instruction for the elderly in the current nursing action scenario is marked as unexecutable, and the voice feedback module is invoked to provide a friendly reminder to the elderly, such as "Please say it again" or "I did not hear your instruction clearly," to ensure that the voice interaction has a feedback loop and improve the efficiency of human-computer communication. Other methods can also be used in other embodiments, which are not limited here.
[0101] It should be noted that the turning action control command in this application refers to the executable command directed to the underlying control module of the assisted turning device, which is used to guide the assisted turning device to perform specific nursing turning actions.
[0102] In another aspect, in some embodiments, this application provides an assistive turning device for elderly care, the device including a voice control unit, see reference. Figure 4The figure is a schematic diagram of the structure of a voice control unit according to some embodiments of this application. The voice control unit 400 includes: an acquisition module 401, a processing module 402, and an execution module 403, which are described below:
[0103] The acquisition module 401 in this application is mainly used to acquire a unique acoustic feature vector containing the breathy voice characteristics, weak voice characteristics and dialect characteristics of the elderly.
[0104] Processing module 402, in this application, is mainly used to determine the robustness weight of speech signals under different nursing action scenarios by combining the exclusive acoustic feature vector with the sudden noise changes of the elderly in different nursing environments.
[0105] The processing module 402 described in this application is further configured to determine the semantic state features of the elderly in the current nursing action scenario based on the nursing voice instructions of the elderly in the current nursing action scenario and the exclusive acoustic feature vector, and to determine the effective recognition value of the nursing voice instructions of the elderly in the current nursing action scenario based on the semantic state features and the robustness weight of the voice signals in different nursing action scenarios.
[0106] The processing module 402 described in this application is also used to determine the action category of the nursing voice commands of the elderly in the current nursing action scenario, obtain the turning action category of the elderly in the current nursing action scenario, and then determine the action logic association between the turning action category and the current turning action state of the auxiliary turning device.
[0107] The execution module 403 in this application is mainly used to determine the credibility of the elderly person's voice control command execution in the current nursing action scenario by the recognition valid value and the action logic association, and then generate the turning action control command of the assisted turning device based on the command execution credibility.
[0108] The foregoing has detailed examples of assisted turning devices for elderly care and their voice control methods provided in the embodiments of this application. It is understood that, in order to achieve the above functions, the corresponding devices include hardware structures and / or software modules for executing each function. Those skilled in the art should readily recognize that, based on the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0109] In some embodiments, this application also provides a computer device, the computer device including a memory and a processor, the memory for storing a computer program, and the processor for calling and running the computer program from the memory, causing the computer device to execute the above-described voice control method for an assistive turning device for elderly care.
[0110] In some embodiments, reference Figure 5 The dashed lines in the figure indicate that the unit or module is optional. This figure is a schematic diagram of the structure of a computer device implementing the voice control method for the assistive turning device for elderly care according to this application. The voice control method for the assistive turning device for elderly care in the above embodiments can be achieved through… Figure 5 The computer device 500 shown is used to implement this, and the computer device 500 includes at least one processor 501, a memory 502 and at least one communication unit 505. The computer device 500 may be a terminal device, a server or a chip.
[0111] The processor 501 can be a general-purpose processor or a special-purpose processor. For example, the processor 501 can be a central processing unit (CPU). The CPU can be used to control the computer device 500, execute software programs, and process data from the software programs. The computer device 500 may also include a communication unit 505 for inputting (receiving) and outputting (transmitting) signals.
[0112] For example, computer device 500 may be a chip, communication unit 505 may be the input and / or output circuit of the chip, or communication unit 505 may be the communication interface of the chip, and the chip may be a component of terminal device, network device or other device.
[0113] For example, computer device 500 may be a terminal device or a server, and communication unit 505 may be a transceiver of the terminal device or the server, or communication unit 505 may be a transceiver circuit of the terminal device or the server.
[0114] The computer device 500 may include one or more memories 502 storing a program 504. The program 504 can be executed by a processor 501 to generate instructions 503, causing the processor 501 to perform the methods described in the above method embodiments according to the instructions 503. Optionally, the memory 502 may also store data (such as a target audit model). Optionally, the processor 501 may also read data stored in the memory 502, which may be stored at the same storage address as the program 504, or the data may be stored at a different storage address than the program 504.
[0115] The processor 501 and memory 502 can be configured separately or integrated together, for example, integrated on the system-on-chip (SOC) of the terminal device.
[0116] It should be understood that each step of the above method embodiment can be completed by hardware logic circuits or software instructions in the processor 501. The processor 501 can be a CPU, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, such as discrete gate, transistor logic devices, or discrete hardware components.
[0117] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0118] For example, in some embodiments, this application also provides a computer-readable storage medium storing instructions or code that, when executed on a computer, cause the computer to implement the above-described voice control method for an assistive turning device for elderly care.
[0119] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.
[0120] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A voice control method for an assistive turning device for elderly care, characterized in that, The method includes the following steps: Obtain a unique acoustic feature vector that includes the breathy voice characteristics, weak voice characteristics, and dialect characteristics of the elderly; The robustness weights of speech signals under different nursing action scenarios are determined by combining the specific acoustic feature vector with the sudden noise changes of the elderly in different nursing environments. The robustness weights represent the stability score of speech command recognition of the elderly under different nursing action scenarios. Based on the nursing voice commands of the elderly in the current nursing action scenario and the exclusive acoustic feature vector, the semantic state features of the elderly in the current nursing action scenario are determined. Based on the semantic state features and the robustness weight of the voice signals in different nursing action scenarios, the effective value of the nursing voice commands of the elderly in the current nursing action scenario is determined. The nursing voice commands of the elderly in the current nursing action scenario are classified into action categories to obtain the corresponding turning action category of the elderly in the current nursing action scenario, and then the action logic relationship between the turning action category and the current turning action state of the auxiliary turning device is determined. The credibility of the elderly person's voice control command execution in the current nursing action scenario is determined by the recognition valid value and the action logic association, and then the turning action control command of the assistive turning device is generated based on the command execution credibility. Specifically, determining the effective value of the nursing voice command for the elderly in the current nursing action scenario based on the semantic state features and the robustness weights of the voice signals under different nursing action scenarios includes: The semantic state features are combined with the robustness weights of the speech signals under different nursing action scenarios for evaluation to obtain the collaborative score between the semantic state features and the robustness weights. The effective value of the recognition of nursing voice commands for the elderly in the current nursing action scenario is determined based on the collaborative score.
2. The method as described in claim 1, characterized in that, The robustness weights of speech signals under different nursing action scenarios are determined by combining the specific acoustic feature vectors with sudden noise changes in elderly people under different nursing environments. Specifically, this includes: The noise interference experienced by elderly individuals in different care environments was determined by sudden changes in noise levels. The elderly’s voice signals collected under different nursing action scenarios are compared with the exclusive acoustic feature vector to obtain the acoustic feature matching error of the elderly’s voice signals under the noise interference. The robustness weights of speech signals under different nursing action scenarios are determined by the acoustic feature matching error.
3. The method as described in claim 1, characterized in that, The nursing voice commands given by the elderly in the current nursing action scenario, combined with the specific acoustic feature vector, specifically include: Acquire the nursing voice commands from the elderly in the current nursing action scenario; The nursing voice commands are enhanced based on the specific acoustic feature vector; The semantic state features of elderly people in the current nursing action scenario are determined by the nursing voice commands with enhanced understanding.
4. The method as described in claim 1, characterized in that, The nursing voice commands given to the elderly in the current nursing action scenario are classified into action categories. The specific turning action categories corresponding to the elderly in the current nursing action scenario include: The nursing voice commands given by the elderly in the current nursing action scenario are converted into text to obtain the written expression of the elderly's voice; Based on the textual expression of the elderly person's speech, keywords are extracted and semantics are matched for the nursing voice instructions to obtain the core action words of the elderly person in the current nursing action scenario; By identifying the meaning of the elderly person's actions in the current nursing scenario using the core action words, the corresponding turning-over action category for the elderly person in the current nursing scenario can be obtained.
5. The method as described in claim 1, characterized in that, Determining the action logic association between the type of turning over and the current turning over state of the assisted turning over device specifically includes: Read the current turning motion status of the assisted turning device; Determine the association between the current rolling-over action state and the rolling-over action category; The associated state establishes a logical association between the rolling-over action category and the current rolling-over action state.
6. The method as described in claim 1, characterized in that, Determining the reliability of voice control command execution by the elderly in the current nursing action scenario through the identification of valid values and the association of action logic specifically includes: The executable strength of voice control for the elderly in the current nursing action scenario is determined by the action logic association. The reliability of the elderly person's voice control command execution in the current nursing action scenario is determined based on the executable strength and the recognized valid value.
7. The method as described in claim 1, characterized in that, The specific instructions for generating the turning-over action control commands for the assisted turning-over device based on the reliability of the instructions include: Preset trusted execution threshold; When the credibility of the instruction execution is greater than or equal to the credibility execution threshold, the nursing voice instruction of the elderly in the current nursing action scenario is determined to be executable, and then the turning action control instruction of the assistive turning device is generated. When the credibility of the instruction execution is less than the credibility execution threshold, the nursing voice instruction for the elderly in the current nursing action scenario is marked as unexecutable, and the elderly are prompted to reconfirm through voice feedback.
8. The method as described in claim 1, characterized in that, The sudden noise changes in the nursing environment refer to background noise sources that suddenly appear and have unstable spectral characteristics during the nursing process.
9. An assistive turning device for elderly care, comprising a voice control unit, wherein the device includes a voice control unit, characterized in that, The voice control unit includes: The acquisition module is used to acquire exclusive acoustic feature vectors containing the breathy voice characteristics, weak voice characteristics, and dialect characteristics of the elderly; The processing module is used to determine the robustness weights of the speech signal under different nursing action scenarios by combining the dedicated acoustic feature vector with the sudden noise changes of the elderly in different nursing environments. The processing module is further configured to determine the semantic state features of the elderly in the current nursing action scenario based on the nursing voice commands of the elderly in the current nursing action scenario and the exclusive acoustic feature vector, and to determine the effective recognition value of the nursing voice commands of the elderly in the current nursing action scenario based on the semantic state features and the robustness weight of the voice signals in different nursing action scenarios. The processing module is also used to determine the action category of the nursing voice commands of the elderly in the current nursing action scenario, obtain the turning action category of the elderly in the current nursing action scenario, and then determine the action logic association between the turning action category and the current turning action state of the auxiliary turning device. The execution module is used to determine the credibility of the elderly person's voice control command execution in the current nursing action scenario by the recognition valid value and the action logic association, and then generate the turning action control command of the assisted turning device based on the command execution credibility.
Citation Information
Patent Citations
Voice recognition method and device
CN110875039A
Intelligent health robot combining intelligent conversation and health intervention
CN120183386A