Auxiliary turning-over equipment for nursing old people and voice control method thereof
By obtaining the feature vectors of the elderly's breathy voice, weak voice and dialect, and combining them with the noise in the nursing environment, fault-tolerant execution of elderly voice control is achieved, solving the problem of low voice recognition accuracy in nursing scenarios and improving the intelligence and safety of the equipment.
Patent Information
- Application Number
- CN202511300698.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-12
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-09-12
AI Technical Summary
Existing speech recognition technology is difficult to adapt to the voice characteristics of the elderly, especially in the case of noise interference in nursing scenarios, resulting in low recognition accuracy and the risk of misoperation.
By obtaining the exclusive acoustic feature vectors of the elderly's breathy voice, weak voice and dialect characteristics, combined with the noise changes in the nursing environment, the robustness weight of the voice signal is determined, and the semantic state features and device state logic are integrated to achieve fault-tolerant voice command execution.
It improves the accuracy and stability of voice control for the elderly, avoids misoperation, and enhances the intelligence and safety of the nursing process.
Smart Images

Figure CN120808768A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of voice control, more particularly, the present application relates to an auxiliary turning-over device for the care of the elderly and a voice control method thereof. BACKGROUND
[0002] Voice control refers to a control method in which a user issues an action command to a device through a voice instruction (such as "turn left", "stop", "increase angle"), and a natural language is converted into a control signal executable by the device through a voice recognition method; an auxiliary turning-over device refers to an electric or mechanical structure installed on a hospital bed or a nursing bed, which is used to help the elderly automatically perform left and right turning-over, adjust the lying angle, and the like, so as to reduce the risk of bedsores and the workload of nursing personnel.
[0003] The voice control method of the auxiliary turning-over device for the care of the elderly refers to an interactive method in which a voice recognition (converting a voice into a machine-recognizable instruction) is used to realize the operation control of the auxiliary turning-over device for the care of the elderly, so that the elderly with limited mobility can trigger the device to complete the turning-over action or adjust the related functions only by speaking without manually pressing a key, touching a screen, or relying on the assistance of others. It is a humanized function designed to improve the self-care ability of the elderly, simplify the operation process, and realize the self-control of the turning-over device by the elderly when they are conscious, so as to improve the care experience and enhance the intelligence and humanization of the care process. It is particularly suitable for the elderly with limited mobility or language expression difficulties, so that the elderly can control the device in the most natural "speaking" way without complex body movements. However, in the prior art, voice recognition is mainly targeted at ordinary adults and is not optimized for the voice characteristics of the elderly. The existing voice recognition model is difficult to adapt to the problems of voice distortion and unclear pronunciation caused by physiological characteristics (such as tooth loss, respiratory diseases, slow speech speed, and heavy accent) of the elderly. It lacks the ability to adapt to the common dialect expression, weak tone, or gas sound pronunciation of the elderly. At the same time, in the nursing scene, the environment is often accompanied by sudden interference sound (such as monitoring alarm, nursing staff conversation, bed movement noise, etc.), which is easy to cause voice recognition errors or missed judgments, resulting in low recognition accuracy of the voice of the elderly and the risk of misoperation. Therefore, how to realize fault-tolerant voice instruction execution in the care scene of the elderly based on the fusion control of voice recognition confidence and device state logic, so as to reduce the misrecognition rate when the elderly control the turning-over device by voice, has become a difficult problem in the industry. SUMMARY
[0004] The present application provides an auxiliary turning-over device for the care of the elderly and a voice control method thereof, which can realize fault-tolerant voice instruction execution in the care scene of the elderly based on the fusion control of voice recognition confidence and device state logic.
[0005] In a first aspect, the application provides a voice control method for an assisted turning-over device for the elderly, the method comprising the following steps: obtaining an exclusive acoustic feature vector containing the acoustic characteristics, weak voice characteristics and dialect characteristics of the elderly; determining the robustness weight of the voice signal in different nursing action scenarios from the exclusive acoustic feature vector combined with the noise changes of the elderly in different nursing environments; determining the semantic state feature of the elderly in the current nursing action scenario from the exclusive acoustic feature vector according to the nursing voice instruction of the elderly in the current nursing action scenario, and determining the recognition effective value of the nursing voice instruction of the elderly in the current nursing action scenario according to the semantic state feature and the robustness weight of the voice signal in different nursing action scenarios; determining the corresponding turning-over action category of the elderly in the current nursing action scenario through action category determination of the nursing voice instruction of the elderly in the current nursing action scenario, and then determining the action logic association between the turning-over action category and the current turning-over action state of the assisted turning-over device; determining the instruction execution credibility of the voice control of the elderly in the current nursing action scenario through the recognition effective value and the action logic association, and then generating the turning-over action control instruction of the assisted turning-over device according to the instruction execution credibility.
[0006] In some embodiments, determining the robustness weight of the voice signal in different nursing action scenarios from the exclusive acoustic feature vector combined with the noise changes of the elderly in different nursing environments specifically comprises: determining the noise interference of the elderly in different nursing environments from the noise changes of the elderly in different nursing environments; comparing the voice signal of the elderly collected in different nursing action scenarios with the exclusive acoustic feature vector to obtain the acoustic feature matching error of the voice signal of the elderly under the noise interference; determining the robustness weight of the voice signal in different nursing action scenarios through the acoustic feature matching error.
[0007] In some embodiments, determining the semantic state feature of the elderly in the current nursing action scenario from the exclusive acoustic feature vector according to the nursing voice instruction of the elderly in the current nursing action scenario specifically comprises: obtaining the nursing voice instruction of the elderly in the current nursing action scenario; enhancing the understanding of the nursing voice instruction based on the exclusive acoustic feature vector; determining the semantic state feature of the elderly in the current nursing action scenario from the nursing voice instruction after understanding enhancement.
[0008] In some embodiments, the recognition effective value of the nursing voice instruction of the elderly in the current nursing action scene is determined according to the semantic state feature and the robustness weight of the voice signal in different nursing action scenes, specifically comprising: Combining the semantic state feature with the robustness weight of the voice signal in different nursing action scenes for combined evaluation to obtain the synergy score between the semantic state feature and the robustness weight; Determine the recognition effective value of the nursing voice instruction of the elderly in the current nursing action scene based on the synergy score.
[0009] In some embodiments, the action category of the nursing voice instruction of the elderly in the current nursing action scene is determined to obtain the corresponding turning-over action category of the elderly in the current nursing action scene, specifically comprising: Text conversion is performed on the nursing voice instruction of the elderly in the current nursing action scene to obtain the written expression of the elderly voice; Based on the written expression of the elderly voice, keyword extraction and semantic matching are performed on the nursing voice instruction to obtain the core action word of the elderly in the current nursing action scene; The meaning of the elderly in the current nursing action scene is determined through the core action word to obtain the corresponding turning-over action category of the elderly in the current nursing action scene.
[0010] In some embodiments, the action logic association between the turning-over action category and the current turning-over action state of the auxiliary turning-over device is determined, specifically comprising: Read the current turning-over action state of the auxiliary turning-over device; Determine the association state of the current turning-over action state and the turning-over action category; The action logic association between the turning-over action category and the current turning-over action state is constructed through the association state.
[0011] In some embodiments, the instruction execution credibility of the voice control of the elderly in the current nursing action scene is determined through the recognition effective value and the action logic association, specifically comprising: Determine the executable intensity of the voice control of the elderly in the current nursing action scene through the action logic association; Determine the instruction execution credibility of the voice control of the elderly in the current nursing action scene based on the executable intensity and the recognition effective value.
[0012] In some embodiments, the turning-over action control instruction of the auxiliary turning-over device is generated according to the instruction execution credibility, specifically comprising: Pre-set a trusted execution threshold; When the instruction execution credibility is greater than or equal to the credible execution threshold, the care voice instruction of the old person in the current care action scene is determined as executable, and then a turning-over action control instruction of the auxiliary turning-over device is generated; When the instruction execution credibility is less than the credible execution threshold, the care voice instruction of the old person in the current care action scene is marked as unexecutable, and the old person is prompted to reconfirm through voice feedback.
[0013] In some embodiments, the sudden noise change in the care environment refers to a background noise source that suddenly appears in the care process and has unstable spectral characteristics.
[0014] In a second aspect, the application provides an auxiliary turning-over device for old people care, comprising a voice control unit, wherein the voice control unit comprises: The acquisition module is configured to acquire an exclusive acoustic feature vector containing the voice characteristics, the weak voice characteristics and the dialect characteristics of the old person. The processing module is configured to determine the robustness weight of the voice signal in different care action scenes by combining the exclusive acoustic feature vector with the sudden noise change of the old person in different care environments. The processing module is further configured to determine the semantic state feature of the old person in the current care action scene by combining the exclusive acoustic feature vector with the care voice instruction of the old person in the current care action scene, and determine the recognition effective value of the care voice instruction of the old person in the current care action scene according to the semantic state feature and the robustness weight of the voice signal in different care action scenes. The processing module is further configured to determine the turning-over action category corresponding to the old person in the current care action scene by performing action category determination on the care voice instruction of the old person in the current care action scene, and then determine the action logic association between the turning-over action category and the current turning-over action state of the auxiliary turning-over device. The execution module is configured to determine the instruction execution credibility of the voice control of the old person in the current care action scene by the recognition effective value and the action logic association, and then generate the turning-over action control instruction of the auxiliary turning-over device according to the instruction execution credibility.
[0015] The technical scheme provided by the embodiments of the application has the following beneficial effects: In the present application, by acquiring a special acoustic feature vector containing the acoustic characteristics of the elderly, weak voice characteristics and dialect characteristics; the robustness weight of the voice signal in different nursing action scenes is determined by the special acoustic feature vector combined with the noise change of the elderly in different nursing environments; according to the nursing voice instruction of the elderly in the current nursing action scene combined with the special acoustic feature vector, the semantic state feature of the elderly in the current nursing action scene is determined, and the recognition effective value of the nursing voice instruction of the elderly in the current nursing action scene is determined according to the semantic state feature and the robustness weight of the voice signal in different nursing action scenes; the action category of the nursing voice instruction of the elderly in the current nursing action scene is determined, and the corresponding turning-over action category of the elderly in the current nursing action scene is obtained, and then the action logic association between the turning-over action category and the current turning-over action state of the auxiliary turning-over equipment is determined; the instruction execution credibility of the voice control of the elderly in the current nursing action scene is determined through the recognition effective value and the action logic association, and then the turning-over action control instruction of the auxiliary turning-over equipment is generated according to the instruction execution credibility.
[0016] It can be seen that, in the present application, first, the recognition effective value of the nursing voice instruction of the elderly in the current nursing action scene is determined according to the semantic state feature and the robustness weight of the voice signal in different nursing action scenes, which effectively quantifies the recognizable degree of the current voice instruction in a complex environment, can avoid errors caused by insufficient confidence or ambiguous semantics, significantly enhances the controllability and fault tolerance of the voice recognition process of the elderly, and is the basic guarantee for subsequent voice instruction execution logic judgment; secondly, the action logic association between the turning-over action category and the current turning-over action state of the auxiliary turning-over device is determined, and whether the voice request of the elderly has execution significance (such as avoiding repeated left turning, direction conflict, etc.) is judged through the action logic, which can effectively compensate for the misjudgment caused by auditory interference or unclear expression in the voice recognition process, avoid the direct transmission of voice recognition errors to the auxiliary turning-over device, and reflect the constraint fusion of the voice control system on the behavior logic of the device, so as to realize fault-tolerant control and error operation avoidance; then, the instruction execution credibility of the voice control of the elderly in the current nursing action scene is determined through the recognition effective value and the action logic association, which can effectively avoid the judgment of whether to execute only according to the recognition effectiveness, improve the intelligence and stability of the voice control strategy, for example, when the voice expression is relatively ambiguous but the instruction is highly consistent with the current device state, it can also be judged as "credible", and the actual fault-tolerant control ability of the voice interaction process of the elderly in the nursing scene under multi-source interference is significantly improved through the decision mechanism of the instruction execution credibility; finally, the turning-over action control instruction of the auxiliary turning-over device is generated according to the instruction execution credibility, which can effectively block the false triggering when facing the weak voice, panting speech or reversed speech sequence of the elderly, avoid the disconnection between the action of the auxiliary turning-over device and the intention of the elderly user, and cause the risk of injury, further perfect the fault-tolerant control closed loop, and ensure that the intelligent response process of the auxiliary turning-over device is more stable, safe and humanized; in summary, the scheme can realize fault-tolerant voice instruction execution in the nursing scene of the elderly based on the fusion control of voice recognition confidence and device state logic. BRIEF DESCRIPTION OF DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and those skilled in the art can also obtain other drawings according to these drawings without creative labor.
[0018] Figure 1 is an exemplary flow chart of a voice control method for an auxiliary turning-over device for elderly care according to some embodiments of the present application; Figure 2 is an exemplary flow chart of determining a semantic state feature according to some embodiments of the present application; Figure 3 is an exemplary flow chart for implementing action category determination according to some embodiments of the present application; Figure 4 is a schematic structural diagram of a voice control unit according to some embodiments of the present application; Figure 5 It is a structural diagram of a computer device for implementing a voice control method for an auxiliary turning device for elderly care as shown in some embodiments of the present application. DETAILED DESCRIPTION
[0019] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0020] refer to Figure 1 , which is an exemplary flow chart of a method for voice control of an auxiliary turning device for elderly care according to some embodiments of the present application. The method mainly includes the following steps: In step 101, a unique acoustic feature vector including the breath characteristics, weak voice characteristics and dialect characteristics of the elderly is obtained.
[0021] In a specific implementation, the exclusive acoustic feature vector containing the senile vocalization characteristics, weak voice characteristics and dialect characteristics can be obtained in the following manner: first, a voice collection device equipped with an array microphone is used to collect the natural voice instructions of the elderly during the care process in multiple channels, and the collected raw voice data of the elderly is subjected to noise segment elimination and voice segment calibration by a voice activity detection (VAD) method, then a multi-channel voice enhancement technology such as a fusion method based on beamforming and spectral subtraction is applied to the voice segment to suppress background noise and highlight the voice main signal of the elderly, then the short-time Fourier transform (STFT) spectrogram is extracted from the enhanced voice signal, and further acoustic features such as Mel-frequency cepstral coefficients (MFCC), formant distribution, voice energy envelope are calculated; the deep neural network model (such as BiLSTM-Attention) with attention mechanism is constructed to identify the vocalization characteristics (i.e. the low-energy interval feature with high proportion of exhalation component caused by incomplete closure of vocal cords), weak voice characteristics (i.e. the sound intensity distribution feature with low overall voice intensity and narrow dynamic range) and dialect characteristics (i.e. the phonetic substitution, tone shift and language speed change on the phonetic unit level) in the voice of the elderly, and finally the above three types of label features of vocalization, weak voice and dialect are multi-dimensionally fused with the corresponding acoustic features to construct a multi-dimensional vector representing the voice characteristics of the individual elderly, which is used as the exclusive acoustic feature vector of the elderly for subsequent semantic state modeling and recognition reliability evaluation; in other embodiments, other methods can also be used for obtaining, which are not limited here.
[0022] It should be noted that the vocalization characteristics in the present application refer to the phenomenon of voice with significant exhalation noise and high-frequency energy diffusion in the spectrum, which reflects the weak pronunciation of the elderly; the weak voice characteristics in the present application refer to the feature of low overall sound pressure level, which reflects the weak tone of the sick elderly; the dialect characteristics in the present application refer to the phonetic or lexical variation characteristics produced in a specific regional background, which has interference to the standard Mandarin model; in addition, the exclusive acoustic feature vector in the present application is a multi-dimensional numerical representation for representing the pronunciation mode, voice signal structure and variation characteristics of the individual elderly, which is used as the input key data of the individualized voice recognition ability of the elderly model, and through the exclusive acoustic feature vector, the adaptability and accuracy of subsequent semantic analysis and action instruction judgment under the input of the voice of the elderly can be enhanced.
[0023] In step 102, the robustness weight of the speech signal in different nursing action scenarios is determined by the specific acoustic feature vector in combination with the noise change suddenly occurring to the elderly in different nursing environments.
[0024] In some embodiments, the robustness weight of the speech signal in different nursing action scenarios can be determined by the specific acoustic feature vector in combination with the noise change suddenly occurring to the elderly in different nursing environments, which can be achieved by the following steps: The noise interference of the elderly in different nursing environments is determined by the noise change suddenly occurring to the elderly in different nursing environments. The speech signal of the elderly collected in different nursing action scenarios is compared with the specific acoustic feature vector to obtain the acoustic feature matching error of the speech signal of the elderly under the noise interference. The robustness weight of the speech signal in different nursing action scenarios is determined by the acoustic feature matching error.
[0025] In specific implementation, the noise interference of the elderly in different nursing environments can be determined by sudden noise changes in different nursing environments. This can be achieved in the following way: the acoustic environment perception technology can be used to monitor in real time the sudden noise sources that may occur during the nursing process of the elderly, such as the noise of nursing staff operating equipment, bed movement, oxygen machine working, etc., by setting up multiple microphone arrays for omnidirectional sound sampling, and using the short-time energy detection method (Short-Time Energy) combined with the spectral entropy analysis method (Spectral Entropy Analysis) to detect the noise interference of the elderly in different nursing environments. Entropy) performs sudden feature detection on the collected audio to calibrate the noise intensity and frequency range of sudden noise changes, and use it as noise interference for the elderly in different nursing environments; the elderly's voice signals collected in different nursing action scenarios are compared with the exclusive acoustic feature vector to obtain the acoustic feature matching error of the elderly's voice signal under the noise interference. This can be achieved in the following way, namely: collecting the nursing instruction voice actually issued by the elderly in different nursing action scenarios (such as getting up, turning over, and lying flat), and matching and comparing it with the exclusive acoustic feature vector frame by frame, and then using the dynamic time warping algorithm and the spectral difference measurement method (such as the logarithmic spectrum distance) in combination with the noise intensity and frequency range calibrated in the noise interference to evaluate the acoustic feature matching error between the elderly's voice signal and the exclusive acoustic feature vector; The robustness weight of the speech signal in different nursing action scenarios determined by the acoustic feature matching error can be achieved in the following manner, namely: the acoustic feature matching error under different noise interferences can be used as an input indicator through a multi-channel robustness evaluation function, and the energy retention and spectral morphology stability of the speech signal can be combined to comprehensively output the robustness score of the speech signal in different nursing action scenarios, and then the score is normalized to generate a numerical value between 0 and 1 as the robustness weight of the speech signal in different nursing action scenarios, which can be used to measure the stability and recognizability of the speech signal in different nursing action scenarios under sudden noise interference, and can be used as an adaptive compensation factor for subsequent nursing scenarios to improve the stability and execution reliability of voice command recognition for the elderly under high noise conditions; other methods can also be used for determination in other embodiments, which are not limited here.
[0026] It should be noted that the sudden noise changes in the nursing environment in this application refer to background noise sources that suddenly appear and have unstable spectral characteristics during the nursing process; the nursing action scenes in this application represent the operating states at different stages of the elderly care process, such as before turning over, during turning over, and during nursing intervention; the noise interference in this application refers to non-target sound components that interfere with speech recognition in the elderly care environment, which is used to identify the external pressure conditions faced by the elderly voice command recognition; the acoustic feature matching error in this application represents the degree of deviation between the elderly voice and the standard features in the time or frequency domain, which is used to evaluate the impact of nursing environment noise on the stability of the elderly voice expression; the robustness weight in this application represents the stability score of the elderly voice command recognition under different nursing action scenarios, and its role is to improve the accuracy and reliability of the subsequent elderly voice command recognition stage in non-ideal environments.
[0027] In step 103, the semantic state characteristics of the elderly in the current nursing action scenario are determined based on the nursing voice instructions of the elderly in the current nursing action scenario combined with the exclusive acoustic feature vector, and the recognition effective value of the nursing voice instructions of the elderly in the current nursing action scenario is determined based on the semantic state characteristics and the robustness weights of the voice signals in different nursing action scenarios.
[0028] In some embodiments, reference Figure 2 As shown in FIG, this figure is an exemplary flow chart for determining semantic state features in some embodiments of the present application. In this embodiment, the semantic state features of the elderly in the current nursing action scenario are determined based on the nursing voice instructions of the elderly in the current nursing action scenario in combination with the exclusive acoustic feature vector, which can be achieved by the following steps: First, in step 1031, the nursing voice instructions of the elderly in the current nursing action scenario are obtained; Next, in step 1032 , the nursing voice instruction is enhanced in understanding based on the exclusive acoustic feature vector; Finally, in step 1033, the semantic state characteristics of the elderly person in the current nursing action scenario are determined based on the nursing voice instructions after understanding enhancement.
[0029] In a specific implementation, the obtaining of the care voice instruction of the elderly in the current care action scene can be implemented in the following manner: the care voice instruction actively issued by the elderly in the current care action (such as turning over or assisting in lifting) is obtained through a voice collection device in the care scene, for example, “help me turn over”, “a little to the left”, etc.; the understanding enhancement of the care voice instruction based on the exclusive acoustic feature vector can be implemented in the following manner: the care voice instruction can be divided into short frames, the basic acoustic parameters (such as MFCC, zero-crossing rate, energy spectrum, etc.) and the frame-level time information of the elderly care voice sequence are extracted, then a pre-trained acoustic model (such as a Conformer model or a Factorized Time-Delay Neural Network (TDNN-F structure)) and a language model (such as a Bidirectional Encoder Representations from Transformers (BERT) or a Long Short-Term Memory (LSTM)) joint decoder are called to fuse the context semantics and voice sequence signals of the care voice instruction to obtain a preliminary semantic vector, the preliminary semantic vector is cross-aligned with the exclusive acoustic feature vector of the elderly (including the air sound frequency band, weak sound intensity, and dialect pronunciation mode, etc.) and attention weighted modeling is performed, the acoustic consistency attention mechanism is used to enhance the understanding ability of the air sound fuzzy boundary, unstable intonation or dialect expression, so as to correct the features of the care voice instruction, reduce the deviation caused by the speed of pronunciation, and strengthen the key semantic information in the voice content of the elderly, so as to realize the understanding enhancement of the care voice instruction. The attention weighted modeling refers to using the attention mechanism to strengthen the mapping between the key voice frame and the feature vector, and improve the response sensitivity of semantic analysis to complex voice features; the determination of the semantic state feature of the elderly in the current care action scene from the care voice instruction after understanding enhancement can be implemented in the following manner: the semantic prior knowledge of the care voice instruction after understanding enhancement and the current care action scene label (such as “turning over before sleeping” or “sudden call”) can be extracted through a graph convolutional neural network or a semantic scene reasoning network to extract the semantic direction and state emotion of the care voice instruction in the current care action scene, for example, the intention category, context emotion, semantic coherence, tone intensity, expression confidence, etc. of the care voice instruction of the elderly, and the finally extracted result is taken as the semantic state feature of the elderly in the current care action scene; in other embodiments, other methods can also be used for implementation, which are not limited here.
[0030] It should be noted that the nursing voice instruction in the current nursing action scenario in the present application represents the voice content that the elderly actively express their needs in the nursing process under the current nursing operation; the understanding enhancement in the present application refers to the process of adjusting and enhancing the acoustic and semantic content of the current nursing voice instruction of the elderly by combining personalized acoustic characteristics, so as to improve the recognizability and semantic expression clarity of the nursing voice instruction; the semantic state feature in the present application is used to represent the voice intention that the elderly want to express under the current nursing action scenario, which is an important basis for subsequent nursing action logic judgment and execution control.
[0031] In some embodiments, the determination of the recognition effective value of the nursing voice instruction of the elderly under the current nursing action scenario according to the semantic state feature and the robustness weight of the voice signal under different nursing action scenarios can be implemented by the following steps: combining and evaluating the semantic state feature and the robustness weight of the voice signal under different nursing action scenarios to obtain a synergy score between the semantic state feature and the robustness weight; determining the recognition effective value of the nursing voice instruction of the elderly under the current nursing action scenario based on the synergy score.
[0032] In specific implementation, the combination and evaluation of the semantic state feature and the robustness weight of the voice signal under different nursing action scenarios to obtain a synergy score between the semantic state feature and the robustness weight can be implemented in the following manner, i.e., the semantic state feature and the robustness weight of the voice signal under the current corresponding nursing action scenario can be input into a fusion evaluation model, which can adopt a Multi-Layer Perceptron (MLP) structure or a weighted scoring model based on an attention mechanism (such as a transformer encoder) to combine and process the input semantic information intensity and anti-interference ability; in the combination and evaluation stage, the model comprehensively scores the similarity, complementarity and dynamic coordination relationship between the semantic state feature and the robustness weight according to a pre-trained voice semantic alignment function, so as to calculate the synergy score between the semantic state feature and the robustness weight; the determination of the recognition effective value of the nursing voice instruction of the elderly under the current nursing action scenario based on the synergy score can be implemented in the following manner, i.e., the synergy score can be input into a linear mapping layer or a sigmoid normalization function to map the synergy score to the interval [0, 1], and the finally obtained value is taken as the recognition effective value of the nursing voice instruction of the elderly under the current nursing action scenario, wherein the higher the value, the more reliable and suitable the nursing voice instruction is under the current nursing action scenario as an execution input of voice uprising control; in other embodiments, other methods can also be used for determination, which is not limited here.
[0033] It should be noted that the collaborative score in this application reflects the degree to which the elderly care voice instructions have both semantic clarity and anti-interference stability in the current nursing action scenario; the combined evaluation in this application refers to the process of jointly processing the semantic state features and the robustness weights, and quantitatively scoring their joint reliability through a fusion evaluation model, which can be used to comprehensively judge the credibility of the elderly care voice instructions; the recognition effective value in this application is used to reflect the reliability of the elderly care voice instructions being accurately recognized and semantically judged in the current nursing action scenario, which is the key threshold judgment basis before executing the voice control logic.
[0034] In step 104, the action category of the nursing voice instructions of the elderly in the current nursing action scenario is determined to obtain the turning action category corresponding to the elderly in the current nursing action scenario, and then determine the action logical association between the turning action category and the current turning action state of the auxiliary turning device.
[0035] In some embodiments, reference Figure 3 As shown in FIG, this figure is an exemplary flow chart for implementing action category determination in some embodiments of the present application. Action category determination is performed on the nursing voice instructions of the elderly in the current nursing action scenario, and the corresponding turning action category of the elderly in the current nursing action scenario can be obtained by the following steps: Convert the elderly's nursing voice instructions in the current nursing action scenario into text to obtain the text expression of the elderly's voice; Perform keyword extraction and semantic matching on the nursing voice instructions based on the text expression of the elderly person's voice to obtain the core action words of the elderly person in the current nursing action scenario; The meaning of the elderly person in the current nursing action scenario is judged by the core action words, and the corresponding turning action category of the elderly person in the current nursing action scenario is obtained.
[0036] In a specific implementation, the text conversion of the nursing voice instruction of the elderly in the current nursing action scene can be achieved by the following manner: a voice recognition model (such as an acoustic-linguistic integrated model based on Connectionist Temporal Classification (CTC) or Transformer structure) is used to transcribe the nursing voice instruction of the elderly in the current nursing action scene into a standard text expression sentence; the keyword extraction and semantic matching of the nursing voice instruction based on the text expression of the elderly voice can be achieved by the following manner: a pre-constructed action semantic mapping word table is used to extract the keyword and match the semantic of the transcribed text expression sentence of the voice of the elderly, and the core action words with instructional meaning are extracted, such as "turn over", "turn left", "turn right", "turn back", "turn back", "turn a little more", "turn left 30 degrees", etc.; the meaning of the elderly in the current nursing action scene is determined by the core action words, and the corresponding turning over action category of the elderly in the current nursing action scene is obtained by the following manner: a trained action classification model, such as a multi-classifier based on BERT or BiLSTM, is used to input the core action words into the action classification model to determine the meaning of the elderly in the current nursing action scene, and determine the corresponding turning over action category of the elderly in the current nursing action scene, such as "turn over from left side", "turn over from right side", "turn back from supine position", etc.; in other embodiments, other methods can also be used for implementation, which are not limited here.
[0037] It should be noted that the action category determination in the present application refers to the process of mapping the recognized voice information of the elderly to a specific action type according to a preset classification standard; the text expression in the present application refers to the result of transcribing the nursing voice instruction of the elderly into a natural language text by voice recognition technology, which is a readable sentence or phrase; the core action word in the present application refers to the key verb or verb phrase extracted from the transcribed text, which represents the core control intention in the nursing voice instruction of the elderly; the turning over action category in the present application refers to the standardized turning over action corresponding to the nursing voice instruction in the current nursing action scene, which is the final execution target of the auxiliary turning over device control instruction; in addition, the semantic alignment and context awareness of this step improves the accuracy of voice control in real nursing scene and the fault tolerance of instruction recognition.
[0038] In some embodiments, the determination of the action logic association between the turning over action category and the current turning over action state of the auxiliary turning over device can be achieved by the following steps: reading a current turning-over action state of the turning-over assisting device; judging an association state between the current turning-over action state and the turning-over action category; constructing an action logic association between the turning-over action category and the current turning-over action state through the association state.
[0039] In a specific implementation, reading a current turning-over action state of the turning-over assisting device can be implemented in the following manner: the current turning-over action state of the turning-over assisting device can be acquired in real time by a built-in action state reading module (such as a position sensor, an angle encoder, or an execution mechanism feedback system) of the turning-over assisting device, which includes the current turning-over direction (such as left lateral recumbency, right lateral recumbency, supine position) and the execution progress (such as completed, executing, standby) of the device, thereby reading the current turning-over action state of the turning-over assisting device. Judging an association state between the current turning-over action state and the turning-over action category can be implemented in the following manner: a state-action mapping model bound to the action structure of the turning-over assisting device can be used to judge whether the target turning-over action (i.e., the turning-over action category corresponding to the old person in the current nursing action scene) conflicts with the current device state (such as the current device state is left lateral recumbency and the user issues a "left turning" instruction again), to obtain the determination results of the current turning-over action state and the turning-over action category as consistent, conflicting, continuous, and repetitive, as the association state, to judge whether there is a reasonable logic relationship between the turning-over action category and the current turning-over action state of the device. Constructing an action logic association between the turning-over action category and the current turning-over action state through the association state can be implemented in the following manner: if the association state indicates that the target turning-over action and the current device state can be logically continuously executed, it is determined that the action logic association between the turning-over action category and the current turning-over action state is valid, which indicates that the feasibility and rationality degree of the turning-over action conversion is higher (i.e., the association weight is higher), and if it is conflicting or redundant, a feasible transition action (such as "first return to the home position, then right turning", "the current device has turned right by 30 degrees, and still needs to turn right by 15 degrees") can be adjusted by an action scheduling logic module (such as a state machine judgment or a rule engine), to construct a logic relationship diagram between the turning-over action category and the current turning-over action state as the action logic association therebetween, and to assign an association weight (the weight indicates the feasibility and rationality degree of the turning-over action conversion, and takes a value between 0 and 1), to form execution a priori information for reference by a subsequent control module, to improve the rationality and continuity of the execution action, and to effectively prevent the execution of repeated, conflicting, or invalid control instructions. In other embodiments, other methods can also be used for implementation, which are not limited here.
[0040] It should be noted that the current turning-over action state in the present application refers to the specific physical state description of the actual execution position and motion phase of the auxiliary turning-over device; the associated state in the present application represents the logical relationship type between the target turning-over action and the current turning-over state of the auxiliary turning-over device, such as consistent, conflict, continuous, repetition, etc. The action logic association in the present application refers to the execution dependency logic between the target turning-over action in the voice instruction of the elderly care and the current turning-over state of the auxiliary turning-over device, which is the intermediate bridge from the voice recognition action result of the elderly to the execution action of the control device, and is used to ensure that the action connection of the auxiliary turning-over device during operation is reasonable, safe and efficient.
[0041] In step 105, the execution confidence of the voice control instruction of the elderly in the current care action scene is determined by the identified effective value and the action logic association, and then the turning-over action control instruction of the auxiliary turning-over device is generated according to the execution confidence of the instruction.
[0042] In some embodiments, the execution confidence of the voice control instruction of the elderly in the current care action scene can be determined by the identified effective value and the action logic association by the following steps: The executable intensity of the voice control of the elderly in the current care action scene is determined by the action logic association; The execution confidence of the voice control instruction of the elderly in the current care action scene is determined based on the executable intensity and the identified effective value.
[0043] In a specific implementation, the executable intensity of the voice control of the elderly in the current nursing action scene can be determined based on the action logic association, for example, if the turning-over assisting device is in the process of turning over to the left and the nursing voice instruction of the elderly in the current nursing action scene is "continue to turn over to the left", the executable intensity can be set to between 0.8 and 1 (indicating that the action logic is reasonable), and if the nursing voice instruction at this time is "turn over to the right", the executable intensity can be set to between 0.4 and 0.5 (indicating that the action logic is somewhat conflicting); if the nursing voice instruction of the elderly is "turn over to the right" but the turning-over assisting device is currently at the right turning-over end point, the executable intensity is 0 (indicating that the action logic is conflicting), wherein the executable intensity is between 0 and 1; the instruction execution confidence of the voice control of the elderly in the current nursing action scene can be determined based on the executable intensity and the recognized effective value, which can be implemented in the following manner: the executable intensity and the recognized effective value can be weighted and fused by linear combination, fuzzy reasoning or a confidence fusion function (such as product fusion or softmax normalization, a fusion network based on attention weight), wherein the weighting coefficients can be set according to historical experience data (and the action safety is prioritized), and finally a confidence index representing whether the nursing voice instruction of the elderly in the current nursing action scene is worth being executed by the device is obtained, and it is ensured that the fusion output result is between 0 and 1 and has discriminability, and then the output confidence index is taken as the instruction execution confidence of the voice control of the elderly in the current nursing action scene, which is used to judge whether the current nursing voice instruction of the elderly is adopted in the subsequent voice control process; in other embodiments, other methods can also be used for determination, which is not limited here.
[0044] It should be noted that the executable intensity in the present application represents a quantitative value of the execution feasibility of the nursing voice instruction action in the current device turning-over state, which is used to constrain whether the voice recognition result can be reasonably executed in the current context to embody the state perception ability of the turning-over assisting device; the instruction execution confidence in the present application refers to a comprehensive score index of whether the current voice control instruction of the elderly is worth being executed, which is used to reflect the acceptability of the current voice instruction of the elderly and the coordination of the state of the turning-over assisting device, and it is used as the basis for finally judging whether the nursing voice instruction is executed, which can improve the reliability and accurate response ability of the voice control of the elderly.
[0045] In some embodiments, the turning-over action control instruction of the turning-over assisting device can be generated according to the instruction execution confidence in the following steps: a preset confidence execution threshold is set; When the instruction execution credibility is greater than or equal to the credible execution threshold, the nursing voice instruction of the old person in the current nursing action scene is determined as executable, and a turning-over action control instruction of the auxiliary turning-over device is generated; When the instruction execution credibility is less than the credible execution threshold, the nursing voice instruction of the old person in the current nursing action scene is marked as unexecutable, and the old person is prompted to reconfirm through voice feedback.
[0046] In a specific implementation, the preset credible execution threshold can be implemented in the following manner: based on sample training of a large number of real nursing scenes, the control instruction execution success rate distribution under different instruction execution credibilities is counted, and then a reasonable credible execution threshold (such as 0.7) is preset. In other embodiments, the threshold can also be set by other manners, which is not limited here. The turning-over action control instruction of the auxiliary turning-over device can be implemented in the following manner: when the instruction execution credibility is greater than or equal to the credible execution threshold, the nursing voice instruction of the old person in the current nursing action scene is determined as executable, and the corresponding turning-over action category is taken as an action type index. The current turning-over action state and the device action parameters (such as motion angle, speed, and duration) of the auxiliary turning-over device are matched with the preset action execution instruction to generate a specific turning-over action control instruction. The instruction includes a control target turning-over action (such as "left turning"), an execution parameter (such as "45 degrees, 30 seconds of slow speed"), and a control timing to drive the turning-over to complete the precise action. When the instruction execution credibility is less than the credible execution threshold, the nursing voice instruction of the old person in the current nursing action scene is marked as unexecutable, and the old person is prompted to reconfirm through voice feedback. When the instruction execution credibility is less than the credible execution threshold, the nursing voice instruction of the old person in the current nursing action scene is marked as unexecutable, and a friendly reminder is given to the old person by calling the voice feedback module, such as "please repeat your instruction" or "I did not hear your instruction clearly". The reconfirmation feedback voice is used to ensure that the voice interaction has a feedback loop and improve the human-machine communication efficiency. In other embodiments, other methods can also be used for implementation, which is not limited here.
[0047] It should be noted that the turning-over action control instruction in the present application refers to an executable command for the bottom control module of the auxiliary turning-over device, which is used to guide the auxiliary turning-over device to execute a specific nursing turning-over action.
[0048] In addition, another aspect of the present application provides an auxiliary turning-over device for old people nursing in some embodiments. The device includes a voice control unit, which is used to receive the nursing voice instruction of the old person in the current nursing action scene and determine the instruction execution credibility of the old person in the current nursing action scene. Figure 4, the figure is a structural schematic diagram of a voice control unit shown according to some embodiments of the present application, the voice control unit 400 comprises: an acquisition module 401, a processing module 402 and an execution module 403, which are described as follows: The acquisition module 401 is mainly used for acquiring a special acoustic feature vector containing the sound characteristics of the old people, the weak voice characteristics and the dialect characteristics in the present application; The processing module 402 is mainly used for determining the robustness weight of the voice signal in different nursing action scenes from the special acoustic feature vector combined with the noise change of the old people in different nursing environments in the present application; The processing module 402 is also used for determining the semantic state feature of the old people in the current nursing action scene according to the nursing voice instruction of the old people in the current nursing action scene combined with the special acoustic feature vector, and determining the recognition effective value of the nursing voice instruction of the old people in the current nursing action scene according to the semantic state feature and the robustness weight of the voice signal in different nursing action scenes; The processing module 402 is also used for determining the action category of the nursing voice instruction of the old people in the current nursing action scene, obtaining the corresponding turning-over action category of the old people in the current nursing action scene, and then determining the action logic association between the turning-over action category and the current turning-over action state of the auxiliary turning-over equipment. The execution module 403 is mainly used for determining the instruction execution credibility of the voice control of the old people in the current nursing action scene through the recognition effective value and the action logic association, and then generating the turning-over action control instruction of the auxiliary turning-over equipment according to the instruction execution credibility in the present application.
[0049] The above describes in detail the examples of the auxiliary turning-over equipment for old people nursing and the voice control method thereof provided by the embodiments of the present application. It can be understood that the corresponding device contains the corresponding hardware structure and / or software module for executing each function in order to realize the above functions. Those skilled in the art should easily realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in the present application can be realized in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed by hardware or computer software driven hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0050] In some embodiments, the present application also provides a computer device comprising a memory and a processor, the memory is used to store a computer program, and the processor is used to call and run the computer program from the memory, so that the computer device executes the voice control method of the auxiliary turning-over device for the elderly care described above.
[0051] In some embodiments, with reference to Figure 5 The dashed line in the figure indicates that the unit or the module is optional, and the figure is a structural schematic diagram of a computer device implementing the voice control method of the auxiliary turning-over device for the elderly care of the present application. The voice control method of the auxiliary turning-over device for the elderly care in the above-mentioned embodiments can be implemented by the computer device shown in the figure, which comprises at least one processor 501, a memory 502, and at least one communication unit 505. The computer device 500 can be a terminal device or a server or a chip. Figure 5
[0052] The processor 501 can be a general-purpose processor or a special-purpose processor. For example, the processor 501 can be a central processing unit (CPU), which can be used to control the computer device 500, execute software programs, and process data of the software programs. The computer device 500 can further comprise a communication unit 505 to realize input (reception) and output (transmission) of signals.
[0053] For example, the computer device 500 can be a chip, and the communication unit 505 can be an input and / or output circuit of the chip, or the communication unit 505 can be a communication interface of the chip. The chip can be a component of a terminal device or a network device or other devices.
[0054] For another example, the computer device 500 can be a terminal device or a server, and the communication unit 505 can be a transceiver of the terminal device or the server, or the communication unit 505 can be a transceiver circuit of the terminal device or the server.
[0055] The computer device 500 can comprise one or more memories 502, which have programs 504 stored thereon. The programs 504 can be run by the processor 501 to generate instructions 503, so that the processor 501 executes the methods described in the above-mentioned method embodiments according to the instructions 503. Optionally, the memory 502 can also store data (such as a target audit model). Optionally, the processor 501 can also read the data stored in the memory 502. The data can be stored in the same storage address as the programs 504, or the data can be stored in different storage addresses from the programs 504.
[0056] The processor 501 and the memory 502 can be separately arranged or integrated together, for example, on a system on chip (SOC) of the terminal device.
[0057] It should be understood that each step of the above method embodiments can be completed by a logic circuit in the form of hardware or instructions in the form of software in the processor 501, and the processor 501 can be a CPU, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, for example, discrete gates, transistor logic, or discrete hardware components.
[0058] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, etc.) containing computer-usable program code.
[0059] For example, in some embodiments, the present application also provides a computer-readable storage medium, which stores instructions or codes, when the instructions or codes are run on a computer, cause the computer to perform the above voice control method for the auxiliary turning-over device for the elderly care.
[0060] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications to the embodiments once they know the basic inventive concept. Therefore, the appended claims are intended to be interpreted as including all the preferred embodiments and all the changes and modifications falling within the scope of the present application.
[0061] Obviously, those skilled in the art can make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application also intends to include these modifications and variations.
Claims
1. A voice control method for an auxiliary turning device for elderly care, characterized in that: The method comprises the following steps: Obtaining a unique acoustic feature vector that includes the breathy voice characteristics, weak voice characteristics, and dialect characteristics of the elderly; Determine the robustness weight of the speech signal in different nursing action scenarios based on the exclusive acoustic feature vector and sudden noise changes of the elderly in different nursing environments; Determine the semantic state features of the elderly in the current nursing action scenario based on the nursing voice instructions of the elderly in the current nursing action scenario in combination with the exclusive acoustic feature vector, and determine the recognition effective value of the nursing voice instructions of the elderly in the current nursing action scenario based on the semantic state features and the robustness weights of the voice signals in different nursing action scenarios; Determine the action category of the nursing voice instruction of the elderly in the current nursing action scenario, obtain the turning action category corresponding to the elderly in the current nursing action scenario, and then determine the action logic association between the turning action category and the current turning action state of the auxiliary turning device; The recognition effective value and the action logic association are used to determine the execution credibility of the voice-controlled instructions for the elderly in the current nursing action scenario, and then the turning action control instructions of the auxiliary turning device are generated based on the instruction execution credibility.
2. The method according to claim 1, wherein The robustness weights of the speech signals under different nursing action scenarios are determined by combining the exclusive acoustic feature vector with sudden noise changes in different nursing environments for the elderly, specifically including: The noise disturbance of the elderly in different care environments is determined by the sudden noise changes in different care environments; Comparing the elderly's voice signals collected in different nursing action scenarios with the exclusive acoustic feature vectors to obtain the acoustic feature matching error of the elderly's voice signals under the noise interference; The robustness weight of the speech signal in different nursing action scenarios is determined by the acoustic feature matching error.
3. The method according to claim 1, wherein The nursing voice instructions for the elderly in the current nursing action scenario combined with the exclusive acoustic feature vector specifically include: Obtaining the elderly's nursing voice instructions in the current nursing action scenario; enhancing understanding of the nursing voice instruction based on the exclusive acoustic feature vector; The semantic state characteristics of the elderly in the current nursing action scenario are determined by the nursing voice instructions after understanding enhancement.
4. The method according to claim 1, wherein Determining the recognition effective value of the nursing voice instruction for the elderly in the current nursing action scenario based on the semantic state features and the robustness weights of the voice signals in different nursing action scenarios specifically includes: Combining and evaluating the semantic state feature with the robustness weight of the speech signal in different nursing action scenarios to obtain a synergy score between the semantic state feature and the robustness weight; The recognition effectiveness value of the nursing voice instructions for the elderly in the current nursing action scenario is determined based on the collaboration score.
5. The method according to claim 1, wherein The action category of the nursing voice instructions of the elderly in the current nursing action scenario is determined, and the corresponding turning action categories of the elderly in the current nursing action scenario are specifically as follows: Convert the elderly's nursing voice instructions in the current nursing action scenario into text to obtain the text expression of the elderly's voice; Perform keyword extraction and semantic matching on the nursing voice instructions based on the text expression of the elderly person's voice to obtain the core action words of the elderly person in the current nursing action scenario; The meaning of the elderly person in the current nursing action scenario is judged by the core action words, and the corresponding turning action category of the elderly person in the current nursing action scenario is obtained.
6. The method according to claim 1, wherein Determining the action logic association between the turning action category and the current turning action state of the turning auxiliary device specifically includes: Read the current turning action status of the auxiliary turning device; Determining the association between the current turning action state and the turning action category; An action logic association between the turning action category and the current turning action state is established through the association state.
7. The method according to claim 1, wherein Determining the execution credibility of the voice-controlled instruction of the elderly in the current nursing action scenario through the identification effective value and the action logic association specifically includes: Determining the executable strength of the voice control of the elderly in the current nursing action scenario through the action logic association; The execution credibility of the voice-controlled instructions of the elderly in the current nursing action scenario is determined based on the executable strength and the recognition validity value.
8. The method according to claim 1, wherein Generating a turning action control instruction for the turning assisting device according to the instruction execution credibility specifically includes: Preset trusted execution threshold; When the instruction execution credibility is greater than or equal to the trusted execution threshold, the nursing voice instruction of the elderly in the current nursing action scenario is determined to be executable, and then a turning action control instruction of the auxiliary turning device is generated; When the instruction execution credibility is less than the trusted execution threshold, the nursing voice instruction of the elderly in the current nursing action scenario is marked as unexecutable, and the elderly is prompted to reconfirm through voice feedback.
9. The method according to claim 1, wherein The sudden noise changes in the nursing environment refer to background noise sources that suddenly appear during the nursing process and have unstable spectral characteristics.
10. An auxiliary turning device for elderly care, the device includes a voice control unit, characterized in that: The voice control unit includes: An acquisition module is used to obtain a unique acoustic feature vector containing the breathy voice characteristics, weak voice characteristics and dialect characteristics of the elderly; A processing module, configured to determine the robustness weights of speech signals in different nursing action scenarios based on the exclusive acoustic feature vectors combined with sudden noise changes of the elderly in different nursing environments; The processing module is further configured to determine the semantic state features of the elderly in the current nursing action scenario based on the nursing voice instructions of the elderly in the current nursing action scenario in combination with the exclusive acoustic feature vector, and determine the recognition effective value of the nursing voice instructions of the elderly in the current nursing action scenario based on the semantic state features and the robustness weights of the voice signals in different nursing action scenarios; The processing module is further configured to determine the action category of the nursing voice instruction of the elderly person in the current nursing action scenario, obtain the turning action category corresponding to the elderly person in the current nursing action scenario, and then determine the action logical association between the turning action category and the current turning action state of the auxiliary turning device; An execution module is used to determine the execution credibility of the voice-controlled instructions of the elderly in the current nursing action scenario through the identification valid value and the action logic association, and then generate a turning action control instruction for the auxiliary turning device based on the instruction execution credibility.
Citation Information
Patent Citations
Guide robot with voice and image recognizing function
CN110070865A
Voice recognition method and device
CN110875039A
Instruction receiving method and system, electronic equipment, cloud server and storage medium
CN113889102A
Voice interaction method and device, equipment and storage medium
CN113963687A
Old people voice emergency help seeking method and system based on Internet of Things
CN119152844A