Intelligent health care robot voice recognition control method based on large model

By extracting user and background speech features to build a pointing text data set, and using a large model to convert speech text and keyword matching, the problem of insufficient speech recognition accuracy in intelligent health robots in noisy environments is solved, and more efficient speech control is achieved.

CN120544559AInactive Publication Date: 2025-08-26天津仁爱学院
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510669843.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-08-26
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In noisy environments or individual differences in accent and speech speed, the accuracy of speech recognition is difficult to maintain, and misunderstandings and error responses are prone to occur.

Method used

By obtaining user and background speech signals, extracting speech features and building a pointing text data set, using a large model to convert speech text and keyword matching, combining conflict detection and priority sorting, accurate control instructions are generated.

Benefits of technology

Improves the accuracy and response efficiency of speech recognition, reduces the impact of environmental noise and accent factors, and ensures that the robot performs the correct operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544559A_ABST
    Figure CN120544559A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voice recognition control, in particular to an intelligent health care robot voice recognition control method based on a large model, which comprises the following steps: acquiring a user voice signal and a background voice signal, and determining a user voice feature and a background voice feature; determining a pointing text data set corresponding to the user voice text, and converting the user voice signal into the user voice text; determining a plurality of important instruction keywords; determining a target instruction index set based on the instruction attributes of the instruction keywords, and determining a plurality of important instruction keyword combinations and probabilities corresponding to the important instruction keyword combinations; determining a plurality of candidate instruction keyword combinations, and generating corresponding candidate control instructions based on the candidate instruction keyword combinations; and performing conflict detection on each candidate control instruction, and determining a target control instruction based on a conflict detection result. The accuracy of voice recognition control can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech recognition control technology, and in particular to a speech recognition control method for an intelligent health care robot based on a large model. Background Art

[0002] The rapid development of speech recognition technology has greatly improved the user experience and is finding widespread application in mobile internet, wearable devices, healthcare, call centers, in-car navigation, smart home appliances, education, and other fields. Speech recognition is primarily used in small, intelligent embedded devices, both domestically and internationally, because these devices are more mobile and in greater demand than large internet devices. Intelligent robotics is an interdisciplinary field that combines electronics, mechanical engineering, computer science, and artificial intelligence to design, manufacture, and operate intelligent robots.

[0003] As the population ages, the healthcare needs of the elderly continue to grow, putting significant pressure on traditional healthcare models. Intelligent healthcare robots have become a hot research and development area. Intelligent healthcare robots focus on providing healthcare services for the elderly, specifically designed to address their physiological and psychological characteristics, with functions more tailored to their needs. As assistive devices, intelligent healthcare robots are gaining significant importance for the elderly, expanding their capabilities in areas such as companionship, care, rehabilitation training, and health monitoring. Intelligent healthcare robots, specifically designed for voice recognition and control, integrate voice recognition and control technologies, designed to operate and interact via voice commands. This aims to achieve a more natural and intuitive human-machine interaction, allowing users to control the robot through verbal commands, thereby improving operational convenience and efficiency. Traditional intelligent healthcare robots have limitations in voice recognition accuracy, semantic understanding, and interaction complexity. They struggle to maintain high recognition rates in noisy environments or when faced with individual differences in accent and speaking speed, leading to misunderstandings and incorrect responses.

[0004] Chinese patent application publication number CN117894311A discloses a bionic robot for speech recognition and control, including a speech recognition module, a user intent analysis module, a behavior pattern recognition module, a security enhancement module, a task priority adjustment module, an abnormal state recognition module, a decision support module, and a response execution module. It uses convolutional neural networks and long short-term memory networks to deeply analyze speech signals and improve recognition accuracy. The bidirectional encoder represents the output from the transformer model, deeply analyzes the text semantics, accurately grasps the user intent, and uses hidden Markov models and voiceprint recognition technology to optimize behavior pattern recognition and security. Genetic algorithms and isolation forest algorithms are used to achieve resource optimization and anomaly detection.

[0005] The existing technology has the following problems: it does not take into account individual differences in noisy environments or accents, speaking speed, etc., and speech recognition control is difficult to maintain high recognition accuracy, which may lead to misunderstandings and incorrect responses. Summary of the Invention

[0006] To this end, the present invention provides a large-model-based intelligent health care robot speech recognition control method to overcome the problem in the prior art that speech recognition control is difficult to maintain high accuracy in noisy environments or due to individual differences in accent, speaking speed, etc.

[0007] To achieve the above objectives, the present invention provides a large-scale model-based intelligent health care robot voice recognition control method, comprising:

[0008] Step S1, obtaining a user voice signal and a background voice signal, and determining the user voice features and the background voice features;

[0009] Step S2, determining a directional text data set corresponding to the user voice text based on the user voice feature and the background voice feature, and converting the user voice signal into user voice text, wherein each directional text data in the directional text data set has a corresponding directional text keyword and an instruction keyword;

[0010] Step S3, determining a number of voice text keywords based on the user voice text, and determining a number of important instruction keywords based on a comparison result of each voice text keyword with each pointing text keyword in the pointing text dataset;

[0011] Step S4, determining a target instruction index set based on the instruction attributes of each of the important instruction keywords, and inputting each of the important instruction keywords into a target recommendation model corresponding to the target instruction index set to determine a number of important instruction keyword combinations and corresponding probabilities of each important instruction keyword combination;

[0012] Step S5, determining a number of candidate instruction keyword combinations based on the corresponding probabilities of the important instruction keyword combinations, and generating corresponding candidate control instructions based on the candidate instruction keyword combinations;

[0013] Step S6: performing conflict detection on each of the candidate control instructions, and determining a target control instruction based on the conflict detection result.

[0014] Furthermore, in step S2, determining the pointed text data set corresponding to the user voice text includes:

[0015] Step S21, determining a user feature matching value based on a comparison result of the user voice feature and a standard user voice feature;

[0016] Step S22, determining a background feature matching value based on a comparison result between the background speech feature and a standard background speech feature;

[0017] Step S23: determining the pointed text data set corresponding to the user speech text based on the user feature matching value and the background feature matching value.

[0018] Furthermore, in step S2, converting the user voice signal into user voice text includes:

[0019] Step S24, determining a user speech recognition model based on the user feature matching value;

[0020] Step S25: input the user voice signal into the user voice recognition model to obtain the user voice text.

[0021] Furthermore, step S3 includes:

[0022] Step S31, performing text preprocessing on the user's voice text, and determining a number of voice text keywords based on a preset keyword extraction algorithm;

[0023] Step S32, determining a matching degree between each of the voice text keywords and each of the pointing text keywords in the pointing text dataset based on a comparison result of each of the voice text keywords and each of the pointing text keywords in the pointing text dataset;

[0024] Step S33 , determining a number of important instruction keywords according to the matching degree between each of the voice text keywords and each of the pointing text keywords in the pointing text data set.

[0025] Furthermore, in step S4, determining the target instruction index set includes:

[0026] Step S41, determining a target instruction type based on instruction attributes of each of the important instruction keywords;

[0027] Step S42: determining a target instruction index set corresponding to each of the important instruction keywords based on the target instruction type.

[0028] Furthermore, in the step S5, it includes:

[0029] Determining a number of candidate instruction keyword combinations based on a comparison result of the probability corresponding to each of the important instruction keyword combinations and a preset probability;

[0030] Among these, the instruction keyword combination whose probability corresponding to each of the important instruction keyword combinations is greater than a preset probability is determined as a candidate instruction keyword combination.

[0031] Furthermore, in step S5, a candidate instruction sentence corresponding to each candidate instruction keyword combination is determined based on a preset instruction rule library, and a corresponding candidate control instruction is determined based on each candidate instruction sentence.

[0032] Furthermore, step S6 includes:

[0033] Step S61, performing conflict detection on each of the candidate control instructions based on a preset conflict detection rule to obtain a number of key control instructions, wherein if there are conflicting candidate control instructions, determining the candidate control instructions to be screened out based on the instruction priorities of the conflicting candidate control instructions;

[0034] Step S62 , sorting the key control instructions obtained after the conflict detection according to the instruction priority, and determining the target control instruction based on the sorting result.

[0035] Furthermore, the step S1 includes:

[0036] Step S11, obtaining a user voice signal and a background voice signal, and preprocessing the user voice signal and the background voice signal;

[0037] Step S12: extracting features from the pre-processed user voice signal and background voice signal to obtain user voice features and background voice features, wherein the voice features include spectrum features, time domain features, fundamental frequency features, and vocal tract features.

[0038] Furthermore, in step S62, if the number of key control instructions obtained after the conflict detection is unique, the corresponding key control instructions are determined as target control instructions.

[0039] Compared to the prior art, the present invention offers the following advantages: by simultaneously collecting both user and background voice signals, it provides a data foundation for subsequent accurate user voice recognition and background environment analysis. By determining a targeted text dataset based on user and background voice features, the targeted text dataset can be accurately determined based on the current voice environment and user voice characteristics. Converting the user voice signal to text in conjunction with the corresponding targeted text dataset can reduce recognition errors caused by factors such as ambient noise and accents, improve the accuracy of speech-to-text conversion, and thus enhance the accuracy of subsequent speech recognition control. By determining several speech-to-text keywords from the user's speech text, irrelevant information can be filtered out, improving the efficiency of understanding user intent. By comparing the targeted text keywords with the targeted text keywords in the targeted text dataset, command keywords related to user commands can be further precisely located, improving speech recognition accuracy. By determining a target command index set based on the command attributes of the command keywords, recognition accuracy can be improved. Each command keyword is input into a target recommendation model corresponding to the target command index set to determine several important command keyword combinations and the corresponding probabilities of each important command keyword combination. Based on the corresponding probabilities of each important command keyword combination, several candidate command keyword combinations are determined, improving the accuracy of command generation and further enhancing the accuracy of speech recognition control. By generating corresponding candidate control instructions based on the keyword combinations of each candidate instruction, and performing conflict detection and priority sorting on each candidate control instruction to determine the target control instruction, it is possible to promptly discover and eliminate contradictory or unreasonable instruction combinations, and avoid the robot from performing erroneous or invalid operations. Through priority sorting, the execution order of instructions is reasonably arranged according to the importance and urgency of the instructions, so that the robot can give priority to performing key tasks, improve work efficiency, and further improve the accuracy of voice recognition control.

[0040] Furthermore, the present invention determines the pointed text data set corresponding to the user voice text by determining the user feature matching value and the background feature matching value, which can comprehensively consider the individual differences of the user voice signal and the scene differences of the background voice signal, thereby improving the accuracy of determining the pointed text data set and improving the voice recognition control accuracy and response efficiency in complex scenarios.

[0041] Furthermore, the present invention determines the user speech recognition model based on the user feature matching value, accurately matches the user features, and adaptively selects the user speech recognition model, which can effectively reduce recognition errors caused by individual differences, improve the accuracy of speech recognition, and improve the accuracy of text output.

[0042] Furthermore, the present invention pre-processes the user's voice text and determines the voice text keywords using a preset keyword extraction algorithm, thereby removing irrelevant information and redundant vocabulary, improving the accuracy and efficiency of keyword extraction. By determining the degree of match between each voice text keyword and each pointing text keyword in the pointing text dataset, the present invention accurately measures the similarity between the voice text keywords and the pointing text keywords in the pointing text dataset, accurately screening out the command keywords that are highly relevant to the user's intent, and thus improving the accuracy of command recognition.

[0043] Furthermore, the present invention determines the target instruction type through the instruction attributes of the instruction keyword, accurately classifying the instruction keyword into a specific instruction type, ensuring the pertinence and accuracy of subsequent instruction processing. In this way, the target instruction index set corresponding to each instruction keyword is determined, which can improve the efficiency and accuracy of instruction retrieval. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 This is a flow chart of a large-model-based intelligent health care robot voice recognition control method according to an embodiment of the present invention;

[0045] Figure 2 A schematic diagram of a process for determining a text data set corresponding to a user's voice text according to an embodiment of the present invention;

[0046] Figure 3 Schematic diagram of the process of step S3 of the embodiment of the present invention;

[0047] Figure 4 A schematic diagram of a process for determining a target instruction index set according to an embodiment of the present invention. DETAILED DESCRIPTION

[0048] In order to make the objects and advantages of the present invention more clearly understood, the present invention is further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are merely used to explain the present invention and are not intended to limit the present invention.

[0049] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.

[0050] It should be noted that, in the description of the present invention, terms such as "up", "down", "left", "right", "inside", and "outside" indicating directions or positional relationships are based on the directions or positional relationships shown in the accompanying drawings. This is only for the convenience of description and does not indicate or imply that the device or element must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it cannot be understood as a limitation on the present invention.

[0051] Furthermore, it should be noted that, in the description of the present invention, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood in a broad sense. For example, they may refer to fixed connections, detachable connections, or integral connections; mechanical connections or electrical connections; direct connections or indirect connections through an intermediate medium; and internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on specific circumstances.

[0052] See also Figures 1-4 As shown, Figure 1 This is a flow chart of a large-model-based intelligent health care robot voice recognition control method according to an embodiment of the present invention; Figure 2 A schematic diagram of a process for determining a text data set corresponding to a user's voice text according to an embodiment of the present invention; Figure 3 Schematic diagram of the process of step S3 of the embodiment of the present invention; Figure 4 This is a flow chart of determining a target instruction index set according to an embodiment of the present invention. An embodiment of the present invention provides a large-model-based intelligent health care robot voice recognition control method, including:

[0053] Step S1, obtaining a user voice signal and a background voice signal, and determining the user voice features and the background voice features;

[0054] Specifically, step S1 includes:

[0055] Step S11, obtaining a user voice signal and a background voice signal, and preprocessing the user voice signal and the background voice signal;

[0056] In practice, the specific structure for acquiring user voice signals and background voice signals is not limited. User voice signals are voice signals that the user sends to convey instructions or information to the robot, such as "Please bring me something I left on the table in my room" or "It's a bit cold today, please get me a coat." While the user sends the command, the robot also collects voice signals from the environment, such as the conversations of other elderly people in the nursing home room, the footsteps of caregivers in the corridor, and traffic noise.

[0057] It is understandable that the pre-processing of the user voice signal and the background voice signal includes filtering and denoising, voice enhancement, etc. This is an existing technology and will not be described in detail here.

[0058] Step S12: extracting features from the pre-processed user voice signal and background voice signal to obtain user voice features and background voice features, wherein the voice features include spectrum features, time domain features, fundamental frequency features, and vocal tract features.

[0059] In implementation, the spectrum features reflect the energy distribution characteristics of the speech signal in the frequency domain, the time domain features reflect the amplitude change characteristics of the speech signal over time, the fundamental frequency features reflect the pitch characteristics of the speech signal, and the vocal tract features reflect the clarity of the speech signal.

[0060] Step S2, determining a directional text data set corresponding to the user voice text based on the user voice feature and the background voice feature, and converting the user voice signal into user voice text, wherein each directional text data in the directional text data set has a corresponding directional text keyword and an instruction keyword;

[0061] Specifically, in step S2, determining the pointed text data set corresponding to the user voice text includes:

[0062] Step S21, determining a user feature matching value based on a comparison result of the user voice feature and a standard user voice feature;

[0063] Step S22, determining a background feature matching value based on a comparison result between the background speech feature and a standard background speech feature;

[0064] Step S23: determining the pointed text data set corresponding to the user speech text based on the user feature matching value and the background feature matching value.

[0065] In implementation, according to the user voice features YT1, YT2, ..., YT j ,…,YT m Compared with standard user voice features BT1, BT2, ..., BT j ,…,BT m Determine the user feature matching value YP, where j = 1, 2, ..., m; YP = (∑ m j=1 YT j ×BT j ) / (sqrt(∑ m j=1 (YT j ) 2 )×sqrt(∑ m j=1 (BT j ) 2 )); According to the background speech features ET1, ET2, ..., ET j ,…,ET m Compared with the standard background speech features BE1, BE2, ..., BE j ,…,BE m Determine the background feature matching value EP, where EP = (∑ m j=1 ETj ×BE j ) / (sqrt(∑ m j=1 (ET j ) 2 )×sqrt(∑ m j=1 (BE j ) 2 )); m is the number of speech features, YT j is the jth user’s speech feature, BT j is the jth standard user speech feature, ET j is the jth background speech feature, BE j is the jth standard background speech feature, and sqrt() is a preset square root determination function.

[0066] During implementation, the comprehensive feature matching value is determined based on the average of the user feature matching value and the background feature matching value, and the corresponding pointed text data set is determined based on the comprehensive feature matching value comparison table, where each pointed text data set has a corresponding comprehensive feature matching value range. The actual implementer can set it based on the actual situation. For example, the comprehensive feature matching values ​​corresponding to the user voice signals and background voice signals of different users in different scenarios are comprehensively calculated, and corresponding pointed text data sets are constructed for the voice text data of different users in different scenarios.

[0067] The present invention determines the pointed text data set corresponding to the user's voice text by determining the user feature matching value and the background feature matching value. It can comprehensively consider the individual differences of the user's voice signal and the scene differences of the background voice signal, improve the accuracy of determining the pointed text data set, and improve the voice recognition control accuracy and response efficiency in complex scenarios.

[0068] Specifically, in step S2, converting the user voice signal into user voice text includes:

[0069] Step S24, determining a user speech recognition model based on the user feature matching value;

[0070] Step S25: input the user voice signal into the user voice recognition model to obtain the user voice text.

[0071] During implementation, the intelligent health care robot has built-in multiple user voice recognition models, covering different accents, speaking speeds and common vocabulary combinations, and sets a corresponding user feature matching value range for each user voice recognition model to build a user feature matching value comparison table, thereby determining the corresponding user voice recognition model based on the user feature matching value.

[0072] It can be understood that a dataset can be constructed based on the speech data and corresponding speech texts of users with different characteristics such as different accents, speech rates, ages, genders, etc. in different environments (such as a quiet indoor environment, a noisy street, a room with background music, etc.), and a large model architecture can be constructed based on recurrent neural networks (such as long short-term memory networks and gated recurrent units) and convolutional neural networks, etc., and the large model can be trained based on the dataset to obtain a user speech recognition model.

[0073] In this invention, by determining the user speech recognition model based on the user feature matching value, accurately matching the user features, and adaptively selecting the user speech recognition model, it can effectively reduce the recognition errors caused by individual differences, improve the accuracy of speech recognition, and enhance the accuracy of text output.

[0074] Step S3: Determine a number of speech text keywords based on the user speech text, and determine a number of important instruction keywords based on the comparison results between each speech text keyword and each pointing text keyword in the pointing text dataset;

[0075] Specifically, the step S3 includes:

[0076] Step S31: Perform text preprocessing on the user speech text, and determine a number of speech text keywords based on a preset keyword extraction algorithm;

[0077] Step S32: Determine the matching degree between each speech text keyword and each pointing text keyword in the pointing text dataset based on the comparison results between each speech text keyword and each pointing text keyword in the pointing text dataset;

[0078] Step S33: Determine a number of important instruction keywords according to the matching degree between each speech text keyword and each pointing text keyword in the pointing text dataset.

[0079] In implementation, the text preprocessing includes removing punctuation marks, stop words (such as "de", "le", "shi", etc.), numbers, and special characters, etc. in the text, and performing word form reduction (such as restoring the past tense of a verb to its original form).

[0080] It can be understood that the preset keyword extraction algorithm can be the TF-IDF algorithm, the TextRank algorithm, the LDA topic model, etc.

[0081] It can be understood that each voice text keyword and each pointing text keyword are represented as a vector respectively, and the matching degree of each voice text keyword and the pointing text keyword is determined based on the cosine value of the angle between the two vectors. The closer the cosine value is to 1, the greater the matching degree between the voice text keyword and the pointing text keyword. If the cosine value of any voice text keyword and the pointing text keyword is greater than the preset threshold, the instruction keyword corresponding to the pointing text keyword is determined as the important instruction keyword corresponding to the voice text keyword.

[0082] It is understandable that the number of important instruction keywords corresponding to any voice text keyword may be zero or may not be unique. The actual implementation personnel can set the preset threshold based on the actual situation. Preferably, the preset threshold is set in the range of 0.7 to 0.8.

[0083] The present invention preprocesses user speech text and determines speech text keywords using a preset keyword extraction algorithm, thereby removing irrelevant information and redundant vocabulary, improving the accuracy and efficiency of keyword extraction. The invention also determines instruction keywords by determining the degree of match between each speech text keyword and each pointing text keyword in the pointing text dataset. This accurately measures the similarity between speech text keywords and pointing text keywords in the pointing text dataset, precisely screening out instruction keywords that are highly relevant to user intent, and thus improving instruction recognition accuracy.

[0084] Step S4, determining a target instruction index set based on the instruction attributes of each of the important instruction keywords, and inputting each of the important instruction keywords into a target recommendation model corresponding to the target instruction index set to determine a number of important instruction keyword combinations and corresponding probabilities of each important instruction keyword combination;

[0085] Specifically, in step S4, determining the target instruction index set includes:

[0086] Step S41, determining a target instruction type based on instruction attributes of each of the important instruction keywords;

[0087] Step S42: determining a target instruction index set corresponding to each of the important instruction keywords based on the target instruction type.

[0088] In implementation, each important instruction keyword has corresponding instruction attributes, including function type (such as motion control, equipment operation, information query, interactive communication, etc.), operation object (such as robot's own parts, external environment equipment, virtual information, etc.) and priority (such as emergency instructions, ordinary instructions, low priority instructions). The instruction attributes of each important instruction keyword are analyzed, and the instruction type is determined according to the function type in the instruction attributes of each important instruction keyword. The instruction type includes motion control instruction type, equipment operation instruction type, etc., and each instruction type has a corresponding instruction index set.

[0089] In implementation, an instruction keyword training sample set can be constructed based on instruction keywords that have passed the qualification test in historical data, and each instruction keyword combination can be used as a sample label to train the initial recommendation model to obtain a target recommendation model.

[0090] It is understandable that those skilled in the art know that any recommendation model in the prior art that can output the probability of an instruction keyword combination falls within the scope of protection of the present invention, for example, the deepfm model, the multi-armed bandit model, etc., which will not be elaborated here.

[0091] The present invention determines the target instruction type by the instruction attribute of the instruction keyword, accurately classifies the instruction keyword into a specific instruction type, and ensures the pertinence and accuracy of subsequent instruction processing. In this way, the target instruction index set corresponding to each instruction keyword is determined, which can improve the efficiency and accuracy of instruction retrieval.

[0092] Step S5, determining a number of candidate instruction keyword combinations based on the corresponding probabilities of the important instruction keyword combinations, and generating corresponding candidate control instructions based on the candidate instruction keyword combinations;

[0093] Specifically, step S5 includes:

[0094] Determining a number of candidate instruction keyword combinations based on a comparison result of the probability corresponding to each of the important instruction keyword combinations and a preset probability;

[0095] Among these, the instruction keyword combination whose probability corresponding to each of the important instruction keyword combinations is greater than a preset probability is determined as a candidate instruction keyword combination.

[0096] During implementation, actual implementers can set the preset probability based on actual conditions. Preferably, the preset probability value range is set to 0.5 to 0.6.

[0097] Specifically, in step S5, the candidate instruction sentences corresponding to each candidate instruction keyword combination are determined based on a preset instruction rule library, and the corresponding candidate control instructions are determined based on each candidate instruction sentence.

[0098] During implementation, the actual implementers can set up a preset instruction rule library based on actual conditions. The preset instruction rule library includes several instruction statement templates, covering different instruction statement structures, for example, "Please [operate] [location] [equipment]", and fill in the candidate instruction keywords in each candidate instruction keyword combination into the corresponding positions in the template to obtain the corresponding candidate instruction statements.

[0099] It is understandable that each candidate instruction statement in the preset instruction rule library has a corresponding candidate control instruction.

[0100] Step S6: performing conflict detection on each of the candidate control instructions, and determining a target control instruction based on the conflict detection result.

[0101] Specifically, step S6 includes:

[0102] Step S61, performing conflict detection on each of the candidate control instructions based on a preset conflict detection rule to obtain a number of key control instructions, wherein if there are conflicting candidate control instructions, determining the candidate control instructions to be screened out based on the instruction priorities of the conflicting candidate control instructions;

[0103] During implementation, the actual implementer can set preset conflict detection rules based on actual conditions. For example, "forward" and "backward" are conflicting instructions with opposite movement directions, and "turn on the air conditioner" and "turn off the air conditioner" are conflicting instructions with opposite device operations. Candidate control instructions are compared pairwise and the preset conflict detection rules are applied to determine whether there is a conflict. If a conflict exists, the candidate control instructions with lower priority in the conflict are screened out. For example, if "emergency brake" (high priority) conflicts with "continue forward" (low priority), "continue forward" is screened out.

[0104] Step S62 , sorting the key control instructions obtained after the conflict detection according to the instruction priority, and determining the target control instruction based on the sorting result.

[0105] During implementation, the actual implementers can set preset priority rules based on actual conditions, and set them according to factors such as the urgency of the instructions and the function type. For example, emergency instructions (such as "emergency braking" and "call emergency") have the highest priority; core health care function instructions (such as "health monitoring start") are second; general function instructions (such as "check the weather" and "play music") have a lower priority. Sort the key control instructions by priority, with the highest priority key control instructions at the front, and execute the instructions at the top in turn; if not all can be executed, the instruction at the front will be executed first.

[0106] Specifically, in step S62, if the number of key control instructions obtained after the conflict detection is unique, the corresponding key control instructions are determined as target control instructions.

[0107] The present invention simultaneously collects user voice signals and background voice signals, providing a data foundation for subsequent accurate recognition of user voice and analysis of the background environment. By determining a targeted text dataset based on user voice and background voice features, the targeted text dataset can be accurately determined based on the current voice environment and user voice characteristics. Converting the user voice signal to text in conjunction with the corresponding targeted text dataset can reduce recognition errors caused by factors such as ambient noise and accents, improve the accuracy of voice-to-text conversion, and thus enhance the accuracy of subsequent voice recognition control. By determining a number of voice-to-text keywords from the user voice text, irrelevant information can be filtered out, improving the efficiency of understanding user intent. By comparing the targeted text keywords with the targeted text keywords in the targeted text dataset, command keywords related to user commands can be further precisely located, improving voice recognition accuracy. By determining a target command index set based on the command attributes of the command keywords, recognition accuracy can be improved. Each command keyword is input into a target recommendation model corresponding to the target command index set to determine several important command keyword combinations and the corresponding probabilities of each important command keyword combination. Based on the corresponding probabilities of each important command keyword combination, several candidate command keyword combinations are determined, improving the accuracy of command generation and further enhancing the accuracy of voice recognition control. By generating corresponding candidate control instructions based on the keyword combinations of each candidate instruction, and performing conflict detection and priority sorting on each candidate control instruction to determine the target control instruction, it is possible to promptly discover and eliminate contradictory or unreasonable instruction combinations, and avoid the robot from performing erroneous or invalid operations. Through priority sorting, the execution order of instructions is reasonably arranged according to the importance and urgency of the instructions, so that the robot can give priority to performing key tasks, improve work efficiency, and further improve the accuracy of voice recognition control.

[0108] Thus far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art may make equivalent changes or substitutions to the relevant technical features, and the technical solutions after such changes or substitutions will fall within the scope of protection of the present invention.

Claims

1. A large-scale model-based intelligent health care robot voice recognition control method, characterized in that: include: Step S1, obtaining a user voice signal and a background voice signal, and determining the user voice features and the background voice features; Step S2, determining a directional text data set corresponding to the user voice text based on the user voice feature and the background voice feature, and converting the user voice signal into user voice text, wherein each directional text data in the directional text data set has a corresponding directional text keyword and an instruction keyword; Step S3, determining a number of voice text keywords based on the user voice text, and determining a number of important instruction keywords based on a comparison result of each voice text keyword with each pointing text keyword in the pointing text dataset; Step S4, determining a target instruction index set based on the instruction attributes of each of the important instruction keywords, and inputting each of the important instruction keywords into a target recommendation model corresponding to the target instruction index set to determine a number of important instruction keyword combinations and corresponding probabilities of each important instruction keyword combination; Step S5, determining a number of candidate instruction keyword combinations based on the corresponding probabilities of the important instruction keyword combinations, and generating corresponding candidate control instructions based on the candidate instruction keyword combinations; Step S6: performing conflict detection on each of the candidate control instructions, and determining a target control instruction based on the conflict detection result.

2. The large model-based intelligent health care robot voice recognition control method according to claim 1 is characterized in that: In step S2, determining the pointed text data set corresponding to the user voice text includes: Step S21, determining a user feature matching value based on a comparison result of the user voice feature and a standard user voice feature; Step S22, determining a background feature matching value based on a comparison result between the background speech feature and a standard background speech feature; Step S23: determining the pointed text data set corresponding to the user speech text based on the user feature matching value and the background feature matching value.

3. The large-model-based intelligent health care robot voice recognition control method according to claim 2 is characterized in that: In step S2, converting the user voice signal into user voice text includes: Step S24, determining a user speech recognition model based on the user feature matching value; Step S25: input the user voice signal into the user voice recognition model to obtain the user voice text.

4. The large-model-based intelligent health care robot voice recognition control method according to claim 3 is characterized in that: The step S3 comprises: Step S31, performing text preprocessing on the user's voice text, and determining a number of voice text keywords based on a preset keyword extraction algorithm; Step S32, determining a matching degree between each of the voice text keywords and each of the pointing text keywords in the pointing text dataset based on a comparison result of each of the voice text keywords and each of the pointing text keywords in the pointing text dataset; Step S33 , determining a number of important instruction keywords according to the matching degree between each of the voice text keywords and each of the pointing text keywords in the pointing text data set.

5. The large-model-based intelligent health care robot voice recognition control method according to claim 4 is characterized in that: In step S4, determining the target instruction index set includes: Step S41, determining a target instruction type based on instruction attributes of each of the important instruction keywords; Step S42: determining a target instruction index set corresponding to each of the important instruction keywords based on the target instruction type.

6. The large-model-based intelligent health care robot voice recognition control method according to claim 5 is characterized in that: In the step S5, it includes: Determining a number of candidate instruction keyword combinations based on a comparison result of the probability corresponding to each of the important instruction keyword combinations and a preset probability; Among these, the instruction keyword combination whose probability corresponding to each of the important instruction keyword combinations is greater than a preset probability is determined as a candidate instruction keyword combination.

7. The large-model-based intelligent health care robot voice recognition control method according to claim 6 is characterized in that: In step S5, candidate instruction sentences corresponding to each candidate instruction keyword combination are determined based on a preset instruction rule library, and corresponding candidate control instructions are determined based on each candidate instruction sentence.

8. The large model-based intelligent health care robot voice recognition control method according to claim 7 is characterized in that: The step S6 comprises: Step S61, performing conflict detection on each of the candidate control instructions based on a preset conflict detection rule to obtain a number of key control instructions, wherein if there are conflicting candidate control instructions, determining the candidate control instructions to be screened out based on the instruction priorities of the conflicting candidate control instructions; Step S62 , sorting the key control instructions obtained after the conflict detection according to the instruction priority, and determining the target control instruction based on the sorting result.

9. The large model-based intelligent health care robot voice recognition control method according to claim 8 is characterized in that: The step S1 comprises: Step S11, obtaining a user voice signal and a background voice signal, and preprocessing the user voice signal and the background voice signal; Step S12: extracting features from the pre-processed user voice signal and background voice signal to obtain user voice features and background voice features, wherein the voice features include spectrum features, time domain features, fundamental frequency features, and vocal tract features.

10. The large model-based intelligent health care robot voice recognition control method according to claim 9 is characterized in that: In step S62 , if the number of key control instructions obtained after conflict detection is unique, the corresponding key control instructions are determined as target control instructions.

Citation Information

Patent Citations

  • Bionic robot in aspect of voice recognition control

    CN117894311A