Voice instruction recognition method and system for composite robot
By fusing multi-source information and dynamically evaluating the confidence level of composite robot voice commands, the accuracy and security issues of traditional speech recognition in complex environments are solved, enabling efficient and safe decision-making under noisy and multi-user operation conditions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- WUXI INSTITUTE OF TECHNOLOGY
- Filing Date
- 2026-01-28
- Publication Date
- 2026-05-12
AI Technical Summary
In complex industrial production environments, traditional speech recognition methods struggle to accurately recognize the voice commands of composite robots, especially in the presence of noise, multiple operators, and ambiguous command semantics. This can lead to misjudgment or execution delays, impacting operational efficiency and safety.
By acquiring and preprocessing multiple speech signals, extracting speech features from independent speech command signals and calibrating them in conjunction with environmental acoustic features, generating command text and parsing its semantics, assessing confidence, and employing a dynamic confidence-weighted fusion method to process multi-source priority information, the accuracy and security of decision-making are ensured.
It improves the accuracy and robustness of voice emotion recognition, enabling it to make quick and accurate decisions in complex environments that conform to the actual situation, ensuring operational safety and task continuity, and enhancing the robot's decision-making accuracy and safety under multi-source conflicting commands.
Smart Images

Figure CN122024718A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of industrial control technology, and more specifically, to a method and system for recognizing voice commands of a composite robot. Background Technology
[0002] In modern intelligent manufacturing environments, composite robots are key equipment for performing various tasks such as material handling, component assembly, and equipment inspection. To achieve efficient and flexible operation, these robots are typically equipped with voice command recognition systems, allowing on-site operators to schedule tasks via verbal commands. However, the actual production workshop environment is complex and variable, often accompanied by noise from equipment operation, differences in accents among operators, and potential semantic ambiguity in the commands themselves. These factors all pose significant challenges to the accurate recognition of voice commands. In particular, the system often encounters difficulties in recognizing key action parameters contained in complex commands, such as target location, operational force, or the urgency of execution. Traditional voice recognition methods have insufficient parsing capabilities when processing unstructured spoken expressions, easily leading to misinterpretation of commands or execution delays, thereby affecting the robot's operational efficiency and production safety. Summary of the Invention
[0003] To address the aforementioned deficiencies, the present invention provides a composite robot voice command recognition method and system, aiming to not only understand the literal meaning of the command but also perceive the real urgency behind it, thereby enabling quick and accurate decision-making in emergency situations that aligns with the actual context, ensuring operational safety and task continuity.
[0004] The first aspect of this invention provides a method for recognizing voice commands in a composite robot, wherein the composite robot is used to receive multiple voice signals and perform action tasks corresponding to the voice signals, including: Step S1: Acquire the multiple voice signals and preprocess the voice signals to obtain independent voice command signals corresponding to different voice sources.
[0005] Step S2: Extract the speech features of the independent speech command signal, and calibrate the speech features in combination with environmental acoustic features to obtain calibrated speech emotion feature information.
[0006] Step S3: Generate instruction text by performing speech recognition based on independent speech command signals, parse the instruction semantics of the instruction text, match the instruction semantics with pre-stored action task priority information, and obtain the priority information of the action task corresponding to the instruction semantics.
[0007] Step S4: Based on the calibrated speech emotion feature information, instruction semantic information, action task priority information, and environmental safety perception information, evaluate the corresponding confidence levels respectively, and calculate the multi-source priority information fusion score through dynamic confidence weighted fusion.
[0008] Step S5: Save the current task state and the physical state of the robot, select and execute the task based on the multi-source priority information fusion score.
[0009] According to one embodiment of the present invention, the environmental acoustic features include interfering acoustic features such as high-frequency noise or mechanical noise that affect the accuracy of speech emotion recognition; and the dynamic calibration includes adjusting the judgment threshold or feature weight of the emotion features according to the intensity of the interfering acoustic features.
[0010] According to an embodiment of the present invention, the confidence level in step S4 is evaluated respectively, including: based on the emotional priority information corresponding to the voice emotion features, the semantic priority information corresponding to the instruction semantics, the action task priority information, and the safety warning priority information triggered by the environmental perception information, the confidence level of the voice emotion features, the confidence level of the instruction semantics, the confidence level of the action task priority, and the confidence level of the environmental perception results are evaluated respectively.
[0011] According to one embodiment of the present invention, the confidence level is dynamically adjusted according to at least one of the following: the confidence level of the voice emotion feature is adjusted according to the intensity of environmental interference.
[0012] The confidence level of the instruction semantics is adjusted based on the accuracy of instruction semantic recognition and the explicitness of keyword matching; The confidence level of the action task priority is adjusted based on the stability of the source of the action task priority.
[0013] The confidence level of the environmental perception results is adjusted based on whether the accuracy of physical security risk identification exceeds a set threshold.
[0014] According to one embodiment of the present invention, the environmental perception information includes: when a potential physical security risk is predicted, generating a self-triggered security warning without receiving a voice command, and using the self-triggered security warning as the highest priority information in the multi-source priority information fusion.
[0015] According to one embodiment of the present invention, the task state storage includes current task execution progress information, the physical state storage includes robot pose, actuator state and carried material state, and the task state and physical state are marked as recoverable breakpoint states.
[0016] According to one embodiment of the present invention, step S5 further includes: when the instruction information corresponding to the multi-source priority information fusion score is incomplete, entering a clarification pending state, actively detecting potential risk areas, and simultaneously sending clarification inquiry information. Based on the received clarification instruction or autonomous perception result, the interrupted action task can be resumed from the recoverable power-off state.
[0017] A second aspect of the present invention provides a composite robot voice command recognition system, the system comprising: a voice acquisition and preprocessing module, used to acquire the plurality of voice signals and preprocess the voice signals to obtain independent voice command signals corresponding to different voice sources.
[0018] The emotion feature perception and calibration module is used to extract the speech features of the independent speech command signal, and calibrate the speech features in combination with environmental acoustic features to obtain calibrated speech emotion feature information.
[0019] The semantic parsing task mapping module generates instruction text by performing speech recognition based on independent speech command signals, parses the instruction semantics of the instruction text, matches the instruction semantics with pre-stored action task priority information, and obtains the priority information of the action task corresponding to the instruction semantics.
[0020] The dynamic confidence-weighted fusion module is used to evaluate the corresponding confidence levels based on the calibrated speech emotion feature information, instruction semantic information, action task priority information, and environmental safety perception information, and to calculate the multi-source priority information fusion score through dynamic confidence-weighted fusion.
[0021] The interaction module is used to save the task state of the current action task and the physical state of the robot, and select and execute the action task based on the multi-source priority information fusion score.
[0022] According to an embodiment of the present invention, the execution interaction module is further configured to enter a state pending clarification when the instruction information corresponding to the multi-source priority information fusion score is incomplete, actively detect potential risk areas, send clarification inquiry information, and continue to execute the interrupted action task from the recoverable breakpoint state according to the clarification instruction or autonomous perception result.
[0023] A third aspect of the present invention provides a composite robot, the composite robot including the composite robot voice command recognition system described above.
[0024] The beneficial effects provided by this invention are as follows: First, the method acquires multiple voice signals and performs preprocessing to obtain independent voice command signals corresponding to different voice sources, which effectively solves the problem that it is difficult to separate mixed audio streams from different operators when the robot receives them simultaneously or within a short period of time in a multi-person collaborative scenario.
[0025] Secondly, the speech features of independent speech command signals are extracted and calibrated by combining them with environmental acoustic features to obtain calibrated speech emotion feature information. This enables the system to overcome noise interference in complex production workshop environments and improve the accuracy of speech emotion recognition.
[0026] Furthermore, the system generates instruction text through speech recognition and parses the semantics of the instructions. At the same time, it matches the text with pre-stored action task priority information to obtain the corresponding priority information. This enables the system to perform in-depth analysis of instruction semantics, distinguish the urgency and importance of different instructions, and thus flexibly determine the execution priority of the instructions.
[0027] Finally, based on the calibrated speech emotion features, instruction semantics, action task priority, and environmental safety perception, the corresponding confidence levels are evaluated. A multi-source priority information fusion score is then calculated using a dynamic confidence-weighted fusion method, and the action task is selected and executed accordingly. This multi-source information fusion mechanism, particularly the introduction of dynamic confidence weighting, effectively addresses instruction conflicts and incomplete information, ensuring the robot's decision-making accuracy and safety in complex and unexpected situations.
[0028] In summary, the method of this application overcomes the shortcomings of existing technologies in handling unstructured spoken expressions, conflicting multi-source instructions, and incomplete information by independently processing multi-source speech commands, calibrating emotion features adapted to the environment, performing deep semantic parsing and priority matching, and fusing multi-source information with dynamic confidence weighting. This significantly improves the accuracy, robustness, and intelligent decision-making capabilities of composite robot speech command recognition, ensuring the robot's operational efficiency and production safety in complex industrial environments. Attached Figure Description
[0029] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0030] Figure 1 This is a flowchart of the composite robot voice command recognition method disclosed in an embodiment of the present invention; Figure 2 This is a block diagram of a composite robot voice command recognition system disclosed in an embodiment of the present invention.
[0031] The accompanying drawings have illustrated specific embodiments of this disclosure, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concepts of this disclosure to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0032] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0033] Obviously, the above specific implementation examples are merely illustrative of the application of this method and not intended to limit the implementation. Those skilled in the art can make other variations and modifications based on the above description to study other related issues. Therefore, the scope of protection of this invention should be limited to the scope of the claims.
[0034] This invention deeply integrates and analyzes the emotional information contained in human voice commands with the literal content of the commands to achieve intelligent judgment of the command's intent and urgency. Compared to the limitations of traditional speech recognition that only focuses on semantic content, this method captures non-verbal features such as the operator's speech rate, tone, and volume in real time and incorporates them as key weighting factors into the command priority evaluation system. This approach enables robots, when faced with multi-source, conflicting, and potentially emotionally charged commands in intelligent manufacturing workshops, to not only understand the literal meaning of the commands but also perceive the underlying urgency, just like experienced humans. This allows them to make quick and accurate decisions in emergency situations, ensuring operational safety and task continuity.
[0035] This invention aims to address the challenge of accurately identifying, flexibly prioritizing, and proactively handling subsequent uncertain tasks when faced with conflicting commands from multiple operators in the complex environment of a smart manufacturing workshop. It integrates and analyzes the emotional information contained in the operator's voice with the literal content of the commands to more intelligently assess the true intent and urgency of the instructions.
[0036] like Figure 1 As shown, this application proposes a method for recognizing voice commands in a composite robot. The composite robot is used to receive multiple voice signals and execute action tasks corresponding to the voice signals, including: Step S1: Acquire the multiple voice signals and preprocess the voice signals to obtain independent voice command signals corresponding to different voice sources.
[0037] Step S2: Extract the speech features of the independent speech command signal, and calibrate the speech features in combination with environmental acoustic features to obtain calibrated speech emotion feature information.
[0038] Step S3: Generate instruction text by performing speech recognition based on independent speech command signals, parse the instruction semantics of the instruction text, match the instruction semantics with pre-stored action task priority information, and obtain the priority information of the action task corresponding to the instruction semantics.
[0039] Step S4: Based on the calibrated speech emotion feature information, instruction semantic information, action task priority information, and environmental safety perception information, evaluate the corresponding confidence levels respectively, and calculate the multi-source priority information fusion score through dynamic confidence weighted fusion.
[0040] Step S5: Save the current task state and the physical state of the robot, select and execute the task based on the multi-source priority information fusion score.
[0041] Specifically, speech signal acquisition can be achieved through a microphone array integrated on the composite robot. These microphones can be distributed at different locations on the robot to capture speech from different directions. After acquiring the raw speech signal, preprocessing is required. Preprocessing can include techniques such as noise suppression, echo cancellation, and voice activity detection (VAD) to improve the quality of the speech signal. For example, spectral subtraction or deep learning noise reduction algorithms can be used to remove environmental noise. Subsequently, to separate independent speech command signals from different speech sources from the mixed speech stream, source separation techniques can be used, such as methods based on independent component analysis (ICA) or deep clustering. These techniques can decompose the mixed speech signal into multiple independent speech streams, each corresponding to a speaker.
[0042] In step S2, it is necessary to extract the speech features of independent speech command signals and calibrate the speech features in combination with environmental acoustic features to obtain calibrated speech emotion feature information.
[0043] Speech feature extraction can employ various mature speech processing techniques, such as Mel-frequency cepstral coefficients (MFCC), linear predictive cepstral coefficients (LPCC), or parameters like fundamental frequency and formants. These features effectively characterize speech timbre, pitch, loudness, and other information. To obtain speech emotion features, pre-trained emotion recognition models can be used. These models classify or regress the speaker's emotion based on the extracted speech features, for example, determining whether the speaker is in a state of emergency, calmness, anger, or anxiety.
[0044] Environmental acoustic features can be obtained through additional environmental microphones or by performing non-speech component analysis on the main microphone signal. For example, the intensity and frequency distribution of interfering acoustic features such as high-frequency noise and mechanical noise in the environment can be analyzed.
[0045] Calibrating speech emotion features is crucial in this step. For example, the presence of high-intensity mechanical noise in the environment may cause the emotion recognition model to misinterpret calm speech as anxiety or tension. To address this issue, the judgment threshold or feature weights of emotion features can be dynamically adjusted based on the intensity of the environmental acoustic features. For instance, when high-intensity noise is detected, the threshold for judging urgency can be appropriately increased, or the weight of noise-sensitive emotion features can be decreased, thereby reducing misjudgments.
[0046] In step S3, speech recognition is performed based on independent speech command signals to generate command text, the command semantics of the command text are parsed, and the command semantics are matched with pre-stored action task priority information to obtain the priority information of the action task corresponding to the command semantics.
[0047] Speech recognition can convert individual speech commands into a processable text form. This can be achieved using existing automatic speech recognition (ASR) engines, such as models based on deep neural networks (DNNs) or recurrent neural networks (RNNs).
[0048] After the instruction text is generated, semantic parsing is required. Semantic parsing aims to understand the deeper meaning and operational intent of the instruction. For example, for the instruction "move that box to area A," semantic parsing needs to identify "moving" as the action, "box" as the target, and "area A" as the destination. This can be achieved through natural language processing (NLP) techniques, such as rule-based parsers, statistical models, or deep learning models (such as BERT, GPT, etc.).
[0049] After parsing the semantics of the instruction, it needs to be matched against pre-stored action task priority information. This priority information can be a predefined database or knowledge graph containing various action tasks and their corresponding priorities. For example, "emergency stop" might be assigned the highest priority, "material handling" might have a medium priority, and "equipment inspection" might have a low priority. Through matching, the priority information of the action task corresponding to the current instruction can be obtained.
[0050] In step S4, based on the calibrated speech emotion feature information, instruction semantic information, action task priority information, and environmental safety perception information, the corresponding confidence levels are evaluated respectively, and the multi-source priority information fusion score is calculated through dynamic confidence weighted fusion.
[0051] First, the confidence level of each piece of information needs to be evaluated separately. The confidence level of voice emotion features can be determined based on the output probability of the emotion recognition model or the degree of matching with the preset emotion pattern. For example, when the emotion recognition model has a high probability of recognizing the emotion of "emergency," its confidence level is also correspondingly high.
[0052] The semantic confidence of an instruction can be assessed based on the accuracy of speech recognition, the completeness of semantic parsing, and the explicitness of keyword matching. For example, if the accuracy of instruction text recognition is high and all key semantic elements are clearly identified, then the semantic confidence is high.
[0053] The confidence level of action task priorities can be determined based on the stability of the source of the priority information or a preset reliability level. For example, if the priority information comes from a verified system configuration, the confidence level is high.
[0054] The confidence level of environmental perception results can be assessed based on the accuracy and real-time nature of the environmental safety perception information, as well as whether it exceeds a set safety threshold. For example, when the visual sensor clearly identifies an obstacle, and the obstacle is located on the robot's path, the confidence level of the environmental perception result is high.
[0055] After assessing the confidence levels of each element, a dynamic confidence-weighted fusion method is used to calculate the multi-source priority information fusion score. This means that the weights of different information sources are not fixed but dynamically adjusted based on their confidence levels. For example, when the confidence level of environmental safety perception information is extremely high (such as when an urgent safety risk is detected), its weight will be significantly increased, potentially even overriding the influence of other information sources, thereby ensuring safety priority. The fusion algorithm can employ weighted averages, fuzzy logic, or machine learning models, among others.
[0056] In step S5, the task state of the current action task and the physical state of the composite robot are saved, and the action task is selected and executed based on the multi-source priority information fusion score.
[0057] Saving task status can include information such as the current task's execution progress, completed subtasks, and subtasks yet to be executed. For example, if the task is "moving boxes", the task status can be recorded as "boxes have been grabbed, heading to area A".
[0058] Saving physical state can include the robot's current pose (position and orientation), the state of each actuator (such as the robotic arm and the mobile chassis), and whether it is carrying materials. For example, it can record the joint angles of the robotic arm, the speed and direction of the mobile chassis, and whether it successfully grasped the box.
[0059] These state information are marked as recoverable breakpoint states, which means that after a task is interrupted, the robot can continue to perform the task from these saved state points without having to start from scratch.
[0060] Finally, based on the multi-source priority information fusion score, the robot selects and executes the action task. The task with the highest fusion score will be executed first. For example, if the fusion score shows that "emergency stop" has the highest priority, the robot will immediately stop all current actions. If the fusion score shows that "carrying boxes" has the highest priority, the robot will perform the carrying task according to the predetermined path.
[0061] The composite robot voice command recognition method proposed in this application significantly improves the accuracy, robustness, and safety of composite robots in processing voice commands in complex industrial environments by introducing multi-source information fusion and dynamic confidence evaluation mechanisms.
[0062] The environmental acoustic features include interfering acoustic features such as high-frequency noise or mechanical noise that affect the accuracy of speech emotion recognition; and the dynamic calibration includes adjusting the judgment threshold or feature weight of the emotion features according to the intensity of the interfering acoustic features.
[0063] Specifically, environmental acoustic features can be understood as non-speech sounds present in the working environment of a composite robot that may interfere with the acquisition and processing of speech command signals. These interfering acoustic features, such as high-frequency noise or mechanical noise, can significantly reduce the accuracy of speech emotion recognition. High-frequency noise may originate from equipment operation, ventilation systems, etc., while mechanical noise is commonly found in industrial production lines and the robot's own moving parts. These noises can mask the emotion-related features in the speech signal, leading to misjudgments by emotion recognition algorithms.
[0064] Dynamic calibration refers to an adaptive adjustment mechanism designed to counteract or mitigate the impact of interfering acoustic features on speech emotion recognition. Specifically, when the system detects specific types of interfering acoustic features, such as high-frequency noise or mechanical noise, it dynamically adjusts the judgment threshold or feature weights of emotion features based on the real-time intensity of these interfering acoustic features. For example, when the intensity of high-frequency noise is high, the judgment threshold for recognizing specific emotions (such as anger or anxiety) can be appropriately increased to avoid misinterpreting noise-induced speech fluctuations as emotional changes; alternatively, the weights of speech features susceptible to high-frequency noise can be reduced while the weights of less affected features can be increased, thus enabling more accurate extraction and recognition of speech emotions even in noisy environments.
[0065] By explicitly identifying and quantifying the interfering acoustic features that affect the accuracy of speech emotion recognition, such as high-frequency noise or mechanical noise, the system can be calibrated specifically. It is precisely because these specific interference sources are identified that subsequent dynamic calibration becomes possible. By dynamically adjusting the judgment threshold or feature weight of emotion features based on the intensity of the interfering acoustic features, the system can adaptively adapt to different noise environments. When the noise intensity increases, the calibration mechanism adjusts the parameters accordingly to filter out the interference of noise on emotion feature extraction, thereby ensuring that speech emotion feature information can still be accurately acquired and utilized in complex and changing environments. This dynamic adjustment mechanism avoids the problem of fixed calibration parameters performing poorly under different noise levels, significantly improving the robustness of emotion recognition.
[0066] Through the aforementioned technical solutions, the composite robot can more effectively cope with the interference of complex and ever-changing environmental noise on speech emotion recognition. By identifying and specifically processing interfering acoustic features such as high-frequency noise or mechanical noise, and dynamically adjusting the judgment threshold or feature weight of emotion features according to their intensity, the accuracy and reliability of speech emotion feature information are significantly improved. As a result, the composite robot can more accurately understand the emotional information contained in the user's voice commands, and thus make more appropriate and human-like responses, improving the naturalness and efficiency of human-computer interaction.
[0067] Imagine a composite robot performing material handling tasks in a factory workshop. The workshop is filled with continuous mechanical noise from operating machinery and high-frequency noise from motors. When the operator issues a "Stop!" command to the robot, if the operator's voice contains obvious anxiety or urgency, but the ambient noise is high, traditional voice emotion recognition methods may fail to accurately identify this emotion due to noise interference.
[0068] In this application, the composite robot first detects and identifies high-intensity mechanical noise and high-frequency noise within the workshop using its environmental acoustic sensors. The emotion feature perception and calibration module dynamically adjusts the judgment threshold for voice emotion features based on the real-time intensity of these interfering acoustic features. For example, the recognition threshold for voice features expressing "urgency" may be appropriately increased. Simultaneously, the weight of voice features susceptible to mechanical noise is reduced, while the weight of voice features less affected by noise is increased. Through this dynamic calibration, even in noisy mechanical and high-frequency noise environments, the system can more accurately extract the urgency emotion from the operator's "Stop!" command, enabling the composite robot to respond faster and more decisively, such as immediately stopping its current action and entering a safe standby state to avoid potential dangers.
[0069] By refining the confidence assessment into confidence levels for voice emotion features, instruction semantics, action task priority, and environmental perception results, the composite robot can more comprehensively and meticulously consider the reliability of different information sources. Specifically, by evaluating these confidence levels separately, the system can quantify the impact of each information dimension on the final decision. For example, when the voice emotion recognition result is uncertain, the confidence level for voice emotion features will be low, thus reducing its weight in subsequent fusion scoring; conversely, when the environmental perception system issues a high-confidence safety warning, the confidence level for the environmental perception result will be high, ensuring that safety factors dominate the decision-making process. Therefore, this detailed confidence assessment mechanism provides more accurate input for subsequent dynamic confidence-weighted fusion, ensuring the accuracy and robustness of multi-source priority information fusion scoring.
[0070] Through the aforementioned technical solutions, the composite robot can perform more refined reliability assessments of priority information from different sources, avoiding decision-making biases that may result from single or vague confidence assessments. This meticulous confidence assessment mechanism enables the system to more accurately identify and respond to critical information when facing complex and ever-changing working environments and user instructions. Especially in terms of safety warnings, it ensures that high-confidence safety information is prioritized, thereby significantly improving the accuracy, safety, and intelligence level of the composite robot in performing its tasks.
[0071] A second aspect of the present invention provides a composite robot voice command recognition system, the system comprising: a voice acquisition and preprocessing module, used to acquire the plurality of voice signals and preprocess the voice signals to obtain independent voice command signals corresponding to different voice sources.
[0072] The emotion feature perception and calibration module is used to extract the speech features of the independent speech command signal, and calibrate the speech features in combination with environmental acoustic features to obtain calibrated speech emotion feature information.
[0073] The semantic parsing task mapping module generates instruction text by performing speech recognition based on independent speech command signals, parses the instruction semantics of the instruction text, matches the instruction semantics with pre-stored action task priority information, and obtains the priority information of the action task corresponding to the instruction semantics.
[0074] The dynamic confidence-weighted fusion module is used to evaluate the corresponding confidence levels based on the calibrated speech emotion feature information, instruction semantic information, action task priority information, and environmental safety perception information, and to calculate the multi-source priority information fusion score through dynamic confidence-weighted fusion.
[0075] The interaction module is used to save the task state of the current action task and the physical state of the robot, and select and execute the action task based on the multi-source priority information fusion score.
[0076] According to an embodiment of the present invention, the execution interaction module is further configured to enter a state pending clarification when the instruction information corresponding to the multi-source priority information fusion score is incomplete, actively detect potential risk areas, send clarification inquiry information, and continue to execute the interrupted action task from the recoverable breakpoint state according to the clarification instruction or autonomous perception result.
[0077] A third aspect of the present invention provides a composite robot, the composite robot including the composite robot voice command recognition system described above.
[0078] The beneficial effects provided by this invention are as follows: First, the method acquires multiple voice signals and performs preprocessing to obtain independent voice command signals corresponding to different voice sources, which effectively solves the problem that it is difficult to separate mixed audio streams from different operators when the robot receives them simultaneously or within a short period of time in a multi-person collaborative scenario.
[0079] Secondly, the speech features of independent speech command signals are extracted and calibrated by combining them with environmental acoustic features to obtain calibrated speech emotion feature information. This enables the system to overcome noise interference in complex production workshop environments and improve the accuracy of speech emotion recognition.
[0080] Furthermore, the system generates instruction text through speech recognition and parses the semantics of the instructions. At the same time, it matches the text with pre-stored action task priority information to obtain the corresponding priority information. This enables the system to perform in-depth analysis of instruction semantics, distinguish the urgency and importance of different instructions, and thus flexibly determine the execution priority of the instructions.
[0081] Finally, based on the calibrated speech emotion features, instruction semantics, action task priority, and environmental safety perception, the corresponding confidence levels are evaluated. A multi-source priority information fusion score is then calculated using a dynamic confidence-weighted fusion method, and the action task is selected and executed accordingly. This multi-source information fusion mechanism, particularly the introduction of dynamic confidence weighting, effectively addresses instruction conflicts and incomplete information, ensuring the robot's decision-making accuracy and safety in complex and unexpected situations.
[0082] In summary, the method of this application overcomes the shortcomings of existing technologies in handling unstructured spoken expressions, conflicting multi-source instructions, and incomplete information by independently processing multi-source speech commands, calibrating emotion features adapted to the environment, performing deep semantic parsing and priority matching, and fusing multi-source information with dynamic confidence weighting. This significantly improves the accuracy, robustness, and intelligent decision-making capabilities of composite robot speech command recognition, ensuring the robot's operational efficiency and production safety in complex industrial environments.
Claims
1. A method for recognizing voice commands in a composite robot, characterized in that, The composite robot is used to receive multiple voice signals and perform action tasks corresponding to the voice signals, including: Step S1: Acquire the multiple voice signals and preprocess the voice signals to obtain independent voice command signals corresponding to different voice sources; Step S2: Extract the speech features of the independent speech command signal, and calibrate the speech features in combination with environmental acoustic features to obtain calibrated speech emotion feature information; Step S3: Generate instruction text by performing speech recognition based on independent speech command signals, parse the instruction semantics of the instruction text, match the instruction semantics with pre-stored action task priority information, and obtain the priority information of the action task corresponding to the instruction semantics; Step S4: Based on the calibrated speech emotion feature information, instruction semantic information, action task priority information, and environmental safety perception information, evaluate the corresponding confidence levels respectively, and calculate the multi-source priority information fusion score through dynamic confidence weighted fusion method; Step S5: Save the current task state and the physical state of the robot, select and execute the task based on the multi-source priority information fusion score.
2. The method according to claim 1, characterized in that, The environmental acoustic features include interfering acoustic features such as high-frequency noise or mechanical noise that affect the accuracy of speech emotion recognition; and the dynamic calibration includes adjusting the judgment threshold or feature weight of the emotion features according to the intensity of the interfering acoustic features.
3. The method according to claim 1, characterized in that, The confidence levels evaluated in step S4 include: Based on the emotional priority information corresponding to the voice emotion features, the semantic priority information corresponding to the instruction semantics, the action task priority information, and the safety warning priority information triggered by the environmental perception information, the confidence of the voice emotion features, the confidence of the instruction semantics, the confidence of the action task priority, and the confidence of the environmental perception results are evaluated respectively.
4. The method according to claim 3, characterized in that, The confidence level is dynamically adjusted according to at least one of the following: the confidence level of the voice emotion feature is adjusted according to the intensity of environmental interference. The confidence level of the instruction semantics is adjusted based on the accuracy of instruction semantic recognition and the explicitness of keyword matching; The confidence level of the action task priority is adjusted based on the stability of the source of the action task priority; The confidence level of the environmental perception results is adjusted based on whether the accuracy of physical security risk identification exceeds a set threshold.
5. The method according to claim 1, characterized in that, The environmental perception information includes: when a potential physical security risk is predicted, generating a self-triggered security warning without receiving a voice command, and using the self-triggered security warning as the highest priority information in the multi-source priority information fusion.
6. The method according to claim 1, characterized in that, The task status is saved, including the current task execution progress information. The physical status is saved, including the robot pose, the actuator status, and the status of the carried materials. The task status and physical status are marked as recoverable breakpoint states.
7. The method according to any one of claims 1 to 6, characterized in that, Step S5 further includes: When the instruction information corresponding to the multi-source priority information fusion score is incomplete, it enters a state awaiting clarification and actively detects potential risk areas while sending clarification inquiry information. Based on the received clarification instructions or autonomous sensing results, the interrupted action task can be resumed from the recoverable power-off state.
8. A composite robot voice command recognition system, characterized in that, The system includes: The voice acquisition and preprocessing module is used to acquire the multiple voice signals and preprocess the voice signals to obtain independent voice command signals corresponding to different voice sources; The emotion feature perception and calibration module is used to extract the speech features of the independent speech command signal, and calibrate the speech features in combination with environmental acoustic features to obtain calibrated speech emotion feature information. The semantic parsing task mapping module generates instruction text by performing speech recognition based on independent speech command signals, parses the instruction semantics of the instruction text, matches the instruction semantics with pre-stored action task priority information, and obtains the priority information of the action task corresponding to the instruction semantics. The dynamic confidence-weighted fusion module is used to evaluate the corresponding confidence based on the calibrated speech emotion feature information, instruction semantic information, action task priority information and environmental safety perception information, and calculate the multi-source priority information fusion score through dynamic confidence-weighted fusion. The interaction module is used to save the task state of the current action task and the physical state of the robot, and select and execute the action task based on the multi-source priority information fusion score.
9. The system according to claim 8, characterized in that, The execution interaction module is also used to enter a pending clarification state when the instruction information corresponding to the multi-source priority information fusion score is incomplete, actively detect potential risk areas, and send clarification inquiry information; and continue to execute the interrupted action task from the recoverable breakpoint state according to the clarification instruction or autonomous perception result.
10. A composite robot, characterized in that, The composite robot includes the composite robot voice command recognition system as described in claim 8.