Control command response method, device, robot and storage medium

By collecting audio signals for semantic recognition and large language model agent discrimination, the problem of single response methods of traditional robots is solved, the diversification and intelligence of robot responses are realized, and the interactive experience and accuracy are improved.

CN119993152BActive Publication Date: 2025-08-15SHANGHAI FOURIER INTELLIGENCE CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510473589.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-08-15
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

Traditional robots have single response methods, poor interaction and low intelligence. Especially in humanoid robot applications, users hope that the robot can respond to voice commands anthropomorphically.

Method used

Acquire audio signals and perform semantic recognition processing, distinguish response methods through large language model agents, including performing actions and/or answering questions, and evaluate the results through rehearsals to determine the robot's feedback, ensuring the accuracy and diversity of the response.

Benefits of technology

It realizes the enrichment and intelligence of robot response methods, improves the interactive experience, reduces power consumption, and improves the accuracy and anthropomorphism of the response.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993152B_ABST
    Figure CN119993152B_ABST
Patent Text Reader

Abstract

The present application relates to a control instruction response method, device, robot and storage medium. The method includes: when a preset trigger condition is met, collecting an audio signal; performing semantic recognition processing on the audio signal to obtain corresponding text information; using a pre-established large language model intelligent agent to discriminate the text information and determine a response method, the response method including: executing an action and / or answering a question; rehearsing the response method based on a constructed robot model, and evaluating the rehearsal result of the response method through a large language model intelligent agent to obtain a corresponding evaluation result; when the evaluation result meets the requirements, controlling the robot to broadcast the answer to the question generated by the robot model during the rehearsal, and / or executing the action sequence generated by the robot model during the rehearsal. In this way, the response method can be adaptively determined according to the text content corresponding to the audio signal, making the response method more diversified, the response feedback more intelligent, and improving the interactive experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a control instruction response method, apparatus, computer equipment, computer-readable storage medium, and computer program product. Background Art

[0002] With the development of artificial intelligence technology, more and more electronic devices can support human-computer interaction functions. Among them, the human-computer interaction function supports users to wake up voice assistants, send voice commands to electronic devices, have conversations and ask questions with electronic devices, etc., so that users can quickly acquire knowledge and control devices.

[0003] Traditionally, robots respond to voice commands and other forms of control by simply providing a spoken response. For example, after recognizing the voice, robots retrieve the corresponding answer from a knowledge base and then announce it. This response model lacks intelligence, making it particularly effective for humanoid robots, where users prefer anthropomorphic responses.

[0004] However, when applying traditional technology and robot Q&A, the robot's response method is single and the interactivity is poor. Summary of the Invention

[0005] Based on this, it is necessary to provide a control instruction response method, device, computer equipment, computer-readable storage medium and computer program product that can enrich the response mode and enhance interactivity in response to the above technical problems.

[0006] In a first aspect, the present application provides a control instruction response method, comprising:

[0007] When the preset trigger conditions are met, the audio signal is collected;

[0008] Performing semantic recognition processing on the audio signal to obtain corresponding text information;

[0009] The text information is judged by a pre-established large language model agent to determine a response method, wherein the response method includes: performing an action and / or answering a question;

[0010] Previewing the response method based on the constructed robot model, and evaluating the preview result of the response method through the large language model agent to obtain a corresponding evaluation result;

[0011] When the evaluation result meets the requirements, the robot is controlled to broadcast the answer to the question generated by the robot model during the preview, and / or execute the action sequence generated by the robot model during the preview.

[0012] In one embodiment, the preset trigger condition includes at least one of the following:

[0013] An object is detected entering the preset detection range and / or a human face is detected;

[0014] The decibel value of the audio signal in the environment is detected to be greater than a preset value;

[0015] Detecting that the audio signal contains the target wake-up word;

[0016] The pre-established visual model collects continuous video frames containing the lip area of the human face, and the image recognition results of the continuous video frames indicate that there is a lip opening and closing movement.

[0017] In one embodiment, performing semantic recognition processing on the audio signal to obtain corresponding text information includes:

[0018] Using an end-to-end speech recognition model to segment the audio signal according to a preset duration, and converting the segmented audio signal into a feature vector;

[0019] Decode the feature vector and output text information.

[0020] In one embodiment, the large language model agent includes: a thinking model, a reasoning and action model, wherein:

[0021] The thinking model is used to determine a response method based on the text information;

[0022] The reasoning and action model is used to generate a control instruction for the robot model according to the response mode, and evaluate the preview result of the robot model executing the control instruction to obtain a corresponding evaluation result;

[0023] The reasoning and action model is further configured to regenerate control instructions for the robot model when the evaluation result does not meet the requirements, and evaluate preview results of the robot model executing the control instructions until the evaluation result meets the requirements;

[0024] The reasoning and action model is further configured to feed back response completion indication information to the thinking model when the evaluation result meets the requirements.

[0025] In one embodiment, when the response includes answering a question, the large language model agent evaluates the preview result of the response to obtain a corresponding evaluation result, including:

[0026] The reasoning and action model generates a control instruction based on the text information, and sends the control instruction to the robot model, wherein the control instruction is used to instruct the robot model to broadcast the answer to the question;

[0027] The answer to the question is evaluated by the reasoning and action model to determine whether the answer to the question is complete and / or correct.

[0028] In one embodiment, when the response includes executing an action, evaluating the preview result of the response by the large language model agent to obtain a corresponding evaluation result includes:

[0029] The reasoning and action model generates a control instruction based on the text information, and sends the control instruction to the robot model, wherein the control instruction is used to instruct the robot model to execute an action sequence;

[0030] The reasoning and action model evaluates the execution results of the action sequence to determine whether the actions corresponding to the action sequence are completed and / or correct.

[0031] In one embodiment, before controlling the robot to announce the answers to the questions generated by the robot model during the rehearsal and to execute the action sequences generated by the robot model during the rehearsal, the method further includes:

[0032] Setting the answers to the questions generated by the robot model during the preview and the action sequences generated by the robot model during the preview to a number of timestamps;

[0033] Through the correspondence between the timestamps, the text content of the answer to the question and the action sequence are aligned, so that the robot can synchronously execute the corresponding action sequence when broadcasting the answer to the question.

[0034] In a second aspect, the present application further provides a control instruction response device, comprising:

[0035] The acquisition module is used to collect audio signals when a preset trigger condition is met;

[0036] A recognition module, configured to perform semantic recognition processing on the audio signal to obtain corresponding text information;

[0037] a determination module, configured to identify the text information using a pre-established large language model agent and determine a response method, wherein the response method includes: performing an action and / or answering a question;

[0038] An evaluation module is used to preview the response mode based on the constructed robot model, and evaluate the preview result of the response mode through the large language model agent to obtain a corresponding evaluation result;

[0039] The control module is used to control the robot to broadcast the answer to the question generated by the robot model during the preview and / or execute the action sequence generated by the robot model during the preview when the evaluation result meets the requirements.

[0040] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0041] When the preset trigger conditions are met, the audio signal is collected;

[0042] Performing semantic recognition processing on the audio signal to obtain corresponding text information;

[0043] The text information is judged by a pre-established large language model agent to determine a response method, wherein the response method includes: performing an action and / or answering a question;

[0044] Previewing the response method based on the constructed robot model, and evaluating the preview result of the response method through the large language model agent to obtain a corresponding evaluation result;

[0045] When the evaluation result meets the requirements, the robot is controlled to broadcast the answer to the question generated by the robot model during the preview, and / or execute the action sequence generated by the robot model during the preview.

[0046] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the following steps are implemented:

[0047] When the preset trigger conditions are met, the audio signal is collected;

[0048] Performing semantic recognition processing on the audio signal to obtain corresponding text information;

[0049] The text information is judged by a pre-established large language model agent to determine a response method, wherein the response method includes: performing an action and / or answering a question;

[0050] Previewing the response method based on the constructed robot model, and evaluating the preview result of the response method through the large language model agent to obtain a corresponding evaluation result;

[0051] When the evaluation result meets the requirements, the robot is controlled to broadcast the answer to the question generated by the robot model during the preview, and / or execute the action sequence generated by the robot model during the preview.

[0052] In a fifth aspect, the present application further provides a computer program product, comprising a computer program, which, when executed by a processor, implements the following steps:

[0053] When the preset trigger conditions are met, the audio signal is collected;

[0054] Performing semantic recognition processing on the audio signal to obtain corresponding text information;

[0055] The text information is judged by a pre-established large language model agent to determine a response method, wherein the response method includes: performing an action and / or answering a question;

[0056] Previewing the response method based on the constructed robot model, and evaluating the preview result of the response method through the large language model agent to obtain a corresponding evaluation result;

[0057] When the evaluation result meets the requirements, the robot is controlled to broadcast the answer to the question generated by the robot model during the preview, and / or execute the action sequence generated by the robot model during the preview.

[0058] The control command response method, apparatus, computer device, computer-readable storage medium, and computer program product described above collect audio signals when preset trigger conditions are met. This allows for proactive audio signal collection when the preset trigger conditions are met, avoiding prolonged monitoring of voice signals in the environment and reducing power consumption. The audio signals are semantically recognized and processed to obtain corresponding text information. This text information is then interpreted by a pre-established large language model agent to determine a response method, including performing an action and / or answering a question. This allows for a more diverse and intelligent response method based on the text corresponding to the voice information. The response method is previewed based on a constructed robot model, and the previewed response method is evaluated by the large language model agent to obtain a corresponding evaluation result. This allows for pre-rehearsing and evaluation of the response method results, ensuring more accurate control command execution by the robot. When the evaluation result meets the requirements, the robot is controlled to announce the answer to the question generated by the previewed robot model and / or execute the action sequence generated by the previewed robot model. In this way, the response method can be adaptively determined according to the text content corresponding to the audio signal, making the response method more diversified, the response feedback more intelligent, and improving the interactive experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments of the present application or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying any creative work.

[0060] Figure 1 1. A diagram illustrating an application environment of a control instruction response method in an embodiment;

[0061] Figure 2 1 is a flow chart of a control instruction response method according to an embodiment;

[0062] Figure 3 Schematic diagram of the working principle of a large language model agent in one embodiment;

[0063] Figure 4 Schematic diagram of a flow chart of a control instruction response method in another embodiment;

[0064] Figure 5 is a structural block diagram of a control instruction response device in one embodiment;

[0065] Figure 6 is a structural block diagram of a control instruction response device in another embodiment;

[0066] Figure 7 FIG. 4 is a diagram showing the internal structure of a processing system of a robot in one embodiment. DETAILED DESCRIPTION

[0067] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0068] The control instruction response method provided in the embodiment of the present application can be applied to Figure 1In the application environment shown, robot 101 communicates with a server via a network. An offline knowledge base is stored in the local memory of robot 101, and robot 101 can also obtain data from a cloud knowledge base from the server via the network. Exemplarily, each robot 101 has its own control range (as indicated by the dashed box). When a person enters this control range and a preset trigger condition is met, robot 101 collects an audio signal; performs semantic recognition processing on the audio signal to obtain corresponding text information; uses a pre-established large language model agent to identify the text information and determine a response method, which includes: executing an action and / or answering a question; rehearses the response method based on the constructed robot model, and evaluates the rehearsal result of the response method using the large language model agent to obtain a corresponding evaluation result; and when the evaluation result meets the requirements, controls robot 101 to broadcast the answer to the question generated by the robot model during the rehearsal and / or execute the action sequence generated by the robot model during the rehearsal.

[0069] In an exemplary embodiment, Figure 2 As shown, a control instruction response method is provided, which is applied to Figure 1 The robot in FIG is taken as an example to illustrate the method, which includes the following steps 201 to 205. Among them:

[0070] Step 201: When a preset trigger condition is met, an audio signal is collected.

[0071] In this embodiment, audio signal collection is started only when a preset trigger condition is met, thereby reducing the power consumption of the robot and preventing the robot from being in the audio signal collection state for a long time.

[0072] Exemplarily, the preset trigger condition includes at least one of the following:

[0073] 1) An object is detected entering the preset detection range and / or a human face is detected.

[0074] 2) The decibel value of the audio signal in the environment is detected to be greater than the preset value.

[0075] 3) Detect that the audio signal contains the target wake-up word.

[0076] 4) The pre-established visual model captures continuous video frames containing the lip area of the human face, and the image recognition results of the continuous video frames indicate the presence of a lip opening and closing movement.

[0077] For method 1), various distance sensors can be used to detect whether an object enters a preset detection range. For example, ultrasonic sensors, infrared sensors, laser sensors, and radar sensors are all distance sensors that detect objects within a certain range of the robot.

[0078] For method 1), a visual sensor can also be used to detect whether a face image appears within a preset detection range. If a face image exists, it means that someone is approaching.

[0079] For example, when a robot uses a visual sensor, it can also be configured to recognize a target person. Specifically, the robot will only wake up when a target person enters the robot's control range. For example, when the visual sensor detects a person entering the control range, it captures a facial image of the person and compares it with a reference target facial image. If the comparison is successful, the robot wakes up to collect audio signals.

[0080] Regarding method 2), when the robot collects audio signals from the surrounding environment, it does not analyze the content of the audio signals, but only judges the decibel value of the audio signals. When the decibel value is greater than the preset value, the robot is triggered to collect audio signals for a long period of time. At this time, the collected audio signals need to be semantically recognized and processed to obtain the corresponding text information.

[0081] For method 3), the robot will wake up to collect audio signals only when the target wake-up word is detected.

[0082] Regarding method 4), compared with other methods, this is a more accurate detection method. It uses a pre-established visual model to collect continuous video frames containing the lip area of the face, and then recognizes and processes these continuous video frames to determine whether there is any lip opening and closing movement. If there is any lip opening and closing movement, it uses a narrow beam to more accurately receive the audio signal.

[0083] Step 202: Perform semantic recognition processing on the audio signal to obtain corresponding text information.

[0084] Exemplarily, an end-to-end speech recognition model may be used to segment the audio signal according to a preset duration, and convert the segmented audio signal into a feature vector; the feature vector is decoded and output as text information.

[0085] In this embodiment, the latest end-to-end Paraformer model can be used to achieve efficient and accurate speech recognition. It converts audio signals into feature vectors, processes and decodes them using acoustic and language models, and ultimately outputs text. This semantic recognition method offers multilingual support and real-time inference capabilities, making it suitable for efficient speech recognition in low-latency environments.

[0086] Step 203: The text information is judged by the pre-established large language model agent to determine the response method.

[0087] The response method includes: performing an action and / or answering a question.

[0088] In this embodiment, the large language model intelligent agent includes: a thinking model and a reasoning and action model, wherein: the thinking model is used to determine the response method based on the text information; the reasoning and action model is used to generate control instructions for the robot model based on the response method, and evaluate the preview results of the robot model executing the control instructions to obtain corresponding evaluation results; the reasoning and action model is also used to regenerate control instructions for the robot model when the evaluation results do not meet the requirements, and evaluate the preview results of the robot model executing the control instructions until the evaluation results meet the requirements; the reasoning and action model is also used to feedback response completion indication information to the thinking model when the evaluation results meet the requirements.

[0089] For example, Figure 3 The figure below illustrates the working principle of the large language model agent. Assume that the text message corresponding to the audio signal is "wave your hand." The large language model agent's thought model (Thought) then determines whether to perform an action or answer a question. If the judgment result is "perform an action," the large language model agent's reasoning and action model (ReAct) then performs the following steps:

[0090] ReAct1 sends a control instruction to the robot model: wave your hand.

[0091] Obs1 determines whether the waving action is complete / correct.

[0092] If it is not completed / incorrect, ReAct2 sends a control instruction to the robot model: wave your hand.

[0093] Obs2 determines whether the waving action is complete / correct.

[0094] If it is not completed / incorrect, ReAct3 sends a control instruction to the robot model: wave your hand.

[0095] Obs3 determines whether the waving action is complete or correct.

[0096]

[0097] Until the waving action is completed / correct, the robot model gives feedback to Thought and the action is completed.

[0098] The above ReAct1, ReAct2, ReAct3, and Obs1, Obs2, and Obs3 respectively represent the three reasoning and dynamic decision-making processes of the reasoning and action model. Among them, ReAct is an intelligent agent model that combines reasoning and action, which aims to enable the model to reason when performing tasks and take corresponding actions based on the reasoning results. It can handle complex problems through reasoning, generate reasonable intermediate reasoning steps, and help intelligent agents make wise decisions. When performing tasks, ReAct can dynamically adjust decision-making strategies according to the progress of the task and changes in the environment to adapt to different situations. In addition, ReAct can perform self-optimization through reinforcement learning, gradually improve the effectiveness of decisions and actions, and break down complex tasks into multiple small tasks, and complete the overall goal by gradually solving these subtasks.

[0099] Optionally, by combining multiple information sources such as text, voice or images, ReAct can perceive the environment more comprehensively and make more appropriate decisions based on information from different modalities.

[0100] Similarly, the large language model agent's thought model (Thought) judges the text information and determines whether to perform an action or answer the question. In this case, the judgment result is "answer the question." The large language model agent's reasoning and action model (ReAct) then performs the following steps:

[0101] ReAct1 answers related questions.

[0102] Obs1 determines whether the answer is complete / correct.

[0103] If incomplete / incorrect, ReAct2 will re-answer the relevant questions.

[0104] Obs2 determines whether the answer is complete / correct.

[0105] If incomplete / incorrect, ReAct3 will ask you to answer the relevant questions again.

[0106] Obs3 determines whether the answer is complete / correct.

[0107]

[0108] The answer is complete / correct when it is complete.

[0109] Step 204 : Preview the response method based on the constructed robot model, and evaluate the preview result of the response method through the large language model intelligent agent to obtain a corresponding evaluation result.

[0110] Exemplarily, when the response method includes answering a question, the reasoning and action model generates a control instruction based on the text information and sends the control instruction to the robot model, and the control instruction is used to instruct the robot model to broadcast the answer to the question; the reasoning and action model evaluates the answer to the question to determine whether the answer to the question is complete and / or correct.

[0111] Exemplarily, when the response method includes executing an action, the reasoning and action model generates a control instruction based on the text information and sends the control instruction to the robot model, and the control instruction is used to instruct the robot model to execute an action sequence; the reasoning and action model evaluates the execution result of the action sequence to determine whether the action corresponding to the action sequence is completed and / or correct.

[0112] In this embodiment, the large language model agent can continuously check whether it has obtained a complete / correct answer to the question, or whether it has obtained a correct action sequence. If the answer to the question is incomplete / incorrect, or the action sequence is incorrect, it can continue to loop until a complete / correct answer to the question or a correct action sequence is generated. This makes the robot's response more accurate.

[0113] It should be noted that when it comes to the two response methods of answering questions and performing actions, the above two methods can be combined to achieve answering questions and performing actions at the same time, so that the robot's response results are more humanized and the interactivity is greatly enhanced.

[0114] Step 205 : When the evaluation result meets the requirements, the robot is controlled to broadcast the answer to the question generated by the robot model during the preview, and / or to execute the action sequence generated by the robot model during the preview.

[0115] In this embodiment, the large language model agent can accurately generate complete / correct answers to questions, or generate a sequence of correct actions, so that the robot can determine different response methods based on different speech semantics, making the response methods more flexible and varied.

[0116] In the control command response method described above, audio signals are collected when preset trigger conditions are met. This allows for proactive audio signal collection when the preset trigger conditions are met, avoiding long-term monitoring of voice signals in the environment and reducing power consumption. Semantic recognition processing is performed on the audio signals to obtain corresponding text information. A pre-established large language model agent interprets the text information and determines a response method, including performing an action and / or answering a question. This allows for a more diverse and intelligent response method based on the text corresponding to the voice information. The response method is previewed based on a constructed robot model, and the previewed response method is evaluated by the large language model agent to obtain a corresponding evaluation result. This allows for the pre-preview and evaluation of the response method results, ensuring more accurate control command execution by the robot. When the evaluation result meets the requirements, the robot is controlled to announce the answer to the question generated by the preview model and / or execute the action sequence generated by the preview model. This allows for adaptively determining the response method based on the text content corresponding to the audio signal, resulting in more diverse response methods, more intelligent feedback, and an enhanced interactive experience.

[0117] In another exemplary embodiment, Figure 4 As shown, a control instruction response method is provided, which is applied to Figure 1 The robot in the example is used to illustrate the process, including the following steps 401 to 406. Among them:

[0118] Step 401: When a preset trigger condition is met, an audio signal is collected.

[0119] Step 402: Perform semantic recognition processing on the audio signal to obtain corresponding text information.

[0120] Step 403: The text information is judged by the pre-established large language model agent to determine the response method.

[0121] Step 404: preview the response method based on the constructed robot model, and evaluate the preview result of the response method through the large language model intelligent agent to obtain a corresponding evaluation result.

[0122] For the specific implementation process and technical effects of steps 401 to 404 in this embodiment, please refer to Figure 2 The descriptions of steps 201 to 204 in the illustrated method embodiment are not repeated here.

[0123] Step 405 : setting the answers to the questions generated by the robot model during the preview and the action sequences generated by the robot model during the preview to a plurality of timestamps.

[0124] In this embodiment, the number of timestamps to add can be determined based on the length of the answer text and the length of the action sequence. For example, the robot's speaking speed can be empirically determined, and the estimated duration of the answer can be determined based on that speed. Similarly, the estimated duration of the robot's action sequence can be empirically determined. Timestamps can then be added based on the duration of the answer and the action sequence, for example, adding a timestamp every two seconds.

[0125] Step 406 , aligning the text content of the answer to the question and the action sequence through the correspondence between the timestamps, so that the robot synchronously executes the corresponding action sequence when announcing the answer to the question.

[0126] In this embodiment, after step 405, a question answer including N timestamps and an action sequence including N timestamps can be obtained, where N is a natural number greater than 1. Furthermore, by aligning the same timestamps, the aligned question answer and action sequence can be obtained.

[0127] In this embodiment, the robot can complete the synchronous execution of the answer to the question and the action sequence according to the aligned answer to the question, so that the robot can synchronously execute the corresponding action sequence in the process of answering the question.

[0128] Step 407 : When the evaluation result meets the requirements, the robot is controlled to broadcast the answer to the question generated by the robot model during the preview, and to execute the action sequence generated by the robot model during the preview.

[0129] For the specific implementation process and technical effects of step 407 in this embodiment, please refer to Figure 2 The description of step 205 in the illustrated method embodiment will not be repeated here.

[0130] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0131] Based on the same inventive concept, the present application also provides a control instruction response device for implementing the control instruction response method mentioned above. The implementation solution provided by this device is similar to the implementation solution described in the above method. Therefore, the specific limitations of one or more control instruction response device embodiments provided below can be found in the above-mentioned limitations of the control instruction response method and will not be repeated here.

[0132] In an exemplary embodiment, Figure 5 As shown, a control instruction response device is provided, including: an acquisition module 501, an identification module 502, a determination module 503, an evaluation module 504 and a control module 505, wherein:

[0133] The acquisition module 501 is used to acquire audio signals when a preset trigger condition is met;

[0134] The recognition module 502 is used to perform semantic recognition processing on the audio signal to obtain corresponding text information;

[0135] A determination module 503 is configured to identify the text information using a pre-established large language model agent and determine a response method, wherein the response method includes: performing an action and / or answering a question;

[0136] An evaluation module 504 is configured to preview the response mode based on the constructed robot model, and evaluate the preview result of the response mode through the large language model agent to obtain a corresponding evaluation result;

[0137] The control module 505 is used to control the robot to broadcast the answer to the question generated by the robot model during the preview and / or execute the action sequence generated by the robot model during the preview when the evaluation result meets the requirements.

[0138] Exemplarily, the preset trigger condition includes at least one of the following:

[0139] An object is detected entering the preset detection range and / or a human face is detected;

[0140] The decibel value of the audio signal in the environment is detected to be greater than a preset value;

[0141] Detecting that the audio signal contains the target wake-up word;

[0142] The pre-established visual model collects continuous video frames containing the lip area of the human face, and the image recognition results of the continuous video frames indicate that there is a lip opening and closing movement.

[0143] Exemplarily, the recognition module 502 is specifically configured to: segment the audio signal according to a preset duration using an end-to-end speech recognition model, and convert the segmented audio signal into a feature vector; decode the feature vector and output text information.

[0144] Exemplarily, the large language model agent includes: a thinking model, a reasoning and action model, wherein:

[0145] The thinking model is used to determine a response method based on the text information;

[0146] The reasoning and action model is used to generate a control instruction for the robot model according to the response mode, and evaluate the preview result of the robot model executing the control instruction to obtain a corresponding evaluation result;

[0147] The reasoning and action model is further configured to regenerate control instructions for the robot model when the evaluation result does not meet the requirements, and evaluate preview results of the robot model executing the control instructions until the evaluation result meets the requirements;

[0148] The reasoning and action model is further configured to feed back response completion indication information to the thinking model when the evaluation result meets the requirements.

[0149] Exemplarily, the evaluation module 504 is specifically used to: generate a control instruction based on the text information by the reasoning and action model, and send the control instruction to the robot model, wherein the control instruction is used to instruct the robot model to broadcast the answer to the question; and evaluate the answer to the question by the reasoning and action model to determine whether the answer to the question is complete and / or correct.

[0150] Exemplarily, the evaluation module 504 is specifically used to: generate a control instruction based on the text information by the reasoning and action model, and send the control instruction to the robot model, wherein the control instruction is used to instruct the robot model to execute an action sequence; and evaluate the execution result of the action sequence by the reasoning and action model to determine whether the action corresponding to the action sequence is completed and / or correct.

[0151] In another exemplary embodiment, Figure 6 As shown, a control instruction response device is provided. Figure 5 Based on the device shown, it can also include:

[0152] A setting module 506 is used to set the answers to the questions generated by the robot model during the preview and the action sequences generated by the robot model during the preview to a number of timestamps;

[0153] The alignment module 507 is used to align the text content of the answer to the question and the action sequence through the correspondence between the timestamps, so that the robot can synchronously execute the corresponding action sequence when broadcasting the answer to the question.

[0154] Each module in the control instruction response device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.

[0155] In an exemplary embodiment, a processing system of a robot is provided, and the internal structure diagram of the processing system can be shown as follows: Figure 7 As shown. The processing system includes a processor, memory, an input / output interface, a communication interface, a display unit, and an input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are connected to the system bus via the input / output interface. The processor of the processing system is used to provide computing and control capabilities. The memory of the processing system includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The input / output interface of the processing system is used to exchange information between the processor and external devices. The communication interface of the processing system is used to communicate with external terminals via wired or wireless means, and the wireless means can be implemented via Wi-Fi, a mobile cellular network, near field communication (NFC), or other technologies. When executed by the processor, the computer program implements a method for generating robot motions. The display unit of the processing system is used to form a visually visible image, and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the processing system can be a touch layer covering the display screen, or a button, trackball or touchpad set on the robot shell, or an external keyboard, touchpad or mouse.

[0156] Those skilled in the art will understand that Figure 7 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0157] In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.

[0158] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0159] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0160] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0161] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), quantum computing-based data processing logic devices, artificial intelligence (AI) processors, and the like.

[0162] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0163] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A control instruction response method, characterized in that: The method comprises: When the preset trigger conditions are met, the audio signal is collected; An end-to-end Paraformer model is used to segment the audio signal according to a preset duration, and the segmented audio signal is converted into a feature vector; Decoding the feature vector and outputting text information; The text information is judged by a pre-established large language model agent to determine a response method, which includes: performing an action and answering a question; Previewing the response method based on the constructed robot model, and evaluating the preview result of the response method through the large language model agent to obtain a corresponding evaluation result; Setting a number of timestamps for the answers to the questions and the action sequences generated by the robot model during the preview; wherein the number of timestamps is related to the text length of the answers to the questions and the length of the action sequences; By using the correspondence between the timestamps, the text content of the answer to the question and the action sequence are aligned so that the robot synchronously executes the corresponding action sequence when announcing the answer to the question; When the evaluation result meets the requirements, the robot is controlled to broadcast the answers to the questions generated by the robot model during the preview and to execute the action sequences generated by the robot model during the preview.

2. The method according to claim 1, characterized in that The preset triggering condition includes at least one of the following: An object is detected entering the preset detection range and / or a human face is detected; The decibel value of the audio signal in the environment is detected to be greater than a preset value; Detecting that the audio signal contains the target wake-up word; The pre-established visual model collects continuous video frames containing the lip area of the human face, and the image recognition results of the continuous video frames indicate that there is a lip opening and closing movement.

3. The method according to claim 1 or 2, characterized in that The large language model agent includes: a thinking model, a reasoning and action model, wherein: The thinking model is used to determine a response method based on the text information; The reasoning and action model is used to generate a control instruction for the robot model according to the response mode, and evaluate the preview result of the robot model executing the control instruction to obtain a corresponding evaluation result; The reasoning and action model is further configured to regenerate control instructions for the robot model when the evaluation result does not meet the requirements, and evaluate preview results of the robot model executing the control instructions until the evaluation result meets the requirements; The reasoning and action model is further configured to feed back response completion indication information to the thinking model when the evaluation result meets the requirements.

4. The method according to claim 3, characterized in that When the response mode includes answering a question, the large language model agent evaluates the preview result of the response mode to obtain a corresponding evaluation result, including: The reasoning and action model generates a control instruction based on the text information, and sends the control instruction to the robot model, wherein the control instruction is used to instruct the robot model to broadcast the answer to the question; The answer to the question is evaluated by the reasoning and action model to determine whether the answer to the question is complete and / or correct.

5. The method according to claim 3, characterized in that When the response mode includes executing an action, the large language model agent evaluates the preview result of the response mode to obtain a corresponding evaluation result, including: The reasoning and action model generates a control instruction based on the text information, and sends the control instruction to the robot model, wherein the control instruction is used to instruct the robot model to execute an action sequence; The reasoning and action model evaluates the execution results of the action sequence to determine whether the actions corresponding to the action sequence are completed and / or correct.

6. A control instruction response device, characterized in that: The device comprises: The acquisition module is used to collect audio signals when a preset trigger condition is met; A recognition module is configured to segment the audio signal according to a preset duration using an end-to-end Paraformer model, convert the segmented audio signal into a feature vector, decode the feature vector, and output text information; A determination module is used to identify the text information through a pre-established large language model agent and determine a response method, wherein the response method includes: performing an action and answering a question; An evaluation module is used to preview the response mode based on the constructed robot model, and evaluate the preview result of the response mode through the large language model agent to obtain a corresponding evaluation result; A setting module, configured to set the answers to the questions and the action sequences generated by the robot model during the rehearsal to a plurality of timestamps; wherein the number of timestamps is related to the text length of the answers to the questions and the length of the action sequences; An alignment module, configured to align the text content of the answer to the question with the action sequence based on the correspondence between the timestamps, so that the robot synchronously executes the corresponding action sequence when announcing the answer to the question; The control module is used to control the robot to broadcast the answers to the questions generated by the robot model during the preview and to execute the action sequence generated by the robot model during the preview when the evaluation result meets the requirements.

7. The device according to claim 6, characterized in that The preset triggering condition includes at least one of the following: An object is detected entering the preset detection range and / or a human face is detected; The decibel value of the audio signal in the environment is detected to be greater than a preset value; Detecting that the audio signal contains the target wake-up word; The pre-established visual model collects continuous video frames containing the lip area of the human face, and the image recognition results of the continuous video frames indicate that there is a lip opening and closing movement.

8. The device according to claim 6, characterized in that The large language model agent includes: a thinking model, a reasoning and action model, wherein: The thinking model is used to determine a response method based on the text information; The reasoning and action model is used to generate a control instruction for the robot model according to the response mode, and evaluate the preview result of the robot model executing the control instruction to obtain a corresponding evaluation result; The reasoning and action model is further configured to regenerate control instructions for the robot model when the evaluation result does not meet the requirements, and evaluate preview results of the robot model executing the control instructions until the evaluation result meets the requirements; The reasoning and action model is further configured to feed back response completion indication information to the thinking model when the evaluation result meets the requirements.

9. A robot comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Robot control method, device and equipment and storage medium

    CN114227698A

  • Robot operation self-correction method and system based on multi-dimensional data space

    CN118254190A

  • Interaction method and device, electronic equipment, storage medium and program product

    CN119227795A

  • Task execution method based on large language model and API gateway system

    CN119597504A