An Information Interaction Method for Embodied Intelligent Robots
Through the combination of contextual understanding modules and multimodal sensors, real-time monitoring and adjustment of execution strategies is solved, and the task execution and user interaction problems of embodied intelligent robots in complex environments is achieved, achieving efficient and natural human-computer interaction.
Patent Information
- Application Number
- CN202411299785.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-18
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2044-09-18
AI Technical Summary
Existing embodied intelligent robots have challenges in understanding complex contexts, environmental perception and multimodal fusion. They are difficult to accurately execute user instructions in highly dynamic scenarios, and lack personalized interaction capabilities, resulting in poor user experience.
The context understanding module is used to analyze user instructions, combine environmental perception and multimodal sensors, optimize the robot's interactive capabilities through deep learning and reinforcement learning, monitor environmental parameters in real time and adjust execution strategies, and use natural language processing and multiple sensor data for comprehensive perception and feedback.
It improves the task execution efficiency and user experience of robots in high dynamic environments, enhances their understanding of complex instructions and environmental adaptability, provides a more natural human-computer interaction experience, and reduces user learning costs.
Smart Images

Figure CN119260754B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and specifically relates to an information interaction method for an embodied intelligent robot. Background Art
[0002] In the current intelligent robot technology, the information interaction method of embodied intelligent robots is a key research area. This technology utilizes the principle of embodied cognition to directly associate behaviors and perceptions in the physical environment, enabling the robot to obtain environmental information through its own sensors and understand and execute instructions from users in combination with the situation it is in. However, in practical applications, there are still many challenges in accurately parsing and following human commands in different or special scenarios. The existing technology mainly relies on the continuous improvement of human language understanding algorithms, the enhancement of environmental context awareness, and the integration of multi-modal sensing technologies. These technological advancements enable the robot to not only accurately understand speech but also combine non-verbal cues, such as body movements and emotional expressions, for more precise task execution.
[0003] However, the current solutions still have obvious limitations in some specific situations: Limitations in language understanding: Although natural language processing technology is constantly advancing, for complex contexts or instructions with multiple meanings, existing algorithms often struggle to comprehensively and accurately parse them, which may lead to understanding deviations or execution errors when the robot faces ambiguity, advanced reasoning, and implicit commands; Insufficient environmental perception ability: The environmental perception of intelligent robots depends on the accuracy of sensors and data processing capabilities. However, existing sensors have insufficient perception accuracy and response speed in complex environments (such as strong light, low light, noisy backgrounds) or when facing rapidly changing scenarios, which limits the adaptability and execution efficiency of the robot; Challenges in multi-modal fusion: Although the fusion of multi-modal sensors helps improve the robot's understanding of non-verbal signals (such as gestures, facial expressions), in practical applications, there are still technical bottlenecks in the synchronization, association, and processing of different modal data. Especially in high-dynamic scenarios, how to achieve real-time and accurate integration between different data sources is a difficult problem; Lack of personalized interaction: Existing embodied intelligent robots lack sufficient personalized interaction capabilities when dealing with different users, and it is difficult to adapt to the unique instruction expression methods and individual differences of users, which may lead to a poor user experience in practical applications and limit the popularity and application effect of the robot. Therefore, although embodied intelligent robots show great potential in the field of information interaction, the existing technology still faces significant challenges in understanding, perception, and execution. Summary of the Invention
[0004] The purpose of the present invention is to provide an information interaction method for an embodied intelligent robot, so as to solve the problem that the existing technology still faces significant challenges in understanding, perception, and execution.
[0005] To achieve the above object, the present invention provides the following technical solution: An information interaction method for an embodied intelligent robot, the method comprising:
[0006] S1: Receive the information containing the task instruction sent by the user;
[0007] S2: Parse the information based on the context understanding module to identify the specific meaning of the task instruction and the environmental conditions;
[0008] S3: Execute the operation corresponding to the specific meaning and adjust the execution parameters in a specific environment to ensure the accurate implementation of the task instruction;
[0009] Wherein, the step S2 includes:
[0010] S21: Extract the key phrases and words expressed by the user in the information as keywords;
[0011] S22: Analyze the correlation degree between the keywords and the task type;
[0012] S23: Identify the environmental conditions based on the currently perceived environmental features;
[0013] S24: Use the context understanding algorithm to parse out the user's intention;
[0014] The step S3 includes:
[0015] S31: Real-time monitor the current environmental parameters, including temperature, light intensity, sound level;
[0016] S32: If the current environmental parameters exceed the expected working range, adjust the robot's action parameters;
[0017] S33: Based on the environmental adaptation adjustment factor Adjust the task execution speed If Less than the standard value Then prompt that a strategy change is required;
[0018] S34: Update the parameter adjustment result for reference use and feedback whether the adjustment is necessary.
[0019] Preferably, the step S1 includes:
[0020] S11: Capture the user's voice signal through a microphone array;
[0021] S12: Convert the captured voice signal into a digital audio file;
[0022] S13: Use natural language processing technology to extract text from the converted digital audio file to obtain the task instruction in text form;
[0023] S14: Separate the voice information from the received multimedia information and convert it into text;
[0024] S15: Perform voice command activation word recognition to start the information receiving process.
[0025] Preferably, the step S13 includes:
[0026] Adopt a deep learning model to perform semantic analysis on the text, use the multi-level embedding vectors of the pre-trained language model, understand the complex task instruction context through the attention mechanism, and extract the key task words. The specific formula is: ,
[0027] where, represents mechanism, which is used to improve the focusing ability of the model on the input information, represents the query matrix, represents the key matrix, represents the value matrix, is a normalization function, represents the dimension size of the key vector, represents the key matrix transpose.
[0028] Preferably, the step S22 includes:
[0029] Construct a term frequency vector for the key phrase provided by the user ;
[0030] Obtain the key phrases of the preset type tasks from the task library to form a term frequency vector ;
[0031] Calculate the similarity between and based on the cosine similarity formula , and the specific formula is: ,
[0032] where, is the calculated cosine similarity score, represents the term frequency vector constructed for the key phrase provided by the user, represents the term frequency vector formed by obtaining the key phrases of the preset type tasks from the task library, represents the vector in the th element, represents the vector in the th element, represents the vector and The serial number of the element represents the length of the vector;
[0033] If is higher than a predetermined threshold, the keyword is considered to be associated with a certain task type.
[0034] Preferably, the step S24 includes:
[0035] Introduce a dynamic context window mechanism. When parsing the user's intention, adaptively adjust the length and depth of the context information according to the complexity of the current task. Use the sequence-to-sequence model combined with the attention mechanism to deeply analyze the user's intention. Combine the user's historical interaction information with the current instruction through context encoding to form a comprehensive understanding of the user's intention. The parsing formula of the user's intention is: ,
[0036] wherein, is the user's intention, represents mechanism, represents the encoding matrix of the past context information, represents the encoding vector of the current instruction, is a normalization function, represents the transpose of the current instruction encoding, represents the dimension size of the key vector.
[0037] Preferably, the step S31 includes:
[0038] Real-time monitor the current environmental parameters through various sensors built in the robot. Use exponential weighted moving average filtering to process the sensor data to remove noise and obtain stable and accurate environmental parameter values. Use these data as the basis for subsequent step judgment and adjustment. The specific formula is:
[0039] ,
[0040] wherein, represents the estimated value at the current time , represents the time, represents the measured value of the sensor at the current time, represents the estimated value at the previous time, represents the smoothing factor, and the value range is 0 ≤ ≤ 1.
[0041] Preferably, the step S32 includes:
[0042] Set the working range of environmental parameters, and use the non-linear constraint optimization method to adjust the robot's action parameters. The optimization function takes into account various requirements for task execution. Specifically, the Lagrange multiplier method is adopted, and by introducing the Lagrange multiplier , the constraint conditions are integrated into the objective function, which is transformed into an unconstrained optimization problem. The specific formula is: ,
[0043] where, is the Lagrangian function, is the objective function, represents the robot's action parameters, is the Lagrange multiplier, represents the constraint function, ensuring that the action parameters are within the working range of the environmental parameters, is the serial number of the constraint function.
[0044] Preferably, the step S33 includes:
[0045] The environmental adaptation adjustment factor Adjusts the task execution speed , and makes dynamic adjustments according to the changes in environmental conditions. The adjustment factor is updated in real time through the recursive least squares method , and a standard value is set. When is lower than , the robot will issue a prompt, suggesting changing the task execution strategy or pausing the task. The specific formula is: ,
[0046] where, represents the adjustment factor at the current moment, represents the moment, represents the adjustment factor at the previous moment, represents the gain matrix, represents the current desired output or reference value, represents the current environmental feature vector, represents the predicted value of the environmental feature and the adjustment factor at the previous moment;
[0047] Set a standard value to evaluate the adaptability of the adjustment factor. This standard value is set according to the specific requirements of the task, including ensuring the stability of the task or avoiding excessive adjustment. When < , it indicates that the environment is too complex or not conducive to the current task execution, and the robot will issue a prompt, suggesting changing the execution strategy or pausing the task.
[0048] Preferably, the step S34 includes:
[0049] The adjusted results are evaluated using reinforcement learning methods, and the robot records the adjusted results in the experience library to form a value table, and continuously optimizes its decision-making model through a reward mechanism. After each task is completed, the effect of parameter adjustment is evaluated, and it is feedback whether further adjustment or optimization is required.
[0050] Preferably, the value table is adjusted using the reinforcement learning method update formula. According to the updated value table, continuously optimize the action strategy of the robot in different states, and select the action that can maximize the value;
[0051] The reward mechanism includes giving a higher positive reward when the task is completed, giving a negative reward when an invalid action is executed or the task fails, and obtaining an additional reward by saving resources or improving efficiency during the task execution process;
[0052] After each task is completed, the robot evaluates the effect of the current parameter adjustment. The evaluation criteria include task success rate, execution time, energy consumption, and execution accuracy. The robot will feedback the evaluation results to the experience library for further adjustment and optimization. the value table and strategy;
[0053] If the evaluation results show that there are deficiencies in the current adjustment strategy, the robot will adjust the learning parameters or modify the reward mechanism to further optimize its strategy;
[0054] The robot also recommends adjusting the task execution strategy based on the evaluation results, including recommending pausing the task or reducing the task complexity when the conditions are unfavorable.
[0055] From the above technical solutions, the present invention has the following beneficial effects:
[0056] The information interaction method of the embodied intelligent robot receives information containing task instructions sent by the user, analyzes the information based on the context understanding module to identify the specific meaning of the task instructions and environmental conditions, performs operations corresponding to the specific meaning, and adjusts the execution parameters in a specific environment to ensure the accurate implementation of the task instructions. This enhances the environmental perception and adaptation capabilities, can integrate various perceptual information such as vision, hearing, and touch, realizes the comprehensive perception of the environment, improves the accuracy of instruction understanding and execution. The embodied intelligent robot can more accurately parse complex user instructions, perform real-time multimodal data fusion and decision optimization, enabling the robot to achieve real-time integration of different data sources in high-dynamic scenarios, improving the task execution efficiency and user experience, enabling the robot to provide a more natural and intuitive human-computer interaction experience, reducing the user's learning cost and operation complexity, thereby enhancing the practicality and usability of the robot in actual applications, promoting the wide application of intelligent robots in multiple fields, and solving the problem that the existing technology still faces significant challenges in multiple aspects such as understanding, perception, and execution. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] In the figure: Attached Figure 1 is the control flow chart of the information interaction method of the embodied intelligent robot of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0058] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0059] As Figure 1 shown, an information interaction method of an embodied intelligent robot, the method includes:
[0060] S1: Receive information containing task instructions sent by the user;
[0061] S2: Analyze the information based on the context understanding module to identify the specific meaning of the task instructions and environmental conditions;
[0062] S3: Perform operations corresponding to the specific meaning and adjust the execution parameters in a specific environment to ensure the accurate implementation of the task instructions;
[0063] In this embodiment, the embodied intelligent robot receives information containing task instructions issued by the user and uses its internal context understanding module to parse the information in order to identify the specific meaning of the task instructions and the current environmental conditions. This parsing process includes extracting the key phrases and vocabulary expressed by the user, analyzing the relevance of these keywords to the task type, and identifying the environmental conditions in combination with the current environmental characteristics (such as temperature, light, sound, etc.). Finally, the user's intention is parsed through the context understanding algorithm. When performing a task, the robot monitors the environmental parameters in real time. If it is found that the current environmental parameters exceed the expected working range, corresponding adjustments are made. According to the environmental adaptation adjustment factor, the task execution speed is adjusted. If the speed is less than the standard value, it is prompted that a strategy change is needed. The parameters after each adjustment will be updated and used as a reference for subsequent tasks, and at the same time, feedback is given to determine whether further adjustments are needed. The beneficial effect of the present invention is to achieve the accuracy and flexibility of information interaction of the embodied intelligent robot. Through in-depth parsing of the user's task instructions and real-time monitoring and adjustment of environmental parameters, this method can ensure that the robot can dynamically adjust the execution strategy when performing tasks in a complex and changeable environment, thereby improving the success rate of the robot in performing tasks. In addition, by using the parsing ability of the context understanding algorithm, the user's intention can be more accurately identified, reducing errors in the task execution process and improving the user experience. The monitoring and adjustment mechanism of the robot for environmental parameters also increases the robustness, enabling it to operate efficiently under different environmental conditions.
[0064] In an alternative embodiment, different context understanding algorithms can be used, such as a neural network model based on deep learning to replace the traditional rule matching algorithm, so as to improve the parsing ability for complex contexts. In addition, in terms of environmental parameter monitoring, more sensors (such as humidity sensors, barometric pressure sensors) can be introduced to expand the perception range of environmental conditions and enhance the adaptability. In the adjustment of task execution parameters, fuzzy logic control or reinforcement learning algorithms can be adopted to more precisely adjust the behavior parameters of the robot. In terms of the operation method, different environmental adaptation adjustment factors can also be used, such as adding a learning module for user personal preferences, so that more detailed adjustments can be made according to the usage habits of different users. Finally, for the feedback mechanism in the information interaction process, by introducing natural language generation technology, the robot can feedback its operation status and adjustment suggestions in a more natural and user-friendly way, further improving the human-computer interaction experience.
[0065] Step S1 includes:
[0066] S11: Capture the user's voice signal through a microphone array;
[0067] S12: Convert the captured voice signal into a digital audio file;
[0068] S13: Use natural language processing technology to extract text from the converted digital audio file to obtain a task instruction in text form;
[0069] S14: Separate the voice information from the received multimedia information and convert it into text;
[0070] S15: Perform voice command activation word recognition to start the information reception process.
[0071] In this embodiment, the microphone array captures the user's voice signal and converts the captured signal into a digital audio file. The converted audio file is subjected to text extraction through natural language processing technology to obtain a task instruction in text form. During the processing, the voice part in the multimedia information is separated and converted into text, and then voice command activation word recognition is performed to start information reception. These steps ensure that the robot can accurately receive and parse the user's instructions. This method improves the capture accuracy of the voice signal through the microphone array and uses natural language processing technology to accurately extract the task instruction text, effectively improving the accuracy of instruction reception. The recognition mechanism of the voice command activation word ensures that the user's instructions can be responded to in a timely manner when needed, avoiding unnecessary interference and improving the interaction efficiency.
[0072] In different implementations, different microphone array layouts can be adopted or a single microphone can be used in combination with a noise reduction algorithm to capture the voice signal. In addition, various models can be used for natural language processing technology, such as rule-based parsers or more complex neural network models such as BERT, GPT, etc. For the recognition of voice command activation words, a customized voice model can be used to improve the recognition accuracy.
[0073] Step S13 includes:
[0074] Adopt a deep learning model to perform semantic analysis on the text, use the multi-level embedding vectors of the pre-trained language model, and understand the complex task instruction context through the attention mechanism, and extract the key task words. The specific formula is: ,
[0075] where represents mechanism, which is used to improve the focusing ability of the model on the input information, represents the query matrix, represents the key matrix, represents the value matrix, is a normalization function, represents the dimension size of the key vector, represents the key matrix transpose.
[0076] This method uses a deep learning model to parse text, generates multi-level embedding vectors using a pre-trained language model, and focuses on the parts related to the task instructions through an attention mechanism. Specifically, the attention mechanism extracts the key information of complex task instructions by calculating the relationships between the query matrix, key matrix, and value matrix, ensuring that the robot can accurately understand and respond to complex instruction contexts. By adopting the pre-trained language model and the attention mechanism, this implementation can significantly improve the robot's ability to understand complex instructions. The attention mechanism allows the model to focus on the key parts related to the task when processing long texts or complex sentences, thereby improving the accuracy of semantic parsing. This method can adapt to more complex instructions and has strong generalization ability.
[0077] Different deep learning models and pre-trained language models can be selected, such as replacing them with variants of the GPT, BERT, or Transformers series. At the same time, when using the attention mechanism, self-attention or multi-head attention mechanisms can be tried to further enhance the ability to parse complex instructions. In addition, the algorithm can be adjusted between different computational precisions and efficiencies according to task requirements.
[0078] Step S22 includes:
[0079] Construct a term frequency vector for the key phrases provided by the user ;
[0080] Obtain the key phrases of the preset type tasks from the task library to form a term frequency vector ;
[0081] Calculate based on the cosine similarity formula and the similarity between , and the specific formula is: ,
[0082] where is the calculated cosine similarity score, represents the term frequency vector constructed for the key phrases provided by the user, represents the term frequency vector formed by obtaining the key phrases of the preset type tasks from the task library, represents the th element in the vector represents the th element in the vector represents the and element serial numbers of the vectors, represents the length of the vector;
[0083] If If it is higher than a predetermined threshold, it is considered that the keyword is associated with a certain task type.
[0084] By constructing word frequency vectors for the key phrases provided by the user and the preset phrases in the task library respectively, and using cosine similarity to quantify the similarity between the two, this method effectively evaluates the association degree between the key phrases and the task types, thereby realizing the automatic discrimination of the task types of the user instructions. Cosine similarity provides an efficient and accurate method to evaluate the association degree between the user input and the preset tasks in the task library, so that the user's intention can be quickly recognized. This method has strong adaptability to identify polysemous words and similar phrases, and the calculation process is simple and efficient, suitable for real-time application scenarios. Other similarity calculation methods, such as Euclidean distance and Manhattan distance, can be used to replace the cosine similarity calculation. In addition, it can be considered to replace the word vectors with more complex word embeddings to improve the accuracy and richness of the vocabulary representation.
[0085] Step S24 includes:
[0086] Introduce a dynamic context window mechanism. When parsing the user's intention, adaptively adjust the length and depth of the context information according to the complexity of the current task. Use the sequence-to-sequence model combined with the attention mechanism to deeply parse the user's intention. Through context encoding, combine the user's historical interaction information with the current instruction to form a comprehensive understanding of the user's intention. The user's intention The parsing formula is: ,
[0087] where, is the user's intention, represents mechanism, represents the encoding matrix of the past context information, represents the encoding vector of the current instruction, is a normalization function, represents the transpose of the current instruction encoding, represents the dimension size of the key vector.
[0088] This embodiment uses the dynamic context window mechanism to adaptively adjust the context length and depth in the parsing process according to the task complexity, and uses the sequence-to-sequence model combined with the attention mechanism to deeply parse the user's intention. By combining the historical interaction information with the current instruction, a comprehensive understanding of the user's intention is formed, effectively handling multi-round conversations or complex task instructions. The dynamic context window mechanism enables flexible response to task parsing of different complexities, ensuring the accuracy and efficiency of the parsing. The use of the sequence-to-sequence model combined with the attention mechanism can capture the detailed information in the user's instruction and improve the depth of understanding of the user's intention through the combination of historical interaction information.
[0089] Other types of sequence-to-sequence models can be introduced, such as long short-term memory networks, gated recurrent units, or variants based on the Transformer architecture. Context information encoding can also be replaced by pre-trained language models such as BERT, GPT, etc. The adjustment strategy of the dynamic context window can be further optimized through reinforcement learning algorithms, enabling the model to automatically learn the optimal context length and depth in different task scenarios.
[0090] Step S31 includes:
[0091] Real-time monitor the current environmental parameters through various sensors built into the robot, and use exponential weighted moving average filtering to process the sensor data to remove noise and obtain stable and accurate environmental parameter values. These data serve as the basis for subsequent step judgment and adjustment. The specific formula is:
[0092] ,
[0093] where, represents the estimated value at the current time of, represents the time, represents the measured value of the sensor at the current time, represents the estimated value at the previous time, represents the smoothing factor, and the value range is 0 ≤ ≤ 1.
[0094] This implementation uses the built-in sensors to monitor environmental parameters (such as temperature, light, sound, etc.) in real time. Through exponential weighted moving average filtering, the sensor data is smoothed to reduce the influence of noise, thereby obtaining more stable and accurate environmental parameter values. These parameters serve as the basis for subsequent task judgment and adjustment to ensure the reliability of task execution. Using exponential weighted moving average filtering can effectively remove random noise in the sensor data, and the smoothed data can better reflect the real environmental situation, thus improving the accuracy and stability of subsequent parameter adjustment and task execution. This method is computationally simple and easy to implement in real time. It can be replaced by other filtering methods such as Kalman filtering, mean filtering, or median filtering to process the sensor data, and the most suitable filtering method can be selected according to actual needs. In addition, the value range of the smoothing factor can also be adjusted and dynamically adjusted according to different environmental conditions to balance the requirements of response speed and stability.
[0095] Step S32 includes:
[0096] Set the working range of environmental parameters, and use the nonlinear constraint optimization method to adjust the robot's action parameters. The optimization function takes into account various requirements for task execution. Specifically, the Lagrange multiplier method is adopted by introducing the Lagrange multiplier , integrate the constraint conditions into the objective function, and transform it into an unconstrained optimization problem. The specific formula is: ,
[0097] where, is the Lagrangian function, is the objective function, represents the robot's action parameters, is the Lagrange multiplier, represents the constraint function to ensure that the action parameters are within the working range of the environmental parameters, is the constraint function serial number.
[0098] When performing a task, the robot adjusts its action parameters through the set working range of environmental parameters. This method uses the Lagrange multiplier method to integrate the constraint conditions into the objective function to achieve optimization of various requirements, ensuring that the robot adjusts its action parameters under different environmental conditions to meet the task execution requirements. By transforming the complex constrained problem into an unconstrained optimization problem, this method can adjust the action parameters more flexibly. The Lagrange multiplier method can handle optimization problems under multiple constraint conditions, enabling the robot to automatically adjust its action parameters in a complex environment, thereby improving the efficiency and accuracy of task execution. This method not only ensures the stability of the robot's actions but also can adjust its execution strategy in real time according to the dynamic changes of the environment.
[0099] Other nonlinear constraint optimization methods, such as quadratic programming, penalty function method, or interior point method, can be used to replace the Lagrange multiplier method. According to different task requirements and computing resources, the complexity and accuracy of the optimization algorithm can be adjusted, or different objective functions and constraint conditions can be set according to the specific task environment.
[0100] Step S33 includes an environmental adaptation adjustment factor to adjust the task execution speed , dynamically adjust according to the changes in environmental conditions, and update the adjustment factor in real time through the recursive least squares method , set the standard value , when is lower than , the robot will issue a prompt to suggest changing the task execution strategy or pausing the task. The specific formula is: ,
[0101] where, represents the adjustment factor at the current moment, represents the moment, represents the adjustment factor at the previous moment, represents the gain matrix, represents the current desired output or reference value, represents the current environmental feature vector, represents the predicted value of the environmental feature and the adjustment factor at the previous moment;
[0102] Set a standard value to evaluate the adaptability of the adjustment factor. This standard value is set according to the specific requirements of the task, including to ensure the stability of the task or avoid over-adjustment. When < it indicates that the environment is too complex or not conducive to the current task execution, and the robot will issue a prompt, suggesting to change the execution strategy or pause the task.
[0103] The adjustment factor is updated in real time through the recursive least squares (RLS) method. The robot can adjust the task execution speed according to the changes in environmental conditions. If the adjustment factor is lower than the preset standard value, it indicates that the current environment is not conducive to the continued execution of the task, and a prompt will be issued, suggesting that the user change the task strategy or pause the task. The RLS algorithm enables the update of the adjustment factor to have high precision and response speed. The recursive least squares method provides an efficient real-time update mechanism, enabling the execution speed to be adjusted quickly according to environmental changes, ensuring that the task is carried out under suitable conditions. By setting a standard value to evaluate the adjustment factor, this method can identify adverse environments in a timely manner and provide corresponding suggestions for strategy adjustment, thereby improving the success rate and robustness of the task. Other real-time adjustment algorithms, such as adaptive filtering, gradient descent, or optimization methods based on genetic algorithms, can be used to replace the recursive least squares method. For the evaluation criteria of the adjustment factor, dynamic optimization can also be carried out through experiments or historical data to better adapt to different task requirements and environmental conditions.
[0104] Step S34 includes evaluating the adjusted result using the reinforcement learning method. The robot records the adjustment result in the experience library and continuously optimizes its decision-making model through the reward mechanism. After each task is completed, the effect of the parameter adjustment will be evaluated, and feedback will be provided on whether further adjustment or optimization is required.
[0105] In this embodiment, the parameters adjusted by the robot are evaluated through reinforcement learning. After each task is completed, the adjustment effect is evaluated according to criteria such as task success rate, execution time, energy consumption, and execution accuracy, and the results are recorded in the experience library. The robot optimizes its decision-making model based on the evaluation results, continuously adjusts the strategy through positive and negative rewards, making the adjustment process more efficient and accurate. The reinforcement learning method provides an adaptive optimization mechanism for the robot, which can continuously improve its adjustment strategy according to the execution effect of the task. Through the reward mechanism, the robot can learn the optimal adjustment method, thereby improving the task success rate and execution efficiency. At the same time, the establishment of the experience library also provides important reference information for subsequent tasks. Different reinforcement learning algorithms, such as deep Q-network, policy gradient method, or proximal policy optimization, can be selected to replace the current reinforcement learning method. In addition, different reward mechanisms can be designed according to the requirements of specific tasks, or a multi-objective optimization strategy can be adopted to enable the robot to obtain better execution effects under various conditions.
[0106] Adjust using the reinforcement learning update formula the value table, and according to the updated value table, continuously optimize the action strategy of the robot in different states, and select the action that can maximize the value;
[0107] The reward mechanism includes giving a higher positive reward when the task is completed, giving a negative reward when an invalid action is performed or the task fails, and obtaining an additional reward by saving resources or improving efficiency during the task execution process;
[0108] After each task is completed, the robot will evaluate the effect of the current parameter adjustment. The evaluation criteria include task success rate, execution time, energy consumption, and execution accuracy. The robot will feedback the evaluation results to the experience library for further adjustment and optimization the value table and the strategy;
[0109] If the evaluation results show that there are deficiencies in the current adjustment strategy, the robot will adjust the learning parameters or modify the reward mechanism to further optimize its strategy;
[0110] The robot also recommends adjusting the task execution strategy according to the evaluation results, including recommending pausing the task or reducing the task complexity when the conditions are unfavorable.
[0111] In this embodiment, the robot continuously updates its value table through a reinforcement learning algorithm to optimize its action strategies in different states. A reward mechanism is adopted, where a positive reward is given when the task is completed, a negative reward is given when an invalid action is executed or the task fails, and additional rewards are obtained by saving resources or improving efficiency during the task execution. After each task is completed, the robot evaluates the effect of parameter adjustment and feeds the results back to the experience library. If the evaluation shows that the current strategy has deficiencies, the robot will adjust the learning parameters or modify the reward mechanism to further optimize the strategy. In addition, the robot can provide strategy suggestions based on the evaluation results, such as pausing the task or reducing the complexity under adverse conditions. By reinforcement learning and dynamically adjusting the value table, the robot can adaptively optimize its execution strategy, improve the task success rate and efficiency. The multi-dimensional reward mechanism ensures that while pursuing the task completion rate, attention is also paid to the effective utilization of resources and the optimization of energy consumption. This adaptive optimization ability enables the robot to perform more stably and efficiently in a changing environment. In addition, the ability to provide real-time adjustment suggestions during task execution further enhances the flexibility and safety of the task.
[0112] Different reinforcement learning strategies can be adopted, such as deep reinforcement learning, trust region policy optimization, etc., to further enhance the optimization ability. The reward mechanism can be customized according to different task requirements, such as introducing multi-objective optimization strategies or using hierarchical reward mechanisms. In addition, for value table updates, an approximate model can be constructed in combination with neural networks to handle more complex state spaces and decision-making processes.
[0113] This method enables the robot to execute tasks in a complex and changing environment through multi-step information parsing and environmental parameter adjustment. This method mainly involves multiple links such as information reception, parsing, execution, and feedback, ensuring that the robot can accurately identify the user's intention and efficiently execute tasks, including:
[0114] Information reception and preprocessing. In this process, the robot first captures the user's voice signal through a microphone array, converts it into a digital audio file, and uses natural language processing technology to extract the task instructions in text form. Through the separation of multimedia information and activation word recognition, the robot can initiate the information reception process to ensure accurate acquisition of the user's instructions.
[0115] Semantic parsing and intent recognition. The robot uses a context understanding module to deeply analyze the received information. This process includes extracting the key phrases and words expressed by the user, analyzing the relevance of these keywords to the task type, and parsing the user's intent through a context understanding algorithm. Specifically, deep learning models are adopted for semantic parsing, and through pre-trained language models and attention mechanisms, the understanding of complex task instructions and the extraction of key task words are realized. In addition, the dynamic context window mechanism can adaptively adjust the parsing length and depth according to the task complexity, so as to enhance the comprehensive understanding of the user's intent.
[0116] Environmental parameter monitoring and adjustment. During the task execution process, the robot real-time monitors environmental parameters (such as temperature, light, sound, etc.), and uses exponential weighted moving average filtering to process the sensor data to remove noise and ensure the stability and accuracy of the data. If it is detected that the environmental parameters exceed the predetermined working range, the robot will use non-linear constraint optimization methods (such as the Lagrange multiplier method) to adjust its action parameters to ensure that the task is executed under safe and efficient conditions. Through the environmental adaptation adjustment factor, the robot dynamically adjusts the task execution speed according to environmental changes and issues strategy adjustment prompts when necessary to improve the flexibility and robustness of task execution.
[0117] Parameter adjustment and reinforcement learning. In the subsequent process of task execution, the robot continuously evaluates the results after environmental adaptation adjustment and uses reinforcement learning algorithms to optimize its decision-making model. After each task is completed, the robot feeds back indicators such as task success rate, execution time, energy consumption, and execution accuracy to the experience library and updates its value table based on these data. Through the reward mechanism, the execution strategy can be gradually optimized in multiple rounds of adjustments to achieve adaptive learning and performance improvement. Reinforcement learning also allows the robot to provide adjustment suggestions for task execution according to the current evaluation results, such as pausing the task or reducing the task complexity to cope with adverse environments.
[0118] Overall technical advantages and innovations. The core advantage of the present invention lies in its highly intelligent and flexible interaction capabilities. Through multi-level parsing technologies and deep learning models, the robot can accurately identify and understand complex instructions from users. The real-time monitoring and adaptive adjustment mechanism of environmental parameters ensure that the robot can still execute tasks efficiently and safely in a changing environment. The introduction of reinforcement learning enables the robot to continuously optimize its behavior strategy and improve the efficiency and success rate of task execution. These technical features endow the robot with strong task adaptation capabilities and operational flexibility, providing solid technical support for intelligent interaction in practical applications.
[0119] The method supports various alternative or modified implementation manners to enhance its adaptability in different application scenarios. For example, the audio processing module can select different microphone layouts or noise reduction algorithms; semantic parsing can be based on different deep learning models; environmental monitoring can introduce more sensors or more complex filtering methods; the reinforcement learning algorithm can select different strategy optimization models according to requirements. These flexible replacement ways demonstrate the wide applicability and technological innovation of the present invention.
[0120] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made in these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An information interaction method for an embodied intelligent robot, characterized in that, The method includes: S1: Receive the information containing the task instruction sent by the user; S2: Parse the information based on the context understanding module to identify the specific meaning of the task instruction and the environmental conditions; S3: Execute the operation corresponding to the specific meaning and adjust the execution parameters in a specific environment to ensure the accurate implementation of the task instruction; Wherein, the step S2 includes: S21: Extract the key phrases and words expressed by the user in the information as keywords; S22: Analyze the correlation degree between the keywords and the task type; S23: Identify the environmental conditions based on the currently perceived environmental features; S24: Use the context understanding algorithm to parse out the user's intention; The step S3 includes: S31: Real-time monitor the current environmental parameters, including temperature, light intensity, sound level; S32: If the current environmental parameters exceed the expected working range, adjust the robot's action parameters; S33: Adjust the task execution speed based on the environmental adaptation adjustment factor Adjust the task execution speed , if is less than the standard value then prompt that a strategy change is needed; S34: Update the parameter adjustment result for reference use and feedback whether the adjustment is necessary; The step S33 includes: Environmental adaptation adjustment factor Adjust the task execution speed , dynamically adjust according to changes in environmental conditions, and update the adjustment factor in real time through recursive least squares , set the standard value , when is lower than , the robot will issue a prompt, suggesting to change the task execution strategy or pause the task. The specific formula is: , Among them, represents the adjustment factor at the current moment, represents the moment, represents the adjustment factor at the previous moment, represents the gain matrix, represents the current desired output or reference value, represents the current environmental feature vector, represents the predicted value of the environmental feature and the adjustment factor at the previous moment; Set a standard value to evaluate the adaptability of the adjustment factor. This standard value is set according to the specific requirements of the task, including ensuring the stability of the task or avoiding excessive adjustment. When < it indicates that the environment is too complex or not conducive to the execution of the current task, and the robot will issue a prompt, suggesting changing the execution strategy or pausing the task.
2. The information interaction method of an embodied intelligent robot according to claim 1, characterized in that: The step S1 includes: S11: Capture the user's voice signal through the microphone array; S12: Convert the captured voice signal into a digital audio file; S13: Use natural language processing technology to extract text from the converted digital audio file to obtain the task instruction in text form; S14: Separate the voice information from the received multimedia information and convert it into text; S15: Perform voice command activation word recognition to start the information receiving process.
3. The information interaction method of an embodied intelligent robot according to claim 2, wherein: The step S13 includes: Using a deep learning model, semantic analysis is performed on the text. Multilevel embedding vectors of a pre-trained language model are used, and the attention mechanism is employed to understand the context of complex task instructions and extract key task words. The specific formula is as follows: , Among them, represents a mechanism for improving the model's focusing ability on input information, represents the query matrix, represents the key matrix, represents the value matrix, is a normalization function, represents the dimension size of the key vector, represents the key matrix transpose.
4. The information interaction method of an embodied intelligent robot according to claim 1, wherein: The step S22 includes: Construct a term frequency vector for the key phrases provided by the user ; Obtain the key phrase composition word frequency vector of the preset type task according to the task library ; Calculate based on the cosine similarity formula and the similarity between , and the specific formula is: , Among them, is the calculated cosine similarity score, represents constructing a term frequency vector for the key phrases provided by the user, represents constructing a term frequency vector from the key phrases of the preset type tasks obtained from the task library, represents the vector the th element in, represents the vector the th element in, represents the vector and the sequence numbers of the elements, represents the length of the vector; If is higher than a predetermined threshold, the keyword is considered to be associated with a certain task type.
5. The information interaction method of an embodied intelligent robot according to claim 4, wherein: The step S24 includes: Introduce a dynamic context window mechanism. When parsing the user's intention, adaptively adjust the length and depth of the context information according to the complexity of the current task. Utilize the sequence-to-sequence model combined with the attention mechanism to deeply analyze the user's intention. Combine the user's historical interaction information with the current instruction through context encoding to form a comprehensive understanding of the user's intention, user intention The parsing formula is as follows: , Among them, is the user's intention, denotes mechanism, the encoding matrix representing past context information, the encoding vector representing the current instruction, is a normalization function, represents the transpose of the current instruction encoding, represents the dimensionality size of the key vector.
6. The information interaction method of an embodied intelligent robot according to claim 1, characterized in that: The step S31 includes: Real-time monitor the current environmental parameters through a variety of sensors built in the robot, use exponential weighted moving average filtering to process the sensor data, remove noise, and obtain stable and accurate environmental parameter values. Use these data as the basis for judgment and adjustment in the subsequent steps. The specific formula is: , Among them, represents the estimated value at the current moment , represents the moment represents the measured value of the sensor at the current moment, represents the estimated value at the previous moment, represents the smoothing factor, and the value range is 0 ≤ ≤ 1.
7. The information interaction method of an embodied intelligent robot according to claim 1, characterized in that: The step S32 includes: Set the working range of environmental parameters, and use the nonlinear constraint optimization method to adjust the robot's action parameters. The optimization function takes into account various requirements for task execution. Specifically, the Lagrange multiplier method is adopted, and by introducing the Lagrange multiplier , the constraint conditions are integrated into the objective function, which is transformed into an unconstrained optimization problem. The specific formula is: , Among them, is the Lagrangian function, is the objective function, represents the robot's action parameters, is the Lagrange multiplier, represents the constraint function to ensure that the action parameters are within the working range of the environmental parameters, is the constraint function serial number.
8. The information interaction method of an embodied intelligent robot according to claim 1, wherein: The step S34 includes: The adjusted results are evaluated using the reinforcement learning method, and the robot records the adjusted results in the experience library to form a value table, and continuously optimizes its decision-making model through the reward mechanism. After each task is completed, the effect of parameter adjustment will be evaluated, and feedback will be provided on whether further adjustment or optimization is required.
9. The information interaction method of an embodied intelligent robot according to claim 8, characterized in that: Adjust using the update formula of the reinforcement learning method the value table, and according to the updated value table, continuously optimize the action strategy of the robot in different states, and select the action that can maximize the value; The reward mechanism includes giving a higher positive reward when the task is completed, giving a negative reward when an invalid action is executed or the task fails, and obtaining an additional reward by saving resources or improving efficiency during the task execution process; After each task is completed, the robot will evaluate the effect of the current parameter adjustment. The evaluation criteria include task success rate, execution time, energy consumption, and execution accuracy. The robot will feedback the evaluation results to the experience library for further adjustment and optimization. Value tables and strategies; If the evaluation result shows that there are deficiencies in the current adjustment strategy, the robot will adjust the learning parameters or modify the reward mechanism to further optimize its strategy; The robot also recommends adjusting the task execution strategy according to the evaluation result, including recommending pausing the task or reducing the task complexity when the conditions are unfavorable.
Citation Information
Patent Citations
Task-oriented unstructured information intelligent question and answer system construction method
CN109800284A
Human-computer interaction system based on natural language
CN116402057A