Robot control method, computer equipment and computer readable storage medium
Through multimodal data processing and machine learning algorithms, response strategies are generated to control robot motion, solving the problems of single mode of human-computer interaction, data dependence and lack of independent decision-making in the prior art, and achieving more efficient and accurate human-computer interaction.
Patent Information
- Application Number
- CN202510513912.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-04-22
AI Technical Summary
In the prior art, human-computer interaction has the limitations of the single interaction mode, data dependence, and lack of independent decision-making capabilities, resulting in insufficient efficiency and accuracy of interaction.
By obtaining the user's multimodal data, including attitude data, sound data and force control data, feature extraction and fusion processing are performed to generate multimodal fusion features. Combining environmental information and historical interactive data, context information is determined, and a response strategy is generated by imitation learning and reinforcement learning algorithms to control robot motion.
It realizes a more comprehensive and accurate perception of users' interaction intentions and environmental status, and improves the accuracy of response strategies and the naturalness, accuracy and intelligence of human-computer interaction.
Smart Images

Figure CN120095830A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of robotics technology, and in particular to a robot control method, a computer device, and a computer-readable storage medium. Background Art
[0002] With the continuous expansion of robot application fields, the demand for interaction between robots and humans is increasing, and the requirements for human-computer interaction technology are also getting higher and higher.
[0003] In the existing technology, the interaction between robots and humans mostly relies on traditional computer interfaces based on keyboards and mice and traditional human-computer interfaces such as touch screens, which lack naturalness and adaptability, making it difficult to achieve smooth communication and highly personalized responses. In addition, most existing human-computer interaction systems are based on single sensory input, which limits the system's perception ability and the naturalness of the interaction. There are also preliminary multimodal systems that attempt to combine vision and voice systems, but the lack of deep data fusion and advanced decision-making mechanisms leads to insufficient efficiency and accuracy of the interaction.
[0004] Therefore, the human-computer interaction in the existing technology has the disadvantages of single interaction mode limitation, data dependence, and lack of autonomous decision-making ability. Summary of the invention
[0005] The purpose of this application is to provide a robot control method, a computer device and a computer-readable storage medium to address the shortcomings of the above-mentioned prior art, so as to solve the practical problems of the limitations of a single interaction mode, data dependence and lack of autonomous decision-making ability in human-computer interaction in the prior art.
[0006] To achieve the above purpose, the technical solution adopted in the embodiment of the present application is as follows:
[0007] In a first aspect, an embodiment of the present application provides a robot control method, the method comprising:
[0008] Acquiring multimodal data of the user, wherein the multimodal data includes posture data, sound data, and force control data;
[0009] Performing feature extraction and fusion processing on the multimodal data to obtain multimodal fusion features;
[0010] Determining context information of human-machine interaction based on environmental information of the area where the robot is located and historical interaction data of the robot;
[0011] A response strategy is generated according to the multimodal fusion features and the context information, and the robot movement is controlled according to the response strategy.
[0012] As an optional implementation, after controlling the movement of the robot according to the response strategy, the method further includes:
[0013] Acquiring feedback data input by a user based on motion information of the robot;
[0014] The response strategy is modified according to the feedback data.
[0015] As an optional implementation, the acquiring of multimodal data of the user includes:
[0016] Acquiring user's posture data through a visual sensor, wherein the posture data includes the user's facial expression, hand gestures and body posture;
[0017] Acquiring the user's voice data through a voice sensor, wherein the voice data includes the user's voice commands and non-verbal sounds;
[0018] The force control data of the user is acquired through the force sensor, and the force control data includes the touch interaction signal and the push-pull interaction signal of the user.
[0019] As an optional implementation, the extracting and fusing the multimodal data to obtain multimodal fusion features includes:
[0020] Extracting posture features according to the posture data by a visual encoder in the neural network model;
[0021] The sound encoder in the neural network model extracts sound features according to the sound data;
[0022] Extracting force control features according to the force control data by a force encoder in the neural network model;
[0023] The decoder in the neural network model obtains the multimodal fusion feature according to the posture feature, the sound feature and the force control feature.
[0024] As an optional implementation, obtaining the multimodal fusion feature according to the posture feature, the sound feature and the force control feature includes:
[0025] Determining, according to the posture feature, the sound feature, and the force control feature, a first attention weight of the posture feature, a second attention weight of the sound feature, and a third attention weight of the force control feature;
[0026] The multimodal fusion feature is obtained according to the posture feature, the sound feature, the force control feature, the first attention weight, the second attention weight and the third attention weight.
[0027] As an optional implementation manner, generating a response strategy according to the multimodal fusion feature and the context information includes:
[0028] According to the multimodal fusion features, an initial response strategy is generated using an imitation learning algorithm;
[0029] According to the multimodal fusion features and the context information, a reinforcement learning algorithm is used to determine the reward signal corresponding to the initial response strategy, and the initial response strategy is iteratively corrected according to the reward signal, and the initial response strategy at the end of the iteration is used as the response strategy.
[0030] As an optional implementation, controlling the robot movement according to the response strategy includes:
[0031] Determine the robot's action to be performed according to the response strategy, and generate voice data corresponding to the action to be performed;
[0032] A control instruction is generated according to the action to be performed and the voice data, and the robot is controlled according to the control instruction to perform the action to be performed while outputting the voice data.
[0033] As an optional implementation manner, the generating voice data corresponding to the action to be performed includes:
[0034] According to the action to be performed, generating a natural language response matching the action to be performed;
[0035] According to the natural language response, speech data corresponding to the action to be performed is generated through speech synthesis.
[0036] In a second aspect, an embodiment of the present application provides a robot control device, the device comprising:
[0037] An acquisition module, used to acquire multimodal data of the user, wherein the multimodal data includes posture data, sound data and force control data;
[0038] A processing module, used for performing feature extraction and fusion processing on the multimodal data to obtain multimodal fusion features;
[0039] A determination module, used to determine context information of human-machine interaction based on environmental information of the area where the robot is located and historical interaction data of the robot;
[0040] A control module is used to generate a response strategy according to the multimodal fusion features and the context information, and control the movement of the robot according to the response strategy.
[0041] As an optional implementation, the device further includes: a correction module; the correction module is used to:
[0042] Acquiring feedback data input by a user based on motion information of the robot;
[0043] The response strategy is modified according to the feedback data.
[0044] As an optional implementation, the acquisition module is specifically used for:
[0045] Acquiring user's posture data through a visual sensor, wherein the posture data includes the user's facial expression, hand gestures and body posture;
[0046] Acquiring the user's voice data through a voice sensor, wherein the voice data includes the user's voice commands and non-verbal sounds;
[0047] The force control data of the user is acquired through the force sensor, and the force control data includes the touch interaction signal and the push-pull interaction signal of the user.
[0048] As an optional implementation, the processing module is specifically used for:
[0049] Extracting posture features according to the posture data by a visual encoder in the neural network model;
[0050] The sound encoder in the neural network model extracts sound features according to the sound data;
[0051] Extracting force control features according to the force control data by a force encoder in the neural network model;
[0052] The decoder in the neural network model obtains the multimodal fusion feature according to the posture feature, the sound feature and the force control feature.
[0053] As an optional implementation, the processing module is specifically used for:
[0054] Determining, according to the posture feature, the sound feature, and the force control feature, a first attention weight of the posture feature, a second attention weight of the sound feature, and a third attention weight of the force control feature;
[0055] The multimodal fusion feature is obtained according to the posture feature, the sound feature, the force control feature, the first attention weight, the second attention weight and the third attention weight.
[0056] As an optional implementation, the control module is specifically used for:
[0057] According to the multimodal fusion features, an initial response strategy is generated using an imitation learning algorithm;
[0058] According to the multimodal fusion features and the context information, a reinforcement learning algorithm is used to determine the reward signal corresponding to the initial response strategy, and the initial response strategy is iteratively corrected according to the reward signal, and the initial response strategy at the end of the iteration is used as the response strategy.
[0059] As an optional implementation, the control module is specifically used for:
[0060] Determine the robot's action to be performed according to the response strategy, and generate voice data corresponding to the action to be performed;
[0061] A control instruction is generated according to the action to be performed and the voice data, and the robot is controlled according to the control instruction to perform the action to be performed while outputting the voice data.
[0062] As an optional implementation, the control module is specifically used for:
[0063] According to the action to be performed, generating a natural language response matching the action to be performed;
[0064] According to the natural language response, speech data corresponding to the action to be performed is generated through speech synthesis.
[0065] In a third aspect, an embodiment of the present application provides a computer device, comprising: a processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the computer device is running, the processor communicates with the memory via the bus, and the processor executes the machine-readable instructions to perform the steps of the robot control method described in the first aspect above.
[0066] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the robot control method described in the first aspect above are executed.
[0067] The beneficial effects of this application are:
[0068] The present application provides a robot control method, a computer device and a computer-readable storage medium, which obtain multimodal data during the interaction between the user and the robot, including posture data collected under the visual sense, sound data collected under the auditory sense and force control data collected under the tactile sense during the interaction between the user and the robot. Feature extraction is performed on the multimodal data to obtain key features under the multi-sensory senses, and feature fusion is performed on the extracted key features under the multi-sensory senses to obtain fused multimodal fusion features. The multimodal fusion feature integrates the key features under the multi-sensory senses of vision, hearing and touch, and can fully perceive the user's interaction intention. The environmental information of the area where the robot is located and the historical interaction data between the user and the robot are obtained, and the context information of the current human-computer interaction is obtained by combining various external factors related to the current human-computer interaction scene in the environmental information and various historical interaction behaviors in the historical interaction data. Based on the multimodal fusion features and context information, a response strategy is generated, and the robot movement is controlled according to the response strategy. By combining the multi-sensory multimodal fusion features and the current contextual information of human-computer interaction, we can perceive the user's interaction intentions and the current state of the interaction environment more comprehensively and accurately, thereby improving the accuracy of the response strategy and the naturalness, accuracy and intelligence of human-computer interaction. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0070] Figure 1 A schematic diagram of a flow chart of a robot control method provided in an embodiment of the present application;
[0071] Figure 2 Another schematic diagram of a flow chart of a robot control method provided in an embodiment of the present application;
[0072] Figure 3 A schematic diagram of a process of obtaining multimodal data of a user in a robot control method provided in an embodiment of the present application;
[0073] Figure 4 A schematic diagram of a flow chart of obtaining multimodal fusion features of a robot control method provided in an embodiment of the present application;
[0074] Figure 5 Another schematic diagram of a process of obtaining multimodal fusion features of a robot control method provided in an embodiment of the present application;
[0075] Figure 6A schematic diagram of a flow chart of a response strategy generation method for a robot control method provided in an embodiment of the present application;
[0076] Figure 7 A schematic diagram of a flow chart of controlling robot motion of a robot control method provided in an embodiment of the present application;
[0077] Figure 8 A schematic diagram of a flow chart of generating voice data for a robot control method provided in an embodiment of the present application;
[0078] Fig. 9 A module structure diagram of a robot control device provided in an embodiment of the present application;
[0079] Fig.10 A schematic diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0080] To make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. It should be understood that the drawings in the present application only serve the purpose of explanation and description and are not used to limit the scope of protection of the present application. In addition, it should be understood that the schematic drawings are not drawn in real proportion. The flowchart used in this application shows the operations implemented according to some embodiments of the present application. It should be understood that the operations of the flowchart can be implemented out of sequence, and the steps without logical context can be reversed in order or implemented simultaneously. In addition, those skilled in the art can add one or more other operations to the flowchart under the guidance of the content of the present application, or remove one or more operations from the flowchart.
[0081] In addition, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application described and shown in the drawings here can be arranged and designed in various configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application claimed for protection, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work belong to the scope of protection of the present application.
[0082] It should be noted that the term "comprising" will be used in the embodiments of the present application to indicate the existence of the features declared thereafter, but does not exclude the addition of other features.
[0083] In the field of robotics, the demand for interaction between robots and humans is growing, and the requirements for human-computer interaction technology are getting higher and higher. In the existing technology, human-computer interaction has the limitations of a single interaction mode, data dependence, and lack of autonomous decision-making ability, resulting in low efficiency and accuracy of human-computer interaction.
[0084] Based on the above problems, the embodiment of the present application proposes a robot control method, which comprehensively captures and extracts the multimodal fusion features of the user's body posture, voice and tactile force control by integrating multiple sensory inputs such as vision, hearing and touch. This multimodal fusion feature enables a more comprehensive understanding of the user's intentions and environmental status during human-computer interaction, thereby providing a richer and more natural interactive experience. The use of reinforcement learning algorithms allows the robot to learn autonomously during the interaction with the environment, reducing dependence on large amounts of labeled data. Through iterative feedback correction, the response strategy can be automatically adjusted and optimized, thereby improving the generalization ability, adaptability and accuracy of robot control.
[0085] Figure 1 The flowchart of the robot control method provided in the embodiment of the present application is shown in FIG. 1 , and the execution subject of the method can be any computer device with computing processing capabilities. Figure 1 As shown, the method includes:
[0086] S101. Acquire multimodal data of a user, where the multimodal data includes posture data, sound data, and force control data.
[0087] Optionally, multimodal data during the interaction between the user and the robot is collected from multiple senses, including posture data, sound data, and force control data during the interaction between the user and the robot. The posture data may be interaction data collected under visual senses, the sound data may be interaction data collected under auditory senses, and the force control data may be interaction data collected under tactile senses.
[0088] Specifically, gesture data can reflect the gesture information of various parts of the user's body, such as the gesture information of the face and hands. Voice data can reflect the user's language expression and emotional state, and force control data can reflect the interaction information between the user and the object or environment.
[0089] The posture data, sound data and force control data in the collected multimodal data are pre-processed by data cleaning and standardization, where data cleaning includes filtering and denoising the multimodal data. The cleaned multimodal data is normalized to unify the format of the posture data, the format of the sound data and the format of the force control data, and standardized posture data, sound data and force control data are obtained to improve the credibility and accuracy of the multimodal data.
[0090] S102: extract and fuse the multimodal data to obtain multimodal fusion features.
[0091] Optionally, through feature extraction, visual key features, auditory key features and tactile key features are respectively extracted from the preprocessed multimodal data, and feature fusion is performed on the visual key features, auditory key features and tactile key features to obtain fused multimodal fusion features.
[0092] By first extracting features from the preprocessed multimodal data and then performing feature fusion, the information of multiple sensory inputs in the multimodal data is effectively integrated, and more representative and informative multimodal fusion features are obtained. The multimodal fusion features integrate the key features of multiple senses of vision, hearing, and touch, so as to fully perceive the user's interaction intentions and provide data support for the naturalness and accuracy of subsequent interactions with the robot.
[0093] S103: Determine context information of human-machine interaction according to environmental information of the area where the robot is located and historical interaction data of the robot.
[0094] Optionally, the environmental information of the area where the robot is located represents various external factors related to the current human-computer interaction scene to assist in identifying the user's interaction intention. The environmental information may include physical environmental information and social environmental information. The physical environmental information may include the time, location, and status of surrounding objects and devices of the current interaction. The social environmental information may include the interaction occasion and the number of people present at the interaction occasion.
[0095] The historical interaction data between users and robots records each historical interaction behavior. By analyzing the historical interaction data, users’ interaction preferences, habits, and behavior patterns can be determined. For example, users used to focus more on interacting with robots using voice and consulting on the latest developments in the technology field.
[0096] Obtain environmental information of the area where the robot is located and historical interaction data between the user and the robot, combine various external factors related to the current human-computer interaction scenario and various historical interaction behaviors, understand the context of the current human-computer interaction, and obtain the context information of the current human-computer interaction to improve the accuracy of identifying user interaction intentions, enhance human-computer interaction effects and user satisfaction.
[0097] S104: Generate a response strategy based on the multimodal fusion features and context information, and control the robot movement according to the response strategy.
[0098] Optionally, the multimodal fusion feature contains comprehensive information after the posture data, sound data and force control data are extracted and fused, integrating the key features of vision, hearing and touch. The context information of the current human-computer interaction combines the environmental information and historical interaction data to characterize various external factors related to the current human-computer interaction scene and various historical interaction behaviors.
[0099] The multimodal fusion features are used as the main data, and the contextual information of the current human-computer interaction is used as the auxiliary data. They are input into the pre-trained decision model, and the decision model predicts and outputs the response strategy based on the input multimodal fusion features and contextual information. The robot's motion trajectory is planned according to the response strategy, and the movement of each joint of the robot is controlled based on the motion trajectory. Based on the multimodal fusion features and contextual information, a more reasonable response strategy is generated and the robot's motion is accurately controlled, achieving a more intelligent and natural human-computer interaction.
[0100] In this embodiment, multimodal data is obtained during the interaction between the user and the robot, including posture data collected under the visual sense, sound data collected under the auditory sense, and force control data collected under the tactile sense. Feature extraction is performed on the multimodal data to obtain key features under the multi-sensory senses, and feature fusion is performed on the extracted key features under the multi-sensory senses to obtain the fused multimodal fusion features. The multimodal fusion features integrate the key features under the multi-sensory senses of vision, hearing and touch, and can fully perceive the user's interaction intention. The environmental information of the area where the robot is located and the historical interaction data between the user and the robot are obtained, and the context information of the current human-computer interaction is obtained through context understanding by combining various external factors related to the current human-computer interaction scene in the environmental information and various historical interaction behaviors in the historical interaction data. Based on the multimodal fusion features and the context information, a response strategy is generated, and the robot movement is controlled according to the response strategy. By combining the multimodal fusion features under the multi-sensory senses and the context information of the current human-computer interaction, the user's interaction intention and the current interactive environment state can be perceived more comprehensively and accurately, thereby improving the accuracy of the response strategy and the naturalness, accuracy and intelligence of the human-computer interaction.
[0101] Figure 2 Another flowchart of the robot control method provided in the embodiment of the present application is as follows: Figure 2 As shown, the steps after controlling the robot movement according to the response strategy in the above step S104 further include:
[0102] S201. Obtain feedback data input by the user based on the motion information of the robot.
[0103] Optionally, after the current human-computer interaction ends, the feedback data input by the user based on the robot's motion information can be obtained through voice interaction, interface interaction, visual monitoring or wearable devices. Specifically, the user's voice feedback is captured in real time as feedback information through the robot's built-in microphone. Alternatively, feedback information is input by entering text on the evaluation interface and selecting evaluation options (such as star ratings or satisfaction ratings, etc.). Alternatively, the robot's own camera or the camera of the surrounding interactive environment is used for visual monitoring to analyze the user's expression and body language to obtain the user's positive feedback information or negative feedback information. Alternatively, the user can also wear a specific wearable device, such as a smart bracelet, to monitor the user's heart rate and other physiological feedback information.
[0104] For example, the feedback data may be the user's evaluation information on the robot's motion performance in the form of language, text, etc. Alternatively, the feedback information may be obtained by observing the user's behavior. For example, when the robot is performing an action, if the user evades, frowns, shakes his head, etc., it may indicate that the user is dissatisfied with or uncomfortable with the robot's motion performance. On the contrary, if the user nods, smiles, actively approaches, etc., it may indicate that the user approves of the robot's motion performance.
[0105] S202. Modify the response strategy according to the feedback data.
[0106] Optionally, the link to be corrected is determined based on the feedback data input by the user. For example, if the user feedbacks that the robot's action when picking up an object is inaccurate, the action link of picking up the object in the response strategy needs to be taken as the link to be corrected.
[0107] According to the link to be corrected and the feedback data input by the user, the feedback data is integrated into the decision model, and the model parameters in the decision model are updated through phased incremental learning using an online learning algorithm. The decision model corrects the response strategy and outputs the corrected response strategy based on the input multimodal fusion features, context information, and feedback data. By dynamically adjusting the response strategy, the robot's movement can be dynamically adjusted.
[0108] In this embodiment, after the current human-computer interaction ends, feedback data input by the user based on the robot's motion information is obtained through voice interaction, interface interaction, visual monitoring or wearable devices. According to the feedback data input by the user, the link to be corrected is determined, and based on the link to be corrected and the feedback data input by the user, an online learning algorithm is used to update the model parameters in the decision model through phased incremental learning to correct the response strategy. The response strategy is dynamically adjusted based on the feedback data in the interaction process, and then the robot's motion is dynamically adjusted to achieve a more smooth, natural and accurate interaction experience.
[0109] The following is a detailed description of the process of obtaining multimodal data of a user.
[0110] Figure 3 A schematic diagram of a process of obtaining multimodal data of a user in a robot control method provided in an embodiment of the present application, such as Figure 3 As shown, the step of obtaining the multimodal data of the user in the above step S101 includes:
[0111] S301. Acquire user's posture data through a visual sensor, where the posture data includes the user's facial expression, hand gestures, and body posture.
[0112] Optionally, at least one visual sensor is deployed in the human-machine interaction environment, and the user's posture data is obtained through the visual sensor. Specifically, the visual sensor can be a camera, which captures the user's facial expressions, gestures and body postures when interacting with the robot by collecting video streams or images of the user interacting with the robot in real time, and obtains the user's posture data, providing a visual sensory data source for human-machine interaction.
[0113] For example, the visual sensor can be deployed in front of or on the side of the user to clearly and comprehensively capture the user's gesture data. For example, the visual sensor captures a video stream or image of the user pointing to a water cup on the table as the user's gesture data.
[0114] S302: Acquire the user's voice data through a voice sensor, where the voice data includes the user's voice commands and non-verbal sounds.
[0115] Optionally, at least one sound sensor is deployed in the human-computer interaction environment to obtain the user's sound data. Specifically, the sound sensor can be a microphone, which can capture the voice commands and non-verbal sounds of the user interacting with the robot in real time, obtain the user's sound data, and provide a data source for auditory senses for human-computer interaction. Among them, the non-verbal sound can be the sound accompanied by the change of the user's body posture, such as clapping, stomping, etc.
[0116] For example, the sound sensor can be deployed at a location close to the user so as to clearly and effectively capture the user's sound data. For example, the sound sensor captures the user's voice command "help me get a water cup from the table" and the non-verbal sound of the user's arm pointing to the water cup on the table, which is the friction sound of the clothes, as the user's sound data.
[0117] S303 . Obtain force control data of the user through the force sensor, where the force control data includes a touch interaction signal and a push-pull interaction signal of the user.
[0118] Optionally, at least two force sensors are deployed in the human-computer interaction environment, and the force control data of the user is obtained through the force sensors. Specifically, the force sensors may include piezoresistive force sensors and strain gauge force sensors. The piezoresistive force sensors are used to collect touch interaction signals when the user interacts with the robot, and the strain gauge force sensors are used to collect push-pull interaction signals when the user interacts with the robot, to obtain the user's force control data, and to provide a data source for tactile sensations for human-computer interaction.
[0119] For example, the force sensor can be deployed on the object touched by the user to accurately capture the user's force control data. For example, when the user takes a water cup on the table, the force control sensor installed on the water cup detects the force of the user's touching, pressing, pushing, pulling or holding the water cup, and obtains the user's touch interaction signal and push-pull interaction signal as the user's voice data.
[0120] In this embodiment, visual sensors, sound sensors and force sensors are deployed in the human-computer interaction environment. The visual sensors are used to capture the facial expressions, gestures and body postures of the user when interacting with the robot, and the posture data of the user is obtained, which provides a data source for the visual sense in human-computer interaction. The sound sensors are used to capture the voice commands and non-verbal sounds of the user when interacting with the robot, and the voice data of the user is obtained, which provides a data source for the auditory sense in human-computer interaction. The force sensors are used to capture the touch interaction signals and push-pull interaction signals of the user when interacting with the robot, and the force control data of the user is obtained, which provides a data source for the tactile sense in human-computer interaction. By obtaining the data sources of vision, hearing and touch during the human-computer interaction process, the user's interaction intention can be fully and accurately perceived.
[0121] The following describes in detail the process of extracting and fusing multimodal data to obtain multimodal fusion features.
[0122] Figure 4 A schematic diagram of a process for obtaining multimodal fusion features of a robot control method provided in an embodiment of the present application is shown in FIG. Figure 4 As shown, the step of extracting features and fusing multimodal data to obtain multimodal fusion features in the above step S102 includes:
[0123] S401, extracting posture features based on posture data by the visual encoder in the neural network model.
[0124] Optionally, a visual encoder is designed in advance for the posture data of the visual sensor through a deep neural network model, and the posture features of the visual modality are extracted through the visual encoder. The visual encoder can be a convolutional neural network (CNN), which automatically learns the local features of each image or video frame in the posture data and gradually extracts the posture features through a hierarchical structure.
[0125] Specifically, the visual encoder performs facial expression recognition on the posture data, and extracts the user's facial expression features by analyzing the movement of the user's facial muscles during human-computer interaction. The visual encoder performs gesture recognition on the posture data, and extracts the movement trajectory of the user's gesture during human-computer interaction, including gesture features such as gesture shape, speed, acceleration, and gesture duration. The visual encoder performs body posture estimation on the posture data, and extracts the user's body posture features, such as standing, sitting, walking, etc., by identifying the position and angle of each body joint of the user during human-computer interaction. Facial expression features, gesture features, and body posture features are used as posture features under the visual modality.
[0126] S402, the sound encoder in the neural network model extracts sound features based on the sound data.
[0127] Optionally, a sound encoder is designed in advance for the sound data of the sound sensor through a deep neural network model, and the sound features of the auditory modality are extracted through the sound encoder. The sound encoder can be a recurrent neural network (RNN) to process the time series information in the audio signal in the sound data and extract the sound features.
[0128] Specifically, the sound encoder performs speech recognition on the sound data, and identifies the voice command features by extracting features such as the pitch, frequency, and formant of the voice. The sound encoder performs non-language sound recognition on the sound data, and extracts non-language sound features that accompany the user's body posture during human-computer interaction, such as clapping and knocking sounds. These non-language sounds may be associated with the user's body posture to provide additional interaction information. The voice command features and non-language sound features are used as sound features under the auditory modality.
[0129] S403, extracting force control features based on force control data by the force encoder in the neural network model.
[0130] Optionally, a force encoder is designed for the force control data of the force sensor in advance through a deep neural network model, and the force control features of the tactile modality are extracted through the force encoder. The force encoder can also be an RNN, and the input layer of the RNN receives a sequence of force control data, which is processed by one or more RNN layers and then obtained through a fully connected layer to obtain a force control feature representation, so as to better capture the change trend of the force.
[0131] Specifically, the force encoder performs touch recognition on the force control data and extracts touch features such as the touch force, duration, and location of the touch point. The force encoder performs push-pull action recognition on the force control data and extracts push-pull action features by analyzing the force change, speed, and direction of the push-pull action. The force encoder performs interactive mode recognition on the force control data and combines touch and push-pull actions to identify the user's interactive intentions, such as selection, sliding, dragging, etc., to obtain interactive action features. The touch features, push-pull action features, and interactive action features are used as force control features under the tactile mode.
[0132] S404: The decoder in the neural network model obtains multimodal fusion features according to the posture features, the sound features and the force control features.
[0133] Optionally, a decoder based on the attention mechanism is designed in advance through a deep neural network model, and the decoder fuses posture features, sound features, and force control features through the attention mechanism to form a unified multimodal fusion feature.
[0134] Among them, the decoder based on the attention mechanism can dynamically assign attention weights to different modal features, thereby automatically focusing on more important modal features according to the requirements of the interactive task and improving the feature fusion effect.
[0135] In this embodiment, a visual encoder, a sound encoder and a force encoder are designed for the posture data of the visual modality, the sound data of the auditory modality and the force control data of the tactile modality respectively through a deep neural network model, and each encoder extracts the features of the corresponding modality. The visual encoder in the neural network model extracts the posture features under the visual modality according to the posture data, the sound encoder extracts the sound features under the auditory modality according to the sound data, and the force encoder extracts the force control features under the tactile modality according to the force control data. A decoder based on the attention mechanism is designed through a deep neural network model, and the decoder fuses the posture features under the visual modality, the sound features under the auditory modality and the force control features under the tactile modality through the attention mechanism to form a unified multimodal fusion feature. By dynamically allocating the attention weights of each modal feature, it automatically focuses on the more important modal features to improve the feature fusion effect.
[0136] The following is a detailed description of the process of obtaining multimodal fusion features based on posture features, sound features, and force control features.
[0137] Figure 5 Another schematic diagram of a process of obtaining multimodal fusion features of the robot control method provided in the embodiment of the present application is as follows: Figure 5 As shown, the step of obtaining multimodal fusion features according to the posture features, the sound features and the force control features in the above step S404 includes:
[0138] S501. Determine a first attention weight of the posture feature, a second attention weight of the sound feature, and a third attention weight of the force control feature based on the posture feature, the sound feature, and the force control feature.
[0139] Optionally, the attention mechanism-based decoder utilizes a multimodal attention network, takes posture features in the visual modality, sound features in the auditory modality, and force control features in the tactile modality as vector inputs, performs linear transformations on the posture features, sound features, and force control features, and maps the three features to the same dimension.
[0140] In the same dimension, the first attention score of the posture feature, the second attention score of the sound feature, and the third attention score of the force control feature are calculated, and each attention score is converted into an attention weight through the Softmax function to obtain the first attention weight of the posture feature, the second attention weight of the sound feature, and the third attention weight of the force control feature.
[0141] S502: Obtain multimodal fusion features according to the posture features, the sound features, the force control features, the first attention weight, the second attention weight, and the third attention weight.
[0142] Optionally, according to the first attention weight of the posture feature, the second attention weight of the sound feature and the third attention weight of the force control feature, the features of each modality are weighted and summed to obtain a multimodal fusion feature.
[0143] Specifically, the first product of the posture feature and the first attention weight, the second product of the sound feature and the second attention weight, and the third product of the force control feature and the third attention weight are calculated, and the sum of the first product, the second product, and the third product is used as the multimodal fusion feature. Through this dynamic weight allocation method, the fused multimodal features can more accurately reflect the user's interaction intention, so that the robot can more effectively understand the user's needs and make more appropriate interactive responses.
[0144] In this embodiment, the decoder based on the attention mechanism uses a multimodal attention network to obtain the first attention weight of the posture feature, the second attention weight of the sound feature, and the third attention weight of the force control feature according to the posture feature in the visual mode, the sound feature in the auditory mode, and the force control feature in the tactile mode. According to the first attention weight of the posture feature, the second attention weight of the sound feature, and the third attention weight of the force control feature, the features of each modality are weighted and summed to obtain a multimodal fusion feature. The attention weights of each modal feature are dynamically allocated based on the current interaction needs, so that the fused multimodal features more accurately reflect the user's interaction intentions and achieve effective fusion of multimodal features.
[0145] The following is a detailed description of the process of generating a response strategy based on multimodal fusion features and context information.
[0146] Figure 6 A schematic diagram of a flow chart of a response strategy generation method for a robot control method provided in an embodiment of the present application, such as Figure 6 As shown, the step of generating a response strategy according to the multimodal fusion features and context information in the above step S104 includes:
[0147] S601. Generate an initial response strategy using an imitation learning algorithm based on multimodal fusion features.
[0148] Optionally, the multimodal fusion features are input into a pre-trained decision model, which may be a fully connected neural network response strategy model. An imitation learning algorithm is used to map the multimodal fusion features to the response strategy, and the initial response strategy is predicted and output.
[0149] Specifically, the decision-making model adopts an imitation learning algorithm to make autonomous decisions based on multimodal fusion features, and predicts and outputs natural initial response strategies by imitating human behavior.
[0150] S602. According to the multimodal fusion features and context information, a reinforcement learning algorithm is used to determine the reward signal corresponding to the initial response strategy, and the initial response strategy is iteratively corrected according to the reward signal, and the initial response strategy at the end of the iteration is used as the response strategy.
[0151] Optionally, the multimodal fusion features and the contextual information of the current human-computer interaction are input into the decision model, a reinforcement learning algorithm is used to evaluate the expected effect of the initial response strategy, and the initial response strategy is iteratively corrected according to the evaluation results.
[0152] Specifically, an intelligent agent is pre-designed in the decision model. At each time step, the intelligent agent can receive reward signals based on multimodal fusion features, current human-computer interaction context information, and the initial response strategy, and iteratively correct the initial response strategy based on each reward signal. Through continuous trial and error and strategy correction, the iteration is stopped when the optimal initial response strategy is learned, and the initial response strategy at the end of the iteration is used as the response strategy.
[0153] In this embodiment, the decision model uses an imitation learning algorithm to make autonomous decisions based on multimodal fusion features, and predicts and outputs a natural initial response strategy by imitating human behavior. The multimodal fusion features and the context information of the current human-computer interaction are input into the decision model, and the reinforcement learning algorithm is used to determine the reward signal corresponding to the initial response strategy, and the initial response strategy is iteratively corrected according to the reward signal, and the initial response strategy at the end of the iteration is used as the response strategy. Using multimodal fusion features and context information, the optimal response strategy is learned through continuous trial and error, which improves the autonomous decision-making performance during human-computer interaction.
[0154] The following is a detailed description of the process of controlling the robot motion according to the response strategy.
[0155] Figure 7 A schematic diagram of a flow chart of a robot control method for controlling robot motion provided in an embodiment of the present application, such as Figure 7 As shown, the step of controlling the robot movement according to the response strategy in the above step S104 includes:
[0156] S701. Determine the robot's actions to be performed according to the response strategy, and generate voice data corresponding to the actions to be performed.
[0157] Optionally, a predefined action space list is obtained, which includes various actions that the robot may perform. The robot's motion trajectory and action are planned according to the response strategy, and the action index with the highest probability of matching the response strategy is found. Based on the action index, the corresponding action is determined from the action space list as the robot's to-be-performed action, such as reaching out, picking up and placing an object, etc., and the to-be-performed action is converted into voice data to generate voice data corresponding to the to-be-performed action.
[0158] S702: Generate a control instruction according to the action to be executed and the voice data, and control the robot to execute the action to be executed and output the voice data at the same time according to the control instruction.
[0159] Optionally, a control instruction is generated according to the robot's action to be performed and the corresponding voice data, and the control instruction is sent to the robot controller, so that the robot controller controls the robot to output voice data while performing the action to be performed based on the control instruction.
[0160] Specifically, the robot controller ensures the synchronization of the robot's action execution and voice output based on the control instructions. For example, when the robot starts to perform the action of picking up a cup, the voice prompt of "pick up the cup" is played at the same time.
[0161] In this embodiment, according to the response strategy, the robot's pending actions are determined, and the pending actions are converted into voice data to generate voice data corresponding to the pending actions. A control instruction is generated according to the robot's pending actions and the corresponding voice data, and the robot controller controls the robot to output voice data while executing the pending actions based on the control instructions. A more natural, accurate and intuitive human-computer interaction is achieved.
[0162] The following describes in detail the process of generating voice data corresponding to the action to be performed.
[0163] Figure 8 A schematic diagram of a flow chart of generating voice data for a robot control method provided in an embodiment of the present application, such as Figure 8 As shown, the step of generating voice data corresponding to the action to be performed in the above step S701 includes:
[0164] S801. Generate a natural language response matching the action to be executed according to the action to be executed.
[0165] Optionally, the action to be performed is identified and the action intention is obtained by parsing. According to the action intention, a target natural language template matching the action intention is selected from predefined natural language templates. Specific action information such as picking up a cup is filled into the target natural language template to generate an initial natural language response, and the initial natural language response is checked for grammar, semantics, and language style to obtain a natural language response.
[0166] Among them, grammar checking can ensure that the generated natural language response complies with grammatical rules and avoids grammatical errors. Semantic checking can ensure that the semantics are clear and easy to understand, and avoid using overly complex or professional terms. Language style checking based on the current user's identity and interaction scenario can maintain consistency in language style. For example, if the user is a child, a simple and lively language style can be used in a child interaction scenario.
[0167] S802: Generate speech data corresponding to the action to be performed through speech synthesis according to the natural language response.
[0168] Optionally, based on a preset text-to-speech library, the generated natural language response is converted into voice data through speech synthesis technology, and the voice data corresponds to the robot's action to be performed, so that the robot outputs voice data while performing the action to be performed.
[0169] In this embodiment, the robot's pending action is identified and the action intention is obtained by parsing. According to the action intention, a natural language response matching the pending action is generated. The generated natural language response is converted into voice data corresponding to the robot's pending action through speech synthesis technology, so that the robot outputs voice data while executing the pending action, providing a richer human-computer interaction experience.
[0170] Based on the same inventive concept, a robot control device corresponding to the robot control method is also provided in the embodiment of the present application. Since the principle of solving the problem by the device in the embodiment of the present application is similar to the above-mentioned robot control method in the embodiment of the present application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be repeated.
[0171] Fig. 9 The module structure diagram of the robot control device provided in the embodiment of the present application is as follows: Fig. 9 As shown, the device comprises:
[0172] The acquisition module 901 is used to acquire multimodal data of the user, where the multimodal data includes posture data, sound data, and force control data.
[0173] The processing module 902 is used to perform feature extraction and fusion processing on the multimodal data to obtain multimodal fusion features.
[0174] The determination module 903 is used to determine the context information of the human-machine interaction according to the environmental information of the area where the robot is located and the historical interaction data of the robot.
[0175] The control module 904 is used to generate a response strategy according to the multimodal fusion features and context information, and control the robot movement according to the response strategy.
[0176] As an optional implementation, the device further includes: a correction module 905; the correction module 905 is used to:
[0177] Obtain feedback data based on the robot's motion information input by the user.
[0178] Modify the response strategy based on the feedback data.
[0179] As an optional implementation, the acquisition module 901 is specifically used for:
[0180] The user's posture data is obtained through visual sensors, and the posture data includes the user's facial expressions, gestures, and body postures.
[0181] The user's voice data is acquired through the sound sensor, and the sound data includes the user's voice commands and non-verbal sounds.
[0182] The force sensor is used to obtain the user's force control data, which includes the user's touch interaction signal and push-pull interaction signal.
[0183] As an optional implementation manner, the processing module 902 is specifically configured to:
[0184] The visual encoder in the neural network model extracts posture features based on the posture data.
[0185] The sound encoder in the neural network model extracts sound features based on the sound data.
[0186] The force encoder in the neural network model extracts force control features based on the force control data.
[0187] The decoder in the neural network model obtains multimodal fusion features based on posture features, sound features and force control features.
[0188] As an optional implementation manner, the processing module 902 is specifically configured to:
[0189] Determine, according to the posture feature, the sound feature, and the force control feature, a first attention weight of the posture feature, a second attention weight of the sound feature, and a third attention weight of the force control feature;
[0190] According to the posture features, the sound features, the force control features, the first attention weight, the second attention weight and the third attention weight, a multimodal fusion feature is obtained.
[0191] As an optional implementation manner, the control module 904 is specifically configured to:
[0192] According to the multimodal fusion features, the imitation learning algorithm is used to generate the initial response strategy.
[0193] According to the multimodal fusion features and context information, the reinforcement learning algorithm is used to determine the reward signal corresponding to the initial response strategy, and the initial response strategy is iteratively corrected according to the reward signal, and the initial response strategy at the end of the iteration is used as the response strategy.
[0194] As an optional implementation manner, the control module 904 is specifically configured to:
[0195] According to the response strategy, the robot's actions to be performed are determined, and voice data corresponding to the actions to be performed are generated.
[0196] A control instruction is generated according to the action to be performed and the voice data, and the robot is controlled to perform the action to be performed according to the control instruction while outputting the voice data.
[0197] As an optional implementation, the control module 904 is specifically configured to:
[0198] According to the action to be performed, a natural language response matching the action to be performed is generated.
[0199] Based on the natural language response, speech data corresponding to the action to be performed is generated through speech synthesis.
[0200] The present application also provides a computer device, such as Fig.10 1 is a schematic diagram of the structure of a computer device provided in an embodiment of the present application, including: a processor 101, a memory 102 and a bus 103. The memory 102 stores machine-readable instructions executable by the processor 101 (for example, Fig. 9 In the device, the execution instructions corresponding to the acquisition module 901, the processing module 902, the determination module 903, the control module 904 and the correction module 905 are obtained, etc. When the computer device is running, the processor 101 communicates with the memory 102 through the bus 103. When the machine-readable instructions are executed by the processor 101, the steps of the robot control method in the above embodiment are executed.
[0201] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the robot control method in the above embodiment are executed.
[0202] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, the specific working process of the system and device described above can refer to the corresponding process in the method embodiment, and will not be repeated in this application. In the several embodiments provided in this application, it should be understood that the disclosed system, device and method can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the modules is only a logical function division. There may be other division methods in actual implementation. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.
[0203] In addition, each functional unit in each embodiment of the present application can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, RandomAccess Memory), disk or optical disk and other media that can store program code.
[0204] The above are only specific implementation methods of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be covered by the protection scope of the present application.
Claims
1. A robot control method, characterized in that: include: Acquiring multimodal data of the user, wherein the multimodal data includes posture data, sound data, and force control data; Performing feature extraction and fusion processing on the multimodal data to obtain multimodal fusion features; Determining context information of human-machine interaction based on environmental information of the area where the robot is located and historical interaction data of the robot; A response strategy is generated according to the multimodal fusion features and the context information, and the robot movement is controlled according to the response strategy.
2. The method according to claim 1, characterized in that After controlling the movement of the robot according to the response strategy, the method further comprises: Acquiring feedback data input by a user based on motion information of the robot; The response strategy is modified according to the feedback data.
3. The method according to claim 1, characterized in that The obtaining of multimodal data of the user includes: Acquiring user's posture data through a visual sensor, wherein the posture data includes the user's facial expression, hand gestures and body posture; Acquiring the user's voice data through a voice sensor, wherein the voice data includes the user's voice commands and non-verbal sounds; The force control data of the user is acquired through the force sensor, and the force control data includes the touch interaction signal and the push-pull interaction signal of the user.
4. The method according to claim 1, characterized in that: The extracting and fusing the multimodal data to obtain multimodal fusion features includes: Extracting posture features according to the posture data by a visual encoder in the neural network model; The sound encoder in the neural network model extracts sound features according to the sound data; Extracting force control features according to the force control data by a force encoder in the neural network model; The decoder in the neural network model obtains the multimodal fusion feature according to the posture feature, the sound feature and the force control feature.
5. The method according to claim 4, characterized in that The obtaining of the multimodal fusion feature according to the posture feature, the sound feature and the force control feature includes: Determining, according to the posture feature, the sound feature, and the force control feature, a first attention weight of the posture feature, a second attention weight of the sound feature, and a third attention weight of the force control feature; The multimodal fusion feature is obtained according to the posture feature, the sound feature, the force control feature, the first attention weight, the second attention weight and the third attention weight.
6. The method according to claim 1, characterized in that The generating a response strategy according to the multimodal fusion feature and the context information includes: According to the multimodal fusion features, an initial response strategy is generated using an imitation learning algorithm; According to the multimodal fusion features and the context information, a reinforcement learning algorithm is used to determine the reward signal corresponding to the initial response strategy, and the initial response strategy is iteratively corrected according to the reward signal, and the initial response strategy at the end of the iteration is used as the response strategy.
7. The method according to claim 1, characterized in that The controlling the robot movement according to the response strategy comprises: Determine the robot's action to be performed according to the response strategy, and generate voice data corresponding to the action to be performed; A control instruction is generated according to the action to be performed and the voice data, and the robot is controlled according to the control instruction to perform the action to be performed while outputting the voice data.
8. The method according to claim 7, characterized in that The generating of voice data corresponding to the action to be performed includes: According to the action to be performed, generating a natural language response matching the action to be performed; According to the natural language response, speech data corresponding to the action to be performed is generated through speech synthesis.
9. A computer device, characterized in that: include: A processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the computer device is running, the processor communicates with the memory via the bus, and the processor executes the machine-readable instructions to perform the steps of the robot control method according to any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the steps of the robot control method according to any one of claims 1 to 8 are executed.
Citation Information
Patent Citations
Multi-modal fusion natural interaction method and system of intelligent robot and medium
CN114995657A
Method for applying fused multi-modal machine learning to automatic production line balance
CN115222192A
Multi-modal interaction method and apparatus
WO2023216765A1
Cited By
Humanoid robot somatosensory control system and method based on deep learning
CN120395905A
Multi-modal sensing fusion robot humanoid operation method and system
CN120588247A
Self-adaptive strategy optimization method and system for robot with body based on interactive feedback
CN121223793A
Accompanying robot interaction control method and device, electronic equipment and storage medium
CN122018702A
Companion robot interaction control method and device, electronic equipment and storage medium
CN122018702B