Interactive Control Method, Device, Storage Medium and Robot of Robot

By acquiring multiple modal data and performing sensory knowledge and semantic-level environmental characteristics analysis, the problem of poor interaction accuracy of service robots is solved, and higher interaction accuracy is achieved.

CN115213884BActive Publication Date: 2025-08-05CLOUDMINDS BEIJING TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110729750.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-29
Publication Date
2025-08-05
Estimated Expiration
2041-06-29

AI Technical Summary

Technical Problem

When interacting with objects in the environment, service robots have poor interaction accuracy. This is mainly due to the complex and diverse objects in the environment, the robot cannot accurately understand the characteristics of the environment.

Method used

Data in various modes are obtained, including data collected through flexible screen limb components, environmental state sensors, vision sensors, speech sensors and force sensors, to distinguish them in a sense, obtain semantic-level environmental features, and predict target interaction information based on the object semantic relationship network to control the robot to interact.

Benefits of technology

Through the use of semantic-level environmental features, the ambiguity of environmental feature recognition is reduced and the accuracy of robot interaction is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115213884B_ABST
    Figure CN115213884B_ABST
Patent Text Reader

Abstract

This disclosure relates to a robot interaction control method, device, storage medium, and robot. The method comprises: acquiring data in multiple modalities for controlling the robot in a target environment, wherein the data in each modality represents a type of data source data; perceiving and identifying the data in multiple modalities to obtain semantic-level environmental features corresponding to the various modalities in the target environment; predicting target interaction information corresponding to the robot based on the semantic-level environmental features corresponding to the various modalities and the object semantic relationship network corresponding to the target environment; and controlling the robot to interact based on the target interaction information. This method can improve the accuracy of robot interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of robotics technology, and in particular, to a robot interactive control method, device, storage medium, and robot. Background Art

[0002] With the development of robotics technology, service robots are increasingly appearing in our daily lives. Service robots complete their work by interacting with objects in the environment, such as people or objects.

[0003] However, when the service robot in the related art interacts with objects in the environment, due to the complexity and diversity of the objects in the environment, the service robot cannot accurately recognize the environmental characteristics, which in turn leads to the problem of poor interaction accuracy of the service robot. Summary of the Invention

[0004] The purpose of the present disclosure is to provide a robot interaction control method, device, storage medium and robot, which solve the problem of poor interaction accuracy of the robot.

[0005] To achieve the above objectives, in a first aspect, the present disclosure provides a method for interactive control of a robot, the method comprising:

[0006] Acquire multiple modal data for controlling the robot in the target environment, where data of one modality represents a type of data source data;

[0007] Perceive and identify data from multiple modalities to obtain semantic-level environmental features corresponding to various modalities in the target environment;

[0008] Based on the semantic-level environmental features corresponding to various modalities and the object semantic relationship network corresponding to the target environment, the target interaction information corresponding to the robot is predicted;

[0009] Based on the target interaction information, the robot is controlled to interact.

[0010] Optionally, after acquiring data of multiple modalities in the target environment, the method further includes:

[0011] Optimize the data of multiple modes separately to obtain optimized data of multiple modes;

[0012] Perform fusion correction processing on the optimized data of multiple modalities to obtain fusion corrected data of multiple modalities;

[0013] Perceive and identify data from multiple modalities to obtain semantic-level environmental features corresponding to various modalities in the target environment, including:

[0014] The fused and corrected data of multiple modalities are perceived and recognized to obtain the semantic-level environmental features corresponding to various modalities in the target environment.

[0015] Optionally, based on semantic-level environmental features corresponding to various modalities and an object semantic relationship network corresponding to the target environment, target interaction information corresponding to the robot is predicted, including:

[0016] The semantic-level environmental features corresponding to various modalities and the object semantic relationship network corresponding to the target environment are input into the target interaction information prediction model to obtain the target interaction information corresponding to the robot.

[0017] Optionally, based on semantic-level environmental features corresponding to various modalities and an object semantic relationship network corresponding to the target environment, target interaction information corresponding to the robot is predicted, including:

[0018] Align the semantic-level environment features corresponding to various modalities to obtain cross-semantic environment features corresponding to various modalities;

[0019] The cross-semantic environment features corresponding to various modalities are fused to obtain the fused environment features in the target environment;

[0020] The fusion environment features and the object semantic relationship network corresponding to the target environment are input into the target interaction information prediction model to obtain the target interaction information corresponding to the robot.

[0021] Optionally, the training process of the target interaction information prediction model includes:

[0022] Acquire a plurality of sample data, wherein each sample data includes a sample object corresponding to a target environment and an object semantic relationship network;

[0023] The initial neural network model is trained based on multiple sample data until the preset training conditions are met, the training is stopped and the target interaction information prediction network is output.

[0024] Optionally, the robot includes flexible screen limb parts, environmental state sensors, visual sensors, voice sensors and force sensors, and the data of multiple modalities include tactile data collected by the flexible screen limb parts, environmental state data collected by the environmental state sensors, visual data collected by the visual sensors, voice data collected by the voice sensors and force data collected by the force sensors.

[0025] Optionally, the target interaction information includes multimodal target interaction information, and each modal target interaction information carries corresponding timing information. Based on the target interaction information, controlling the robot to interact includes:

[0026] Based on the target interaction information of each modality and the timing information carried by the target interaction information of each modality, the robot is controlled to interact.

[0027] Optionally, the target interaction information includes an image to be displayed, and controlling the robot to interact based on the target interaction information includes:

[0028] Control the flexible screen limb parts to display the image to be displayed.

[0029] Optionally, the target interaction information includes position information of the flexible screen limb component, and controlling the flexible screen limb component to display the image to be displayed includes:

[0030] Acquire the association relationship between each image area included in the image to be displayed and the preset display position;

[0031] Based on the association relationship between each image area included in the image to be displayed and the preset display position, obtaining the target image area corresponding to the position information of each flexible screen limb component;

[0032] Control the flexible screen limb component to display the image of the target image area.

[0033] In a second aspect, the present disclosure further provides an interactive control device for a robot, the device comprising: a multimodal data acquisition module for acquiring data of multiple modalities for controlling the robot in a target environment, wherein data of one modality represents data of a type of data source;

[0034] The perception and recognition module is used to perceive and recognize data of multiple modalities and obtain semantic-level environmental features corresponding to various modalities in the target environment;

[0035] The prediction module is used to predict the target interaction information corresponding to the robot based on the semantic-level environment features corresponding to various modalities and the object semantic relationship network corresponding to the target environment;

[0036] The control module is used to control the robot to interact based on the target interaction information.

[0037] In a third aspect, the present disclosure further provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of the method in the first aspect when executed by a processor.

[0038] In a fourth aspect, the present disclosure further provides an interactive control device for a robot, comprising:

[0039] a memory having a computer program stored thereon;

[0040] A processor is configured to execute a computer program in a memory to implement the steps of the method in the first aspect.

[0041] In a fifth aspect, the present disclosure further provides a robot comprising an actuator, multiple types of sensors, and a processing device connected to the multiple sensors and the actuator, wherein the actuator comprises a flexible screen limb component provided on a robot body;

[0042] The processing device is used to: obtain data of multiple modalities, including tactile data collected by flexible screen limb parts, environmental status data collected by multiple types of sensors, visual data, voice data and a combination of at least two types of force data; perceive and identify data of multiple modalities to obtain semantic-level environmental features corresponding to various modalities in the target environment; predict target interaction information corresponding to the robot based on the semantic-level environmental features corresponding to various modalities and the object semantic relationship network corresponding to the target environment; and control the corresponding actuator to perform interactive operations based on the target interaction information.

[0043] Through the above technical solution, after acquiring multi-modal data used to control the robot in the target environment, the multi-modal data is first perceived and identified to obtain semantic-level environmental features corresponding to the various modalities in the target environment. Then, based on the semantic-level environmental features corresponding to the various modalities and the object semantic relationship network corresponding to the target environment, the target interaction information corresponding to the robot is predicted. Finally, based on the target interaction information, the robot can be controlled to interact. Because the semantic-level environmental features take into account the semantic information carried by the environmental features, they can reduce the ambiguity of the identified environmental features. Therefore, the semantic-level environmental features can improve the accuracy of the predicted target interaction information, thereby improving the accuracy of the robot interaction.

[0044] Other features and advantages of the present disclosure will be described in detail in the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The accompanying drawings are used to provide a further understanding of the present disclosure and constitute a part of the specification. Together with the following detailed description, they are used to explain the present disclosure but do not constitute a limitation of the present disclosure. In the accompanying drawings:

[0046] Figure 1 This is a flow chart of a robot interactive control method provided in an embodiment.

[0047] Figure 2 It is a flow chart of step S13 in the embodiment.

[0048] Figure 3 This is a flowchart of a training process of a target interaction information prediction model provided by an embodiment.

[0049] Figure 4This is a flow chart of controlling a flexible screen limb component to display an image to be displayed, as provided in an embodiment.

[0050] Figure 5 It is a flowchart of another robot interactive control method provided in an embodiment.

[0051] Figure 6 It is a structural diagram of an interactive control device for a robot provided in an embodiment.

[0052] Figure 7 It is a structural diagram of another interactive control device for a robot provided in an embodiment.

[0053] Figure 8 It is a structural diagram of another interactive control device for a robot provided in an embodiment. DETAILED DESCRIPTION

[0054] The following describes the specific embodiments of the present disclosure in detail with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to illustrate and explain the present disclosure and are not intended to limit the present disclosure.

[0055] See also Figure 1 , Figure 1 FIG. 1 is a flow chart of a robot interactive control method according to an exemplary embodiment of the present disclosure. Figure 1 As shown, the robot interactive control method includes steps S11 to S14. Specifically:

[0056] S11, acquiring data of multiple modes for controlling a robot in a target environment, wherein data of one mode represents a type of data source data.

[0057] The target environment refers to the environment in which the robot being controlled is currently located. It's understandable that different service robots serve different environments. For example, a welcome robot primarily serves at the entrance, an educational robot primarily serves in the classroom, and a conference room service robot primarily serves in the meeting room. For example, when a welcome robot serves at the entrance of a venue, the target environment is the environment at the entrance of the venue it serves.

[0058] In this disclosure, each source or form of information can be referred to as a modality. Data from different sources can be considered data of different modalities, or data of one modality can be considered to represent data from one source. For example, data from different sources can be collected using different sensors, and thus data collected by multiple different sensors can be referred to as multi-modal data. It is understood that different robots can carry different types of sensors depending on the tasks they need to complete. Consequently, the multi-modal data collected will also be different.

[0059] In some embodiments, the robot may carry basic sensors such as tactile sensors, environmental state sensors, visual sensors, voice sensors, and force sensors. In this case, the multi-modal data may include tactile data collected by the tactile sensors, environmental state data collected by the environmental state sensors, visual data collected by the visual sensors, voice data collected by the voice sensors, and force data collected by the force sensors. The visual sensors may be sensors such as 3D depth cameras and lidars.

[0060] In other embodiments, in addition to carrying basic sensors such as environmental state sensors, visual sensors, voice sensors, and force sensors, the robot can also carry flexible screen limb parts and use the flexible screen limb parts as tactile sensors. In this case, the multi-modal data can include tactile data collected by the flexible screen limb parts, environmental state data collected by the environmental state sensors, visual data collected by the visual sensors, voice data collected by the voice sensors, and force data collected by the force sensors. Among them, the flexible screen limb parts can be understood as the flexible touch screen shell covering the robot limb parts.

[0061] In this embodiment, the robot can be any humanoid robot or any other form of robot. The humanoid robot's limbs include, but are not limited to, a multi-degree-of-freedom head, neck joints, arms (upper arms, lower arms, elbow joints), humanoid hands (2-5 fingers, each with 2-3 joints), a multi-degree-of-freedom waist, single or dual legs, knee joints, feet, and a wheeled chassis with various walking positions. The outer shells of these robot limbs can all or partially utilize this flexible touch screen to form the robot's "skin."

[0062] It can be understood that by setting up flexible screen limb parts, objects in the environment can collect tactile data by contacting any limb part of the robot, which simplifies the tactile data collection process and makes tactile data collection more convenient.

[0063] In addition, the flexible screen limb component in this embodiment supports but is not limited to single-point touch, multi-point touch, sliding touch, etc.

[0064] There are many ways to obtain data of multiple modes for controlling a robot in a target environment.

[0065] As an implementation method, all multimodal data can be obtained from sensors carried by the robot.

[0066] In addition, considering that the sensors carried by the robot may malfunction or have data collection authority restrictions, it is impossible to obtain all the required data of various modalities from the sensors carried by the robot. In this case, as another implementation method, multimodal data can be obtained partially from the sensors carried by the robot and partially from sensors installed in the target environment.

[0067] In the case of obtaining data from sensors installed in the target environment, specifically, a communication connection may be first established with the sensors installed in the target environment, and then data of the corresponding modality may be obtained from the sensors installed in the target environment.

[0068] S12, perceive and identify data of multiple modalities to obtain semantic-level environmental features corresponding to various modalities in the target environment.

[0069] Semantic-level environmental features refer to environmental features that take semantic information into account. Environmental features can be understood as objects in the environment and their attributes. For example, without considering semantic information, an apple can be considered an edible fruit, but it can also be understood as a brand of mobile phone. Therefore, in order to ensure that the identified environmental features are unambiguous and avoid incorrect interactive behaviors when controlling the robot, in this embodiment, data from multiple modalities can be first perceived and identified to obtain semantic-level environmental features corresponding to each modality in the target environment.

[0070] There are many ways to perceive and identify data of multiple modalities and obtain semantic-level environmental features corresponding to various modalities in the target environment.

[0071] Optionally, a compact hash coding method for multimodal expression can be used to process data of multiple modalities. This method can take into account the correlation constraints within and between modalities. Then, the orthogonal regularization method is used to further process the obtained hash coding features, thereby reducing the redundancy of the hash coding features and finally obtaining the semantic-level environmental features corresponding to various modalities in the target environment.

[0072] Optionally, a partial multimodal sparse coding model based on adaptive similarity structure regularization can be used to process data of multiple modalities to obtain semantic-level environmental features corresponding to various modalities in the target environment.

[0073] S13, based on the semantic-level environmental features corresponding to various modalities and the object semantic relationship network corresponding to the target environment, predict the target interaction information corresponding to the robot.

[0074] The object semantic relationship network corresponding to the target environment can be constructed based on the scene information and prior knowledge of the target environment. Target interaction information refers to the information used to control the robot to interact.

[0075] In some embodiments, the predicted interaction information corresponding to the robot may be only one type. In this case, this interaction information can be directly determined as the target interaction information. In other embodiments, the predicted interaction information corresponding to the robot may be multiple types. In this case, a pre-set comprehensive decision module can be used to make a comprehensive decision on the multiple types of interaction information corresponding to the robot, and determine the optimal interaction information from the multiple types of interaction information, i.e., the target interaction information.

[0076] For example, let's assume the target environment is a conference room. Based on the scene information for the conference room, we know that a conference room typically has a conference table and tea cups, with the cups placed on the table 50 cm apart. When attendees are present, tea should be poured into the cups. Furthermore, based on prior knowledge, we know that the tea should not exceed 80% of the cup's rim and that the cup should be covered after pouring. In this case, based on the scene information and prior knowledge for the target environment, we can construct an object semantic relationship network between objects such as the tea table, tea cups, tea, and cup lids.

[0077] In this embodiment, after obtaining the semantic-level environmental features corresponding to various modalities and the object semantic relationship network corresponding to the target environment, the target interaction information corresponding to the robot can be predicted by semantic reasoning.

[0078] Target interaction information refers to the instructions used to control the robot's interaction with objects in the environment. For example, in a conference room setting, the target interaction information might be a water-adding instruction, specifically, to add tea to the first cup, filling it to 75% of the rim and then replacing the lid. Another example, in a guest reception setting, the target interaction information might be instructions such as smiling, bending over, and extending a hand to shake the guest's hand.

[0079] In some embodiments, target interaction information

[0080] Among them, there are many ways to predict the target interaction information corresponding to the robot based on the semantic-level environmental features corresponding to various modalities and the object semantic relationship network corresponding to the target environment.

[0081] As an implementation method, the semantic-level environmental features corresponding to various modalities and the object semantic relationship network corresponding to the target environment can be directly input into the target interaction information prediction model to obtain the target interaction information corresponding to the robot.

[0082] As another implementation method, we can first process the semantic-level environmental features corresponding to various modalities, and then input the processed environmental features and the object semantic relationship network corresponding to the target environment into the target interaction information prediction model to obtain the target interaction information corresponding to the robot. In this case, please refer to Figure 2 ,like Figure 2 As shown, step S13 may include the following steps:

[0083] S131, aligning the semantic-level environment features corresponding to various modalities to obtain cross-semantic environment features corresponding to various modalities.

[0084] Among them, alignment processing refers to identifying the correspondence between components and elements between different modalities, thereby making the learned cross-semantic environment feature representations of various modalities more accurate and providing more detailed clues for the subsequent target interaction information prediction model.

[0085] Optionally, specific methods of alignment processing include but are not limited to using a maximum margin learning method combined with local alignment (for example, alignment of visual objects and words, or alignment with touched body parts) and global alignment (for example, alignment of pictures and sentences, or alignment with actions corresponding to touch (handshake / hug, etc.)) methods to learn a common embedding representation space. The aligned cross-semantic representation can better improve the prediction quality of the target interaction information prediction model.

[0086] S132: Fusing the cross-semantic environment features corresponding to various modalities to obtain fused environment features in the target environment.

[0087] Fusion processing involves integrating models and features from different modalities. This process can yield more comprehensive features, improve model robustness, and ensure that the model still works effectively even when certain modal information is missing. For example, it can still work effectively even when a user's visual expression is missing, or when the expression is happy but the language expresses frustration. Another example is when voice recognition indicates a person's emotion is happy, but visual recognition indicates frustration. While both describe the person's emotions, contextual information from the current scene must be incorporated. For example, the voice expressing happiness may be from a different person.

[0088] S133, inputting the fusion environment features and the object semantic relationship network corresponding to the target environment into the target interaction information prediction model to obtain the target interaction information corresponding to the robot.

[0089] After obtaining the fused environment features, the fused environment features and the object semantic relationship network corresponding to the target environment can be input into the target interaction information prediction model to obtain the target interaction information corresponding to the robot.

[0090] It can be understood that the semantic-level environmental features and fused environmental features corresponding to various modalities are all semantic-level environmental features, and therefore can be processed using the same target interaction information prediction model.

[0091] Next, combine Figure 3 , the training process of the target interaction information prediction model is introduced. Figure 3 As shown in Figure 2, the training process of the target interaction information prediction model includes the following steps:

[0092] S21, obtaining a plurality of sample data.

[0093] Each sample data includes a sample object corresponding to the target environment and an object semantic relationship network, and the sample object carries corresponding semantic information.

[0094] As an implementation method, the sample objects in the target environment may be manually selected and determined, and semantic information may be annotated for them.

[0095] S22, training the initial neural network model based on multiple sample data until the preset training conditions are met, stopping the training and outputting the target interaction information prediction network.

[0096] Optionally, the preset training condition may be a preset number of iterations. Correspondingly, determining whether the preset training condition is met is as follows: determining whether the current number of iterations is greater than the preset number of iterations, and when the current number of iterations is greater than the preset number of iterations, determining that the preset condition is met.

[0097] As an implementation method, the initial neural network model may be a Markov decision chain based on reinforcement learning.

[0098] S14, based on the target interaction information, controlling the robot to interact.

[0099] It should be noted that the robot interactive control method provided by the embodiment of the present disclosure can be executed only locally on the robot, or only on the server, or partially on the robot and partially on the server.

[0100] As an embodiment, when the interactive control method of the robot is executed locally on the robot, the robot may include an actuator, multiple types of sensors, and a processing device connected to the multiple sensors and actuators, the actuator including a flexible screen limb part provided on the robot body, wherein the flexible screen limb part can be obtained by covering the robot limb part with a flexible screen. In this case, the processing device is specifically used to obtain data of multiple modalities, the multiple modalities of data including a combination of at least two of tactile data collected by the flexible screen limb part, environmental state data collected by multiple types of sensors, visual data, voice data, and force data; perceive and identify the data of multiple modalities to obtain semantic-level environmental features corresponding to various modalities in the target environment; predict the target interaction information corresponding to the robot based on the semantic-level environmental features corresponding to the various modalities and the object semantic relationship network corresponding to the target environment; and control the corresponding actuator to perform interactive operations based on the target interaction information.

[0101] Of course, data of some of the multiple modalities can also be obtained from sensors installed in the target environment.

[0102] As another embodiment, when the robot's interactive control method is executed on the server, the server can execute the above steps S11-S14. In this case, in step S11, the server can obtain data of multiple modalities from the sensors carried by the robot. The server can also obtain data from the sensors carried by the robot and the sensors installed in the target environment. In step S14, the server controls the robot to interact based on the target interaction information. Specifically, the server may first send the target interaction information to the robot, and the local actuator of the robot performs the interactive operation. In this way, the processing of multimodal data can be handed over to a server with stronger computing power. Ultimately, it is only necessary to send the target interaction information to the actuator, and the actuator performs the interactive operation, which can reduce the hardware requirements for the robot.

[0103] As another embodiment, when the robot's interactive control method is partially executed locally by the robot and partially executed by the server, any of the aforementioned steps S11-S14 can be executed by the robot, while the remaining steps can be executed by the server. Furthermore, when two adjacent steps are executed locally by the server and the robot, respectively, data from the intermediate processes can be transmitted via the network between the robot and the server. In this manner, if a processing failure occurs locally on the robot, but the network and actuator functions are normal, the processing can be transferred to the server, which is operating normally and has greater computing power. This allows the robot's interactive control function to be implemented even in the event of a local processing failure.

[0104] In addition, in order to enhance transmission privacy and security, in some embodiments, the network between the local robot and the server can be a dedicated network.

[0105] Using this technical solution, after acquiring multimodal data used to control the robot in the target environment, the multimodal data is first perceived and identified to obtain semantic-level environmental features corresponding to the various modalities in the target environment. Then, based on the semantic-level environmental features corresponding to the various modalities and the object semantic relationship network corresponding to the target environment, the target interaction information corresponding to the robot is predicted. Finally, based on this target interaction information, the robot can be controlled to interact. Because semantic-level environmental features take into account the semantic information carried by environmental features, they can reduce ambiguity in the identified environmental features. Therefore, using semantic-level environmental features can improve the accuracy of the predicted target interaction information, thereby improving the accuracy of robot interaction.

[0106] In combination with the foregoing content, it can be seen that in some embodiments, the robot may include flexible screen limb parts, environmental state sensors, visual sensors, voice sensors and force sensors, and the data of multiple modalities include tactile data collected by flexible screen limb parts, environmental state data collected by environmental state sensors, visual data collected by visual sensors, voice data collected by voice sensors and force data collected by force sensors.

[0107] In this embodiment, the robot can interact in a variety of ways. Optionally, when the robot includes a flexible screen limb component, the interaction can be performed by displaying images on the flexible screen limb, or by changing the shape of the flexible screen limb, or by voice output, or by performing multiple interaction methods simultaneously. In this case, step S14 can include a combination of one or more of the following steps:

[0108] In a case where the target interaction information includes an image to be displayed, controlling the flexible screen limb component to display the image to be displayed; or

[0109] In a case where the target interaction information includes position movement information, controlling the robot to move based on the position movement information; or

[0110] In a case where the target interaction information includes limb movement information, controlling the flexible screen limb component of the robot to move based on the limb movement information; or

[0111] In a case where the target interaction information includes voice information, the robot is controlled to output content corresponding to the voice information in audio form.

[0112] It can be understood that in addition to being used to collect tactile data, the flexible screen limb component can also be used for image display. In this case, if the target interaction information includes the image to be displayed, the flexible screen limb component can be controlled to display the image to be displayed.

[0113] Among them, the flexible screen limb component can display the image to be displayed in multiple ways.

[0114] In some embodiments, each flexible screen limb may independently display the entire image or a portion of the image corresponding to the image to be displayed, or all flexible screen limbs may display the entire image or a portion of the image corresponding to the image to be displayed as a whole. Furthermore, the image to be displayed may be displayed in a picture-in-picture format.

[0115] In this embodiment, the image to be displayed is displayed by the flexible screen limb part, which increases the diversity of image display compared to the related art method of displaying only on the robot's chest position with a display screen.

[0116] In other embodiments, it is considered that the robot can have limb interactions, which may cause the position of the flexible screen limb parts to change, such as waving up and down. In this case, if the flexible screen limb parts of the robot always display the image of the same image area, it may cause the image displayed by a certain flexible screen limb part to not be able to be integrated into a whole image with the images displayed by other flexible screen limb parts before and after the position of the flexible screen limb parts changes, resulting in image misalignment. For example, the flexible screen on the robot's arm displayed the image of the upper image area of the image to be displayed at the previous moment. When the robot makes a downward waving action, if it still displays the image of the upper image area of the image to be displayed, then the image displayed by the flexible screen on the arm and the image displayed by the flexible screen on the leg will be misaligned and cannot form a whole image. In this case, in order to avoid display misalignment and improve the robot's interaction effect, the target interaction information may include the position information of the flexible screen limb parts. In this case, please refer to Figure 4 Controlling the flexible screen limb component to display the image to be displayed may specifically include steps S141 to S143. Specifically:

[0117] S141: Obtain an association relationship between each image area included in the image to be displayed and a preset display position.

[0118] It is understandable that any image to be displayed can be divided into different image areas. For example, it can be divided according to pixels, with one or more adjacent pixels as an image display area.

[0119] In this embodiment, the display position corresponding to each image region in the image to be displayed can be preset when the image to be displayed is fully displayed, thereby obtaining an association relationship between each image region included in the image to be displayed and the preset display position. The display position corresponding to the image region is a position in a two-dimensional or three-dimensional coordinate system.

[0120] S142, based on the association relationship between each image area included in the image to be displayed and the preset display position, obtain the target image area corresponding to the position information of each flexible screen limb component.

[0121] S143, controlling the flexible screen limb component to display the image of the target image area.

[0122] After obtaining the target image area corresponding to the position information of each flexible screen limb component, the flexible screen limb component can be controlled to display the image of the target image area.

[0123] Among them, the target interaction information includes the position information of the flexible screen limb parts and can be in multiple forms.

[0124] Optionally, the position information may be position information before the interaction and position information after the interaction. In this case, the flexible screen limb component of the robot may display different images before and after the interaction.

[0125] Optionally, the position information may be real-time position information during the interaction process. In this case, the flexible screen limbs of the robot can dynamically display different images in real time during the interaction process, further improving the effect of the robot's interaction.

[0126] It is understandable that the above process only describes in detail the process of controlling the robot to interact based on the image to be displayed included in the target interaction information. In combination with the above content, it can be seen that the target interaction information can be multimodal target interaction information, that is, target interaction information in multiple forms. That is, in addition to including the image to be displayed, the target interaction information can also include other forms of information, such as voice information to be output, limb movement information, or position movement information. In this case, in addition to controlling the flexible screen limb component to display the image to be displayed, it is also possible to control the voice output module to output voice information, that is, control the robot to output the content corresponding to the voice information in audio form, or control the flexible screen limb component of the robot to move according to the limb movement information, or control the robot to move based on the position movement information, etc. Among them, the position movement information refers to the position information of the overall movement of the robot. The limb movement information refers to the information about the movement of the flexible screen limb component, for example, how the robot limb moves specifically and to which position, etc., and can be any action that the robot can perform, such as shaking hands, dancing, grabbing, etc.

[0127] In some embodiments, when the target interaction information includes multimodal target interaction information, the target interaction information of each modality may also carry corresponding timing information. At this time, step S14 may specifically include the steps of: controlling the robot to interact based on the target interaction information of each modality and the timing information carried by the target interaction information of each modality.

[0128] The timing information refers to the execution time of the target interaction information. In this embodiment, when the target interaction information of each modality carries the timing information, the robot can be controlled to interact based on the target interaction information of each modality and the timing information carried by the target interaction information of each modality.

[0129] Optionally, the timing information carried by target interaction information of different modalities may be different. For example, the robot may execute target interaction information of different modalities sequentially, for example, first waving, then uttering a greeting voice, and then displaying the consultation question in the form of an image to be displayed.

[0130] Optionally, the timing information carried by target interaction information of different modalities may also be the same. For example, the robot may execute target interaction information of different modalities simultaneously, for example, issuing a greeting voice while waving and displaying consultation questions at the same time.

[0131] Optionally, the timing information carried by target interaction information of different modalities may be partially the same and partially different. For example, a greeting voice is issued while waving, and consultation questions are displayed after waving and the greeting voice.

[0132] By adopting the method of the above embodiment, the interactive behavior of the robot can be made closer to the real human interaction process.

[0133] See also Figure 5 , Figure 5 FIG. 1 is a flow chart of a robot interactive control method according to another exemplary embodiment of the present disclosure. Figure 5 As shown, the method includes steps S31 to S36. Specifically:

[0134] S31, acquiring data of multiple modes for controlling a robot in a target environment, wherein data of one mode represents a type of data source data.

[0135] Step S31 is similar to step S11 and will not be described again here.

[0136] S32, optimizing the data of the multiple modes respectively to obtain optimized data of the multiple modes.

[0137] In this embodiment, one or a combination of denoising, filtering, and optimization algorithms can be used to optimize the data of multiple modalities to obtain optimized data of multiple modalities. In this way, the quality of the optimized data can be improved.

[0138] S33, performing fusion correction processing on the optimized data of multiple modalities to obtain fusion corrected data of multiple modalities.

[0139] In this embodiment, the fusion correction process is mainly to ensure the consistency and complementarity of data of multiple modalities.

[0140] For example, in some cases, different sensors select different origins or reference points, which may lead to deviations or omissions in the data collected by different sensors for the same object. In order to more accurately express the data of the same object, the optimized data of multiple modalities can be fused and corrected to obtain fused and corrected data of multiple modalities.

[0141] S34, performing perception and recognition on the fused and corrected data of multiple modalities to obtain semantic-level environmental features corresponding to various modalities in the target environment.

[0142] In this embodiment, after obtaining the fused and corrected data of multiple modalities, perception and recognition can be performed on the fused and corrected data of multiple modalities to obtain semantic-level environmental features corresponding to various modalities in the target environment.

[0143] S35, based on the semantic-level environmental features corresponding to various modalities and the object semantic relationship network corresponding to the target environment, predict the target interaction information corresponding to the robot.

[0144] S36, controlling the robot to interact based on the target interaction information.

[0145] Among them, steps S34-S36 are similar to steps S12-S14 and will not be repeated here.

[0146] By adopting the method of this embodiment, after obtaining the data of multiple modalities for controlling the robot in the target environment, the data of multiple modalities are first optimized separately to obtain the optimized data of multiple modalities, and then the optimized data of multiple modalities are fused and corrected, and then the fused and corrected data of multiple modalities are sensed and recognized, which can improve the quality of the sensed and recognized data and further improve the accuracy of subsequent robot interaction.

[0147] The interactive control method of the robot of the embodiment of the present application is described below with reference to a specific example in a business reception or consulting environment, wherein the method is applied to the robot.

[0148] In a public place, such as an airport, train station or subway station, a robot is used as a receptionist and business consulting service. A user approaches the robot and greets it with a gesture, saying "Hello, robot". The robot sees someone approaching in the video captured by the visual sensor, accompanied by gestures, and receives voice signals, that is, multi-modal data; the robot perceives and recognizes the multi-modal data and obtains the semantic-level environmental features corresponding to various modalities, that is, a male (visual perception recognition gender) around 25 years old (visual perception recognition age) is walking towards the robot happily (visual perception recognition expression), and makes a "shake hands" gesture with his right hand (visual recognition posture), and says "Hello, robot, I have a question to ask." "Question" (speech recognition); based on the semantic-level environmental features corresponding to various modalities and the object semantic relationship network corresponding to the target environment, the robot predicts the target interaction information corresponding to the robot, that is, the target interaction information is the voice output "Good morning, happy to help", the body movement "extending hand to cater to the user's handshake", and the flexible screen body component displays "images with themes such as friendliness, warmth and professional service". Then, the robot can perform the following behaviors: say "Good morning, happy to help", and at the same time extend hand to cater to the user's handshake, and at the same time, the robot's body skin displays images with themes such as friendliness, warmth and professional service.

[0149] After the robot hand and the user's hand hold each other, the robot continues to capture the face of the person seen in the video through the visual sensor, and collects palm prints through the flexible screen limb parts of the hand. Then, it continues to perceive and identify data from multiple modalities, and obtains the semantic-level environmental features corresponding to various modalities, that is, which user is specifically identified by face recognition and palm print recognition; then, based on the semantic-level environmental features corresponding to various modalities and the object semantic relationship network corresponding to the target environment, the robot predicts the corresponding target interaction information of the robot, and the robot can interact based on the target interaction information.

[0150] For example, the flexible display skin can display the user's favorite color theme (for example, the holding hand can be displayed as a delicate hand, or a cartoon hand, or a monster hand), and the flexible screen skin can display a list of questions he may need to consult at the moment (answer question display), and voice output to ask questions (for example, do you want to know the weather conditions at your destination?).

[0151] After completing the inquiry, the robot can release the handshake and stand in a natural standing service posture while continuing to receive voice data (for example, yes), then perceive and recognize the voice data, and predict the target interaction information, and then interact based on the target interaction information.

[0152] For example, the robot can answer the user's question, "It will rain when you reach your destination," by voice output. The flexible screen on its chest displays various umbrellas for purchase, while its waist displays a rainy scene at the destination. It then voice outputs the question, "Is there anything else you can help me with?"

[0153] Finally, the robot continues to receive voice data (for example, the user says: No more) and continues to interact, that is, the robot says: Have a nice day, "Goodbye", and makes a gesture of waving goodbye.

[0154] See also Figure 6 An exemplary embodiment of the present disclosure further provides an interactive control device 400 for a robot, the device comprising:

[0155] The multimodal data acquisition module 410 is used to acquire data of multiple modalities for controlling the robot in the target environment, where data of one modality represents a type of data source data.

[0156] The perception and recognition module 420 is used to perceive and recognize data of multiple modalities to obtain semantic-level environmental features corresponding to various modalities in the target environment.

[0157] The prediction module 430 is used to predict the target interaction information corresponding to the robot based on the semantic-level environment features corresponding to various modalities and the object semantic relationship network corresponding to the target environment.

[0158] The control module 440 is used to control the robot to interact based on the target interaction information.

[0159] Optionally, the device 400 further includes: an optimization processing module for optimizing the data of multiple modalities separately to obtain optimized data of the multiple modalities. A correction processing module for fusing and correcting the optimized data of the multiple modalities to obtain fused and corrected data of the multiple modalities. In this case, the perception and recognition module 420 is further configured to perform perception and recognition on the fused and corrected data of the multiple modalities to obtain semantic-level environmental features corresponding to the various modalities in the target environment.

[0160] Optionally, the prediction module 430 is further configured to input semantic-level environmental features corresponding to various modalities and an object semantic relationship network corresponding to the target environment into a target interaction information prediction model to obtain target interaction information corresponding to the robot.

[0161] Optionally, the prediction module 430 is also used to align the semantic-level environmental features corresponding to various modalities to obtain cross-semantic environmental features corresponding to various modalities; fuse the cross-semantic environmental features corresponding to various modalities to obtain fused environmental features under the target environment; input the fused environmental features and the object semantic relationship network corresponding to the target environment into the target interaction information prediction model to obtain the target interaction information corresponding to the robot.

[0162] Optionally, the device 400 also includes a training module for obtaining multiple sample data, wherein each sample data includes a sample object corresponding to the target environment and an object semantic relationship network; the initial neural network model is trained based on the multiple sample data until the preset training conditions are met, the training is stopped and the target interaction information prediction network is output.

[0163] Optionally, the robot includes flexible screen limb parts, environmental state sensors, visual sensors, voice sensors and force sensors, and the data of multiple modalities include tactile data collected by the flexible screen limb parts, environmental state data collected by the environmental state sensors, visual data collected by the visual sensors, voice data collected by the voice sensors and force data collected by the force sensors.

[0164] Optionally, the target interaction information includes multimodal target interaction information, and the target interaction information of each modality carries corresponding timing information. In this case, the control module 440 is also used to control the robot to interact based on the target interaction information of each modality and the timing information carried by the target interaction information of each modality.

[0165] Optionally, the target interaction information includes an image to be displayed. In this case, the control module 440 is further configured to control the flexible screen limb component to display the image to be displayed.

[0166] Optionally, the target interaction information includes position information of a flexible screen limb component. In this case, the control module 440 includes a first acquisition submodule, a second acquisition submodule, and a control submodule, wherein:

[0167] The first acquisition submodule is configured to acquire an association relationship between each image area included in the image to be displayed and a preset display position.

[0168] The second acquisition submodule is used to acquire the target image area corresponding to the position information of each flexible screen limb component based on the association relationship between each image area included in the image to be displayed and the preset display position.

[0169] The control submodule is used to control the flexible screen limb component to display the image of the target image area.

[0170] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0171] Figure 7 FIG. 5 is a block diagram of an interactive control device 500 for a robot according to an exemplary embodiment. The interactive control device 500 for a robot may be a part of the robot, for example. Figure 7 As shown, the robot interactive control device 500 may include: a processor 501 and a memory 502. The robot interactive control device 500 may also include one or more of a multimedia component 503, an input / output (I / O) interface 504, and a communication component 505.

[0172] The processor 501 is used to control the overall operation of the robot's interactive control device 500 to complete all or part of the steps in the above-mentioned robot interactive control method. The memory 502 is used to store various types of data to support the operation of the robot's interactive control device 500. This data may include, for example, instructions for any application or method operating on the robot's interactive control device 500, as well as application-related data such as contact information, sent and received messages, images, audio, video, etc. The memory 502 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The multimedia component 503 may include a screen and an audio component. The screen may be, for example, a touch screen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signal may be further stored in the memory 502 or transmitted via the communication component 505. The audio component also includes at least one speaker for outputting audio signals. The I / O interface 504 provides an interface between the processor 501 and other interface modules, and the above-mentioned other interface modules can be a keyboard, a mouse, buttons, etc. These buttons can be virtual buttons or physical buttons. The communication component 505 is used for wired or wireless communication between the robot's interactive control device 500 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, 4G, NB-IOT, eMTC, or other 5G, etc., or a combination of one or more of them, is not limited here. Therefore, the corresponding communication component 505 may include: Wi-Fi module, Bluetooth module, NFC module, etc.

[0173] In an exemplary embodiment, the robot's interactive control device 500 can be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to execute the above-mentioned robot's interactive control method.

[0174] In another exemplary embodiment, a computer-readable storage medium including program instructions is also provided. When executed by a processor, the program instructions implement the steps of the above-described robot interactive control method. For example, the computer-readable storage medium may be the aforementioned memory 502 including the program instructions. The program instructions may be executed by the processor 501 of the robot interactive control device 500 to implement the above-described robot interactive control method.

[0175] Figure 8 is a block diagram of a robot interactive control device 600 according to an exemplary embodiment. For example, the robot interactive control device 600 can be provided as a server. Figure 8 The robot interactive control device 600 includes a processor 622, which may be one or more, and a memory 632 for storing a computer program executable by the processor 622. The computer program stored in the memory 632 may include one or more modules, each corresponding to a set of instructions. In addition, the processor 622 may be configured to execute the computer program to perform the above-mentioned robot interactive control method.

[0176] In addition, the interactive control device 600 of the robot may further include a power supply component 626 and a communication component 650. The power supply component 626 may be configured to perform power management of the interactive control device 600 of the robot, and the communication component 650 may be configured to implement communication of the interactive control device 600 of the robot, such as wired or wireless communication. In addition, the interactive control device 600 of the robot may further include an input / output (I / O) interface 658. The interactive control device 600 of the robot may operate based on an operating system stored in the memory 632, such as Windows Server 200. TM , Mac OS X TM, Unix TM , Linux TM etc.

[0177] In another exemplary embodiment, a computer-readable storage medium including program instructions is also provided. When executed by a processor, the program instructions implement the steps of the above-described robot interactive control method. For example, the computer-readable storage medium may be the aforementioned memory 632 including the program instructions. The program instructions may be executed by the processor 622 of the robot interactive control device 600 to implement the above-described robot interactive control method.

[0178] In another exemplary embodiment, a computer program product is further provided. The computer program product includes a computer program executable by a programmable device, and has a code portion for executing the above-mentioned interactive control method of a robot when executed by the programmable device.

[0179] In another exemplary embodiment, a robot is also provided, which includes an actuator, multiple types of sensors, and a processing device connected to the multiple sensors and actuators, the actuator including a flexible screen limb part arranged on the robot body; the processing device is used to: obtain data of multiple modalities, the data of multiple modalities including tactile data collected by the flexible screen limb part, environmental state data collected by multiple types of sensors, visual data, voice data and a combination of at least two of force data; perceive and identify the data of multiple modalities to obtain semantic-level environmental features corresponding to various modalities in the target environment; predict target interaction information corresponding to the robot based on the semantic-level environmental features corresponding to the various modalities and the object semantic relationship network corresponding to the target environment; and control the corresponding actuator to perform interactive operations based on the target interaction information.

[0180] The preferred embodiments of the present disclosure are described in detail above in conjunction with the accompanying drawings. However, the present disclosure is not limited to the specific details of the above embodiments. Within the technical concept of the present disclosure, various simple modifications can be made to the technical solutions of the present disclosure, and these simple modifications all fall within the scope of protection of the present disclosure.

[0181] It should also be noted that the various specific technical features described in the above specific embodiments can be combined in any suitable manner without contradiction. To avoid unnecessary repetition, the present disclosure will not further describe various possible combinations.

[0182] In addition, the various embodiments of the present disclosure may be arbitrarily combined, and as long as they do not violate the concept of the present disclosure, they should also be regarded as the contents disclosed by the present disclosure.

Claims

1. A robot interactive control method, characterized in that: The method comprises: Acquire multiple modal data for controlling the robot in the target environment, where data of one modality represents a type of data source data; Optimizing the data of the multiple modes respectively to obtain optimized data of the multiple modes; Performing fusion correction processing on the optimized data of the multiple modalities to obtain fusion corrected data of the multiple modalities; Perceiving and identifying the data of the multiple modalities to obtain semantic-level environmental features corresponding to the various modalities in the target environment; Predicting target interaction information corresponding to the robot based on semantic-level environmental features corresponding to the various modalities and an object semantic relationship network corresponding to the target environment; Based on the target interaction information, controlling the robot to interact; The perceiving and identifying the data of the multiple modalities to obtain semantic-level environmental features corresponding to the various modalities in the target environment includes: The fused and corrected data of the multiple modalities are subjected to perception and recognition to obtain semantic-level environmental features corresponding to the various modalities in the target environment.

2. The robot interactive control method according to claim 1, characterized in that: The predicting of target interaction information corresponding to the robot based on the semantic-level environment features corresponding to the various modalities and the object semantic relationship network corresponding to the target environment includes: The semantic-level environmental features corresponding to the various modalities and the object semantic relationship network corresponding to the target environment are input into a target interaction information prediction model to obtain target interaction information corresponding to the robot.

3. The robot interactive control method according to claim 1, characterized in that: The predicting of target interaction information corresponding to the robot based on the semantic-level environment features corresponding to the various modalities and the object semantic relationship network corresponding to the target environment includes: Aligning the semantic-level environment features corresponding to the various modalities to obtain cross-semantic environment features corresponding to the various modalities; Fusing the cross-semantic environment features corresponding to the various modalities to obtain fused environment features under the target environment; The fusion environment features and the object semantic relationship network corresponding to the target environment are input into a target interaction information prediction model to obtain target interaction information corresponding to the robot.

4. The robot interactive control method according to claim 2 or 3, characterized in that: The training process of the target interaction information prediction model includes: Acquire a plurality of sample data, wherein each sample data includes a sample object corresponding to the target environment and an object semantic relationship network; The initial neural network model is trained based on the multiple sample data until a preset training condition is met, the training is stopped, and the target interaction information prediction network is output.

5. The robot interactive control method according to any one of claims 1 to 3, characterized in that: The robot includes a flexible screen limb part, an environmental state sensor, a visual sensor, a voice sensor and a force sensor. The data of multiple modalities include tactile data collected by the flexible screen limb part, environmental state data collected by the environmental state sensor, visual data collected by the visual sensor, voice data collected by the voice sensor and force data collected by the force sensor.

6. The robot interactive control method according to claim 5, characterized in that: The target interaction information includes multi-modal target interaction information, each modal target interaction information carries corresponding timing information, and controlling the robot to interact based on the target interaction information includes: Based on the target interaction information of each modality and the timing information carried by the target interaction information of each modality, the robot is controlled to interact.

7. The robot interactive control method according to claim 5, characterized in that: Controlling the robot to interact based on the target interaction information comprises one or more of the following steps: In a case where the target interaction information includes an image to be displayed, controlling the flexible screen limb component to display the image to be displayed; or In a case where the target interaction information includes position movement information, controlling the robot to move based on the position movement information; or In a case where the target interaction information includes limb movement information, controlling the flexible screen limb component of the robot to move based on the limb movement information; or In a case where the target interaction information includes voice information, the robot is controlled to output content corresponding to the voice information in audio form.

8. The robot interactive control method according to claim 7, characterized in that: The target interaction information includes position information of the flexible screen limb component, and when the target interaction information includes an image to be displayed, controlling the flexible screen limb component to display the image to be displayed includes: Acquire the association relationship between each image area included in the image to be displayed and the preset display position; Based on the association relationship between each image area included in the image to be displayed and the preset display position, obtaining the target image area corresponding to the position information of each flexible screen limb component; The flexible screen limb component is controlled to display the image of the target image area.

9. A robot interactive control device, characterized in that: include: A multimodal data acquisition module is used to acquire data of multiple modalities for controlling the robot in the target environment, where data of one modality represents a type of data source data; An optimization processing module is used to optimize the data of multiple modes respectively to obtain optimized data of multiple modes; A correction processing module is used to perform fusion correction processing on the optimized data of multiple modalities to obtain fused and corrected data of multiple modalities; A perception and recognition module, configured to perceive and recognize the data of the multiple modalities to obtain semantic-level environmental features corresponding to the various modalities in the target environment; A prediction module, configured to predict target interaction information corresponding to the robot based on semantic-level environmental features corresponding to the various modalities and an object semantic relationship network corresponding to the target environment; a control module, configured to control the robot to interact based on the target interaction information; The perception and recognition module is also used to perceive and recognize the fused and corrected data of multiple modalities to obtain the semantic-level environmental features corresponding to various modalities in the target environment.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

11. A robot interactive control device, characterized in that: include: a memory having a computer program stored thereon; A processor, configured to execute the computer program in the memory to implement the steps of the method according to any one of claims 1 to 8.

12. A robot, characterized in that: The robot includes an actuator, multiple types of sensors, and a processing device connected to the multiple sensors and the actuator. The actuator includes a flexible screen limb part arranged on the robot body, and the processing device is used to implement the steps of the method described in any one of claims 1-8.

Citation Information

Patent Citations

  • Interaction method and device based on multiple modes, storage medium and intelligent screen equipment

    CN111966212A