Method and device for processing supervised embodied task data in unstructured environments
By selecting the target device and interactive information source based on the network connection status in the humanoid robot and processing visual information in combination with the image and video segmentation model, the problem of humanoid robots identifying and responding to user goals in an unstructured environment is solved, and faster response and higher intelligence are achieved.
Patent Information
- Application Number
- CN202510570492.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2045-05-06
AI Technical Summary
Humanoid robots are difficult to quickly identify and respond to user's temporary choice goals in an unstructured environment, and are less intelligent.
By determining the target task mode based on the network connection status, selecting the target device and interactive information source, combining the image and video segmentation model and user interaction terminal, visually perceived picture information in real time, generating operation instructions, and controlling robot postures and operations.
It improves the response speed and intelligence of humanoid robots in unstructured environments, ensuring more coherent and efficient interactions with users.
Smart Images

Figure CN120080326B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of robotics technology, and in particular to a method and device for processing supervised embodied task data in an unstructured environment. Background Art
[0002] Robotics has made significant progress in recent years, evolving from robotic arms to humanoid robots with voice communication and image recognition capabilities, finding use in a wide range of fields. Due to their anthropomorphic appearance, flexible joints, and friendly appearance, humanoid robots are suitable for performing tasks in place of humans.
[0003] At present, in the task execution scenarios of embodied intelligent robots in unstructured environments, the targets that humanoid robots need to identify are variable. When users temporarily select temporary targets for interaction, humanoid robots may not be able to respond quickly and have low intelligence. Summary of the Invention
[0004] The embodiments of the present application provide a method and device for processing data of supervised embodied tasks in an unstructured environment, which can make the supervised task control of a humanoid robot in an unstructured environment more efficient and accurate.
[0005] In a first aspect, embodiments of the present application provide a method for processing data for a supervised embodied task in an unstructured environment, the method being applied to a controller of a humanoid robot in a supervised task control system, the supervised task control system also including a server and a user interaction terminal, the method comprising:
[0006] Determining a target task mode based on a network connection state, the target task mode being used to constrain selection of a target device and constrain sources of interaction information performed by the humanoid robot, the sources of the interaction information including a server and / or a user interaction terminal, the target device including a server, and the network connection state including connection to the user interaction terminal via a wide area network;
[0007] Acquiring visual perception image information of an external environment, including an unstructured environment, and sending the visual perception image information to a target device;
[0008] receiving a target object and an operation instruction from a server, where the operation instruction is an instruction generated by the server based on the target object for controlling the humanoid robot, and the target object is a designated object represented by an image region obtained by the target device based on visually perceived image information and an image and video segmentation model, and the image and video segmentation model is used to perform image segmentation on the visually perceived image information to obtain the image region;
[0009] The humanoid robot is controlled to reach a target posture based on an operation instruction and initiate a target operation to a target object. The target operation includes interaction based on interaction information from a server or from a user interaction terminal. The target posture includes a posture facing the target object.
[0010] In a possible embodiment, the network connection state further includes being connected to a user interaction terminal via a local area network, the target device further includes the user interaction terminal, and determining the target task mode based on the network connection state includes:
[0011] If the network status network connection status is connected to the user interaction terminal through the wide area network, the target task mode restricts the target device to include the server;
[0012] If the network connection state is connected to the user interaction terminal through a local area network, the target task mode restricts the target device to include the user interaction terminal.
[0013] In a possible embodiment, the target device includes a server, the target object includes a first target object, the operation instruction includes a first operation instruction, and receiving the target object and the operation instruction from the server includes:
[0014] receiving, from the server, a first target object and a first operation instruction obtained according to historical target object information and user-specified feature information;
[0015] Among them, the process of determining the first target object includes the following steps: obtaining multiple potential target objects based on visual perception picture information and an image video segmentation model, the multiple potential target objects being multiple image regions segmented by a server from the visual perception picture information; screening based on the multiple potential target objects and the stored multiple historically determined target objects to determine whether a historical target object is obtained, the historical target object being an object that matches the potential target object and the historically determined target object; if a historical target object exists, sending a prompt message to the user interaction terminal, the prompt message being used to inform the user of the existence of an interactive target object and to inquire about the user's willingness to initiate interaction; using the historical target object as the first target object, or performing a feature matching operation on the historical target object based on user-specified feature information, the user-specified feature information being descriptive information of the intended interactive object entered by the user after receiving the prompt message;
[0016] The process of determining the first operation instruction includes the following steps: performing orientation analysis based on the first target object and visual perception image information to generate an operation instruction, where the operation instruction is a control instruction for controlling the humanoid robot to turn toward the first target object.
[0017] In a possible embodiment, the process of obtaining the user-specified characteristic information includes the following steps:
[0018] Pop up a feature information input box, which is used for user input; obtain the input information entered by the user in the feature information input box, and use the input information as the user-specified feature information;
[0019] In a possible embodiment, the target device includes a user interaction terminal, the target object includes a second target object, the operation instruction includes a second operation instruction, and receiving the target object and the operation instruction from the server includes:
[0020] Receiving a second target object and a second operation instruction obtained from the server according to the designated target object, wherein the designated target object is a target image area generated by the user interaction terminal according to the user's operation, and the second operation instruction is a control instruction generated by the server according to the second target object;
[0021] The method for generating a designated target object includes the following steps: receiving visual perception image information from a humanoid robot and displaying the visual perception image information; detecting a trajectory operation performed by a user based on the displayed visual perception image information, wherein the trajectory operation includes clicking or selecting a frame; and obtaining the designated target object based on the trajectory operation, the visual perception image information, and an image and video segmentation model;
[0022] Among them, the method for generating the second operation instruction includes the following steps: performing orientation analysis based on the second target object and visual perception image information to generate the second operation instruction, and the second operation instruction is a control instruction for controlling the humanoid robot to turn towards the second target object.
[0023] In one possible embodiment, the humanoid robot includes a voice interaction device and a picture interaction device;
[0024] The interaction information from the server or from the user interaction terminal includes: first interaction information preset by the server or second interaction information input by the user through the user interaction terminal;
[0025] The first interaction information includes preset voice information and / or preset expression information; the second interaction information includes voice information and / or expression information; the voice interaction device of the humanoid robot is used to output voice information or preset voice, and the picture interaction device is used to output expression information or preset expression information.
[0026] In one possible embodiment, the user interaction terminal is equipped with an application for processing supervised embodied task data in an unstructured environment. The user interaction terminal includes a display screen, which is used to display visual perception images and to perform human-computer interaction with the user through the application and the display screen.
[0027] In a second aspect, embodiments of the present application provide a device for processing data for a supervised embodied task in an unstructured environment, which is applied to a controller of a humanoid robot in a supervised task control system. The supervised task control system also includes a server and a user interaction terminal. The device includes:
[0028] a mode determination module, which determines a target task mode based on a network connection state, wherein the target task mode is used to constrain selection of a target device and a source of interaction information executed by the humanoid robot, wherein the source of the interaction information includes a server and / or a user interaction terminal, the target device includes a server, and the network connection state includes connection to the user interaction terminal via a wide area network;
[0029] The image acquisition module acquires visual perception image information of the external environment and sends the visual perception image information to the target device, wherein the external environment includes an unstructured environment;
[0030] An information receiving module receives a target object and an operation instruction from a server. The operation instruction is an instruction generated by the server based on the target object for controlling the humanoid robot. The target object is a designated object represented by an image region obtained by the target device based on visually perceived image information and an image and video segmentation model. The image and video segmentation model is used to perform image segmentation on the visually perceived image information to obtain the target object.
[0031] The operation control module controls the humanoid robot to reach the target posture based on the operation instructions and initiate the target operation to the target object. The target operation includes interaction based on the interaction information from the server or from the user interaction terminal.
[0032] In a third aspect, an embodiment of the present application provides a computer-readable storage medium on which a data processing program for a supervised embodied task in an unstructured environment is stored. The data processing program for a supervised embodied task in an unstructured environment includes program instructions, which, when executed by a processor, enable the processor to perform some or all of the steps described in the first aspect.
[0033] In a fourth aspect, an embodiment of the present application provides an electronic device comprising a processor, a memory, a communication interface, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the processor, and the program includes instructions for executing some or all of the steps described in the first aspect of the embodiment of the present application.
[0034] In a fifth aspect, embodiments of the present application provide a computer program product, wherein the computer program product includes a non-transitory computer-readable storage medium storing a computer program, wherein the computer program is operable to cause a computer to perform some or all of the steps described in the first aspect of the embodiments of the present application. The computer program product may be a software installation package.
[0035] By implementing the embodiments of the present application, a target task mode is determined based on the network connection status, the target task mode is used to constrain the selection of the target device, and the source of the interaction information executed by the humanoid robot, the source of the interaction information includes a server and / or a user interaction terminal, the target device includes a server, and the network connection status includes connecting to the user interaction terminal through a wide area network; obtaining visual perception picture information of the external environment, and sending the visual perception picture information to the target device, the external environment includes an unstructured environment; receiving a target object and an operation instruction from the server, the operation instruction is an instruction generated by the server based on the target object for controlling the humanoid robot, the target object is a specified object represented by an image area obtained by the target device based on the visual perception picture information and the image and video segmentation model, the image and video segmentation model is used to perform image segmentation on the visual perception picture information to obtain an image area; controlling the humanoid robot to reach a target posture based on the operation instruction and initiating a target operation to the target object, the target operation includes interacting based on the interaction information from the server or from the user interaction terminal, and the target posture includes a posture facing the target object. Compared with the existing technology in which the robot accompanies the pet according to the preset escort route, actions and audio sent by the wireless terminal, this solution adds real-time control of the robot by the user through the terminal, and the humanoid robot automatically switches to different target devices for transmitting signals when encountering different network connection conditions. This makes the interaction between the humanoid robot and the user in an unstructured environment more coherent, improves the humanoid robot's response speed to user instructions or environment during the embodied service process, and improves the intelligence of the humanoid robot. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the background technology, the drawings required for use in the embodiments of the present invention or the background technology will be described below.
[0037] Figure 1 This is a schematic diagram of the architecture of a supervised task control system provided by an embodiment of the present application;
[0038] Figure 2 This is a block diagram of the functional modules of a humanoid robot provided in an embodiment of the present application;
[0039] Figure 3 1 is a flowchart of a method for processing supervised embodied task data in an unstructured environment provided by an embodiment of the present application;
[0040] Figure 4 This is a schematic diagram of information flow in a method for processing supervised embodied task data in an unstructured environment provided by an embodiment of the present application;
[0041] Figure 5This is a schematic diagram of information flow of another method for processing data of supervised embodied tasks in an unstructured environment provided by an embodiment of the present application;
[0042] Figure 6 This is a schematic diagram of an interface of a user interaction terminal provided in an embodiment of the present application;
[0043] Figure 7 Schematic diagram of a scenario in which an embodiment of the present application provides a method for processing supervised embodied task data in an unstructured environment;
[0044] Figure 8 is a schematic diagram of an interface of a user interaction terminal provided in an embodiment of the present application;
[0045] Figure 9 1 is a schematic diagram of the structure of a device for processing data of supervised embodied tasks in an unstructured environment proposed in an embodiment of the present application;
[0046] Figure 10 This is a schematic structural diagram of a humanoid robot proposed in an embodiment of the present application. DETAILED DESCRIPTION
[0047] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0048] The terms "first," "second," and the like in the specification and claims of this application and the accompanying drawings are used to distinguish between different objects, not to describe a specific order. Furthermore, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or electronic device comprising a series of steps or units is not limited to the listed steps or units, but may, in an optional example, also include steps or units not listed, or may, in an optional example, include other steps or units inherent to the process, method, product, or electronic device.
[0049] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0050] The unstructured environment supervised embodied task data processing method and device provided by the embodiment of the present application can be applied to Figure 1 In the supervised task control system shown, see Figure 1 , Figure 1 This is an architectural diagram of a supervised task control system provided in an embodiment of the present application. The supervised task control system 100 includes a terminal 110, a controller 120 of a humanoid robot, and a server 130. The terminal 110 can transmit data with the controller 120 of the humanoid robot through a network. The terminal 110 refers to a device used by a user, corresponding to the user interaction device in this solution, such as a smart phone, a computer, and other electronic devices on which a control application can be installed.
[0051] In this embodiment, terminal 110 has a display screen and is installed with a user-accessible application. The application features a user interface that allows users to easily view the images transmitted by the humanoid robot and input commands. Other human-machine interaction operations can also be performed, such as clicking buttons or clicking on screens. Terminal 110 can also perform operations such as image segmentation using an image and video segmentation model, and the resulting data can be used to control the humanoid robot. Data can be transmitted between terminal 110 and server 130, as well as between terminal 110 and the humanoid robot's controller 120. The humanoid robot's controller 120 determines different interaction partners based on different network connection states. The humanoid robot can choose to transmit data to either the humanoid robot's controller 120 or the terminal 110. The humanoid robot's controller 120 can control the humanoid robot based on data transmitted by server 130. Furthermore, the humanoid robot's controller 120 can output information, such as voice, based on data transmitted from server 130 or terminal 110.
[0052] Among them, the terminal 110 may include a smart phone (such as an Android phone, an iOS phone, a Windows Phone phone, etc.), a tablet computer, a PDA, a driving recorder, a vehicle-mounted electronic device, a server, a laptop computer, a mobile Internet electronic device (MID, Mobile Internet Devices) or a wearable electronic device (such as a smart watch, a Bluetooth headset), etc. The above are only examples and not exhaustive, including but not limited to the above electronic devices.
[0053] See also Figure 2 , Figure 2 This is a block diagram of the functional modules of a humanoid robot provided in an embodiment of the present application. Figure 2As shown, the functional modules of the humanoid robot may include: a control module, a camera module, a data acquisition module, a data transceiver module, a display module, a sound playback module, etc., which are not limited here.
[0054] The control module, the "brain" of the humanoid robot, is responsible for coordinating the work of the other modules. This corresponds to the controller of the humanoid robot in this solution. It processes data from the various modules and makes decisions based on pre-set programs and algorithms, directing the humanoid robot's movements, behaviors, and task execution. For example, it determines the robot's movements, sounds, and images, as well as processing data acquired by the data acquisition module.
[0055] The camera module is used to capture images and video of the surrounding environment. This visual information can be used for viewing by the user and for screening target objects. For example, when processing data for supervised embodied tasks in unstructured environments, the camera module captures the unstructured environment and transmits it to the terminal for the user to select the target object they want to interact with. The camera module can be installed in the robot's head structure or elsewhere. This solution only describes the case where it is installed in the head structure, but it can only be installed in the head without limitation.
[0056] Among them, the data acquisition module is like the "sensory organ" of the humanoid robot. For example, it can be a sound collector, which can be responsible for collecting external voice data, etc.
[0057] The signal transceiver module is the interface through which the humanoid robot communicates with external control sources or other devices. It receives commands, data, and information from a terminal or server, enabling the humanoid robot to perform tasks according to the control requirements of the server or terminal. The humanoid robot can communicate via a local area network or a wide area network.
[0058] Among them, the display module is a component used by the humanoid robot to display images. It can be set on the head, which is equivalent to the "face" of the humanoid robot. Specifically, it can display facial expressions, making the humanoid robot more friendly and the emotions expressed more three-dimensional.
[0059] Among them, the sound playback module is a component of the humanoid robot used to play sound, just like the "mouth" of the humanoid robot, used to make sounds. Specifically, it can be used to play voice information and interact with the target object.
[0060] Through the mutual collaboration between these modules, the humanoid robot is able to perceive the environment, receive instructions, make decisions, perform corresponding actions and interact with target objects to complete the user's instructions.
[0061] Based on this, the present application provides a method and device for processing supervised embodied task data in an unstructured environment. The present application is described in detail below with reference to the accompanying drawings.
[0062] See also Figure 3 , Figure 3 This is a flow chart of a method for processing data of a supervised embodied task in an unstructured environment provided by an embodiment of the present application. The method is applied to a controller of a humanoid robot in a supervised task control system. The supervised task control system also includes a server and a user interaction terminal. Figure 3 As shown, the method includes the following steps:
[0063] S310, determining the target task mode based on the network connection status, the target task mode is used to constrain the selection of the target device and the source of the interactive information executed by the humanoid robot, the source of the interactive information includes the server and / or the user interaction terminal, the target device includes the server, and the network connection status includes connection with the user interaction terminal through the wide area network.
[0064] Among them, the wide area network includes the World Wide Web communication network and the Global System for Mobile Communications network through the SIM card, or other communication networks through the SIM card, or other wide area network networks, which are not limited here; the network connection status also includes connection with the user interaction terminal through the local area network, and the local area network includes WiFi network, Bluetooth, etc., which are not limited here.
[0065] The network connection status can be determined by prioritizing the LAN connection and, when the LAN connection is poor, switching the network connection to the WAN connection. Specifically, the poor LAN connection status can be determined based on network-related parameters and set thresholds, and the specific method for determining the poor network connection status is not limited here. Of course, the preferred network connection can also be the WAN connection status, which is determined based on actual application and is not limited here.
[0066] Among them, interactive information refers to the information received for interaction with the outside world after the humanoid robot determines the target task mode and obtains the external visual perception picture information and sends it to the target device. The interactive information can be diverse, including but not limited to voice information and expression information. The target task mode is used to constrain which device the humanoid robot transmits data to in real time, and the source of the device that constrains the interactive information executed by the humanoid robot. In this solution, the target device is the device to which the humanoid robot transmits data in real time. Under different network connection states, the target device indicated by the target task mode is different, and the source of the indicated interactive information is also different. The target device can be a server or a user interaction terminal. The network connection state includes connecting to the user interaction terminal through a wide area network. Specifically, the switching of the network connection state may select different networks according to different environments. For example, when the local area network cannot be covered or the signal strength is poor, the wide area network can be switched to communicate. This is not limited here.
[0067] In a possible embodiment, the network connection status also includes connection with the user interaction terminal through a local area network, and the target device also includes the user interaction terminal. Determining the target task mode based on the network connection status includes: if the network connection status is connection with the user interaction terminal through a wide area network, then the target task mode limits the target device to include the server; if the network connection status is connection with the user interaction terminal through a local area network, then the target task mode limits the target device to include the user interaction terminal.
[0068] Specifically, when the network connection status is connected to the user interaction terminal through a wide area network, the target task mode constrains the target device of the humanoid robot to be the server, and the source of the interaction information can be the server or the user interaction terminal; when the network connection status is connected to the user interaction terminal through a local area network, the target task mode constrains the target device of the humanoid robot to be the user interaction terminal, and the source of the interaction information can be the user interaction terminal.
[0069] It can be seen that in this embodiment, by judging the network connection status of the environment in which the humanoid robot is located, the different target devices and interactive information sources of the humanoid robot are determined, so as to achieve a more efficient transmission path and interactive information transmission effect for switching between different network connection states, thereby improving the response efficiency and speed of the humanoid robot.
[0070] S320: Obtain visual perception picture information of an external environment, and send the visual perception picture information to a target device, where the external environment includes an unstructured environment.
[0071] The visually perceived image information is information obtained by the humanoid robot through image and video capture using a camera module, etc. The external environment can be an unstructured environment, which refers to an environment where environmental information is non-fixed and unknowable. Of course, a structured environment can also be used. The specific application environment is determined based on actual conditions and is not limited here. After the humanoid robot collects the visually perceived image information, the controller transmits the acquired visually perceived image information of the external environment to a target device, specifically, to a server or a user interaction terminal.
[0072] Specifically, when connected through a wide area network, the controller of the humanoid robot sends the visual perception picture information to the server; when connected through a local area network, the controller of the humanoid robot sends the visual perception picture information to the user interaction terminal.
[0073] Among them, after receiving the visual perception picture information, the server can first cache it. When the network connection status switches to a local area network connection, the visual perception picture information is synchronized to the user interaction terminal or the cloud for the user to view.
[0074] S330, receiving the target object and operation instructions from the server, where the operation instructions are instructions generated by the server based on the target object for controlling the humanoid robot, and the target object is a designated object represented by an image area obtained by the target device based on visually perceived picture information and an image and video segmentation model, and the image and video segmentation model is used to perform image segmentation on the visually perceived picture information to obtain an image area.
[0075] The image and video segmentation model can be the Segment Anything Model 2 (SAM 2), an image and video segmentation model. SAM 2 can quickly and accurately select and segment objects in any video or image. It not only segments objects but also tracks them in videos, even if they were never seen during training. SAM 2 also features a memory attention module that stores memories of previously identified target objects. Segmentation refers to identifying a specific object in an image or video and distinguishing its pixels from the background. Object segmentation in videos is real-time, allowing the model to rapidly process each frame during video playback, instantly identifying, segmenting, and tracking specific objects. Objects can be specified by clicking on them or by selecting a frame, with no specific restrictions. The image and video segmentation model can be deployed on user interaction terminals and servers. Depending on the target device, different devices will perform image segmentation processing.
[0076] When the target device is a user interaction terminal, the terminal uses the image and video segmentation model to obtain the target object and sends it to the server. When the target device is a server, the server obtains the target object based on the image and video segmentation model. The server generates operation instructions for controlling the humanoid robot based on the obtained target object. The operation instructions can be used to control the humanoid robot to turn, move, change posture, etc. based on the location of the target object, and are not limited here.
[0077] In a possible embodiment, the target device includes a server, the target object includes a first target object, and the operation instruction includes a first operation instruction. Receiving the target object and the operation instruction from the server includes: receiving the first target object and the first operation instruction obtained from the server based on historical target object information and user-specified feature information; wherein a process of determining the first target object includes the following steps:
[0078] 3311. The server obtains a plurality of potential target objects based on the visually perceived picture information and the image and video segmentation model, where the plurality of potential target objects are a plurality of image regions segmented by the server from the visually perceived picture information;
[0079] 3312. The server screens the multiple potential target objects and the multiple stored historically determined target objects to determine whether a historical target object is obtained, where the historical target object is an object that matches the potential target object and the historically determined target object.
[0080] 3313. If a historical target object exists, the server sends a prompt message to the user interaction terminal, which is used to inform the user that an interactive target object exists and inquire about the user's willingness to initiate interaction;
[0081] 3314. The server uses the historical target object as the first target object, or performs a feature matching operation on the historical target object based on the user-specified feature information, where the user-specified feature information is the descriptive information of the intended interactive object input by the user after receiving the prompt information; wherein, the process of determining the first operation instruction includes the following steps: the server generates an operation instruction based on the orientation analysis of the first target object and the visual perception picture information, where the operation instruction is a control instruction for controlling the humanoid robot to turn toward the first target object.
[0082] Among them, see Figure 4 , Figure 4This is a schematic diagram of the information flow of a method for processing data for a supervised embodied task in an unstructured environment provided by an embodiment of the present application. When the humanoid robot selects the target device as the server based on the network connection status and the target task mode, the controller of the humanoid robot sends the visual perception image information to the server after obtaining it. After the server receives the visual perception image information, it uses an image segmentation model to segment the visual perception image information to obtain multiple image regions, each of which represents a potential target object, that is, an object that may be the target object. The segmented potential target objects are then matched and screened based on historically determined target objects to determine whether there are historically determined target objects that match the potential target objects. The above-mentioned historically determined target objects can be obtained based on the memory attention module of SAM 2 or stored in the storage space, which is not limited here. If a historically determined target object that matches the potential target object is screened, a prompt message is sent to the user interaction terminal to inform the user that there is an object that can be interacted with and to ask whether the interaction is to be performed. The prompt message can also be used to prompt the user to enter the feature information of the object they want to interact with. The server then determines the first target object. If it receives feedback from the user confirming the interaction, it uses the historical target object as the first target object. If it receives user-specified feature information from the user, it matches the historical target object with the user-specified feature information. If a matching historical target object is found, it uses the matching historical target object as the first target object. After the server determines the first target object, it performs an orientation analysis based on the first target object and the visually perceived image information to determine the position of the first target object relative to the humanoid robot. It then generates an operational instruction for controlling the humanoid robot to turn toward the first target object and sends the operational instruction and the first target object to the humanoid robot's controller.
[0083] Specifically, when a historical target object is stored, historical target object feature information for the historical target object is also stored. The historical target object feature information includes clothing type, hairstyle, age, gender, facial features, and appearance features. The above historical target object feature information may also include features of non-human objects, including but not limited to object shape, object color, object material, etc. More specifically, the historical target object feature information may be a name or alias entered by the user, such as a person's name, an object name, or a nickname, etc., without limitation. The user can retrieve the stored historical target object and assign features to the historical target object. A matching operation is performed with the historical target object based on the user-specified feature information to obtain a matching historical target object, thereby obtaining the first target object.
[0084] For example, when reviewing historical target objects, the user assigns the feature of "Xiao Ming" to a historical target object. Then, when the user subsequently inputs and matches the specified feature information, the user can enter "Xiao Ming" to match the historical target object corresponding to "Xiao Ming".
[0085] Specifically, when the user does not input user-specified feature information but is willing to initiate interaction, the server directly uses the historical target object as the first target object.
[0086] Specifically, the process of obtaining user-specified characteristic information includes the following steps: the user interaction terminal pops up a characteristic information input box, which is used for the user to input information; the user interaction terminal obtains the input information entered by the user in the characteristic information input box and uses the input information as the user-specified characteristic information.
[0087] The input box may also be a label selection box, for example, providing multiple selectable characteristic word labels, such as "female," "male," "long hair," "middle-aged," and the like. It may also be a typed input or handwritten input box, and the specific input form is not limited here. After receiving the input information, the user interaction terminal packages the user input information as user-specified characteristic information and sends it to the server.
[0088] It can be seen that in this embodiment, when the humanoid robot is in a network connection state of being connected through a wide area network, the server performs screen segmentation based on the current visual perception screen information and matches it with the historical target object, and sends a prompt message to the user interaction terminal to inquire about the user's instructions. The server determines the first target object according to the instructions given by the user. This can achieve the goal of determining the object the user wants to interact with without going through the screen when the network connection state is connected through a wide area network, reducing the time delay caused by a long loading time of the screen due to poor network, making the interaction between the humanoid robot and the user in an unstructured environment more coherent, improving the response speed of the humanoid robot to user instructions or environment during the embodied service process, and improving the intelligence of the humanoid robot.
[0089] In a possible embodiment, the target device includes a user interaction terminal, the target object includes a second target object, and the operation instruction includes a second operation instruction. Receiving the target object and the operation instruction from the server includes: receiving the second target object and the second operation instruction from the server based on a specified target object, wherein the specified target object is a target image area generated by the user interaction terminal based on a user operation, and the second operation instruction is a control instruction generated by the server based on the second target object. The method for generating the specified target object includes the following steps:
[0090] 3321. The server receives visual perception image information from the humanoid robot and displays the visual perception image information;
[0091] 3322. The server detects a user's trajectory operation based on the displayed visual perception screen information, where the trajectory operation includes clicking or selecting a box;
[0092] 3323. The server obtains a designated target object based on trajectory operation, visual perception picture information, and image and video segmentation model; wherein, the method for generating a second operation instruction comprises the following steps: the server generates a second operation instruction by performing orientation analysis based on the second target object and visual perception picture information, and the second operation instruction is a control instruction for controlling the humanoid robot to turn toward the second target object.
[0093] Among them, see Figure 5 , Figure 5 This is an information flow diagram of another method for processing data for supervised embodied tasks in unstructured environments provided by an embodiment of the present application. When a humanoid robot selects a target device as a user interaction terminal based on the network connection status and the target task mode, the humanoid robot's controller transmits the acquired visually perceived image information to the user interaction terminal. The user interaction terminal then displays the visually perceived image information. When the user interaction terminal detects a user's trajectory operation on the real visually perceived image information, such as a click or a selection, it calls an image and video segmentation model, i.e., SAM2, to segment the user's trajectory operation and the visually perceived image information, and to obtain a target object corresponding to the user's trajectory operation. The user interaction terminal then transmits feedback information or user-specified feature information containing the designated target object to a server. After receiving the designated target object, the server determines a second target object and sets the designated target object as the second target object. The server then performs an orientation analysis based on the second target object and the visually perceived image information to determine the position of the second target object relative to the humanoid robot, and generates an operation instruction for controlling the humanoid robot to turn toward the second target object. The server then transmits the operation instruction and the second target object to the humanoid robot's controller.
[0094] Specifically, see Figure 6 , Figure 6 This is a schematic diagram of an interface of a user interaction terminal provided in an embodiment of the present application. Figure 6 As shown, the visual perception picture information displays target objects including A, B, C, D, and E. The visual perception picture information is received at the user interaction terminal, and the visual perception picture is displayed on the interface for the user to watch. When it is detected that the user clicks on D in the displayed visual perception picture information, the image segmentation model is called to identify D, and D is determined to be the designated target object.
[0095] It can be seen that in this embodiment, when the humanoid robot is in a LAN network connection, the user interaction terminal quickly segments the designated target object according to the user's trajectory operation and visual perception picture information and the image and video segmentation model, and the server uses the designated target object given by the user as the second target object. It can achieve the goal of using the image and video segmentation model to quickly respond to the object that the user wants to interact with under the network LAN network connection, making the interaction between the humanoid robot and the user in an unstructured environment more coherent, improving the response speed of the humanoid robot to user instructions or environment during the embodied service process, and improving the intelligence of the humanoid robot.
[0096] S340, controlling the humanoid robot to reach a target posture based on the operation instruction and initiate a target operation to the target object, the target operation including interaction based on interaction information from the server or from the user interaction terminal, and the target posture including a posture facing the target object.
[0097] After receiving an operation instruction from the server, the humanoid robot controller controls the humanoid robot to achieve a target posture according to the operation instruction. Specifically, the target posture can be the humanoid robot facing the target object, or, when the target object is not human, a specific posture, such as picking up the object. The target operation includes outputting interaction information from the server or the user interaction terminal. Specifically, the interaction can be voice interaction, action interaction, or expression interaction, which is not limited here.
[0098] In a possible embodiment, the humanoid robot includes a voice interaction device and a screen interaction device; the interaction information from the server or from the user interaction terminal includes: the first interaction information preset by the server or the second interaction information input by the user through the user interaction terminal; the first interaction information includes preset voice information and / or preset expression information; the second interaction information includes voice information and / or expression information; the voice interaction device of the humanoid robot is used to output voice information or preset voice, and the screen interaction device is used to output expression information or preset expression information.
[0099] Among them, the voice interaction device can be a combination of a speaker and a sound collection device, or it can be a device with playback and collection functions, which is not limited here; the picture interaction device can be a display screen, and the picture interaction device can be set on the head of the humanoid robot to display images or text according to the interactive information, specifically, it can be an expression image.
[0100] When the humanoid robot determines, based on the network connection status, that the target task mode constraint interaction information originates from a server or a user interaction terminal, it may obtain first interaction information from the server, including a preset voice message, or a preset voice message and preset expression message, or a preset expression message. The first interaction information may be a voice message and expression message preset in the server, for interaction. For example, the voice message may be "hello," and the expression message may be "smile" or "surprise." The user may also input second interaction information, including a voice message, a voice message and expression message, or an expression message as interaction message, in real time through the user interaction terminal and send it to the humanoid robot. The controller of the humanoid robot controls the humanoid robot to initiate the target operation, i.e., controls the humanoid robot to output the interaction information. Specifically, the voice message may be played through a speaker and the user's expression message may be displayed on a display screen. When the humanoid robot determines, based on the network connection status, that the target task mode constraint interaction information originates from the user interaction terminal, it obtains the second interaction information from the user interaction terminal. For example, the user may input a voice message, or a voice message and expression message, or an expression message as interaction message, in real time through the user interaction terminal and send it to the humanoid robot. The controller of the humanoid robot controls the humanoid robot to initiate a target operation, that is, controls the humanoid robot to output interactive information, specifically, plays voice information through a speaker, and displays the user's expression information through a display screen.
[0101] The acquisition of expression information may be performed by collecting the user's facial expression through a camera component of the user interaction terminal to obtain facial expression information, and processing the facial expression information into robot expressions. The specific processing process is not limited here.
[0102] It can be seen that in this embodiment, when the humanoid robot is in a local area network connection, it can directly obtain the second interaction information transmitted by the user through the user interaction terminal; when it is in a wide area network connection, it can obtain the second interaction information transmitted by the user through the interaction terminal or the preset first interaction information transmitted from the server, so that the humanoid robot can output interaction information faster in different network connection states without having to wait all the time, making the interaction between the humanoid robot and the user in an unstructured environment more coherent, improving the response speed of the humanoid robot to user instructions or environment during the embodied service process, and improving the intelligence of the humanoid robot.
[0103] For examples, see Figure 7 , Figure 7This is a scenario diagram of a method for processing supervised embodied task data in an unstructured environment provided by an embodiment of the present application. Assuming that a user controls a humanoid robot to go out to perform a task through a user interaction terminal, the user can view the visual perception image information collected by the humanoid robot in real time through the user interaction terminal. However, during the execution of the task, the humanoid robot may need to respond to the unstructured environment of the outside world. For example, when the humanoid robot is in an elevator environment, the network connection status is connected through a wide area network. The visual perception image information viewed by the user through the user interaction terminal may be delayed. The humanoid robot sends the visual perception image information to the server. The server analyzes the existence of a historical target object in the elevator based on the visual perception image information, that is, an object that has been interacted with before. The server then sends a prompt message to the user interaction terminal, asking the user whether to interact and waiting for the user's response. If the user only responds that he can interact but does not specify a target user, the server uses the historical target object as the target object, analyzes the position of the target object in the visual perception image information, outputs a control instruction, and controls the humanoid robot to turn so that the humanoid robot faces the target user. At this point, the system can wait for the user to input voice and emoticon information through the user interaction terminal, or retrieve preset voice and emoticon information from the server, output the voice information through the speaker, and output the emoticon information through the display screen. Assume that the target user is identified as "Mr. Li", and the user transmits the interactive information as the voice message "Hello, Mr. Li" and the emoticon is a smile. Then, the speaker outputs "Hello, Mr. Li" and the display screen shows a smile.
[0104] It can be seen that in this embodiment, the target task mode is determined based on the network connection status, and the target task mode is used to constrain the selection of the target device and the source of the interactive information executed by the humanoid robot. The source of the interactive information includes the server and / or the user interaction terminal, the target device includes the server, and the network connection status includes connecting to the user interaction terminal through a wide area network; obtaining visual perception picture information of the external environment and sending the visual perception picture information to the target device, the external environment includes an unstructured environment; receiving the target object and operation instructions from the server, the operation instructions are instructions generated by the server for controlling the humanoid robot based on the target object, the target object is a specified object represented by the image area obtained by the target device based on the visual perception picture information and the image and video segmentation model, the image and video segmentation model is used to perform image segmentation on the visual perception picture information to obtain the image area; controlling the humanoid robot to reach the target posture based on the operation instruction and initiate the target operation to the target object, the target operation includes interaction based on the interactive information from the server or from the user interaction terminal, and the target posture includes a posture facing the target object. Compared with the existing technology in which the robot accompanies the pet according to the preset escort route, actions and audio sent by the wireless terminal, this solution adds real-time control of the robot by the user through the terminal, and the humanoid robot automatically switches to different target devices for transmitting signals when encountering different network connection conditions. This makes the interaction between the humanoid robot and the user in an unstructured environment more coherent, improves the humanoid robot's response speed to user instructions or environment during the embodied service process, and improves the intelligence of the humanoid robot.
[0105] In one possible embodiment, see Figure 8 , Figure 8 This is a schematic diagram of the interface of the user interaction terminal provided in the embodiment of the present application. Figure 8 As shown, the user interaction terminal is equipped with an application for processing supervised embodied task data in an unstructured environment. The user interaction terminal includes a display screen, which is used to display visual perception images and to perform human-computer interaction with the user through the application and the display screen.
[0106] Among them, the display screen is also used to control the movement of the humanoid robot through a mobile control component. The mobile control component can be a physical device or a virtual button, which is not limited here; the display screen is also used to display multiple function control components, such as emergency shutdown, remote intercom, pause, etc. When a user operation on a function control component is detected, such as a click operation, the function corresponding to the function control component can be executed.
[0107] It can be seen that in this embodiment, controlling the humanoid robot through a user interaction terminal with a display screen can enable the user to better understand the environment in which the humanoid robot is located, and can interact with the target object according to the environment in which the humanoid robot is located, thereby improving the user's control efficiency over the humanoid robot, thereby improving the humanoid robot's response speed to user instructions or the environment during the embodied service process, and improving the intelligence of the humanoid robot.
[0108] See Figure 9 , Figure 9 This is a schematic diagram of the structure of a device for processing data for supervised embodied tasks in an unstructured environment proposed in an embodiment of the present application. The device is applied to the controller of a humanoid robot in a supervised task control system. The supervised task control system also includes a server and a user interaction terminal. The device 900 for processing data for supervised embodied tasks in an unstructured environment includes: a mode determination module 910, an image acquisition module 920, an information receiving module 930, and an operation control module 940, wherein:
[0109] A mode determination module 910 determines a target task mode based on the network connection status. The target task mode is used to constrain the selection of a target device and the source of interaction information performed by the humanoid robot. The source of interaction information includes a server and / or a user interaction terminal. The target device includes a server. The network connection status includes connection to the user interaction terminal via a wide area network.
[0110] The image acquisition module 920 acquires visually perceived image information of the external environment and sends the visually perceived image information to the target device, wherein the external environment includes an unstructured environment;
[0111] Information receiving module 930 receives a target object and an operation instruction from the server. The operation instruction is an instruction generated by the server based on the target object for controlling the humanoid robot. The target object is a designated object represented by an image region obtained by the target device based on visually perceived image information and an image and video segmentation model. The image and video segmentation model is used to perform image segmentation on the visually perceived image information to obtain the target object.
[0112] The operation control module 940 controls the humanoid robot to reach a target posture based on the operation instruction and initiate a target operation to the target object. The target operation includes interaction based on interaction information from the server or from the user interaction terminal.
[0113] In one possible implementation, the information receiving module 930 is specifically configured to:
[0114] receiving, from the server, a first target object and a first operation instruction obtained according to historical target object information and user-specified feature information;
[0115] Among them, the process of determining the first target object includes the following steps: obtaining multiple potential target objects based on visual perception picture information and an image video segmentation model, the multiple potential target objects being multiple image regions segmented by a server from the visual perception picture information; screening based on the multiple potential target objects and the stored multiple historically determined target objects to determine whether a historical target object is obtained, the historical target object being an object that matches the potential target object and the historically determined target object; if a historical target object exists, sending a prompt message to the user interaction terminal, the prompt message being used to inform the user of the existence of an interactive target object and to inquire about the user's willingness to initiate interaction; using the historical target object as the first target object, or performing a feature matching operation on the historical target object based on user-specified feature information, the user-specified feature information being descriptive information of the intended interactive object entered by the user after receiving the prompt message;
[0116] The process of determining the first operation instruction includes the following steps: performing orientation analysis based on the first target object and visual perception image information to generate an operation instruction, where the operation instruction is a control instruction for controlling the humanoid robot to turn toward the first target object.
[0117] In a possible implementation, the information receiving module 930 is further configured to:
[0118] Receiving a second target object and a second operation instruction obtained from the server according to the designated target object, wherein the designated target object is a target image area generated by the user interaction terminal according to the user's operation, and the second operation instruction is a control instruction generated by the server according to the second target object;
[0119] The method for generating a designated target object includes the following steps: receiving visual perception image information from a humanoid robot and displaying the visual perception image information; detecting a trajectory operation performed by a user based on the displayed visual perception image information, wherein the trajectory operation includes clicking or selecting a frame; and obtaining the designated target object based on the trajectory operation, the visual perception image information, and an image and video segmentation model;
[0120] Among them, the method for generating the second operation instruction includes the following steps: performing orientation analysis based on the second target object and visual perception image information to generate the second operation instruction, and the second operation instruction is a control instruction for controlling the humanoid robot to turn towards the second target object.
[0121] It is worth noting that the specific functional implementation of the unstructured environment supervised embodied task data processing device 900 can be found in the description of the unstructured environment supervised embodied task data processing method described above, for example, the pattern determination module 910 is used to implement the relevant content of executing S310. The various units or modules in the unstructured environment supervised embodied task data processing device 900 can be individually or completely merged into one or several other units or modules to form a structure, or one (or some) of the units or modules can be further divided into multiple functionally smaller units or modules to form a structure, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present invention. The above-mentioned units or modules are divided based on logical functions. In actual applications, the functions of one unit (or module) are implemented by multiple units (or modules), or the functions of multiple units (or modules) are implemented by one unit (or module).
[0122] It can be seen that the unstructured environment described in the embodiment of the present application has a supervised embodied task data processing device, which determines the target task mode based on the network connection status. The target task mode is used to constrain the selection of the target device and the source of the interactive information executed by the humanoid robot. The source of the interactive information includes the server and / or the user interaction terminal. The target device includes the server. The network connection status includes connecting to the user interaction terminal through a wide area network; obtaining visual perception picture information of the external environment and sending the visual perception picture information to the target device. The external environment includes an unstructured environment; receiving the target object and operation instructions from the server. The operation instructions are instructions generated by the server for controlling the humanoid robot based on the target object. The target object is a specified object represented by the image area obtained by the target device based on the visual perception picture information and the image and video segmentation model. The image and video segmentation model is used to perform image segmentation on the visual perception picture information to obtain the image area; controlling the humanoid robot to reach the target posture based on the operation instruction and initiate the target operation to the target object. The target operation includes interaction based on the interactive information from the server or from the user interaction terminal. The target posture includes a posture facing the target object. Compared with the existing technology in which the robot accompanies the pet according to the preset escort route, actions and audio sent by the wireless terminal, this solution adds real-time control of the robot by the user through the terminal, and the humanoid robot automatically switches to different target devices for transmitting signals when encountering different network connection conditions. This makes the interaction between the humanoid robot and the user in an unstructured environment more coherent, improves the humanoid robot's response speed to user instructions or environment during the embodied service process, and improves the intelligence of the humanoid robot.
[0123] See also Figure 10 , Figure 10 This is a schematic diagram of the structure of a humanoid robot proposed in an embodiment of the present application. Figure 10As shown, the humanoid robot 1000 includes a processor 1010 , a memory 1020 , a communication interface 1030 , and one or more programs 1021 . The one or more programs 1021 are stored in the memory 1020 and are configured to be executed by the processor 1010 .
[0124] The processor 1010, the memory 1020, and the communication interface 1030 are interconnected and perform communication between them.
[0125] The memory 1020 can be a volatile memory such as a dynamic random access memory (DRAM) or a non-volatile memory such as a mechanical hard disk. The memory 1020 is used to store a set of executable program codes, and the processor 1010 is used to call one or more programs 1021 stored in the memory 1020 to execute some or all of the steps of any of the methods for processing data for supervised embodied tasks in unstructured environments as described in the above embodiments.
[0126] An embodiment of the present application also provides a computer storage medium, wherein the computer storage medium stores a computer program for electronic data exchange, and the computer program enables a computer to execute part or all of the steps of any method described in the above method embodiments, and the above computer includes an electronic device.
[0127] The present application also provides a computer program product comprising a non-transitory computer-readable storage medium storing a computer program, wherein the computer program is operable to cause a computer to perform some or all of the steps of any of the methods described in the above method embodiments. The computer program product may be a software installation package, and the computer may comprise an electronic device.
[0128] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.
[0129] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0130] In the several embodiments provided in this application, it should be understood that the disclosed devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical or other forms.
[0131] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0132] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0133] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, or the part that contributes to the existing technology, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a memory and includes a number of instructions for enabling a computer electronic device (which can be a personal computer, electronic device, or network electronic device, etc.) to execute all or part of the steps of the above-mentioned methods in each embodiment of the present application. The aforementioned memory includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk, etc. Various media that can store program code.
[0134] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments can be completed by a program instructing related hardware. The program can be stored in a computer-readable memory, which may include a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0135] The above is a detailed introduction to the embodiments of the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core idea of the present application. At the same time, for those skilled in the art, based on the ideas of the present application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A method for processing supervised embodied task data in an unstructured environment, characterized in that: A controller for a humanoid robot in a supervised task control system, wherein the supervised task control system further includes a server and a user interaction terminal, and the method includes: Determining a target task mode based on a network connection state, the target task mode being used to constrain selection of a target device, the target device being a device to which the controller of the humanoid robot sends data in real time, the target device including the server, and the network connection state including being connected to the user interaction terminal via a wide area network; Acquiring visual perception picture information of an external environment, and sending the visual perception picture information to the target device, wherein the external environment includes an unstructured environment; receiving a target object and an operation instruction from the server, wherein the operation instruction is an instruction generated by the server based on the target object for controlling the humanoid robot, the target object being a designated object represented by an image region obtained by the target device based on visually perceived image information and an image and video segmentation model, the image and video segmentation model being used to perform image segmentation on the visually perceived image information to obtain the image region; The humanoid robot is controlled to reach a target posture based on the operation instruction and initiate a target operation to the target object, wherein the target operation includes interacting based on interaction information from the server or from the user interaction terminal, and the target posture includes a posture facing the target object.
2. The method according to claim 1, characterized in that The network connection state further includes being connected to the user interaction terminal via a local area network, the target device further includes a user interaction terminal, and determining the target task mode based on the network connection state includes: If the network connection state is connected to the user interaction terminal via a wide area network, the target task mode restricts the target device to include the server; If the network connection status is connected to a user interaction terminal via a local area network, the target task mode limits the target device to include the user interaction terminal.
3. The method according to claim 2, characterized in that The target device includes the server, the target object includes a first target object, the operation instruction includes a first operation instruction, and receiving the target object and the operation instruction from the server includes: receiving, from the server, the first target object and the first operation instruction obtained according to historical target object information and user-specified feature information; Among them, the process of determining the first target object includes the following steps: obtaining multiple potential target objects based on the visual perception picture information and the image video segmentation model, and the multiple potential target objects are multiple image areas segmented by the server from the visual perception picture information; screening according to the multiple potential target objects and the stored multiple historically determined target objects, and judging whether a historical target object is obtained, and the historical target object is an object that matches the potential target object and the historically determined target object; if the historical target object exists, sending a prompt message to the user interaction terminal, and the prompt message is used to inform the user that there is an interactive target object and to inquire about the user's willingness to initiate interaction; using the historical target object as the first target object, or, performing a feature matching operation on the historical target object based on user-specified feature information, to obtain the first target object, and the user-specified feature information is descriptive information of the intended interactive object entered by the user after receiving the prompt message; The process of determining the first operation instruction includes the following steps: performing orientation analysis based on the first target object and the visual perception image information to generate the operation instruction, and the operation instruction is a control instruction for controlling the humanoid robot to turn toward the first target object.
4. The method according to claim 3, characterized in that The process of obtaining the user-specified feature information includes the following steps: A feature information input box pops up, where the user can input information. The input information input by the user in the feature information input box is acquired, and the input information is used as the user-specified feature information.
5. The method according to claim 2, characterized in that The target device includes the user interaction terminal, the target object includes a second target object, the operation instruction includes a second operation instruction, and the receiving of the target object and the operation instruction from the server includes: receiving, from the server, the second target object and the second operation instruction obtained according to the designated target object, wherein the designated target object is a target image area generated by the user interaction terminal according to the user's operation, and the second operation instruction is a control instruction generated by the server according to the second target object; The method for generating the designated target object includes the following steps: receiving the visual perception image information from the humanoid robot and displaying the visual perception image information; detecting a trajectory operation performed by the user based on the displayed visual perception image information, wherein the trajectory operation includes clicking or selecting a frame; and obtaining the designated target object based on the trajectory operation, the visual perception image information, and the image and video segmentation model. Among them, the method for generating the second operation instruction includes the following steps: performing orientation analysis based on the second target object and the visual perception picture information to generate the second operation instruction, and the second operation instruction is a control instruction for controlling the humanoid robot to turn toward the second target object.
6. The method according to claim 1, characterized in that The humanoid robot includes a voice interaction device and a picture interaction device; The interaction information from the server or from the user interaction terminal includes: first interaction information preset by the server or second interaction information input by the user through the user interaction terminal; The first interaction information includes preset voice information and / or preset expression information; the second interaction information includes voice information and / or expression information; the voice interaction device of the humanoid robot is used to output the voice information or the preset voice, and the screen interaction device is used to output the expression information or the preset expression information.
7. The method according to any one of claims 1 to 6, characterized in that The user interaction terminal is equipped with an application for processing supervised embodied task data in an unstructured environment. The user interaction terminal includes a display screen, which is used to display the visual perception picture and to perform human-computer interaction with the user through the application and the display screen.
8. A device for processing data of supervised embodied tasks in an unstructured environment, characterized in that: A controller for a humanoid robot in a supervised task control system, wherein the supervised task control system further includes a server and a user interaction terminal, and the device includes: a mode determination module, which determines a target task mode based on a network connection state, wherein the target task mode is used to constrain selection of a target device, wherein the target device is a device to which the controller of the humanoid robot sends data in real time, wherein the target device includes the server, and wherein the network connection state includes connection with the user interaction terminal via a wide area network; A picture acquisition module, which acquires visually perceived picture information of an external environment and sends the visually perceived picture information to the target device, wherein the external environment includes an unstructured environment; an information receiving module, receiving a target object and an operation instruction from the server, wherein the operation instruction is an instruction generated by the server based on the target object for controlling the humanoid robot, and the target object is a designated object represented by an image region obtained by the target device based on visually perceived image information and an image and video segmentation model, wherein the image and video segmentation model is used to perform image segmentation on the visually perceived image information to obtain the target object; An operation control module controls the humanoid robot to reach a target posture based on the operation instruction and initiate a target operation to the target object, wherein the target operation includes interacting based on interaction information from the server or from the user interaction terminal.
9. A computer-readable storage medium, characterized in that A data processing program for supervised embodied tasks in an unstructured environment is stored, and includes execution instructions. When a processor of an electronic device executes the execution instructions, the processor performs the method according to any one of claims 1 to 7.
10. An electronic device, characterized in that: The method comprises a processor, a memory, a communication interface, and one or more programs, wherein the one or more programs are stored in the memory; when the processor executes the one or more programs stored in the memory, the processor executes the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Robot assisted interaction system and method thereof
CN110136499A
Video target segmentation method and device, equipment and medium
CN113763385A