Robot control method and device, control equipment and storage medium
By obtaining the robot's task commands and static scene images, and using the visual-spatial reasoning model and task generator to generate action sequence instructions, the control fragmentation problem of multimodal models in complex scenarios is solved, and efficient and accurate robot control is achieved.
Patent Information
- Application Number
- CN202511128834.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-08-13
AI Technical Summary
Existing multimodal models in robot control overly rely on implicit spatial knowledge, resulting in fragmented reasoning chains and making it difficult to achieve efficient and accurate multi-robot collaborative control in complex scenarios.
By obtaining the task commands and static scene images of the target robot, a preset visual-spatial reasoning model is used to generate the task action sequence and its object space state information, and the preset task generator is combined to generate action sequence instructions to ensure that the actions performed by the robot are consistent with the real physical environment.
It improves the efficiency and accuracy of robot control, solves the limitations of traditional models that rely on pre-training data, avoids the accumulation of decision-making errors, and ensures the physical consistency of multi-step decisions.
Smart Images

Figure CN120620236A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of intelligent control technology, and in particular to a control method, apparatus, control device, and storage medium for a robot. Background Art
[0002] With the rapid development of robotics technology, robot application scenarios are becoming increasingly complex, gradually expanding from structured industrial environments to open dynamic environments. Single modal information is difficult to cope with environmental uncertainty and task diversity. Therefore, multimodal model robot control technology has become a research hotspot.
[0003] At present, multimodal models in the field of robot control focus on integrating multi-source perception data such as vision, touch, voice, and inertial measurement, and using intelligent algorithms such as deep learning and reinforcement learning to achieve adaptive human-computer interaction and autonomous decision-making capabilities, thereby promoting robots to perform tasks more intelligently in complex environments.
[0004] However, when dealing with spatial reasoning tasks, this model suffers from a significant drawback: an overreliance on implicit spatial knowledge. This often leads to reasoning that is divorced from the underlying geometric relationships between objects, resulting in a fragmented reasoning chain. When faced with tasks that consider both object functionality and topological constraints, this model's sequential decision-making capabilities easily break down, ultimately resulting in low efficiency and accuracy in multi-robot collaborative control, making it difficult to meet the practical application requirements in complex scenarios. Summary of the Invention
[0005] The purpose of this application is to address the deficiencies in the above-mentioned prior art and provide a robot control method, device, control equipment and storage medium to improve the control efficiency and control accuracy of the robot to meet the control requirements in complex scenarios.
[0006] To achieve the above objectives, the technical solutions adopted in the embodiments of the present application are as follows: In a first aspect, an embodiment of the present application provides a method for controlling a robot, the method comprising: Obtaining an input task command for a target robot and a static scene image of a task scene captured by the target robot based on the task command; Based on the static scene image and the task command, a preset visual-spatial reasoning model is used to generate a task action sequence corresponding to the task command, and object spatial state information corresponding to each action in the task action sequence; wherein the object spatial state information corresponding to each action is used to represent spatial state change information of each object in the static scene image under each action; According to the task action sequence and the object space state information corresponding to each action, a preset task generator is used to generate an action sequence instruction corresponding to the task command; According to each action instruction in the action sequence instruction, the target robot is controlled in sequence to perform the corresponding action.
[0007] Optionally, the preset visual-spatial reasoning model includes: a graphic node extractor, a visual aid extractor, a scene image generator, and a spatial reasoning module. The preset visual-spatial reasoning model is used to generate, based on the static scene image and the task command, a task action sequence corresponding to the task command, and object space state information corresponding to each action in the task action sequence, including: Using the graphic node extractor, extracting nodes from the static scene image to obtain node identification information in the static scene image; Using the visual aid extractor, extracting features from the node identification information to obtain scene visual features corresponding to the static scene image; The scene image generator is used to generate a dynamic structured scene image corresponding to the static scene image according to the scene visual features, wherein the dynamic structured scene image includes: attribute information of each object and the spatial position relationship between the objects; The spatial reasoning module is used to perform spatial reasoning on the dynamic structured scene image according to the task command to obtain the task action sequence and the object spatial state information corresponding to each action in the task action sequence.
[0008] Optionally, the visually assisted extractor includes: an object detection network and a depth estimation network; and using the visually assisted extractor to perform feature extraction on the node identification information to obtain scene visual features corresponding to the static scene image includes: Using the object detection network, performing object detection based on the node identification information to generate attribute information of each object in the static scene image; The depth estimation network is used to perform depth estimation according to the node identification information to obtain depth estimation information of the task scene, and the scene visual features include: attribute information of each object and the depth estimation information.
[0009] Optionally, the using the graphic node extractor to extract nodes from the static scene image to obtain node identification information in the static scene image includes: If the task command is a reorganization task command, the graphic node extractor is used to extract nodes from the static scene image to obtain identification information of each node in the static scene image. The identification information of each node includes: the space occupancy information of each node. The node identification information includes: the attribute information of each node and the space occupancy information of each node.
[0010] Optionally, the method further includes: Obtaining an input first question for the target robot and a first scene image captured by the target robot based on the first question; Based on the first question and the first scene image, using the preset visual-spatial reasoning model, generating first visual-spatial reasoning information of the first scene image under the first question; the first visual-spatial reasoning information includes: spatial state information of each object in the first scene image; Generate first answer information corresponding to the first question using the preset task generator according to the first visual-spatial reasoning information; Control the target robot to play the first answer information.
[0011] Optionally, the method further includes: Obtaining sample video data and preset question-answer pairs corresponding to the sample video data in an input preset spatial data set; generating, based on the second question in the preset question-answer pair and the second scene image corresponding to the second question in the sample video data, second visual-spatial reasoning information of the second scene image under the second question using the preset visual-spatial reasoning model; the second visual-spatial reasoning information including spatial state information of each object in the second scene image; Generate second answer information corresponding to the second question using the preset task generator according to the second visual-spatial reasoning information; The preset visual-spatial reasoning model is tested according to the second answer information and the preset answer corresponding to the second question in the preset question-answer pair.
[0012] Optionally, before obtaining the sample video data in the input preset spatial data set and the preset question-answer pairs corresponding to the sample video data, the method further includes: Acquire the sample video data and the preset task category corresponding to the sample video data; Based on the preset task category and the sample video data, a preset question-answer pair corresponding to the sample video data is generated.
[0013] In a second aspect, another embodiment of the present application provides a control device for a robot, the device comprising: An acquisition module is used to acquire an input task command for a target robot and a static scene image of the task scene in which the target robot is located, which is collected based on the task command; a generation module, configured to generate, based on the static scene image and the task command, a task action sequence corresponding to the task command and object spatial state information corresponding to each action in the task action sequence using a preset visual-spatial reasoning model; wherein the object spatial state information corresponding to each action is used to represent spatial state change information of each object in the static scene image under each action; A generating module, configured to generate an action sequence instruction corresponding to the task command using a preset task generator according to the task action sequence and the object space state information corresponding to each action; The control module is used to control the target robot to perform corresponding actions in sequence according to each action instruction in the action sequence instruction.
[0014] In the third aspect, another embodiment of the present application provides a robot control device, comprising: a processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor. When the robot control device is used, the processor and the memory communicate through the bus, and the processor executes the machine-readable instructions to perform the steps of the robot control method as described in any one of the first aspects above.
[0015] In a fourth aspect, another embodiment of the present application provides a storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the robot control method as described in any one of the above-mentioned first aspects are executed.
[0016] The beneficial effects of this application are: The present application provides a robot control method, apparatus, control device, and storage medium. These methods utilize a task command input for a target robot and a static scene image of the target robot's task scene captured based on the task command. The method utilizes multiple task commands, combined with static images, to achieve cross-modal alignment, ensuring the acquisition of task-related scene information, avoiding redundant computations associated with full-scene scanning, and improving processing efficiency. Based on the static scene image and task command, a preset visual-spatial reasoning model is employed to generate a task action sequence corresponding to the task command, as well as object spatial state information corresponding to each action in the task action sequence. By explicitly encoding object states and spatial relationships through a dynamic scene graph, zero-shot spatial reasoning is achieved, addressing the limitations of traditional models that rely on pre-trained data. Based on the task action sequence and the object spatial state information corresponding to each action, a preset task generator is employed to generate action sequence instructions corresponding to the task command. The method converts the abstract action sequence into specific instructions executable by the robot, ensuring that the instructions are consistent with the real physical environment. Based on each action instruction in the action sequence instruction, the target robot is sequentially controlled to perform the corresponding action. This application uses a preset task generator through the object space state information and the object space state information corresponding to each action to generate action sequence instructions corresponding to the task command, solve the problem of reasoning chain fragmentation, ensure the physical consistency of multi-step decision-making, ensure that subsequent actions correct early spatial constraint conflicts, avoid error accumulation, and thus improve the control efficiency and accuracy of the robot. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.
[0018] Figure 1 A robot control scene diagram provided for an embodiment of the application; Figure 2 A schematic flow chart of a robot control method provided in an embodiment of the present application; Figure 3 A schematic diagram of a flow chart for determining object spatial state information in a robot control method provided in an embodiment of the present application; Figure 4 A schematic diagram of a process for determining scene visual features in a robot control method provided in an embodiment of the present application; Figure 5 A schematic flow chart of another robot control method provided in an embodiment of the present application; Figure 6A schematic diagram of a flow chart of testing a preset space model in a robot control method provided in an embodiment of the present application; Figure 7 A process for determining a preset spatial data set in a robot control method provided in an embodiment of the present application; Figure 8 A schematic diagram of a robot control architecture provided in an embodiment of the present application; Figure 9 A schematic structural diagram of a robot control device provided in an embodiment of the present application; Figure 10 A schematic structural diagram of a robot control device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0019] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. It should be understood that the drawings in the present application only serve the purpose of illustration and description and are not used to limit the scope of protection of the present application. In addition, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate the operations implemented according to some embodiments of the present application. It should be understood that the operations of the flowcharts can be implemented out of sequence, and steps without logical context can be reversed or implemented simultaneously. In addition, those skilled in the art, under the guidance of the contents of this application, can add one or more other operations to the flowchart, or remove one or more operations from the flowchart.
[0020] In addition, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application generally described and shown in the drawings here can be arranged and designed in various configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present application.
[0021] It should be noted that the term "comprising" will be used in the embodiments of the present application to indicate the existence of the features declared thereafter, but does not exclude the addition of other features.
[0022] At present, a large multimodal language model is usually used to control a robot based on the recognized multimodal information, so that the robot can base language reasoning on visual perception, thereby controlling the robot. Specifically, the multimodal model needs to interpret the scene semantics from the visual data or generate a plan based on object detection. However, although the existing multimodal model has the ability to encode visual data, it will incorrectly present the relationship between objects without a clear geometric basis. When it is necessary to consider the functional attributes and topological constraints of the objects at the same time, the sequential decision-making usually fails. To this end, the present application provides a robot control method, which obtains an input task command for a target robot and a static scene image of the task scene in which the target robot is located based on the task command. According to the static scene image and the task command, a preset visual space reasoning model is used to generate a task action sequence corresponding to the task command and the object space state information corresponding to each action in the task action sequence; according to the task action sequence and the object space state information corresponding to each action, a preset task generator is used to generate an action sequence instruction corresponding to the task command. According to each action instruction in the action sequence instruction, the target robot is sequentially controlled to perform the corresponding action. This application significantly improves the robot's understanding ability through object space state information, reduces decision-making errors in the robot control process through a preset task generator, and thus improves the accuracy of robot control.
[0023] As follows, with reference to a plurality of drawings, a use scenario of a robot control method provided by an embodiment of the present application is described. Figure 1 A robot control scene diagram provided for the application embodiment, such as Figure 1 As shown, the robot can be an articulated robot, including a base, an upper arm, a lower arm, and a gripper provided at the front end of the lower arm for operation. The upper arm, the lower arm and the gripper are connected by joints to form a multi-degree-of-freedom structure similar to a human arm. The robot can also be a mobile robot with a mobile chassis that can move autonomously or semi-autonomously in the environment. Or a robot of any other structure, the embodiment of the present application does not limit this. When the target robot obtains the input task command for the target robot, it collects the static scene graph of the scene based on the task command through the visual sensor, and generates a task action sequence and object space state information according to the task command and the static scene graph, and generates an action sequence instruction according to the task action sequence and the object space state information corresponding to each action; according to each action instruction in the action sequence instruction, the target robot is controlled to perform the corresponding action in sequence. Among them, the object space state information is used to characterize the spatial state change information of each object in the static scene image under each action, for example, Figure 1The objects in the image may include objects such as tables, cups, plates or bowls on the desktop, and the spatial state change information of the objects may include the attribute information, position information and relationship information of the objects. For example, the object is a green cup, the coordinates of the cup are (3, 5, 6), and the attributes of the cup may be green, the cup is on a red plate, etc. When the task command is "put the green cup on the pink plate", the corresponding action sequence instruction is generated to control the target robot to perform the corresponding action so that the green cup is placed on the pink plate. It is worth noting that this application only uses Figure 1 For example, the specific robot usage scenario can also be any scene of moving objects such as Lego splicing, so Figure 1 The tableware can also be Legos of different shapes and colors, and this embodiment of the present application does not limit this.
[0024] As follows, the robot control method provided in the embodiment of the present application is described in conjunction with the accompanying drawings. The robot control method is applied to the control device of the robot. The control device of the robot includes a processor and a memory. The processor can be a local controller of the robot or a cloud server. The processor is used to execute the robot motion control model training method and the robot control method. Figure 2 A flow chart of a robot control method provided in an embodiment of the present application is shown as follows: Figure 2 As shown, the method includes: Step 201: Obtain an input task command for a target robot and a static scene image of the task scene collected by the target robot based on the task command.
[0025] Among them, the task command can be input through text or obtained through voice recognition of input text. The static scene graph is the content of a fixed scene and cannot reflect the changes of entities in the scene. The static scene graph can include: RGB (Red, Green, Blue) image and depth image. The RGB image records the color, texture and other visual information of the task scene, and the depth image records the depth value of each point in the task scene to the camera to form spatial geometric information. Specifically, the target robot is equipped with a color image and depth information (RGB-Depth, abbreviated as RGB-D) camera, and the RGB image and depth image are obtained through the RGB-D camera.
[0026] Optionally, the input text or voice information determines the task command for the target robot. When the task command is obtained, the target entity corresponding to the question or instruction in the task command is determined according to the task command, and a static scene image of the task scene where the target entity is located is obtained.
[0027] For example, when the input voice is recognized as "What are the cutlery below the blue cup?", the task command is determined to be "Identify the parameters below the blue cup", the task scene centered on the blue cup is determined from the current scene, and a static scene map of the task scene is obtained through the RGB-D camera installed on the robot.
[0028] Step 202: Based on the static scene image and the task command, a preset visual-spatial reasoning model is used to generate a task action sequence corresponding to the task command and object space state information corresponding to each action in the task action sequence.
[0029] The preset visual-spatial reasoning model is a pre-trained model obtained by processing multimodal information. It is used to convert static scene images and task commands into executable action sequences. The object spatial state information corresponding to each action is used to represent the spatial state change information of each object in the static scene image under each action. The task action sequence can be the steps of operating the objects in the task scene under the task command, for example, it can be "grab the green cup, move it above the pink plate, and put it down vertically." The object spatial state information is the spatial state change information of each object under each action, including the type, color, position, and positional relationship of each object with other objects.
[0030] Optionally, based on the static scene image and task command, a preset visual-spatial reasoning model is used to determine the object spatial state information corresponding to each node, and based on the spatial state information, a task action sequence corresponding to the task command and the object spatial state information corresponding to each action in the task action sequence are generated.
[0031] Step 203: Based on the task action sequence and the object space state information corresponding to each action, a preset task generator is used to generate an action sequence instruction corresponding to the task command.
[0032] Among them, the preset task generator is used to sort the task action sequence according to the task action sequence and the object space state information corresponding to each action, so that each action after sorting can be carried out smoothly. Each step of reasoning is verified based on geometric rules to avoid generating hallucination actions and improve the success rate of task execution. For example, if the task action sequence is directly executed through the preset task generator, it will cause a certain node to form a hallucination, for example, Figure 1 After a certain action is executed, the green cup floats in the air without support.
[0033] Optionally, based on the task action sequence and the object space state information corresponding to each action, it is determined whether the object space state information corresponding to each action in the task action sequence is continuous state information. If it is continuous state information, an action sequence instruction is generated. If it is not continuous sequence information, an action sequence determined from the object space state information corresponding to each action according to a preset task planner is used as the action sequence instruction.
[0034] Optionally, based on the task action sequence and the object space state information corresponding to each action, a preset task generator is used to infer the task action sequence and the object space state information corresponding to each action, so that the task action sequence is consistent with the logic, thereby generating action sequence instructions corresponding to the task command.
[0035] For example, when the task command is "Move the green cup from the red plate to the pink plate", the task action sequence in the task scene is "Grab the green cup, move it above the pink plate, and put it down vertically". The object space state information corresponding to grabbing the green cup is "red plate A (position (0,0)), green cup B (position (0,1), grab), pink plate C (position (2,0))". The object space state information corresponding to moving it above the pink plate is "red plate A (position (0,0)), green cup B (position (2,1), grab), pink plate C (position (2,0))". The object space state information corresponding to putting it down vertically is "red plate A (position (0,0)), green cup B (position (2,0)), pink plate C (position (2,0))".
[0036] Step 204: According to each action instruction in the action sequence instruction, control the target robot to perform the corresponding action in sequence.
[0037] Optionally, according to each action instruction in the action sequence instruction, the target robot is controlled in sequence to perform the corresponding action, and the actual spatial state change information of each object after performing each action is fed back to judge whether the actual spatial state change information of each object after performing each action is consistent with the spatial state information corresponding to the action instruction. If inconsistent, re-planning is performed based on the task planner. If consistent, the action instructions in the action sequence instruction continue to be executed.
[0038] In an embodiment of the present application, an input task command for a target robot and a static scene image of the task scene in which the target robot is located, collected based on the task command, are obtained; the multi-form task command of the present application is combined with the static image to achieve cross-modal alignment, which can ensure the acquisition of scene information related to the task, avoid redundant calculations of full-scene scanning, and improve processing efficiency. Based on the static scene image and the task command, a preset visual-spatial reasoning model is used to generate a task action sequence corresponding to the task command, as well as the object space state information corresponding to each action in the task action sequence. The object state and spatial relationship are explicitly encoded through a dynamic scene graph, and zero-sample spatial reasoning can be achieved without task-specific fine-tuning, thus addressing the limitations of traditional models that rely on pre-trained data. Based on the task action sequence and the object space state information corresponding to each action, a preset task generator is used to generate action sequence instructions corresponding to the task command; the present application converts the abstract action sequence into specific instructions that can be executed by the robot, ensuring that the instructions are consistent with the real physical environment. According to each action instruction in the action sequence instruction, the target robot is controlled to perform the corresponding action in sequence. This application uses a preset task generator through the object space state information and the object space state information corresponding to each action to generate action sequence instructions corresponding to the task command, solve the problem of reasoning chain fragmentation, ensure the physical consistency of multi-step decision-making, ensure that subsequent actions correct early spatial constraint conflicts, avoid error accumulation, and thus improve the control efficiency and accuracy of the robot.
[0039] Based on the above embodiment, the preset visual space reasoning model includes: a graphic node extractor, a visual aid extractor, a scene image generator and a spatial reasoning module. To this end, the present application provides a process for determining the spatial state information of an object in a robot control method. Figure 3 A schematic diagram of a flow chart for determining object space state information in a robot control method provided in an embodiment of the present application, such as Figure 3 As shown, in the above step 202, based on the static scene image and the task command, a preset visual-spatial reasoning model is used to generate a task action sequence corresponding to the task command, and object space state information corresponding to each action in the task action sequence, including: Step 301: Use a graphic node extractor to extract nodes from a static scene image to obtain node identification information in the static scene image.
[0040] Node identification information may include: action, color, type, size, quantity and other information.
[0041] Optionally, a graphic node extractor is used to extract entities from static scene images, treating the entities in the static image as nodes and identifying the node's action, color, type, size, quantity and other information to obtain node identification information. Entities are specific objects or areas that can be perceived in static scene graphs and are digital representations of physical existence in the real world. For example, Figure 1 Cups, plates and other objects in the room.
[0042] Step 302: Use a visually assisted extractor to perform feature extraction on the node identification information to obtain scene visual features corresponding to the static scene image.
[0043] Optionally, a visually assisted extractor is used to perform feature extraction on the node identification information to obtain attribute information and depth information corresponding to the static scene image as scene visual features.
[0044] Step 303: Using a scene image generator, a dynamic structured scene image corresponding to the static scene image is generated according to the scene visual features.
[0045] The dynamic structured scene image includes: attribute information of each object and the spatial position relationship between the objects.
[0046] Optionally, a scene image generator is used to construct a dynamic structured scene graph based on the visual information features and depth information features in the scene visual features, so that the attribute information between multiple objects and the spatial position relationship between each object can be determined based on the dynamic structured scene graph.
[0047] Step 304: Using a spatial reasoning module, perform spatial reasoning on the dynamic structured scene image according to the task command to obtain a task action sequence and object spatial state information corresponding to each action in the task action sequence.
[0048] The system decomposes the logical order of task commands into multiple atomic reasoning steps, each of which includes a task action sequence. Specifically, it determines the objects corresponding to the task commands in a dynamic structured scene based on the task commands, and generates ordered atomic reasoning steps based on the object attributes and the positional relationships between the objects.
[0049] Optionally, a spatial reasoning module is used to determine the target object corresponding to the task command from the dynamic scene graph according to the task command, perform spatial reasoning, obtain the task action sequence, and determine, based on the task action sequence, the object spatial state information corresponding to each action in the dynamic scene graph after each action is executed.
[0050] Optionally, a spatial reasoning module is used to perform spatial reasoning on dynamic structured scene images according to task commands. It is necessary to ensure the consistency of task execution and verify the feasibility of actions through physical rules.
[0051] In an embodiment of the present application, a graphic node extractor is used to obtain node identification information in a static scene image; a visual aid extractor is used to obtain scene visual features corresponding to the static scene image; a scene image generator is used to generate a dynamic structured scene image corresponding to the static scene image; and a spatial reasoning module is used to obtain a task action sequence and object spatial state information corresponding to each action in the task action sequence. The present application can accurately mine key node information from a static scene image, and by generating a dynamic structured scene image, the scene is understood as dynamic and related, thereby improving the accuracy of processing complex task processes, adapting to the needs of robots and other devices performing embodied tasks, better coping with complex situations in real environments, and improving the versatility and practicality of robot control.
[0052] Based on the above embodiment, the visual aid extractor includes: an object detection network and a depth estimation network. To this end, the present application also provides a process for determining scene visual features in a robot control method, Figure 4 A schematic diagram of a process for determining scene visual features in a robot control method provided in an embodiment of the present application is shown in FIG. Figure 4 As shown, in the above step 302, a visual auxiliary extractor is used to extract features from the node identification information to obtain scene visual features corresponding to the static scene image, including: Step 401: Use an object detection network to perform object detection based on node identification information to generate attribute information of each object in a static scene image.
[0053] The attribute information may include visual attribute features such as appearance, color, shape, and texture, and the object detection network may be an open vocabulary detection model.
[0054] Optionally, the specific area corresponding to the node identification information is extracted from the static scene image by the object detection network in the visual assisted extractor, so as to accurately determine the visual attribute features such as appearance, color, shape, texture, etc. of each node, thereby generating the attribute information of each object in the static scene image.
[0055] Step 402: Use a depth estimation network to perform depth estimation based on the node identification information to obtain depth estimation information of the task scene.
[0056] The visual features of a scene include: attribute information of each object and depth estimation information. The depth estimation information is used to describe the distance relationship between objects in the scene and the camera, as well as the position and layout of objects in three-dimensional space.
[0057] Optionally, the depth value of each pixel in the task scene is determined by the depth estimation network in the visual aid extractor, the three-dimensional coordinates of the node are generated according to the node identification information and the depth value of each pixel in the task scene, the physical position of the node in the task scene is determined, and the depth estimation information of the task scene is obtained according to the three-dimensional coordinates of the node.
[0058] In the embodiments of the present application, an object detection network is used to detect objects based on node identification information, generating attribute information for each object in a static scene image. This attribute information allows for precise regional location of nodes, splitting adjacent or overlapping objects of the same color into independent nodes, thereby avoiding node recognition errors due to chromatic similarity. A depth estimation network is also used to perform depth estimation based on node identification information, obtaining depth estimation information for the task scene. This avoids the erroneous reasoning of unsupported objects by traditional models, ensures that node states conform to physical laws, and prevents collisions during robot control.
[0059] Based on the above embodiment, the present application further provides a process for determining node identification information in a robot control method. In the above step 302, a graphic node extractor is used to extract nodes from a static scene image to obtain node identification information in the static scene image, including: If the task command is a reorganization task command, a graphic node extractor is used to extract nodes from the static scene image to obtain identification information of each node in the static scene image.
[0060] Among them, the attribute information of each node and the space occupied by each node. The attribute information of each node may include the color, specifications, category and other information of each node. The space occupied information of each node may be the position information and occupied space information of each node in the task scene. For example, Figure 1 The green cup in the image is a node in the current task scene. Its attributes are: green, cylindrical, and a water cup. Its spatial occupancy information is: it is located at the second layer (5, 1) in three-dimensional space, occupying a volume of 2 × 2 × 4.
[0061] Optionally, the reorganization task is to reorganize the nodes in the task scenario, combining Figure 1 The reorganized task can be "pick up the green cup from the red plate and put it on the pink plate", or "put the yellow cup next to the orange bowl", etc., which is a task that reorganizes the nodes in the task scene.
[0062] Alternatively, if the task command is a reorganization task command, nodes in the task scene need to be adjusted. A graphical node extractor is then used to extract nodes from the static scene image to obtain identification information for each node in the static scene image. Based on the identification information for each node in the static scene image, the corresponding reorganization task command is executed.
[0063] In an embodiment of the present application, by extracting node attribute information and node space occupancy information in static scene images, a comprehensive description of objects in static scene images can be achieved, and scattered scene information can be integrated into structured knowledge, thereby enhancing task reasoning efficiency and improving execution accuracy.
[0064] Based on the above embodiment, this application also provides another process of a robot control method: Figure 5 A flow chart of another robot control method provided in an embodiment of the present application is shown as follows: Figure 5 As shown, based on the above steps 201 to 204, the method further includes: Step 501: Obtain a first question inputted to a target robot and a first scene image captured by the target robot based on the first question.
[0065] The first question is a spatial question and answer, not a question about driving the robot. For example, it can be a question like "What is the closest object to the robot?" or "What is the yellow object?" This embodiment of the application does not limit this. In this case, the target robot can be any robot, and there are no requirements for the robot's structure.
[0066] Optionally, the first question can be a question converted from an input voice into text, or a text question input directly. According to the input first question, the target robot is driven to perform image acquisition based on the first question, thereby obtaining a first scene image acquired by the target robot based on the first question.
[0067] Step 502: Based on the first question and the first scene image, a preset visual-spatial reasoning model is used to generate first visual-spatial reasoning information of the first scene image under the first question.
[0068] The first visual space reasoning information includes: spatial state information of each object in the first scene image. The first scene image is a static scene image.
[0069] Optionally, based on the first question and the first scene image, multiple nodes are extracted from the first question and the first scene image to obtain node identification information of the multiple nodes, feature extraction is performed on the node identification information to obtain scene visual features in the first scene image, a dynamic structured scene graph corresponding to the first scene image is generated based on the scene visual features in the first scene image, and first visual spatial reasoning information of the first scene image is determined from the dynamic structured scene graph corresponding to the first scene image.
[0070] Step 503: Generate first answer information corresponding to the first question using a preset task generator based on the first visual-spatial reasoning information.
[0071] Optionally, based on the first visual-spatial reasoning information, a preset task generator is used to determine the spatial state information of the corresponding node in the first visual-spatial reasoning information from the first question, and determine the first answer information corresponding to the first question.
[0072] For example, combined with Figure 1 To illustrate, when the first question is "What is the closest object to the robot?", the position of the robot is determined from the first visual space reasoning information, and the spatial state information of the robot and the spatial state information of other nodes except the robot are determined from the first visual space, thereby determining that the answer to the first question is "yellow water cup".
[0073] Step 504: Control the target robot to play the first answer information.
[0074] For example, at this time, the target robot plays the first answer information "yellow water cup".
[0075] In an embodiment of the present application, a first question directed to a target robot and a first scene image captured by the target robot based on the first question are input; based on the first question and the first scene image, a preset visual-spatial reasoning model is used to generate first visual-spatial reasoning information of the first scene image under the first question; based on the first visual-spatial reasoning information, a preset task generator is used to generate first answer information corresponding to the first question, and the target robot is controlled to play the first answer information. The present application can combine the first question and the first scene image to determine the first visual-spatial reasoning information in the current scene, thereby avoiding the illusion of spatial relationships and improving the accuracy of the first answer information.
[0076] Based on the above embodiments, the present application also provides a process for testing a preset space model in a robot control method. Figure 6 A schematic diagram of a flow chart of testing a preset space model in a robot control method provided in an embodiment of the present application, such as Figure 6 As shown, based on the above steps 201 to 204, the method further includes: Step 601: Obtain sample video data and preset question-answer pairs corresponding to the sample video data in the input preset spatial data set.
[0077] The preset spatial dataset includes multiple categories of preset question-answer pairs, including reachability questions, adjacency questions, and distance questions. Reachability questions are used to determine whether a subject has achieved a specific goal. Examples of preset question-answer pairs include: "Can the robot reach the blue cup?" and "Yes." Adjacency questions are used to determine the spatial proximity between objects. Examples of preset question-answer pairs include: "What object is next to the yellow cup?" and "Pink plate." Distance questions ask and answer questions about the distance between an object and a reference object. Examples of preset question-answer pairs include: "Which object is closest to the yellow plate?" and "Purple bowl." Dynamic changes in real or simulated scenes in sample video data. The preset spatial dataset can be data collected in advance based on the current scene or a public dataset.
[0078] Optionally, sample video data and preset question-answer pairs corresponding to the sample video data are obtained from the input preset spatial data set. The sample video data and the preset question-answer pairs corresponding to the sample video data are questions that the target robot can answer or process.
[0079] Step 602: Based on the second question in the preset question-answer pair and the second scene image corresponding to the second question in the sample video data, a preset visual-spatial reasoning model is used to generate second visual-spatial reasoning information of the second scene image under the second question.
[0080] The second visual space reasoning information includes: spatial state information of each object in the second scene image. The second scene image is a static scene image.
[0081] Optionally, based on the second question and the sample video data, the sample video data is identified to obtain a second scene image corresponding to the second question in the sample video data, multiple nodes are extracted from the second question and the second scene image to obtain node identification information of the multiple nodes, feature extraction is performed on the node identification information to obtain scene visual features in the second scene image, a dynamic structured scene graph corresponding to the second scene image is generated based on the scene visual features in the second scene image, and second visual spatial reasoning information of the second scene image is determined from the dynamic structured scene graph corresponding to the second scene image.
[0082] Step 603: Generate second answer information corresponding to the second question using a preset task generator based on the second visual-spatial reasoning information.
[0083] Optionally, based on the first visual-spatial reasoning information, a preset task generator is used to determine the spatial state information of the corresponding node in the first visual-spatial reasoning information from the first question, and determine the first answer information corresponding to the first question.
[0084] For example, combined with Figure 1 To illustrate, when the second question is "Which object is closest to the yellow plate?", the position of the yellow plate is determined from the second visual space reasoning information, and the spatial state information of the yellow plate and the spatial state information of other nodes except the yellow plate are determined from the second visual space, thereby determining that the answer to the second question is "purple bowl".
[0085] Step 604: Test the preset visual-spatial reasoning model based on the second answer information and the preset answer corresponding to the second question in the preset question-answer pair.
[0086] For example, when the second answer information is "purple bowl" and the preset answer corresponding to the second question in the preset question and answer pair is "purple bowl", the reasoning ability of the preset visual-spatial reasoning model is strong; when the second answer information is "pink plate" and the preset answer corresponding to the second question in the preset question and answer pair is "purple bowl", the reasoning ability of the preset visual-spatial reasoning model is poor, then the preset visual-spatial reasoning model is trained according to the preset spatial data set to improve the reasoning ability of the preset visual-spatial reasoning model.
[0087] On the basis of the above embodiment, sample video data in the input preset spatial data set and the preset question-answer pair corresponding to the sample video data are obtained; based on the second question in the preset question-answer pair and the second scene image corresponding to the second question in the sample video data, a preset visual spatial reasoning model is used to generate second visual spatial reasoning information of the second scene image under the second question; based on the second visual spatial reasoning information, a preset task generator is used to generate second answer information corresponding to the second question; based on the second answer information and the preset answer corresponding to the second question in the preset question-answer pair, the preset visual spatial reasoning model is tested. This application can strengthen the spatial relationship reasoning of the preset visual spatial reasoning model in a targeted manner and improve the scene generalization and robustness of the preset visual spatial reasoning model.
[0088] Based on the above embodiments, the present application further provides a process for determining a preset spatial data set in a robot control method. Figure 7 The process of determining a preset spatial data set in a robot control method provided in an embodiment of the present application is as follows: Figure 7 As shown, before obtaining the sample video data and the preset question-answer pairs corresponding to the sample video data in the input preset spatial data set in the above step 601, the method further includes: Step 701: Obtain sample video data and a preset task category corresponding to the sample video data.
[0089] Preset task categories include spatial reasoning problems and assembly tasks. Spatial reasoning tasks assess the robot's dynamic interactions, and can include tasks such as quantity, color, direction and orientation, object overlap, reachability, task success, manipulator feasibility, and distance or proximity relationships. Assembly tasks assess the robot's ability to understand static spatial attributes, focusing on the attributes and compositional relationships of objects, such as quantity, color, relative relationships, and size.
[0090] Optionally, the sample video data can be sample video data pre-constructed for the target robot's task, or relevant video data can be obtained from a dataset in an open source library, or sample video data can be constructed according to preset task categories, etc. This application does not impose any restrictions on this.
[0091] Optionally, when obtaining relevant video data from the dataset of the open source library, it is necessary to remove the interference data, obtain sample video data from the continuous video, and retain the scene fragments related to reasoning. The interference data can be task interruption data, video or image fragmentation error data.
[0092] Step 702: Generate a preset question-answer pair corresponding to the sample video data based on the preset task category and the sample video data.
[0093] Optionally, when the preset task category is a spatial reasoning problem, the multimodal large model and the sample video data will be labeled based on the sample video data to generate question-answer pairs corresponding to the sample video data; when the preset task category is an assembly task, it will be determined whether the assembly task is completed based on the sample video data.
[0094] In an embodiment of the present application, sample video data and a preset task category corresponding to the sample video data are obtained; based on the preset task category and the sample video data, preset question-answer pairs corresponding to the sample video data are generated. This application filters and categorizes the sample video data by preset task categories, ensuring that the generated question-answer pairs closely focus on specific task objectives and avoid interference from irrelevant information. By obtaining sample data and preset question-answer pairs, the coverage capability of the preset visual-spatial reasoning model data for complex scenarios can be enhanced, thereby improving the accuracy of controlling the target robot.
[0095] The robot control method provided by the embodiment of the present application is described below with reference to the accompanying drawings. Figure 8 A schematic diagram of a robot control architecture provided in an embodiment of the present application is shown in FIG. Figure 8As shown, the preset visual-spatial reasoning model includes: a graph node extractor, a visual aid extractor, a scene image generator, and a spatial reasoning module. The visual aid extractor includes: an object detection network and a depth estimation network.
[0096] Optionally, an input task command for a target robot and a static scene image of the task scene captured by the target robot based on the task command are obtained. A graphical node extractor is used to extract nodes from the static scene image to obtain node identification information in the static scene image. A visual aid extractor is used to extract features from the node identification information to obtain scene visual features corresponding to the static scene image. A scene image generator is used to generate a dynamic structured scene image corresponding to the static scene image based on the scene visual features. The dynamic structured scene image includes attribute information of each object and the spatial positional relationships between the objects. A spatial reasoning module is used to perform spatial reasoning on the dynamic structured scene image based on the task command to obtain a task action sequence and object spatial state information corresponding to each action in the task action sequence. Based on the task action sequence and the object spatial state information corresponding to each action, a preset task generator is used to generate action sequence instructions corresponding to the task command. Based on each action instruction in the action sequence instructions, the target robot is sequentially controlled to perform the corresponding action.
[0097] Based on the same inventive concept, the embodiment of the present application also provides a robot control device corresponding to the robot control method. Since the principle of solving the problem by the device in the embodiment of the present application is similar to the control method of the robot in the embodiment of the present application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be repeated.
[0098] Figure 9 This is a schematic diagram of the structure of a robot control device provided in an embodiment of the present application. The device includes: an acquisition module 901, a generation module 902, and a control module 903. The acquisition module 901 is used to acquire an input task command for a target robot and a static scene image of the task scene in which the target robot is located, which is collected based on the task command. A generation module 902 is configured to generate, based on the static scene image and the task command, a task action sequence corresponding to the task command and object spatial state information corresponding to each action in the task action sequence using a preset visual-spatial reasoning model; wherein the object spatial state information corresponding to each action is used to represent the spatial state change information of each object in the static scene image under each action; A generation module 902 is used to generate an action sequence instruction corresponding to the task command using a preset task generator according to the task action sequence and the object space state information corresponding to each action; The control module 903 is used to control the target robot to perform corresponding actions in sequence according to each action instruction in the action sequence instruction.
[0099] Optionally, the preset visual-spatial reasoning model includes: a graphic node extractor, a visual aid extractor, a scene image generator, and a spatial reasoning module, wherein the generation module 902 is specifically configured to: use the graphic node extractor to extract nodes from the static scene image to obtain node identification information in the static scene image; A visually assisted extractor is used to extract features from node identification information to obtain scene visual features corresponding to static scene images. A scene image generator is used to generate a dynamic structured scene image corresponding to the static scene image according to the scene visual features. The dynamic structured scene image includes: attribute information of each object and the spatial position relationship between the objects; The spatial reasoning module is used to perform spatial reasoning on the dynamic structured scene image according to the task command to obtain the task action sequence and the object spatial state information corresponding to each action in the task action sequence.
[0100] Optionally, the visual aid extractor includes: an object detection network and a depth estimation network; a generation module 902, specifically configured to: use the object detection network to perform object detection based on node identification information, and generate attribute information of each object in the static scene image; A depth estimation network is used to perform depth estimation based on node identification information to obtain depth estimation information of the task scene. The scene visual features include: attribute information of each object and depth estimation information.
[0101] Optionally, the generation module 902 is specifically used to: if the task command is a reorganization task command, use a graphic node extractor to extract nodes from the static scene image to obtain identification information of each node in the static scene image, and each node identification information includes: attribute information of each node and space occupancy information of each node.
[0102] Optionally, the device further includes: a question-answering module, the question-answering module being specifically configured to: obtain an input first question directed to the target robot, and a first scene image captured by the target robot based on the first question; Based on the first question and the first scene image, a preset visual-spatial reasoning model is used to generate first visual-spatial reasoning information of the first scene image under the first question; the first visual-spatial reasoning information includes: spatial state information of each object in the first scene image; Generate first answer information corresponding to the first question using a preset task generator according to the first visual-spatial reasoning information; Control the target robot to play the first answer information.
[0103] Optionally, the device further includes: a testing module, the testing module being specifically configured to: obtain sample video data and preset question-answer pairs corresponding to the sample video data in the input preset spatial data set; Based on the second question in the preset question-answer pair and the second scene image corresponding to the second question in the sample video data, a preset visual-spatial reasoning model is used to generate second visual-spatial reasoning information of the second scene image under the second question; the second visual-spatial reasoning information includes: spatial state information of each object in the second scene image; Generate second answer information corresponding to the second question using a preset task generator based on the second visual-spatial reasoning information; The preset visual-spatial reasoning model is tested according to the second answer information and the preset answer corresponding to the second question in the preset question-answer pair.
[0104] Optionally, the testing module is further configured to: obtain sample video data and a preset task category corresponding to the sample video data; Based on the preset task category and sample video data, a preset question-answer pair corresponding to the sample video data is generated.
[0105] For descriptions of the processing flow of each module in the device and the interaction flow between each module, reference can be made to the relevant descriptions in the above method embodiment, which will not be described in detail here.
[0106] The embodiment of the present application also provides a robot, Figure 10 A schematic diagram of the structure of a robot control device provided in an embodiment of the present application is shown in FIG. Figure 10 As shown, the control device of the robot includes: a processor 1001, a memory 1002, and optionally, a bus 1003. The memory 1002 stores machine-readable instructions that can be executed by the processor 1001 (for example, Figure 9 In the device, the acquisition module 901, the generation module 902, the execution instructions corresponding to the control module 903, etc. are obtained). When the control device is running, the processor 1001 communicates with the memory 1002 through the bus 1003. When the machine-readable instructions are executed by the processor 1001, the steps of the above-mentioned robot control method are performed.
[0107] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned robot control method are executed.
[0108] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the system and device described above can refer to the corresponding process in the method embodiment, and will not be repeated in this application. In the several embodiments provided in this application, it should be understood that the disclosed system, device and method can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the modules is only a logical function division. There may be other division methods in actual implementation. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.
[0109] In addition, the functional units in the various embodiments of the present application can be integrated into a single processing unit, each unit can exist physically separately, or two or more units can be integrated into a single unit. If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in the various embodiments of the present invention. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0110] The above is only a specific implementation method of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in this application, which should be covered by the protection scope of the present application.
Claims
1. A robot control method, characterized in that: The method comprises: Obtaining an input task command for a target robot and a static scene image of a task scene captured by the target robot based on the task command; Based on the static scene image and the task command, a preset visual-spatial reasoning model is used to generate a task action sequence corresponding to the task command, and object spatial state information corresponding to each action in the task action sequence; wherein the object spatial state information corresponding to each action is used to represent spatial state change information of each object in the static scene image under each action; According to the task action sequence and the object space state information corresponding to each action, a preset task generator is used to generate an action sequence instruction corresponding to the task command; According to each action instruction in the action sequence instruction, the target robot is controlled in sequence to perform the corresponding action.
2. The method according to claim 1, characterized in that The preset visual-spatial reasoning model includes: a graphic node extractor, a visual aid extractor, a scene image generator, and a spatial reasoning module. The preset visual-spatial reasoning model is used to generate a task action sequence corresponding to the task command and object space state information corresponding to each action in the task action sequence based on the static scene image and the task command, including: Using the graphic node extractor, extracting nodes from the static scene image to obtain node identification information in the static scene image; Using the visual aid extractor, extracting features from the node identification information to obtain scene visual features corresponding to the static scene image; The scene image generator is used to generate a dynamic structured scene image corresponding to the static scene image according to the scene visual features, wherein the dynamic structured scene image includes: attribute information of each object and the spatial position relationship between the objects; The spatial reasoning module is used to perform spatial reasoning on the dynamic structured scene image according to the task command to obtain the task action sequence and the object spatial state information corresponding to each action in the task action sequence.
3. The method according to claim 2, characterized in that The visually assisted extractor includes: an object detection network and a depth estimation network; the visually assisted extractor is used to perform feature extraction on the node identification information to obtain scene visual features corresponding to the static scene image, including: Using the object detection network, performing object detection based on the node identification information to generate attribute information of each object in the static scene image; The depth estimation network is used to perform depth estimation according to the node identification information to obtain depth estimation information of the task scene, and the scene visual features include: attribute information of each object and the depth estimation information.
4. The method according to claim 2, characterized in that The step of extracting nodes from the static scene image using the graphic node extractor to obtain node identification information in the static scene image includes: If the task command is a reorganization task command, the graphic node extractor is used to extract nodes from the static scene image to obtain identification information of each node in the static scene image, and the identification information of each node includes: attribute information of each node and space occupancy information of each node.
5. The method according to claim 1, wherein The method further comprises: Obtaining an input first question for the target robot and a first scene image captured by the target robot based on the first question; Based on the first question and the first scene image, using the preset visual-spatial reasoning model, generating first visual-spatial reasoning information of the first scene image under the first question; the first visual-spatial reasoning information includes: spatial state information of each object in the first scene image; Generate first answer information corresponding to the first question using the preset task generator according to the first visual-spatial reasoning information; Control the target robot to play the first answer information.
6. The method according to claim 1, characterized in that The method further comprises: Obtaining sample video data and preset question-answer pairs corresponding to the sample video data in an input preset spatial data set; generating, based on the second question in the preset question-answer pair and the second scene image corresponding to the second question in the sample video data, second visual-spatial reasoning information of the second scene image under the second question using the preset visual-spatial reasoning model; the second visual-spatial reasoning information including spatial state information of each object in the second scene image; Generate second answer information corresponding to the second question using the preset task generator according to the second visual-spatial reasoning information; The preset visual-spatial reasoning model is tested according to the second answer information and the preset answer corresponding to the second question in the preset question-answer pair.
7. The method according to claim 6, characterized in that Before obtaining the sample video data and the preset question-answer pairs corresponding to the sample video data in the input preset spatial data set, the method further includes: Acquire the sample video data and the preset task category corresponding to the sample video data; Based on the preset task category and the sample video data, a preset question-answer pair corresponding to the sample video data is generated.
8. A robot control device, characterized in that: The device comprises: An acquisition module is used to acquire an input task command for a target robot and a static scene image of the task scene in which the target robot is located, which is collected based on the task command; a generation module, configured to generate, based on the static scene image and the task command, a task action sequence corresponding to the task command and object spatial state information corresponding to each action in the task action sequence using a preset visual-spatial reasoning model; wherein the object spatial state information corresponding to each action is used to represent spatial state change information of each object in the static scene image under each action; The generating module is configured to generate an action sequence instruction corresponding to the task command using a preset task generator according to the task action sequence and the object space state information corresponding to each action; The control module is used to control the target robot to perform corresponding actions in sequence according to each action instruction in the action sequence instruction.
9. A robot control device, characterized in that: include: A processor and a memory, wherein the memory stores machine-readable instructions executable by the processor, and when the control device of the robot is running, the processor executes the machine-readable instructions to perform the steps of the robot control method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, executes the steps of the robot control method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Space environment identification method and system for robot intelligent service
CN111496784A
Robot grabbing method and system in unknown environment
CN113664826A
Fish and vegetable symbiosis scene image visual question and answer method and device and electronic equipment
CN117668169A
Intelligent mechanical arm operation method and system based on multi-mode large visual language model
CN119567268A
Indoor mobile service robot interaction task execution method and device, storage medium and indoor mobile service robot system
CN119772883A
Cited By
Robot control method and device and storage medium
CN120921403A
Public network communication area-free multi-nest cooperative unmanned aerial vehicle autonomous flight patrol method and system
CN121455215A
Robot action generation system and method based on diffusion model and asynchronous inference
CN122735763A