Robot operation control method, device and equipment and readable storage medium

By acquiring task operation instructions and image recognition, combined with a vision-language-action mapping method, the accuracy problem of robot operation in high-risk environments was solved, and high-precision power operation control was achieved.

CN121374615APending Publication Date: 2026-01-23SHENZHEN POWER SUPPLY BUREAU
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511773902.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-28
Publication Date
2026-01-23

AI Technical Summary

Technical Problem

When performing electrical work in high-risk and complex environments such as distribution boxes, manual operation faces high safety risks and low accuracy in task completion. Existing technologies are insufficient to achieve high-precision robot operation control.

Method used

By acquiring task instructions and images, identifying the task object and execution component information, and combining visual instructions with other information, the robot's task progress is determined and the next action is generated. The vision-language-action mapping method is used to improve the accuracy of task completion.

Benefits of technology

It enables real-time and accurate progress determination and action execution for robot operations, improving the accuracy of task completion and reducing the safety risks of manual operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121374615A_ABST
    Figure CN121374615A_ABST
Patent Text Reader

Abstract

The invention relates to a robot operation control method and device, computer equipment, a computer readable storage medium and a computer program product. The method comprises the steps of obtaining a task operation instruction and an operation image corresponding to a task operation environment; identifying the operation image to obtain a target identification result; the target recognition result comprises operation object information and execution component information of the robot; determining the task operation progress of the robot based on the operation object information and the execution component information; fusing the task operation instruction and the target identification result to obtain visual instruction joint information; and according to the execution component information, the task operation progress and the visual instruction joint information, determining the next action of the robot until a target operation task corresponding to the task operation instruction is completed. By adopting the method, the task completion accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of power automation, in particular to a robot operation control method and device, computer equipment, computer readable storage medium and computer program product. BACKGROUND

[0002] In high-risk and complex environments such as distribution boxes, manual operation faces inherent safety risks such as electric shock and electric arc, and the labor cost is high. Therefore, introducing autonomous robots with high-precision positioning and reliable stage determination capability is an inevitable trend to ensure operation safety and efficiency. To reduce task-specific rule engineering and enhance cross-device / task migration, a unified modeling "vision-language-action" (VLA, Vision-Language-Action) mapping method is adopted to enable robots to understand natural language instructions and generate operations based on environmental visual information, which is the current research direction.

[0003] In the traditional technology, the phenomenon of task "false completion" is easy to occur, resulting in low task completion accuracy. SUMMARY

[0004] Therefore, it is necessary to provide a robot operation control method, device, computer equipment, computer readable storage medium and computer program product capable of improving task completion accuracy in view of the above technical problems.

[0005] In a first aspect, the present application provides a robot operation control method, comprising:

[0006] obtaining a task operation instruction and an operation image corresponding to a task operation environment;

[0007] identifying the operation image to obtain a target recognition result; the target recognition result includes operation object information and robot execution component information;

[0008] determining a task operation progress of the robot based on the operation object information and the execution component information;

[0009] fusing the task operation instruction and the target recognition result to obtain visual instruction joint information;

[0010] determining a next action of the robot according to the execution component information, the task operation progress and the visual instruction joint information until a target operation task corresponding to the task operation instruction is completed.

[0011] In one embodiment, the determination of the task operation progress of the robot based on the operation object information and the execution component information comprises:

[0012] obtaining an initial position of an execution component of the robot;

[0013] determining a reference distance between the execution component and the work object based on a work object position in the work object information and the initial position of the execution component;

[0014] determining an actual distance between the execution component and the work object based on the work object position in the work object information and an actual position of the execution component in the execution component information;

[0015] determining a task work progress of the robot based on the actual distance and the reference distance.

[0016] In one of the embodiments, the determining the task work progress of the robot based on the actual distance and the reference distance comprises:

[0017] calculating a ratio between the actual distance and the reference distance, and determining a difference between 1 and the ratio as the task work progress of the robot.

[0018] In one of the embodiments, the fusing the task work instruction and the target recognition result to obtain visual instruction joint information comprises:

[0019] encoding the task work instruction through a text encoder to obtain work instruction features;

[0020] encoding the target recognition result through a visual encoder to obtain multi-scale spatial features;

[0021] fusing the work instruction features and the multi-scale spatial features through a cross-modal attention mechanism to obtain visual instruction joint features.

[0022] In one of the embodiments, the fusing the task work instruction and the target recognition result to obtain visual instruction joint information comprises:

[0023] annotating the work image through the target recognition result to obtain a task annotation image;

[0024] establishing a semantic correspondence between the task annotation image and the task work instruction, and aligning the task work instruction and the task annotation image according to the semantic correspondence to obtain visual instruction joint information.

[0025] In one of the embodiments, the method further comprises:

[0026] obtaining a multi-view image dataset corresponding to the task work environment;

[0027] annotating the multi-view image dataset to obtain an annotated image dataset;

[0028] training an initial detection model based on the annotated image dataset to obtain a target detection model;

[0029] The identifying the work image to obtain a target recognition result comprises:

[0030] identifying the work image based on the target detection model to obtain a target recognition result.

[0031] In a second aspect, the present application further provides a robot work control device, comprising:

[0032] a data acquisition module configured to acquire a task work instruction and a work image corresponding to a task work environment;

[0033] an image recognition module configured to identify the work image to obtain a target recognition result, wherein the target recognition result comprises work object information and robot execution component information;

[0034] a progress determination module configured to determine a task work progress of the robot based on the work object information and the execution component information;

[0035] an information fusion module configured to fuse the task work instruction and the target recognition result to obtain visual instruction joint information;

[0036] an action determination module configured to determine a next action of the robot according to the execution component information, the task work progress and the visual instruction joint information, until a target work task corresponding to the task work instruction is completed.

[0037] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the robot work control method provided in the first aspect when executing the computer program.

[0038] In a fourth aspect, the present application further provides a computer readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the steps of the robot work control method provided in the first aspect.

[0039] In a fifth aspect, the present application further provides a computer program product comprising a computer program, wherein the computer program is executed by a processor to implement the steps of the robot work control method provided in the first aspect.

[0040] The robot operation control method, device, computer equipment, computer readable storage medium and computer program product can obtain a task operation instruction and an operation image corresponding to a task operation environment, identify the operation image to obtain a target identification result, the target identification result includes operation object information and robot execution component information, determine a task operation progress of the robot based on the operation object information and the execution component information, fuse the task operation instruction and the target identification result to obtain visual instruction joint information, and determine a next action of the robot according to the execution component information, the task operation progress and the visual instruction joint information, until a target operation task corresponding to the task operation instruction is completed. The robot operation control method, device, computer equipment, computer readable storage medium and computer program product can realize real-time determination of the task progress of the robot based on actual operation object information and execution component information, accurate determination of the subsequent action of the robot based on the execution component information, the task operation progress and the visual instruction joint information, and accurate completion of the corresponding operation task, thereby improving the task completion accuracy of the robot. BRIEF DESCRIPTION OF DRAWINGS

[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the drawings needed to be used in the description of the embodiments of the present application or the related art will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other related drawings can be obtained by those skilled in the art without creative labor.

[0042] Figure 1 An application environment diagram of the robot operation control method in an embodiment;

[0043] Figure 2 A flowchart of the robot operation control method in an embodiment;

[0044] Figure 3 A flowchart of the robot operation control method in another embodiment;

[0045] Figure 4 A structural block diagram of the robot operation control device in an embodiment;

[0046] Figure 5 An internal structure diagram of the computer equipment in an embodiment. DETAILED DESCRIPTION

[0047] In order to make the purpose, technical solutions and advantages of the present application more clear, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.

[0048] It should be noted that the terms "first", "second", etc. used in the present application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "include" and "have" and any variations thereof used in the present application are intended to cover non-exclusive inclusion. The term "multiple" used in the present application refers to two or more. The term "and / or" used in the present application refers to one of the options or any combination of multiple options.

[0049] The robot operation control method provided by the embodiments of the present application can be applied to an application environment as shown in Figure 1 The terminal 102 communicates with the server 104 through a network. The data storage system can store data required to be processed by the server 104. The data storage system can be integrated on the server 104, or placed on a cloud or other network server. The terminal 102 obtains task operation instructions and operation images corresponding to a task operation environment, and sends the obtained task operation instructions and operation images corresponding to the task operation environment to the server 104. The server 104 identifies the operation images to obtain target identification results, which include operation object information and robot execution component information. Based on the operation object information and the execution component information, the task operation progress of the robot is determined. The task operation instructions and the target identification results are fused to obtain visual instruction joint information. According to the execution component information, the task operation progress and the visual instruction joint information, the next action of the robot is determined, until the target operation task corresponding to the task operation instructions is completed. The terminal 102 can be, but is not limited to, various personal computers, notebook computers, smart phones, tablet computers, unmanned aerial vehicles, low-altitude flying vehicles, Internet of Things devices and portable wearable devices. The Internet of Things device can be a smart speaker, a smart television, a smart air conditioner, a smart vehicle device, a projection device, etc. The portable wearable device can be a smart watch, a smart bracelet, a head-mounted device, etc. The head-mounted device can be a virtual reality (VR) device, an augmented reality (AR) device, smart glasses, etc. The server 104 can be a standalone physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. It should be noted that the robot operation control method provided by the embodiments of the present application is not limited to the application scenario of interaction between the above terminal and the server, but is also applicable to the application scenarios of single terminal, single server, server-server interaction or terminal-terminal interaction.

[0050] In an exemplary embodiment, as shown in Figure 2 A robot operation control method is provided. The method is applied to Figure 1Taking the server in the example, the explanation includes the following steps 202 to 210. Wherein:

[0051] Step 202: Obtain the task operation instructions and the corresponding operation image of the task operation environment.

[0052] The task operation instruction refers to the instruction to achieve or complete the target task. Task operation instructions can be obtained in text or voice format. The target task is a task to be completed by the robot. For example, the target task might be "press the yellow button," "turn on the indicator light," or "tighten the knob." The task operation environment refers to the operating environment corresponding to the task operation instruction. The operation image can be images from multiple perspectives corresponding to the task operation environment. For example, images of the task operation environment can be simultaneously captured by cameras in different locations.

[0053] For example, images of the task's working environment can be simultaneously acquired by cameras installed on the robot's head, left wrist, and right wrist, and the images from different perspectives can be adjusted to the same resolution.

[0054] Step 204: Recognize the work image to obtain the target recognition result; wherein, the target recognition result includes the work object information and the robot's execution component information.

[0055] Target recognition results are derived from feature recognition of the task image. It's easy to understand that a robot typically operates within a task environment, and the task image corresponding to this environment can include information about the work objects and actuators within that environment. Therefore, target recognition results include information about the work objects and the robot's actuators. The work object information refers to the information about the work object specified in the task command. For example, a work object is typically a component of a distribution box, such as a button, knob, pressure plate, or indicator light. Work object information includes, for example, the type and location of the work object. The actuator information refers to the information about the actuators themselves. An actuator is the component that the robot actually uses to manipulate the work object, such as the gripper or "hand" at the end of a robot arm. Actuator information includes, for example, the position and orientation of the actuator.

[0056] For example, a trained object detection model can be used to identify the target in the work image to obtain the target recognition result. The object detection model can be trained using a dataset of feature-annotated work images. Alternatively, an object detection algorithm can be used to identify the target in the work image to obtain the target recognition result.

[0057] Step 206: Determine the robot's task progress based on the task object information and the execution component information.

[0058] It is easy to understand that the target job task corresponding to the task job instruction usually needs the robot to perform a series of continuous actions to complete. The task job progress of the robot is used to represent the progress of the robot in completing the target job task. The task job progress can be represented by a task phase label. The task phase label includes, for example, approach, contact, execution, verification, and exit.

[0059] For example, the server can determine the distance between the job object and the execution component based on the job object information and the execution component information, and determine the task job progress of the robot according to the distance between the job object and the execution component. For example, assuming that the distance between the job object and the execution component is the largest in the initial state, the smaller the distance between the job object and the execution component, the closer the corresponding task job progress is to the completion state; on the contrary, the larger the distance between the job object and the execution component, the closer the corresponding task job progress is to the initial state, that is, the farther from the completion state.

[0060] In an exemplary embodiment, the state of the job object can be determined based on the job object information, and the state of the execution component can be determined based on the execution component information. The state of the job object and the state of the execution component are matched with the corresponding reference state respectively, and the job progress corresponding to the matched reference state is determined as the task job progress of the robot in the corresponding case. There is a one-to-one correspondence between the reference state and the job progress.

[0061] Step 208, fusing the task job instruction and the target recognition result to obtain visual instruction joint information.

[0062] The visual instruction joint information refers to the information corresponding to the visual perception result and the task job instruction. It is easy to understand that the visual perception result is an identification result obtained by identifying the job image, and the task job instruction is a language instruction input in the form of text or voice. The task job instruction and the target recognition result belong to different modalities of information. Fusing the task job instruction and the target recognition result of different modalities can obtain the visual instruction joint information.

[0063] For example, the task job instruction in the form of text can be encoded to obtain job instruction features, and the target recognition result can be visually encoded to obtain multi-scale spatial features. The job instruction features and the multi-scale spatial features are fused to obtain visual instruction joint features.

[0064] Step 210, determining the next action of the robot according to the execution component information, the task job progress and the visual instruction joint information, until the target job task corresponding to the task job instruction is completed.

[0065] Exemplarily, the action generation model trained can determine the next action of the robot according to the execution component information, the task operation progress and the visual instruction joint information. The execution component information includes the execution component information identified from the obtained current frame operation image, and the execution component information identified from K frame operation images before the current frame. In other words, the execution component information includes the pose (position and attitude) information of the execution component in the current frame operation image and the historical pose information of the execution component in the K frame operation images before the current frame.

[0066] In an actual application scenario, the server can obtain a task operation instruction, and obtain an operation image corresponding to a task operation environment every interval of a preset time length. According to the operation image obtained at each time and the task operation instruction, the next action of the robot corresponding to the next time can be determined until the target operation task corresponding to the task operation instruction is completed, and the operation image corresponding to the task operation environment is stopped. Exemplarily, the server can detect the task operation progress in real time, and if the task operation progress is greater than or equal to a progress threshold, it is determined that the corresponding target operation task is completed. Alternatively, the server can detect the state of the operation object, and if the state of the operation object reaches the expected state corresponding to the target operation task, it is determined that the corresponding target operation task is completed.

[0067] In the above robot operation control method, by obtaining a task operation instruction and an operation image corresponding to a task operation environment, the operation image is identified to obtain a target identification result, the target identification result includes operation object information and execution component information of the robot, the task operation progress of the robot is determined based on the operation object information and the execution component information, the task operation instruction and the target identification result are fused to obtain visual instruction joint information, and the next action of the robot is determined according to the execution component information, the task operation progress and the visual instruction joint information, until the target operation task corresponding to the task operation instruction is completed. It can realize real-time determination of the task progress of the robot based on the actual operation object information and the execution component information, accurate determination of the subsequent action of the robot combined with the execution component information, the task operation progress and the visual instruction joint information, and accurate completion of the corresponding operation task, thereby improving the task completion accuracy of the robot.

[0068] In some embodiments, the task operation progress of the robot is determined based on the operation object information and the execution component information in step 206, including:

[0069] obtain an initial position of the execution component of the robot; determine a reference distance between the execution component and the work object based on the work object position in the work object information and the initial position of the execution component; determine an actual distance between the execution component and the work object based on the work object position in the work object information and the actual position of the execution component in the execution component information; and determine the task work progress of the robot based on the actual distance and the reference distance.

[0070] The initial position refers to the position of the execution component of the robot before the robot performs the target work task. It is easy to understand that the execution component of the robot will be placed at the initial position after completing each work task. The initial position may be, for example, a fixed position in front of the chest of the robot, or a position in which the robot arm is vertically lowered, and the initial position may be set according to the actual application scenario.

[0071] It is easy to understand that the positions in the embodiment can be represented by coordinates. The distance can be represented by the Euclidean distance or the Manhattan distance. The work object information includes the work object position, and the execution component information includes the position of the execution component. For example, a work image corresponding to the task work environment in the initial state can be obtained, and the execution component information of the robot can be obtained by identifying the work image. At this time, the execution component information of the robot includes the initial position. Alternatively, the initial position of the execution component of the robot can be determined based on the position of the robot in the work task environment and the positional relationship between the robot and the execution component.

[0072] In an actual application scenario, the server can calculate the distance between the work object position in the work object information and the initial position of the execution component, and take the distance as the reference distance between the execution component and the work object. The distance between the work object position in the work object information and the actual position of the execution component in the execution component information is calculated, and the distance is taken as the actual distance between the execution component and the work object. The task work progress of the robot is determined based on the actual distance and the reference distance.

[0073] In an exemplary embodiment, the ratio of the actual distance to the reference distance can be determined as the task work progress of the robot. The larger the ratio of the actual distance to the reference distance, the closer the task work progress to the initial state of the task. The smaller the ratio of the actual distance to the reference distance, the closer the task work progress to the completion state of the task. For example, if the ratio of the actual distance to the reference distance is 1, it means that the task work progress is 0, i.e., the target work task has not started. If the ratio of the actual distance to the reference distance is 0, it means that the task work progress is 1, i.e., the target work task is completed.

[0074] In the embodiment, the reference distance between the execution component and the work object is determined based on the work object position in the work object information and the initialization position of the execution component, the actual distance between the execution component and the work object is determined based on the work object position in the work object information and the actual position of the execution component in the execution component information, and the task work progress of the robot is determined based on the actual distance and the reference distance, so that the task work progress of the robot can be accurately determined based on the work object position and the execution component position, and the accuracy of the task work progress is improved.

[0075] In some embodiments, the task work progress of the robot is determined based on the actual distance and the reference distance, including:

[0076] A ratio between the actual distance and the reference distance is calculated, and a difference between 1 and the ratio is determined as the task work progress of the robot.

[0077] Exemplarily, the reference distance D0 between the execution component and the work object can be calculated by the following formula (1).

[0078] Formula (1)

[0079] wherein u0 represents the initialization position of the execution component; c t represents the work object position; represents the calculation of the Euclidean distance. In other words, the distance between the initialization position of the execution component and the work object position is taken as the reference distance.

[0080] The actual distance D t between the execution component and the work object can be calculated by the following formula (2).

[0081] Formula (2)

[0082] wherein u t represents the actual position of the execution component in the execution component information.

[0083] The task work progress s can be calculated by the following formula (3).

[0084] Formula (3)

[0085] wherein s = 0 represents that the target work task has not started, and s = 1 represents that the execution component reaches the work object position and the target work task is completed.

[0086] In an exemplary embodiment, the task work progress s corresponding to each frame of work image is determined, and if the task work progress s corresponding to a continuous preset number of frames of work image is greater than or equal to the progress threshold s donedetermine the target task completion.

[0087] In this embodiment, by calculating the ratio between the actual distance and the reference distance, and determining the difference between 1 and the ratio as the task progress of the robot, the quantitative representation of the task progress can be conveniently realized, and the calculation complexity of the subsequent task progress participation is reduced.

[0088] In some embodiments, the task instruction and the target recognition result are fused in step 208 to obtain visual instruction joint information, which includes:

[0089] The task instruction is encoded by a text encoder to obtain instruction features; the target recognition result is encoded by a visual encoder to obtain multi-scale spatial features; and the instruction features and the multi-scale spatial features are fused by a cross-modal attention mechanism to obtain visual instruction joint features.

[0090] The visual instruction joint information includes visual instruction joint features. The task instruction is usually a natural language instruction describing the task, such as "rotate the red knob clockwise" and "press the green button". The instruction features are features obtained by encoding the task instruction, which are used to represent the task intent and the corresponding action type. The text encoder is, for example, a Transformer encoder or a CLIP Text Encoder. The multi-scale spatial features can be used to represent the positional relationship between the structure of the task environment and the operating objects. The visual encoder is, for example, a ViT encoder, a Swin Transformer, or a DeiT. The cross-modal attention mechanism can be used to fuse features of different modalities to align corresponding features between different modalities and generate unified multi-modal feature representation. The cross-modal attention mechanism includes, for example, a cross-attention mechanism and a bidirectional attention mechanism.

[0091] In actual application scenarios, the task instruction in natural language form can be encoded by a text encoder in a visual language model (VLM) to obtain instruction features. The task image is labeled by the target recognition result, and the task labeled image is encoded by a ViT visual encoder to obtain multi-scale spatial features. The instruction features and the multi-scale spatial features are aggregated by a multi-layer cross-attention mechanism to obtain visual instruction joint features.

[0092] In this embodiment, the task operation instruction is encoded by the text encoder to obtain operation instruction features, the target recognition result is encoded by the visual encoder to obtain multi-scale spatial features, and the operation instruction features and the multi-scale spatial features are fused by the cross-modal attention mechanism to obtain visual instruction joint features, so that the visual features and the language features of different modalities can be fused to obtain unified multi-modal features, thereby laying a solid foundation for subsequent determination of the action of the robot execution component.

[0093] In some embodiments, the task operation instruction and the target recognition result are fused in step 208 to obtain visual instruction joint information, including:

[0094] The task annotation image is obtained by labeling the operation image according to the target recognition result, the semantic correspondence relationship between the task annotation image and the task operation instruction is established, and the task operation instruction and the task annotation image are aligned according to the semantic correspondence relationship to obtain the visual instruction joint information.

[0095] The task annotation image is an operation image labeled by the target recognition result. By labeling the corresponding operation image according to the target recognition result, the semantic representation in the operation image can be more clear.

[0096] For example, the task operation instruction can be divided into multiple task nodes, and the task annotation image can be divided into multiple image blocks. Each task node is sequentially matched with each image block to determine the image block matched with each task node, and the visual instruction joint information between each task node and the matched image block is obtained. In other words, the visual instruction joint information fuses the instruction semantics corresponding to the task node and the semantic information in the image block matched with the task node. It should be noted that each task node usually represents a task meaning. For example, if the task operation instruction is "press the red button", it can be divided into multiple task nodes such as "press", "red", and "button". It is easy to understand that if the semantic similarity between the task node and the image block is greater than or equal to a similarity threshold, it is determined that the task node is matched with the image block. The image block matched with each task node can include one or more. The similarity threshold can be set according to the actual application scenario.

[0097] In this embodiment, the task annotation image is obtained by labeling the operation image according to the target recognition result, the semantic correspondence relationship between the task annotation image and the task operation instruction is established, and the task operation instruction and the task annotation image are aligned according to the semantic correspondence relationship to obtain the visual instruction joint information, so that the visual perception features and the natural language instructions can be better associated.

[0098] In some embodiments, the above method further includes:

[0099] obtain a multi-view image dataset corresponding to the task working environment; label the multi-view image dataset to obtain a labeled image dataset, train the initial detection model based on the labeled image dataset to obtain a target detection model;

[0100] In step 204, the working image is recognized to obtain a target recognition result, which includes:

[0101] The working image is recognized based on the target detection model to obtain a target recognition result.

[0102] In an actual application scenario, images corresponding to the task working environment can be collected from multiple perspectives to obtain a multi-view image dataset. The multi-view image dataset is labeled by manual labeling or automatic labeling to obtain a labeled image dataset. Each labeled image in the labeled image dataset is input into an initial detection model in turn, and the features in the labeled image are detected based on the initial detection model to obtain a prediction result. The initial detection model is trained based on the difference between the prediction result and the labeled label until the convergence condition is met (for example, the difference between the prediction result and the labeled label is less than a threshold value), and a target detection model is obtained. In the process of controlling the robot to perform the target task, the working image corresponding to the task working environment can be recognized by the target detection model to obtain a target recognition result, which can improve the accuracy of the target recognition result.

[0103] In an exemplary embodiment, in order to achieve robust recognition of multiple types of components (such as buttons, knobs, pressure plates, and indicator lights) in the distribution box, a multi-view image dataset is collected under different lighting, angles, and distances, and the multi-view image dataset is labeled by manual labeling to obtain a labeled image dataset. Each image in the labeled image dataset includes a class label and a bounding box of the component. The YOLOv11 model is trained based on the labeled image dataset to obtain a target detection model. Specifically, the labeled image dataset can be input into the YOLOv11 model to obtain a prediction result. Based on the prediction result and the labeled label, a bounding box regression loss, a class classification loss, and a target existence loss are determined. Based on the bounding box regression loss, the class classification loss, and the target existence loss, a target loss is determined. The parameters of the YOLOv11 model are adjusted based on the target loss until the target loss is less than a loss threshold value, and a target detection model is obtained. The target detection model can stably recognize the panel components under conditions of multiple lighting, occlusion, and perspective changes.

[0104] Exemplarily, the collected multi-view image dataset can be data enhanced, such as increasing brightness disturbance, random cropping, rotation, affine transformation, etc., to obtain an enhanced multi-view image dataset. The enhanced multi-view image dataset is labeled to obtain a labeled image dataset.

[0105] In this embodiment, the initial detection model is trained through the labeled multi-view image dataset to obtain a target detection model. The target detection model is used to recognize the work image to obtain a target recognition result. The multi-view work image can be accurately recognized, and the accuracy of the target recognition result is improved.

[0106] In one example, as shown in Figure 3 , a robot carries a camera to obtain multi-view images (work images) of a distribution box scene, such as left view, right view, and main view. The YOLO model obtained by training is used to recognize the multi-view images to obtain a target recognition result. The target recognition result can include the component type, component bounding box, and component center point in the multi-view images. The multi-view images are labeled based on the target recognition result to obtain task labeled images. The task labeled images and task work instructions (i.e., user instructions) are input into the pre-trained VLM. The VLM generates a feature representation F VLM (i.e., visual instruction joint information) that characterizes the scene and task context. According to the positions of the components in the target recognition result and the positions of the robot execution components, the task work progress (i.e., task execution progress s) of the robot is determined. The historical pose of the robot, F VLM , the task work progress are input into the action expert model to generate the next action of the robot. The next action execution instruction is sent to the robot execution component, and the component executes the next action. The above process is repeated until the target work task corresponding to the task work instruction is completed.

[0107] The action expert model can take the Transformer structure as the core, fuse the historical pose, F VLM , and the task work progress to obtain a state vector F state , and perform a push decoding on the state vector F state to output a continuous control instruction vector . Wherein, represents the position increment of the execution component in the Cartesian space, represents the quaternion form of the output pose, which is used to describe the direction and rotation state of the execution component at the end of the mechanical part in space.

[0108] In the action output process, a limiting and dynamic constraint mechanism can be introduced to constrain physical quantities such as joint speed, acceleration, and actuator torque, to ensure that the control instruction meets the safe working range of the robot arm, and to avoid impact, shock, or overdrive phenomenon, thereby ensuring the smoothness and repeatability of the execution process. As can be easily understood, the action expert model can be trained in a supervised manner. The continuous control instruction output by the action expert model is converted into a format recognizable by the robot controller through the motion control module. The robot controller drives the joints to move, so that the end effector reaches the specified pose and performs the corresponding operation (such as pressing, rotating, or dialing, etc.). After the corresponding operation is completed, the camera can repeatedly capture multi-view images of the distribution box scene, and identify the new pose data of the execution component through the YOLO model, to determine the next action and implement the next action, until the task is completed. That is, the robot action data corresponding to different time points can be obtained .

[0109] In the above embodiments, by setting an independent task process to estimate the output task progress, and jointly executing the historical pose of the execution component to assist the action expert model in determining the action, it can be avoided that the task is determined as completed due to insignificant appearance changes, thereby improving the accuracy of task completion.

[0110] It should be understood that, although each step in the flowchart involved in each embodiment as described above is displayed in sequence according to the arrow, these steps are not necessarily executed in sequence according to the arrow. Unless otherwise specified herein, there is no strict order limitation for the execution of these steps, and these steps can be executed in other orders. Moreover, at least part of the steps in the flowchart involved in each embodiment as described above can include multiple steps or stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but can be alternately executed with at least part of other steps or steps or stages in other steps. It can be understood that the steps in different embodiments can be freely combined as needed, and various non-contradictory schemes formed by the combination are within the scope of protection of the present application.

[0111] Based on the same inventive concept, the embodiments of the present application also provide a robot operation control device for implementing the above-mentioned robot operation control method. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme described in the above method, so the specific limitations in one or more robot operation control device embodiments provided below can refer to the limitations of the robot operation control method described above, which will not be repeated here.

[0112] In one exemplary embodiment, asFigure 4 As shown, a robot operation control device 400 is provided, comprising a data acquisition module 402, an image recognition module 404, a progress determination module 406, an information fusion module 408, and an action determination module 410, wherein:

[0113] The data acquisition module 402 is configured to acquire task operation instructions and operation images corresponding to a task operation environment.

[0114] The image recognition module 404 is configured to recognize the operation images to obtain target recognition results; the target recognition results include operation object information and robot execution component information.

[0115] The progress determination module 406 is configured to determine a task operation progress of the robot based on the operation object information and the execution component information.

[0116] The information fusion module 408 is configured to fuse the task operation instructions and the target recognition results to obtain visual instruction joint information.

[0117] The action determination module 410 is configured to determine a next action of the robot according to the execution component information, the task operation progress, and the visual instruction joint information, until a target operation task corresponding to the task operation instructions is completed.

[0118] In some embodiments, the progress determination module 406 is further configured to acquire an initial position of an execution component of the robot; determine a reference distance between the execution component and an operation object based on an operation object position in the operation object information and the initial position of the execution component; determine an actual distance between the execution component and the operation object based on the operation object position in the operation object information and an actual position of the execution component in the execution component information; and determine the task operation progress of the robot based on the actual distance and the reference distance.

[0119] In some embodiments, the progress determination module 406 is further configured to calculate a ratio between the actual distance and the reference distance, and determine a difference between 1 and the ratio as the task operation progress of the robot.

[0120] In some embodiments, the information fusion module 408 is further configured to encode the task operation instructions through a text encoder to obtain operation instruction features; encode the target recognition results through a visual encoder to obtain multi-scale spatial features; and fuse the operation instruction features and the multi-scale spatial features through a cross-modal attention mechanism to obtain visual instruction joint features.

[0121] In some embodiments, the information fusion module 408 is further configured to label the task image by the target recognition result to obtain a task labeled image; establish a semantic correspondence relationship between the task labeled image and the task operation instruction, and align the task operation instruction and the task labeled image according to the semantic correspondence relationship to obtain visual instruction joint information.

[0122] In some embodiments, the device described above further comprises a model training module configured to obtain a multi-view image dataset corresponding to the task operation environment; label the multi-view image dataset to obtain a labeled image dataset; and train the initial detection model based on the labeled image dataset to obtain the target detection model.

[0123] The image recognition module 404 is further configured to recognize the task image based on the target detection model to obtain a target recognition result.

[0124] The modules in the robot operation control device described above can be realized by software, hardware, and combinations thereof, in whole or in part. The modules described above can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory in the computer device in software form, so as to be called and executed by the processor to perform the operations corresponding to the modules.

[0125] In an exemplary embodiment, a computer device is provided, which can be a server, and an internal structure diagram thereof can be as shown in Figure 5 The computer device comprises a processor, a memory, an input / output interface (I / O), and a communication interface. The processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device comprises a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The database of the computer device is configured to store related data of a robot operation control method. The input / output interface of the computer device is configured to exchange information between the processor and external devices. The communication interface of the computer device is configured to communicate with external terminals through network connection. The computer program is executed by the processor to implement a robot operation control method.

[0126] Those skilled in the art can understand that Figure 5 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can comprise more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0127] In an example embodiment, a computer device is provided, comprising a memory and a processor, the memory storing a computer program, and the processor implementing the steps in the above method embodiments when executing the computer program.

[0128] In an example embodiment, a computer readable storage medium is provided, storing a computer program, and the computer program implementing the steps in the above method embodiments when executed by a processor.

[0129] In an example embodiment, a computer program product is provided, comprising a computer program, and the computer program implementing the steps in the above method embodiments when executed by a processor.

[0130] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant regulations.

[0131] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when executed, can include the processes of the above-mentioned embodiment methods. Any reference to memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. The non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. The volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration but not limitation, the RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., without being limited thereto.

[0132] The technical features of the above embodiments can be combined in any manner. To make the description concise, all possible combinations of the technical features in the above embodiments are not described, but as long as the combinations of the technical features do not exist, they should be considered as the scope of the present application.

[0133] The above-described embodiments are merely illustrative of several embodiments of the present application, and the description is relatively specific and detailed, but should not be understood as a limitation on the scope of the patent. It should be noted that for those skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are all within the scope of the present application. Therefore, the scope of protection of the present application should be subject to the appended claims.

Claims

1. A robot operation control method, characterized in that, The method includes: Obtain the task operation instructions and the corresponding operation image of the task operation environment; The task image is identified to obtain a target recognition result; the target recognition result includes information about the task object and information about the robot's execution components. Based on the task object information and the execution component information, the task progress of the robot is determined; The task operation instructions and the target recognition results are fused to obtain joint visual instruction information; Based on the combined information of the execution components, the task progress, and the vision instructions, the robot's next action is determined until the target task corresponding to the task instruction is completed.

2. The method according to claim 1, characterized in that, Determining the robot's task progress based on the task object information and the execution component information includes: Obtain the initialization position of the robot's execution components; Based on the position of the work object in the work object information and the initialization position of the execution component, a reference distance between the execution component and the work object is determined; Based on the position of the work object in the work object information and the actual position of the execution component in the execution component information, the actual distance between the execution component and the work object is determined; The robot's task progress is determined based on the actual distance and the reference distance.

3. The method according to claim 2, characterized in that, Determining the robot's task progress based on the actual distance and the reference distance includes: Calculate the ratio between the actual distance and the reference distance, and determine the difference between 1 and the ratio as the robot's task progress.

4. The method according to claim 1, characterized in that, The step of fusing the task operation instructions and the target recognition results to obtain joint visual instruction information includes: The task instructions are encoded using a text encoder to obtain task instruction features; The target recognition results are encoded by a visual encoder to obtain multi-scale spatial features; The task instruction features and the multi-scale spatial features are fused by a cross-modal attention mechanism to obtain joint visual instruction features.

5. The method according to claim 1, characterized in that, The step of fusing the task operation instructions and the target recognition results to obtain joint visual instruction information includes: The task image is labeled using the target recognition results to obtain a task-labeled image; Establish a semantic correspondence between the task-annotated image and the task operation instruction, and align the task operation instruction and the task-annotated image according to the semantic correspondence to obtain visual instruction joint information.

6. The method according to any one of claims 1 to 5, characterized in that, The method further includes: Obtain the multi-view image dataset corresponding to the task's operating environment; The multi-view image dataset is labeled to obtain a labeled image dataset; The initial detection model is trained based on the labeled image dataset to obtain the target detection model; The process of recognizing the task image to obtain the target recognition result includes: The target detection model is used to identify the task image to obtain the target recognition result.

7. A robot operation control device, characterized in that, The device includes: The data acquisition module is used to acquire task operation instructions and corresponding operation images of the task operation environment; An image recognition module is used to recognize the work image and obtain a target recognition result; the target recognition result includes information about the work object and information about the robot's execution components; The progress determination module is used to determine the task progress of the robot based on the task object information and the execution component information; The information fusion module is used to fuse the task operation instructions and target recognition results to obtain joint visual instruction information; The action determination module is used to determine the robot's next action based on the combined information of the execution component, the task progress, and the vision instruction, until the target task corresponding to the task instruction is completed.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.