A body intelligent robot end-edge collaborative task execution method and device, electronic equipment, and storage medium
Patent Information
- Application Number
- CN202611150976.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-31
- Publication Date
- 2026-09-25
AI Technical Summary
端侧推理响应快但受限于模型容量而无法处理复杂语义,边缘侧虽具备强智能水平却因依赖网络传输而难以满足工业现场毫秒级确定性时延要求;任务分配被静态固化后,端侧在执行中遭遇视觉遮挡或力控扰动等不确定性时无法主动寻求辅助,边缘侧也无法感知端侧执行质量
[0015]本发明实施例带来了以下有益效果:本申请提供一种具身智能机器人端边协同任务执行方法、装置及电子设备、存储介质,该方法应用于包括端侧、边缘侧和协同控制器的协同控制系统,该方法包括:端侧接收作业指令,并向协同控制器发送当前的能力信息;边缘侧接收作业指令,通过边缘侧的垂直模型将作业指令分解为子任务序列,并为每个子任务生成能力需求信息,得到作业指令的任务分解结果,任务分解结果包括子任务序列以及每个子任务对应的能力需求信息;协同控制器获取边缘侧对作业指令的任务分解结果,针对每个子任务,将子任务对应的能力需求信息与端侧能力注册表进行匹配以生成任务契约,并将任务契约下发至端侧;端侧能力注册表基于端侧的能力信息建立;端侧按照任务契约执行各子任务,并在执行过程中实时计算用于表征当前执行可靠性的置信度信息,将置信度信息发送至协同控制器;协同控制器根据置信度信息判定当前执行状态,触发端侧执行相应模式,模式包括自主执行模式、端边协同执行模式或边缘侧接管执行模式中的至少一种。
Smart Images

Figure CN122807904A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of industrial robot control technology, and in particular to a method, apparatus, electronic device, and storage medium for end-to-end collaborative task execution of an embodied intelligent robot. Background Technology
[0002] Currently, industrial embodied intelligent robots in dynamic manufacturing scenarios generally adopt a layered collaborative control architecture of edge-cloud or edge-device. The robot body (edge side) deploys a lightweight AI model, responsible for high-frequency real-time perception and motion control; the workshop edge server (edge side) deploys a large-parameter model, responsible for scene understanding and global task planning. In this existing architecture, the task division and collaboration mode between the edge side and the device side are usually statically configured based on preset rules, and there is only a one-way data flow or command flow between the two.
[0003] However, in existing edge-end collaborative architectures, lightweight models are deployed on the edge side for real-time perception and control, while large-parameter models are deployed on the edge side for scene understanding and task planning. Typically, there is only a one-way data or instruction flow between the two, and task allocation and collaborative modes are statically configured before task execution. While edge-side inference has a fast response time, it is limited by model capacity and cannot handle complex semantics. Although the edge side possesses strong intelligence, its reliance on network transmission makes it difficult to meet the millisecond-level deterministic latency requirements of industrial environments. Furthermore, with task allocation statically fixed, the edge side cannot proactively seek assistance when encountering uncertainties such as visual occlusion or force control disturbances during execution, and the edge side cannot perceive the execution quality of the edge side. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a method, apparatus, electronic device, and storage medium for performing edge-to-edge collaborative tasks with an embodied intelligent robot.
[0005] In a first aspect, embodiments of the present invention provide a method for performing edge-side collaborative tasks with an embodied intelligent robot. This method is applied to a collaborative control system including an edge side, a collaborative controller, and the method includes: The terminal receives the operation command and sends the current capability information to the collaborative controller; The edge side receives the work instruction, decomposes the work instruction into a sequence of sub-tasks through the vertical model of the edge side, and generates capacity requirement information for each sub-task to obtain the task decomposition result of the work instruction. The task decomposition result includes the sequence of sub-tasks and the capacity requirement information corresponding to each sub-task. The collaborative controller obtains the task decomposition results of the operation instructions from the edge side. For each sub-task, it matches the capability requirement information corresponding to the sub-task with the capability registry of the edge side to generate a task contract and sends the task contract to the edge side. The capability registry of the edge side is established based on the capability information of the edge side. The endpoint executes each subtask according to the task contract, and calculates confidence information in real time to characterize the reliability of the current execution during the execution process, and sends the confidence information to the collaborative controller; The collaborative controller determines the current execution status based on confidence information and triggers the corresponding execution mode on the edge. The mode includes at least one of autonomous execution mode, edge collaborative execution mode, or edge takeover execution mode.
[0006] In conjunction with the first aspect, the steps of matching the capability requirement information corresponding to the subtask with the end-side capability registry to generate a task contract include: Get the matching results; If the matching result is a complete match, the task contract is configured to allow the subtask to be executed autonomously by the client. If the matching result is a partial match, the task contract is configured as a collaborative execution mode in which the subtask performs basic operations on the end side and provides auxiliary information from the edge side; If the matching result is a mismatch, the task contract is configured to a takeover execution mode where the subtask is executed directly by the edge side.
[0007] In conjunction with the first aspect, the capability information should include at least the currently available model types, model inference confidence, computing load status, and sensor health status on the edge side; the capability requirement information should include at least the sensing type, accuracy requirements, response latency limit, and computing power requirements required to execute the corresponding sub-task.
[0008] In conjunction with the first aspect, the confidence information is the comprehensive execution confidence, which integrates three dimensions: perception confidence, tracking confidence, and anomaly confidence. The steps involved in real-time calculation of confidence information characterizing the current execution reliability during execution include: The perception confidence is determined based on the information entropy value of the category probability distribution output by the perception model and the spatial overlap of the detection boxes in the detection results of consecutive frames. The information entropy value is used to characterize the uncertainty of the perception model's output results, and the spatial overlap is used to characterize the stability of the detection boxes between adjacent frames. The tracking confidence level is determined based on the positional offset between the robot's actual trajectory and the planned trajectory. The degree of anomaly confidence is determined based on the deviation between the real-time feedback value of the sensor and the expected reference value under normal operating conditions. The execution confidence is obtained by weighted summation of the perception confidence, tracking confidence, and anomaly confidence; wherein the sum of the preset weight coefficients corresponding to each confidence is equal to one.
[0009] In conjunction with the first aspect, the collaborative controller determines the current execution state based on confidence information and triggers the execution of the corresponding mode on the endpoint, including the following steps: When the overall execution confidence level is greater than the first threshold, the edge-side autonomous execution mode is triggered, while the edge side only maintains monitoring. When the overall execution confidence is greater than the second threshold and less than or equal to the first threshold, the edge-end collaborative execution mode is triggered, and the edge requests semantic assistance or path replanning from the edge side. When the overall execution confidence is less than or equal to the second threshold, the edge side takeover execution mode is triggered, the end side suspends execution and the edge side issues control commands. The first threshold is greater than the second threshold.
[0010] In conjunction with the first aspect, the method also includes: monitoring the communication link quality at the edge side, and when the predicted communication quality is lower than a preset threshold, pre-configured motion primitives are pushed to the end side for local storage, so that the end side can call the motion primitives to perform tasks after the network is interrupted.
[0011] In conjunction with the first aspect, the steps of monitoring the communication link quality at the edge and pre-pushing pre-configured motion primitives to the end side include: The edge side uses extended Kalman filtering to predict the robot's pose in the next time window and simultaneously monitors the signal quality indicators of the current communication link. When it is predicted that the signal quality index will be lower than the communication quality threshold after a preset time window, the edge side will push the motion primitives corresponding to the current subtask and subsequent subtasks to the end side for local storage. After a network interruption, the device switches to autonomous mode without network access and continues execution by calling locally stored motion primitives. After the network is restored, the endpoint uploads the execution logs from the interruption period to the edge, and the edge performs offline analysis and incremental model updates based on the execution logs.
[0012] Secondly, this application provides an end-to-end collaborative task execution device for an embodied intelligent robot. This device is applied to a collaborative control system including an end-side component, an edge-side component, and a collaborative controller. The device includes: The capability negotiation module enables the endpoint to receive work instructions, enables the edge side to receive work instructions, enables the endpoint to send its capability information to the collaborative controller, enables the edge side to decompose the work instructions into a sequence of sub-tasks using its vertical model and generate capability requirement information for each sub-task, and enables the collaborative controller to obtain the task decomposition results from the edge side. For each sub-task, the module matches the capability requirement information corresponding to the sub-task with the endpoint capability registry to generate a task contract and send it to the endpoint. The endpoint capability registry is built based on the endpoint capability information. The execution and feedback module is used to enable the edge to execute each sub-task according to the task contract, and to calculate the confidence information to characterize the current execution reliability in real time during the execution process, and send the confidence information to the collaborative controller; The gating arbitration module is used to enable the collaborative controller to determine the current execution status based on confidence information and trigger the corresponding mode among the end-side autonomous execution mode, end-edge collaborative execution mode, or edge-side takeover execution mode. The prediction degradation module is used to enable the edge side to monitor the quality of the communication link. When the predicted communication quality is lower than a preset threshold, the pre-configured motion primitives are pushed to the edge side for local storage, so that the edge side can call the motion primitives to perform tasks after the network is interrupted.
[0013] Thirdly, this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor runs the computer program to cause the electronic device to perform the above-described method.
[0014] Fourthly, this application provides a readable storage medium storing computer program instructions, which are read and executed by a processor to perform the above-described method.
[0015] The embodiments of the present invention bring the following beneficial effects: This application provides a method, device, electronic device, and storage medium for edge-side collaborative task execution of an embodied intelligent robot. The method is applied to a collaborative control system including an edge side, an edge side, and a collaborative controller. The method includes: the edge side receiving a work instruction and sending current capability information to the collaborative controller; the edge side receiving the work instruction, decomposing the work instruction into a sequence of sub-tasks through the vertical model of the edge side, and generating capability requirement information for each sub-task to obtain a task decomposition result of the work instruction, the task decomposition result including the sub-task sequence and the capability requirement information corresponding to each sub-task; the collaborative controller obtaining the task decomposition result of the work instruction from the edge side, matching the capability requirement information corresponding to each sub-task with the capability registry of the edge side to generate a task contract for each sub-task, and sending the task contract to the edge side; the edge-side capability registry is established based on the capability information of the edge side; the edge side executes each sub-task according to the task contract, and calculates confidence information to characterize the reliability of the current execution in real time during the execution process, and sends the confidence information to the collaborative controller; the collaborative controller determines the current execution state according to the confidence information and triggers the edge side to execute the corresponding mode, the mode including at least one of autonomous execution mode, edge-side collaborative execution mode, or edge side takeover execution mode.
[0016] This invention achieves bidirectional dynamic collaboration between the end-side and the edge-side through a closed-loop technology that involves the end-side sending capability information to the collaborative controller, the collaborative controller matching and generating task contracts, the end-side executing and providing confidence feedback, the collaborative controller implementing hierarchical gating arbitration, and the edge-side predictively pushing motion primitives. This simultaneously meets the stringent requirements of industrial manufacturing scenarios in terms of real-time performance, intelligence, and robustness.
[0017] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention are realized and obtained in accordance with the structures particularly pointed out in the description, claims and drawings.
[0018] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0019] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0020] Figure 1 A schematic diagram of a three-layer edge-edge collaborative architecture used in an embodiment of the present invention for an embodied intelligent robot edge-edge collaborative task execution method; Figure 2 This is a flowchart of the bidirectional capability contract generation and gated arbitration in the method provided in the embodiments of the present invention; Figure 3 This is a state machine switching diagram for the five cooperative modes in the method provided in the embodiments of the present invention; Figure 4 This is a timing diagram of the terminal-edge-physical interaction in the method provided in the embodiments of the present invention; Figure 5 A flowchart illustrating an edge-to-edge collaborative task execution method for an embodied intelligent robot, provided as an embodiment of the present invention; Figure 6 This is a schematic diagram of the structure of an edge-to-edge collaborative task execution device for an embodied intelligent robot, provided in an embodiment of the present invention. Figure 7 This is a schematic diagram of the electronic device structure provided in an embodiment of the present invention.
[0021] Figure label: 10-Capability Negotiation Module, 20-Execution and Feedback Module, 30-Gated Arbitration Module, 40-Predictive Degradation Module; 130 - Processor, 131 - Memory, 132 - Bus, 133 - Communication interface. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0023] To facilitate understanding of this embodiment, the technical terms used in this application will be briefly introduced below.
[0024] Embossed intelligent robots: refer to intelligent robot systems that have a physical body, can interact with the physical environment in real time through sensors, and continuously optimize their own behavioral decisions in the process of interaction.
[0025] End-side: This refers to the robot body side of the embodied intelligent robot, where a cluster of small AI models is deployed to perform high-frequency real-time perception and motion control, and to evaluate the local execution confidence in real time.
[0026] Edge side: This refers to the edge server side of the workshop, where vertical models are deployed to perform scene semantic understanding, task decomposition, dynamic skill anchor management, and communication quality prediction.
[0027] The bidirectional collaborative controller (hereinafter referred to as the collaborative controller) is deployed at the edge or independently of the edge but communicatively connected to it. It is used to perform capability matching, task contract generation, confidence-gated arbitration, and predictive scheduling decisions. Internally, it includes a capability registration and contract generation module, a confidence-gated arbitration module, a skill anchor management module, and a predictive scheduling module.
[0028] Capability Information: Information sent from the edge device to the collaborative controller, describing the sensing and computing resources available at the current moment, including at least the currently available model type, model inference confidence, computing load status, and sensor health status.
[0029] Capability requirement information: Information generated by the edge side for each subtask during the task decomposition process, used to describe the capability conditions required to execute the corresponding subtask, including at least the perception type, accuracy requirements, response latency limit, and computing power requirements required to execute the subtask.
[0030] End-side capability registry: A data structure established and maintained by the collaborative controller based on the capability information reported by the end-side. It is used to store and manage real-time capability data from the end-side, so that the collaborative controller can query and compare it when performing capability matching.
[0031] Task contract: A data structure generated by the collaborative controller based on the capability matching results, which defines the execution subject and collaborative mode of each sub-task. It is sent to the edge side and serves as the basis for the execution of each sub-task on the edge side.
[0032] Confidence information: This information is calculated in real time by the edge device and sent to the collaborative controller during the execution of subtasks. It is used to characterize the reliability of the current execution on the edge device. This information is a comprehensive execution confidence score, which integrates three dimensions: perception confidence score, tracking confidence score, and anomaly confidence score. The higher the comprehensive execution confidence score, the higher the execution reliability.
[0033] Perception confidence: A sub-dimension of comprehensive execution confidence, used to characterize the output reliability of the edge perception model in the current operating environment. It is determined based on the information entropy value of the category probability distribution output by the perception model and the spatial overlap of the detection boxes in the detection results of multiple consecutive frames.
[0034] Tracking confidence: A sub-dimension of comprehensive execution confidence, used to characterize the tracking accuracy of the planned trajectory by the end-side during motion control execution, and determined based on the positional offset between the robot's actual motion trajectory and the planned trajectory.
[0035] Anomaly confidence: A sub-dimension of overall execution confidence, used to characterize whether each sensor and actuator on the end side is in normal working condition. It is determined based on the degree of deviation between the real-time feedback value of the sensor and the expected reference value under normal operating mode.
[0036] Motion primitives: These are pre-trained motion primitives (skill anchors) stored in the skill anchor management module on the edge side. They are pre-configured control parameter sequences for the robot to perform specific actions. After obtaining the motion primitives, the edge side can autonomously generate control commands based on them without real-time guidance from the edge side.
[0037] Vertical Model: A scene model deployed at the edge. It is a large language model or visual language model that is finely tuned for a specific manufacturing scenario. The number of parameters ranges from 1B to 7B. It is used for semantic understanding and task decomposition of work instructions.
[0038] AI Small Model Cluster: A collection of lightweight AI models deployed on the edge, including but not limited to lightweight object detection models, local obstacle avoidance models, force control compliant models, and pose estimation models, with a total number of parameters of less than 100 megabytes and an inference latency of less than 10 milliseconds.
[0039] Extended Kalman Filter (EPF): A recursive state estimation algorithm used at the edge to predict the signal quality index of the communication link. Based on the current signal quality observation and its changing trend, it predicts the signal quality state at future times, providing a predictive basis for early triggering of motion element push.
[0040] After introducing the technical terms used in this application, the application scenarios and design concepts of the embodiments of this application will be briefly described below.
[0041] Existing end-edge collaboration solutions suffer from problems such as static unidirectional task allocation, the inability of the edge side to perceive the execution quality of the end side, and a sharp degradation of intelligent capabilities when the network is interrupted, making it difficult to simultaneously satisfy real-time performance, intelligence, and robustness.
[0042] Based on this, this application provides a method, apparatus, electronic device, and storage medium for performing edge-to-edge collaborative tasks with an embodied intelligent robot.
[0043] Example 1 This application provides a method for edge-side collaborative task execution of an embodied intelligent robot. This method is applied to a collaborative control system including an edge-side controller, a terminal controller, and a collaborative controller (such as...). Figure 1 As shown in the figure, the collaborative control system adopts a three-layer end-edge collaborative architecture.
[0044] The first layer is the robot itself (edge device). As an embodied intelligent agent, the edge device continuously performs perception, motion control, and task execution in the physical environment. During execution, it generates confidence information to assess its own execution reliability, driving dynamic collaboration with the edge device. The edge device deploys a cluster of AI mini-models, including a lightweight object detection model, a local obstacle avoidance model, a force control compliant model, and a pose estimation model, used for high-frequency real-time perception and motion control. The edge device is used for high-frequency real-time perception (response time less than 10 milliseconds), motion control, and real-time evaluation of execution confidence. The total number of parameters in the AI mini-model cluster is less than 100 megabytes, and the inference latency is less than 10 milliseconds, ensuring that the edge device can meet the millisecond-level response requirements for control commands in industrial scenarios. Furthermore, the edge device is also used to send capability information to the cooperative controller and, during the execution of the task contract, calculates confidence information in real time to characterize the current execution reliability, sending the confidence information to the cooperative controller.
[0045] The second layer is the workshop edge side (edge side). The edge side deploys a vertical model, which is a scenario model fine-tuned for a specific manufacturing scenario, with parameters ranging from 1B to 7B. The core responsibilities of the edge side include scenario semantic understanding, task decomposition, dynamic skill anchor management, and communication quality prediction. The edge side receives work instructions, decomposes them into a sequence of sub-tasks using the vertical model, and generates capability requirement information for each sub-task. Furthermore, the edge side monitors communication link quality; when the predicted communication quality falls below a preset threshold, it pre-pushes pre-configured motion primitives (i.e., skill anchors) to the edge side for local storage, so that the edge side can access them after a network outage.
[0046] The third layer is a bidirectional collaborative controller, deployed at the edge or independently of the edge but communicating with it. The bidirectional collaborative controller comprises four sub-modules: a capability registration and contract generation module, a confidence-gated arbitration module, a skill anchor management module, and a predictive scheduling module. The capability registration and contract generation module maintains the edge-side capability registry, obtains the task decomposition results of job instructions from the edge side, matches the task decomposition results with the edge-side capability information to generate a task contract, and then sends the task contract to the edge side. The confidence-gated arbitration module receives confidence information reported by the edge side, determines the current execution state based on the confidence information, and triggers the corresponding mode among edge-side autonomous execution mode, edge-edge collaborative execution mode, or edge-side takeover execution mode. The skill anchor management module stores and manages pre-trained motion primitives. The predictive scheduling module executes the position prediction algorithm and triggers the advance push of skill anchors.
[0047] In this collaborative control system, the key data flows between the endpoint, edge, and collaborative controller include: the endpoint sending capability information to the collaborative controller (step S110 capability reporting); the collaborative controller issuing task contracts to the endpoint (step S120 contract issuance); the endpoint sending confidence information to the collaborative controller (step S130 confidence feedback); and the edge pushing motion primitives to the endpoint (step S150 anchor pushing). Through the above three-layer architecture and data flow interaction, bidirectional dynamic collaboration between the endpoint and edge is achieved.
[0048] Combination Figure 5 As shown, the method includes: S110: The terminal receives the work instruction and sends the current capability information to the cooperative controller.
[0049] S120, the edge side receives the operation instruction, decomposes the operation instruction into a sequence of sub-tasks through the vertical model of the edge side, and generates capacity requirement information for each sub-task; the task decomposition result includes the sequence of sub-tasks and the capacity requirement information corresponding to each sub-task.
[0050] S130, the collaborative controller obtains the task decomposition results of the operation instructions from the edge side, matches the capability requirement information and capability information corresponding to each sub-task to generate a task contract, and sends the task contract to the edge side.
[0051] S140, the terminal executes each sub-task according to the task contract, and calculates confidence information in real time to characterize the current execution reliability during the execution process, and sends the confidence information to the cooperative controller.
[0052] S150, the collaborative controller determines the current execution status based on the confidence information and triggers the corresponding mode among the end-side autonomous execution mode, end-edge collaborative execution mode, or edge-side takeover execution mode.
[0053] S160 monitors the quality of the communication link at the edge. When the predicted communication quality is lower than a preset threshold, it pushes the pre-configured motion primitives to the edge for local storage, so that the edge can call the motion primitives to perform tasks after the network is interrupted.
[0054] This application achieves bidirectional dynamic collaboration between the end-side and the edge-side by transmitting capability information from the end-side, generating task contracts through matching by the collaborative controller, executing and feeding back confidence from the end-side, gating arbitration by the collaborative controller, and predictively pushing motion primitives from the edge-side. It can obtain semantic assistance from the edge-side on demand under high real-time requirements, automatically trigger hierarchical responses when execution confidence decreases, and pre-caching motion primitives to maintain core functions before network interruption. Thus, it simultaneously meets the stringent requirements of industrial manufacturing scenarios in terms of real-time performance, intelligence, and robustness.
[0055] Step S110 is the start step of the entire edge-edge collaborative task execution process. The end side (robot body) and the edge side respectively receive the operation instructions from the scheduling system. At the same time, the end side sends capability information to the collaborative controller to trigger the subsequent edge side task decomposition and collaborative controller contract generation.
[0056] Specifically, combined Figure 2 As shown, the endpoint receives a job instruction from the scheduling system. This job instruction contains the task content to be executed and related parameters. Upon receiving the job instruction, the endpoint confirms the existence of a task to be executed, and the process begins. The edge side also receives job instructions from the scheduling system, providing a data foundation for subsequent semantic understanding and task decomposition of the job instructions by the edge side using a vertical model.
[0057] Simultaneously, the edge device sends its current capability information to the cooperative controller. This capability information describes the perception and computing resources available to the edge device at the current moment, including at least the currently available model type (e.g., lightweight object detection model, local obstacle avoidance model, force control compliant model, or pose estimation model), model inference confidence, computing load status, and sensor health status. These multiple pieces of information collectively form the basis for the cooperative controller's subsequent judgments regarding capability matching and task contract generation.
[0058] It should be noted that there is no strict sequential constraint between the actions of receiving the work instruction and sending capability information to the collaborative controller; both can be executed concurrently after the work instruction is received. After completing the above actions, the terminal enters a waiting state, waiting for the collaborative controller to issue a task contract so that the execution entity and collaborative mode specified in the contract can be used to initiate the execution of the sub-task.
[0059] Through the above step S110, the terminal side, as the process initiator, receives the work instruction and sends its own capability information to the collaborative controller to drive contract generation. The edge side receives the work instruction to drive task decomposition, providing complete information input for the collaborative controller to perform capability matching and contract generation in the subsequent step S120.
[0060] After receiving the job instruction (which is directly issued to the edge by the scheduling system), the edge side performs semantic understanding and task decomposition on the job instruction through the vertical model deployed on the edge side, providing structured task decomposition results for the subsequent matching of the execution capabilities of the collaborative controller.
[0061] Specifically, the edge side first receives the job instruction, which is a raw task description issued by the scheduling system or operator. It is usually in the form of natural language or structured instructions, and its content is relatively macroscopic, making it difficult to directly map to specific execution actions on the edge side. Therefore, the edge side needs to call the vertical model deployed on the edge side to perform semantic understanding and structured decomposition of the job instruction.
[0062] The vertical model deployed at the edge is a scenario model fine-tuned for specific manufacturing scenarios, with a parameter count ranging from 1B to 7B. Compared to the cluster of small AI models deployed at the end, this vertical model possesses stronger semantic understanding and logical reasoning capabilities. It can deeply analyze work instructions and break them down into several sub-tasks that can be executed independently or collaboratively by the end or edge. For example, for the work instruction "go to workstation A to pick up a workpiece and move it to workstation B", the vertical model can decompose it into multiple sub-task sequences such as "navigate to workstation A", "target recognition and localization", "force-controlled gripping", "navigate to workstation B", and "force-controlled placement".
[0063] After the edge side completes the decomposition of the sub-task sequence, it further generates corresponding capability requirement information for each sub-task (combined with...). Figure 2 The capability requirement vector is generated within the context of the system. This capability requirement information describes the capability conditions required to execute the corresponding subtask, including at least the perception type, accuracy requirement, response latency limit, and computing power requirement. Specifically, the perception type indicates the type of perception capability required to execute the subtask, such as target detection, local obstacle avoidance, or force perception; the accuracy requirement indicates the specific indicators of perception results or control accuracy for the subtask; the response latency limit indicates the real-time requirement for control command feedback from the edge device when executing the subtask, directly related to whether the edge device executes autonomously or requires edge assistance; and the computing power requirement indicates the computing resource requirements of the edge device for the subtask. These multiple pieces of information collectively constitute the matching benchmark for the collaborative controller when performing capability matching.
[0064] After the edge side completes the above decomposition and generation operations, it obtains the task decomposition result of the job instruction. The task decomposition result contains a complete sequence of sub-tasks and the capability requirement information corresponding to each sub-task.
[0065] In conjunction with the first aspect, before the collaborative controller obtains the task decomposition results of the job instructions from the edge side in step S120, the following steps are also included: S1201, the edge side receives the operation instruction, decomposes the operation instruction into a sequence of sub-tasks through the vertical model of the edge side, and generates capacity requirement information for each sub-task.
[0066] Specifically, step S120 involves the task decomposition of the operation instructions on the edge side, the capability matching of the collaborative controller, and the generation and issuance of task contracts.
[0067] Before the collaborative controller obtains the task decomposition results from the edge side, the edge side performs semantic understanding and task decomposition on the job instruction received in step S110 using a vertical model deployed on the edge side.
[0068] After acquiring the work instructions at the edge, a vertical model deployed at the edge (a scenario model finely tuned for a specific manufacturing scenario, with 1B to 7B parameters) performs semantic understanding and task decomposition on the work instructions. The work instructions are broken down into several sub-task sequences, and corresponding capability requirement information is generated for each sub-task. This capability requirement information describes the capability conditions required to execute the corresponding sub-task, including at least the type of perception required (such as object detection, obstacle avoidance, force control, etc.), accuracy requirements, response latency limit, and computing power requirements.
[0069] The edge side completes the above task decomposition and generates capability requirement information (combined with...) Figure 3 After generating the capability requirement vector, the collaborative controller obtains the task decomposition result. The collaborative controller matches the capability requirement information corresponding to each subtask in the task decomposition result with the capability information stored in the edge capability registry to determine whether the current capabilities of the edge meet the execution requirements of each subtask.
[0070] The matching results include three scenarios: complete match, partial match, and no match. Specifically, when the capability information in the edge capability registry completely covers the capability requirements of the subtask (i.e., the currently available model types, model inference confidence, computing load status, and sensor health status on the edge can all meet the capability conditions required by the subtask), the matching result is a complete match; when the capability information can cover the basic part of the capability requirements but cannot meet all the requirements, the matching result is a partial match; when the capability information cannot meet the capability requirements at all, the matching result is a no match.
[0071] The collaborative controller generates a task contract for each subtask based on the matching results. The task contract defines the execution subject and collaborative mode of each subtask. In the task contract, a complete match corresponds to a mode in which the edge side executes autonomously; a partial match corresponds to a collaborative execution mode in which the edge side performs basic operations and provides auxiliary information (such as scene semantic interpretation, local path replanning, etc.); and a non-match corresponds to a takeover execution mode in which the edge side executes directly.
[0072] After the task contract is generated, the coordination controller sends the task contract to the edge device. Once the edge device receives the task contract, it can start the execution process of each sub-task according to the execution subject and coordination mode specified in the task contract.
[0073] Through the above step S120, a two-way capability contract is established between the endpoint, the edge, and the collaborative controller. The endpoint no longer passively waits for instructions, and the edge no longer statically fixes task allocation. Instead, it dynamically matches the real-time capability information of the endpoint with the task decomposition results of the edge, realizing a technical transformation from "one-way instruction transmission" to "two-way contract negotiation".
[0074] In conjunction with the first aspect, step S130 involves matching the capability requirement information corresponding to the subtask with the end-side capability registry to generate a task contract, specifically including: S131, obtain the matching result.
[0075] S132, if the matching result is a complete match, the task contract is configured to a mode in which the subtask is executed autonomously by the end side.
[0076] S133, if the matching result is a partial match, the task contract is configured as a collaborative execution mode in which the subtask performs basic operations on the end side and provides auxiliary information on the edge side.
[0077] S134, if the matching result is a mismatch, the task contract is configured to a takeover execution mode in which the subtask is executed directly by the edge side.
[0078] Step S130 is based on the fact that the edge side has sent capability information to the collaborative controller in step S110, and that the edge side has completed the decomposition of job instructions and generated task decomposition results in step S120. The collaborative controller generates a task contract that defines the execution subject of each sub-task by matching and comparing the actual capabilities of the edge side with the capability requirements of each sub-task item by item, and sends it to the edge side, thereby completing the establishment and negotiation of the two-way capability contract.
[0079] Specifically, the collaborative controller establishes an edge-side capability registry based on the acquired capability information. This registry is a data structure maintained internally by the collaborative controller to store and manage real-time capability data from the edge. Upon receiving capability information from the edge, the collaborative controller records it in the registry and maintains and refreshes it based on periodic updates or status changes from the edge, ensuring that the capability data in the registry remains consistent with the actual state of the edge. The establishment of the edge-side capability registry enables the collaborative controller to quickly query the various capability indicators currently possessed by the edge using a standardized data structure, providing an efficient query and comparison basis for subsequent capability matching.
[0080] After establishing the edge-side capability registry, the collaborative controller further obtains the task decomposition results generated on the edge side in step S120. These results include a complete sequence of subtasks and the capability requirements for each subtask. For each subtask, the collaborative controller compares the capability requirements with the capability information stored in the edge-side capability registry item by item.
[0081] The capability requirement information includes at least the perception type, accuracy requirements, response latency limit, and computing power requirements required to execute the subtask. The capability information in the edge capability registry includes at least the currently available model type, model inference confidence, computing power load status, and sensor health status.
[0082] The collaborative controller compares the corresponding fields of the two sets of information one by one. For example, it compares the required perception type with the available model type, the accuracy requirement with the model inference confidence, the computing power requirement with the computing power load status, and the response latency limit with the comprehensive capability assessment, in order to determine whether the current capability of the edge side meets the execution requirements of each sub-task.
[0083] The matching results are divided into three categories: complete match, partial match, and no match.
[0084] When the matching result is a perfect match, it means that the capability information in the edge capability registry completely covers the capability requirements of the subtask. Specifically, the currently available model types on the edge cover the perception types required by the subtask, the model inference confidence meets the accuracy requirements of the subtask, and the computing power load status meets the computing power requirements of the subtask and can meet the response latency limit. In this case, the task contract in step S132 configures the execution subject of the subtask as the edge, and the execution mode is edge autonomous execution. After receiving the contract, the edge independently completes all perception and control operations of the subtask, while the edge only monitors without actively interfering.
[0085] When the matching result is a partial match, it means that the capability information in the edge capability registry can cover the basic part of the subtask capability requirement information, but cannot meet all the requirements. Specifically, the edge currently has some capabilities required to execute the subtask, such as being able to perform basic perception and control actions, but it cannot meet the accuracy requirements or response latency limits of the subtask in terms of accuracy or semantic understanding. In this case, step S133, the task contract, configures the execution entity of the subtask to be jointly executed by the edge and the edge, that is, the edge performs basic perception and control actions and is responsible for high-frequency real-time response; the edge provides auxiliary information when the edge makes a request, including but not limited to scene semantic interpretation and local path replanning. This achieves complementarity between the real-time performance of the edge and the intelligence level of the edge.
[0086] When the matching result is a mismatch, it indicates that the capability information in the edge-side capability registry completely fails to meet the capability requirements of the subtask. Specifically, one or more of the currently available model types, model inference confidence, computing power load status, and sensor health status on the edge side cannot meet the basic execution conditions of the subtask. In this case, step S134, the task contract configures the execution subject of the subtask as the edge side, and the execution mode is direct execution by the edge side (i.e., takeover execution). The edge side suspends execution and waits for the edge side to generate and issue control commands to take over motion control. The edge side directly generates control commands and issues them to the edge-side actuator, realizing direct control of the robot's motion by the edge side.
[0087] After the task contract is generated, the coordination controller sends the task contract to the edge device. Upon receiving the task contract, the edge device learns the execution subject and coordination mode of each subtask, and then starts the execution process of each subtask according to the execution method specified in the contract.
[0088] Through step S130 above, the collaborative controller establishes the capability information reported by the endpoint as an endpoint capability registry, and matches the capability requirement information generated by the edge side with the endpoint capability registry. The matching result is then solidified into the execution mode of each sub-task in the form of a task contract, thus realizing a complete contract establishment closed loop from endpoint capability reporting, edge requirement generation to execution mode determination. This step ensures that the execution subject of each sub-task is determined based on the dynamic matching result between the current actual capability of the endpoint and the actual requirement of the sub-task. This avoids the technical defects in existing technologies where task allocation is pre-statically fixed, leading to the endpoint's inability to actively seek assistance when encountering uncertainty during execution, and the edge side's inability to perceive the endpoint's execution quality. It achieves a technological shift from one-way instruction transmission to two-way contract negotiation.
[0089] In conjunction with the first aspect, the capability information should include at least the currently available model types, model inference confidence, computing load status, and sensor health status on the edge side; the capability requirement information should include at least the sensing type, accuracy requirements, response latency limit, and computing power requirements required to execute the corresponding sub-task.
[0090] In this application, capability information is generated by the edge device and sent to the cooperative controller, describing the sensing and computing resources available on the edge device at the current moment. Specifically, the capability information includes at least the following: The currently available model type on the edge indicates the types of AI models currently deployed and running on the edge, such as lightweight object detection models, local obstacle avoidance models, force control compliant models, or pose estimation models. This information is used to enable the collaborative controller to know what specific perception and processing capabilities the edge has. Model inference confidence is used to characterize the reliability of each model's output in the current operating environment. This information is used to enable the collaborative controller to assess the trust level of the edge-side perception results. The computing load status reflects the current computing resource occupancy on the edge side. This information is used by the collaborative controller to determine whether the edge side has the remaining computing power to undertake additional subtasks. Sensor health status indicates whether various sensors configured on the edge (such as vision sensors, force sensors, odometers, etc.) are functioning properly. This information is used by the collaborative controller to confirm whether the execution conditions on the edge are complete. These various pieces of information together constitute the basis for the collaborative controller to determine what the edge can do.
[0091] Capability requirement information is generated by the edge computing side for each subtask during task decomposition, and describes the capability conditions required to execute the corresponding subtask. Specifically, capability requirement information includes at least the following: The required perception type indicates the type of perception capability required to perform the subtask, such as object detection, local obstacle avoidance, or force perception. This information is used to enable the collaborative controller to determine whether the end side has a matching model type. Accuracy requirements indicate specific metrics for the perception results or control accuracy of the subtask. This information is used to enable the collaborative controller to determine whether the confidence level of the model inference on the end side meets the task requirements. The upper limit of response latency indicates the real-time requirement of the control command feedback when the edge executes the subtask. This information is used to enable the collaborative controller to determine whether the computing power load status of the edge can meet the latency constraint. The computing power requirement indicates the computing resources required by the subtask on the edge side. This information is used by the collaborative controller to determine whether the current remaining computing power on the edge side is sufficient to support the execution of the subtask. All of the above information together constitute the basis for the collaborative controller to determine what capabilities the subtask requires.
[0092] After acquiring capability information and capability requirement information, the collaborative controller compares them item by item to obtain matching results (complete match, partial match, or no match). Based on the matching results, it determines the execution entity and collaborative mode for each subtask. Therefore, the fields contained in capability information and capability requirement information are not isolated parameter lists, but rather corresponding matching criteria. Capability information describes the support that the endpoint can provide, while capability requirement information describes the capabilities required by the subtask. These two information form a multi-dimensional correspondence between perception type and model type, accuracy requirements and model inference confidence, computing power requirements and computing power load status, and response latency limits and computing power load status, jointly supporting the collaborative controller in making accurate capability matching decisions.
[0093] In conjunction with the first aspect, the confidence information is the comprehensive execution confidence, which integrates three dimensions: perception confidence, tracking confidence, and anomaly confidence. In step S140, confidence information characterizing the current execution reliability is calculated in real time during execution, specifically including: S141, determine the perception confidence based on the information entropy value of the category probability distribution output by the perception model and the spatial overlap of the detection boxes in the detection results of consecutive frames; wherein, the information entropy value is used to characterize the degree of uncertainty of the perception model in its output results, and the spatial overlap is used to characterize the stability of the detection boxes between adjacent frames.
[0094] S142, determine the tracking confidence level based on the positional offset between the robot's actual motion trajectory and the planned trajectory.
[0095] S143, determine the anomaly confidence level based on the degree of deviation between the real-time feedback value of the sensor and the expected reference value under normal operating mode.
[0096] S144, the perception confidence, tracking confidence and anomaly confidence are weighted and summed to obtain the execution confidence; wherein, the sum of the preset weight coefficients corresponding to each confidence is equal to one.
[0097] Perception confidence (C_perc) is used to characterize the reliability of the edge perception model's output in the current operating environment. The determination of perception confidence comprehensively considers two dimensions: the uncertainty of the perception model's output itself and the temporal consistency of the perception results.
[0098] Regarding the uncertainty dimension, when a perceptual model performs object detection or scene recognition on an input image, its output is a probability distribution for each category. When the model's output probability for a certain category is significantly higher than that for other categories, it indicates that the model's judgment of that category has high certainty; when the probability distributions for each category are relatively dispersed, it indicates that the model's judgment of that category has significant uncertainty. This application quantifies the aforementioned degree of uncertainty by calculating the information entropy value of this probability distribution. The smaller the information entropy value, the more concentrated the model's output probability distribution, and the lower the uncertainty of the model's output results; the larger the information entropy value, the more dispersed the model's output probability distribution, and the higher the uncertainty of the model's output results.
[0099] Regarding temporal consistency, the spatial overlap of detection boxes output by the perceptual model for the same target across consecutive frames reflects the stability of the perception results. When the positions of detection boxes output by the perceptual model for the same target are basically consistent across adjacent frames, it indicates good stability of the perception results; when there are large jumps in the positions of detection boxes between adjacent frames, it indicates that the perception results are jittery or have false detections. This application determines the aforementioned spatial overlap by calculating the intersection-over-union ratio or center point offset of the detection boxes in consecutive frames. The higher the overlap, the higher the stability of the detection boxes between adjacent frames; the lower the overlap, the greater the fluctuation between adjacent frames.
[0100] The aforementioned information entropy value and spatial location overlap together constitute the basis for determining perceptual confidence. That is, the lower the information entropy value and the higher the spatial location overlap, the higher the perceptual confidence; conversely, the higher the information entropy value and the higher the spatial location overlap, the lower the perceptual confidence.
[0101] Step S142, tracking confidence (C_track), characterizes the tracking accuracy of the edge device on the planned trajectory during motion control. When executing sub-tasks, the edge device's motion control system generates control commands according to the planned trajectory and drives the robot's movement. Simultaneously, it estimates the robot's actual pose in real time using sensors such as odometry and inertial measurement units. The positional offset between the actual trajectory and the planned trajectory reflects the execution accuracy of the edge device's motion control; that is, the smaller the offset, the closer the robot's actual movement matches the planned path, and the higher the motion control accuracy; conversely, the larger the offset, the greater the deviation of the robot's actual movement from the planned path, and the lower the motion control accuracy. This application determines the tracking confidence by calculating the positional offset (e.g., root mean square error) between the robot's actual trajectory and the planned trajectory; that is, the smaller the offset, the higher the tracking confidence; conversely, the larger the offset, the lower the confidence.
[0102] Step S143, Anomaly Confidence (C_anom), is used to characterize whether the sensors and actuators on the end-side are in normal working condition. During the execution of sub-tasks, various sensors on the end-side, such as force sensors, vision sensors, and joint status sensors, provide real-time feedback data. Each sensor has a corresponding expected reference value range under normal operating mode. When the real-time feedback value of the sensor continuously deviates from the expected reference value, it indicates that the sensor or actuator may have an abnormal condition (e.g., abnormal force control, loss of visual signal, joint overload, etc.). This application determines the anomaly confidence by calculating the degree of deviation (e.g., difference or ratio) between the real-time feedback value of the sensor and the expected reference value under normal operating mode. That is, the smaller the deviation, the more normal the system operation, and the higher the anomaly confidence; the larger the deviation, the higher the probability of an anomaly in the system, and the lower the anomaly confidence.
[0103] After determining the perception confidence, tracking confidence, and anomaly confidence respectively, the edge device performs a fusion calculation to obtain the comprehensive execution confidence (C). Specifically, the perception confidence, tracking confidence, and anomaly confidence are each multiplied by their respective preset weight coefficients and then summed. The calculation formula is as follows: C=α C_perc+β C_track+γ C_anom.
[0104] In this context, the preset weighting coefficients corresponding to each confidence level are all pre-defined constants, and the sum of the three is equal to one. The specific values of each weighting coefficient can be adjusted according to the actual application scenario and task type. For example, in perception-intensive tasks, the weighting coefficient of perception confidence can be increased, and in motion precision-intensive tasks, the weighting coefficient of tracking confidence can be increased. Through the above weighted summation, the sub-confidences of the three dimensions are merged into a comprehensive execution confidence, which fully reflects the overall reliability level of the edge device in executing the current sub-task.
[0105] After calculating the overall execution confidence level on the endpoint, the endpoint sends it as confidence information to the collaborative controller. The overall execution confidence level ranges from 0 to 1, meaning that a higher value indicates higher execution reliability on the endpoint, and a lower value indicates lower execution reliability. Upon receiving the overall execution confidence level, the collaborative controller can trigger the corresponding collaborative response mode based on the range of its value (see step S150 for details).
[0106] In conjunction with the first aspect, the collaborative controller determines the current execution state based on confidence information and triggers the execution of the corresponding mode on the endpoint, including the following steps: S151, when the overall execution confidence is greater than the first threshold, the edge-side autonomous execution mode is triggered, and the edge side only maintains monitoring.
[0107] S152, when the comprehensive execution confidence is greater than the second threshold and less than or equal to the first threshold, an end-edge collaborative execution mode is triggered, and the end device requests semantic assistance or path re-planning from the edge side.
[0108] S153, when the comprehensive execution confidence is less than or equal to the second threshold, an edge-side takeover execution mode is triggered, the end device pauses execution, and the edge side issues control commands.
[0109] Wherein, the first threshold is greater than the second threshold.
[0110] In this embodiment, a three-level gating mechanism is preset:[ the end device independently executes, and only reports the status periodically; Green gating (C>C1): the end device independently executes, and only reports status periodically; Yellow gating (C2<C≤C1): the end device actively requests assistance from the edge side; Red gating (C≤C2): the end device pauses execution and requests the edge side for emergency takeover.
[0111] Wherein, C1 is the first threshold, and C2 is the second threshold.
[0112] When C>C1, it indicates that the end device has reliable perception, accurate tracking, normal system operation and high overall execution reliability in the current execution process, and does not require intervention from the edge side. In this case, the cooperative controller triggers the end device independent execution mode (corresponding to the "green gating" interval). The end device continues to independently execute all perception and control operations of the current subtask in accordance with the task contract, and control instructions are independently generated by the end device and issued to the actuator. The edge side only keeps monitoring the state of the end device, records the confidence information and execution state information reported by the end device, but does not actively send auxiliary information or control instructions. In the independent execution mode, the end device still sends updated confidence information to the cooperative controller according to a preset period or when its own state changes significantly, so that the cooperative controller can continuously monitor the execution quality of the end device and trigger mode switching in time when the confidence changes across the threshold.
[0113] When C2<C≤C1, it indicates that the end device still has certain execution capability on the whole during execution, but there is local uncertainty in perception reliability, tracking accuracy or abnormal system status (corresponding to the "yellow gating" interval). In this case, if the end device continues to maintain completely independent execution, task failure or degradation of execution quality may be caused by the accumulation of local uncertainty. Therefore, the cooperative controller triggers the end-edge collaborative execution mode.
[0114] In the edge-edge collaborative execution mode, the edge continues to perform basic perception and control actions, responsible for high-frequency real-time response, while actively sending auxiliary requests to the edge. The edge responds to the edge's requests, providing auxiliary information. This auxiliary information includes, but is not limited to: semantic interpretation of the current scene (e.g., fine-grained identification and description of the category, position, and posture of objects in complex scenes to assist the edge's perception model in eliminating ambiguity), and replanning of the current local path (e.g., when a new obstacle appears on the currently planned path, the edge generates an alternative path based on global information and sends it to the edge). Through the collaborative cooperation of the above-mentioned real-time response from the edge and intelligent assistance from the edge, the real-time performance of the edge and the intelligence level of the edge are complemented, enabling the system to maintain reliable task execution even when the overall execution confidence is in the middle range.
[0115] When C≤C2, it indicates that a serious anomaly has occurred on the edge during execution, such as a failure of the perception model, excessive tracking error, or sensor malfunction (corresponding to...). Figure 2 (The "red-gated" zone in the diagram). In this situation, the edge device no longer has the conditions to continue executing tasks independently or collaboratively; continuing execution could lead to task failure or even a security incident. Therefore, the collaborative controller triggers the edge-side takeover execution mode.
[0116] In edge-side takeover execution mode, the edge device suspends the generation and issuance of autonomous control commands and sends a takeover request to the cooperative controller. The edge device directly generates control commands and issues them to the actuators on the edge device, taking over the robot's motion control. The actuators on the edge device execute actions according to the control commands issued by the edge device; the edge device itself only acts as an execution channel and no longer participates in perception and decision-making. Through this direct edge-side takeover, it is ensured that the edge device can maintain basic control capabilities even when the overall execution confidence is severely insufficient, preventing the system from completely failing due to edge device anomalies.
[0117] Through step S150 above, the collaborative controller, based on the comparison result between the comprehensive execution confidence and the preset threshold, triggers the corresponding mode among the three modes: autonomous execution on the end side, collaborative execution on the end side, or edge takeover execution. This achieves hierarchical response of the system under normal conditions and a smooth transition from autonomous execution to edge takeover. Through the above three-level gating mechanism, the system can dynamically adjust the collaborative mode between the end side and the edge side according to the real-time changes in the reliability of end side execution. Based on this, the system dynamically switches between the following five modes under two scenarios: normal network connection and network anomaly. The switching conditions and the respective behaviors of the end side and the edge side are shown in Table 1 below.
[0118] Table 1 shows a diagram illustrating the dynamic switching between the five modes.
[0119]
[0120] Among the above five modes, the first three (end-side autonomous mode, collaborative assistance mode, and edge takeover mode) correspond to the operating state when the network connection is normal. The system dynamically switches between these modes according to the comprehensive execution confidence C. When the comprehensive execution confidence changes across the threshold, the collaborative controller triggers the corresponding mode to achieve a smooth transition from independent execution on the end-side to complete takeover on the edge-side. The latter two modes (off-network autonomous mode and degraded safety mode) correspond to the degraded operation state after network interruption. The system decides whether to continue execution or shut down safely according to whether there are motion primitives locally on the end-side. For the complete relationship and triggering conditions of the above mode switching, please refer to Figure 3 , in this embodiment, C1=0.85; C2=0.5.
[0121] combined with Figure 3 shows the state machine switching relationship between the above five collaborative modes driven by the dual drive of network status and comprehensive execution confidence. The entire state machine is divided into two regions: Normal mode region ( Figure 3 upper half): When the network connection is normal, the system dynamically switches among the end-side autonomous mode, collaborative assistance mode and edge takeover mode according to the gating interval corresponding to the real-time calculated comprehensive execution confidence. When C>0.85, it is in the green gating interval, and the system is in the end-side autonomous mode; when 0.5<C≤0.85, it is in the yellow gating interval, and the system switches to the collaborative assistance mode; when C≤0.5, it is in the red gating interval, and the system switches to the edge takeover mode. The above switching is triggered by the confidence changing across the threshold, realizing a smooth transition from independent execution on the end-side to complete takeover on the edge-side.
[0122] Degraded mode region ( Figure 3 lower half): When the network is interrupted, the system exits the normal mode region and enters the degraded mode region. At this time, the system decides to switch to the off-network autonomous mode or the degraded safety mode according to whether motion primitives are stored locally on the end-side. If motion primitives are stored locally, the end-side switches to the off-network autonomous mode and calls the locally stored motion primitives to continue executing core tasks; if no motion primitive is stored locally, the end-side switches to the degraded safety mode, controls the robot to stop moving and wait for network recovery.
[0123] Switching between various modes is driven by two types of events: one is switching driven by the change of comprehensive execution confidence across thresholds, that is, mode switching triggered when confidence changes between different gating intervals (such as switching between autonomous mode and collaborative assistance mode, and between collaborative assistance mode and edge takeover mode); the other is switching driven by network status changes, that is, switching between the normal mode region and the degraded mode region triggered by network interruption or recovery events.
[0124] Through the state machine design described above, the system has a hierarchical response capability based on execution reliability when the network is normal, and achieves smooth degradation and autonomous recovery when the network is abnormal, ensuring that the system meets the stringent requirements of industrial manufacturing scenarios in terms of real-time performance, intelligence and robustness.
[0125] In conjunction with the first aspect, the method also includes: S160 monitors the quality of the communication link at the edge. When the predicted communication quality is lower than a preset threshold, it pushes the pre-configured motion primitives to the edge for local storage, so that the edge can call the motion primitives to perform tasks after the network is interrupted.
[0126] Step S160 runs parallel to the main execution flow consisting of steps S110 to S150, and can be continuously executed throughout the entire process of task execution on the edge. The edge monitors the quality of the communication link in real time and predicts its changing trend. Before the communication quality degrades to a preset threshold, motion primitives are pushed to the edge for local storage in advance. This allows the edge to still call the locally stored motion primitives to maintain the execution of core tasks after network interruption, thereby achieving predictive and smooth degradation of intelligent capabilities in network anomaly scenarios.
[0127] Specifically, during the execution of tasks on the end side, the edge continuously monitors the signal quality indicators of the current communication link. These indicators include, but are not limited to, parameters that reflect the quality of the communication link, such as signal strength, signal-to-noise ratio, transmission delay, and packet loss rate. By monitoring these indicators in real time, the edge acquires the current status data of the communication link, providing foundational data for predicting future trends in communication quality.
[0128] In conjunction with the first aspect, step S160 includes: S161, the edge side predicts the robot's pose in the next time window through extended Kalman filtering, and simultaneously monitors the signal quality indicators of the current communication link.
[0129] S162, when it is predicted that the signal quality index will be lower than the communication quality threshold after a preset time window, the edge side pushes the motion primitives corresponding to the current subtask and subsequent subtasks to the end side for local storage.
[0130] S163, after a network interruption, the terminal switches to the offline autonomous mode and continues execution by calling locally stored motion primitives.
[0131] S164 After the network is restored, the terminal side uploads the execution logs during the interruption to the edge side, and the edge side performs offline analysis and incremental model updates based on the execution logs.
[0132] After acquiring the current communication link status data, the edge device further predicts the signal quality indicators of the communication link using an Extended Kalman Filter (EKF). EKF, a recursive state estimation algorithm, can predict the signal quality status in the future based on the current signal quality observations and their trends. Specifically, the edge device uses the current communication link's signal quality indicators as state variables and the historical change sequences of each indicator as the basis for state transitions. It then uses EKF to predict the trend of signal quality indicators within the next time window, obtaining an estimated value of the communication quality at the predicted time point and its confidence interval. Through this prediction, the edge device can understand the trend of communication quality changes over a future period, providing a predictive basis for early triggering of motion primitive pushes.
[0133] When the edge side predicts that the communication quality will fall below a communication quality threshold after a preset time window, it triggers the pre-push of motion primitives. Specifically, the edge side packages and pushes the motion primitives (i.e., pre-trained motion primitives, such as pre-trained force-compliant primitives, obstacle avoidance motion primitives, grasping primitives, placement primitives, and other standard motion sequences) corresponding to the current subtask and subsequent possible subtasks to the edge side for local storage. These motion primitives are stored in the skill anchor management module of the edge side in the form of pre-trained motion parameters, which contain the control parameter sequence for the robot to perform specific actions. After obtaining the motion primitives, the edge side can autonomously generate control commands based on the motion primitives without real-time guidance from the edge side. The push is triggered when the estimated communication quality after the prediction time window is lower than a preset communication quality threshold. The length of the preset time window is preset by the system based on the fluctuation characteristics of the network environment, and the communication quality threshold is the minimum signal quality index that can maintain effective communication between the edge side and the edge side. After triggering the push, the edge side sends the motion primitives to the edge side through the currently available communication link.
[0134] After receiving and storing the motion primitives pushed by the edge side, the edge device saves them in its local storage. When the network connection is normal, the edge device still executes subtasks according to the execution mode specified in the task contract. The locally stored motion primitives serve as backup motion data caches and do not participate in the current execution process. When a network interruption occurs, the edge device detects that the communication link is unavailable and switches to the network-disconnected autonomous mode. In the network-disconnected autonomous mode, the edge device cannot report confidence information to the edge side, cannot request semantic assistance or path replanning from the edge side, and cannot be taken over by the edge side. Therefore, the edge device continues to execute the current subtask and subsequent subtasks based on the locally stored motion primitives. Specifically, the edge device reads the motion primitives corresponding to the current subtask from local storage, autonomously generates control commands based on the pre-stored control parameter sequences in the motion primitives, and sends them to the execution mechanism, thereby maintaining the execution of the core task and avoiding a precipitous drop in intelligence capabilities or production line stagnation due to network interruption.
[0135] After network recovery, the endpoint detects that the communication link is available again and uploads the execution logs generated during the network outage to the edge side. These execution logs include at least the sub-task identifiers executed by the endpoint during the outage, actual motion trajectories, sensor feedback data, and records of abnormal events. Upon receiving the execution logs from the endpoint, the edge side compares and analyzes the actual execution data with the expected execution results to identify deviations and anomalies in the endpoint's execution under the autonomous mode without network access. Based on this, the parameters of the vertical model are incrementally updated to improve the semantic understanding and task decomposition accuracy of the vertical model in future tasks.
[0136] Through step S160, the edge device pushes motion primitives to local storage before the communication quality degrades to a threshold. This allows the edge device to maintain the execution of core tasks based on local motion primitives after a network outage. Once the network is restored, the execution logs from the outage are uploaded to the edge device for offline analysis and incremental model updates. This mechanism enables the system to smoothly degrade its intelligent capabilities during network anomalies and to autonomously recover and optimize the model after network restoration. This effectively avoids the technical defects of existing solutions where system stagnation or degradation into blind execution due to network outages.
[0137] Combination Figure 4 As shown, Figure 4 It fully demonstrates the interaction process between the edge side (the robot's computational brain), the collaborative controller / external support system, and the physical environment (the robot's actuators and sensors) from a time perspective. Figure 4 Using the horizontal axis as the time axis, the complete interaction process is divided into three consecutive time stages.
[0138] Phase 1 corresponds to the initialization and contract establishment phase (steps S110 to S130 and S131 to S134). The edge device receives job instructions from the scheduling system and sends its capability information to the collaborative controller. The edge device receives job instructions from the scheduling system, decomposes the job instructions into a sequence of sub-tasks using a vertical model, and generates capability requirement information for each sub-task. For each sub-task, the collaborative controller matches the capability requirement information with the edge device's capability registry to generate a task contract and sends the task contract to the edge device. The above interaction corresponds to... Figure 4 The process involves a two-way interaction between the mid-level (brain) and the peripheral (brain) sides, where the brain reports capability information to the peripheral side (brain trust) and receives information from the contract.
[0139] Phase Two corresponds to the execution and confidence feedback loop phase (steps S140 to S150 and five mode switching). The edge device sends control commands to the physical environment according to the task contract, and the physical environment feeds back sensor data to the edge device. During execution, the edge device calculates the overall execution confidence in real time and reports it to the collaborative controller. The collaborative controller triggers the corresponding collaborative mode based on the confidence level, dynamically switching between three modes: edge-side autonomous execution, edge-side collaborative execution, and edge-side takeover execution. The above interaction corresponds to... Figure 4 The process of two-way interaction between the control command issuance and sensor feedback closed loop between the mid-end and the physical environment, and between the confidence feedback and auxiliary requests between the end and edge sides.
[0140] Phase 3 corresponds to the network outage degradation and recovery phase (step S160). The edge continuously monitors the communication link quality and pushes motion primitives to the terminal when it predicts the communication quality will fall below a threshold. After a network outage, the terminal switches to an autonomous mode and continues to send control commands to the physical environment based on locally stored motion primitives. After the network recovers, the terminal uploads the execution logs from the outage period to the edge. The above interaction corresponds to... Figure 4 The process includes: pushing anchor points from the edge to the end, the autonomous execution loop between the end and the physical environment during network outages, and the log upload interaction process after network recovery.
[0141] The interaction process of the above three stages together constitutes a complete end-edge collaborative closed loop from contract establishment, execution feedback, gating arbitration to predictive degradation.
[0142] Secondly, embodiments of this application provide an end-to-end collaborative task execution device for an embodied intelligent robot. This device is applied to a collaborative control system including an end-side component, an edge-side component, and a collaborative controller, and is combined with... Figure 6 As shown, the device includes: a capability negotiation module 10, an execution and feedback module 20, a gating arbitration module 30, and a prediction degradation module 40.
[0143] The capability negotiation module 10 is used to enable the terminal side to receive the work instruction, enable the edge side to receive the work instruction, enable the terminal side to send the terminal side's capability information to the collaborative controller, enable the edge side to decompose the work instruction into a sequence of sub-tasks through the edge side's vertical model and generate capability requirement information for each sub-task, and enable the collaborative controller to obtain the task decomposition results from the edge side. For each sub-task, the capability requirement information corresponding to the sub-task is matched with the terminal side's capability registry to generate a task contract and send it to the terminal side; wherein, the terminal side's capability registry is established based on the terminal side's capability information.
[0144] The execution and feedback module 20 is used to enable the end side to execute each sub-task according to the task contract, and to calculate the confidence information to characterize the current execution reliability in real time during the execution process, and send the confidence information to the collaborative controller; The gated arbitration module 30 is used to enable the collaborative controller to determine the current execution status based on the confidence information and trigger the corresponding mode among the end-side autonomous execution mode, end-edge collaborative execution mode, or edge-side takeover execution mode.
[0145] The prediction degradation module 40 is used to enable the edge side to monitor the quality of the communication link. When the predicted communication quality is lower than a preset threshold, the pre-configured motion primitives are pushed to the edge side for local storage so that the edge side can call the motion primitives to perform tasks after the network is interrupted.
[0146] Thirdly, embodiments of this application provide an electronic device, combined with Figure 7 As shown, the electronic device includes a memory 131 and a processor 130. The memory 131 stores a computer program, and the processor 130 runs the computer program to make the electronic device perform the above-described method.
[0147] Furthermore, combined Figure 7 The electronic device shown also includes a bus 132 and a communication interface 133, with the processor 130, the communication interface 133 and the memory 131 connected via the bus 132.
[0148] The memory 131 may include high-speed random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage device. Communication between this system network element and at least one other network element is achieved through at least one communication interface 133 (which can be wired or wireless), such as the Internet, wide area network, local area network, metropolitan area network, etc. The bus 132 may be an ISA bus, PCI bus, or EISA bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 7 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.
[0149] Processor 130 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of processor 130 or by instructions in software form. Processor 130 may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this invention can be directly manifested as execution by a hardware decoding processor, or execution by a combination of hardware and software modules in the decoding processor. The software module can reside in a mature storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory 131, and processor 130 reads the information in memory 131 and, in conjunction with its hardware, completes the steps of the method described in the foregoing embodiments.
[0150] Fourthly, embodiments of this application provide a readable storage medium storing computer program instructions, which are read and executed by a processor to perform the above-described method.
[0151] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the system and apparatus described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0152] Furthermore, in the description of the embodiments of the present invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in the present invention based on the specific circumstances.
[0153] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0154] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0155] Finally, it should be noted that the above embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for edge-to-edge collaborative task execution using an embodied intelligent robot, characterized in that, The method is applied to a collaborative control system including an end-side, an edge-side, and a collaborative controller, and the method includes: The terminal receives the operation command and sends the current capability information to the cooperative controller; The edge side receives the operation instruction, decomposes the operation instruction into a sequence of sub-tasks through the vertical model of the edge side, and generates capability requirement information for each sub-task to obtain the task decomposition result of the operation instruction. The task decomposition result includes the sequence of sub-tasks and the capability requirement information corresponding to each sub-task. The collaborative controller obtains the task decomposition results of the operation instructions from the edge side, and for each sub-task, matches the capability requirement information corresponding to the sub-task with the edge-side capability registry to generate a task contract, and sends the task contract to the edge side; the edge-side capability registry is established based on the capability information of the edge side. The endpoint executes each subtask according to the task contract, and calculates confidence information in real time to characterize the current execution reliability during the execution process, and sends the confidence information to the collaborative controller; The collaborative controller determines the current execution state based on the confidence information and triggers the corresponding execution mode on the edge. The mode includes at least one of autonomous execution mode, edge-edge collaborative execution mode, or edge-side takeover execution mode.
2. The method according to claim 1, characterized in that, The step of matching the capability requirement information corresponding to the subtask with the end-side capability registry to generate a task contract includes: Get the matching results; If the matching result is a complete match, the task contract is configured to a mode in which the subtask is executed autonomously by the terminal side; If the matching result is a partial match, the task contract is configured as a collaborative execution mode in which the subtask performs basic operations on the terminal side and provides auxiliary information on the edge side; If the matching result is a mismatch, the task contract is configured to a takeover execution mode in which the subtask is directly executed by the edge side.
3. The method according to claim 1, characterized in that, The capability information includes at least the currently available model types, model inference confidence, computing load status, and sensor health status on the edge; the capability requirement information includes at least the sensing type, accuracy requirements, response latency limit, and computing power requirements required to execute the corresponding sub-task.
4. The method according to claim 1, characterized in that, The confidence information is a comprehensive execution confidence score, which integrates three dimensions: perception confidence score, tracking confidence score, and anomaly confidence score. The steps involved in real-time calculation of confidence information characterizing the current execution reliability during execution include: The perception confidence is determined based on the information entropy value of the category probability distribution output by the perception model and the spatial overlap of the detection boxes in the detection results of consecutive multiple frames; wherein, the information entropy value is used to characterize the degree of uncertainty of the perception model in its output results, and the spatial overlap is used to characterize the degree of stability of the detection boxes between adjacent frames. The tracking confidence level is determined based on the positional offset between the robot's actual trajectory and the planned trajectory. The anomaly confidence level is determined based on the degree of deviation between the real-time feedback value of the sensor and the expected reference value under normal operating conditions. The execution confidence is obtained by weighted summation of the perception confidence, the tracking confidence, and the anomaly confidence; wherein the sum of the preset weight coefficients corresponding to each confidence is equal to one.
5. The method according to claim 4, characterized in that, The steps of the collaborative controller determining the current execution state based on the confidence information and triggering the execution of the corresponding mode on the endpoint include: When the overall execution confidence level is greater than the first threshold, the edge-side autonomous execution mode is triggered, and the edge side only maintains monitoring. When the overall execution confidence is greater than the second threshold and less than or equal to the first threshold, the edge-end collaborative execution mode is triggered, and the edge requests semantic assistance or path replanning from the edge side. When the overall execution confidence is less than or equal to the second threshold, the edge-side takeover execution mode is triggered, the end-side execution is suspended, and the edge side issues control commands. Wherein, the first threshold is greater than the second threshold.
6. The method according to claim 1, characterized in that, The method further includes: the edge side monitors the communication link quality, and when the predicted communication quality is lower than a preset threshold, it pushes the pre-configured motion primitives to the end side for local storage, so that the end side can call the motion primitives to perform tasks after the network is interrupted.
7. The method according to claim 5, characterized in that, The step of monitoring the communication link quality at the edge and pre-pushing pre-configured motion primitives to the end side includes: The edge side uses extended Kalman filtering to predict the robot's pose in the next time window and simultaneously monitors the signal quality indicators of the current communication link. When it is predicted that the signal quality index will be lower than the communication quality threshold after a preset time window, the edge side will push the motion primitives corresponding to the current subtask and subsequent subtasks to the end side for local storage. After a network interruption, the terminal switches to a network-disconnected autonomous mode and continues execution by calling the locally stored motion primitives; After the network is restored, the terminal side uploads the execution logs during the interruption to the edge side, and the edge side performs offline analysis and incremental model updates based on the execution logs.
8. A device for edge-to-edge collaborative task execution of an embodied intelligent robot, characterized in that, The device is applied to a collaborative control system including an end-side, an edge-side, and a collaborative controller, and the device includes: The capability negotiation module is used to enable the endpoint to receive the job instruction, enable the edge side to receive the job instruction, enable the endpoint to send its capability information to the collaborative controller, enable the edge side to decompose the job instruction into a sequence of sub-tasks using its vertical model and generate capability requirement information for each sub-task, and enable the collaborative controller to obtain the task decomposition result from the edge side, and for each sub-task, match the capability requirement information corresponding to the sub-task with the endpoint capability registry to generate a task contract and send it to the endpoint; wherein, the endpoint capability registry is established based on the endpoint capability information; The execution and feedback module is used to enable the terminal to execute each sub-task according to the task contract, and to calculate confidence information in real time during the execution process to characterize the current execution reliability, and send the confidence information to the collaborative controller; The gated arbitration module is used to enable the collaborative controller to determine the current execution state based on the confidence information and trigger the corresponding mode among the end-side autonomous execution mode, end-edge collaborative execution mode, or edge-side takeover execution mode. The prediction degradation module is used to enable the edge side to monitor the communication link quality. When the predicted communication quality is lower than a preset threshold, the pre-configured motion primitives are pushed to the edge side for local storage, so that the edge side can call the motion primitives to perform tasks after the network is interrupted.
9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program and the processor running the computer program to cause the electronic device to perform the method of any one of claims 1 to 7.
10. A storage medium, characterized in that, The storage medium stores computer program instructions, which, when read and executed by a processor, perform the method described in any one of claims 1 to 7.