Data processing method and system of intelligent agent, medium, equipment and product
By integrating configuration and declarative templates, a data processing pipeline for embodied intelligent agents is dynamically constructed, solving the problems of multimodal data compatibility and low resource utilization in embodied intelligent agent systems, and achieving rapid adaptation and efficient processing.
Patent Information
- Application Number
- CN202510860150.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-10-17
AI Technical Summary
Existing embodied agent data processing systems struggle to achieve unified compatibility with multimodal data, resulting in high development and maintenance costs, low resource utilization, and a lack of robust state tracking and backtracking mechanisms, making it difficult to quickly adapt to combinations of different embodied agents and end effectors.
By processing the ontological capability information, end effector description information, and verification rule information of the embodied intelligent agent into integrated configuration information, a target data processing platform is constructed, configuration resources are dynamically loaded, a data processing pipeline is built, multimodal data processing is supported, and type recognition and declarative templates are introduced to achieve automated configuration and resource scheduling.
It achieves cross-device compatibility and scalability of embodied intelligent agent systems, reduces development and maintenance costs, improves resource utilization and response speed, and supports plug-and-play and flexible processing of multimodal data.
Smart Images

Figure CN120803700A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to the technical field of embodied intelligence, and in particular to a data processing method and system of embodied intelligent agent, medium, equipment and product. BACKGROUND
[0002] Under the background of continuous development of embodied intelligent agent technology, according to different combinations of embodied intelligent agent and end effector, various data including visual point cloud, speech stream, and force signal will be generated in the running process. These data have significant differences in format, structure and processing method, which makes it difficult for related data processing systems to achieve unified compatibility, and requires customized processing for each type of data, significantly increasing the development and maintenance cost of the system.
[0003] Based on this, embodiments of the present application provide a data processing method and system of embodied intelligent agent, medium, equipment and product to improve related technologies. SUMMARY
[0004] The purpose of the embodiments of the present application is to provide a data processing method and system of embodied intelligent agent, medium, equipment and product, which is suitable for processing multi-modal data generated by combinations of various embodied intelligent agents and various end effectors, and has strong universality.
[0005] The purpose of the embodiments of the present application is achieved by using the following technical solutions:
[0006] In a first aspect, the embodiments of the present application provide a data processing method of embodied intelligent agent, which comprises: processing ontology capability information, end effector description information and verification rule information of a target embodied intelligent agent into integrated configuration information; based on the integrated configuration information, constructing integrated configuration resources recognizable by a target data processing platform; in the target data processing platform, generating a target data processing job mounted with the integrated configuration resources; in the process of executing the target data processing job, dynamically loading the integrated configuration resources to generate internal representations corresponding to the ontology capability information, the end effector description information and the verification rule information; constructing a data processing pipeline according to the internal representations, the data processing pipeline being used to process multi-modal data collected by the target embodied intelligent agent; the multi-modal data comprising at least two of visual data, radar data, speech data and force data.
[0007] In some embodiments, before processing the ontology capability information, the end effector description information and the verification rule information of the target embodied intelligent agent into integrated configuration information, the method further comprises: in the case of detecting type information of the target embodied intelligent agent, acquiring the ontology capability information, the end effector description information and the verification rule information corresponding to the type information.
[0008] In some embodiments, the verification rule information is determined based on at least one of the ontology capability information and the end effector description information.
[0009] In some embodiments, the constructing the data processing pipeline according to the internal representation comprises: constructing the data processing pipeline according to the internal representation and a declarative template; the declarative template is used to define a data processing procedure of the multi-modal data in a plurality of processing modules, the plurality of processing modules comprising at least two of a scene understanding module, a speech processing module and a force perception processing module.
[0010] In some embodiments, the determining process of the declarative template comprises: receiving a template configuration operation through a visual interface, and configuring at least two of the scene understanding module, the speech processing module and the force perception processing module to define the declarative template.
[0011] In some embodiments, the determining process of the declarative template comprises: defining the declarative template based on a preset configuration file.
[0012] In some embodiments, the method further comprises: scheduling a computing resource of the target data processing job based on a feature extraction result of the multi-modal data and / or a task priority type corresponding to the target data processing job; the feature extraction result comprises a computing power type and / or delay requirement information, and the feature extraction result is obtained by analyzing metadata of the multi-modal data.
[0013] In some embodiments, the process of scheduling the computing resource of the target data processing job comprises: in a case where a number of tasks with a first priority type included in the target data processing job is greater than a specified number, increasing a number of GPU nodes available to the target data processing job, and / or suspending access of the target data processing job to at least part of CPU nodes.
[0014] In some embodiments, the method further comprises: in a case where a CPU utilization of a target container group included in the target data processing job exceeds a first utilization and / or a memory utilization exceeds a second utilization, increasing an upper limit of a CPU quota of the target container group and / or an upper limit of a memory quota.
[0015] In some embodiments, the method further comprises: during execution of the target data processing job, creating a tracking identifier for tracking an intermediate state of multi-modal data processing, the tracking identifier being used to indicate an agent identifier, an environment context identifier and a current timestamp; and saving the tracking identifier, a current processing state and a current multi-modal fusion indicator in association.
[0016] In a second aspect, the embodiments of the present application provide a data processing system of embodied agent, the system comprising a plurality of embodied agents, a plurality of end effectors, and a data processing engine module, the data processing engine module being configured to perform any of the above methods.
[0017] In a third aspect, the embodiments of the present application provide a computer-readable storage medium storing a computer program, the computer program being configured to implement any of the above methods when executed by a processor.
[0018] In a fourth aspect, the embodiments of the present application provide a computer device comprising a memory and a processor, the memory storing a computer program, the processor being configured to implement any of the above methods when executing the computer program.
[0019] In a fifth aspect, the embodiments of the present application provide a computer program product comprising a computer program, the computer program being configured to implement any of the above methods when executed by a processor.
[0020] The embodiments of the present application provide a data processing method and system of embodied agent, a medium, a device, and a product, the method comprising: processing ontology capability information, end effector description information, and verification rule information of a target embodied agent into integrated configuration information; based on the integrated configuration information, constructing integrated configuration resources recognizable by a target data processing platform; in the target data processing platform, generating a target data processing job mounted with the integrated configuration resources; in the process of executing the target data processing job, dynamically loading the integrated configuration resources to generate internal representations corresponding to the ontology capability information, the end effector description information, and the verification rule information; and constructing a data processing pipeline according to the internal representations, the data processing pipeline being configured to process multi-modal data collected by the target embodied agent. The multi-modal data comprises at least two of visual data, radar data, voice data, and force sensation data. The embodiments of the present application can decouple the ontology capability information, the end effector description information, and the verification rule information of the embodied agent from specific data processing logic, so that the system can quickly adapt based on the integrated configuration information when facing multiple embodied agents of different brands, models, or functional combinations, multiple end effectors, and has good cross-device compatibility and scalability, and is highly versatile. Through the way of dynamic loading configuration and on-demand pipeline construction, the running efficiency and resource utilization of the system are improved, the multi-modal data generated by the combination of multiple embodied agents and multiple end effectors can be processed, and the system has wide application value. BRIEF DESCRIPTION OF DRAWINGS
[0021] The embodiments of the present application will be further described below in conjunction with the accompanying drawings and specific embodiments.
[0022] Figure 1is a flow diagram of a data processing method of a body-equipped intelligent agent provided by an embodiment of the present application.
[0023] Figure 2 is a unified modeling framework diagram of a multi-body multi-end provided by an embodiment of the present application.
[0024] Figure 3 is a configuration flow diagram of a declarative template provided by an embodiment of the present application.
[0025] Figure 4 is a horizontal expansion and contraction diagram provided by an embodiment of the present application.
[0026] Figure 5 is a zero-downtime hot update flow diagram provided by an embodiment of the present application.
[0027] Figure 6 is a stream log analysis diagram provided by an embodiment of the present application.
[0028] Figure 7 is a structural block diagram of a data processing system of a body-equipped intelligent agent provided by an embodiment of the present application.
[0029] Figure 8 is a pluggable data processing architecture diagram provided by an embodiment of the present application.
[0030] Figure 9 is a structural block diagram of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0031] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.
[0032] In the description of the embodiments of the present application, it should be understood that the terms "first", "second" are used only for the purpose of description, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features limited by "first", "second" can explicitly or implicitly include one or more of the features. In the description of the embodiments of the present application, the meaning of "multiple" is two or more, unless otherwise specifically limited.
[0033] Under the background of the continuous development of embodied agent technology, the system will generate multi-modal heterogeneous data including visual point cloud, speech stream, and force signal during operation. These data have significant differences in format, structure, and processing methods, making it difficult for related cloud processing systems to achieve uniform compatibility. Customized adaptation is needed for each type of data, significantly increasing the development and maintenance cost of the system, and the adaptation cost can increase by 60%.
[0034] At the same time, embodied agents often face the problem of severe data flow fluctuations in practical applications, with peak traffic exceeding 10 times the valley. If a static resource allocation mode is used, it is difficult to balance the computing power during peak processing and resource conservation during the trough, resulting in low overall resource utilization, with an average GPU utilization rate of less than 35%. In addition, related systems generally lack a perfect state tracking and backtracking mechanism. During multi-modal data processing, abnormal data is difficult to be accurately located to specific links, causing about 75% of abnormal problems to be unable to be effectively diagnosed and solved.
[0035] Related embodied intelligent data processing technology involves three types of solutions. The first is a static cloud resource allocation mode, which uses a pre-configured virtual machine cluster (such as a fixed number of CPU / GPU instances) to allocate computing resources according to pre-set rules. Representative platforms such as AWS Batch and Azure Batch are suitable for offline data analysis tasks, but have limitations in handling dynamic loads and multi-modal requirements. The second is an offline multi-modal processing pipeline built based on tools such as Apache Airflow, which processes each type of data independently according to the stage (such as completing visual processing before performing speech analysis). This is commonly used in robot simulation data analysis, and the coordination between modules is low. The third is a centralized data processing center, where all data collected by embodied agents is uploaded to a centralized cloud server for unified processing by human-configured rules. This approach is typically used in early remote monitoring systems for industrial robots.
[0036] However, there are some problems with the above technologies that need to be addressed. First, the module coupling degree is high, and different data processing flows and business logic are deeply bound. If a new algorithm module is introduced, the service often needs to be restructured, with a long iteration cycle and heavy development burden. Second, the data island problem is serious, and the data collected by different embodied agents lacks unified knowledge sedimentation, leading to repeated learning and processing in the same scenario, resulting in a waste of computing power of more than 45%. Third, related solutions are mostly black-box operation modes, lacking visualization management of intermediate states in data processing. Once the system configuration is changed, it is difficult to accurately assess the scope of impact, and the overall controllability is poor.
[0037] Reference Figure 1 , Figure 1FIG. 1 is a flowchart of a data processing method of a body-equipped intelligent agent according to an embodiment of the present application.
[0038] With the wide application of body-equipped intelligent agents in manufacturing, service, medical treatment and other scenarios, the multi-modal data (such as visual, voice, force sense, etc.) generated by the body-equipped intelligent agents is of complex types and has obvious differences in processing requirements. In actual deployment, the capabilities of different manufacturers and different models of body-equipped intelligent agents differ, and it is difficult to quickly adapt to a unified data processing framework, and the plug-and-play capability is poor. In order to improve the related technology, an embodiment of the present application provides a data processing method of a body-equipped intelligent agent, which comprises steps S101-S105.
[0039] Step S101: processing ontology capability information, end effector description information and verification rule information of a target body-equipped intelligent agent into integrated configuration information.
[0040] Step S102: based on the integrated configuration information, constructing integrated configuration resources recognizable by a target data processing platform.
[0041] Step S103: in the target data processing platform, generating a target data processing job mounted with the integrated configuration resources.
[0042] Step S104: in the process of executing the target data processing job, dynamically loading the integrated configuration resources to generate internal representations corresponding to the ontology capability information, the end effector description information and the verification rule information.
[0043] Step S105: constructing a data processing pipeline according to the internal representations, the data processing pipeline being used for processing multi-modal data collected by the target body-equipped intelligent agent; the multi-modal data comprising at least two of visual data, radar data, voice data and force sense data.
[0044] In some embodiments, the above method can be applied to a data processing system of a body-equipped intelligent agent. As an example, the above method can run on a data processing engine module in the system.
[0045] An embodied agent refers to an intelligent system that can physically interact with the environment through its own sensors and end effectors, such as a robotic system with components such as visual sensors (e.g., depth cameras), voice microphones, force sensors, grippers, robotic arms, and mobile chassis. As an example, an embodied agent can collect multimodal data and perform real-time perception and action decision-making tasks, and has a certain degree of autonomous behavior capabilities. The "target" in the target embodied agent serves as a prefix to distinguish, and the target embodied agent can be one of multiple embodied agents. In some embodiments, the embodied agent includes one or more of an industrial robot, a collaborative robot, a humanoid robot (also known as a humanoid robot or a bipedal robot), a quadruped robot, and a wheeled robot.
[0046] Integrated configuration information refers to the integrated capability description data associated with the target embodied intelligent agent. It is obtained by processing the target embodied intelligent agent's ontological capability information, end-effector description information, and verification rule information. This integrated configuration information can be encapsulated in a unified configuration resource structure to form an integrated configuration resource for unified management and dynamic injection into the target data processing platform. Ontological capability information refers to the description of capability parameters associated with the target embodied intelligent agent itself, including, for example, degrees of freedom (DoF), maximum motion speed, maximum payload, and connected end-effectors. End-effector description information refers to the capability parameters and constraints associated with the end-effector connected to the target embodied intelligent agent, including, for example, the gripper's opening and closing range, the visual sensor's sensing range, and the force sensor's maximum force limit. Verification rule information refers to the various judgment logic and rule sets used to constrain the data processing process, including, for example, collision force detection, safety threshold determination, and accuracy error tolerance during the task.
[0047] An integrated configuration resource refers to a structure that constructs the above-mentioned integrated configuration information in the form of native resources of the target data processing platform (such as Kubernetes' ConfigMap), which can be recognized by the target data processing platform (such as Kubernetes).
[0048] The internal representation corresponding to the ontology capability information, end-effector description information, and verification rule information, for example, refers to a structured representation generated based on the ontology capability information, end-effector description information, and verification rule information, which can be recognized by the target data processing platform and used to drive data processing logic. This internal representation may include configuration objects (e.g., Python objects or other data structures) generated at runtime to express the semantic content corresponding to the ontology capability information, end-effector description information, and verification rule information.
[0049] The target data processing job refers to a data processing task unit created and run for the target embodied agent in the target data processing platform, for example, deployed and executed in a container group manner. The target data processing job mounts the integrated configuration resource, and dynamically loads the integrated configuration resource in the running process, so as to generate an internal representation, and then constructs a data processing pipeline based on the internal representation.
[0050] The data processing pipeline can be used to analyze, fuse, judge and output decision results of the multi-modal data collected by the target embodied agent.
[0051] The embodied agent data processing method provided by the embodiment of the present application first integrates the ontology capability information, end effector description information and verification rule information of the embodied agent, to form integrated configuration information. The integrated configuration information is converted into integrated configuration resources recognizable by the target data processing platform, and is mounted in the generated target data processing job. In the job execution phase, the integrated configuration resources can be dynamically loaded, and parsed into internal representations for program execution. Then, based on these internal representations, a data processing pipeline is dynamically constructed. The data processing pipeline supports processing of multi-modal data such as visual data, radar data, voice data and force sensation data. In the above manner, the embodiment of the present application can decouple the ontology capability information, end effector description information, verification rule information of the embodied agent and the specific data processing logic, so that the system can quickly adapt based on the integrated configuration information when facing different brands, models or functional combinations of multiple embodied agents and multiple end effectors, and has good cross-device compatibility and scalability. Through the dynamic loading configuration and on-demand pipeline construction, the system's running efficiency and resource utilization are improved, supporting processing of high-frequency, real-time or asynchronous collected multi-modal data, and having wide application value.
[0052] As the types of embodied agents become more diversified and the scene tasks change constantly, the system needs to manually configure a large amount of ontology capabilities, end effector characteristics and task rules when processing different embodied agents, which is tedious, error-prone, and difficult to quickly adapt when deploying new devices (referring to embodied agents) or changing tasks (at this time, the end effector used to execute the task often needs to be replaced), which seriously restricts the ability of multi-device concurrent access and large-scale deployment. In addition, since the processing flow is often tightly coupled with the type of embodied agent, it requires the involvement of developers every time a new type of embodied agent is added, resulting in large configuration and debugging workload and long update cycle, and the system's generality and intelligence level need to be improved.
[0053] In some embodiments, before processing the ontology capability information, the end effector description information and the verification rule information of the target embodied agent into integrated configuration information, the method can further include: in the case that the type information of the target embodied agent is detected, obtaining the ontology capability information, the end effector description information and the verification rule information corresponding to the type information.
[0054] The type information of the target embodied agent, for example, is a specific model or other category of agent identifier. As an example, after detecting the type information, the ontology capability information, the end effector description information and the verification rule information corresponding thereto can be automatically retrieved from a preset embodied agent capability database to form integrated configuration information, which can be subsequently processed into integrated configuration resources and injected into the target data processing platform for mounting and calling by the target data processing job. The above embodiments can realize rapid identification and configuration synchronization of the embodied agent, improve the cumbersome process of manually searching and inputting various configuration information, and improve the automation level and response speed of the system. At the same time, the mechanism supports the diversity and dynamic expansion of the type of embodied agent, which is conducive to the system to maintain good adaptability and maintainability in the face of heterogeneous devices or multi-task concurrent environment. The above embodiments introduce a type identification and configuration mapping mechanism, when the type information of the target embodied agent is detected, the ontology capability information, the end effector description information and the verification rule information corresponding to the type information are automatically obtained, so as to process them into integrated configuration information for subsequent automated data processing process calling. The technical solution can significantly reduce the configuration threshold, realize plug and play of new devices (i.e. new embodied agents), and improve the rapid response capability and maintainability of the cloud processing system in multi-modal and multi-scene environments.
[0055] The above embodiments do not limit the way of obtaining the verification rule information, which can be directly obtained according to the type information, or can be determined based on the ontology capability information and the end effector description information after obtaining the ontology capability information and the end effector description information according to the type information. In some embodiments, the verification rule information can be determined based on at least one of the ontology capability information and the end effector description information.
[0056] For example, upon detecting the type information (e.g., robot_type) of the target embodied agent, a set of deployment procedures can be automatically executed to complete the data processing configuration matching the type of the target embodied agent. The deployment procedures can be implemented by invoking the following function: deploy_robot_system(robot_type, end_effectors), where robot_type represents the detected type information of the target embodied agent, and end_effectors represents, for example, a list of end effectors (including one or more end effectors) connected to the target embodied agent.
[0057] In the deploy_robot_system(robot_type, end_effectors) function, the agent specification corresponding to robot_type can be obtained through RobotDB.get_spec(robot_type), which describes the ontology capability information of the target embodied agent. All end effector description information can be obtained through [EndEffectorDB.get(e) for e in end_effectors]. Then, the matching verification rule information can be dynamically compiled by invoking RuleEngine.compile_rules(robot_type, end_effectors) based on the type information of the embodied agent and the end effector description information. The above three types of information can be packaged into an integrated configuration information config in the form of a dictionary, which can include key-value pairs 'robot', 'end_effectors', and 'rules', corresponding to ontology capability information, end effector description information, and verification rule information, respectively.
[0058] Subsequently, the system can invoke the API (Application Programming Interface) of the Kubernetes platform to create a configuration resource object (ConfigMap) named "robot-cluster-{uuid}" through the k8s.create_configmap function. The name field uses a unique cluster identifier, the data field carries the integrated configuration information described above, and the labels field sets the resource label "robot-cluster":"active" to mark the current resource as active.
[0059] Finally, the system can call create_processing_job(config) to automatically generate a corresponding target data processing job (Job) based on the integrated configuration information (i.e., config) and deploy it to a target data processing platform (e.g., a cloud platform). When generating the Job, the corresponding ConfigMap is automatically bound, and the agent configuration, end effector configuration, and verification rule configuration in the integrated configuration resource are loaded at runtime to drive subsequent data processing pipeline construction and execution.
[0060] Through the above processing mode, automatic recognition of the target embodied agent type, automatic generation of integrated configuration resources, and automatic deployment of the target data processing job are achieved, with high universality and automation capability.
[0061] The multi-modal data collected by the embodied agent can include visual data, radar data, voice data, and force sensation data, and other multi-source heterogeneous information. Related data processing procedures lack universality and flexibility, and most systems rely on statically configured pipelines, which greatly increases development and maintenance costs as each new embodied agent or end effector is added.
[0062] To achieve automation of data processing pipeline construction and reduce development pressure, in some embodiments, constructing a data processing pipeline according to the internal representation can include constructing a data processing pipeline according to the internal representation and a declarative template. The declarative template is used to define the data processing procedure of the multi-modal data in multiple processing modules, including at least two of a scene understanding module, a voice processing module, and a force sensation processing module.
[0063] In some embodiments, the multiple processing modules can include a decision generation module in addition to at least two of a scene understanding module, a voice processing module, and a force sensation processing module.
[0064] The declarative template refers to a data processing task orchestration template defined in a structured manner such as YAML (a human-readable data serialization language) or JSON (JavaScript Object Notation), which is used to describe the corresponding processing module call order, resource usage strategy, inter-module dependency relationship, etc. of the multi-modal data (such as visual data, radar data, voice data, force sensation data, etc.) collected by the target embodied agent. Decoupling the declarative template from the integrated configuration information and data processing logic enables flexible combination of data processing procedures and module reuse.
[0065] The above embodiment first converts the ontology capability information of the embodied agent, the end device description information and the verification rule information into programmable internal representation, such as Python object, and then automatically constructs a data processing pipeline containing multiple processing modules (such as scene understanding module, speech processing module, force perception processing module and decision generation module) in combination with the declarative template (such as module sequence and resource policy defined in YAML format). The declarative template allows developers to describe the data flow path, module dependency relationship and resource scheduling strategy in a modular manner without deep programming, thereby significantly improving the generality, visibility and maintainability of the data processing flow. By introducing the technical solution, the automation and templating of the data processing pipeline construction are realized, so that the system can quickly adapt to different types of embodied agents and end effectors. Compared with the static configuration mode, the method greatly reduces the manual cost required for system reconstruction and debugging, improves the process deployment efficiency, and enhances the ability of the system in process extension, module reuse and cross-embodied agent and cross-end effector adaptation. The integrated configuration resource and the declarative template are used in combination to realize the plug-and-play data processing pipeline deployment, without the need for manual scripting or service logic reconstruction, thereby greatly reducing the system development and maintenance cost.
[0066] The integrated configuration (corresponding to integrated configuration information and integrated configuration resource) mode contains static capability description (such as ontology capability information, end effector description information and verification rule information) related to the target embodied agent, and is independent of the specific data processing logic (such as visual recognition, speech analysis, force feedback processing and decision generation). This integrated configuration mode improves the cumbersome configuration of each type of algorithm module that needs to reconstruct the service, and realizes the decoupling of the capability embodied in the configuration and the data processing logic defined in the template. In actual application, according to the actual needs of the target embodied agent, the appropriate combination and execution order of processing modules can be selected to realize the processing of multi-modal data. For example, different combinations of embodied agents and end effectors of different brands only need to provide corresponding ontology capability information, end effector description information and verification rule information, and can run on the same data processing platform according to the declarative template, without the need for additional customized code to reconstruct the service, thereby realizing universal adaptation across brands and heterogeneous terminals. In this way, when a new embodied agent or end effector is added, the service does not need to be restarted, and the code does not need to be recompiled. Instead, the integrated configuration information is used to automatically mount resources, construct pipelines and start running, thereby realizing the plug-and-play function.
[0067] In the actual deployment of embodied agent systems, the construction of data processing flow usually involves multiple heterogeneous modules such as scene understanding, speech processing, and force perception processing. The related configuration process highly depends on professional engineers to manually write underlying configuration code, and non-technical users are difficult to participate in flow construction. Moreover, template construction lacks flexibility, making it difficult to quickly dynamically adjust module combination and processing order according to specific tasks, resulting in rigid processing flow, poor universality, and severely limited system adaptation capability.
[0068] In some embodiments, the determining process of the declarative template can include receiving a template configuration operation through a visual interface, configuring at least two of a scene understanding module, a speech processing module, and a force perception processing module to define the declarative template. Alternatively, in other embodiments, the determining process of the declarative template can include defining the declarative template based on a preset configuration file.
[0069] Template configuration operation refers to module selection, parameter setting, and execution flow design operations completed by a user through a visual interface (e.g., a graphical user interface, GUI). As an example, template configuration operation allows a user to drag and drop processing modules such as scene understanding, speech processing, and force perception processing, and define dependencies and data flow between them, thereby dynamically constructing an embodied agent data processing pipeline. Scene understanding module refers to a functional unit for processing visual sensor data such as images, videos, and depth maps, for example, including object detection, three-dimensional reconstruction, image segmentation, and other functions, suitable for visual perception tasks. Speech processing module is used to process audio signals collected by the embodied agent, for example, supporting speech recognition, semantic understanding, speech synthesis, and other processing functions, typical applications including natural language interaction, speech command execution, etc. Force perception processing module can process haptic data, force data, or torque data from haptic sensors, force sensors, or torque sensors, and can be used for tasks such as pose adjustment, contact state determination, and grasp stability analysis.
[0070] The preset configuration file is a standardized declarative template defined in advance, for example, pre-configured and generated by a system administrator or engineer for common application scenarios (such as carrying, grasping, and speech interaction). This configuration file can be directly used for deployment without additional user configuration, facilitating quick system initialization or large-scale deployment.
[0071] In the above embodiments, the user can flexibly configure the relationship and processing logic between the scene understanding module, the speech processing module, and the force perception module through the visualization interface, and the system automatically converts it into a structured declarative template, reducing the configuration threshold. Alternatively, the user can quickly generate a standardized declarative template through a preset configuration file to facilitate reuse in batch deployment or general tasks. In possible implementation manners, the above scheme can be flexibly switched between visualization and automation according to specific needs, taking into account ease of use and engineering efficiency. This scheme significantly reduces the technical threshold for constructing data processing flow, enabling non-professional users to complete complex pipeline definition through a visualization interface, improving the configurability and scope of application of the system. By introducing a preset template mechanism, standardized deployment and rapid reuse are supported, further improving the efficiency and consistency of pipeline construction and updating.
[0072] Since the multi-modal data collected by embodied agents has significant differences in processing complexity and real-time requirements, it is difficult to balance response speed and resource utilization efficiency for different tasks using fixed resource allocation or predefined scheduling strategies, especially in the case of sudden increase in task volume or resource shortage, which may lead to system lag, high delay, or resource waste.
[0073] To dynamically analyze data characteristics and task types during operation and schedule computing resources accordingly, in some embodiments, the method can further include: based on the feature extraction result of the multi-modal data and / or the task priority type corresponding to the target data processing job, scheduling the computing resources of the target data processing job; the feature extraction result includes computing power type and / or latency requirement information, and the feature extraction result is obtained by analyzing the metadata of the multi-modal data.
[0074] Metadata refers to additional information related to the ontology of multi-modal data, used to describe its structure, source, acquisition parameters, and content characteristics, etc. For example, the resolution, frame rate, and acquisition timestamp of an image, the sampling rate and duration of a voice, as well as the sensor type, data acquisition frequency, etc. Feature extraction result refers to the processing requirement information obtained by analyzing the metadata of the multi-modal data collected by the target embodied agent. As an example, the feature extraction result includes but is not limited to the required computing power type (such as whether it is a GPU-accelerated task or a CPU-intensive task) and the latency requirement information (Latency SLA) of the task.
[0075] In some embodiments, the metadata field of the multi-modal data can be parsed by a predefined feature extraction function, and the processing features can be automatically labeled according to the data type (such as whether the visual data is a high-resolution video, whether it comes from a force sensor, etc.). For example, if the metadata indicates that the multi-modal data contains a high-resolution video, the feature extraction result can include {compute_type: gpu_accelerated, latency_sla: 200ms}, indicating a GPU-accelerated task with a latency requirement of 200ms; if it is force sensation data, it is labeled as {compute_type: cpu_intensive, latency_sla: 500ms}, indicating a CPU-intensive task with a latency requirement of 500ms. This process can be achieved in a programmed manner, such as by calling the extract_features(data) function to extract the processing features according to the data.metadata field. Wherein, sla is the abbreviation of Service Level Agreement, which means service level agreement.
[0076] The task priority type refers to the relative importance of the target data processing job in the overall data processing task of the embodied intelligent agent. The task priority type can be set based on the business urgency, real-time safety, or external user specified strategy of the task. The task priority type can be used as a scheduling basis together with the feature extraction result to guide the scheduling system to prioritize the allocation of computing resources required by high-priority tasks in the case of limited resources, ensuring that critical system functions are executed first.
[0077] Scheduling computing resources refers to the process of dynamically allocating CPU, GPU, memory, and other computing resources for the target data processing job according to the feature extraction result and / or task priority type. The scheduling strategy can be matched with the real-time demand and computing power adaptation demand of the data processing task for matching scheduling, automatically determining the type and quantity of hardware resources of the node mounted by the container group required by the deployed task. The dynamic resource scheduling mechanism can significantly improve the response ability and overall computing efficiency of the system, supporting automatic adaptation and execution of data processing processes in a heterogeneous hardware environment.
[0078] The above embodiments analyze the metadata of the multi-modal data collected by the embodied agent, obtain feature extraction results, such as required computing power types (for example, whether a GPU is required) or delay requirements, and use the feature extraction results and / or the priority types of the tasks as the basis for computing resource scheduling. On this basis, the system allocates appropriate computing resources (such as CPUs, GPUs, memory resources, etc.) to ensure that the needs of real-time processing are met while the overall performance of the system is taken into account. On the one hand, by intelligently scheduling data features and task priorities, on-demand dynamic allocation of computing resources can be achieved, improving resource utilization efficiency. On the other hand, in resource-constrained or task-intensive scenarios, the performance of critical tasks is prioritized, significantly reducing task waiting time and system response delay, thereby enhancing the stability and real-time performance of the embodied agent in complex environments.
[0079] One data processing job may contain multiple tasks with different priorities, such as visual recognition, speech understanding, and real-time action planning. These tasks differ significantly in terms of resource consumption and response time limits, especially high-priority tasks often have a stronger dependence on GPU acceleration resources. If resource scheduling is based on static allocation or simple average strategies, it is difficult to flexibly adjust resource allocation, resulting in blocked execution of high-priority tasks, affecting overall task response efficiency and agent system stability.
[0080] In some embodiments, the process of scheduling computing resources for the target data processing job can include: in the case where the number of tasks with a first priority type included in the target data processing job is greater than a specified number, increasing the number of GPU nodes available to the target data processing job, and / or suspending the target data processing job's access to at least part of the CPU nodes.
[0081] The first priority type is, for example, a high-priority type, indicating that the priority of the task is relatively high. The specified number is not limited in the above embodiments and can be selected or set according to actual needs. In the above embodiments, when it is detected that the number of high-priority tasks included in the target data processing job exceeds the specified number, the number of GPU nodes available to the target data processing job is dynamically increased to meet its computing intensity requirements. In addition, the system can selectively suspend the target data processing job's access to part of the CPU nodes to avoid resource conflicts caused by mixing GPUs and CPUs, thereby optimizing resource allocation. On the one hand, high-priority tasks can obtain more adequate computing resource support, significantly reducing execution delay and improving response efficiency. On the other hand, by suspending access to part of the CPU nodes, the overall utilization and scheduling flexibility of the cluster can be improved. This solution effectively addresses issues such as task preemption imbalance and rigid resource allocation in multi-task mixed scenarios, providing a foundation for efficient operation of embodied agents in complex dynamic environments.
[0082] Target data processing jobs, for example in the form of a pod, run on cluster nodes for performing tasks such as scene understanding, speech processing, haptics processing, decision generation, etc. Due to the dynamic nature of resource demand of target data processing jobs during running, especially when processing high-frequency sensor data or performing complex neural network inference, the CPU or memory utilization may quickly rise. If the system fails to timely perceive and respond to these resource pressures, it is likely to result in limited or even abnormal termination of job running within the pod, thereby affecting the continuity of the entire data processing pipeline.
[0083] In some embodiments, the method can further include: in response to detecting that the CPU utilization of the target pod exceeds the first utilization and / or the memory utilization exceeds the second utilization, increasing the upper limit of the CPU quota and / or the upper limit of the memory quota of the target pod.
[0084] In the above embodiments, the target pod may, for example, correspond to a pod in a Kubernetes platform, as an example of a schedulable computing unit. A pod can contain one or more containers that share a network, storage volume, and life cycle. In a data processing scenario, the target pod may, for example, run a processing module in a data processing pipeline of a situated agent.
[0085] The CPU utilization and the memory utilization respectively refer to the percentage of the currently used processor resources and memory resources of the pod relative to the allocated resources. As an example, the resource usage of each pod can be obtained through a resource monitoring mechanism as a basis for determining whether to trigger elastic adjustment.
[0086] The above embodiments do not limit the first utilization and the second utilization, which may, for example, be selected or set according to actual needs. The first utilization may, for example, be 80%, and the second utilization may, for example, be 70%. When the resource usage of the target pod exceeds the corresponding threshold, the system regards it as a signal of resource shortage.
[0087] The upper limit of the CPU quota and the upper limit of the memory quota respectively refer to the maximum available CPU resources (e.g., 2 cores) and the maximum available memory resources (e.g., 4 Gi) that can be allocated to the target pod. These upper limits of the quotas can be increased through a dynamic adjustment mechanism during running to meet the immediate performance needs of the task.
[0088] In the above embodiments, the CPU and memory utilization of each container group in the target data processing job is continuously monitored, and when it is detected that the CPU utilization of the target container group exceeds the set first utilization, and / or the memory utilization exceeds the second utilization, the CPU quota upper limit and / or the memory quota upper limit of the target container group is automatically increased, thereby realizing the elastic adjustment of the container group resources. Thus, the data processing interruption caused by resource bottleneck can be effectively reduced, and the job stability is improved. Secondly, by real-time sensing and responding to the change of resource demand, the adaptability of the system to high dynamic data scenarios is enhanced.
[0089] Since the multi-modal data processing flow usually involves multiple heterogeneous processing modules, when performing data processing jobs in complex environments, there is a lack of unified identification mechanism to track the intermediate state of the processing process, which leads to significant difficulties in task debugging, abnormal source tracing, performance evaluation, etc. For example, when the task processing fails or the result is abnormal, engineers have difficulty in accurately locating the specific processing steps, input data sources or environmental background, affecting the maintainability and expansion ability of the system.
[0090] In some embodiments, the method can further include: during the execution of the target data processing job, creating a tracking identifier for tracking the intermediate state of multi-modal data processing, the tracking identifier being used to indicate an agent identifier, an environment context identifier and a current timestamp; saving the tracking identifier, the current processing stage and the current multi-modal fusion indicator in association.
[0091] The tracking identifier refers to a unique identifier generated for identifying and tracking the intermediate state of multi-modal data processing during the execution of the embodied agent data processing job. The tracking identifier can be generated by fusing multiple elements, for example, including: an agent identifier (robot_id), an environment context identifier reflecting the state of the environment in which the robot is located (such as the result of hashing scene_context), and a current timestamp, such as “R001_5A3B_202310201430”, where “R001” is the agent identifier, “5A3B” is the environment abstract prefix (as an example of the environment context identifier), and “202310201430” is the UTC (Coordinated Universal Time) time. The tracking identifier can be used as an index key for the intermediate state of multi-modal data processing, enabling accurate tracking and result backtracking of the intermediate state.
[0092] The current processing stage refers to the current stage of the data processing job during execution, such as "data preprocessing", "feature extraction", "multimodal fusion", "decision generation", etc. As an example, the system can divide the processing process into multiple stages based on the pipeline structure defined in the declarative template, and automatically record the intermediate state and related performance indicators at each stage for subsequent analysis, abnormal diagnosis and dynamic optimization.
[0093] The multimodal fusion indicator refers to a set of indicators used to measure the quality and performance of different modal data fusion during the execution of the multimodal data fusion operation. As an example, the multimodal fusion indicator can include fields such as vision accuracy (e.g. vision_accuracy), force sensor error (e.g. force_sensor_error), synchronization delay (e.g. sync_delay), etc., which can be stored in JSON format. Associating and persisting the tracking identifier, the current processing stage and the current multimodal fusion indicator can achieve accurate recording of the intermediate state of multimodal data processing and comprehensive evaluation of the quality of the agent behavior.
[0094] The above embodiments can automatically generate a tracking identifier for the current intermediate state during the execution of the target data processing job. The tracking identifier can indicate, for example, the unique identifier of the embodied agent, the context hash value of the environment to which the collection task belongs, and the current timestamp, thereby uniquely identifying the running context of the processing task. At the same time, the system associates and saves the tracking identifier with the current processing stage (e.g. multimodal fusion stage) and the current multimodal fusion indicator, providing accurate and traceable intermediate state records for subsequent debugging, monitoring and performance evaluation. The above embodiments enhance the observability and operability of the embodied agent multimodal processing system, allowing developers to quickly locate problems and bottlenecks in complex processing flows. In addition, it provides a support foundation for building a dynamic optimization mechanism for the system, enabling subsequent scheduling strategies and module adjustments to achieve more intelligent resource allocation and performance control based on tracking information.
[0095] Referring to Figure 2 , Figure 2 is a unified modeling framework diagram of a multi-body multi-end provided by an embodiment of the present application.
[0096] In one specific application scenario, the above-mentioned data processing method of the embodied agent is used to build a unified modeling framework of multi-body multi-end, such as Figure 2The capability matrix (as an example of ontology capability information) of each robot body (as an example of embodied agent) is obtained, the operation constraint information (as an example of end effector description information) of each end device (i.e., end effector) is obtained, and a data processing pipeline is constructed based on the information. During data processing, corresponding verification rules are dynamically obtained.
[0097] The above framework supports plug-and-play and intelligent adaptation of cross-brand heterogeneous devices. In addition, the above framework allows intelligent collaborative adaptation engine. Specifically, dynamic configuration injection is allowed. The ontology capability information, end effector description information, and verification rule information are processed into integrated configuration information, and the integrated configuration information is converted into Kubernetes native resources (as an example of integrated configuration resources) to realize version control and gray release. As an example, when a scheduling node in the system detects a new robot_type, the step of generating integrated configuration information is automatically executed. Then, an integrated configuration resource (e.g., ConfigMap) in a unified format is created. After that, a target data processing job is created in an adaptive manner, and the corresponding ConfigMap is automatically bound when the Job is generated. During the execution of the Job, the integrated configuration resource is dynamically loaded to generate internal representations (e.g., Python objects) corresponding to the ontology capability information, end effector description information, and verification rule information. Based on the generated internal representations, a data processing pipeline is constructed in an adaptive manner to process the multi-modal data collected by the embodied agent.
[0098] The above data processing method of embodied agent can use a declarative template to construct a cloud data processing pipeline for processing multi-modal data. The declarative template allows users to define a multi-modal data processing flow through YAML, and supports dynamic combination of scene understanding modules, speech processing modules, force perception modules, decision generation modules, etc. As an example, the input of the scene understanding module can include depth camera data (as an example of visual data) and radar data. The decision generation module may, for example, rely on the scene understanding module and have a high priority type.
[0099] Referring to Figure 3 , Figure 3is a configuration flowchart of a declarative template provided by an embodiment of the present application. In order to solve the problem of multi-module version conflict caused by redeployment of a complete set of algorithm image due to adaptation to a new scene and business rule change, dynamic arrangement of the declarative template can be performed. By introducing a visual low-code configuration platform, a user is allowed to configure a template through a visual interface, so as to realize rapid combination and flexible adjustment of a complex multi-modal processing flow. As an example, the user can complete flow arrangement of each processing module (such as scene understanding, speech processing, force perception processing, decision generation, etc.) through a drag-and-drop flow designer using a user interface, and set running parameters, resource request strategies and dependency relationships of each module in a parameter configuration panel. The system converts the user configuration into a structured declarative template, and stores and manages it through a versioning mechanism, which can support rollback, comparison and reuse. In addition, with the help of a real-time configuration taking effect engine, hot loading or rapid deployment of the template at runtime can be realized, so as to ensure that the configuration change can take effect in real time.
[0100] The above data processing pipeline depends on dependency analysis and resource allocation. By analyzing the DAG (Directed Acyclic Graph) dependency relationship, the critical path (such as the environment understanding module→the decision generation module) can be automatically identified. Secondly, the dynamic allocation of a heterogeneous computing resource pool (such as a GPU exclusive cluster and a CPU elastic pool) can be performed for high-priority tasks. Moreover, dynamic loading and gray release are allowed. For example, a new version of a plug-in image is loaded in real time through the register_flow method of the Prefect API (an application programming interface provided by a workflow arrangement platform). For another example, the new and old version images are allowed to run in parallel, the requests are shunted according to a preset proportion (such as 5% new version + 95% old version), and full switching is performed after monitoring that there is no exception.
[0101] The data processing method of the embodied agent can introduce a feature-driven elastic scheduling mechanism to optimize the computing performance of the data processing method of the embodied agent. To achieve elastic scheduling and optimization of computing performance, a task-aware monitoring mechanism can be introduced to collect CPU usage, client request queue length, and other indicators in real time through a monitoring collector. Second, cloud resource usage features can be collected, and an elastic scaling decision model can be used to achieve task feature-driven elastic scaling, including vertical scaling and horizontal scaling. For example, by extracting features from metadata, the data features of multi-modal data are analyzed and processing requirements are labeled, such as including computing power types and delay requirement information. As an example, the embodied_scaler(task_type, resource_usage) function can be used to generate an embodied task resource demand matrix. Then, based on the task queue depth and feature labels corresponding to the target data processing job, the resources are adjusted, for example, scaling strategies are customized for embodied tasks such as crawling and navigation. In addition, self-healing is achieved when a cloud fault occurs. For example, when a node anomaly is detected (such as GPU ECC error rate > 5%), the task is automatically migrated to a healthy node. ECC is the abbreviation of Error Checking and Correcting, which means error checking and correction. For example, if a Pod is restarted 3 times in a row, the faulty node is automatically isolated and a standby machine is triggered to start the operation.
[0102] As an example, the calculate_vscale(resource_usage) function can be used to implement vertical scaling. For example, if the CPU utilization of the target container is greater than 85% and the memory utilization is greater than 75%, the upper limit of the target container's quota is increased to 2-core CPU and 4Gi memory.
[0103] Referring to Figure 4 , Figure 4 is a schematic diagram of horizontal scaling provided by an embodiment of the present application. By monitoring the index value of the monitoring index (for example, the number of high-priority tasks), a threshold value judgment is performed. When the index value is greater than the corresponding upper threshold value, the number of instances (for example, GPU nodes) is increased; when the index value is less than the corresponding lower threshold value, the number of instances is decreased. Through the automatic increase and decrease of instances, the purpose of load balancing is achieved, thereby achieving management of the service cluster. When increasing or decreasing instances, the control target corresponding to the instances (i.e., the target number of instances) can be calculated by the decision engine according to the preset strategy. As an example, the K8s Controller or other instance orchestrator can be used to perform instance creation and destruction operations.
[0104] The data processing method of the above-mentioned embodied agent can construct a full-link state management system. For example, a context tracking identification mechanism is introduced on the embodied agent. As an example, the function gen_embodied_trace(robot_id, scene_context) can be used to generate a global tracking ID (as an example of tracking identification) that fuses the robot environment context for each piece of multi-modal data. Then, an incremental snapshot technology can be used to establish a multi-modal processing state snapshot to persist the intermediate state containing multi-modal fusion indicators. For example, the intermediate state can be persisted to a designated object storage system (for example, an S3 system) every 5 minutes, and CRC32 is used to check the data integrity. Wherein, CRC is the abbreviation of Cyclic Redundancy Check, which means cyclic redundancy check. The CRC32 algorithm uses a 32-bit polynomial as the divisor, and the remainder obtained by dividing the data by this polynomial is the CRC check code. The above method can also record key operations, for example, the following information in the processing process can be automatically captured: operation type (such as cleaning, calculation, merging), operation time, processing logic version (such as code Git CommitID). In addition, a bloodline map can be constructed to build a cross-robot data processing knowledge network. In addition, a visual query method is supported to provide an interactive interface to trace the data path. By establishing a full-link state management system, the traceability of the intermediate state of data processing is improved, the time required for fault location is shortened, and the risk of interruption of incomplete data processing tasks during system upgrade is significantly reduced.
[0105] Reference Figure 5 , Figure 5 is a zero-downtime hot update flowchart provided by an embodiment of the present application. As an example, after the traffic router receives a client request, the existing request is routed to the old version V1 Pod based on the version label of the Pod, and the new request is routed to the new version V2 Pod. V1 Pod triggers graceful exit after completing all inventory tasks it has received, and no longer receives new requests; V2 Pod continuously receives and processes new requests from the client, realizing version switching without interrupting service. Wherein, the existing request refers to the client request that the system has received and is processing before the hot update starts, and the new request refers to the new client request received by the system after the hot update starts.
[0106] To address the problem that action execution anomalies are difficult to associate with specific data processing links and cross-agent group behavior pattern analysis is difficult, a unified log fingerprint technology can be used for improvement. As an example, the function generate_log_fingerprint(task_id) can be used to generate a unified tracking ID across the system. In addition, a streaming log analysis engine can be used for streaming analysis of logs.
[0107] Referring to Figure 6 , Figure 6 is a flow log analysis schematic diagram provided by an embodiment of the present application. As an example, an agent (which can be a specific implementation of embodied agent) is responsible for collecting system operation logs and sending log data to a designated platform (such as Kafka, an open source distributed stream processing platform) in a streaming manner. The log stream is processed by a real-time processing module, which can detect abnormal behavior in the log based on pre-set abnormal rules or machine learning models and timely trigger alarm operations. In addition, the real-time processing module can also continuously calculate statistical indicators such as log event frequency, module response delay, etc. for generating a dashboard view as part of system state visualization. At the same time, the original log data can be written to a long-term storage system (such as a log archiving platform based on object storage or distributed file system) for subsequent audit, backtracking analysis or model training, etc.
[0108] Referring to Figure 7 , Figure 7 is a structural block diagram of a data processing system of an embodied agent provided by an embodiment of the present application. Industrial robots and collaborative robots as examples of embodied agents can correspond to one ontology node respectively. The control core of the industrial robot is, for example, a robot domain controller. The industrial robot can be connected to end device A, and the collaborative robot can be connected to end device B. End device A and end device B can be different types of end effectors.
[0109] Embodiments of the present application also provide a data processing system of an embodied agent, which includes a plurality of embodied agents, a plurality of end effectors and a data processing engine module, and the data processing engine module is used to execute any of the above methods.
[0110] Currently, the configuration of multi-brand robots in industrial sites is fragmented, and parameter management relies on manual documentation, which is prone to errors and inefficient. The combination of embodied agents and end effectors lacks systematic verification, and physical constraint conflicts (such as workspace overlap and load overlimit) frequently occur. The above system adopts a declarative collaborative management mechanism. Specifically, a three-dimensional configuration model of ontology capability matrix (as an example of ontology capability information) x end description library (also known as end constraint space, as an example of end effector description information) x verification rule library (as an example of verification rule information) is constructed, supporting the unified abstraction of multiple embodied agents and multiple end effectors, and saving in the form of integrated configuration resources, providing capability parameters (including ontology capability information, end effector description information) and verification rule information for subsequent data processing procedures. The system can realize versioned storage and gray release of related configurations through Kubernetes ConfigMap, and automatically trigger security pre-check when the configuration is updated. In order to achieve plug and play, the end effector description information can be standardized to realize the registration of multiple end effectors such as force control gripper and 3D vision within seconds. In addition, a development compatibility pre-check algorithm is provided to automatically intercept end effector combinations that exceed the ontology load or precision threshold.
[0111] As shown in Figure 7 , the data processing engine module can include a dynamic arranger, a rule verifier, and a model trainer.
[0112] The data processing engine module can receive real-time data streams (including multi-modal data) from embodied agents (including industrial robots, collaborative robots). The data processing engine module can use integrated configuration resources and declarative templates to build data processing pipelines to process multi-modal data collected by embodied agents. For example, using multi-modal data to train a specified model, and using the trained model to update the existing model, and the model training can adopt incremental training or online learning closed loop mode to achieve the purpose of continuous optimization of the model.
[0113] Referring to Figure 8 , Figure 8: This is a pluggable data processing architecture diagram provided by an embodiment of the present application. The data processing engine module can be used to provide a pluggable data processing architecture, which includes, for example, a service layer, a user interface layer, and an execution layer. The service layer is provided with an atomic capability registration center, a template configurator (i.e., a dynamic orchestrator), and a task scheduling engine. The user interface layer includes, for example, an API gateway, an image repository, and a task console. The execution layer includes, for example, a Job execution cluster, which is used to perform data processing operations such as data cleaning, desensitization, and modal conversion. Among them, the template configurator is used to configure and generate declarative templates (also known as Job templates). The task scheduling engine is used to schedule tasks for the Job execution cluster, and the task console is used to monitor the status of the task scheduling engine. The API gateway and the atomic capability registration center can interact with data to submit or update images. The atomic capability registration center can interact with the image repository and store Docker images (as an example of an image, a lightweight, executable, independent software package) in the image repository.
[0114] The relevant environmental perception algorithms are deeply coupled with the core decision-making system. When adding new algorithm modules, the service needs to be shut down and updated. In addition, the processing flow of multi-source heterogeneous data streams (vision / speech / force) is solidified, making it difficult to dynamically adapt to diverse task scenarios. The above-mentioned dynamic orchestrator is used to orchestrate declarative templates, allowing users to define data processing processes through a visual interface or configuration files. The data processing engine module can automatically resolve the dependencies between atomic capabilities and optimize resource allocation strategies (such as parallel computing node scheduling and memory pre-allocation). In addition, hot loading and version rollback mechanisms are introduced. For example, dynamic registration / uninstallation of plug-ins through the Prefect API is allowed, and dual-version images are allowed to run in parallel to achieve grayscale releases.
[0115] In response to the problem that the relevant resource pools have difficulty distinguishing between CPU-intensive (path planning data) and GPU-intensive (3D point cloud data) processing requirements, as well as the problem of processing delays due to resource competition for high-priority data streams, the above system can adopt an embodied task-driven elastic scheduling mechanism to improve it. As an example, a dynamic Pod lifecycle management algorithm based on K8s Worker Pool can be used to implement container group-level monitoring and management. In addition, a priority preemptive scheduling mechanism is introduced to automatically identify processing requirements (for example, computing power type, latency requirement information) based on the metadata of multimodal data, and reserve exclusive computing nodes (for example, GPU nodes) for high-priority critical data streams.
[0116] This system can be applied to the field of cloud-based embodied intelligence, providing a cloud-based, end-to-end processing system for multi-entity, multi-terminal intelligent cluster systems (such as industrial robot production lines and autonomous driving fleets). This system integrates core capabilities such as real-time processing of multimodal perception data, dynamic resource scheduling and multimodal data fusion, and end-to-end traceability and control. It establishes an intelligent optimization closed loop from heterogeneous device data access to group collaborative decision-making, and has broad application prospects.
[0117] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the data processing method of any embodied intelligent body in the above embodiments.
[0118] An embodiment of the present application also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the data processing method of any one of the embodied intelligent bodies in the above embodiments.
[0119] The computer program product may be a portable compact disc read-only memory (CD-ROM) and include program code, and may be run on a terminal device, such as a personal computer. However, the computer program product of the embodiments of the present application is not limited thereto, and the computer program product may be any combination of one or more computer-readable media.
[0120] See also Figure 9 , Figure 9 This is a structural block diagram of a computer device provided in an embodiment of the present application.
[0121] An embodiment of the present application also provides a computer device, which includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, it implements any of the above-mentioned data processing methods for embodied intelligent bodies.
[0122] The embodiments of the present application do not limit the computer device, which may be, for example, a local computer device, a cloud computer device, a distributed computer device, etc.
[0123] The computer device may include: a memory 110, a processor 120, and a communication interface 130. The memory 110, the processor 120, and the communication interface 130 are connected via an internal connection path.
[0124] The memory 110 is used to store computer programs. In some implementations, the computer programs may include codes for implementing the methods of the embodiments of the present application.
[0125] The processor 120 is configured to execute the computer program stored in the memory 110 to control the communication interface 130 to receive input data and information, output operation results and the like. In some embodiments, the computer program for implementing the solutions of the embodiments of the present application can be stored in the processor 120 and executed by the processor 120 when the solutions of the embodiments of the present application are implemented by software or firmware.
[0126] The memory 110 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. The non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM). It should be noted that the memory 110 described herein is intended to include, but not limited to, any memory of these and other suitable types. As an example, the memory 110 includes a random access memory (RAM), a cache memory, and a read-only memory (ROM). The memory 110 stores a computer program, which can be executed by the processor 120, so that the processor 120 implements the steps of any of the above methods.
[0127] The processor 120 can be a central processing unit (CPU), and can also be other general-purpose processors, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, etc. The general-purpose processor can be a microprocessor, or the processor 120 can also be any conventional processor.
[0128] In the implementation process, each step of the above method can be completed by the integrated logic circuit of hardware in the processor 120 or the instruction in the form of software. The method disclosed in combination with the embodiments of the present application can be directly embodied as hardware processor execution completion, or executed by hardware and software modules in the processor 120. The software module can be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, register, etc. The storage medium is located in the memory 110, and the processor 120 reads the information in the memory 110, and combines the hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.
[0129] In some implementations, in addition to the hardware units introduced above, the computer device can also include software modules, where the software modules can be, for example, an operating system, a basic input and output system (BIOS), application software, etc.
[0130] The operating system is used to manage hardware and / or software resources of the computer device, and is the kernel and cornerstone of the computer device. The operating system needs to handle basic transactions such as managing and configuring memory, determining the priority of system resource supply and demand, controlling input and output devices, operating network and managing file system, etc. In order to facilitate user operation, most operating systems will provide an operation interface for users to interact with the system.
[0131] The BIOS is used to run hardware initialization in the power-on boot stage, and provides runtime services for the operating system and application programs. In some implementations, the BIOS can also monitor the processor temperature and perform temperature protection strategies, etc.
[0132] The application software, also known as application program, can be understood as software written for a certain special application purpose of the user, and is one of the main classifications of computer software. For example, the application software can be a program for realizing power control, temperature management, etc.
[0133] It can be understood that the specific examples in the present application are only to help those skilled in the art better understand the embodiments of the present application, and do not limit the protection scope of the present application.
[0134] It can be understood that in various embodiments of the present application, the size of the serial number of each process does not mean the order of execution, and the execution order of each process should be determined by its function and inherent logic, and should not constitute any limitation on the implementation process of the present application.
[0135] It can be understood that the various embodiments described in the present application can be implemented alone or in combination, and the present application is not limited thereto.
[0136] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in the description of the application herein is for describing particular embodiments only and is not intended to be limiting of the application. All technical and scientific terms used herein are intended to have the same meaning as commonly understood by one of ordinary skill in the art unless otherwise specifically defined herein. As used in the description of the application and the appended claims, the singular forms "a", "an" and "the" are intended to include plural forms as well, unless the context clearly indicates otherwise.
[0137] Those skilled in the art can appreciate that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether the functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0138] Those skilled in the art can clearly understand that, for the convenience and brevity of the description, the specific working processes of the above-described embodiments can refer to the corresponding processes in other embodiments, which will not be repeated here.
[0139] In the embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interface, device or unit, and can be electrical, mechanical or other forms.
[0140] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, i.e. they can be located in one place or distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the technical solutions of the present application.
[0141] In addition, the functional units in each embodiment of the present application can be integrated into one processing unit, or each unit can exist physically, or two or more units can be integrated into one unit.
[0142] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application or the parts of the related art that essentially contribute or the parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0143] The above merely describes the specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A data processing method for an embodied intelligent agent, characterized in that: The method comprises: Processing the ontology capability information, end effector description information and verification rule information of the target embodied intelligent body into integrated configuration information; Based on the integrated configuration information, construct an integrated configuration resource that can be identified by the target data processing platform; In the target data processing platform, generating a target data processing job with the integrated configuration resource mounted thereon; During the execution of the target data processing job, dynamically loading the integrated configuration resource to generate an internal representation corresponding to the ontology capability information, the end effector description information, and the verification rule information; A data processing pipeline is constructed based on the internal representation, and the data processing pipeline is used to process multimodal data collected by the target embodied intelligent body; the multimodal data includes at least two of visual data, radar data, voice data and force data.
2. The data processing method of embodied intelligent body according to claim 1, characterized in that: Before processing the ontology capability information, the end effector description information, and the verification rule information of the target embodied intelligent body into integrated configuration information, the method further includes: When the type information of the target embodied intelligent body is detected, the ontology capability information, the end effector description information and the verification rule information corresponding to the type information are acquired.
3. The data processing method of embodied intelligent body according to claim 1, characterized in that: The verification rule information is determined based on at least one of the ontology capability information and the end effector description information.
4. The data processing method of an embodied intelligent body according to claim 1, characterized in that: The constructing of a data processing pipeline according to the internal representation comprises: A data processing pipeline is constructed based on the internal representation and the declarative template; the declarative template is used to define the data processing flow of the multimodal data in multiple processing modules, and the multiple processing modules include at least two of a scene understanding module, a speech processing module, and a force processing module.
5. The data processing method of embodied intelligent body according to claim 4, characterized in that: The process of determining the declarative template includes: A template configuration operation is received through a visual interface, and at least two of the scene understanding module, the speech processing module, and the force processing module are configured to define the declarative template.
6. The data processing method of embodied intelligent body according to claim 4, characterized in that: The process of determining the declarative template includes: Based on a preset configuration file, the declarative template is defined.
7. The data processing method of an embodied intelligent body according to claim 1, characterized in that: The method further comprises: Based on the feature extraction results of the multimodal data and / or the task priority type corresponding to the target data processing job, the computing resources of the target data processing job are scheduled; the feature extraction results include computing power type and / or delay requirement information, and the feature extraction results are obtained by analyzing the metadata of the multimodal data.
8. The data processing method of embodied intelligent body according to claim 7, characterized in that: The process of scheduling computing resources for the target data processing job includes: When the number of tasks of the first priority type included in the target data processing job is greater than a specified number, the number of GPU nodes available for the target data processing job is increased, and / or the access of the target data processing job to at least some CPU nodes is suspended.
9. The data processing method of embodied intelligent body according to claim 7, characterized in that: The method further comprises: When it is detected that the CPU utilization of the target container group included in the target data processing job exceeds the first utilization and / or the memory utilization exceeds the second utilization, the CPU quota upper limit and / or the memory quota upper limit of the target container group are increased.
10. The data processing method of embodied intelligent body according to claim 1, characterized in that: The method further comprises: During the execution of the target data processing job, a tracking identifier is created for tracking the intermediate state of the multimodal data processing, wherein the tracking identifier is used to indicate the agent identifier, the environment context identifier, and the current timestamp; The tracking identifier, the current processing state and the current multimodal fusion index are associated and saved.
11. A data processing system for an embodied intelligent agent, characterized in that: The system includes multiple embodied intelligent agents, multiple end effectors and a data processing engine module, and the data processing engine module is used to execute the method according to any one of claims 1 to 10.
12. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 10 is implemented.
13. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 10 when executing the computer program.
14. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 10 is implemented.
Citation Information
Cited By
FT6678-based multi-core radar target detection method and device
CN121578240A
An embodied intelligent device interaction control method and system
CN122363939A