Operation control method and device based on object flow sequence, equipment and medium

By generating a target object flow sequence independent of the operator's form and establishing a mapping relationship in the simulation environment, the problem of insufficient cross-domain applicability in robot operation learning is solved, achieving low-cost, high-precision operation control and improving the robot's adaptability and accuracy in different environments.

CN120928733APending Publication Date: 2025-11-11PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202511051762.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing robot operation learning technologies face several challenges in cross-domain applications, including difficulty in eliminating subject differences, reliance on high-cost real data for the learning process, limited transfer effects between simulation and reality, and insufficient system applicability and task generalization capabilities. In particular, they struggle to achieve low-cost, rapid deployment, and high-precision operation control in the fields of healthcare and fintech.

Method used

By acquiring demonstration data and task descriptions representing the task, a target object flow sequence independent of the operator's specific form is generated. A mapping relationship between the object flow sequence and the execution device action sequence is established in the simulation environment, a flow-conditional action strategy is generated, and the execution device action sequence is dynamically adjusted by comparing the object flow sequence in real time during execution.

Benefits of technology

It achieves cross-entity, universal, and low-cost dynamic operation control, improves the versatility, accuracy, and adaptability of the robot operating system, and reduces system development and deployment costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120928733A_ABST
    Figure CN120928733A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, can be applied to business scenes such as robot agent decision making, financial science and technology and medical health, and discloses an operation control method, device, equipment and medium based on an object flow sequence. The method comprises the following steps: exploring a mapping relation between an object flow sequence and an execution equipment action sequence in a simulation environment, generating a flow conditional action strategy, inputting a target object flow sequence to generate an initial execution equipment action sequence, tracking a real-time motion state of a target object to obtain a real-time object flow sequence, and comparing the real-time object flow sequence with the target object flow sequence. And dynamically adjusting the action sequence of the execution equipment based on the deviation. According to the method, the object flow and execution equipment action sequence mapping is constructed, and real-time object flow tracking and dynamic adjustment are combined, so that the influence of operator difference and environment uncertainty is reduced, and the operation accuracy and universality of the robot are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to an operation control method, apparatus, device, and storage medium based on object stream sequences. Background Technology

[0002] Existing robot manipulation learning technologies generally suffer from several problems in cross-domain applications, including difficulty in eliminating subject differences, reliance on high-cost real-world data for the learning process, limited transfer effects between simulation and reality, and insufficient system applicability and task generalization ability. These problems severely restrict the widespread implementation of robot manipulation learning in multiple industries and complex environments. In particular, under different object types and subject conditions, existing technologies struggle to achieve low-cost, rapid deployment, and high-precision manipulation control systems. Furthermore, existing methods lack object behavior representations independent of the operator's specific form during task execution, resulting in generally insufficient interpretability and system transparency in manipulation learning, further exacerbating application barriers in high-safety and highly regulated industries.

[0003] In the healthcare sector, with the widespread adoption of intelligent systems such as surgical robots and rehabilitation aids, robots need to adapt to diverse scenarios involving different patients' physiological structures and dynamic state changes. Existing operational learning technologies, heavily reliant on manual instruction and struggling to effectively distinguish between the operator's actions and the independent movement of the target object, result in weak system generalization capabilities, long training cycles, and high deployment costs, failing to meet the demands for rapid clinical adjustments and high-standard operational safety. Furthermore, the complexity of simulation environment construction and its discrepancies with real-world environments lead to low strategy transfer efficiency, impacting the reliability and practicality of robots in actual medical operations.

[0004] In the fintech sector, the demand for smart terminals, service robots, and business automation equipment in areas such as document delivery, identity verification, and data processing is constantly growing. Existing technologies generally rely on collecting large amounts of actual robot operation data for specific business scenarios. Operation templates and strategy design are highly dependent on human experience, resulting in a lack of system versatility and rapid adaptability, severely impacting business expansion efficiency and system stability. Furthermore, due to the differences in perception and execution between simulation and actual business environments, the accuracy during strategy migration is insufficient, failing to guarantee stable robot performance in high-frequency, dynamically changing financial business scenarios, increasing the difficulty and cost of system maintenance and upgrades.

[0005] In other applications of intelligent robot manipulation technology, such as industrial manufacturing, logistics handling, and daily services, the types of manipulated objects are diverse, including rigid bodies, joint structures, and flexible and deformable objects. Existing methods struggle to uniformly adapt to the characteristics of different objects, resulting in poor system versatility, low learning efficiency, and heavy reliance on task-specific simulation environment construction and strategy fine-tuning. Furthermore, the lack of object flow representation independent of the operator's specific form during the manipulation learning process makes it difficult for robots to generate stable and accurate manipulation strategies based on cross-subject data, limiting the application scope and promotion effectiveness of robot operating systems. Summary of the Invention

[0006] The main objective of this invention is to provide an operation control method, apparatus, device, and storage medium based on object flow sequences, aiming to solve the technical problem that the prior art lacks the ability to achieve dynamic operation control that is universal across subjects, low-cost, and does not rely on a large amount of real robot data, based on object flow sequences independent of the specific form of the operator.

[0007] To achieve the above objectives, the present invention provides an operation control method based on an object stream sequence, comprising:

[0008] Obtain demonstration data and task description for the representation task;

[0009] Based on the demonstration data and task description, a target object stream sequence independent of the operator's specific form is generated;

[0010] In the simulation environment, by exploring preset action primitives, a mapping relationship between object flow sequences and execution device action sequences is established, and flow conditional action strategies are generated.

[0011] The target object stream sequence is input into the stream conditionalization action strategy to generate an initial execution device action sequence;

[0012] The initial execution device action sequence is executed, and during the execution process, a real-time object stream sequence is obtained by tracking the real-time motion state of the target object;

[0013] The real-time object stream sequence is compared with the target object stream sequence, and the execution device action sequence during the execution process is dynamically adjusted based on the deviation of the comparison.

[0014] Furthermore, to achieve the above objectives, the present invention provides an operation control device based on an object stream sequence, comprising:

[0015] The data acquisition module is used to acquire demonstration data and task descriptions for representing the task;

[0016] The object stream generation module is used to generate a sequence of target object streams that are independent of the specific form of the operator based on the demonstration data and task description.

[0017] The mapping strategy training module is used to explore and establish the mapping relationship between the object flow sequence and the execution device action sequence in the simulation environment through preset action primitives, and generate flow conditional action strategies.

[0018] An action sequence generation module is used to input the target object stream sequence into the stream conditional action strategy to generate an initial execution device action sequence;

[0019] The real-time tracking and acquisition module is used to execute the initial execution device action sequence and, during the execution process, acquire a real-time object stream sequence by tracking the real-time motion state of the target object;

[0020] The action adjustment control module is used to compare the real-time object stream sequence with the target object stream sequence, and dynamically adjust the action sequence of the execution device during the execution process based on the deviation of the comparison.

[0021] Furthermore, to achieve the above objectives, the present invention also provides a computer device, the computer device including a memory, a processor, and an operation control program based on an object stream sequence stored in the memory and executable on the processor, wherein when the operation control program based on the object stream sequence is executed by the processor, it implements the steps of the operation control method based on the object stream sequence as described above.

[0022] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing an operation control program based on an object stream sequence, wherein the operation control program based on the object stream sequence, when executed by a processor, implements the steps of the operation control method based on the object stream sequence as described above.

[0023] Beneficial Effects: This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as robot intelligent agent decision-making, fintech, and healthcare. It discloses an operation control method, apparatus, device, and medium based on object flow sequences, including: acquiring demonstration data and task description representing the task; generating a target object flow sequence independent of the operator's specific form based on the demonstration data and task description; exploring and establishing a mapping relationship between the object flow sequence and the execution device action sequence in a simulation environment through preset action primitives and generating a flow-conditionalized action strategy; inputting the target object flow sequence into the flow-conditionalized action strategy to generate an initial execution device action sequence; tracking the real-time motion state of the target object to obtain a real-time object flow sequence during the execution of the initial execution device action sequence; comparing the real-time object flow sequence with the target object flow sequence and dynamically adjusting the execution device action sequence during the execution process based on the comparison deviation. This invention combines a target object flow sequence independent of the operator's specific form with a flow-conditionalized action strategy, unifying cross-subject video data and simulation exploration data for operation control processes. This avoids reliance on large amounts of real robot data, improves the universality, accuracy, and adaptability of cross-domain operations, and reduces system development and deployment costs. Attached Figure Description

[0024] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings:

[0025] Figure 1 This is a schematic diagram of an application environment for an operation control method based on object stream sequence according to an embodiment of the present invention;

[0026] Figure 2 This is a flowchart illustrating an embodiment of the operation control method based on object stream sequence of the present invention;

[0027] Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the operation control device based on object stream sequence of the present invention;

[0028] Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention;

[0029] Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation

[0030] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.

[0031] The operation control method based on object stream sequence provided in this invention can be applied to, for example... Figure 1In this application environment, the user terminal communicates with the server via a network. The server can obtain demonstration data and task description representing the task from the user terminal. Based on the demonstration data and task description, it generates a target object flow sequence independent of the operator's specific form. In the simulation environment, it explores and establishes a mapping relationship between the object flow sequence and the execution device action sequence through preset action primitives and generates a flow-conditional action strategy. The target object flow sequence is input into the flow-conditional action strategy to generate an initial execution device action sequence. During the execution of the initial execution device action sequence, the real-time motion state of the target object is tracked to obtain the real-time object flow sequence. The real-time object flow sequence is compared with the target object flow sequence, and the execution device action sequence is dynamically adjusted based on the comparison deviation. This invention combines the target object flow sequence independent of the operator's specific form with the flow-conditional action strategy, unifying cross-subject video data and simulation exploration data for the operation control process. This avoids reliance on a large amount of real robot data, improves the universality, accuracy, and adaptability of cross-domain operations, and reduces system development and deployment costs. The user terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The present invention will now be described in detail through specific embodiments.

[0032] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the operation control method based on object flow sequence provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.

[0033] like Figure 2 As shown, the operation control method based on object stream sequence proposed in this invention includes the following steps:

[0034] S10, Obtain demonstration data and task description representing the task;

[0035] In this embodiment, acquiring demonstration data and task description for the characterization task comprises two parts: demonstration data and task description. Both are used together to subsequently construct an object flow sequence independent of the operator, enabling cross-subject migration of the operation control process. Demonstration data refers to a set of data reflecting the dynamic information of the operator during task execution. Specifically, it originates from demonstrations of humans or other operating subjects captured by video acquisition devices. The video data possesses visual temporal continuity, facilitating the extraction of action information related to object operations. This type of video data is acquired in real-time through video acquisition devices. These devices can be any image sensing device with imaging capabilities, including cameras, depth cameras, 3D reconstruction equipment, and sensing systems with visual perception functions; the specific device type is not limited. The video data acquisition process can be configured with parameters such as resolution, frame rate, and viewing angle to adapt to the data quality requirements of different demonstration scenarios, ensuring the accuracy of subsequent information extraction.

[0036] The task description is an operational intent information expressed in linguistic form, obtained through a text input interface. The text input interface refers to an interactive device capable of accessing text data, such as a speech recognition terminal, graphical user interface, natural language processing system, or other hardware or software modules with text input functionality. The task description text specifically reflects the operator's expected operational goals and task content, such as "pour water into a cup" or "open the box to retrieve an item," providing clear task context information to assist in semantic understanding of video data and task logic modeling.

[0037] After acquiring the demonstration data, it needs to undergo standardization processing. Standardization operations include, but are not limited to, resolution standardization, frame rate standardization, color space adjustment, and image distortion correction. The purpose is to eliminate data inconsistencies caused by differences in equipment, environment, and shooting parameters, ensuring the accuracy of subsequent information extraction. Standardized video data, after processing, possesses a unified data format and structure, making it suitable for efficient processing by different backend systems.

[0038] Upon receiving the task description text, semantic parsing is performed. This process relies on natural language processing (NLP) technology to identify key semantic units such as action verbs, object nouns, and spatial location words, generating a structured task description. The structured task description is a data representation format designed for computing systems, transforming natural language information into a clearly structured and semantically explicit data structure, such as tree structures, graph structures, and sequence structures, facilitating joint analysis with video data.

[0039] After the demonstration data and task description have been processed, their consistency is verified using a temporal alignment algorithm. Specifically, this involves calculating the temporal relationship between the object motion trajectories in the video data and the expected actions in the task description to determine if there is a logical match. The temporal alignment algorithm can be based on dynamic time warping, sequence similarity analysis, semantic matching, etc., to ensure the validity of the final collected data and filter out irrelevant, erroneous, or data segments that do not conform to the task objectives.

[0040] The validated data was used as both demonstration data and task descriptions in the subsequent object stream sequence generation process. The acquisition and processing of the demonstration data and task descriptions ensured data quality, semantic integrity, and structural consistency, providing a data foundation for the cross-agent generalized application of operational control.

[0041] Video acquisition equipment can utilize multi-view synchronous camera systems to cover complex demonstration scenarios and improve the spatial information integrity of the data. For low-light environments or high dynamic range scenarios, infrared imaging equipment and HDR imaging technology can be combined to ensure clear and usable video data. The text input interface can integrate a speech recognition module to convert speech signals into text data, enabling task input in natural language form and improving human-computer interaction efficiency. During the standardized video data generation process, convolutional neural networks can be used to perform denoising and image enhancement on video frames, further improving video quality. The semantic parsing stage can be based on pre-trained language models such as BERT and GPT to improve the accuracy of semantic understanding of task descriptions. The temporal alignment algorithm can be combined with multimodal information fusion strategies, utilizing the joint features of visual and linguistic information to improve the robustness and accuracy of the alignment process, ensuring the reliability of the final selected data.

[0042] Example Description: In a robotic agent decision-making scenario, the system records the entire process of a human demonstrating the operation of a robotic arm to move an object using video capture equipment. The demonstration data captures the relative positional changes between the robotic arm's end effector and the object, the operation path, and obstacle information in the environment. The task description received by the text input interface is "move the target object to the designated location." The system standardizes the video data to ensure a consistent data format across different demonstration devices. Combined with the semantically parsed task description, it accurately extracts the operational intent and object relationships. Through temporal alignment, it verifies the consistency between the human demonstration process and the task requirements, selecting valid data as input for subsequent object flow generation and policy learning. This enhances the robot's autonomous decision-making ability and cross-agent transfer capability when facing different environments and tasks.

[0043] In fintech business scenarios, intelligent robots need to perform the handling and retrieval of high-value items during vault management. Staff record demonstrations of the item handling process via a video system. The video data details the item's specifications, weight, storage location, and handling route. Simultaneously, staff input a task description, "Move designated valuable items to a secure storage area," through a graphical interface. The system jointly processes the video and text data, standardizing data structures from different sources. Combining semantic analysis and temporal alignment, it ensures a high degree of consistency between the demonstration process and the operational objective. This avoids operational errors caused by unclear task information or inaccurate demonstration data, guaranteeing a high level of security and automation in the flow of goods within financial transactions.

[0044] In healthcare settings, rehabilitation assistive robots need to learn patient positioning or daily auxiliary operations based on demonstrations by nursing staff. The system records the entire process of nursing staff assisting patients with turning over and sitting up using multi-angle video equipment. The video data accurately captures changes in the patient's physical state and the coordinated operation of the assistive devices. Simultaneously, the nursing staff provides a task description via voice input, such as "assist the patient in safely turning over." The system performs resolution and frame rate standardization on the video data, combines it with a structured task description after semantic parsing, and verifies the consistency of the content through temporal alignment. Valid demonstration data is then selected to ensure that the robot can accurately understand operational needs and the patient's state in different patients and nursing environments, thus improving the safety and generalization capabilities of intelligent assistive devices in medical care.

[0045] This embodiment eliminates the heterogeneity of data under different demonstration subjects, different acquisition devices and environmental conditions by jointly acquiring and processing demonstration data and task descriptions. It ensures the consistency between the expression of operational intent and object behavior data, improves the data foundation quality of subsequent operation control processes, reduces the need for human intervention, and enhances the system's generalization ability and adaptability.

[0046] S20, Generate a target object stream sequence independent of the operator's specific form based on the demonstration data and task description;

[0047] In this embodiment, the demonstration data is acquired through a front-end video capture device and contains complete video information of the operator performing a specified task. The data reflects the spatial location, appearance features, dynamic changes, and environmental background factors of the operated object. The demonstration data source is not limited to a single perspective and can be acquired through a combination of multiple cameras and angles to ensure that data from different operators and under different shooting conditions has a unified expression format. The task description is structured semantic information received based on a text input interface, including operation instructions, object attributes, and target requirements. The task description can be input in various forms such as natural language, key phrases, and icon labels. The system performs semantic parsing and structural extraction on the task description to extract the operation intention, object category, and action target.

[0048] The system uses demonstration data and task description as joint inputs to an object stream generation network, which works collaboratively through a feature encoding layer and a temporal decision layer. The feature encoding layer extracts spatial structure and temporal dynamic features from the demonstration data and combines them with the operational target information in the task description to construct a multi-dimensional fusion representation. The feature encoding layer can achieve efficient extraction and compression of spatial and temporal features based on 3D convolutional networks, graph neural networks, or attention mechanisms. The temporal decision layer, through a sequence modeling structure and combined with the encoded multimodal information, identifies the target object in the video based on the task description, further locates the key points of the target object during the operation process, and extracts the key point trajectory of the target object.

[0049] "Operator-independent" means that during the generation of the target object flow sequence, information related to the operator's own structure, movement, and external features in the demonstration data is removed, retaining only data related to the target object's state changes, spatial position, and posture adjustments. This processing method ensures that the expression of the target object flow sequence is unaffected by the operator's specific physiological structure, operating style, body size, or behavioral patterns, thus possessing cross-subject and cross-device generalization capabilities. Through this method, the system can extract object motion features highly relevant to the specific operation task but independent of the demonstrator's physical condition from demonstration data from any source, avoiding information interference caused by operator differences and achieving a more stable and universal expression of object behavior. For example, in the field of robot intelligent agent operation, the system analyzes video data of human demonstrations, automatically excluding specific motion information of human arms, fingers, and other parts during the demonstration, extracting only the trajectory changes, posture adjustments, and positional movements of the manipulated object during the demonstration, ultimately generating a target object flow sequence that does not contain human body structure information. This allows robots, even with structures completely different from humans, to still complete consistent object operation tasks based on this sequence.

[0050] During keypoint trajectory extraction, the system employs background separation and instance segmentation techniques to eliminate the operator's own body structure, motion information, and background interference factors, retaining only the independent motion features of the target object within the scene. This ensures that the extracted motion state information does not depend on the operator's specific form and body structure. Instance segmentation can utilize deep segmentation models such as Mask R-CNN and DeepLab. Keypoint tracking technology combines optical flow estimation, image feature matching, and multi-frame fusion algorithms to achieve continuous and stable dynamic tracking of the target object.

[0051] Based on the key point trajectory change information of the target object, the system constructs a motion state change sequence of the target object. This sequence reflects the spatial position changes, motion direction, velocity characteristics, and deformation information of the target object throughout the demonstration process. As a target object flow sequence, this sequence has the ability to express consistent information across subjects and environments. The target object flow sequence only expresses the motion changes of the object itself and does not contain operator action characteristics or redundant environmental information, which facilitates the unified processing of subsequent policy learning and control mapping.

[0052] The feature encoding layer of the object stream generation network can extract spatial structural features and temporal dynamic features of video sequences based on 3D convolutional networks, or establish topological relationships between target objects and other elements in the environment using graph neural network structures, thereby improving the segmentation and tracking capabilities of target objects in complex scenes. The temporal decision layer can employ sequence modeling structures such as recurrent neural networks and Transformers to achieve dynamic capture and trend prediction of target object motion changes over long periods, ensuring accurate identification of target objects and extraction of stable motion trajectories under different operator demonstration data.

[0053] In instance segmentation and keypoint tracking, depth image segmentation and multimodal fusion techniques can be combined to improve the accuracy of target object recognition and motion information extraction under low light, occlusion, or complex background conditions. For highly dynamic environments or situations with multiple target interference, segmentation and tracking parameters can be dynamically adjusted to ensure the continuity and accuracy of the target object stream sequence, adapting to the application needs of different industries and tasks.

[0054] Example description: In the robot intelligent agent decision-making scenario, the industrial robot obtains the process video of the robotic arm operating the object through human demonstration. Combined with the task description "to move the target object to the designated position", the system excludes the operator's arm structure and action information, and only extracts the independent motion changes of the target object during the operation. It generates a target object flow sequence that is independent of the operator, which improves the robot's adaptability to different environments and different demonstration data.

[0055] In fintech business scenarios, security robots record the demonstration process of security personnel handling valuable items based on the monitoring system. The system combines the task description of "moving high-value items to a safe location", extracts the independent movement trajectory of the target items, eliminates the interference of the security personnel's physical characteristics, and generates a target object flow sequence, ensuring that the robot can accurately perform item handling and security tasks under different operators and different monitoring environments.

[0056] In healthcare scenarios, assistive nursing robots acquire video footage of healthcare workers demonstrating patient turning over, sitting up, and other nursing procedures. Combined with the task description of "assisting patients in turning over safely," the system extracts the trajectory of the patient's body movements, excludes the characteristics of the caregiver's actions, and generates a target object flow sequence. This ensures that the robot can accurately understand and execute assistive operations under different caregivers and patient conditions, improving the practicality and reliability of intelligent nursing equipment in changing medical environments.

[0057] This embodiment generates an object stream network by jointly inputting demonstration data and task description into the network. The system can stably acquire target object stream sequences independent of the operator's specific form under different operators and environments. This avoids the limitations of traditional operator action mapping and effectively solves the migration barriers caused by differences in operator body structure and changes in environmental conditions. It also improves the generalization ability and cross-domain adaptability in subsequent strategy learning and task execution.

[0058] S30: In the simulation environment, by exploring preset action primitives, a mapping relationship between object flow sequences and execution device action sequences is established, and a flow conditional action strategy is generated.

[0059] In this embodiment, the simulation environment is an operational experimental platform constructed using a virtual 3D scene and a physics engine. It possesses controllable object, environment, and equipment models, capable of simulating spatial structures, dynamic characteristics, and interactive feedback in the real world. It is suitable for strategy exploration and data generation under risk-free and low-cost conditions. Preset action primitives refer to a set of basic operational units covering common spatial displacement, rotation, grasping, and releasing operation modes. Action primitives can be flexibly defined according to the structural characteristics, motion capabilities, and operational requirements of different execution devices, forming an action combination space suitable for the target task range.

[0060] The system controls execution devices in a simulation environment to execute preset action primitives in different combinations, simulating the dynamic interaction between the devices, target objects, and the environment. During each action execution, the system records the changes in the target object's motion state in the simulation environment, forming an object flow sequence, and records the sequence of control commands issued by the execution devices, forming an execution device action sequence. The object flow sequence expresses the spatial position changes, motion trajectory, and state adjustments of the target object caused by the actions of the execution devices, while the execution device action sequence expresses the specific control commands and operating modes of the devices.

[0061] Based on multi-round, comprehensive exploration of pre-defined action primitive combinations, the system collects a large amount of data corresponding to object flow sequences and execution device action sequences, constructing a mapping dataset. In this dataset, each data point records the correspondence between a specific execution device action sequence and its corresponding target object flow sequence, reflecting the specific impact of different operations on the change in the target object's motion state.

[0062] The system trains a policy network based on a mapping dataset. This policy network can employ supervised learning, deep reinforcement learning, or sequence prediction structures, and is capable of generating dynamically matching execution device action sequences from input object stream sequences. The trained policy network becomes a stream-conditionalized action policy, capable of dynamically inferring and outputting adapted execution device action sequences based on the input target object stream sequence. The policy output can be directly used for action control or policy transfer in real-world environments.

[0063] The simulation environment can be built based on physics engines such as Gazebo, MuJoCo, or Unity. The environment integrates a multi-dimensional model of the target object, execution devices, and environmental interference factors, ensuring that the interaction process possesses realistic physical properties and dynamic responses. Preset motion primitives can be configured for different types of execution devices. For example, for robotic arms, motion primitives include single-joint rotation, multi-joint linkage, end effector displacement, and grasping / releasing operation modes; for mobile robots, motion primitives include forward movement, turning, acceleration / deceleration, and path adjustment operation modes. The design scope of motion primitives can be flexibly expanded according to task complexity and device structure.

[0064] During the construction of the mapping dataset, the system explores large-scale and diverse combinations of action primitives to enhance data coverage and diversity, ensuring that the policy network possesses stable mapping capabilities across different object types and operational scenarios. The training of the policy network can incorporate sequence-to-sequence mapping structures, utilizing models such as temporal convolutional networks, Transformers, and recurrent neural networks to improve the ability to capture and express the correspondence between long-term object stream sequences and action sequences. For different types of target objects, such as rigid bodies, jointed structures, or deformable objects, the system can dynamically adjust the policy network structure and training parameters to optimize policy inference and action output effects under different operational scenarios.

[0065] Example description: In the decision-making scenario of robot intelligent agents, the robotic arm device performs grasping and handling tasks in a simulation environment through different angles and path combinations. The system records the changes in the motion state of the target object under each action combination, constructs the mapping relationship between the object flow sequence and the action sequence, and the trained flow conditional action strategy can adapt to different object shapes and placement environments, improving the robot's adaptability to complex tasks and unknown environments.

[0066] In fintech business scenarios, intelligent logistics robots perform high-value asset handling tasks in simulation systems. The preset action primitives cover operation modes such as path planning, obstacle avoidance, and stable control of items. The system generates a mapping dataset and flow-conditional action strategies to ensure that the robot can accurately perform handling and safe operation tasks in actual high-risk and high-value environments, reducing safety risks and economic costs in the operation process.

[0067] In healthcare scenarios, assistive nursing robots simulate patient movement, posture adjustment, and rehabilitation training through simulation systems. The preset action primitives include assistive operations with different intensities and paths. The system trains strategies based on mapping relationship datasets to generate flow-conditional action strategies, enabling the robot to flexibly adjust assistive actions in actual medical environments according to different patient body types and different operational needs, thereby improving patient comfort and the safety of the nursing process.

[0068] This embodiment executes a preset action primitive combination in the simulation environment, enabling the system to efficiently obtain the mapping relationship between the object flow sequence and the execution device action sequence. This reduces the risk and cost of relying on real environment data acquisition, bridges the subject gap problem in cross-subject migration and simulation strategy transformation, and improves the applicability, versatility, and reliability of flow conditional action strategies in multi-task, multi-device, and cross-environment operations. It also avoids the problem of repeatedly building high-cost simulation systems for a single task or device structure.

[0069] S40, input the target object stream sequence into the stream conditionalization action strategy to generate an initial execution device action sequence;

[0070] In this embodiment, the target object flow sequence is generated based on demonstration data and task description, representing the change process of the target object's motion state in space. It includes time-continuous position information, posture information, velocity information, and trajectory features, reflecting an ideal operational target that is independent of the specific operator's body structure. The flow-conditional action strategy is a mapping model established based on previous simulation exploration. It has the ability to infer the matching execution device action sequence based on the input object flow sequence. The strategy integrates dynamic mapping parameters for different object types and different task requirements.

[0071] The system inputs the target object stream sequence into a feature encoding structure for the conditional action strategy. The feature encoding structure can use a sequence encoding network, a spatiotemporal feature extraction network, or a multidimensional feature processing module that combines graph structures to extract multidimensional information such as spatial position changes, posture evolution, velocity distribution, and trajectory trends in the target object stream sequence, forming a spatiotemporal feature representation that expresses the dynamic characteristics of the target object.

[0072] The spatiotemporal feature representation is input into the temporal decision structure of the flow conditional action strategy. The temporal decision structure can adopt a dynamic decision network based on sequence prediction, sequence generation or reinforcement learning inference. Combined with the input spatiotemporal feature representation, it predicts and generates the corresponding action vector sequence of the execution device. The action vector sequence contains the device control parameters, motion planning information and operation instruction features in a continuous time period, reflecting the spatial operation and dynamic adjustment process that the device needs to perform.

[0073] The system performs kinematic constraint processing on the motion vector sequence, taking into account the structural parameters, motion capabilities, joint limits, and dynamic characteristics of the specific execution device. It corrects unreasonable operational parts in the motion vector sequence to ensure that the motion output conforms to the actual reachability, safety, and physical constraints of the device. Kinematic constraint processing can be implemented using methods such as inverse kinematics solving, collision detection, and dynamic balance control.

[0074] The motion vector sequence after kinematic constraint processing is converted into an initial execution device motion sequence. The initial execution device motion sequence is expressed in a standard control instruction format supported by the device control interface, which has the function of directly driving the device to perform operation tasks, ensuring that the output motion sequence is compatible with the device control protocol and meets the needs of subsequent control and dynamic adjustment.

[0075] The target object flow sequence can include spatial 3D coordinates, attitude Euler angles, velocity vectors, and acceleration information, suitable for describing the state of different types of target objects in multi-task operation scenarios. The feature encoding structure of the flow conditional action strategy can be based on 3D convolutional networks, temporal graphical neural networks, or Transformer structures to extract global and local dynamic features of the target object flow sequence. The temporal decision structure can employ recurrent neural networks, attention mechanism networks, or time-step-based hierarchical decision networks to dynamically generate multi-step continuous device action vector sequences in combination with different task requirements.

[0076] During kinematic constraint processing, constraint rules can be dynamically adjusted based on the physical parameters of different actuators. For example, for robotic arms, the angle limits, speed limits, and torque constraints of each joint are considered; for mobile robots, the turning radius, acceleration smoothness, and path accessibility are considered; and for flexible devices, the deformation range, structural stability, and operational safety boundaries are considered. The initial actuator motion sequence can generate pulse width modulation signals, position control commands, or speed control curves based on different device control interfaces, ensuring the physical executability and logical coherence of the motion output.

[0077] Example description: In the decision-making scenario of robot intelligent agents, the robotic arm automatically generates an initial sequence of multi-joint linkage execution equipment actions based on the input target object flow sequence. The action sequence is optimized by kinematic constraints to ensure that the equipment has a smooth motion path and structural stability in handling, assembly or sorting tasks, thereby improving the operational reliability and adaptability under complex tasks.

[0078] In fintech business scenarios, data center operation and maintenance robots generate initial execution device action sequences based on the input target object flow sequence. These sequences are used to automate server hardware maintenance, module replacement, or line inspection operations. The action sequences are dynamically optimized by combining the high-density equipment layout of the data center with the robot's kinematic parameters. This ensures that the robot accurately executes the specified operation path in a confined space, reducing the need for manual intervention, improving the operation and maintenance efficiency and security of financial information infrastructure, and enhancing system stability and business continuity.

[0079] In healthcare scenarios, rehabilitation robots combine the patient's target object flow sequence to generate an optimized initial execution device action sequence, assisting the patient in completing rehabilitation training, position adjustment, or movement operations. The action sequence is processed through kinematic constraints to avoid operational discomfort and safety hazards caused by differences in patient body size and equipment structure limitations, thereby improving the comfort, adaptability, and intelligence level of medical operations.

[0080] This embodiment enables the system to efficiently and accurately generate initial execution device action sequences that conform to the dynamic characteristics of the target object and the physical limitations of the device by conditionalizing the target object stream sequence input stream action strategy. This solves the problem of difficult action migration caused by differences in operators and inconsistent device structures in traditional strategies, improves the system's operational versatility, control accuracy and dynamic adaptability in multi-device and multi-task environments, and reduces the risk of task failure and human dependence in the operation process.

[0081] S50, execute the initial execution device action sequence, and during the execution process, obtain the real-time object stream sequence by tracking the real-time motion state of the target object;

[0082] In this embodiment, executing the initial execution device action sequence refers to controlling a device with execution capabilities to perform specific physical actions based on previously generated action instructions. The execution device can be a robotic arm, a mobile platform, an execution tool, or a robot system with operational functions. The action sequence is a structured instruction set containing information such as position parameters, posture parameters, and velocity parameters, ensuring that the device completes the operation according to a predetermined path and method. During execution, a real-time object flow sequence is obtained by tracking the real-time motion state of the target object. The target object is the object being manipulated or observed, and its category is not limited to rigid bodies, joint bodies, or deformable bodies; it can include medical devices, financial terminals, production components, etc. The real-time motion state refers to the dynamic position changes, posture changes, and structural deformations of the target object under the influence of the execution device's actions. Obtaining the real-time object flow sequence means forming a data sequence that reflects the trend of the target object's state changes through continuous state perception. This sequence can be expressed as a set of key point locations, a three-dimensional spatial trajectory, or an encoded result of an image sequence. The specific tracking process is implemented through a multimodal sensing system, which includes, but is not limited to, a visual camera, a depth sensor, an inertial measurement unit, and a force feedback device. Through information fusion and data synchronization, the real-time state changes of the target object are accurately captured. The data acquisition process is real-time and continuous, ensuring the integrity and coherence of the object stream sequence to meet the needs of subsequent dynamic adjustment or control optimization.

[0083] Binocular vision combined with an inertial measurement unit (IMU) can be used to track the spatial position and attitude of a target object. Image recognition algorithms can be used to extract key feature points on the target object's surface in real time. By calculating the displacement of feature points between consecutive frames, real-time motion state information of the target object can be formed, and an object flow sequence can be further constructed. The sequence information can be compressed and encoded to reduce data transmission latency. Alternatively, LiDAR and a depth camera can work together to acquire 3D point cloud data of the target object in real time. Through dynamic point cloud analysis, the spatial structure and motion state change information of the target object can be extracted to form a high-precision object flow sequence, suitable for highly complex and space-constrained operation scenarios. For different types of execution devices, the specific parameter configuration of the motion sequence can be dynamically adjusted according to the structural characteristics of the device. For example, in a multi-joint robotic arm system, the motion sequence can be reflected as a combination of joint angles; in a mobile platform system, the motion sequence can be reflected as path nodes and velocity distribution.

[0084] Example description: In the robot intelligent agent decision-making scenario, the industrial assembly robot executes the initial execution equipment action sequence to complete the automated assembly of complex parts. At the same time, it tracks the real-time motion state of the parts through the vision and force feedback system, obtains the real-time object flow sequence, and assists the system to dynamically correct the assembly path and force control strategy, thereby improving assembly efficiency and accuracy.

[0085] In healthcare scenarios, surgical robots execute the initial sequence of actions of the surgical instruments, accurately position and operate them, track the real-time movement of tissues in the surgical area through endoscopic imaging and force sensing devices, obtain real-time object flow sequences, dynamically adjust operating strategies, reduce the risk of accidental injury, and improve the safety and stability of minimally invasive surgery.

[0086] In fintech business scenarios, intelligent operation and maintenance robots execute the initial sequence of equipment actions, complete the automatic detection and replacement of data center server modules, track the real-time changes in the position and structural status of server modules through a high-precision vision system, obtain real-time object flow sequences, achieve precise action docking and dynamic compensation during equipment maintenance, improve the level of unmanned operation and maintenance of data centers, and ensure the high reliability and business continuity of financial information infrastructure.

[0087] Through the above steps, this embodiment can achieve real-time linkage between the operation behavior of the execution device and the state of the target object, avoid target deviation caused by external interference, differences in object attributes or execution errors, obtain highly timely and accurate object stream sequence data, and ensure the stability, flexibility and adaptability of the operation process.

[0088] S60, compare the real-time object stream sequence with the target object stream sequence, and dynamically adjust the execution device action sequence during the execution process based on the comparison deviation.

[0089] In this embodiment, comparing the real-time object stream sequence with the target object stream sequence refers to performing dynamic matching and difference analysis operations on the correspondence between the real-time object stream sequence and the pre-generated target object stream sequence after acquiring the real-time motion state change data of the target object. The target object stream sequence is an ideal motion state change sequence generated based on demonstration data and task description reasoning, reflecting the standard behavior pattern of the target object under the expected task. The real-time object stream sequence is the dynamic state change process of the target object under the influence of the execution device actions during actual operation. The comparison operation analyzes the deviation information of the two types of sequences on a time-node and spatial-dimensional basis through means such as time alignment, spatial transformation correction, and state parameter mapping. The deviations may include position deviation, attitude deviation, velocity deviation, structural deformation deviation, etc. Dynamic adjustment of the execution device action sequence based on comparison deviation refers to calculating error compensation parameters in real time based on the comparison results during operation, and applying the compensation parameters to the update process of the current and subsequent action sequences. The adjustment strategy may include parameter correction, path replanning, speed optimization, and enhanced stability constraints, so as to ensure that the actual state of the target object gradually approaches the ideal state reflected by the target object flow sequence, thereby improving the operation accuracy and task completion quality.

[0090] A dynamic time warping algorithm can be used to align the real-time object stream sequence with the target object stream sequence, eliminating comparison deviations caused by inconsistent execution speeds. A spatial attitude registration algorithm is employed to address spatial correspondence offsets caused by errors in observation angle or initial position. During the comparison process, a real-time state decoding module extracts position, attitude, and structural change features from the sequence to construct a multi-dimensional deviation vector. During dynamic adjustment, a proportional-integral-derivative (PID) control algorithm calculates the motion compensation amount in real-time based on the deviation vector, and this compensation is then added to the corresponding node parameters of the current execution device's motion sequence to form an updated motion sequence. Furthermore, an adaptive path replanning strategy can be used to dynamically optimize the motion sequence structure based on real-time deviation trends, ensuring the continuity and safety of the operation path in complex environments.

[0091] Through the above steps, this embodiment can construct a highly real-time error feedback mechanism for the operation process, realize the dynamic adjustment and precise compensation of the action sequence of the execution device, effectively cope with execution errors, external interference and changes in target object attributes, and ensure the stability, accuracy and adaptability of the operation process.

[0092] This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as robot intelligent agent decision-making, fintech, and healthcare. It discloses an operation control method, apparatus, device, and medium based on object flow sequences, including: acquiring demonstration data and task description representing the task; generating a target object flow sequence independent of the operator's specific form based on the demonstration data and task description; exploring and establishing a mapping relationship between the object flow sequence and the execution device action sequence in a simulation environment through preset action primitives and generating a flow-conditionalized action strategy; inputting the target object flow sequence into the flow-conditionalized action strategy to generate an initial execution device action sequence; tracking the real-time motion state of the target object to obtain a real-time object flow sequence during the execution of the initial execution device action sequence; comparing the real-time object flow sequence with the target object flow sequence and dynamically adjusting the execution device action sequence during the execution process based on the comparison deviation. This invention combines a target object flow sequence independent of the operator's specific form with a flow-conditionalized action strategy, unifying cross-subject video data and simulation exploration data for operation control processes. This avoids reliance on large amounts of real robot data, improves the universality, accuracy, and adaptability of cross-domain operations, and reduces system development and deployment costs.

[0093] In one embodiment, step S10 includes:

[0094] S101, acquires raw demonstration video data through video capture equipment;

[0095] S102, receive the original task description text through the text input interface;

[0096] S103, The original demonstration video data is processed to unify the resolution and standardize the frame rate to generate standardized video data;

[0097] S104, Semantic parsing and key verb extraction are performed on the original task description text to generate a structured task description;

[0098] S105, verify the consistency between the standardized video data and the structured task description using a temporal alignment algorithm;

[0099] S106, Use the verified standardized video data as demonstration data;

[0100] S107, the validated structured task description is used as the task description.

[0101] In this embodiment, acquiring demonstration data and task descriptions for the representation task refers to providing a standardized and structured information foundation for subsequent operation control. This involves comprehensively collecting and processing both video and text data sources to form input information that can be used for object stream generation and task reasoning. The video acquisition device is a device with image data acquisition capabilities, specifically a fixed high-definition camera, a mobile terminal's camera module, or an industrial vision sensing unit. The device parameters support multiple resolutions and frame rates to meet the visual information acquisition needs of different task scenarios. The original demonstration video data is a continuous image sequence acquired through the video acquisition device, containing dynamic visual information about the presenter, the target object, and the environment. The data format can be common video encoding formats such as MP4, AVI, and MOV.

[0102] The text input interface is an interactive channel for receiving task description information. It can be a text input interface based on a natural language processing platform, a speech-to-text system, or a structured command input terminal, supporting the input and parsing of multilingual and multi-format task descriptions. The original task description text is language information about the operation task obtained through the text input interface, including the operation purpose, key actions, object names, environmental conditions, etc.

[0103] Standardizing the resolution and frame rate of the original demonstration video data involves adjusting the spatial and temporal resolution of the original video data using image resampling and temporal interpolation algorithms. This generates standardized video data, ensuring consistency in spatial scale and temporal step between video information from different sources and under different shooting conditions. It eliminates data imbalances caused by differences in acquisition equipment or operating habits. Standardized video data is a visual data sequence conforming to a unified resolution and frame rate standard, facilitating a consistent input for subsequent object stream generation networks.

[0104] Semantic parsing and key verb extraction of the original task description text refers to using a natural language understanding model to analyze the grammatical structure, semantic relationships, and keyword information in the text, focusing on extracting verbs, object entities, and action logic closely related to the task, and generating a structured task description. A structured task description is a collection of task information with clear semantic boundaries and logical structure, oriented towards machine processing, facilitating subsequent object recognition and operational strategy reasoning.

[0105] Verifying the consistency between standardized video data and structured task descriptions using time-series alignment algorithms involves analyzing whether there is a temporal and semantic correspondence between the action sequences in standardized video data and the action information in structured task descriptions, based on time-series matching, semantic mapping, and multimodal alignment techniques. If the comparison results meet the preset consistency criteria, the verification is deemed successful, ensuring the reliability of the input data and the accuracy of the task information.

[0106] Using validated standardized video data as demonstration data means that, while ensuring a high degree of match between the video information and the task description, the standardized video data is confirmed to have the effectiveness and completeness in expressing the operational task. This demonstration data will be used for subsequent object stream sequence generation and operational reasoning. Using validated structured task descriptions as task descriptions means, based on consistency verification results, the structured task descriptions are confirmed to accurately reflect the operational purpose and action logic. These task descriptions will be used to guide subsequent object recognition and strategy planning.

[0107] Through the above steps, this embodiment can efficiently obtain clearly structured and semantically accurate demonstration data and task descriptions, eliminate the impact of differences in data sources on the operational reasoning process, improve the standardization, structuring and reliability of data, provide a unified and stable multimodal information foundation for object stream generation and cross-domain general operation control, and enhance the system's task adaptability and operational effectiveness in complex environments.

[0108] In one embodiment, step S20 above includes:

[0109] S201, Generate a network by combining the demonstration data with the task description input object stream;

[0110] S202, the target object is identified based on the task description through the object stream generation network;

[0111] S203, extract the key point trajectory of the target object from the demonstration data, and exclude the operator's motion information during the extraction process;

[0112] S204, Generate a motion state change sequence of the target object based on the changes in the trajectory of the key points, and use the motion state change sequence as the target object flow sequence.

[0113] In this embodiment, a target object stream sequence independent of the operator's specific form is generated based on demonstration data and task description. This belongs to the action information transformation and reconstruction process for cross-subject operation tasks, aiming to obtain a standardized object action expression form that does not depend on the demonstrator's body structure and movement pattern through structured data extraction and object separation. The demonstration data is a sequence of visual information that has undergone resolution unification and frame rate standardization in the previous steps, possessing clear image details and a stable time step. The task description is a structured set of semantic information, containing key content such as operation objectives, action requirements, and task logic.

[0114] The Object Stream Generation Network is a deep learning model for the joint parsing of video data and task semantics. Its structure may include a visual encoding module, a semantic embedding module, and a multimodal fusion decision module. The network input receives demonstration data and task description information, and based on the fused representation of visual features and task semantics, it identifies target objects in the video. Target objects are manipulated entities in the operational scene, and their specific types can be rigid structures, joint structures, or deformable structures, covering typical everyday items, industrial components, medical devices, etc.

[0115] Key point trajectories of the target object are extracted. Based on the spatial structure and motion characteristics of the target object, image key point detection, pose estimation, and trajectory tracking techniques are used to extract the target object's position changes, pose adjustments, and deformation information over time. Key points can be geometric feature points, structural connection nodes, or functional operation parts on the object's surface. Operator motion information is excluded. Image segmentation, background modeling, or human pose recognition methods are used to remove the demonstrator's body contours, limb movements, and auxiliary equipment movements contained in the demonstration data, ensuring that only the independent motion information of the target object itself is retained, preventing individual differences in operators from affecting the universality of the object's motion expression.

[0116] The motion state change sequence generated based on the changes in keypoint trajectories refers to the expression sequence describing the motion pattern, operation path, and dynamic characteristics of an object in the time dimension, based on the changing trends of the spatial coordinates, attitude angles, or structural parameters of key points within continuous time steps. The sequence content includes position changes, velocity information, acceleration information, and spatial rotation parameters. Ultimately, this motion state change sequence is defined as the target object flow sequence. The target object flow sequence is a standardized, clearly structured, and operator-independent expression of object actions, possessing transferable, interpretable, and cross-subject universal characteristics, serving as the core data input for subsequent control strategy reasoning and operation execution.

[0117] This embodiment utilizes an object stream generation network built based on demonstration data and task description. By combining target object identification, key point trajectory extraction, and operator motion information exclusion, it can extract and reconstruct the independent motion state change sequence of the target object itself without relying on the operator's specific body structure, posture characteristics, and motion patterns. This forms a standardized target object stream sequence, enabling effective conversion of operation action information between different subjects and different execution devices. It solves the migration barrier problem caused by subject differences and inconsistent motion expression, and improves the universality of action expression, data adaptation efficiency, and control strategy reliability of the operating system in cross-subject and cross-environment application scenarios.

[0118] In one embodiment, step S30 above includes:

[0119] S301, executes a preset combination of action primitives in the simulation environment;

[0120] S302, record the object stream sequence generated after executing the preset action primitive combination, and record the execution device action sequence of executing the preset action primitive combination;

[0121] S303, Construct a dataset showing the correspondence between the object stream sequence and the execution device action sequence;

[0122] S304, Train a policy network based on the corresponding relationship dataset, establish a mapping relationship between object flow sequence and execution device action sequence, and obtain the trained policy network;

[0123] S305 uses the trained policy network as the stream-conditionalized action policy.

[0124] In this embodiment, the simulation environment is used to virtually reproduce the state and interaction conditions of the object to be operated, the execution device, and the external space environment. It supports dynamic physical characteristic simulation, operation feedback process simulation, and flexible adjustment of scene parameters. A virtual simulation platform with high-precision dynamic modeling and multi-sensor data output capabilities is often used. Preset action primitives refer to the set of basic operation actions designed for the target task scene. Each primitive contains a separately defined operation posture, motion parameters, and control logic, such as basic operation modules like clamping, translation, rotation, and release. Preset action primitive combinations are formed by combining different primitives in a time sequence to create diverse operation action chains.

[0125] The object flow sequence reflects the state change process of the target object after the execution of the preset action primitive combination in the simulation environment. Specifically, it is a set of dynamic changes in parameters such as object position, attitude, and deformation over time. The execution device action sequence refers to the control command sequence or motion trajectory data of the execution device in the simulation environment during the execution of the preset action primitive combination. It usually includes parameters such as multi-joint angles, end effector pose, or velocity.

[0126] The correspondence dataset forms a one-to-one mapping sample by synchronously recording object stream sequences and execution device action sequences, possessing time synchronization, operational consistency, and cross-sample scalability. The policy network, as a parameterized mapping structure, combined with a deep neural network architecture or a temporal decision model, receives object stream sequences as input and outputs corresponding execution device action sequences. During training, it continuously optimizes the mapping performance by minimizing the error between predicted and actual actions, ultimately forming a trained policy network with generalization and policy transfer capabilities, used as a stream-conditional action policy for object stream input.

[0127] This embodiment constructs a dataset containing the correspondence between dynamic changes in object states and control commands from executing devices. Combined with a pre-defined action primitive exploration process based on a simulation environment, it efficiently learns the mapping relationship between object flow sequences and execution device action sequences, avoiding the high cost associated with repeated trials with real robots and large-scale physical data collection. The flow-conditional action policy formed after the policy network training possesses the ability to generate actions driven by the input object flow. This supports flexible adaptation of the operation control system to different physical environments and task requirements, improving the versatility and autonomous decision-making level of cross-subject and cross-device operation control, reducing the impact of environmental differences during the migration from simulation to reality, and enhancing the overall system's stability and applicability.

[0128] In one embodiment, step S40 above includes:

[0129] S401, input the target object stream sequence into the feature encoding layer of the stream conditional action strategy;

[0130] S402, extract the spatiotemporal feature representation of the target object stream sequence through the feature encoding layer;

[0131] S403, input the spatiotemporal feature representation into the temporal decision layer of the flow conditional action strategy;

[0132] S404, Multi-step action planning is performed through the timing decision layer to generate a sequence of action vectors for the execution device;

[0133] S405, Perform kinematic constraint processing on the action vector sequence of the execution device;

[0134] S406, convert the kinematically constrained sequence of execution device motion vectors into the initial sequence of execution device motions.

[0135] In this embodiment, the target object flow sequence is a continuous data set characterizing the dynamic changes of the target object over time. It typically includes the object's position, posture, motion trend, and state change information in three-dimensional space, exhibiting continuity, temporal sequence, and structured characteristics. This sequence is input into the flow conditional action strategy, aiming to infer matching execution device control actions based on the object's state change trend.

[0136] The streaming conditional action strategy has a hierarchical structure. The feature encoding layer is responsible for compressing information and representing high-dimensional features of the input target object stream sequence. The feature encoding layer can adopt a spatiotemporal fusion structure, combining one-dimensional convolution, temporal recurrent network or self-attention mechanism to model the spatial state information and temporal dynamic features in the target object stream sequence at the same time, generating a stable and expressive spatiotemporal feature representation. This feature representation preserves the overall motion law of the target object, local key changes and dynamic dependencies within the sequence.

[0137] The generated spatiotemporal feature representations are further input into the temporal decision layer of the flow-conditional action strategy. The temporal decision layer is used to infer multi-step temporal structured action outputs based on the spatiotemporal features, possessing multi-step prediction and action sequence generation capabilities. The temporal decision layer can employ sequence-modeling-based network structures, such as recurrent neural networks, gated recurrent units, long short-term memory networks, or Transformer-based temporal coding modules. It supports multi-step recursive reasoning, enabling continuous, multi-step action planning for target object flow sequences, and outputting an action vector sequence for the execution device. This sequence can contain sets of control parameters corresponding to multiple action moments.

[0138] To ensure the physical executability and system stability of the generated action sequence, kinematic constraints need to be applied to the action vector sequence of the execution device. Specifically, this may include limiting the range of motion of the execution device's main structure, checking joint limits, smoothing motion continuity, constraining velocity and acceleration, and collision avoidance analysis. Constraint optimization algorithms are then used to adjust the action outputs that do not meet the constraints, ensuring that the generated actions conform to the structural limitations and motion capabilities of the actual hardware platform.

[0139] After kinematic constraint processing, the motion vector sequence is finally converted into the initial execution device motion sequence. The initial execution device motion sequence has a specific executable format, such as joint control commands, end effector trajectory, operation parameters or control signal sequence, which is directly used to drive the execution device to carry out subsequent operation control process.

[0140] This embodiment, through high-dimensional feature encoding and temporal decision reasoning of the target object flow sequence, combined with kinematic constraint mechanisms, can quickly generate an initial execution device action sequence that meets the physical structure requirements, motion parameter constraints, and operational task requirements of the actual execution device based on the dynamic change process of the target object. This avoids the high-cost problem of relying on manual experience to design actions or repeated physical trial and error verification in traditional control processes, and improves the efficiency, stability, and adaptability of system action generation, which is conducive to flexible deployment and autonomous control of different devices and different task scenarios.

[0141] In one embodiment, step S50 above includes:

[0142] S501, control the execution device to execute the initial execution device action sequence;

[0143] S502, during the execution of the initial execution device action sequence, real-time motion state data of the target object is collected through a multi-sensor fusion system;

[0144] S503, Process the real-time motion state data to identify key points of the target object;

[0145] S504, track the continuous motion trajectory of the key points;

[0146] S505, Generate a real-time motion state change sequence of the target object based on the continuous motion trajectory;

[0147] S506, the real-time motion state change sequence is encoded into a real-time object stream sequence.

[0148] In this embodiment, the initial execution device action sequence is a set of control instructions generated by combining the target object flow sequence and flow-conditional action strategy reasoning to guide the operation of the execution device. It typically includes phased motion parameters, joint control signals, or end-effector path information. Controlling the execution device to execute the initial execution device action sequence refers to the execution device control system parsing and issuing the corresponding action instructions at each moment, driving the execution device to complete specific spatial movements, operational adjustments, or interactive behaviors. This process involves real-time hardware control, system feedback response, and action status monitoring.

[0149] During the execution of the initial execution device action sequence, in order to obtain the actual dynamic performance of the target object, it is necessary to use a multi-sensor fusion system to carry out external environment and object state perception. The multi-sensor fusion system includes vision sensors, depth sensors, inertial measurement units, force and tactile sensors or other auxiliary positioning devices. The system obtains high-precision, continuous real-time motion state data of the target object in the actual environment through synchronous acquisition, information fusion and dynamic calibration of data from different types of sensors. The motion state data can include multi-dimensional information such as position, attitude, velocity, acceleration, and force changes.

[0150] Based on real-time motion state data, the system identifies key points of the target object through feature recognition and data parsing. Key points are spatial feature locations that characterize the structural changes, motion trajectory, or operational effects of the target object. They are usually obtained through image processing, 3D reconstruction, depth estimation, edge detection, semantic segmentation, or key point detection networks to ensure that the core information reflecting the target object itself and its interaction with the environment is captured.

[0151] Tracking the continuous motion trajectory of key points of a target object refers to using time series analysis, target tracking algorithms, and trajectory modeling methods to monitor the spatial path and dynamic state of key points over time in real time, generating a complete continuous motion trajectory. This trajectory can reflect the actual motion pattern, state deviation, or abnormal changes of the target object during the operation process.

[0152] Based on the continuous motion trajectory, the system further extracts the real-time motion state change sequence of the target object. The motion state change sequence, combined with information such as position, attitude, velocity, and acceleration, reflects the dynamic process of the object's evolution over time and retains the full-process data characteristics of the operation result.

[0153] The real-time motion state change sequence is encoded into a real-time object stream sequence. The encoding process uses structured data representation, sequence normalization and information compression to generate standardized sequence data that is consistent with the target object stream sequence format and can be used for comparison analysis or strategy optimization. This ensures that the real-time object stream sequence has the same structure, scale and semantics as the initial expected object stream sequence, which facilitates subsequent error analysis and dynamic adjustment.

[0154] This embodiment controls the execution device to execute an initial action sequence and combines it with a multi-sensor fusion system to track the target object's state in real time. It can dynamically acquire the actual motion performance of the target object without relying on external human intervention. Through continuous key point tracking and state change modeling, it generates a structured and standardized real-time object flow sequence, providing a stable and reliable data foundation for precise monitoring, error judgment, and action adjustment of the operation process. This improves the overall system's operational accuracy, environmental adaptability, and autonomous adjustment capabilities.

[0155] In one embodiment, step S60 above includes:

[0156] S601, Perform dynamic time warping on the real-time object stream sequence and the target object stream sequence to generate a time-aligned real-time object stream sequence and a time-aligned target object stream sequence;

[0157] S602, determine the motion state difference value between the time-aligned real-time object stream sequence and the time-aligned target object stream sequence at the same timestamp;

[0158] S603, aggregate the motion state difference values ​​of all timestamps to generate a deviation vector;

[0159] S604, the motion compensation amount is determined by the proportional-integral-derivative control module based on the deviation vector;

[0160] S605, the action compensation amount is superimposed on the corresponding action node of the current execution device action sequence to generate an updated execution device action sequence and apply it to the execution process.

[0161] In this embodiment, the real-time object stream sequence and the target object stream sequence respectively reflect the actual observed changes in the target object's motion state during the operation and the ideal state changes generated before the operation based on the task settings. Both are expressed through a standardized coding structure and have the same temporal dimension, spatial structure, and state parameter dimension. The process of comparing the two types of object stream sequences requires first aligning the two sets of sequences in the temporal dimension through dynamic time warping. Dynamic time warping employs an algorithm framework based on distance metrics and optimal path matching to automatically adjust the mapping relationship between the time series, resolving the inconsistency in time scales caused by speed fluctuations, rhythm changes, or external interference during execution, and generating a time-aligned real-time object stream sequence and a target object stream sequence with the same timestamp distribution.

[0162] After completing time alignment, for each timestamp position, the system calculates the motion state difference value at the corresponding timestamp based on the motion state parameters contained in the object stream sequence. The motion state difference value reflects the dynamic offset between the actual operation effect and the expected operation target, and often includes multi-dimensional error information such as position error, attitude deviation, speed difference, and acceleration change.

[0163] By aggregating the motion state differences across all timestamp locations, the system generates a deviation vector. Mathematically, the deviation vector represents the overall expression of multi-dimensional error signals, enabling comprehensive quantification of global deviation trends, local error concentration areas, or dynamic error change patterns during the current operation. The deviation vector provides a quantitative and structured reference for subsequent action compensation.

[0164] Based on the deviation vector, the system uses a proportional-integral-derivative (PID) control module to calculate the motion compensation amount. The PID control module integrates proportional, integral, and derivative components. By analyzing the amplitude, rate of change, and cumulative deviation trend of the deviation vector in real time, it dynamically generates adaptive motion compensation amounts. The motion compensation amounts are usually expressed as control parameters such as angle adjustment, position fine-tuning, or speed correction, which are used to optimize the real-time output state of the execution device's motion sequence.

[0165] The system adds the action compensation amount to the corresponding action node of the current execution device action sequence. By adjusting the operation instructions node by node, the system locally optimizes the key action parameters in the current execution process, ensuring that the updated execution device action sequence has dynamic compensation capability on the basis of overall coherence. This is then applied to the execution process to form a real-time closed-loop adjustment mechanism.

[0166] Example Description: In the field of robotic agent decision-making, for a multi-functional collaborative robotic arm with autonomous operation capabilities, the overall system is used to perform flexible and high-precision assembly and operation tasks in a production workshop. First, the system acquires raw demonstration video data of a human operator at the assembly workstation through a configured video acquisition device, and simultaneously receives assembly task description text input by the operator through a text input interface. The system performs resolution and frame rate standardization processing on the raw demonstration video data to generate video data with a unified format. Simultaneously, it performs semantic parsing and key verb extraction on the task description text to form structured task description data. Using a temporal alignment algorithm, the standardized video data and structured task description are verified for consistency. After ensuring that the video content matches the task semantics, the standardized video data is used as demonstration data, and the structured task description is used as the task description, and then input into the downstream process.

[0167] Subsequently, the system generates a target object flow sequence independent of the operator's specific form based on the demonstration data and task description. Specifically, it parses the demonstration data through an object flow generation network to identify target objects such as screws, parts, and tools involved in the operation. During the extraction of keypoint trajectories for these target objects, it eliminates interference from information such as the operator's hands and body, retaining only the spatial trajectory and posture changes related to the target object's own movement. Combining these keypoint trajectory changes, the system generates a sequence of motion state changes for the target objects, forming the final target object flow sequence used to guide robot operations.

[0168] In the simulation environment, the system loads a virtual robot model and executes various basic operations composed of preset action primitives, such as grasping, transporting, and rotating. It records the target object flow sequence and the corresponding robot execution device action sequence during the operation, and constructs a mapping dataset between the two. Based on this dataset, a policy network is trained to learn the corresponding mapping logic from the target object flow sequence to the execution device action sequence. Finally, the trained policy network is used as a flow-conditional action policy for use in actual operation.

[0169] In the formal execution phase, the system inputs the generated target object flow sequence into the flow-conditionalized action policy. The policy network first extracts the spatiotemporal feature representation of the target object flow sequence through a feature encoding layer, and then inputs this feature into the temporal decision layer for multi-step action planning, generating a sequence of action vectors for the robot's execution device. The system performs kinematic constraint processing on this sequence of action vectors to ensure that the output motion scheme conforms to the structural constraints and safety specifications of the robot body, and finally converts it into the initial sequence of action for the execution device.

[0170] The system controls the robot to perform actual operations based on the initial sequence of actions of the execution device. At the same time, it collects motion state data of the target object in real time through a multi-sensor fusion system. Combining visual, force and position sensing information, it identifies key points of the target object and tracks continuous motion trajectory. Based on the trajectory changes, it generates a real-time motion state change sequence of the target object and encodes it into a real-time object stream sequence.

[0171] The system dynamically warps the real-time object stream sequence with the previously generated target object stream sequence to achieve consistent alignment in the time dimension. It calculates the motion state difference between the two sequences at the same timestamp and aggregates the difference data from all timestamps to form a deviation vector. Based on this deviation vector, the system applies a proportional-integral-derivative (PID) control algorithm to calculate motion compensation in real time. This compensation is then added to the corresponding motion node of the currently executing device's motion sequence to generate an updated motion sequence. This dynamically optimizes the robot's operation, ensuring that the robot can adapt to error changes during operation and improving the overall stability, accuracy, and autonomous decision-making capability of task execution.

[0172] In healthcare scenarios, for the operational learning of mobile nursing robots within hospitals, task learning data is constructed by acquiring demonstration videos of nurses performing bed-making tasks and corresponding task descriptions. The entire process of nurses making beds is recorded using video capture devices, while text input interfaces record nurses' written descriptions of the operation sequence, action objectives, and precautions. This video data undergoes resolution and frame rate standardization to form standardized video data. The task descriptions, after semantic parsing and key verb extraction, form structured operational step descriptions. Timing alignment is used to verify whether there are any missing steps or time discrepancies between the standardized video data and the task descriptions, ensuring consistency between the video action sequence and the task logic description.

[0173] The network generates an object stream by inputting validated standardized video data and structured task descriptions. Based on the task descriptions, it extracts target objects during the bed-tidying process, such as sheets, blankets, and bed rails, while excluding the movement information of nursing staff and retaining only the movement trajectories of the target objects. By tracking the changes in the movement state of these target objects during the demonstration, it generates a target object stream sequence for the bed-tidying task.

[0174] In a simulation environment, by combining preset action primitives such as adjusting bed sheets, placing pillows, and raising or lowering bed rails, the motion changes of the target object are repeatedly explored, and the object flow sequence and robot action sequence generated during the simulation are recorded. By constructing a mapping dataset between the object flow sequence and the action sequence, a policy network is trained to associate the motion state of the target object with the robot's execution actions, and to generate a flow-conditional action policy adapted to the bed-tidying task.

[0175] The target object flow sequence is input into the flow-conditionalized action strategy. Spatiotemporal features of the target object flow sequence are extracted through feature encoding and then input into the temporal decision layer to execute multi-step action planning, resulting in the robot's action vector sequence. Kinematic constraints are applied to the generated action vector sequence to adjust action nodes that exceed the robot's joint range of motion or pose safety risks. The processed action vector sequence is then converted into an initial execution device action sequence that the robot can directly execute.

[0176] The robot is controlled to perform bed-making actions, such as laying out sheets and adjusting bed rails, according to the initial sequence of actions from the execution equipment. During the execution, the robot uses vision sensors, force sensors, and depth sensors to collect real-time data on the motion status of target objects such as sheets, blankets, and bed rails. Key point detection technology is used to continuously track the continuous motion trajectories of these target objects, and a real-time motion state change sequence is generated based on the tracking results. This sequence is then encoded into a real-time object stream sequence.

[0177] The real-time object stream sequence is dynamically time-warped with the expected target object stream sequence to ensure alignment in execution rhythm and action timing. The motion states of the two aligned sequences at the same timestamp are compared, motion state differences are calculated, and aggregated to generate a deviation vector. Based on the deviation vector, a proportional-integral-derivative (PID) control algorithm is used to dynamically calculate motion compensation amounts. These compensation amounts are then superimposed on the current robot's motion nodes, adjusting the motion trajectory in real time and generating an updated sequence of execution device actions. This sequence is applied to the robot's execution process, ensuring the bed-tidying task continuously follows the target object's motion path, effectively improving tidying efficiency and execution safety.

[0178] In the fintech field, to address the cross-domain task control needs of intelligent agents for the maintenance of bank self-service equipment, the system acquires video data and accompanying text instructions from experienced equipment maintenance engineers performing tasks such as self-service terminal repair and cash module replacement in a simulated environment. This forms the data foundation for representing complex maintenance tasks. The video data undergoes standardized processing to unify resolution and frame rate, eliminating data differences between different recording environments. The text data is semantically parsed and key verbs are extracted to structurally express operational goals and action logic. Combined with temporal alignment verification, this ensures consistency between video and text information in terms of task steps and timing.

[0179] The processed demonstration data and task description are input into the object stream generation network. Based on the task description, the network identifies key target objects in the self-service equipment, such as equipment door locks, banknote boxes, and coin slots. It excludes the operator's body movement information and extracts only the key point trajectories of the target objects, accurately reflecting the actual operation process of the physical objects and generating a target object stream sequence independent of the operator's form.

[0180] In the simulation environment of self-service equipment, action primitives for different models and operation modules are configured, such as opening the equipment door, taking out the banknote box, and resetting the hardware module. The system explores by combining different action primitives, records the object flow sequence of the target object in the simulation environment and the action sequence of the corresponding execution device, builds a mapping relationship dataset, trains the policy network, realizes the mapping relationship between the target object flow sequence and the robot action sequence, and forms a flow conditional action policy with task adaptability.

[0181] The target object stream sequence is input into the conditional action strategy. The feature encoding layer extracts its spatiotemporal features, and the temporal decision layer performs multi-step action planning to generate the action vector sequence of the execution device. Combined with the physical structure and safety specifications of the maintenance robot, kinematic constraint processing is performed to generate the initial execution device action sequence suitable for the self-service equipment environment.

[0182] The control and maintenance robot completes equipment operation based on the initial execution sequence of equipment actions. During the execution process, it uses multi-sensor fusion technology to collect motion state data of target objects such as equipment doors, banknote boxes, and coin slots in real time, extract key points and track continuous motion trajectories, generate a real-time motion state change sequence of the target objects, and encode it as a real-time object stream sequence.

[0183] The real-time object flow sequence is dynamically time-aligned with the target object flow sequence. The differences in motion state are compared point by point in time. The difference values ​​are aggregated to generate a deviation vector. The proportional-integral-derivative control algorithm is applied to calculate the motion compensation amount in real time. The compensation amount is superimposed on the current execution device motion sequence to dynamically adjust the robot's operation process, ensuring that the robot can stably, safely and accurately complete cross-domain operation and maintenance tasks in complex equipment environments.

[0184] This embodiment constructs a complete adjustment process of time alignment, error quantification and dynamic compensation. Combined with dynamic time warping, deviation vector generation and proportional-integral-derivative control strategies, it can perceive the state deviation of the target object in real time during operation and accurately adjust the action sequence of the execution device. This improves the real-time performance, accuracy and stability of the overall operation, effectively bridges the gap between expectations and reality in task execution, enhances the system's autonomous adjustment capability and environmental adaptability, and reduces the risk of operation failure due to error accumulation.

[0185] In one embodiment, an operation control device based on an object stream sequence is provided, which corresponds one-to-one with the operation control method based on the object stream sequence described in the above embodiments. (Refer to...) Figure 3 , Figure 3This is a schematic diagram of the functional modules of a preferred embodiment of the operation control device based on object stream sequences of the present invention. The module includes a data acquisition module 10, an object stream generation module 20, a mapping strategy training module 30, an action sequence generation module 40, a real-time tracking acquisition module 50, and an action adjustment control module 60. Detailed descriptions of each functional module are as follows:

[0186] The data acquisition module 10 is used to acquire demonstration data and task descriptions representing the task;

[0187] Object stream generation module 20 is used to generate a target object stream sequence that is independent of the specific form of the operator based on the demonstration data and task description;

[0188] The mapping strategy training module 30 is used to explore and establish the mapping relationship between the object flow sequence and the execution device action sequence in the simulation environment through preset action primitives, and generate flow conditional action strategies.

[0189] Action sequence generation module 40 is used to input the target object stream sequence into the stream conditional action strategy to generate an initial execution device action sequence;

[0190] The real-time tracking and acquisition module 50 is used to execute the initial execution device action sequence and, during the execution process, acquire a real-time object stream sequence by tracking the real-time motion state of the target object.

[0191] The action adjustment control module 60 is used to compare the real-time object stream sequence with the target object stream sequence, and dynamically adjust the action sequence of the execution device during the execution process based on the deviation of the comparison.

[0192] In one embodiment, the data acquisition module 10 is specifically used for:

[0193] Acquire raw demonstration video data using video capture equipment;

[0194] Receive the original task description text via the text input interface;

[0195] The original demonstration video data is processed to unify resolution and standardize frame rate to generate standardized video data;

[0196] Semantic parsing and key verb extraction are performed on the original task description text to generate a structured task description;

[0197] The consistency between the standardized video data and the structured task description was verified using a temporal alignment algorithm.

[0198] Use the validated standardized video data as demonstration data;

[0199] Use the validated structured task description as the task description.

[0200] In one embodiment, the object stream generation module 20 is specifically used for:

[0201] The demonstration data and the task description input object stream are used to generate a network;

[0202] The object stream generation network identifies the target object based on the task description.

[0203] Extract the key point trajectory of the target object from the demonstration data, and exclude operator motion information during the extraction process;

[0204] Based on the changes in the trajectory of the key points, a sequence of motion state changes of the target object is generated, and the sequence of motion state changes is used as the target object flow sequence.

[0205] In one embodiment, the mapping policy training module 30 is specifically used for:

[0206] Execute a pre-defined combination of action primitives in a simulation environment;

[0207] Record the object stream sequence generated after executing the preset action primitive combination, and record the execution device action sequence of executing the preset action primitive combination;

[0208] Construct a dataset showing the correspondence between the object stream sequence and the execution device action sequence;

[0209] A policy network is trained based on the aforementioned correspondence dataset, and a mapping relationship is established between object flow sequences and execution device action sequences to obtain the trained policy network.

[0210] The trained policy network is used as the flow-conditional action policy.

[0211] In one embodiment, the action sequence generation module 40 is specifically used for:

[0212] The target object stream sequence is input into the feature encoding layer of the stream conditional action strategy;

[0213] The spatiotemporal feature representation of the target object stream sequence is extracted through the feature encoding layer;

[0214] The spatiotemporal feature representation is input into the temporal decision layer of the stream conditional action strategy;

[0215] Multi-step action planning is performed through the time-series decision layer to generate a sequence of action vectors for the execution device.

[0216] The motion vector sequence of the execution device is subjected to kinematic constraint processing;

[0217] The sequence of motion vectors of the actuator after kinematic constraint processing is converted into the initial sequence of motion of the actuator.

[0218] In one embodiment, the real-time tracking acquisition module 50 is specifically used for:

[0219] Control the execution device to execute the initial execution device action sequence;

[0220] During the execution of the initial execution device action sequence, real-time motion state data of the target object is collected through a multi-sensor fusion system;

[0221] Process the real-time motion state data to identify key points of the target object;

[0222] Track the continuous motion trajectory of the key points;

[0223] Generate a real-time motion state change sequence of the target object based on the continuous motion trajectory;

[0224] The real-time motion state change sequence is encoded into a real-time object stream sequence.

[0225] In one embodiment, the motion adjustment control module 60 is specifically used for:

[0226] Dynamic time warping is performed on the real-time object stream sequence and the target object stream sequence to generate a time-aligned real-time object stream sequence and a time-aligned target object stream sequence.

[0227] Determine the difference in motion state between the time-aligned real-time object stream sequence and the time-aligned target object stream sequence at the same timestamp;

[0228] Aggregate the motion state differences across all timestamps to generate a deviation vector;

[0229] The motion compensation amount is determined by the proportional-integral-derivative control module based on the deviation vector.

[0230] The action compensation amount is superimposed on the corresponding action node of the current execution device action sequence to generate an updated execution device action sequence and apply it to the execution process.

[0231] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used for communication with external user terminals via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a server-side operation control method based on object stream sequences.

[0232] In one embodiment, a computer device is provided, which may be a user terminal, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides determination and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface is used to communicate with an external server via a network connection. When executed by the processor, the computer program implements the user-side functions or steps of an operation control method based on an object stream sequence.

[0233] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps:

[0234] Obtain demonstration data and task description for the representation task;

[0235] Based on the demonstration data and task description, a target object stream sequence independent of the operator's specific form is generated;

[0236] In the simulation environment, by exploring preset action primitives, a mapping relationship between object flow sequences and execution device action sequences is established, and flow conditional action strategies are generated.

[0237] The target object stream sequence is input into the stream conditionalization action strategy to generate an initial execution device action sequence;

[0238] The initial execution device action sequence is executed, and during the execution process, a real-time object stream sequence is obtained by tracking the real-time motion state of the target object;

[0239] The real-time object stream sequence is compared with the target object stream sequence, and the execution device action sequence during the execution process is dynamically adjusted based on the deviation of the comparison.

[0240] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:

[0241] Obtain demonstration data and task description for the representation task;

[0242] Based on the demonstration data and task description, a target object stream sequence independent of the operator's specific form is generated;

[0243] In the simulation environment, by exploring preset action primitives, a mapping relationship between object flow sequences and execution device action sequences is established, and flow conditional action strategies are generated.

[0244] The target object stream sequence is input into the stream conditionalization action strategy to generate an initial execution device action sequence;

[0245] The initial execution device action sequence is executed, and during the execution process, a real-time object stream sequence is obtained by tracking the real-time motion state of the target object;

[0246] The real-time object stream sequence is compared with the target object stream sequence, and the execution device action sequence during the execution process is dynamically adjusted based on the deviation of the comparison.

[0247] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and user side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0248] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0249] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0250] It should be noted that if any software tools or components not belonging to this company appear in the embodiments of this application, they are merely illustrative examples and do not represent actual use. The embodiments described above are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. An operation control method based on object stream sequences, characterized in that, Includes the following steps: Obtain demonstration data and task description for the representation task; Based on the demonstration data and task description, a target object stream sequence independent of the operator's specific form is generated; In the simulation environment, by exploring preset action primitives, a mapping relationship between object flow sequences and execution device action sequences is established, and flow conditional action strategies are generated. The target object stream sequence is input into the stream conditionalization action strategy to generate an initial execution device action sequence; The initial execution device action sequence is executed, and during the execution process, a real-time object stream sequence is obtained by tracking the real-time motion state of the target object; The real-time object stream sequence is compared with the target object stream sequence, and the execution device action sequence during the execution process is dynamically adjusted based on the deviation of the comparison.

2. The operation control method based on object stream sequence as described in claim 1, characterized in that, Obtain demonstration data and task description for the representation task, including: Acquire raw demonstration video data using video capture equipment; Receive the original task description text via the text input interface; The original demonstration video data is processed to unify resolution and standardize frame rate to generate standardized video data; Semantic parsing and key verb extraction are performed on the original task description text to generate a structured task description; The consistency between the standardized video data and the structured task description was verified using a temporal alignment algorithm. Use the validated standardized video data as demonstration data; Use the validated structured task description as the task description.

3. The operation control method based on object stream sequence as described in claim 1, characterized in that, Based on the demonstration data and task description, a target object flow sequence independent of the operator's specific form is generated, including: The demonstration data and the task description input object stream are used to generate a network; The object stream generation network identifies the target object based on the task description. Extract the key point trajectory of the target object from the demonstration data, and exclude operator motion information during the extraction process; Based on the changes in the trajectory of the key points, a sequence of motion state changes of the target object is generated, and the sequence of motion state changes is used as the target object flow sequence.

4. The operation control method based on object stream sequence as described in claim 1, characterized in that, In the simulation environment, by exploring preset action primitives, a mapping relationship between object flow sequences and execution device action sequences is established, generating flow-conditional action strategies, including: Execute a pre-defined combination of action primitives in a simulation environment; Record the object stream sequence generated after executing the preset action primitive combination, and record the execution device action sequence of executing the preset action primitive combination; Construct a dataset showing the correspondence between the object stream sequence and the execution device action sequence; A policy network is trained based on the aforementioned correspondence dataset, and a mapping relationship is established between object flow sequences and execution device action sequences to obtain the trained policy network. The trained policy network is used as the flow-conditional action policy.

5. The operation control method based on object stream sequence as described in claim 1, characterized in that, The target object stream sequence is input into the stream conditionalization action strategy to generate an initial execution device action sequence, including: The target object stream sequence is input into the feature encoding layer of the stream conditional action strategy; The spatiotemporal feature representation of the target object stream sequence is extracted through the feature encoding layer; The spatiotemporal feature representation is input into the temporal decision layer of the stream conditional action strategy; Multi-step action planning is performed through the time-series decision layer to generate a sequence of action vectors for the execution device. The motion vector sequence of the execution device is subjected to kinematic constraint processing; The sequence of motion vectors of the actuator after kinematic constraint processing is converted into the initial sequence of motion of the actuator.

6. The operation control method based on object stream sequence as described in claim 1, characterized in that, Execute the initial execution device action sequence, and during the execution process, obtain a real-time object stream sequence by tracking the real-time motion state of the target object, including: Control the execution device to execute the initial execution device action sequence; During the execution of the initial execution device action sequence, real-time motion state data of the target object is collected through a multi-sensor fusion system; Process the real-time motion state data to identify key points of the target object; Track the continuous motion trajectory of the key points; A real-time motion state change sequence of the target object is generated based on the continuous motion trajectory; The real-time motion state change sequence is encoded into a real-time object stream sequence.

7. The operation control method based on object stream sequence as described in claim 1, characterized in that, The process of comparing the real-time object stream sequence with the target object stream sequence and dynamically adjusting the execution device action sequence during execution based on the comparison deviation includes: Dynamic time warping is performed on the real-time object stream sequence and the target object stream sequence to generate a time-aligned real-time object stream sequence and a time-aligned target object stream sequence. Determine the difference in motion state between the time-aligned real-time object stream sequence and the time-aligned target object stream sequence at the same timestamp; Aggregate the motion state differences across all timestamps to generate a deviation vector; The motion compensation amount is determined by the proportional-integral-derivative control module based on the deviation vector. The action compensation amount is superimposed on the corresponding action node of the current execution device action sequence to generate an updated execution device action sequence and apply it to the execution process.

8. An operation control device based on an object stream sequence, characterized in that, The operation control device based on object stream sequence includes: The data acquisition module is used to acquire demonstration data and task descriptions for representing the task; The object stream generation module is used to generate a sequence of target object streams that are independent of the specific form of the operator based on the demonstration data and task description. The mapping strategy training module is used to explore and establish the mapping relationship between the object flow sequence and the execution device action sequence in the simulation environment through preset action primitives, and generate flow conditional action strategies. An action sequence generation module is used to input the target object stream sequence into the stream conditional action strategy to generate an initial execution device action sequence; The real-time tracking and acquisition module is used to execute the initial execution device action sequence and, during the execution process, acquire a real-time object stream sequence by tracking the real-time motion state of the target object; The action adjustment control module is used to compare the real-time object stream sequence with the target object stream sequence, and dynamically adjust the action sequence of the execution device during the execution process based on the deviation of the comparison.

9. A computer device, characterized in that, The computer device includes a memory, a processor, and an object stream sequence-based operation control program stored in the memory and executable on the processor. When executed by the processor, the object stream sequence-based operation control program implements the steps of the object stream sequence-based operation control method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The storage medium stores an operation control program based on an object stream sequence, which, when executed by a processor, implements the steps of the operation control method based on an object stream sequence as described in any one of claims 1-7.

Citation Information

Cited By

  • Robot control reconstruction method and system based on PLC replacement

    CN121625172A

  • Robot control reconstruction method and system based on plc replacement

    CN121625172B

  • Industrial enterprise typical fire scene construction system based on multi-source data fusion

    CN122021077A