Action sequence execution and re-planning method and device, equipment and medium

By generating a closed-loop mechanism of task model files, action sequences and feedback information, the inflexibility of task execution in dynamic environments in existing technologies is solved, efficient and autonomous replanning and task optimization are achieved, and the adaptability and reliability of task execution are improved.

CN120806352APending Publication Date: 2025-10-17PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510881714.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing technologies lack a standardized closed-loop system for task models and motion skill mapping for dynamic environments, and are unable to achieve efficient autonomous replanning and task execution optimization based on real-time feedback information.

Method used

Obtain task definition information to generate a task model file, receive current environment status information to generate an action sequence, schedule execution of skills according to the action-skill mapping relationship, and update the action sequence when feedback information is collected during or after execution to meet re-planning conditions.

Benefits of technology

By combining the task model with the real-time environment status to dynamically generate action sequences, the adaptability and reliability of task execution in complex environments are improved, forming a complete closed loop of task modeling, action mapping, execution control and feedback re-planning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120806352A_ABST
    Figure CN120806352A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, can be applied to business scenes of robot agent decision making, financial science and technology, medical treatment and health and the like, and discloses an action sequence execution and replanning method, device, equipment and medium. Generating an action sequence based on the task model file and the current environment state information, mapping the action sequence into an executable skill, scheduling and executing the executable skill to control an execution unit to operate, collecting an execution state and environment perception information to form feedback information, and when the feedback information meets a re-planning triggering condition, performing re-planning. And regenerating the action sequence based on the updated environment state information. According to the method, the environment state feedback is dynamically associated with the action sequence, and the task model file is combined to update the plan in real time, so that the execution flexibility and the environment adaptability are improved, and the task completion effect in a complex scene is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and in particular to a motion sequence execution and re-planning method, device, equipment and storage medium. BACKGROUND

[0002] Existing task control and decision systems generally rely on static and deterministic task models and rule configurations, which can implement basic motion sequence planning and execution in specific environments, but in the face of complexity and dynamic changes in actual application scenarios, the systems generally have limited model expression ability, complex skill mapping configuration, lack of efficient feedback mechanism and real-time re-planning capability, etc., resulting in insufficient autonomous decision flexibility, stability and reliability in the task execution process. These technical shortcomings are particularly prominent in multi-domain applications, severely restricting the wide deployment and efficient operation of intelligent decision systems in complex scenarios.

[0003] In the field of robot agent decision-making business, mainstream systems usually adopt static task definition and single planning mechanism, lacking high-frequency closed-loop perception and real-time task adjustment capability for dynamic environmental changes, execution process feedback and sudden conditions. Existing technologies cannot effectively handle problems such as execution unit state changes, environmental uncertainty and sudden obstacles, and the system lacks complete task model dynamic updating, action skill mapping standardization and autonomous re-planning mechanism, resulting in delayed task response, unstable execution and limited overall autonomous decision-making level of robot agents in complex environments.

[0004] In the field of medical and health business, intelligent devices and assistive robots are widely used in diagnosis and treatment support, rehabilitation nursing and remote collaboration scenarios, but existing systems still rely on static models and manual configuration in terms of task control and environmental perception, making it difficult to respond to patient state changes, environmental dynamic adjustments or operation abnormalities in a timely manner. Lack of dynamic perception, feedback information fusion and task adaptive adjustment mechanism for real-time data in medical environments reduces the ability of medical assistance systems in handling sudden situations, intelligent collaboration and high-reliability task execution, limiting their intelligent decision-making level and application effect in complex medical environments.

[0005] In the field of financial technology business, intelligent decision systems are gradually applied to intelligent outlets, business handling and risk monitoring scenarios, but existing systems generally lack real-time task planning based on dynamic environmental information, feedback information-driven task adjustment and multi-source data fusion mechanism, resulting in a lack of efficient autonomous response and dynamic task optimization capability when the system faces changes in customer behavior, business process abnormalities or environmental emergencies. At the same time, there is a lack of flexible and efficient skill mapping mechanism between task models and system underlying operations, increasing the system configuration and maintenance cost, and reducing the stability, flexibility and autonomous decision-making level of intelligent systems in financial business scenarios. SUMMARY

[0006] The main purpose of the present application is to provide a motion sequence execution and re-planning method, device, equipment and storage medium, aiming at solving the technical problem that the prior art lacks a standardized closed-loop system of task model and action skill mapping facing dynamic environment, and cannot realize efficient autonomous re-planning and task execution optimization based on real-time feedback information.

[0007] To achieve the above purpose, the present application provides a motion sequence execution and re-planning method, comprising:

[0008] Obtaining task definition information, and generating a task model file based on the task definition information;

[0009] Receiving current environment state information, and generating a motion sequence from an initial state to a target state based on the task model file and the current environment state information;

[0010] According to a preset action skill mapping relationship, the actions in the motion sequence are mapped into executable skills;

[0011] Scheduling and executing the executable skills to control the execution unit to perform corresponding operations;

[0012] During or after the execution of the executable skills, the execution state of the execution unit and the perception information of the environment are collected to form feedback information;

[0013] When the feedback information meets a predefined re-planning trigger condition, the feedback information is taken as updated environment state information, and a motion sequence is re-generated based on the task model file and the updated environment state information for re-planning.

[0014] Further, to achieve the above purpose, the present application provides a motion sequence execution and re-planning device, comprising:

[0015] A modeling module for obtaining task definition information and generating a task model file based on the task definition information;

[0016] A planning module for receiving current environment state information and generating a motion sequence from an initial state to a target state based on the task model file and the current environment state information;

[0017] A mapping module for mapping actions in the motion sequence into executable skills according to a preset action skill mapping relationship;

[0018] An execution module for scheduling and executing the executable skills to control the execution unit to perform corresponding operations;

[0019] a perception module, configured to collect perception information of an execution state of the execution unit and an environment during or after execution of the executable skill, and form feedback information;

[0020] a re-planning module, configured to, when the feedback information meets a predefined re-planning trigger condition, take the feedback information as updated environment state information, and re-generate an action sequence based on the task model file and the updated environment state information for re-planning.

[0021] Further, to achieve the above object, the present application also provides a computer device, which comprises a memory, a processor, and an action sequence execution and re-planning program stored in the memory and executable on the processor, and the action sequence execution and re-planning program, when executed by the processor, implements the steps of the action sequence execution and re-planning method as described above.

[0022] Further, to achieve the above object, the present application also provides a computer readable storage medium, which stores an action sequence execution and re-planning program, and the action sequence execution and re-planning program, when executed by a processor, implements the steps of the action sequence execution and re-planning method as described above.

[0023] Beneficial effects: The present application relates to the technical field of artificial intelligence, and can be applied to business scenarios such as robot agent decision-making, financial technology, and medical health. The present application discloses an action sequence execution and re-planning method, device, equipment, and medium, which comprises: obtaining task definition information, generating a task model file based on the task definition information; receiving current environment state information, generating an action sequence from an initial state to a target state based on the task model file and the current environment state information; mapping actions in the action sequence to executable skills according to a preset action skill mapping relationship; scheduling and executing the executable skills to control an execution unit to perform corresponding operations; collecting perception information of an execution state of the execution unit and an environment during or after execution of the executable skill, and forming feedback information; when the feedback information meets a predefined re-planning trigger condition, taking the feedback information as updated environment state information, and re-generating an action sequence based on the task model file and the updated environment state information for re-planning. The present application dynamically generates an action sequence by combining a task model file and real-time environment state information, collects feedback information during execution, triggers a re-planning operation based on the feedback information, forms a complete closed loop of task modeling, action mapping, execution control, and feedback re-planning, and improves the adaptability and reliability of task execution in a complex environment. BRIEF DESCRIPTION OF DRAWINGS

[0024] The present application will be further described below with reference to the accompanying drawings and embodiments. In the drawings:

[0025] Figure 1 A schematic diagram of an application environment of the action sequence execution and re-planning method in an embodiment of the present application;

[0026] Figure 2 A schematic diagram of a flow of the action sequence execution and re-planning method in an embodiment of the present application;

[0027] Figure 3 A schematic diagram of functional modules of the action sequence execution and re-planning device in a preferred embodiment of the present application;

[0028] Figure 4 A schematic diagram of a structure of a computer device in an embodiment of the present application;

[0029] Figure 5 Another schematic diagram of a structure of a computer device in an embodiment of the present application. DETAILED DESCRIPTION

[0030] It should be understood that the specific embodiments described herein are merely exemplary and are not intended to limit the present application.

[0031] The action sequence execution and re-planning method provided by the embodiments of the present application can be applied in an application environment as shown in Figure 1 , in which a user terminal communicates with a server through a network. The server can obtain task definition information through the user terminal, generate a task model file based on the task definition information, receive current environment state information, generate an action sequence from an initial state to a target state based on the task model file and the current environment state information, map actions in the action sequence to executable skills according to a preset action skill mapping relationship, schedule and execute the executable skills to control an execution unit to perform corresponding operations, collect execution state of the execution unit and perception information of the environment during or after execution of the executable skills to form feedback information, and when the feedback information satisfies a predefined re-planning trigger condition, take the feedback information as updated environment state information, and regenerate the action sequence based on the task model file and the updated environment state information to perform re-planning. The present application dynamically generates an action sequence by combining a task model file and real-time environment state information, collects feedback information during execution, triggers a re-planning operation based on the feedback information, forms a complete closed loop of task modeling, action mapping, execution control and feedback re-planning, and improves adaptability and reliability of task execution in a complex environment. The user terminal can be, but is not limited to, various personal computers, notebook computers, smart phones, tablet computers and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers. The present application will be described in detail through specific embodiments.

[0032] Please refer to Figure 2 , Figure 2A flowchart of an embodiment of the action sequence execution and re-planning method provided by the present application is shown. It should be noted that although a logical order is shown in the flowchart, in some cases, the steps shown or described can be performed in an order different from that shown.

[0033] As shown in Figure 2 The action sequence execution and re-planning method provided by the present application includes the following steps:

[0034] S10, obtaining task definition information and generating a task model file based on the task definition information;

[0035] In this embodiment, the process of obtaining task definition information is to collect a set of information expressing specific task requirements, operation constraints and skill composition structure through an external interface or system configuration process for the needs of agent control, complex task planning and execution. The task definition information includes but is not limited to skills, skill parameters, skill preconditions and skill execution effects. Among them, skills refer to the basic abilities or action modules possessed by the agent, which come from the system built-in action library or user-defined function library, skill parameters are used to limit the input data or calling mode of the skill, skill preconditions are used to describe the environmental state, resource state or logical relationship that needs to be met before triggering the skill, and skill execution effects are used to clarify the expected changes to the environmental state, system data or internal state of the agent after the skill execution is completed. The collection of task definition information can be realized through interface access, and the interface includes graphical editing interface, configuration file input, command line input, remote call input and other forms. The specific mode is flexibly configured according to the use scene, and it supports to obtain part of the information or batch import complete structured data.

[0036] The process of generating a task model file based on task definition information is to structure, standardize, and format the collected task definition information, forming a data file that meets the input requirements of the planning system. The task model file is a data expression form for the task planning system, usually containing action identifiers, parameter structures, state constraint expressions, state transition rules, and other content. During the process of generating the task model file, first, the completeness of the skill's parameter content is checked to ensure that all required parameters have a clear data type, value range, and logical relationship, avoiding task parsing failures due to missing or type errors in subsequent processing. Then, the logic of the skill's preconditions is checked, including state expression logic, variable dependency relationships, constraint consistency, and other checks to ensure that the preconditions are theoretically satisfiable. Further, the execution effect of the skill is verified for achievability, including state change logic, impact scope, compatibility with the global task state, and other verification content to ensure that the execution effect does not cause uncontrollable abnormalities or logical conflicts. After verification, the skill is mapped to an action identifier to simplify the task planning system's recognition and invocation process of actions. The skill's parameters are encoded as typed variables to ensure data structure uniformity and system parsability. The skill's preconditions are compiled as state constraints, converted into constraint structures that meet the state space expression specifications. The skill's execution effect is transformed into a state transition strategy, expressing the state change path during skill execution in a structured and standardized data form. The above action identifiers, typed variables, state constraints, and state transition strategies are combined to form a logically closed and complete intermediate model data structure. To improve model reliability and subsequent use safety, verification metadata is added to the intermediate model, including model structure integrity identification, logical consistency labels, parameter constraint ranges, and exception detection markers. Finally, the intermediate model and verification metadata are encapsulated to form a task model file with a structured data format that meets the planning system's invocation requirements.

[0037] The acquisition of task definition information can be achieved in various forms. For structured and large-scale business scenarios, a graphical interface can be used, allowing users to intuitively input task definition information through graphical element combination, parameter configuration, and condition setting, suitable for high reliability and low error rate task configuration requirements. For environments requiring batch import or automatic generation, structured data input mode can be used based on standard format configuration files to quickly load task definition information, suitable for large-scale task deployment and dynamic task updating. For scenarios requiring flexible control and remote invocation, command line interfaces or API interfaces can be provided to allow dynamic setting of task definition information through program calls or remote instructions, suitable for highly dynamic scenarios and remote agent management.

[0038] In the process of generating the task model file, different data expression formats can be selected according to the standards of the task planning system, such as structured text oriented to PDDL language specification, label structured file oriented to XML, or lightweight data format oriented to JSON, to ensure good compatibility and scalability of the task model file between different systems. In the parameter integrity checking process, the parameter checking standard can be dynamically adjusted according to the task complexity to ensure that the basic task passes quickly and the complex task has comprehensive verification. In the logic checking and reachability verification stage, static analysis and dynamic simulation can be combined to improve the logical rigor and actual executability of the task model file. The content of the verification metadata can be extended according to business needs, supporting the introduction of version control information, difference change records, and exception warning labels, to facilitate version management and risk control of the task model file.

[0039] Example: In the medical health business field, for a hospital nursing robot, configure tasks such as patient transfer, material distribution, and environment inspection through a graphical interface, input various skill information, device parameters, and environmental constraints, and the system automatically generates a task model file to achieve high reliability and safety of task planning data preparation in a hospital environment, avoiding omissions and logical errors caused by manual input.

[0040] In the financial technology business field, for an intelligent inspection robot to perform tasks such as machine room inspection and device state verification, configure files to batch import device types, inspection routes, detection standards, and exception handling logic, and the system quickly generates a structured task model file to achieve efficient task deployment and intelligent operation and maintenance in a data center environment.

[0041] In the field of robot agent decision-making business, for multi-agent collaborative transportation tasks, use API interfaces to dynamically obtain the skill capabilities, resource states, and environmental information of each agent, and the system generates a task model file that meets the current environmental state and task requirements in real time, supporting dynamic task planning and efficient collaborative work in complex environments.

[0042] This embodiment can achieve structured expression of task requirements, standardized conversion of action skills and state logic, and input data preparation for task planning systems by obtaining task definition information and generating a task model file, solving the problems of scattered task definition, non-uniform expression, low efficiency and error-prone manual coding, and difficulty for planning systems to efficiently analyze in the prior art, improving the standardization of task configuration and the compatibility of planning systems, and providing a reliable data foundation for subsequent task execution and dynamic adjustment.

[0043] S20, receiving current environmental state information, and generating an action sequence from an initial state to a target state based on the task model file and the current environmental state information;

[0044] In this embodiment, receiving current environment state information means obtaining a data set reflecting the conditions of the environment in which the intelligent agent is located, the state of resources, and external variable information through a system interface, a perception device, or a data transmission module. The environment state information can include, but is not limited to, spatial position, obstacle distribution, target position, resource occupation, sensor feedback, external environment parameters, and the like. The sources involve laser radar, camera, GPS module, ultrasonic sensor, visual system, Internet of Things platform data interface, and the like. The specific content and acquisition method are dynamically configured according to the application scenario to ensure that the environment state information has real-time, accuracy, and structured expression capabilities.

[0045] Using the current environment state information as one of the input data, in combination with the task model file, the action sequence from the initial state to the target state is generated, which belongs to the process of task-oriented execution path planning and action logic reasoning. The task model file contains the expression structure of the task planning domain, the task problem definition, the action identifier, the state constraint information, and the state transition rule, and has the ability to express the task execution logic, the state change path, and the target state constraint. In the process of analyzing the task model file, first, the task planning domain definition is obtained, which is used to limit the action set, parameter type, state variable, and logic structure involved in the task, to ensure that the task generation process is executed within the legal action space, and to avoid generating action combinations that are not executable or do not meet the system capabilities.

[0046] Further, the task planning problem definition is extracted from the task model file, which defines the initial state, target state, constraint condition, and expected result of the current task, constituting the input boundary condition of task planning. The current environment state information is parsed into the initial state expression form required by task planning, ensuring that the environment state information and the parameter structure, state space expression in the task model file are consistent, and avoiding task generation failure or logical errors due to inconsistent data structures.

[0047] Based on the above data, the task planner is called to perform action sequence reasoning and path generation operations. The task planner is a system module for action logic reasoning, path planning, and task sequence optimization, and has the ability to generate an action sequence from the initial state to the target state based on the input state, domain definition, and problem constraint. The task planner can use classical deterministic planning algorithms, heuristic search methods, graph search structures, or extended algorithms combined with uncertainty models, and the specific implementation mode is flexibly configured according to the task complexity, environmental dynamics, and system performance requirements.

[0048] The action sequence is a specific output of the task planning result, containing a logically ordered action list, and each action in the action sequence corresponds to an action identifier in the task model file, ensuring that the system can be parsed in accordance with the task logic. The generation process of the action sequence takes into account the state change path, environmental constraints, resource occupation, and task optimization objectives, ensuring that the final output action sequence has logical rationality, physical feasibility, and environmental adaptability, supporting the stability and reliability of the subsequent task execution process.

[0049] In specific application processes, the acquisition of current environmental state information can be achieved by combining various perception technologies. For example, in indoor navigation scenarios, laser radar and visual SLAM technology can be combined to obtain real-time environmental space maps, obstacle positions, and self-positioning information, forming structured environmental state data. In outdoor mobile platform applications, GPS, inertial navigation units, and remote environmental perception systems can be used to dynamically obtain external environmental information and positioning data. For multi-agent collaboration environments, system-level information sharing mechanisms can be used to integrate the state, resource information, and environmental perception results of each agent to form a unified environmental state representation.

[0050] The structure of the task model file can be adjusted according to application requirements. For complex task planning requirements, a model file containing multi-level action definitions, state hierarchy structures, and composite constraint expressions can be used to support high-complexity task planning and execution. For scenarios with high real-time requirements, lightweight and simplified task model files can be used to improve the efficiency and response speed of task analysis and action sequence generation.

[0051] The generation method of the action sequence is designed according to the differences in task planning system capabilities and application scenarios. For tasks with clear structure and clear goals, classical deterministic planning methods can be used to quickly generate complete action sequences. For tasks with strong environmental dynamics and uncertainties, probability planning, dynamic adjustment, and incremental re-planning techniques can be used to optimize and update action sequences in real time, ensuring adaptability and stability during task execution. The system supports dynamically binding action sequences with environmental states, adjusting action execution paths based on real-time feedback information, and improving the reliability and safety of overall task completion.

[0052] Example: In the medical health business field, for a hospital delivery robot, real-time access to obstacle distribution, personnel dynamic information, and equipment state data in the ward corridor, combined with the task model file of the ward delivery, generates an optimal path action sequence from the current location to the target ward, achieving dynamic obstacle avoidance, path optimization, and safe and stable material delivery.

[0053] In the field of financial technology business, the intelligent inspection robot performs equipment state inspection tasks in the data center. Combined with the space state of the machine room obtained by the environment monitoring system, the device operation parameters and safety alarm information, the inspection task model file is parsed, and the action sequence from the current inspection point to the target device position is generated to ensure a safe and efficient inspection process and to avoid risk areas.

[0054] In the field of robot agent decision-making business, for multi-agent collaborative handling tasks in a warehouse logistics environment, the system obtains real-time environment state information, cargo location and other agent dynamic data in the warehouse area, combines with the handling task model file, generates a collaborative action sequence of multiple agents from the current position to the target stacking area, and realizes an efficient, collaborative and dynamically adaptive handling operation process.

[0055] The embodiment can realize dynamic environment perception before task execution, linkage analysis of task logic structure and environment state, and optimal action path planning for specific environmental conditions by receiving current environment state information and generating an action sequence from an initial state to a target state based on a task model file. The problems of task planning relying on static models, being difficult to cope with environmental changes, and action sequences not having real-time and adaptability in the prior art are solved, and the accuracy and efficiency of task execution and the intelligent decision-making level of the system are improved.

[0056] S30, according to a preset action skill mapping relationship, mapping the actions in the action sequence into executable skills;

[0057] In the embodiment, according to the preset action skill mapping relationship, the actions in the action sequence are mapped into executable skills, involving operation links such as action sequence analysis, mapping relationship retrieval, skill conversion and executable skill generation. The action sequence is a task execution path generated based on the task model file and the environment state information, and contains a set of actions arranged in logical order. Each action corresponds to an action identifier defined in the task planning field, reflecting the task decomposition structure and state transition requirements at the system logic level.

[0058] The preset action skill mapping relationship is a mapping structure between actions and underlying skills defined in the system, which has the ability to express the association between actions and robot underlying control interfaces, execution instructions or function modules. The mapping relationship can be organized in table structure, key-value pair structure, database structure or embedded configuration structure to ensure the scalability, maintainability and efficient retrieval capability of the mapping relationship. Each entry in the mapping relationship associates an action identifier with a corresponding executable skill identifier. The executable skill identifier corresponds to an underlying skill, API interface, function module or control instruction that can actually be called by the system, and has the ability to directly drive the execution unit to complete specific operations.

[0059] The action in the action sequence is mapped to the executable skill. First, the action sequence is parsed, and each action identifier in the action sequence is extracted step by step. For each action identifier, the preset action skill mapping relationship is retrieved, the executable skill identifier corresponding to the current action is obtained, and a one-to-one correspondence between the action and the skill is formed. The mapping process not only requires consistency at the semantic level, but also needs to consider the matching of parameter structure, data type and calling logic to avoid abnormal situations during execution due to inconsistent parameters and incompatible skill interfaces.

[0060] During the mapping relationship retrieval process, the system can combine dynamic verification, version matching and interface availability detection mechanisms to ensure the accuracy and real-time performance of the mapping results. After completing the mapping relationship confirmation, the system generates a structured executable skill set based on the acquired executable skill identifier. The executable skill set reflects the skill sequence that can directly drive the execution unit based on the action sequence logic under the current task, has the characteristics of keeping consistent with the task planning logic structure and being highly compatible with the underlying system interface, and supports the stability and continuity of subsequent task scheduling, skill calling and task execution operations.

[0061] In specific application processes, the preset action skill mapping relationship can be stored in local configuration files, databases or dynamic registration systems. For complex task environments, the mapping relationship supports multi-version switching, dynamic updating and online editing to ensure that the mapping structure is flexibly adjusted with system function upgrades and task demand changes. During the construction process of the mapping relationship, a graphical configuration interface, a semantic analysis system and an automatic verification mechanism can be combined to improve configuration efficiency and accuracy.

[0062] During the action sequence analysis process, the system supports multi-level action identifiers, dynamically adjusts the mapping retrieval logic for composite actions, parameterized actions and environment-dependent actions, and ensures accurate mapping of complex action structures. The skill mapping result can be combined with a skill parameter verification mechanism to ensure that the parameter type, data structure and skill interface are consistent, and to avoid skill calling failures due to parameter errors.

[0063] The generation method of the executable skill set is designed according to the differences in system architecture and task requirements. For single-robot systems, a linear structure skill set is generated to ensure the simplicity and efficiency of the task execution process. For multi-agent systems, a parallel structure, a branch structure or a collaborative structure skill set is generated based on task allocation, collaboration strategy and skill priority structure to improve the overall collaboration efficiency and task completion rate of the system.

[0064] Example explanation: In the medical health business field, for the surgical auxiliary robot system, the system analyzes the action sequence generated by the surgical path planning, combines the preset action skill mapping relationship, maps the logical actions such as navigation, positioning, and mechanical arm adjustment into specific control instructions and skill module calls, and ensures that the robot completes various operation tasks with high precision and safety in the surgical auxiliary process.

[0065] In the financial technology business field, for the data center operation robot, the system generates an action sequence based on the inspection task, uses the mapping relationship to map the logical inspection actions into skill instructions such as device state collection, environment parameter detection, and alarm processing, and realizes intelligent inspection, state perception, and efficient operation of the data center.

[0066] In the robot agent decision-making business field, for the intelligent handling task in the logistics and warehousing scene, the system analyzes the handling path action sequence, combines the mapping relationship to map the handling path, goods grabbing, and putting actions into executable skills, drives the robot agent to efficiently complete goods handling, path obstacle avoidance, and collaborative work, and improves the execution efficiency and intelligence level of the overall logistics system.

[0067] This embodiment maps the actions in the action sequence into executable skills according to the preset action skill mapping relationship, realizes the automatic docking of task logic and bottom layer execution capability, avoids the tediousness and error-prone problem of manual writing of docking code, improves the development efficiency, running reliability and environmental adaptability of the task system, and solves the problem of lack of efficient and stable mapping mechanism between task planning and actual execution in the prior art.

[0068] S40, scheduling and executing the executable skill to control the execution unit to execute the corresponding operation;

[0069] In this embodiment, the executable skill is scheduled and executed to control the execution unit to execute the corresponding operation, which involves skill scheduling strategy, resource management mechanism, execution control process, and state monitoring process. The executable skill is a structured skill set generated by the system based on the task planning result and the skill mapping relationship, and has the ability to directly drive the execution unit to complete specific operation tasks. The execution unit refers to a system component that has specific physical operation capability or control instruction response capability, which can be a mechanical arm, an execution module, an operation terminal, or other system structures with physical or logical execution functions.

[0070] The scheduling process dynamically determines the skill execution order and concurrency strategy based on the system's built-in skill dependency, execution priority, resource usage demand, and system state information. The dependency reflects the logical association between skills, ensuring that the skill execution process meets the task planning logic and system state transition requirements. The execution priority reflects the urgency of the skill, resource occupation, or task importance, and is used to dynamically adjust the execution order in the case of resource competition or system load changes.

[0071] The resource management mechanism combines the system resource monitoring module and resource scheduling strategy to dynamically lock the resources required for skill execution, including hardware resources, interface resources, data resources, or other system resources, to avoid execution failure or system abnormalities caused by resource conflicts. During the resource locking process, the system evaluates the resource occupation state, availability information, and task urgency in real time, and prioritizes the resource acquisition needs of high-priority skills.

[0072] The execution control flow is based on the locked resources and the determined execution order, triggering the actual execution of executable skills in sequence or in parallel. The system drives the execution unit to complete specific operation tasks according to the skill-defined operation logic, parameter configuration, and state requirements through skill invocation interfaces, underlying control modules, or physical drive structures, ensuring that the execution effect and task planning expectations remain consistent.

[0073] The state monitoring process collects operation state information of the execution unit in real time through the system monitoring module, including execution state, task progress, operation feedback, and abnormal signals, ensuring the controllability and transparency of the skill execution process. The system dynamically adjusts the execution strategy based on the state information, responds to abnormal situations or external changes, and improves the stability of task execution and the reliability of the overall system.

[0074] In the specific implementation process, the skill scheduling strategy supports static planning, dynamic adjustment, or autonomous decision-making mechanism design. In the static planning mode, the system determines a fixed skill execution order based on the skill dependency and execution priority generated by the task planning, which is suitable for relatively stable environments and clear task structures. In the dynamic adjustment mode, the system dynamically optimizes the skill execution order and concurrency strategy based on real-time analysis of resource state, environmental changes, and system load, which is suitable for variable environments or complex system structures. In the autonomous decision-making mode, the system dynamically generates skill execution sequences and control strategies based on environmental perception, task feedback, and autonomous reasoning, improving the intelligence level and environmental adaptability of the system.

[0075] The resource management mechanism can combine the system resource monitoring module, intelligent scheduling algorithm, and exception handling mechanism to improve resource utilization and system stability. For multi-robot systems, complex device structures, or resource-constrained environments, the system ensures the continuity of skill execution and stable operation of the system through dynamic resource allocation, resource priority control, and resource conflict detection mechanisms.

[0076] The execution control flow can be implemented through a system call interface, an underlying control bus, an intelligent controller, or a distributed control structure, supports multiple skill types, operation modes, and execution unit structures, ensures the flexibility of skill invocation and system adaptation capability. The state monitoring process combines sensor data, system logs, and operation feedback information, supports real-time anomaly detection, task progress evaluation, and system state analysis, improves the reliability of skill execution and the intelligence level of the system.

[0077] Example: In the medical health business field, for a hospital surgery auxiliary robot system, the system schedules navigation, positioning, and mechanical arm operation skills according to the surgery planning results, controls the execution unit to complete positioning, tissue operation, and auxiliary treatment tasks with high precision, ensures the safety and accuracy of the operation process, and improves the intelligence level of medical operation and the stability of the system.

[0078] In the financial technology business field, for a data center intelligent operation and maintenance robot, the system dynamically schedules state acquisition, environment detection, and emergency response skills based on inspection task planning, drives the execution unit to perform inspection, detection, and fault handling operations, improves the operation efficiency, system safety, and environmental adaptability of the data center, and reduces labor costs and operation risks.

[0079] In the field of robot agent decision-making business, for an intelligent logistics system, the system schedules path navigation, obstacle avoidance, and cargo handling skills, controls the execution unit to efficiently complete path planning, obstacle avoidance, and material handling tasks, improves the intelligent level, resource utilization efficiency, and overall operation efficiency of the logistics system, and supports efficient intelligent handling and collaborative work in complex logistics environments.

[0080] The embodiment schedules and executes executable skills to control the execution unit to perform corresponding operations, and builds an efficient connection mechanism between task planning results and actual operations, avoiding task failure, system instability, or execution abnormalities caused by resource conflicts, improper skill execution order, or lack of state monitoring, improving the execution efficiency, resource utilization, and overall operation reliability of the task system, and solving the problem of lack of intelligent scheduling and stable control mechanism in the task execution stage in the prior art.

[0081] S50, during or after the execution of the executable skill, collecting the execution state of the execution unit and the perception information of the environment to form feedback information;

[0082] In the embodiment, the execution state of the execution unit and the perception information of the environment during or after the execution of the executable skill are collected to form feedback information, involving a data collection mechanism, a state perception structure, an information fusion process, and a feedback data generation process. The executable skill is a skill set determined by the system according to the action sequence mapping relationship, which has the operation function of directly driving the execution unit. The execution unit is a system component with physical operation ability, task execution ability or information feedback ability. The execution state refers to a multidimensional information set reflecting the current operation process, task completion status and system running state of the execution unit. The perception information of the environment is a data set collected by the system through sensors, monitoring devices or data interfaces, reflecting the external environment state, change characteristics and spatial information.

[0083] The data collection mechanism supports two modes of real-time collection during skill execution and centralized collection after skill execution. The data collection process during skill execution continuously acquires operation state, task progress information and environmental perception data of the execution unit through embedded sensors, system monitoring interfaces or remote monitoring devices, ensuring the transparency of the system state and the continuity of environmental perception during task execution. The data collection process after skill execution is based on the system task feedback mechanism, which centrally acquires the final operation state, result information and complete perception data of the environment of the execution unit, improving the data integrity and the accuracy of system state judgment.

[0084] The state perception structure combines multi-source sensor layout, system data interface and state monitoring module to support high-precision, multi-dimensional execution state and environmental information collection, ensuring the comprehensive control of the system over the execution process and environmental changes. The information fusion process unifies the data streams of different sources, formats and time sequences through data synchronization mechanism, data verification logic and multi-source data integration algorithm to generate complete, accurate and usable system feedback information.

[0085] The feedback data generation process is based on the collected execution state data and environmental perception data, and forms feedback information with clear structure, complete content and convenient subsequent processing through data preprocessing, information filtering and structured expression, which is used to support system state evaluation, abnormal detection, task adjustment and environmental adaptability analysis, and to improve the intelligence level and stability of the system.

[0086] In the implementation process, the data acquisition mechanism can adopt a distributed sensor network, an integrated monitoring module combined with a system control bus, and support high-frequency, low-latency, and high-reliability data acquisition requirements. The execution state acquisition can combine position sensors, speed detection modules, operation feedback interfaces, and system log information to fully reflect the running state, task progress, and operation effect of the execution unit. Environmental perception information acquisition can be achieved through visual sensors, laser radars, ultrasonic detectors, environmental monitoring devices, and external data interfaces to realize real-time monitoring of the spatial environment, obstacle distribution, environmental changes, and external interference.

[0087] The information fusion process can combine data timestamp alignment, data format standardization, and redundant data verification logic to improve the accuracy of data fusion and the integrity of system feedback information. The feedback information generation process can use structured data models, data tag systems, and hierarchical information organization methods to ensure clear expression, convenient access, and efficient processing of feedback information, supporting efficient operation of system state analysis, task adjustment, and environmental response mechanisms.

[0088] Example: In the medical and health business field, for a hospital intelligent disinfection robot system, the system continuously acquires robot position state, disinfection progress, and environmental air quality information during the execution of disinfection tasks to form feedback information, supporting real-time path adjustment, operation parameter optimization, and dynamic response to environmental changes, ensuring the comprehensiveness of disinfection tasks, the safety of operations, and the hygiene level of the environment.

[0089] In the financial technology business field, for smart bank outlets, self-service financial terminals, or VIP customer service scenarios, the system continuously acquires running state data of the robot execution unit and perception information of the service environment during tasks such as lobby guidance, teller machine inspection, financial device operation assistance, etc. For example, it collects the voice output state of the intelligent guidance robot during the provision of business consultation services, motion path deviation information, monitors the crowd density in the customer waiting area, the running temperature of the interactive device, the touch screen response, and abnormal sound signals. Through continuous acquisition and centralized collection of the above information, the system can dynamically perceive the service state, device running state, and customer experience status in the financial outlet, form complete feedback information, and serve as a decision basis for subsequent service process dynamic adjustment, customer path optimization, or intelligent terminal abnormality early warning, further ensuring the continuity of the financial service process, the stability of customer experience, and the safety of the overall business environment.

[0090] In the field of robot agent decision-making business, for an intelligent warehouse handling system, the system collects real-time robot operation state, handling progress and warehouse environment information during the execution of the material handling task, forms feedback information, supports the system to dynamically adjust the handling path, optimize the operation strategy and efficiently respond to environmental changes, and improves the intelligent level of the warehouse system, the efficiency of material handling and the reliability of overall operation.

[0091] The embodiment collects the execution state of the execution unit and the perception information of the environment during or after the execution of the executable skill, forms feedback information, and constructs an efficient closed-loop mechanism of system state real-time monitoring, environment change perception and task execution quality evaluation, improves the state transparency, operation reliability and environmental adaptability of the system, and solves the problems of unstable task execution, lagging response to environmental changes and low overall operation efficiency of the system caused by the lack of efficient state perception and environment monitoring mechanism in the existing system.

[0092] S60, when the feedback information meets the predefined re-planning trigger condition, the feedback information is taken as updated environment state information, and an action sequence is regenerated based on the task model file and the updated environment state information for re-planning.

[0093] In the embodiment, when the feedback information meets the predefined re-planning trigger condition, the feedback information is taken as updated environment state information, and an action sequence is regenerated based on the task model file and the updated environment state information for re-planning, involving feedback information determination, environment state updating, task model data calling and action sequence regeneration processes. The feedback information is the combined data of the execution state of the execution unit and the environment perception information collected by the system, the re-planning trigger condition is the set of environmental change indicators, state deviation threshold and task abnormality determination logic preset by the system, the updated environment state information reflects the current environment actual condition, spatial structure and dynamic change, the task model file is the structured task data set pre-constructed by the system, and the action sequence is the operation instruction sequence to achieve the task target.

[0094] The feedback information determination process detects state deviation, environmental change and system abnormality in the feedback information through data comparison logic, condition judgment mechanism and multi-dimensional data analysis, and confirms whether the re-planning trigger condition is met. The environment state updating process extracts environmental change data, state correction information and external interference factors based on the feedback information that meets the condition, constructs updated environment state information, and accurately reflects the current environment state and system operation condition.

[0095] The task model data calling process is based on the task model file, extracts the task definition, operation rules and target information, combines the updated environment state information, comprehensively analyzes the influence of environment change on task execution, adjusts the task parameters, optimizes the operation path and updates the target state. The action sequence regeneration process is based on the adjusted task data and environment information, generates a new action sequence through the system built-in action generation logic, path planning algorithm and operation instruction construction mechanism, ensures the sustainable execution of the task, the environmental adaptability and the stability of the system operation.

[0096] In the specific implementation process, the feedback information determination process can combine multi-source data fusion, environment state comparison and task execution deviation analysis to accurately detect environment changes and state abnormalities and determine whether the re-planning trigger condition is met. The environment state updating process can quickly build accurate updated environment state information by structuring environment perception data, correcting state parameters and updating spatial information, improving the system's environment perception ability and state expression accuracy.

[0097] The task model data calling process can flexibly adapt to environmental changes by task model file parsing, operation logic extraction and target parameter updating, optimizing task execution strategy and operation process. The action sequence regeneration process can dynamically generate a new action sequence based on path re-planning algorithm, action optimization mechanism and task sequence adjustment logic, improving the system's environmental adaptability, task completion efficiency and operation safety level.

[0098] Example: In the medical and health business field, for an intelligent inspection robot system, when the system detects environmental changes, obstacles or task state abnormalities during the inspection process, it updates the environment state based on feedback information, regenerates the inspection action sequence, dynamically adjusts the inspection path and operation strategy, and ensures the continuity of the inspection task, environmental adaptability and operation safety.

[0099] In the financial technology business field, for an intelligent maintenance robot system in data center machine room, financial transaction terminal or self-service equipment maintenance application scenario, the system monitors the equipment running state, environmental safety indicators and abnormal events in real time during the inspection process. When it detects information such as server cabinet temperature abnormalities, financial terminal screen failures or self-service teller machine appearance damage that meets the re-planning trigger condition, the system dynamically updates the environment state based on feedback information, regenerates the maintenance action sequence for the affected area, automatically adjusts the inspection path, operation steps and emergency response strategy, ensures the continuity of the data center maintenance process, the availability of self-service financial services and the stable operation of the overall business system of the financial institution, effectively reduces the need for human intervention, improves the equipment failure response speed and the safety and reliability of the financial business system.

[0100] In the field of robot agent decision-making business, for an intelligent warehouse handling system, when detecting path blockage, environmental changes or task state abnormalities during material handling, the system updates the environmental state based on feedback information, regenerates the handling operation sequence, dynamically adjusts the handling path and operation strategy, and improves the environmental adaptability, handling efficiency and overall operation reliability of the warehouse system.

[0101] The embodiment constructs a dynamic adaptation mechanism, environmental change response capability and task continuous execution guarantee system of the system by taking the feedback information as the updated environmental state information when the feedback information meets the predefined re-planning trigger condition, and regenerating the action sequence based on the task model file and the updated environmental state information, thereby improving the environmental adaptability, task completion reliability and overall operation efficiency of the system, and solving the problems of task interruption, insufficient environmental adaptation and unstable system operation caused by the lack of an efficient re-planning mechanism in the existing system.

[0102] The present application relates to the field of artificial intelligence technology, which can be applied to robot agent decision-making, financial technology and medical health business scenarios, and discloses an action sequence execution and re-planning method, device, equipment and medium, comprising: obtaining task definition information, generating a task model file based on the task definition information; receiving current environmental state information, generating an action sequence from an initial state to a target state based on the task model file and the current environmental state information; mapping the actions in the action sequence to executable skills according to a preset action skill mapping relationship; scheduling and executing the executable skills to control the execution unit to perform corresponding operations; during or after the execution of the executable skills, collecting the execution state of the execution unit and the perception information of the environment to form feedback information; when the feedback information meets a predefined re-planning trigger condition, taking the feedback information as updated environmental state information, and regenerating the action sequence based on the task model file and the updated environmental state information for re-planning. The present application dynamically generates an action sequence by combining a task model file and real-time environmental state information, collects feedback information during execution, triggers a re-planning operation based on the feedback information, forms a complete closed loop of task modeling, action mapping, execution control and feedback re-planning, and improves the adaptability and reliability of task execution in a complex environment.

[0103] In one embodiment, the above step S10 comprises:

[0104] S101, receiving task definition information containing skills, parameters of the skills, preconditions of the skills and execution effects of the skills through a modeling interface;

[0105] S102, checking whether the parameters of the skills are completely defined, verifying whether the precondition logic of the skills is satisfiable, and confirming whether the execution effect state of the skills can be achieved;

[0106] S103, when the parameters of the skill are completely defined, the preconditions of the skill are logically satisfiable, and the execution effect state of the skill is achievable, a verification pass signal is generated;

[0107] S104, based on the verification pass signal, mapping the skill to an action identifier, encoding the parameters of the skill as typed variables, compiling the preconditions of the skill as state constraints, and converting the execution effect of the skill into a state transition strategy;

[0108] S105, combining the action identifier, the typed variables, the state constraints, and the state transition strategy to form an intermediate model;

[0109] S106, adding model verification metadata to the intermediate model, and packaging the intermediate model into a structured data format to generate a task model file.

[0110] In this embodiment, obtaining task definition information involves accessing various types of data describing task structure from upper layer systems, user interfaces or configuration platforms. Skills in the task definition information can be understood as basic operation capabilities or functional modules possessed by robots or intelligent systems, which come from pre-designed functional libraries or capability models. The parameters of the skill include variables, conditions or resource configurations that need to be set during the execution of the skill, such as position parameters, speed parameters, and object recognition parameters, and the specific content depends on the complexity of the designed task and the system capability boundary. The preconditions of the skill are used to limit whether the skill has the basic environment or state for execution, such as accurate object position, system in idle state, completion of the previous operation, etc., to ensure the logical coherence of skill execution and system stability. The execution effect of the skill describes the change of the system state after the skill is completed, which is usually manifested as environment variable update, internal state adjustment or output of perceptible results.

[0111] When receiving task definition information through a modeling interface, flexible data input methods can be provided based on graphical interfaces, script configurations, interface calls or automatic generation mechanisms, combined with actual application requirements. Checking the completeness of the parameters of the skill requires traversing all skill structures to check whether there are missing, conflicting or unconfigured parameter fields, to ensure the effectiveness of the subsequent model. The verification of the precondition logic involves the analysis of logical expressions and the judgment of state reachability, to avoid contradictions, dead loops or unreasonable condition combinations in the configuration. The achievability check of the execution effect state is performed through simulation, deduction or static analysis to ensure that the system can enter the set target state after the skill is executed, and to avoid logical breakage or system instability.

[0112] When the parameters, preconditions and execution effects meet the requirements, the system generates a verification pass signal as the trigger basis for entering the model conversion phase. The skill mapping to action identifier process corresponds the high-level abstract skill name to the specific low-level action code or standard action set, ensuring the consistency and standardization of the internal data structure of the planning system. The parameters of the skill are encoded as typed variables, indicating the data type, value range and scope of each parameter, supporting subsequent reasoning, verification and state transition logic. The preconditions of the skill are compiled as state constraints, defining the limiting conditions of the system state through a standardized expression form, which helps to constrain the boundaries and action space of the task model. The execution effect of the skill is transformed into a state transition strategy, which clearly defines the state change path of the system from before execution to after execution, supporting the planning system to build a complete task state flow model.

[0113] The combination of action identifiers, typed variables, state constraints and state transition strategies forms an intermediate model, which is a structured, standardized and computable internal expression result of task information, supporting the system's general call and analysis under different tasks or different scenarios. The model verification metadata process includes timestamp, version number, data integrity identifier, etc., which facilitates the tracing, comparison and dynamic updating of the task model. The encapsulated structured data format forms a stable, reliable and convenient cross-system transmission and analysis task model file through unified data structure and coding standards.

[0114] The embodiment ensures the accuracy and executability of skill configuration information through parameter integrity verification, precondition logic verification and execution effect reachability check; improves the standardization of task information and the efficiency of system analysis through skill mapping, parameter type coding, precondition state constraint compilation and execution effect state transition strategy generation; enhances the structural standardization and data reliability of the task model file through intermediate model structure combination and model verification metadata encapsulation, realizes the rapid and accurate generation of system model for task planning and execution in complex environments, and improves the stability and automation level of the system task configuration process.

[0115] In one embodiment, the above step S20 comprises:

[0116] S201, receiving current environment state information;

[0117] S202, parsing the task model file to obtain planning domain definition and planning problem definition;

[0118] S203, extracting the target state in the planning problem definition;

[0119] S204, taking the current environment state information as the initial state;

[0120] S205, generating an action sequence from the initial state to the target state by processing the planning domain definition, the initial state and the target state through the planner.

[0121] In the embodiment, the current environment state information is received, which means that the structured description information of the external environment is acquired in real time through a system input interface or a data interaction channel. The environment state information usually includes the spatial layout in the scene, the position of the operable object, the state attribute, the resource occupation condition and the dynamic change factors observable by the system, and the sources can include sensor data, historical environment records, external system synchronization data, etc., and have the characteristics of multi-source fusion and dynamic update.

[0122] The task model file is parsed to obtain the planning domain definition and the planning problem definition, which means that the structural parsing and semantic decoding operations are performed on the generated task model file. The task model file internally contains planning domain information and planning problem information described in a standardized data structure. The planning domain definition includes action sets, state variables, parameter types, state transition rules, etc., which are used to build an executable action logic framework for the system. The planning problem definition includes initial state, target state, constraint conditions, optimization preferences, etc., which are used as input parameters for the planning task.

[0123] The target state in the planning problem definition is extracted, which means that according to the structural specification of the planning problem definition, the data set related to the state desired to be achieved by the system is accurately identified and extracted. The target state is usually expressed in the form of a logical combination of a group of state variables or a state condition table, which is used to constrain the output direction of the final result of the planning process.

[0124] The current environment state information is taken as the initial state, which means that the real-time received environment state information is mapped to the initial state input recognizable by the planner, ensuring that the environment modeling result of the system is consistent with the actual scene, and avoiding the invalidation of the planning result caused by the deviation of the environment state.

[0125] The planning domain definition, the initial state and the target state are processed by the planner to generate an action sequence from the initial state to the target state, which means that based on general planning algorithms, heuristic search, constraint solving or combinatorial optimization methods, the action logic, state constraints and target state requirements in the planning domain are considered comprehensively, and an ordered action execution sequence from the current initial state to the target state is deduced. Each action in the action sequence has clear execution parameters, state influence and preconditions, ensuring that the sequence as a whole has continuity and implementability.

[0126] The embodiment acquires external environment state information in real time, combines the planning domain and planning problem definition in the task model file, dynamically generates a goal-oriented action sequence conforming to the current environment state, improves the task adaptation capability and planning efficiency of the system in a complex and dynamic environment, ensures high matching between the planning result and the actual execution environment, reduces the execution deviation and failure risk caused by environmental changes, and realizes a stable and reliable task planning process for uncertain environments.

[0127] In one embodiment, the above step S30 comprises:

[0128] S301, reading the action sequence and accessing a preset action skill mapping relationship;

[0129] S302, finding the executable skill identifier corresponding to each action in the action sequence in the action skill mapping relationship;

[0130] S303, verifying whether the interface of the executable skill identifier is available, and generating a verification pass signal when the interface is available;

[0131] S304, establishing a binding relationship between the action and the executable skill based on the verification pass signal;

[0132] S305, generating an executable skill set containing the binding relationship of all actions in the action sequence.

[0133] In the embodiment, reading the action sequence refers to that the system obtains the information of each to-be-executed action from the action sequence data generated in the previous step one by one. The action sequence is stored in a structured data format, contains action name, parameter configuration, execution order, state influence and the like, and is used as the basis data source for subsequent skill mapping. Accessing the preset action skill mapping relationship refers to that the system calls the mapping data set stored internally or externally. The data set defines a one-to-one or many-to-many correspondence between the action name and the specific executable skill. The mapping relationship is set in advance based on the task modeling stage and can be dynamically loaded through a static file, a database or an interface.

[0134] Finding the executable skill identifier corresponding to each action in the action sequence in the action skill mapping relationship refers to that the system matches the corresponding item in the mapping relationship for all actions in the action sequence one by one. The skill identifier is a unique data tag for distinguishing different executable skills, usually in the form of a function name, a module number, an instruction code or a service interface path, to ensure accurate association between the action logic and the underlying skill execution unit.

[0135] The interface of the executable skill identification is verified, and the system checks the validity of the corresponding underlying skill interface based on the executable skill identification. The interface availability includes whether the skill implementation module is loaded, whether the interface communication is smooth, whether the dependent resources are met, whether the skill version is compatible, and other check contents, so as to avoid subsequent execution failure caused by interface exception. When the interface verification is passed, the system generates a verification pass signal, which is used to confirm the integrity and reliability of the action to skill mapping link.

[0136] Based on the verification pass signal, the binding relationship between the action and the executable skill is established, that is, the data level mapping structure is formed in the system, and the action information and the corresponding executable skill identification are closely associated. The binding relationship is usually stored in the form of dictionary structure, mapping table or database record, which facilitates efficient retrieval and calling of the scheduling module.

[0137] An executable skill set containing the binding relationship of all actions in the action sequence is generated, that is, the system integrates the above mapping and binding information to construct a complete skill set data. The set covers all the skill identifications and parameter configurations corresponding to the to-be-executed actions, which is used as the standard input for subsequent scheduling execution, ensuring the integrity and accuracy of the skill calling process.

[0138] The embodiment can quickly and accurately convert the high-level action planning result into the underlying executable skill by reading the action sequence and combining the preset action skill mapping relationship. The verification of the skill interface availability improves the safety of the skill calling, the establishment of the binding relationship ensures the completeness and consistency of the data link, and the generated executable skill set provides a clear structure and complete data input condition for the subsequent execution process, which improves the accuracy, stability and system integration efficiency in the task execution process.

[0139] In one embodiment, the above step S40 includes:

[0140] S401, analyzing the dependency relationship in the executable skill, and determining the execution order of the executable skill based on the dependency relationship;

[0141] S402, locking the resources required for the execution unit to perform the operation;

[0142] S403, triggering the execution of the executable skill based on the locked resources and the execution order to control the execution unit;

[0143] S404, monitoring the operation state of the execution unit, and releasing the used resources when the operation is detected to be completed.

[0144] In this embodiment, analyzing the dependency relationship in executable skills means that the system analyzes the correlation of the calling logic, data dependency, resource occupation, time sequence, etc. of each skill in the complete set of executable skills. The dependency relationship usually includes information such as pre-skills, post-skills, mutually exclusive skills, shared resource dependencies, synchronous or asynchronous execution requirements, etc. This process ensures that in the multi-skill collaborative or serial execution scenario, the system can develop a reasonable execution strategy based on the dependency characteristics to avoid abnormal situations caused by skill conflicts, data unpreparedness, or resource competition.

[0145] Determining the execution order of executable skills based on the dependency relationship means that the system forms an ordered and optimized skill execution link by considering the skill priority, task urgency, resource allocation, and overall task flow based on the dependency analysis results. The execution order not only reflects the calling order of skills, but also includes the waiting mechanism between skills, parallel coordination methods, and necessary exception handling nodes to ensure that the skill execution process is orderly, stable, and efficient.

[0146] Locking the resources required for the execution unit to perform operations means that the system completes the occupation and configuration operations in advance for the hardware resources, software modules, network interfaces, or data channels that each skill depends on before skill execution. Common resources include sensors, actuators, communication ports, memory space, computing units, etc. Resource locking is achieved through state markers, resource scheduling modules, or operating system interfaces to prevent different skills or external tasks from competing for critical resources and ensure the exclusivity and safety of skill execution.

[0147] Triggering the execution of executable skills based on the locked resources and execution order to control the execution unit means that the system calls executable skills one by one or in parallel according to the established order, ensuring the stability of skill calling by using the resource locking completed in the previous steps. Skill calling is usually done through interface functions, underlying drivers, API services, or remote instructions to specific execution units. The execution unit is a hardware device or functional module that carries out actual actions, commonly including robot end effectors, control modules, mechanical arms, mobile platforms, production line units, etc. Skill calling controls the execution unit to complete the predetermined operation action.

[0148] Monitoring the operation state of the execution unit and releasing the used resources when detecting that the operation is complete means that the system collects real-time running state data of the execution unit, including position, attitude, sensor feedback, motion feedback signals, or device self-check information, to determine whether the operation action is completed as expected. Operation completion is usually based on state feedback signals, task result information, or time nodes to confirm the skill execution closed loop. Resource release is achieved by removing resource occupation markers, closing interface connections, releasing cache space, and unlocking devices to ensure that subsequent skills or other system tasks can safely and efficiently reuse the corresponding resources.

[0149] The embodiment analyzes the dependency relationship between skills, reasonably formulates the skill execution sequence, combines the resource locking mechanism, improves the coordination and safety in the skill calling process, and through the precise control of the execution unit in the skill execution process, ensures the accurate landing of the operation action, and the real-time monitoring and resource release mechanism guarantees the efficient use of system resources and the stable and reliable execution process, which improves the efficiency, accuracy of task execution and reliability of system operation.

[0150] In one embodiment, the above step S50 comprises:

[0151] S501, during the execution of the executable skill, continuously collecting the execution state of the execution unit and the perception information of the environment;

[0152] S502, after the execution of the executable skill, obtaining the final execution state of the execution unit and the final perception information of the environment;

[0153] S503, fusing the continuously collected execution state and the final execution state to form complete execution state data;

[0154] S504, fusing the continuously collected perception information and the final perception information to form complete environmental perception data;

[0155] S505, combining the complete execution state data and the complete environmental perception data to form feedback information.

[0156] In the embodiment, during the execution of the executable skill, the execution state of the execution unit is continuously collected, which means that after the skill calling starts, the system periodically or in real time acquires the state information of the execution unit through sensors, control interfaces, device self-checking systems and other means. The execution state includes but is not limited to position parameters, attitude angles, motion speeds, system loads, signal feedbacks, device internal states, action execution progress and other data. The continuous collection process adopts data caching, time series recording, queue storage and other methods to ensure the integrity and continuity of the state data, which is convenient for subsequent analysis.

[0157] The continuously collected perception information of the environment means that during the skill execution, the system acquires information data in the working environment in real time through external sensors, vision systems, laser radars, ultrasonic modules, environmental monitoring devices and other devices. The environmental perception information includes obstacle position, dynamic object state, environmental temperature, humidity, noise level, light intensity, spatial structure change and other contents. Continuous collection ensures the dynamic update of environmental information, which provides a basis for the system to understand the changes in the operating environment.

[0158] After the executable skill execution, the final execution state of the execution unit is acquired, that is, at the skill execution action end node, the system acquires the final state information after the operation action is completed through the state feedback mechanism, the action confirmation signal, the device self-checking function or the sensor data. The final execution state reflects the actual result of the skill execution, including the position reaching condition, the action accuracy, the system abnormal information, the execution deviation and the like, which provides a basis for subsequent effect evaluation and system decision.

[0159] The final perception information of the environment is acquired, that is, at the skill execution completion, the system acquires the current latest environment data of the operation environment through the above environment perception device, reflects the final change condition of the environment state, such as the new position of the obstacle, the environment safety, the space structure adjustment and the like, and combines the skill execution result to assist in judging the influence of the environment change on the subsequent task.

[0160] The complete execution state data is formed by fusing the continuously acquired execution state and the final execution state, that is, the system integrates the continuous state data in the skill execution process and the final state information through data merging, time sequence reorganization, redundancy checking, data filtering and the like technical means, forms the state data set covering the whole skill execution process, the accurate data and the complete content.

[0161] The complete environment perception data is formed by fusing the continuously acquired perception information and the final perception information, that is, the system integrates the environment monitoring data in the skill execution process and the environment change information after the execution is completed, forms the environment perception data set with complete structure, accurate time sequence and rich content through data synchronization, space registration, abnormality elimination and the like technical means, so as to facilitate the system to comprehensively understand the environment dynamics in the operation process.

[0162] The feedback information is formed by combining the complete execution state data and the complete environment perception data, that is, the system integrates the above two kinds of data according to the unified data format and the structured standard to form the data basis for the subsequent task adjustment, the system optimization and the re-planning decision. The feedback information includes the skill execution effect, the environment change condition, the system state evaluation and the like, which supports the system closed-loop control, the abnormality detection and the dynamic response.

[0163] In this embodiment, the system comprehensively records the state change and the environment dynamics in the skill execution process by combining the continuous acquisition and the final acquisition, the accuracy and the reliability of the feedback information are improved by fusing the complete data, the data integration process guarantees the structure unification and the content completeness of the feedback information, the feedback information formed provides sufficient and comprehensive data support for the subsequent analysis, the decision and the optimization of the system, and the robustness and the environment adaptability of the task execution are improved.

[0164] In one embodiment, the above step S60 includes:

[0165] S601, parse the feedback information to extract the execution state indicator and the environment perception indicator;

[0166] S602, load a predefined set of re-planning trigger conditions;

[0167] S603, compare the execution state indicator and the environment perception indicator with the threshold range in the set of re-planning trigger conditions;

[0168] S604, when the execution state indicator and the environment perception indicator are not within the threshold range, determine that the re-planning trigger condition is met, and extract the environment state elements including position coordinates and device state variables from the feedback information;

[0169] S605, integrate the environment state elements to construct an environment state object, and convert the environment state object into updated environment state information;

[0170] S606, analyze the change feature based on the updated environment state information;

[0171] S607, verify whether the current action sequence is valid in the updated environment state information based on the change feature, and generate a verification result;

[0172] S608, when the verification result indicates invalid, retain the action segment in the current action sequence that is not affected by the environment change, and reconfigure the action sub-sequence in the current action sequence that is affected by the change feature;

[0173] S609, generate a new action sequence based on the retained action segment and the reconfigured action sub-sequence.

[0174] In this embodiment, the feedback information is parsed to extract the execution state indicator and the environment perception indicator, which means that the system performs structured analysis on the formed feedback information, distinguishes different categories of data in the feedback information, and extracts the execution state indicator for evaluating the execution effect of the skill and the environment perception indicator reflecting the change of the operating environment. The execution state indicator includes but is not limited to position error, action deviation, completion delay, system stability parameter, device load state, etc., and the environment perception indicator includes obstacle dynamics, spatial structure change, environment risk parameter, external interference information, etc. This process ensures the accuracy and reliability of the extracted data through data structure identification, field mapping, parameter analysis and other technical means.

[0175] Load a predefined set of re-planning trigger conditions, which are set in advance by the system according to application requirements and environmental characteristics to determine whether re-planning is needed. The condition set includes multiple independent or combined numerical thresholds, logical rules, and state determination criteria, covering execution state deviation range, environmental change threshold, system safety tolerance, etc. The system dynamically loads and calls this set during task execution as the basis for determining whether to re-plan.

[0176] Compare the execution state indicators and environmental perception indicators with the threshold range in the re-planning trigger condition set, which means the system compares the parsed indicator data with each standard in the predefined condition set one by one, using interval judgment, logical verification, and comprehensive evaluation methods to determine whether the indicators are within the normal range allowed by the system, ensuring the efficiency and accuracy of the comparison process.

[0177] When the execution state indicators and environmental perception indicators are not within the threshold range, it is determined that the re-planning trigger condition is met, which means the system identifies the abnormal indicators or environmental changes beyond the limit based on the comparison results, triggering the internal re-planning logic of the system to ensure the dynamic adjustment capability of task execution and the safety and reliability of the system.

[0178] Extract environmental state elements including position coordinates and device state variables from feedback information, which means the system focuses on extracting key parameters that affect the description of environmental state from feedback information that meets the re-planning condition. Position coordinates reflect spatial position information in the operation scene, and device state variables include indicators such as the working state, abnormal situation, and functionality of execution units, sensors, and control systems. The extraction process ensures data integrity and real-time performance, providing the necessary basic information for updating environmental state.

[0179] Integrate environmental state elements to construct an environmental state object and convert the environmental state object into updated environmental state information, which means the system organizes the extracted environmental state elements into a unified data object through data structure combination, parameter association, and state expression generation. The environmental state object structure is clear and complete, making it easy for the system to handle subsequent processes. The conversion process outputs the environmental state object as updated environmental state information that can be understood and invoked by the system according to system standard formats and interface protocols.

[0180] Analyze the change characteristics based on the updated environmental state information, which means the system identifies the changes in the current environment compared to the existing state through data comparison, state reasoning, spatial mapping, and environmental modeling methods. Change characteristics include position offset, structural change, obstacle addition, and environmental parameter mutation. The analysis process ensures the accuracy and comprehensiveness of change characteristic extraction, facilitating dynamic adjustment of the system.

[0181] The system verifies the validity of the current action sequence in the updated environment state information based on the change feature, generates a verification result, and evaluates whether the existing action sequence can be normally executed in the current environment according to the change feature. The verification content covers action feasibility, path safety, task continuity, system stability and other indicators. The generated verification result is in the form of logical judgment value, executable flag or risk level, reflecting the applicability of the action sequence.

[0182] When the verification result indicates invalidity, the action segments in the current action sequence that are not affected by the environmental change are retained, and the action subsequence in the current action sequence that is affected by the change feature is reconstructed. The system analyzes the action sequence structure according to the change feature and the verification result, identifies the part affected by the environmental change and the part remaining stable, maximizes the retention of the verified safe action segments through sequence segmentation, segment screening, subsequence reconstruction and other technical means, designs new action subsequences for the affected part, and improves the re-planning efficiency.

[0183] Based on the retained action segments and the reconstructed action subsequences, a new action sequence is generated. The system integrates the retained segments and the newly generated subsequences through sequence splicing, structure optimization, parameter adjustment and other methods to form a new action sequence that covers the current task target, adapts to the updated environment state, and meets the system safety requirements, providing a reliable path and operation plan for subsequent task execution.

[0184] Example: In the field of robot agent decision-making, the system can be applied to the task execution process of autonomous transport robot agents in warehouse logistics scenarios. The robot agent receives the transport task demand issued by the warehouse scheduling system. The system receives task definition information through the modeling interface, which includes skills such as transport, obstacle avoidance and path adjustment that can be executed by the agent. The parameters of the skills include the size, weight and target shelf location of the goods. The preconditions of the skills involve the mechanical arm posture before goods grabbing and the obstacle-free state of the operation path. The execution effect of the skills is the successful transport of the goods to the specified location. The system automatically checks whether the skill parameters are complete, verifies whether the precondition logic can be met, confirms whether the execution effect state can be achieved, and generates a verification pass signal. Subsequently, the system maps each skill to an action identifier, encodes the skill parameters as typed variables, compiles the preconditions as state constraints, and converts the execution effect into a state transition strategy to form an intermediate model and encapsulate it as a structured data format, generating a task model file.

[0185] Before the start of the carrying task, the robot agent collects the current environment state information in real time through the environment sensor, including the distribution of obstacles in the field, the position of other mobile units, the situation of the unobstructed passage, the system analyzes and obtains the planning domain definition and planning problem definition based on the task model file and the environment state information, extracts the target state, determines the current environment state as the initial state, calls the path planning module, generates the action sequence from the initial state to the target state, which includes path navigation, obstacle avoidance adjustment, mechanical arm action and other specific operations.

[0186] The system reads the action sequence, accesses the preset action skill mapping relationship, and maps each action in the sequence to a specific executable skill, such as path following control, dynamic obstacle avoidance mechanism, mechanical arm end gripper control, etc. Verify whether the interfaces corresponding to each skill are available, generate a verification signal, and establish a binding relationship between the action and the executable skill, and finally generate a complete set of executable skills.

[0187] During execution, the system analyzes the dependency relationship in the executable skill, such as path planning depending on obstacle avoidance feedback and carrying action depending on navigation position, determines the execution order, locks the necessary system resources and hardware modules, triggers the skills to be executed in turn, and controls the robot agent to complete path navigation, obstacle avoidance, cargo carrying, position calibration and other operations. The system continuously monitors the state of the execution unit, detects the mechanical arm action accuracy, wheel chassis running stability and other parameters, and releases the occupied resources after the task is completed.

[0188] During and after the execution of the executable skill, the system continuously collects execution unit state information and environment perception data, such as cargo position, path deviation, and dynamic obstacle changes, and fuses real-time and final data to form complete feedback information.

[0189] If the feedback information shows that the execution state or environment perception index exceeds the preset threshold range, such as severe path deviation and sudden obstacle occupation of the path, the system determines that the re-planning trigger condition is met, extracts the environment state elements including the current position coordinates, obstacle dynamic information, and device state variables, constructs the environment state object, and updates the environment state information.

[0190] The system analyzes the environment change characteristics based on the updated environment state information, identifies the addition of obstacles and the blocking of the path, evaluates the effectiveness of the current action sequence, and if some actions in the sequence are affected, the system retains the unaffected path segments, reconstructs the affected action subsequence, generates a new complete action sequence, and continues to schedule the robot agent to execute the carrying task, ensuring that the task is dynamically adjusted as needed, the path is reliable, and the overall operation process is safe and stable.

[0191] In the field of medical health, the system can be applied to the auxiliary transport task scenario of autonomous mobile nursing robot agents in hospitals. The hospital information system assigns medicine transport tasks, and the system receives task definition information through a modeling interface. The information includes skills executable by the nursing robot agent, such as medicine taking and placing, path navigation, personnel avoidance, and emergency stop. Skill parameters include medicine type, destination department, and special storage requirements. Skill preconditions involve the robot's location, whether the medicine storage compartment is empty, and whether the channel is unobstructed. The skill execution effect is the safe delivery of the medicine to the designated department. The system automatically checks the integrity of the skill parameters, verifies the logical reasonableness of the preconditions, confirms the execution effect, and generates a verification pass signal. Then, the system maps the skills to action identifiers, encodes the skill parameters as typed variables, compiles the preconditions as state constraints, and translates the execution effect into a state transition strategy. The system combines these to form an intermediate model, encapsulates it as a structured data format, and generates a task model file.

[0192] Before the robot executes the task, it obtains the current environmental state information based on the hospital floor map and the environmental sensing system, including the congestion of the ward corridor, the availability of the elevator, and the sudden obstruction information. The system analyzes the task model file, obtains the planning domain definition and the planning problem definition, extracts the target state, confirms that the current environmental state is the initial state, generates an action sequence from the initial state to the target state through the path planning module, and the sequence includes path navigation, obstacle avoidance actions, medicine taking and placing, and environmental monitoring.

[0193] The system reads the action sequence, accesses the pre-set action skill mapping relationship, maps each action in the sequence to an executable skill, such as route following, dynamic avoidance, automatic hatch opening, and voice prompting, verifies whether the corresponding interface of each skill is available, generates a verification pass signal, establishes the binding relationship between the action and the executable skill, and generates a complete set of executable skills.

[0194] The system analyzes the dependency relationship of the executable skills, such as navigation depending on real-time environmental monitoring and medicine placement depending on accurate position control, determines the execution order of the skills, locks resources such as sensors, control modules, and navigation units, triggers skill execution, controls the robot to complete path driving, obstacle avoidance, cargo compartment control, and state prompting, and releases the resources after the task is completed.

[0195] During and after the execution of the skills, the system continuously collects the execution unit state and environmental perception data, including the robot's position, personnel flow, and obstacle information, and fuses real-time and final data to form complete feedback information.

[0196] If the feedback information shows that the execution state or environmental perception data exceeds the threshold, such as temporary obstruction of the route, elevator unavailability, or personnel density causing path deviation, the system determines that the re-planning trigger condition is met, extracts environmental state elements including the latest position coordinates, dynamic obstacle information, and execution unit state variables, constructs an environmental state object, and updates the environmental state information.

[0197] The system analyzes the environmental change characteristics based on the updated environmental state information, identifies issues such as personnel congestion, elevator delays, and sudden obstacles, evaluates the effectiveness of the current action sequence, and if some action sequences are affected, the system retains unaffected path segments, reconstructs affected action subsequences, generates new action sequences, and ensures that the robot dynamically adapts to route changes in complex hospital environments, safely and efficiently completing the medicine delivery task.

[0198] In the field of financial technology, the system can be applied to the dynamic handling guidance scenario of customer service robots in bank business halls. The bank management system issues customer guidance service tasks, and the system receives task definition information through the modeling interface, which includes skills executable by the agent such as customer identification, business consultation, path navigation, queue management, and risk warning. Skill parameters include customer identity information, pre-handling business type, guidance target location, and business priority. Skill preconditions involve successful customer identity verification, target window idle state, and path unobstructed condition. The skill execution effect is that the customer accurately reaches the business window and completes the corresponding guidance. The system automatically checks the completeness of the skill parameters, verifies whether the preconditions are valid, confirms that the skill execution effect has the implementation conditions, and generates a verification pass signal. Then, the system maps each skill to an action identifier, encodes parameters as typed variables, compiles preconditions as state constraints, and converts execution effects into state transition strategies to form an intermediate model, which is packaged as structured data format and generates a task model file.

[0199] During task execution, the system obtains current environmental state information, including business hall customer flow density, real-time status of each business window, and on-site emergency information. The system analyzes the task model file to obtain planning domain definition and planning problem definition, extracts the target state, determines the current environmental state as the initial state, generates an action sequence from the initial state to the target state based on the planner, and the sequence includes path navigation, personnel avoidance, information interaction, and voice broadcast operations.

[0200] The system reads the action sequence, accesses the pre-set action skill mapping relationship, maps each action in the sequence to executable skills such as face recognition, dynamic path planning, voice prompts, business consultation, and real-time monitoring, verifies the availability of skill interfaces, generates a verification pass signal, establishes the binding relationship between actions and executable skills, and forms an executable skill set.

[0201] The system analyzes the dependencies of executable skills. For example, path navigation depends on real-time environmental monitoring, and customer consultation depends on identity verification. It determines the execution order, locks resources such as the navigation system, voice module, and customer identification unit, triggers the execution of executable skills, and controls the intelligent agent to complete autonomous navigation, customer guidance, business consultation, dynamic obstacle avoidance, information prompts and other operations. It monitors the status of the execution unit in real time, detects navigation accuracy, voice interaction effect, and customer follow-up status, and releases resources after the operation is completed.

[0202] During and after skill execution, the system continuously collects execution unit status and environmental perception information, including customer flow data, path congestion conditions, business window dynamics, and risk event information, integrating real-time and final data to form complete feedback information.

[0203] If the feedback information shows that the execution status or environmental data exceeds the preset threshold, such as customer failure to follow, path obstruction, or temporary business adjustment, the system determines that the re-planning trigger conditions are met, extracts environmental status elements, including the customer's current location, on-site crowd density, and emergency parameters, constructs an environmental status object, and updates the environmental status information.

[0204] Based on the updated environmental status information, the system analyzes the change characteristics, identifies environmental mutations, path failures, business window adjustments, etc., and evaluates the effectiveness of the current action sequence. If the sequence is affected, the unaffected guidance path segments are retained, the action subsequences affected by the environmental changes are reconstructed, and new action sequences are generated to achieve dynamic optimization of the customer guidance process, ensuring that customers can successfully complete business processing instructions in the complex and changing business hall environment, thereby improving overall service efficiency and customer experience.

[0205] Through structured analysis and indicator extraction of feedback information, this embodiment enables the system to accurately monitor task execution status and environmental changes. By comparing with re-planning trigger conditions, the system can efficiently identify abnormal situations that require re-planning. The extraction of environmental status elements and the updating of environmental status information ensure the system's timely response to environmental changes. Change feature analysis and action sequence validity verification improve the accuracy of system decision-making. The mechanism of retaining unaffected action segments and reconstructing affected subsequences greatly improves the efficiency and stability of re-planning. The overall process ensures the continuity of task execution, environmental adaptability and the dynamic adjustment capability of the system.

[0206] In one embodiment, an action sequence execution and re-planning device is provided, which corresponds one-to-one to the action sequence execution and re-planning method in the above embodiment. Figure 3 , Figure 3A function module schematic diagram of a preferred embodiment of the action sequence execution and replanning device of the present application is shown in FIG. 1. A modeling module 10, a planning module 20, a mapping module 30, an execution module 40, a perception module 50, and a replanning module 60 are shown. The function modules are described in detail as follows:

[0207] The modeling module 10 is configured to obtain task definition information and generate a task model file based on the task definition information.

[0208] The planning module 20 is configured to receive current environment state information and generate an action sequence from an initial state to a target state based on the task model file and the current environment state information.

[0209] The mapping module 30 is configured to map actions in the action sequence to executable skills according to a preset action-skill mapping relationship.

[0210] The execution module 40 is configured to schedule and execute the executable skills to control an execution unit to perform corresponding operations.

[0211] The perception module 50 is configured to collect execution state and environment perception information of the execution unit during or after execution of the executable skills to form feedback information.

[0212] The replanning module 60 is configured to, when the feedback information satisfies a predefined replanning trigger condition, take the feedback information as updated environment state information and regenerate an action sequence based on the task model file and the updated environment state information to perform replanning.

[0213] In an embodiment, the modeling module 10 is specifically configured to:

[0214] receive task definition information including skills, parameters of the skills, preconditions of the skills, and execution effects of the skills through a modeling interface;

[0215] check whether the parameters of the skills are completely defined, verify whether the precondition logic of the skills is satisfiable, and confirm whether the execution effect state of the skills is achievable;

[0216] generate a verification pass signal when the parameters of the skills are completely defined, the precondition logic of the skills is satisfiable, and the execution effect state of the skills is achievable;

[0217] map the skills to action identifiers, encode the parameters of the skills to typed variables, compile the preconditions of the skills to state constraints, and convert the execution effects of the skills to state transition strategies based on the verification pass signal;

[0218] combine the action identifiers, the typed variables, the state constraints, and the state transition strategies to form an intermediate model.

[0219] adding model verification metadata to the intermediate model, and packaging the intermediate model into a structured data format to generate a task model file.

[0220] In an embodiment, the planning module 20 is specifically configured to:

[0221] receive current environment state information;

[0222] parse the task model file to obtain a planning domain definition and a planning problem definition;

[0223] extract a target state from the planning problem definition;

[0224] use the current environment state information as an initial state;

[0225] process the planning domain definition, the initial state and the target state through a planner to generate a sequence of actions from the initial state to the target state.

[0226] In an embodiment, the mapping module 30 is specifically configured to:

[0227] read the sequence of actions and access a preset action-skill mapping relationship;

[0228] find an executable skill identifier corresponding to each action in the sequence of actions in the action-skill mapping relationship;

[0229] verify whether an interface of the executable skill identifier is available, and generate a verification pass signal when the interface is available;

[0230] establish a binding relationship between the action and the executable skill based on the verification pass signal;

[0231] generate an executable skill set containing the binding relationship of all actions in the sequence of actions.

[0232] In an embodiment, the execution module 40 is specifically configured to:

[0233] analyze a dependency relationship in the executable skill set and determine an execution order of the executable skill set based on the dependency relationship;

[0234] lock resources required for an operation of an execution unit;

[0235] trigger execution of the executable skill set based on the locked resources and the execution order to control the execution unit;

[0236] monitor an operation state of the execution unit, and release the used resources when detecting that the operation is completed.

[0237] In an embodiment, the perception module 50 is specifically configured to:

[0238] During the execution of the executable skill, continuously collect execution state of the execution unit and perception information of the environment;

[0239] After the execution of the executable skill, obtain the final execution state of the execution unit and the final perception information of the environment;

[0240] Fuse the continuously collected execution state and the final execution state to form complete execution state data;

[0241] Fuse the continuously collected perception information and the final perception information to form complete environment perception data;

[0242] Combine the complete execution state data and the complete environment perception data to form feedback information.

[0243] In an embodiment, the re-planning module 60 is specifically configured to:

[0244] Analyze the feedback information to extract execution state indicators and environment perception indicators;

[0245] Load a predefined set of re-planning trigger conditions;

[0246] Compare the execution state indicators and the environment perception indicators with threshold ranges in the set of re-planning trigger conditions;

[0247] When the execution state indicators and the environment perception indicators are not within the threshold ranges, determine that the re-planning trigger conditions are met, and extract environment state elements including position coordinates and device state variables from the feedback information;

[0248] Integrate the environment state elements to construct an environment state object, and convert the environment state object into updated environment state information;

[0249] Analyze change characteristics based on the updated environment state information;

[0250] Verify whether the current action sequence is valid in the updated environment state information based on the change characteristics, and generate a verification result;

[0251] When the verification result indicates invalidity, retain action segments in the current action sequence that are not affected by the change, and reconfigure action subsequences in the current action sequence that are affected by the change characteristics;

[0252] Generate a new action sequence based on the retained action segments and the reconfigured action subsequences.

[0253] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 4 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide determination and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external user terminal through a network connection. When the computer program is executed by the processor, it realizes the functions or steps of the service side of an action sequence execution and replanning method.

[0254] In one embodiment, a computer device is provided. The computer device may be a user terminal, and its internal structure diagram may be as follows: Figure 5 As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. Among them, the processor of the computer device is used to provide determination and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it realizes the functions or steps of a user-side method for executing and replanning an action sequence.

[0255] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:

[0256] Obtaining task definition information, and generating a task model file based on the task definition information;

[0257] receiving current environment state information, and generating an action sequence from an initial state to a target state based on the task model file and the current environment state information;

[0258] Mapping the actions in the action sequence into executable skills according to a preset action-skill mapping relationship;

[0259] Scheduling and executing the executable skill to control the execution unit to perform the corresponding operation;

[0260] During or after the execution of the executable skill, collecting the execution status of the execution unit and the perception information of the environment to form feedback information;

[0261] When the feedback information meets a predefined re-planning trigger condition, the feedback information is taken as updated environment state information, and an action sequence is re-generated based on the task model file and the updated environment state information for re-planning.

[0262] In an embodiment, a computer readable storage medium is provided, and a computer program is stored on the computer readable storage medium. The computer program is executed by a processor to implement the following steps:

[0263] Task definition information is acquired, and a task model file is generated based on the task definition information;

[0264] Current environment state information is received, and an action sequence from an initial state to a target state is generated based on the task model file and the current environment state information;

[0265] According to a preset action skill mapping relationship, an action in the action sequence is mapped to an executable skill;

[0266] The executable skill is scheduled and executed to control an execution unit to perform a corresponding operation;

[0267] During or after execution of the executable skill, execution state of the execution unit and perception information of an environment are collected to form feedback information;

[0268] When the feedback information meets a predefined re-planning trigger condition, the feedback information is taken as updated environment state information, and an action sequence is re-generated based on the task model file and the updated environment state information for re-planning.

[0269] It should be noted that the functions or steps that the computer readable storage medium or the computer device can implement correspond to the related descriptions of the server side and the user side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0270] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the processes of the above-mentioned embodiments of the methods can be included. Any reference to memory, storage, database or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0271] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is exemplified, and in actual application, the above-mentioned functions can be completed by different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.

[0272] It should be noted that if non-company software tools or components appear in the embodiments of the present application, they are only used for example introduction and do not represent actual use. The above-described embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.

Claims

1. A method for executing and replanning an action sequence, characterized in that: The following steps are involved: Obtaining task definition information, and generating a task model file based on the task definition information; receiving current environment state information, and generating an action sequence from an initial state to a target state based on the task model file and the current environment state information; Mapping the actions in the action sequence into executable skills according to a preset action-skill mapping relationship; Scheduling and executing the executable skill to control the execution unit to perform the corresponding operation; During or after the execution of the executable skill, collecting the execution status of the execution unit and the perception information of the environment to form feedback information; When the feedback information meets a predefined re-planning trigger condition, the feedback information is used as updated environment state information, and an action sequence is regenerated based on the task model file and the updated environment state information for re-planning.

2. The method for executing and replanning an action sequence according to claim 1, wherein: Obtaining task definition information and generating a task model file based on the task definition information includes: Receiving task definition information including skills, skill parameters, skill preconditions, and skill execution effects through a modeling interface; Check whether the skill's parameters are fully defined, verify whether the skill's precondition logic is satisfied, and confirm whether the skill's execution effect state is achievable; When the skill parameters are fully defined, the skill's precondition logic is satisfied, and the skill's execution effect state is achieved, a verification pass signal is generated; Based on the verification pass signal, the skill is mapped into an action identifier, the parameters of the skill are encoded into typed variables, the preconditions of the skill are compiled into state constraints, and the execution effect of the skill is converted into a state transition strategy; Combining the action identifier, typed variables, state constraints, and state transition strategies to form an intermediate model; Add model verification metadata to the intermediate model, encapsulate the intermediate model into a structured data format, and generate a task model file.

3. The method for executing and replanning an action sequence according to claim 1, wherein: Receiving current environment state information and generating an action sequence from an initial state to a target state based on the task model file and the current environment state information, including: Receive current environment status information; Parsing the task model file to obtain a planning domain definition and a planning problem definition; extracting a target state from the planning problem definition; Taking the current environmental state information as the initial state; The planner processes the planning domain definition, the initial state and the target state to generate an action sequence from the initial state to the target state.

4. The method for executing and replanning an action sequence according to claim 1, wherein: Mapping the actions in the action sequence to executable skills according to the preset action-skill mapping relationship includes: Read the action sequence and access the preset action skill mapping relationship; Searching the action-skill mapping relationship for an executable skill identifier corresponding to each action in the action sequence; Verifying whether the interface of the executable skill identifier is available, and generating a verification pass signal when the interface is available; Based on the verification pass signal, establishing a binding relationship between the action and the executable skill; Generate an executable skill set containing binding relationships of all actions in the action sequence.

5. The method for executing and replanning an action sequence according to claim 1, wherein: Scheduling and executing the executable skill to control the execution unit to perform the corresponding operation includes: Analyzing dependencies among the executable skills, and determining an execution order of the executable skills based on the dependencies; Lock the resources required by the execution unit to perform the operation; triggering execution of the executable skill to control an execution unit based on the locked resources and the execution order; The operation status of the execution unit is monitored, and the used resources are released when the operation is detected to be completed.

6. The method for executing and replanning an action sequence according to claim 1, wherein: During or after the execution of the executable skill, the execution status of the execution unit and the perception information of the environment are collected to form feedback information, including: During the execution of the executable skill, continuously collecting the execution status of the execution unit and the perception information of the environment; After the executable skill is executed, obtaining the final execution state of the execution unit and the final perception information of the environment; Fusing the continuously collected execution status and the final execution status to form complete execution status data; Fusing the continuously collected perception information and the final perception information to form complete environmental perception data; The complete execution state data and the complete environment perception data are combined to form feedback information.

7. The method for executing and replanning an action sequence according to claim 1, wherein: When the feedback information satisfies a predefined re-planning trigger condition, the feedback information is used as updated environment state information, and an action sequence is regenerated based on the task model file and the updated environment state information for re-planning, including: Parsing the feedback information to extract execution status indicators and environmental perception indicators; Load a predefined set of replanning trigger conditions; comparing the execution state indicator and the environment perception indicator with a threshold range in the replanning trigger condition set; When the execution status indicator and the environment perception indicator are not within the threshold range, determining that the re-planning trigger condition is met, and extracting the environment status elements including the location coordinates and the device state variables from the feedback information; Integrating the environmental state elements to construct an environmental state object, and converting the environmental state object into updated environmental state information; analyzing change characteristics based on the updated environmental status information; Verifying whether the current action sequence is valid in the updated environmental state information based on the change characteristics, and generating a verification result; When the verification result indicates invalidity, retaining the action segments in the current action sequence that are not affected by the environmental change, and reconstructing the action subsequence in the current action sequence that is affected by the change feature; New action sequences are generated based on the retained action fragments and the reconstructed action subsequences.

8. An action sequence execution and re-planning device, characterized in that: The action sequence execution and re-planning device includes: A modeling module, configured to obtain task definition information and generate a task model file based on the task definition information; A planning module, configured to receive current environment state information and generate an action sequence from an initial state to a target state based on the task model file and the current environment state information; A mapping module, configured to map the actions in the action sequence into executable skills according to a preset action-skill mapping relationship; An execution module, configured to schedule and execute the executable skill to control the execution unit to perform corresponding operations; a perception module, configured to collect perception information of the execution state and environment of the execution unit during or after the execution of the executable skill, and form feedback information; A replanning module is used to use the feedback information as updated environmental state information when the feedback information meets a predefined replanning trigger condition, and regenerate an action sequence based on the task model file and the updated environmental state information for replanning.

9. A computer device, characterized in that: The computer device includes a memory, a processor, and an action sequence execution and replanning program stored in the memory and capable of running on the processor. When the action sequence execution and replanning program is executed by the processor, the steps of the action sequence execution and replanning method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that The storage medium stores an action sequence execution and replanning program, which, when executed by a processor, implements the steps of the action sequence execution and replanning method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Method, system and interface for configurable semiconductor serial production line control

    CN121559998A