Layered task collaboration method and device based on dynamic space-time flow, equipment and medium
By constructing a dynamic spatiotemporal flow graph and a hierarchical task network, the problem of spatiotemporal dynamic evolution of multimodal data is solved, and efficient task decomposition and collaboration between subtasks are achieved, thereby improving the efficiency and accuracy of task execution.
Patent Information
- Application Number
- CN202511100139.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-11-11
AI Technical Summary
Existing technologies struggle to capture the spatiotemporal dynamic evolution of multimodal data when performing vision-language-action tasks, and lack hierarchical task planning and coordination capabilities, resulting in low task execution efficiency and a high susceptibility to errors.
By collecting multimodal data, a dynamic spatiotemporal flow graph is constructed. Complex tasks are broken down into sub-tasks by combining a hierarchical task network, and efficient collaboration between sub-tasks is achieved by using a hierarchical task collaborative reasoning network.
It improves the efficiency and accuracy of complex tasks, and is better able to cope with the dynamic changes of multimodal data and the needs of task collaboration.
Smart Images

Figure CN120929266A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a hierarchical task collaboration method, apparatus, device, and medium based on dynamic spatiotemporal flow. Background Technology
[0002] In existing technologies, when performing Vision-Language-Action (VLA) tasks, the processing of multimodal data generally struggles to capture the dynamic evolution of data over time and space. Traditional models typically process multimodal data statically or discretize it, failing to effectively analyze the intrinsic relationships between dynamic changes in visual scenes, the temporal evolution of language semantics, and spatial trajectory changes in action execution. For example, in intelligent traffic monitoring scenarios, faced with multimodal information such as vehicle trajectories, changes in traffic signs, and voice dispatch commands, traditional solutions cannot accurately grasp the spatiotemporal correlation of this information, leading to misjudgments of abnormal events.
[0003] Furthermore, existing solutions lack hierarchical task planning and coordination capabilities for handling complex tasks. For complex tasks requiring multiple steps and stages, models often employ a single decision-making logic, failing to break down macro-level task objectives into executable micro-level sub-tasks, and struggling to achieve efficient coordination between sub-tasks. This results in low task execution efficiency and a high susceptibility to errors. For example, in the case of robots commonly used in the financial and medical fields, the lack of task coordination can lead to problems such as robot lag and erroneous movements when robot control is required. Summary of the Invention
[0004] In view of the above, it is necessary to provide a hierarchical task collaboration method, device, equipment and medium based on dynamic spatiotemporal flow, which aims to solve the problems of high task execution error rate and low efficiency caused by the lack of analysis of the inherent relationship between multimodal data and the split collaboration of complex tasks.
[0005] A hierarchical task collaboration method based on dynamic spatiotemporal flow, the hierarchical task collaboration method based on dynamic spatiotemporal flow includes:
[0006] In response to a hierarchical task collaboration instruction based on a target task, multimodal data is collected and processed to obtain initial dynamic features of the multimodal data.
[0007] A dynamic spatiotemporal flow graph is constructed based on the aforementioned multimodal initial dynamic features;
[0008] The complexity of the target task is detected by combining the dynamic spatiotemporal flow graph.
[0009] When the target task is determined to be a complex task based on the complexity, the target task is divided into multiple subtasks using a hierarchical task network.
[0010] A hierarchical task structure is constructed based on the dynamic spatiotemporal flow graph and the multiple sub-tasks;
[0011] Obtain a pre-constructed hierarchical task collaborative reasoning network, and input the dynamic spatiotemporal flow graph and the hierarchical task structure into the hierarchical task collaborative reasoning network to obtain the hierarchical task collaborative reasoning result.
[0012] A hierarchical task collaboration device based on dynamic spatiotemporal flow, the hierarchical task collaboration device based on dynamic spatiotemporal flow includes:
[0013] The processing unit is used to respond to hierarchical task collaboration instructions based on the target task, collect multimodal data, and process the multimodal data to obtain initial dynamic features of the multimodality.
[0014] The construction unit is used to construct a dynamic spatiotemporal flow graph based on the initial dynamic features of the multimodal modes;
[0015] A detection unit is used to detect the complexity of the target task in conjunction with the dynamic spatiotemporal flow graph;
[0016] A splitting unit is used to split the target task into multiple subtasks using a hierarchical task network when the target task is determined to be a complex task based on the complexity.
[0017] The construction unit is also used to construct a hierarchical task structure based on the dynamic spatiotemporal flow graph and the multiple subtasks;
[0018] The collaborative unit is used to acquire a pre-constructed hierarchical task collaborative reasoning network, and input the dynamic spatiotemporal flow graph and the hierarchical task structure into the hierarchical task collaborative reasoning network to obtain the hierarchical task collaborative reasoning result.
[0019] A computer device, the computer device comprising:
[0020] Memory, storing at least one instruction; and
[0021] The processor executes the instructions stored in the memory to implement the hierarchical task collaboration method based on dynamic spatiotemporal flow.
[0022] A computer-readable storage medium storing at least one instruction, which is executed by a processor in a computer device to implement the hierarchical task collaboration method based on dynamic spatiotemporal flow.
[0023] As can be seen from the above technical solutions, this invention can process multimodal data to obtain initial dynamic features of multimodality, thereby providing high-quality dynamic spatiotemporal features; construct a dynamic spatiotemporal flow graph based on the initial dynamic features of multimodality to provide various relationships between key events in multimodal data, reflecting the dynamic evolution of multimodal data; combine the dynamic spatiotemporal flow graph to detect the complexity of the target task, and use a hierarchical task network to decompose the high-complexity target task into multiple sub-tasks. Based on the dynamic spatiotemporal flow graph and multiple sub-tasks, a hierarchical task structure is constructed, which can reasonably decompose the macro-level task target into executable micro-level sub-tasks; input the dynamic spatiotemporal flow graph and hierarchical task structure into a hierarchical task collaborative reasoning network to obtain hierarchical task collaborative reasoning results, which can help achieve efficient collaboration between sub-tasks and accurate execution of actions, improve the execution efficiency and accuracy of complex tasks, and better meet the needs of complex task processing. Attached Figure Description
[0024] Figure 1 This is a flowchart of a preferred embodiment of the hierarchical task collaboration method based on dynamic spatiotemporal flow of the present invention.
[0025] Figure 2 This is a functional block diagram of a preferred embodiment of the hierarchical task collaboration device based on dynamic spatiotemporal flow of the present invention.
[0026] Figure 3 This is a schematic diagram of the structure of a computer device that implements a preferred embodiment of the hierarchical task collaboration method based on dynamic spatiotemporal flow according to the present invention. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0028] like Figure 1 The diagram shown is a flowchart of a preferred embodiment of the hierarchical task collaboration method based on dynamic spatiotemporal flow according to the present invention. The order of the steps in this flowchart can be changed, and some steps can be omitted, depending on different requirements.
[0029] The hierarchical task collaboration method based on dynamic spatiotemporal flow is applied to one or more computer devices. The computer device is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0030] The computer device can be any electronic product that can interact with the user, such as a personal computer, tablet computer, smartphone, personal digital assistant (PDA), game console, interactive network television (IPTV), smart wearable device, etc.
[0031] The computer equipment may also include network equipment and / or user equipment. The network equipment includes, but is not limited to, a single network server, a server group consisting of multiple network servers, or a cloud based on cloud computing consisting of a large number of hosts or network servers.
[0032] The server can be a standalone server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.
[0033] Artificial intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.
[0034] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0035] The network in which the computer device is located includes, but is not limited to, the Internet, wide area network, metropolitan area network, local area network, and virtual private network (VPN).
[0036] S10, in response to the hierarchical task collaboration instruction based on the target task, collect multimodal data and process the multimodal data to obtain the initial dynamic features of the multimodality.
[0037] In this embodiment, the target task can be the control task of a robot in a financial hall, the control task of an auxiliary robot in a medical operating room, or the control task of a vehicle in an intelligent traffic monitoring scenario, etc.
[0038] In this embodiment, the hierarchical task collaboration instruction can be triggered when a preset button is selected on a designated interface.
[0039] In this embodiment, the multimodal data may include, but is not limited to, one or more of the following dimensional data combinations:
[0040] Video stream data, language text sequence data, action time series data, etc.
[0041] In this embodiment, processing the multimodal data to obtain the initial dynamic features of the multimodal data includes:
[0042] Extract video stream data, language text sequence data, and action time series data from the multimodal data;
[0043] The Swin Transformer 3D model is used to extract local image features of the video stream data in the spatial dimension using a sliding window mechanism; a temporal self-attention mechanism is used to capture object motion features and scene change features between consecutive frames in the video stream data in the temporal dimension; an optical flow estimation algorithm is combined to calculate the motion vectors of pixels between frames in the video stream data; and the image local features, object motion features and scene change features between consecutive frames in the video stream data, and motion vectors of pixels between frames in the video stream data are combined to obtain a visual dynamic spatiotemporal feature vector.
[0044] Using a pre-trained language model, the language text sequence data is encoded based on a temporal attention mechanism to obtain a dynamic semantic feature vector of the language;
[0045] Temporal features are extracted from the action time series data using a Temporal Convolutional Network (TCN), and action features are analyzed based on the temporal features. The action features are then transformed into a dynamic spatiotemporal feature vector containing spatiotemporal information based on the spatial coordinates of the action execution recorded in the action time series data.
[0046] The visual dynamic spatiotemporal feature vector, the language dynamic semantic feature vector, and the action dynamic spatiotemporal feature vector are integrated to obtain the multimodal initial dynamic features.
[0047] By fusing local image features of the video stream data, object motion features and scene change features between consecutive frames in the video stream data, and motion vectors of pixels between frames in the video stream data, the expression of dynamic visual information can be enhanced.
[0048] Before encoding the language text sequence data based on the temporal attention mechanism, NLTK (Natural Language Toolkit) can be used to preprocess the language text sequence data, such as word segmentation and part-of-speech tagging, to reduce the difficulty of subsequent data processing.
[0049] The language text sequence data can be encoded using a Transformer-based language model.
[0050] By introducing a temporal attention mechanism, we can focus on time-clue words (such as "firstly", "then", "finally") and semantic transition words in the text to extract the evolutionary features of language semantics over time.
[0051] The temporal convolutional network can extract the temporal features of the action, analyze the speed and acceleration variation patterns of the action, and combine the spatial coordinates of the action execution (such as the three-dimensional position of the robot's end effector) to transform the action data into a feature vector containing spatiotemporal information.
[0052] For example, in the field of financial surveillance, dynamic behavioral features of people in surveillance videos (such as abnormal fund transfer actions), dynamic semantic features in transaction instructions (such as descriptions of transaction time and amount changes), and action features of transaction equipment operations (such as the rhythm of keyboard input and the trajectory of mouse clicks) can be extracted. Multimodal feature fusion can then be used to support the identification of abnormal behaviors such as financial fraud. In medical scenarios, dynamic visual features of videos during surgery (such as the movement trajectory of surgical instruments and changes in the state of the patient's organs), verbal instructions between doctors (such as descriptions of surgical steps and communication of precautions), and action features of surgical robots (such as the movement angle and force changes of the robotic arm) can be extracted. By analyzing multimodal feature fusion data, doctors can be assisted in better understanding the surgical process.
[0053] Through the above embodiments, dynamic spatiotemporal features in multiple modal data can be extracted comprehensively and accurately, providing rich and high-quality basic feature data for subsequent model processing and enhancing the model's ability to express multimodal dynamic information.
[0054] S11, construct a dynamic spatiotemporal flow graph based on the initial dynamic features of the multimodal modes.
[0055] In this embodiment, constructing a dynamic spatiotemporal flow graph based on the multimodal initial dynamic features includes:
[0056] Identify events in the initial dynamic features of the multimodal model as nodes, and add spatiotemporal labels to each node;
[0057] Identify the spatiotemporal sequence, causal relationships, and logical dependencies among the events as edges;
[0058] Through the message passing mechanism, each node absorbs the spatiotemporal information of its neighboring nodes to update its own feature representation;
[0059] After the message passing is completed, the currently obtained flow graph is determined as the dynamic spatiotemporal flow graph.
[0060] The events in the initial dynamic features of the multimodal system may include, but are not limited to: the appearance or disappearance of objects in vision, the issuance of language commands, key nodes of actions, etc.
[0061] The spatiotemporal tags may include, but are not limited to, timestamps and spatial coordinates.
[0062] Edges are used to reflect the relationships between events. For example, if a vehicle is detected starting visually, and then a driving command is issued verbally, a causal edge can be established between the "vehicle start" node and the "driving command issuance" node.
[0063] Among them, nodes and edges can be processed by dynamic graph neural networks (Dynamic GNNs). Through message passing mechanism, each node absorbs the spatiotemporal information of its neighboring nodes to update its own feature representation, and the dynamic spatiotemporal flow graph is updated in real time with the input of new multimodal data.
[0064] Through the above embodiments, the constructed dynamic spatiotemporal flow graph can clearly present the various relationships between key events in multimodal data, reflect the dynamic evolution of multimodal data, and enable the models used subsequently to more accurately understand the spatiotemporal state of the current scene.
[0065] S12, Combine the dynamic spatiotemporal flow graph to detect the complexity of the target task.
[0066] In this embodiment, the hierarchical task collaboration instructions can be parsed using natural language processing technology, and the task complexity can be determined by combining the dynamic spatiotemporal flow graph.
[0067] For example, the nodes and edges associated with the hierarchical task collaboration instructions can be detected through the dynamic spatiotemporal flow graph, and the complexity of the target task can be determined by the complexity of the associated nodes and edges.
[0068] S13, when the target task is determined to be a complex task based on the complexity, the target task is split into multiple subtasks using a Hierarchical Task Network (HTN).
[0069] In this embodiment, when the target task is determined to be a simple task based on the complexity, the target task can be executed directly without splitting it.
[0070] In this embodiment, the hierarchical task network is an artificial intelligence technology used for task planning and solving complex problems.
[0071] For example, using the hierarchical task network, the task of "intelligent express sorting" can be broken down into sub-tasks such as "package scanning", "path planning", and "robotic arm grasping and placement"; for the complex task of "annual financial audit" in the financial scenario, it can be broken down into sub-tasks such as "voucher collection", "account reconciliation", and "risk assessment"; for the task of "comprehensive auxiliary tumor detection" in the medical scenario, it can be broken down into sub-tasks such as "imaging examination", "pathological analysis", and "treatment plan formulation".
[0072] The above embodiments enable the reasonable decomposition of complex tasks, transforming macro-level task objectives into executable micro-level sub-tasks.
[0073] S14, construct a hierarchical task structure based on the dynamic spatiotemporal flow graph and the multiple sub-tasks.
[0074] In this embodiment, constructing a hierarchical task structure based on the dynamic spatiotemporal flow graph and the multiple subtasks includes:
[0075] Obtain the spatiotemporal relationships and task logic in the dynamic spatiotemporal flow graph;
[0076] Configure sub-goals for each subtask;
[0077] Based on the spatiotemporal relationship and the task logic, the sub-objectives of each sub-task are prioritized to obtain a hierarchical task structure.
[0078] In this process, prioritizing each sub-target based on the spatiotemporal relationships and task logic in the dynamic spatiotemporal flow graph ensures that the execution order of the sub-tasks conforms to spatiotemporal constraints and task requirements, thereby forming the hierarchical task structure.
[0079] Specifically, the sub-goals configured for each sub-task include:
[0080] Based on the spatiotemporal sequence, causal relationships, and logical dependencies between the events, as well as the spatiotemporal relationships and task logic in the dynamic spatiotemporal flow graph, configure the sub-goals of each sub-task.
[0081] For example, the goal of the "parcel scanning" subtask in the logistics field is to "accurately identify parcel barcodes and destination information".
[0082] The above embodiments clarify the objectives and execution order of subtasks, laying the foundation for the efficient execution of subsequent tasks.
[0083] S15, obtain the pre-constructed hierarchical task collaborative reasoning network, and input the dynamic spatiotemporal flow graph and the hierarchical task structure into the hierarchical task collaborative reasoning network to obtain the hierarchical task collaborative reasoning result.
[0084] In this embodiment, inputting the dynamic spatiotemporal flow graph and the hierarchical task structure into the hierarchical task collaborative reasoning network to obtain the hierarchical task collaborative reasoning result includes:
[0085] Obtain the high-level task coordination network in the hierarchical task collaborative reasoning network, and the low-level subtask execution network corresponding to each subtask;
[0086] The dynamic spatiotemporal flow graph is encoded to obtain encoded features;
[0087] The encoded features and the hierarchical task structure are input into the high-level task coordination network, and the association between the sub-objectives and spatiotemporal information of each sub-task is learned based on the multi-head self-attention mechanism to obtain the cooperative strategy between each currently executed sub-task and other sub-tasks.
[0088] The sub-objectives of each subtask and the corresponding multimodal initial dynamic features are input into the underlying subtask execution network corresponding to each subtask to obtain the execution action sequence of each subtask.
[0089] The hierarchical task collaborative reasoning result is generated based on the collaborative strategy and the execution action sequence of each sub-task.
[0090] The high-level task coordination network can be built based on the Transformer network.
[0091] The coordination strategy can clearly define the execution order of each subtask. For example, based on the positional changes of the package in the dynamic spatiotemporal flow diagram and the working state of the robotic arm, the execution order of the "path planning" subtask and the "robotic arm grasping and placing" subtask can be coordinated.
[0092] Specifically, a low-level subtask execution network can be pre-built for each subtask. This low-level subtask execution network can be a deep learning model, such as a visual object detection network or a motion reinforcement learning policy network. For example, the execution network for the "robotic arm grasping and placing" subtask can determine the package position based on visual features, plan the robotic arm's motion trajectory based on motion features, and finally output joint angle control commands.
[0093] For example, in the "surgical robot-assisted surgery" task in a medical setting, the high-level network coordinates sub-tasks such as "instrument preparation", "lesion localization" and "surgical operation", while the low-level execution networks, such as the "lesion localization" network, determine the location of the lesion based on medical image features, and the "surgical operation" network controls the robotic arm to perform precise surgical operations. Thus, the surgical accuracy and efficiency are improved through multi-task decomposition and collaboration.
[0094] In the above embodiments, the high-level task coordination network controls the start and pause of subtasks and dynamically adjusts the coordination strategy based on feedback information during the execution of subtasks; the low-level subtask execution network is responsible for the execution of specific actions, realizing hierarchical task collaborative execution driven by multimodal information.
[0095] In this embodiment, after obtaining the hierarchical task collaborative reasoning result, the method further includes:
[0096] The execution action sequence of each subtask is scheduled according to the coordination strategy to coordinate the execution of the execution action sequence of each subtask.
[0097] Through the above embodiments, efficient collaboration between subtasks and precise execution of specific actions are achieved, improving the execution efficiency and accuracy of complex tasks and enabling the model to better meet the needs of complex task processing.
[0098] In this embodiment, multi-dimensional feedback information during task execution can also be collected, including sub-task completion status (e.g., whether the package was successfully scanned, whether the robotic arm accurately grasped it), action execution results (e.g., trajectory deviation, execution time), and environmental state changes (e.g., new packages, equipment malfunctions). Simultaneously, the changes in the dynamic spatiotemporal flow graph during task execution and the effectiveness of the hierarchical task collaboration strategy are recorded. Furthermore, the feedback information is combined with the original multimodal dynamic spatiotemporal features, dynamic spatiotemporal flow graph, and hierarchical task structure, and used as training data to optimize the corresponding model. The parameters of dynamic spatiotemporal feature extraction, dynamic spatiotemporal flow graph construction, and hierarchical task collaboration network are updated using the backpropagation algorithm. Based on the feedback, the task decomposition strategy, sub-task goal setting method, and collaborative inference logic are adjusted, enabling the model to process multimodal dynamic information and complex tasks more efficiently in subsequent tasks.
[0099] As can be seen from the above technical solutions, this invention can process multimodal data to obtain initial dynamic features of multimodality, thereby providing high-quality dynamic spatiotemporal features; construct a dynamic spatiotemporal flow graph based on the initial dynamic features of multimodality to provide various relationships between key events in multimodal data, reflecting the dynamic evolution of multimodal data; combine the dynamic spatiotemporal flow graph to detect the complexity of the target task, and use a hierarchical task network to decompose the high-complexity target task into multiple sub-tasks. Based on the dynamic spatiotemporal flow graph and multiple sub-tasks, a hierarchical task structure is constructed, which can reasonably decompose the macro-level task target into executable micro-level sub-tasks; input the dynamic spatiotemporal flow graph and hierarchical task structure into a hierarchical task collaborative reasoning network to obtain hierarchical task collaborative reasoning results, which can help achieve efficient collaboration between sub-tasks and accurate execution of actions, improve the execution efficiency and accuracy of complex tasks, and better meet the needs of complex task processing.
[0100] like Figure 2 The diagram shown is a functional block diagram of a preferred embodiment of the hierarchical task collaboration device based on dynamic spatiotemporal flow of the present invention. The hierarchical task collaboration device 11 based on dynamic spatiotemporal flow includes a processing unit 110, a construction unit 111, a detection unit 112, a splitting unit 113, and a collaboration unit 114. The module / unit referred to in this invention refers to a series of computer program segments that can be executed by a processor and perform a fixed function, and are stored in memory. In this embodiment, the functions of each module / unit will be described in detail in subsequent embodiments.
[0101] The processing unit 110 is used to respond to a hierarchical task collaboration instruction based on a target task, collect multimodal data, and process the multimodal data to obtain initial dynamic features of the multimodality.
[0102] In this embodiment, the target task can be the control task of a robot in a financial hall, the control task of an auxiliary robot in a medical operating room, or the control task of a vehicle in an intelligent traffic monitoring scenario, etc.
[0103] In this embodiment, the hierarchical task collaboration instruction can be triggered when a preset button is selected on a designated interface.
[0104] In this embodiment, the multimodal data may include, but is not limited to, one or more of the following dimensional data combinations:
[0105] Video stream data, language text sequence data, action time series data, etc.
[0106] In this embodiment, the processing unit 110 processes the multimodal data to obtain the initial dynamic features of the multimodal data, including:
[0107] Extract video stream data, language text sequence data, and action time series data from the multimodal data;
[0108] The Swin Transformer 3D model is used to extract local image features of the video stream data in the spatial dimension using a sliding window mechanism; a temporal self-attention mechanism is used to capture object motion features and scene change features between consecutive frames in the video stream data in the temporal dimension; an optical flow estimation algorithm is combined to calculate the motion vectors of pixels between frames in the video stream data; and the image local features, object motion features and scene change features between consecutive frames in the video stream data, and motion vectors of pixels between frames in the video stream data are combined to obtain a visual dynamic spatiotemporal feature vector.
[0109] Using a pre-trained language model, the language text sequence data is encoded based on a temporal attention mechanism to obtain a dynamic semantic feature vector of the language;
[0110] Temporal features are extracted from the action time series data using a Temporal Convolutional Network (TCN), and action features are analyzed based on the temporal features. The action features are then transformed into a dynamic spatiotemporal feature vector containing spatiotemporal information based on the spatial coordinates of the action execution recorded in the action time series data.
[0111] The visual dynamic spatiotemporal feature vector, the language dynamic semantic feature vector, and the action dynamic spatiotemporal feature vector are integrated to obtain the multimodal initial dynamic features.
[0112] By fusing local image features of the video stream data, object motion features and scene change features between consecutive frames in the video stream data, and motion vectors of pixels between frames in the video stream data, the expression of dynamic visual information can be enhanced.
[0113] Before encoding the language text sequence data based on the temporal attention mechanism, NLTK (Natural Language Toolkit) can be used to preprocess the language text sequence data, such as word segmentation and part-of-speech tagging, to reduce the difficulty of subsequent data processing.
[0114] The language text sequence data can be encoded using a Transformer-based language model.
[0115] By introducing a temporal attention mechanism, we can focus on time-clue words (such as "firstly", "then", "finally") and semantic transition words in the text to extract the evolutionary features of language semantics over time.
[0116] The temporal convolutional network can extract the temporal features of the action, analyze the speed and acceleration variation patterns of the action, and combine the spatial coordinates of the action execution (such as the three-dimensional position of the robot's end effector) to transform the action data into a feature vector containing spatiotemporal information.
[0117] For example, in the field of financial surveillance, dynamic behavioral features of people in surveillance videos (such as abnormal fund transfer actions), dynamic semantic features in transaction instructions (such as descriptions of transaction time and amount changes), and action features of transaction equipment operations (such as the rhythm of keyboard input and the trajectory of mouse clicks) can be extracted. Multimodal feature fusion can then be used to support the identification of abnormal behaviors such as financial fraud. In medical scenarios, dynamic visual features of videos during surgery (such as the movement trajectory of surgical instruments and changes in the state of the patient's organs), verbal instructions between doctors (such as descriptions of surgical steps and communication of precautions), and action features of surgical robots (such as the movement angle and force changes of the robotic arm) can be extracted. By analyzing multimodal feature fusion data, doctors can be assisted in better understanding the surgical process.
[0118] Through the above embodiments, dynamic spatiotemporal features in multiple modal data can be extracted comprehensively and accurately, providing rich and high-quality basic feature data for subsequent model processing and enhancing the model's ability to express multimodal dynamic information.
[0119] The construction unit 111 is used to construct a dynamic spatiotemporal flow graph based on the initial dynamic features of the multimodal system.
[0120] In this embodiment, the construction unit 111 constructs a dynamic spatiotemporal flow graph based on the multimodal initial dynamic features, including:
[0121] Identify events in the initial dynamic features of the multimodal model as nodes, and add spatiotemporal labels to each node;
[0122] Identify the spatiotemporal sequence, causal relationships, and logical dependencies among the events as edges;
[0123] Through the message passing mechanism, each node absorbs the spatiotemporal information of its neighboring nodes to update its own feature representation;
[0124] After the message passing is completed, the currently obtained flow graph is determined as the dynamic spatiotemporal flow graph.
[0125] The events in the initial dynamic features of the multimodal system may include, but are not limited to: the appearance or disappearance of objects in vision, the issuance of language commands, key nodes of actions, etc.
[0126] The spatiotemporal tags may include, but are not limited to, timestamps and spatial coordinates.
[0127] Edges are used to reflect the relationships between events. For example, if a vehicle is detected starting visually, and then a driving command is issued verbally, a causal edge can be established between the "vehicle start" node and the "driving command issuance" node.
[0128] Among them, nodes and edges can be processed by dynamic graph neural networks (Dynamic GNNs). Through message passing mechanism, each node absorbs the spatiotemporal information of its neighboring nodes to update its own feature representation, and the dynamic spatiotemporal flow graph is updated in real time with the input of new multimodal data.
[0129] Through the above embodiments, the constructed dynamic spatiotemporal flow graph can clearly present the various relationships between key events in multimodal data, reflect the dynamic evolution of multimodal data, and enable the models used subsequently to more accurately understand the spatiotemporal state of the current scene.
[0130] The detection unit 112 is used to detect the complexity of the target task in conjunction with the dynamic spatiotemporal flow graph.
[0131] In this embodiment, the hierarchical task collaboration instructions can be parsed using natural language processing technology, and the task complexity can be determined by combining the dynamic spatiotemporal flow graph.
[0132] For example, the nodes and edges associated with the hierarchical task collaboration instructions can be detected through the dynamic spatiotemporal flow graph, and the complexity of the target task can be determined by the complexity of the associated nodes and edges.
[0133] The splitting unit 113 is used to split the target task into multiple subtasks using a hierarchical task network (HTN) when the target task is determined to be a complex task based on the complexity.
[0134] In this embodiment, when the target task is determined to be a simple task based on the complexity, the target task can be executed directly without splitting it.
[0135] In this embodiment, the hierarchical task network is an artificial intelligence technology used for task planning and solving complex problems.
[0136] For example, using the hierarchical task network, the task of "intelligent express sorting" can be broken down into sub-tasks such as "package scanning", "path planning", and "robotic arm grasping and placement"; for the complex task of "annual financial audit" in the financial scenario, it can be broken down into sub-tasks such as "voucher collection", "account reconciliation", and "risk assessment"; for the task of "comprehensive auxiliary tumor detection" in the medical scenario, it can be broken down into sub-tasks such as "imaging examination", "pathological analysis", and "treatment plan formulation".
[0137] The above embodiments enable the reasonable decomposition of complex tasks, transforming macro-level task objectives into executable micro-level sub-tasks.
[0138] The construction unit 111 is also used to construct a hierarchical task structure based on the dynamic spatiotemporal flow graph and the multiple sub-tasks.
[0139] In this embodiment, the construction unit 111 constructs a hierarchical task structure based on the dynamic spatiotemporal flow graph and the plurality of subtasks, including:
[0140] Obtain the spatiotemporal relationships and task logic in the dynamic spatiotemporal flow graph;
[0141] Configure sub-goals for each subtask;
[0142] Based on the spatiotemporal relationship and the task logic, the sub-objectives of each sub-task are prioritized to obtain a hierarchical task structure.
[0143] In this process, prioritizing each sub-target based on the spatiotemporal relationships and task logic in the dynamic spatiotemporal flow graph ensures that the execution order of the sub-tasks conforms to spatiotemporal constraints and task requirements, thereby forming the hierarchical task structure.
[0144] Specifically, the sub-goals configured for each sub-task include:
[0145] Based on the spatiotemporal sequence, causal relationships, and logical dependencies between the events, as well as the spatiotemporal relationships and task logic in the dynamic spatiotemporal flow graph, configure the sub-goals of each sub-task.
[0146] For example, the goal of the "parcel scanning" subtask in the logistics field is to "accurately identify parcel barcodes and destination information".
[0147] The above embodiments clarify the objectives and execution order of subtasks, laying the foundation for the efficient execution of subsequent tasks.
[0148] The collaborative unit 114 is used to acquire a pre-constructed hierarchical task collaborative reasoning network, and input the dynamic spatiotemporal flow graph and the hierarchical task structure into the hierarchical task collaborative reasoning network to obtain the hierarchical task collaborative reasoning result.
[0149] In this embodiment, the collaborative unit 114 inputs the dynamic spatiotemporal flow graph and the hierarchical task structure into the hierarchical task collaborative reasoning network to obtain the hierarchical task collaborative reasoning results, including:
[0150] Obtain the high-level task coordination network in the hierarchical task collaborative reasoning network, and the low-level subtask execution network corresponding to each subtask;
[0151] The dynamic spatiotemporal flow graph is encoded to obtain encoded features;
[0152] The encoded features and the hierarchical task structure are input into the high-level task coordination network, and the association between the sub-objectives and spatiotemporal information of each sub-task is learned based on the multi-head self-attention mechanism to obtain the cooperative strategy between each currently executed sub-task and other sub-tasks.
[0153] The sub-objectives of each subtask and the corresponding multimodal initial dynamic features are input into the underlying subtask execution network corresponding to each subtask to obtain the execution action sequence of each subtask.
[0154] The hierarchical task collaborative reasoning result is generated based on the collaborative strategy and the execution action sequence of each sub-task.
[0155] The high-level task coordination network can be built based on the Transformer network.
[0156] The coordination strategy can clearly define the execution order of each subtask. For example, based on the positional changes of the package in the dynamic spatiotemporal flow diagram and the working state of the robotic arm, the execution order of the "path planning" subtask and the "robotic arm grasping and placing" subtask can be coordinated.
[0157] Specifically, a low-level subtask execution network can be pre-built for each subtask. This low-level subtask execution network can be a deep learning model, such as a visual object detection network or a motion reinforcement learning policy network. For example, the execution network for the "robotic arm grasping and placing" subtask can determine the package position based on visual features, plan the robotic arm's motion trajectory based on motion features, and finally output joint angle control commands.
[0158] For example, in the "surgical robot-assisted surgery" task in a medical setting, the high-level network coordinates sub-tasks such as "instrument preparation", "lesion localization" and "surgical operation", while the low-level execution networks, such as the "lesion localization" network, determine the location of the lesion based on medical image features, and the "surgical operation" network controls the robotic arm to perform precise surgical operations. Thus, the surgical accuracy and efficiency are improved through multi-task decomposition and collaboration.
[0159] In the above embodiments, the high-level task coordination network controls the start and pause of subtasks and dynamically adjusts the coordination strategy based on feedback information during the execution of subtasks; the low-level subtask execution network is responsible for the execution of specific actions, realizing hierarchical task collaborative execution driven by multimodal information.
[0160] In this embodiment, after obtaining the hierarchical task collaborative reasoning result, the execution action sequence of each subtask is scheduled according to the collaborative strategy to collaboratively execute the execution action sequence of each subtask.
[0161] Through the above embodiments, efficient collaboration between subtasks and precise execution of specific actions are achieved, improving the execution efficiency and accuracy of complex tasks and enabling the model to better meet the needs of complex task processing.
[0162] In this embodiment, multi-dimensional feedback information during task execution can also be collected, including sub-task completion status (e.g., whether the package was successfully scanned, whether the robotic arm accurately grasped it), action execution results (e.g., trajectory deviation, execution time), and environmental state changes (e.g., new packages, equipment malfunctions). Simultaneously, the changes in the dynamic spatiotemporal flow graph during task execution and the effectiveness of the hierarchical task collaboration strategy are recorded. Furthermore, the feedback information is combined with the original multimodal dynamic spatiotemporal features, dynamic spatiotemporal flow graph, and hierarchical task structure, and used as training data to optimize the corresponding model. The parameters of dynamic spatiotemporal feature extraction, dynamic spatiotemporal flow graph construction, and hierarchical task collaboration network are updated using the backpropagation algorithm. Based on the feedback, the task decomposition strategy, sub-task goal setting method, and collaborative inference logic are adjusted, enabling the model to process multimodal dynamic information and complex tasks more efficiently in subsequent tasks.
[0163] As can be seen from the above technical solutions, this invention can process multimodal data to obtain initial dynamic features of multimodality, thereby providing high-quality dynamic spatiotemporal features; construct a dynamic spatiotemporal flow graph based on the initial dynamic features of multimodality to provide various relationships between key events in multimodal data, reflecting the dynamic evolution of multimodal data; combine the dynamic spatiotemporal flow graph to detect the complexity of the target task, and use a hierarchical task network to decompose the high-complexity target task into multiple sub-tasks. Based on the dynamic spatiotemporal flow graph and multiple sub-tasks, a hierarchical task structure is constructed, which can reasonably decompose the macro-level task target into executable micro-level sub-tasks; input the dynamic spatiotemporal flow graph and hierarchical task structure into a hierarchical task collaborative reasoning network to obtain hierarchical task collaborative reasoning results, which can help achieve efficient collaboration between sub-tasks and accurate execution of actions, improve the execution efficiency and accuracy of complex tasks, and better meet the needs of complex task processing.
[0164] like Figure 3The diagram shown is a schematic representation of the structure of a computer device that implements a preferred embodiment of the hierarchical task collaboration method based on dynamic spatiotemporal flow according to the present invention.
[0165] The computer device 1 may include a memory 12, a processor 13, and a bus (the arrow in the figure represents the bus), and may also include a computer program stored in the memory 12 and executable on the processor 13, such as a hierarchical task coordinating program based on dynamic spatiotemporal flow.
[0166] Those skilled in the art will understand that the schematic diagram is merely an example of computer device 1 and does not constitute a limitation on computer device 1. Computer device 1 can be either a bus topology or a star topology. Computer device 1 may also include more or fewer other hardware or software than shown in the diagram, or different component arrangements. For example, computer device 1 may also include input / output devices, network access devices, etc.
[0167] It should be noted that the computer device 1 described is merely an example. Other existing or future electronic products that are adaptable to this invention should also be included within the scope of protection of this invention and are incorporated herein by reference.
[0168] The memory 12 includes at least one type of readable storage medium, such as flash memory, portable hard drive, multimedia card, card-type memory (e.g., SD or DX memory), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 12 can be an internal storage unit of the computer device 1, such as a portable hard drive of the computer device 1. In other embodiments, the memory 12 can be an external storage device of the computer device 1, such as a plug-in portable hard drive, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the computer device 1. Furthermore, the memory 12 can include both internal storage units and external storage devices of the computer device 1. The memory 12 can be used not only to store application software and various types of data installed on the computer device 1, such as the code of a hierarchical task co-operation program based on dynamic spatiotemporal flow, but also to temporarily store data that has been output or will be output.
[0169] In some embodiments, the processor 13 may be composed of integrated circuits, such as a single packaged integrated circuit or multiple integrated circuits packaged with the same or different functions, including combinations of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 13 is the control unit of the computer device 1, connecting various components of the computer device 1 via various interfaces and lines. It executes programs or modules stored in the memory 12 (e.g., executing hierarchical task co-processing programs based on dynamic spatiotemporal flow) and calls data stored in the memory 12 to perform various functions of the computer device 1 and process data.
[0170] The processor 13 executes the operating system of the computer device 1 and various installed applications. The processor 13 executes these applications to implement the steps in the various embodiments of the hierarchical task collaboration method based on dynamic spatiotemporal flow described above, for example... Figure 1 The steps are shown.
[0171] For example, the computer program may be divided into one or more modules / units, which are stored in the memory 12 and executed by the processor 13 to complete the present invention. The one or more modules / units may be a series of computer-readable instruction segments capable of performing a specific function, which describe the execution process of the computer program in the computer device 1. For example, the computer program may be divided into a processing unit 110, a construction unit 111, a detection unit 112, a splitting unit 113, and a coordination unit 114.
[0172] The integrated unit implemented as a software functional module described above can be stored in a computer-readable storage medium. This software functional module, stored in a storage medium, includes several instructions to cause a computer device (which may be a personal computer, a computer device, or a network device, etc.) or processor to execute portions of the hierarchical task collaboration method based on dynamic spatiotemporal flow described in the various embodiments of this invention.
[0173] If the modules / units integrated in the computer device 1 are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can also be implemented by a computer program instructing related hardware devices. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above.
[0174] The computer program includes computer program code, which may be in the form of source code, object code, executable file, or some intermediate form. The computer-readable medium may include any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory, etc.
[0175] Furthermore, the computer-readable storage medium may primarily include a stored program area and a stored data area, wherein the stored program area may store the operating system, an application program required for at least one function, etc.; and the stored data area may store data created based on the use of blockchain nodes, etc.
[0176] The blockchain referred to in this invention is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include an underlying blockchain platform, a platform product service layer, and an application service layer.
[0177] The bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into address bus, data bus, control bus, etc. For ease of representation, in... Figure 3 The bus is represented by only one straight line, but this does not mean that there is only one bus or one type of bus. The bus is configured to enable communication between the memory 12 and at least one processor 13, etc.
[0178] Although not shown, the computer device 1 may also include a power supply (such as a battery) to power various components. Preferably, the power supply can be logically connected to the at least one processor 13 through a power management device, thereby enabling functions such as charging management, discharging management, and power consumption management. The power supply may also include one or more DC or AC power supplies, recharging devices, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components. The computer device 1 may also include various sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be described in detail here.
[0179] Furthermore, the computer device 1 may also include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a Wi-Fi interface, a Bluetooth interface, etc.), which is typically used to establish a communication connection between the computer device 1 and other computer devices.
[0180] Optionally, the computer device 1 may further include a user interface, which may be a display, an input unit (such as a keyboard), and optionally, a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen, etc. The display may also be appropriately referred to as a screen or display unit, used to display information processed in the computer device 1 and to display a visual user interface.
[0181] It should be understood that the embodiments described are for illustrative purposes only and are not limited to this structure in the scope of the patent application.
[0182] It will be understood by those skilled in the art that Figure 3 The structure shown does not constitute a limitation on the computer device 1, and may include fewer or more components than shown, or combine certain components, or have different component arrangements.
[0183] Combination Figure 1 The memory 12 in the computer device 1 stores multiple instructions to implement a hierarchical task collaboration method based on dynamic spatiotemporal flow, and the processor 13 can execute the multiple instructions to achieve:
[0184] In response to a hierarchical task collaboration instruction based on a target task, multimodal data is collected and processed to obtain initial dynamic features of the multimodal data.
[0185] A dynamic spatiotemporal flow graph is constructed based on the aforementioned multimodal initial dynamic features;
[0186] The complexity of the target task is detected by combining the dynamic spatiotemporal flow graph.
[0187] When the target task is determined to be a complex task based on the complexity, the target task is divided into multiple subtasks using a hierarchical task network.
[0188] A hierarchical task structure is constructed based on the dynamic spatiotemporal flow graph and the multiple sub-tasks;
[0189] Obtain a pre-constructed hierarchical task collaborative reasoning network, and input the dynamic spatiotemporal flow graph and the hierarchical task structure into the hierarchical task collaborative reasoning network to obtain the hierarchical task collaborative reasoning result.
[0190] Specifically, the processor 13's implementation method for the above instructions can be found in [reference needed]. Figure 1 The descriptions of the relevant steps in the corresponding embodiments are not repeated here.
[0191] It should be noted that all data involved in this case was legally obtained. Software tools or components not belonging to this company that appear in the embodiments of this application are merely illustrative examples and do not represent actual use.
[0192] In the several embodiments provided by this invention, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.
[0193] This invention can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This invention can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This invention can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0194] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0195] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.
[0196] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0197] Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be embraced within the invention. No appended diagram markings in the claims should be construed as limiting the scope of the claims.
[0198] Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices described in this invention can also be implemented by a single unit or device through software or hardware. Terms such as "first," "second," etc., are used to indicate names and do not indicate any specific order.
[0199] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A hierarchical task collaboration method based on dynamic spatiotemporal flow, characterized in that, The hierarchical task collaboration method based on dynamic spatiotemporal flow includes: In response to a hierarchical task collaboration instruction based on a target task, multimodal data is collected and processed to obtain initial dynamic features of the multimodal data. A dynamic spatiotemporal flow graph is constructed based on the aforementioned multimodal initial dynamic features; The complexity of the target task is detected by combining the dynamic spatiotemporal flow graph. When the target task is determined to be a complex task based on the complexity, the target task is divided into multiple subtasks using a hierarchical task network. A hierarchical task structure is constructed based on the dynamic spatiotemporal flow graph and the multiple sub-tasks; Obtain a pre-constructed hierarchical task collaborative reasoning network, and input the dynamic spatiotemporal flow graph and the hierarchical task structure into the hierarchical task collaborative reasoning network to obtain the hierarchical task collaborative reasoning result.
2. The hierarchical task collaboration method based on dynamic spatiotemporal flow as described in claim 1, characterized in that, The process of processing the multimodal data to obtain the initial dynamic features of the multimodal data includes: Extract video stream data, language text sequence data, and action time series data from the multimodal data; The Swin Transformer 3D model is used to extract local image features of the video stream data in the spatial dimension using a sliding window mechanism; a temporal self-attention mechanism is used to capture object motion features and scene change features between consecutive frames in the video stream data in the temporal dimension; an optical flow estimation algorithm is combined to calculate the motion vectors of pixels between frames in the video stream data; and the image local features, object motion features and scene change features between consecutive frames in the video stream data, and motion vectors of pixels between frames in the video stream data are combined to obtain a visual dynamic spatiotemporal feature vector. Using a pre-trained language model, the language text sequence data is encoded based on a temporal attention mechanism to obtain a dynamic semantic feature vector of the language; Temporal features are extracted from the action time series data using a temporal convolutional network, and action features are analyzed based on the temporal features; the action features are then transformed into a dynamic spatiotemporal feature vector containing spatiotemporal information based on the spatial coordinates of the action execution recorded in the action time series data. The visual dynamic spatiotemporal feature vector, the language dynamic semantic feature vector, and the action dynamic spatiotemporal feature vector are integrated to obtain the multimodal initial dynamic features.
3. The hierarchical task collaboration method based on dynamic spatiotemporal flow as described in claim 1, characterized in that, The construction of the dynamic spatiotemporal flow graph based on the multimodal initial dynamic features includes: Identify events in the initial dynamic features of the multimodal model as nodes, and add spatiotemporal labels to each node; Identify the spatiotemporal sequence, causal relationships, and logical dependencies among the events as edges; Through the message passing mechanism, each node absorbs the spatiotemporal information of its neighboring nodes to update its own feature representation; After the message passing is completed, the currently obtained flow graph is determined as the dynamic spatiotemporal flow graph.
4. The hierarchical task collaboration method based on dynamic spatiotemporal flow as described in claim 3, characterized in that, The step of constructing a hierarchical task structure based on the dynamic spatiotemporal flow graph and the multiple sub-tasks includes: Obtain the spatiotemporal relationships and task logic in the dynamic spatiotemporal flow graph; Configure sub-goals for each subtask; Based on the spatiotemporal relationship and the task logic, the sub-objectives of each sub-task are prioritized to obtain a hierarchical task structure.
5. The hierarchical task collaboration method based on dynamic spatiotemporal flow as described in claim 4, characterized in that, The sub-goals configured for each sub-task include: Based on the spatiotemporal sequence, causal relationships, and logical dependencies between the events, as well as the spatiotemporal relationships and task logic in the dynamic spatiotemporal flow graph, configure the sub-goals of each sub-task.
6. The hierarchical task collaboration method based on dynamic spatiotemporal flow as described in claim 5, characterized in that, The step of inputting the dynamic spatiotemporal flow graph and the hierarchical task structure into the hierarchical task collaborative reasoning network to obtain the hierarchical task collaborative reasoning result includes: Obtain the high-level task coordination network in the hierarchical task collaborative reasoning network, and the low-level subtask execution network corresponding to each subtask; The dynamic spatiotemporal flow graph is encoded to obtain encoded features; The encoded features and the hierarchical task structure are input into the high-level task coordination network, and the association between the sub-objectives and spatiotemporal information of each sub-task is learned based on the multi-head self-attention mechanism to obtain the cooperative strategy between each currently executed sub-task and other sub-tasks. The sub-objectives of each subtask and the corresponding multimodal initial dynamic features are input into the underlying subtask execution network corresponding to each subtask to obtain the execution action sequence of each subtask. The hierarchical task collaborative reasoning result is generated based on the collaborative strategy and the execution action sequence of each sub-task.
7. The hierarchical task collaboration method based on dynamic spatiotemporal flow as described in claim 6, characterized in that, After obtaining the hierarchical task collaborative reasoning result, the method further includes: The execution action sequence of each subtask is scheduled according to the coordination strategy to coordinate the execution of the execution action sequence of each subtask.
8. A hierarchical task collaboration device based on dynamic spatiotemporal flow, characterized in that, The hierarchical task collaboration device based on dynamic spatiotemporal flow includes: The processing unit is used to respond to hierarchical task collaboration instructions based on the target task, collect multimodal data, and process the multimodal data to obtain initial dynamic features of the multimodality. The construction unit is used to construct a dynamic spatiotemporal flow graph based on the initial dynamic features of the multimodal modes; A detection unit is used to detect the complexity of the target task in conjunction with the dynamic spatiotemporal flow graph; A splitting unit is used to split the target task into multiple subtasks using a hierarchical task network when the target task is determined to be a complex task based on the complexity. The construction unit is also used to construct a hierarchical task structure based on the dynamic spatiotemporal flow graph and the multiple subtasks; The collaborative unit is used to acquire a pre-constructed hierarchical task collaborative reasoning network, and input the dynamic spatiotemporal flow graph and the hierarchical task structure into the hierarchical task collaborative reasoning network to obtain the hierarchical task collaborative reasoning result.
9. A computer device, characterized in that, The computer device includes: Memory, storing at least one instruction; and The processor executes instructions stored in the memory to implement the hierarchical task collaboration method based on dynamic spatiotemporal flow as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one instruction, which is executed by a processor in a computer device to implement the hierarchical task collaboration method based on dynamic spatiotemporal flow as described in any one of claims 1 to 7.