Multi-mode task arrangement and intelligent execution method and system of intelligent car

Through deep learning and reinforcement learning, the task model library and decision tree optimization method are constructed, and the problem of inefficient task execution of smart cars in complex environments is solved, and the efficient and adaptive task execution of smart cars in multi-task scenarios is achieved.

CN120276428APending Publication Date: 2025-07-08BEIJING SUBCUBIC TECH CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510194748.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing smart car task execution system lacks effective utilization of historical task data, cannot dynamically adapt to environmental changes, and is difficult to achieve intelligent scheduling and optimization of multi-task scenarios, resulting in inefficient task execution.

Method used

The task model library is built through deep learning algorithms, and the deep reinforcement learning model is used for online optimization, combining heuristic algorithms and decision tree pruning optimization, generating the optimal task execution path, and adjusting control instructions in real time to adapt to environmental changes.

Benefits of technology

It improves the task execution efficiency and adaptability of smart cars, enhances independent decision-making capabilities, achieves continuous learning and performance improvement, and can complete tasks efficiently in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120276428A_ABST
    Figure CN120276428A_ABST
Patent Text Reader

Abstract

The invention provides a multi-mode task arrangement and intelligent execution method and system for an intelligent car, and relates to the technical field of task arrangement, and the method comprises the steps: carrying out the clustering analysis of historical task data through a deep learning algorithm, and constructing a task model library; based on a new task instruction matching execution mode, constructing and optimizing a decision tree, and generating an optimal execution path; converting the execution path into a control instruction sequence; environment data are collected in real time when a task is executed, and online optimization is carried out based on a deep reinforcement learning model; the execution data is fed back to the task model library, and the follow-up execution efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to task scheduling technology, and in particular to a multi-mode task scheduling and intelligent execution method and system for intelligent vehicles. Background Art

[0002] Existing intelligent vehicle task execution systems lack effective utilization of historical task data. They can usually only execute simple predefined tasks and cannot learn from past experiences to optimize task execution strategies, resulting in low task execution efficiency and difficulty in adapting to complex and changing working environments.

[0003] Secondly, existing technologies lack dynamic adaptability during task planning and execution. Most systems adopt static task execution plans, which are difficult to adjust according to real-time environmental changes once formulated. This rigid execution method is prone to task execution failures or low efficiency when facing emergencies or environmental changes.

[0004] Finally, existing intelligent vehicle systems lack effective management of multi-task scenarios. They often can only execute single tasks sequentially according to a preset order, cannot perform intelligent scheduling based on task priorities and time constraints, and are also difficult to achieve collaborative optimization between multiple tasks, which severely limits the practicality of intelligent vehicles in complex application scenarios. Summary of the Invention

[0005] Embodiments of the present invention provide a multi-mode task scheduling and intelligent execution method and system for intelligent vehicles, which can solve the problems in the prior art.

[0006] In the first aspect of the embodiments of the present invention,

[0007] A multi-mode task scheduling and intelligent execution method for an intelligent vehicle is provided, including:

[0008] Obtaining historical task data of the intelligent vehicle, where the historical task data includes task type, task path, task duration, environmental parameters, and task completion status; performing clustering analysis on the historical task data based on a deep learning algorithm, classifying tasks with similar characteristics into the same task mode, and constructing a task model library containing multiple preset task modes; receiving a new task instruction, where the task instruction includes a task target location, a task execution time, a task priority, and environmental constraint conditions; performing feature matching between the task instruction and the preset task modes in the task model library to determine the execution mode of the current task;

[0009] Construct a task execution decision tree based on the execution mode. Each node of the decision tree contains an action sequence, resource consumption, and time constraints. Use a heuristic algorithm to prune and optimize the decision tree, removing execution paths that do not meet the environmental constraint conditions. Calculate the task switching cost of each node in the decision tree according to the task priority and time constraints, and generate an optimal task execution path. Convert the optimal task execution path into a control instruction sequence for the intelligent vehicle. The control instruction sequence includes a motion trajectory, speed planning, and execution timing.

[0010] The intelligent vehicle starts to execute the task according to the control instruction sequence. During the task execution process, environmental data is collected in real time. The environmental data includes obstacle information, light intensity, temperature, and road condition status. When it is detected that the environment has changed, the control instruction sequence is optimized online based on a deep reinforcement learning model to generate new control instructions adapted to the environmental changes. The state data, environmental data, and optimization results generated during the task execution are fed back to the task model library to update and improve the preset task mode and improve the execution efficiency of subsequent tasks.

[0011] Based on a deep learning algorithm, perform clustering analysis on the historical task data, and classify tasks with similar characteristics into the same task mode. Constructing a task model library containing multiple preset task modes includes:

[0012] Obtain the historical task data of the intelligent vehicle, and construct a task feature vector through a multi-dimensional feature extraction method. The multi-dimensional feature extraction method includes: performing one-hot encoding on the task type to convert it into a numerical feature, extracting the spatial features of the task path to obtain the path length, number of turns, and key node distribution parameters, performing time series analysis on the task duration to obtain the time distribution feature, standardizing the environmental parameters to obtain the numerical representations of light intensity, temperature, and humidity, and converting the task completion status into a task completion rate, task on-time rate, and task quality score.

[0013] Use a deep autoencoder network to perform dimensionality reduction processing on the task feature vector. The number of input layer nodes of the deep autoencoder network is the same as the dimension of the task feature vector. Compress the high-dimensional features to a low-dimensional latent space through multiple hidden layers. Introduce a task correlation loss function during the network training process to ensure that the distances of similar tasks in the latent space are relatively close. Use a density peak clustering algorithm in the low-dimensional latent space to adaptively determine the number of clustering centers to form task clusters.

[0014] Feature extraction and pattern summarization are performed on each of the task clusters, the central vector of the task cluster is calculated as the typical feature representation of this type of task, the path pattern, time window, and resource requirements of the samples within the task cluster are extracted to form a standardized task pattern description, an evaluation index system including execution efficiency, resource utilization rate, and adaptability is established, and the task pattern description and the evaluation index system are stored in the task model library;

[0015] When a new task is received, feature extraction is performed on the new task to obtain a new task feature vector, the similarity between the new task feature vector and the central vectors of each task pattern is calculated, the task pattern with the highest similarity is selected as the reference pattern, and the reference pattern is adjusted in combination with the current environmental constraint conditions to generate an execution plan, and the execution result is fed back to the task model library for incremental update.

[0016] Perform feature matching between the task instruction and the preset task patterns in the task model library to determine the execution mode of the current task, including:

[0017] Receive a task instruction and construct a feature vector. Perform spatial feature encoding on the task target location information to obtain target point coordinates, regional attributes, and spatial constraint conditions. Perform temporal feature analysis on the task execution time to obtain a time window, duration, and time key point information. Convert the task priority into a priority matrix. Perform parameterization processing on the environmental constraint conditions to obtain environmental complexity, dynamic obstacle distribution, and weather condition parameters. Form a task instruction feature vector through feature normalization processing;

[0018] Establish a hierarchical feature matching framework, perform multi-dimensional matching between the task instruction feature vector and the preset task patterns, calculate the matching degree of the task type dimension through cosine similarity, calculate the similarity of the path features using the dynamic time warping algorithm, analyze the matching degree of the execution time using the temporal pattern matching algorithm, and determine the matching degree of the environmental constraints through calculating the environmental parameter similarity;

[0019] Construct a feature importance evaluation model based on historical task execution data, calculate the contribution degree of each dimension feature to the task execution effect, construct an adaptive learning model in combination with the current task characteristics and environmental conditions, and dynamically adjust and online optimize the weight coefficients of each dimension feature;

[0020] Calculate the weighted matching scores of each preset task pattern, where the weighted matching score is the weighted sum of the matching results of each dimension feature and the corresponding weight coefficients. Establish a mode switching cost model to evaluate the resource consumption of state switching, and construct a decision optimization model in combination with the task priority and time constraint;

[0021] When there are multiple candidate execution modes, a fuzzy decision-making method is adopted to comprehensively evaluate the weighted matching score, mode switching cost, and task priority, select the optimal execution mode, and collect execution status data in real time during task execution to evaluate the mode matching accuracy. The matching result and evaluation data are fed back to the feature importance evaluation model for update and optimization.

[0022] Based on the execution mode, a task execution decision tree is constructed. Each node of the decision tree includes an action sequence, resource consumption, and time constraint. A heuristic algorithm is used to prune and optimize the decision tree, and the execution paths that do not meet the environmental constraint conditions are removed, including:

[0023] A multi-layer decision tree is constructed according to the task execution mode. The root node of the decision tree stores the task start state information, the intermediate nodes store the action sequence information, resource consumption information, and time constraint information, the leaf nodes store the target state information, and the connection edges between the nodes represent the state transition relationship.

[0024] Environmental perception data is obtained, including static obstacle position information, dynamic obstacle trajectory information, environmental light information, and temperature information. An environmental constraint model is established, and the environmental constraint model is decomposed into spatial constraint conditions and time constraint conditions.

[0025] A heuristic evaluation function is constructed. The heuristic evaluation function includes a path length evaluation term, an action coherence evaluation term, a resource consumption evaluation term, and a safety margin evaluation term, and each execution path in the decision tree is evaluated and scored.

[0026] A pruning threshold is set, and the decision tree is hierarchically pruned based on the environmental constraint model and the heuristic evaluation function. First, the execution paths that violate the spatial constraint conditions are removed, second, the execution paths that violate the time constraint conditions are removed, and third, the execution paths with evaluation scores lower than the pruning threshold are removed.

[0027] Path optimization is performed on the pruned decision tree. The comprehensive cost of each feasible execution path is calculated. The comprehensive cost is weighted and calculated by combining the path length, the number of actions, the resource consumption, and the safety margin. The execution path with the minimum comprehensive cost is selected as the optimal execution path.

[0028] During the task execution process, environmental change information is obtained in real time. When it is detected that the environmental constraint conditions change, the optimal execution path is adjusted online, and the decision tree pruning and path optimization are performed again to ensure that the execution path continuously meets the dynamic environmental constraint conditions.

[0029] According to the task priority and time constraint, the task switching cost of each node of the decision tree is calculated, and the optimal task execution path is generated, including:

[0030] Obtain task priority information and time constraint information, divide the task priority into multiple levels and assign different weight coefficients, and convert the time constraint information into time window parameters and deadline parameters;

[0031] Construct a task switching cost calculation model, which includes state transition cost, resource reconfiguration cost, and time delay cost. Among them, the state transition cost is calculated based on the state difference between two task nodes, the resource reconfiguration cost is calculated based on the change in resource occupancy, and the time delay cost is calculated based on the matching degree between the task execution time and the time window;

[0032] Calculate the task switching cost between adjacent nodes in the decision tree. First, calculate the state transition matrix between the nodes, and calculate the state transition cost based on the state transition matrix. Secondly, calculate the change in resource configuration between the nodes to obtain the resource reconfiguration cost, and then calculate the overlap degree between the task execution time and the time window to obtain the time delay cost;

[0033] Establish a path optimization objective function, which comprehensively considers the total task switching cost, task priority weight, and time constraint satisfaction degree. Use the dynamic programming algorithm to traverse the decision tree and calculate the cumulative cost from the root node to each leaf node;

[0034] According to the calculation result of the path optimization objective function, select the path with the minimum cumulative cost as the candidate optimal path, and perform time constraint verification on the candidate optimal path to ensure that all task nodes on the path meet their respective deadline constraints;

[0035] During the task execution process, monitor the task state in real time. When it is detected that the task priority changes or the time constraint is adjusted, recalculate the task switching cost of the affected nodes and update the optimal execution path to achieve dynamic optimization and adjustment of the path.

[0036] When it is detected that the environment changes, online optimize the control instruction sequence based on the deep reinforcement learning model to generate new control instructions adapted to the environmental changes, including:

[0037] Construct a deep reinforcement learning model, which includes a state space, an action space, and a reward function. The state space includes environmental state information and task execution state information. The action space includes a set of optional control instructions. The reward function is designed based on the execution effect of the control instructions and environmental adaptability;

[0038] Obtain real-time environmental data through the environmental perception module, parse the environmental data into obstacle information, light information, and temperature information, construct an environmental feature vector, and combine the environmental feature vector with the task execution state information to form a state input;

[0039] Adopt a dual-network structure, including a policy network and a value network. The policy network is responsible for generating a sequence of control instructions, and the value network is responsible for evaluating the execution effect of the control instructions. The two networks use a convolutional layer with shared parameters to extract state features;

[0040] Based on the experience replay mechanism, store historical interaction data, including state transition sequences, control instruction sequences, and reward sequences. Randomly sample training data batches from the experience pool, use the temporal difference algorithm to update the parameters of the value network, and use the policy gradient algorithm to optimize the parameters of the policy network;

[0041] Set an environmental change detection threshold. When the change amplitude of the environmental feature vector exceeds the detection threshold, trigger online optimization. Input the current state into the deep reinforcement learning model, and generate multiple candidate control instruction sequences through the policy network;

[0042] Conduct simulation evaluation on the candidate control instruction sequences, calculate the expected reward value of each sequence in the new environment, select the control instruction sequence with the highest expected reward value as the optimization result, and decompose the optimized control instruction sequence into individual control instructions for step-by-step execution.

[0043] In the second aspect of the embodiments of the present invention,

[0044] Provide a multi-mode task scheduling and intelligent execution system for an intelligent vehicle, including:

[0045] The first unit is used to obtain the historical task data of the intelligent vehicle. The historical task data includes task type, task path, task duration, environmental parameters, and task completion status; perform clustering analysis on the historical task data based on deep learning algorithms, classify tasks with similar characteristics into the same task mode, and construct a task model library containing multiple preset task modes; receive a new task instruction, and the task instruction includes a task target location, a task execution time, a task priority, and environmental constraint conditions; perform feature matching between the task instruction and the preset task modes in the task model library to determine the execution mode of the current task;

[0046] The second unit is used to construct a task execution decision tree based on the execution mode. Each node of the decision tree includes an action sequence, resource consumption, and time constraints; use a heuristic algorithm to prune and optimize the decision tree, and remove execution paths that do not meet the environmental constraint conditions; calculate the task switching cost of each node of the decision tree according to the task priority and time constraints, and generate an optimal task execution path; convert the optimal task execution path into a control instruction sequence of the intelligent vehicle, and the control instruction sequence includes a motion trajectory, speed planning, and execution timing;

[0047] The third unit is used for the intelligent vehicle to start executing tasks according to the control instruction sequence, and to collect environmental data in real time during the task execution. The environmental data includes obstacle information, light intensity, temperature, and road condition status. When it detects that the environment has changed, it online optimizes the control instruction sequence based on the deep reinforcement learning model to generate new control instructions adapted to the environmental changes. It feeds back the state data, environmental data, and optimization results generated during the task execution to the task model library for updating and improving the preset task patterns and enhancing the execution efficiency of subsequent tasks.

[0048] In the third aspect of the embodiments of the present invention,

[0049] a kind of electronic device is provided, including:

[0050] a processor;

[0051] a memory for storing instructions executable by the processor;

[0052] wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.

[0053] In the fourth aspect of the embodiments of the present invention,

[0054] a computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.

[0055] The beneficial effects of this application are as follows:

[0056] 1. Improve the intelligence and adaptability of task execution:

[0057] This method performs clustering analysis on historical task data through deep learning algorithms to construct a task model library containing multiple preset task patterns. When executing a new task, it can quickly match a suitable execution pattern and construct a task execution decision tree based on this. This method can make full use of historical experience, improve the intelligence level of task planning, and enable the intelligent vehicle to better adapt to different types of task requirements.

[0058] 2. Optimize task execution efficiency and flexibility:

[0059] The heuristic algorithm is used to prune and optimize the decision tree, and the optimal execution path is calculated considering task priorities and time constraints. This method can effectively remove the execution paths that do not meet the constraint conditions and achieve optimal task switching and resource allocation in the case of multiple tasks. Therefore, this method can significantly improve the task execution efficiency and scheduling flexibility of the intelligent vehicle.

[0060] 3. Achieve adaptive optimization in a dynamic environment:

[0061] During the task execution process, this method collects environmental data in real time and uses a deep reinforcement learning model to optimize the control instruction sequence online. This enables the intelligent vehicle to adjust the execution strategy in a timely manner according to environmental changes, improving the adaptability and robustness of the system in dynamic and complex environments. At the same time, the data during the execution process is fed back to the task model library to continuously update and improve the preset task patterns, achieving the continuous learning and performance improvement of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 is a schematic flowchart of the multi-mode task scheduling and intelligent execution method for the intelligent vehicle according to an embodiment of the present invention;

[0063] Figure 2 is a schematic structural diagram of the multi-mode task scheduling and intelligent execution system for the intelligent vehicle according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0064] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0065] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0066] Figure 1 is a schematic flowchart of the multi-mode task scheduling and intelligent execution method for the intelligent vehicle according to an embodiment of the present invention, as Figure 1 shown, the method includes:

[0067] S101. Obtain the historical task data of the intelligent vehicle, where the historical task data includes task type, task path, task duration, environmental parameters, and task completion status; perform clustering analysis on the historical task data based on a deep learning algorithm, classify tasks with similar characteristics into the same task mode, and construct a task model library containing multiple preset task modes; receive a new task instruction, where the task instruction includes a task target location, a task execution time, a task priority, and environmental constraint conditions; perform feature matching between the task instruction and the preset task modes in the task model library to determine the execution mode of the current task;

[0068] S102. Construct a task execution decision tree based on the execution mode. Each node of the decision tree contains an action sequence, resource consumption, and time constraint. Use a heuristic algorithm to prune and optimize the decision tree, removing execution paths that do not meet the environmental constraint conditions. Calculate the task switching cost of each node of the decision tree according to the task priority and time constraint to generate an optimal task execution path. Convert the optimal task execution path into a control instruction sequence for the intelligent vehicle, and the control instruction sequence includes a motion trajectory, speed planning, and execution timing.

[0069] S103. The intelligent vehicle starts to execute the task according to the control instruction sequence. During the task execution process, it collects environmental data in real time. The environmental data includes obstacle information, light intensity, temperature, and road condition status. When it detects that the environment has changed, it performs online optimization on the control instruction sequence based on the deep reinforcement learning model to generate new control instructions that adapt to the environmental changes. Feed back the state data, environmental data, and optimization results generated during the task execution to the task model library for updating and improving the preset task mode to improve the execution efficiency of subsequent tasks.

[0070] The specific implementation manner of the multi-mode task scheduling and intelligent execution method of the intelligent vehicle is as follows:

[0071] First, obtain the historical task data of the intelligent vehicle. These data include task types (such as transportation, inspection, cleaning, etc.), task paths (such as starting point, ending point, passing points, etc.), task duration, environmental parameters (such as temperature, humidity, light, etc.), and task completion status (success, failure, or partial completion). For example, a piece of historical task data may be: task type - transportation, task path - from warehouse A to workshop B, task duration - 30 minutes, environmental parameters - temperature 25°C, humidity 60%, light 1000 lux, task completion status - success.

[0072] Next, use a deep learning algorithm to perform clustering analysis on the historical task data. The autoencoder combined with the K-means algorithm can be used for clustering. First, use the autoencoder to reduce the dimension of the task data and extract key features. Then, use the K-means algorithm to cluster the data after dimension reduction, and group similar tasks into the same category. Through multiple iterations and adjustment of the number of clusters, a series of task modes are finally obtained. For example, multiple preset task modes such as "short-distance light-load transportation mode", "long-distance heavy-load transportation mode", "indoor inspection mode", and "outdoor inspection mode" may be obtained to construct a task model library.

[0073] When a new task instruction is received, first parse the instruction content to extract information such as the task target location, execution time, priority, and environmental constraints. For example, a new task instruction might be: target location - production line C, execution time - 9:00 - 10:00, priority - high, environmental constraint - avoid crowded areas. Then, perform feature matching between this information and the preset task patterns in the task model library. Algorithms such as cosine similarity can be used to calculate the similarity between the new task and each preset pattern, and select the pattern with the highest similarity as the execution pattern for the current task.

[0074] After determining the execution pattern, construct a task execution decision tree. The root node of the decision tree is the starting state, the leaf nodes are the target states, and the intermediate nodes represent different execution stages. Each node contains an action sequence (such as going straight, turning, accelerating, decelerating, etc.), resource consumption (such as power, time, etc.), and time constraints. For example, the information contained in a node might be: action sequence - go straight 50 meters and then turn right, resource consumption - 2% power, time constraint - complete within 5 minutes.

[0075] Next, use a heuristic algorithm to prune and optimize the decision tree. First, check whether each node meets the environmental constraint conditions, and if not, prune the node and its subtrees. For example, if the path of a certain node passes through a crowded area, violating the environmental constraint, then prune that node. Then, calculate the heuristic values of the remaining nodes, and factors such as path length, time consumption, and safety factor can be comprehensively considered. Select the path with the optimal heuristic value to further optimize the decision tree.

[0076] According to the task priority and time constraints, calculate the task switching cost for each node of the decision tree. The task switching cost can consider factors such as path change, resource reallocation, and time delay. For example, if switching from one node to another requires changing the path, re-planning the power usage, and may cause task delay, then the cost of this switch is relatively high. Through the dynamic programming algorithm, comprehensively consider the heuristic values and switching costs of each node to generate the optimal task execution path.

[0077] Convert the optimal task execution path into a control instruction sequence for the intelligent vehicle. The control instruction sequence includes motion trajectories (such as specific parameters for going straight and turning), speed planning (such as the timing and amplitude of accelerating, decelerating, and driving at a constant speed), and execution timings (the execution order and time points of each instruction). For example, a section of the control instruction sequence might be: t = 0s, start the motor; t = 2s, accelerate to 1m / s; t = 10s, keep going straight at 1m / s; t = 20s, turn right 30°; t = 25s, decelerate to 0.5m / s, etc.

[0078] The intelligent vehicle starts to execute tasks according to the control instruction sequence. During the execution process, the vehicle collects environmental data in real time through various sensors. These data include obstacle information obtained by lidar or camera, light intensity obtained by photoresistors, temperature data obtained by temperature sensors, road conditions obtained by acceleration sensors, etc. For example, the data collected may be: there is a moving obstacle 5 meters ahead, the current light intensity is 800 lux, the temperature is 28 °C, and the road bumpiness is medium.

[0079] When a significant change in the environment is detected, such as the appearance of an unexpected obstacle or a sudden change in light intensity, the online optimization of the control instruction sequence is triggered. Using a deep reinforcement learning model, such as Deep Q-Network (DQN), based on the current state and environmental data, the optimal action is predicted. The input of the model is the current state and environmental data, and the output is the Q-values of different actions. The action with the highest Q-value is selected as the new control instruction. For example, when an obstacle is detected ahead, the model may output a new instruction of "decelerate and detour to the right".

[0080] Finally, the state data (such as position, speed, power, etc.), environmental data, and optimization results generated during the task execution process are fed back to the task model library. Using an incremental learning algorithm, such as online stochastic gradient descent, the parameters of the task model are updated. For example, if a certain task model performs poorly in a new environment, its parameters can be adjusted or new branches can be added. By continuous learning and optimization, the adaptability and accuracy of the preset task model are improved, thereby enhancing the execution efficiency of subsequent tasks.

[0081] The beneficial effects of this method include the following three aspects:

[0082] 1. It improves the task execution efficiency and adaptability of the intelligent vehicle. By constructing a multi-mode task model library and using deep learning algorithms for task matching, it can quickly select a suitable execution mode for new tasks, reducing the task planning time. At the same time, using deep reinforcement learning for real-time optimization enables the vehicle to flexibly respond to environmental changes and improves the success rate of task completion.

[0083] 2. It enhances the autonomous decision-making ability of the intelligent vehicle. The task execution path generation method based on decision trees and heuristic algorithms enables the vehicle to autonomously plan the optimal path in complex environments. By considering various factors such as task priority, time constraints, and environmental limitations, a more intelligent and reasonable decision-making process is achieved.

[0084] 3. The continuous learning and performance improvement of the intelligent vehicle are realized. By feeding back the data during the task execution process to the task model library and using the incremental learning algorithm to continuously update and improve the preset task patterns, the system can learn from experience and gradually improve the accuracy and efficiency of task execution. This continuous learning mechanism enables the performance of the intelligent vehicle to continuously improve over time.

[0085] In an optional implementation manner, based on the deep learning algorithm, clustering analysis is performed on the historical task data, and tasks with similar characteristics are classified into the same task pattern. The construction of the task model library including multiple preset task patterns includes:

[0086] Obtain the historical task data of the intelligent vehicle, and construct a task feature vector through a multi-dimensional feature extraction method. The multi-dimensional feature extraction method includes: performing one-hot encoding on the task type to convert it into a numerical feature, extracting the spatial features of the task path to obtain the path length, the number of turns, and the key node distribution parameters, performing time series analysis on the task duration to obtain the time distribution feature, performing standardization processing on the environmental parameters to obtain the numerical representations of the light intensity, temperature, and humidity, and converting the task completion status into the task completion rate, task on-time rate, and task quality score;

[0087] Use a deep autoencoder network to perform dimensionality reduction processing on the task feature vector. The number of input layer nodes of the deep autoencoder network is the same as the dimension of the task feature vector. The high-dimensional features are compressed into a low-dimensional latent space through multiple hidden layers. During the network training process, a task correlation loss function is introduced to ensure that the distances of similar tasks in the latent space are relatively close. In the low-dimensional latent space, the density peak clustering algorithm is used to adaptively determine the number of clustering centers to form task clusters;

[0088] Perform feature extraction and pattern summary on each task cluster, calculate the center vector of the task cluster as the typical feature representation of this type of task, extract the path pattern, time window, and resource requirements of the samples within the task cluster to form a standardized task pattern description, establish an evaluation index system including execution efficiency, resource utilization rate, and adaptability, and store the task pattern description and the evaluation index system in the task model library;

[0089] When a new task is received, perform feature extraction on the new task to obtain a new task feature vector, calculate the similarity between the new task feature vector and the center vectors of each task pattern, select the task pattern with the highest similarity as the reference pattern, and adjust the reference pattern in combination with the current environmental constraint conditions to generate an execution plan, and feed back the execution result to the task model library for incremental update.

[0090] This implementation manner provides a method for constructing an intelligent vehicle task pattern based on deep learning, and the specific steps are as follows:

[0091] First, obtain the historical task data of the intelligent vehicle, including information such as task type, task path, task duration, environmental parameters, and task completion status. Perform multi-dimensional feature extraction on this raw data to construct a task feature vector:

[0092] Perform one-hot encoding on the task type. For example, encode task types such as "delivery", "inspection", "cleaning" into numerical features such as [1, 0, 0], [0, 1, 0], [0, 0, 1], etc.

[0093] Extract the spatial features of the task path, and calculate the path length (such as 5000 meters), the number of turns (such as 10 times), and the key node distribution parameters (such as [0.2, 0.5, 0.8] representing the relative positions of 3 key nodes on the path).

[0094] Perform time series analysis on the task duration to obtain time distribution features, such as the distribution of task start times (60% from 8:00 - 10:00, 40% from 10:00 - 12:00), the mean and variance of the task duration (mean 2 hours, variance 0.5 hours).

[0095] Standardize the environmental parameters to obtain numerical representations of light intensity (such as 0.8), temperature (such as 25 °C), and humidity (such as 60%).

[0096] Convert the task completion status into task completion rate (such as 95%), task on-time rate (such as 90%), and task quality score (such as 4.5 points out of 5).

[0097] Through the above feature extraction, a high-dimensional task feature vector can be obtained, such as [1, 0, 0, 5000, 10, 0.2, 0.5, 0.8, 0.6, 0.4, 2, 0.5, 0.8, 25, 60, 0.95, 0.9, 4.5].

[0098] Next, use a deep autoencoder network to perform dimensionality reduction on the task feature vector. Design a multi-layer autoencoder network with the number of input layer nodes being the same as the dimension of the feature vector (such as 18), and compress the high-dimensional features into a 4-dimensional latent space through 3 hidden layers (with the number of nodes being 12, 8, 4 respectively). During the network training process, in addition to the reconstruction error, introduce a task correlation loss function to ensure that the distances of similar tasks in the latent space are closer. For example, cosine similarity can be used as a measure of task correlation. For any two tasks i and j, calculate the cosine similarity S_ij of their original feature vectors and the cosine similarity S'_ij of their latent space representations, and use the mean of |S_ij - S'_ij| as an additional loss term.

[0099] In a 4D latent space, the density peak clustering algorithm is used to adaptively determine the number of clustering centers. This algorithm first calculates the local density of each point and the minimum distance to high-density points, and then identifies density peak points as clustering centers through a decision graph. For example, three clustering centers may be obtained, corresponding to three main task patterns.

[0100] Feature extraction and pattern summarization are performed on each task cluster. The central vector of the task cluster is calculated as the typical feature representation of this type of task, such as [0.2, 0.8, -0.5, 0.1]. The path pattern (such as "circular inspection path"), time window (such as "9:00 - 17:00 on weekdays"), and resource requirements (such as "20% single power consumption") within the task cluster are extracted to form a standardized task pattern description. An evaluation index system including execution efficiency (such as average completion time), resource utilization rate (such as average power utilization rate), and adaptability (such as completion rate under different weather conditions) is established. The task pattern description and the evaluation index system are stored in the task model library.

[0101] When a new task is received, feature extraction is performed on the new task to obtain a new task feature vector. This vector is input into the trained autoencoder network to obtain its representation in the latent space, such as [0.3, 0.7, -0.4, 0.2]. The Euclidean distance or cosine similarity between this vector and the central vectors of each task pattern is calculated, and the task pattern with the highest similarity is selected as the reference pattern.

[0102] The reference pattern is adjusted in combination with the current environmental constraints (such as weather conditions and traffic conditions) to generate an execution plan. For example, if the reference pattern is "circular inspection path", but a certain section of the road is under construction at present, the path needs to be adjusted to bypass the construction area. The generated execution plan includes specific path planning, time arrangement, and resource allocation.

[0103] The execution result is fed back to the task model library for incremental update. For example, if the actual execution efficiency of the new task is lower than expected, the evaluation index of this task pattern can be appropriately reduced; if new task features appear, the dimension of the feature vector can be expanded and the autoencoder network can be retrained.

[0104] The beneficial effects of this method are mainly reflected in the following three aspects:

[0105] 1. It improves the accuracy and efficiency of task pattern recognition. Nonlinear dimensionality reduction of high-dimensional features is achieved through the deep autoencoder network, retaining the similarity relationship between tasks, which is beneficial for subsequent clustering analysis. The density peak clustering algorithm can adaptively determine the number of clusters, avoiding the limitation of traditional clustering methods that require pre-specifying the number of clusters.

[0106] 2. The adaptability and robustness of the task execution plan are enhanced. The task model library constructed based on historical data contains rich task pattern information, providing a reliable reference for the execution of new tasks. By dynamically adjusting the reference pattern in combination with the current environmental constraints, the generated execution plan is more flexible and can adapt to complex and ever-changing actual situations.

[0107] 3. The continuous optimization of task models and knowledge accumulation are realized. By feeding back the results of each task execution to the model library for incremental updates, the system can continuously learn and optimize task patterns, improving its adaptability to new situations. This closed-loop learning mechanism enables the performance of the system to continuously improve with the increase in usage time and has strong practical value.

[0108] In an optional implementation manner, determining the execution mode of the current task by performing feature matching between the task instruction and a preset task pattern in the task model library includes:

[0109] Receiving a task instruction and constructing a feature vector, performing spatial feature encoding on the task target location information to obtain target point coordinates, regional attributes, and spatial constraint conditions, performing temporal feature analysis on the task execution time to obtain a time window, duration, and time key point information, converting the task priority into a priority matrix, parameterizing the environmental constraint conditions to obtain environmental complexity, dynamic obstacle distribution, and weather condition parameters, and forming a task instruction feature vector through feature normalization processing;

[0110] Establishing a hierarchical feature matching framework, performing multi-dimensional matching between the task instruction feature vector and the preset task pattern, calculating the matching degree of the task type dimension through cosine similarity, calculating the similarity of the path feature using the dynamic time warping algorithm, analyzing the matching degree of the execution time using the temporal pattern matching algorithm, and determining the matching degree of the environmental constraint through calculating the environmental parameter similarity;

[0111] Constructing a feature importance evaluation model based on historical task execution data, calculating the contribution degree of each dimension feature to the task execution effect, constructing an adaptive learning model in combination with the current task characteristics and environmental conditions, and dynamically adjusting and online optimizing the weight coefficients of each dimension feature;

[0112] Calculating the weighted matching scores of each preset task pattern, where the weighted matching score is the weighted sum of the matching results of each dimension feature and the corresponding weight coefficients, establishing a mode switching cost model to evaluate the resource consumption of state switching, and constructing a decision optimization model in combination with the task priority and time constraint;

[0113] When there are multiple candidate execution modes, a fuzzy decision-making method is adopted to comprehensively evaluate the weighted matching score, mode switching cost, and task priority, select the optimal execution mode, collect execution status data in real time during task execution to evaluate the mode matching accuracy, and feedback the matching result and evaluation data to the feature importance evaluation model for update and optimization.

[0114] The implementation details are described as follows:

[0115] Receive a task instruction and construct a feature vector. First, perform spatial feature encoding on the task target location information, including the following steps:

[0116] (1) Convert the target location coordinates to coordinate values in a standard three-dimensional coordinate system, such as (x, y, z) = (10.5, 20.3, 5.2) meters;

[0117] (2) Extract the target area attributes, such as indoor / outdoor, flat slope, etc., and represent them with a 0-1 vector. For example, [1, 0, 1, 0] represents an outdoor flat area;

[0118] (3) Analyze the spatial constraint conditions, such as the maximum activity range, restricted access areas, etc., and represent them with a rectangular bounding box. For example, [(0, 0, 0), (100, 100, 10)] represents the activity range.

[0119] Conduct temporal feature analysis on the task execution time, including:

[0120] (1) Extract the time window, such as the start and end times "2023-05-01 08:00:00" to "2023-05-01 18:00:00";

[0121] (2) Calculate the duration, such as 36000 seconds;

[0122] (3) Identify the key time points, such as the task start, end, midpoint checkpoints, etc., and represent them with a list of timestamps.

[0123] Convert the task priority into a priority matrix. For example, for 3 priorities, a 3x3 matrix can be used to represent the relative importance between different priorities.

[0124] Parametrize the environmental constraint conditions:

[0125] (1) Calculate the environmental complexity, considering factors such as the number of obstacles and distribution density, and represent it with a score from 0 to 100;

[0126] (2) Analyze the distribution of dynamic obstacles, such as pedestrians and vehicles, and represent it with a density heat map;

[0127] (3) Extract the weather condition parameters, such as temperature, humidity, visibility, etc., and convert them into normalized values.

[0128] Through feature normalization processing, the above-mentioned features in each dimension are combined to form a task instruction feature vector.

[0129] Establish a hierarchical feature matching framework. First, perform multi-dimensional matching between the task instruction feature vector and the preset task patterns:

[0130] (1) Calculate the matching degree in the task type dimension: Extract the keywords of the task instruction and the preset pattern, construct word vectors, and obtain the matching score by calculating the cosine similarity between the word vectors;

[0131] (2) Analyze the similarity of the path features: Convert the task path and the preset pattern path into time series, and use the dynamic time warping algorithm to calculate the path similarity;

[0132] (3) Evaluate the matching degree of the execution time: Align the task time series with the preset pattern, and use the time series pattern matching algorithm to analyze the similarity degree of the time features;

[0133] (4) Determine the matching degree of the environmental constraints: Normalize the task environment parameters and the preset pattern, and calculate the Euclidean distance to obtain the environmental similarity.

[0134] Construct a feature importance evaluation model based on historical task execution data. The specific steps are as follows:

[0135] (1) Collect historical task data, including task features, execution patterns, completion status, etc.;

[0136] (2) Use the random forest algorithm to construct a feature importance model and calculate the contribution degree of each feature to the task execution effect;

[0137] (3) Combine the characteristics of the current task and the environmental conditions to construct an adaptive learning model, such as using the online gradient descent algorithm to dynamically adjust the feature weights.

[0138] Calculate the weighted matching scores of each preset task pattern:

[0139] (1) Multiply the matching results of each dimension feature by the corresponding weight coefficients and sum them to obtain the weighted matching score;

[0140] (2) Establish a mode switching cost model, consider factors such as switching time and energy consumption, and evaluate the resource consumption of state switching;

[0141] (3) Combine the task priority and time constraints to construct a decision optimization model, such as using a multi-objective optimization algorithm.

[0142] When there are multiple candidate execution patterns, use the fuzzy decision method for comprehensive evaluation:

[0143] (1) Fuzzify the indicators such as weighted matching scores, pattern switching costs, and task priorities;

[0144] (2) Establish a fuzzy rule base, such as "If the matching degree is high and the switching cost is low, then select this pattern";

[0145] (3) Adopt a fuzzy reasoning method, such as Mamdani reasoning, to obtain the comprehensive scores of each candidate pattern;

[0146] (4) Select the pattern with the highest comprehensive score as the optimal execution pattern.

[0147] During the task execution process:

[0148] (1) Real-time collect execution status data, such as position, speed, energy consumption, etc.;

[0149] (2) Calculate the deviation between the actual execution trajectory and the expected trajectory, and evaluate the pattern matching accuracy;

[0150] (3) Feed back the matching results and evaluation data to the feature importance evaluation model;

[0151] (4) Adopt an incremental learning method, such as online Bayesian update, to optimize the model parameters in real time.

[0152] The beneficial effects of this method include:

[0153] 1. Improve the accuracy and adaptability of task execution pattern selection. Through multi-dimensional feature matching and adaptive weight adjustment, it can more accurately identify the execution pattern most suitable for the current task, improving the quality and efficiency of task completion.

[0154] 2. Enhance the robustness and adaptive ability of the system. By adopting fuzzy decision-making and online learning methods, the system can better cope with complex and changing task environments, improving the reliability and stability of the system.

[0155] 3. Optimize the system resource utilization efficiency. By establishing a pattern switching cost model and a decision optimization model, while ensuring the task execution effect, it minimizes unnecessary pattern switching, reducing the system energy consumption and resource consumption.

[0156] In an alternative embodiment, a task execution decision tree is constructed based on the execution pattern, and each node of the decision tree includes an action sequence, resource consumption, and time constraints; a heuristic algorithm is used to prune and optimize the decision tree, removing the execution paths that do not meet the environmental constraint conditions, including:

[0157] Construct a multi - layer decision tree according to the task execution mode. The root node of the decision tree stores the task start state information, the intermediate nodes store the action sequence information, resource consumption information, and time constraint information, the leaf nodes store the target state information, and the connecting edges between the nodes represent the state transition relationship;

[0158] Obtain environmental perception data, including static obstacle position information, dynamic obstacle trajectory information, environmental light information, and temperature information, establish an environmental constraint model, and decompose the environmental constraint model into spatial constraint conditions and time constraint conditions;

[0159] Construct a heuristic evaluation function. The heuristic evaluation function includes a path length evaluation term, an action coherence evaluation term, a resource consumption evaluation term, and a safety margin evaluation term, and evaluate and score each execution path in the decision tree;

[0160] Set a pruning threshold, and perform hierarchical pruning on the decision tree based on the environmental constraint model and the heuristic evaluation function. First, prune the execution paths that violate the spatial constraint conditions, secondly, prune the execution paths that violate the time constraint conditions, and thirdly, prune the execution paths whose evaluation scores are lower than the pruning threshold;

[0161] Optimize the paths of the pruned decision tree, calculate the comprehensive cost of each feasible execution path. The comprehensive cost is calculated by weighted combination of path length, number of actions, resource consumption, and safety margin, and select the execution path with the minimum comprehensive cost as the optimal execution path;

[0162] During the task execution process, obtain environmental change information in real - time. When it is detected that the environmental constraint conditions change, perform online adjustment on the optimal execution path, and re - perform decision tree pruning and path optimization to ensure that the execution path continuously meets the dynamic environmental constraint conditions.

[0163] This embodiment provides a method for constructing a task execution decision tree based on the execution mode and optimizing it. This method first constructs a multi - layer decision tree according to the task execution mode. The root node of the decision tree stores the task start state information, the intermediate nodes store the action sequence information, resource consumption information, and time constraint information, the leaf nodes store the target state information, and the connecting edges between the nodes represent the state transition relationship.

[0164] For example, for a task of a robot grasping an object, the root node of the decision tree can store the initial position and posture information of the robot, the intermediate nodes can store action sequences such as moving and grasping, as well as the energy consumption and time required for each action, and the leaf nodes store the final position and posture information of the robot after successfully grasping the object. The connecting edges between the nodes represent the conversion process from one state to another.

[0165] Next, obtain environmental perception data, including static obstacle position information, dynamic obstacle trajectory information, environmental light information, temperature information, etc., and establish an environmental constraint model. Decompose the environmental constraint model into spatial constraint conditions and time constraint conditions.

[0166] For example, for the above grasping task, three-dimensional coordinate information of fixed obstacles in the workspace, the motion trajectory of moving obstacles such as conveyor belts, environmental light intensity, temperature and other data can be obtained. The spatial constraint conditions can include prohibited areas that the robot needs to avoid during movement, and the time constraint conditions can include the requirement to complete the task within a certain time window.

[0167] Then construct a heuristic evaluation function, including a path length evaluation term, an action coherence evaluation term, a resource consumption evaluation term, and a safety margin evaluation term, and evaluate and score each execution path in the decision tree.

[0168] The path length evaluation term can calculate the path length from the start state to the target state, the action coherence evaluation term can consider the smooth transition between adjacent actions, the resource consumption evaluation term can calculate the total energy consumption required for the execution path, and the safety margin evaluation term can consider the minimum distance between the robot and the obstacles.

[0169] Set a pruning threshold, and perform hierarchical pruning on the decision tree based on the environmental constraint model and the heuristic evaluation function. First, prune the execution paths that violate the spatial constraint conditions, such as the paths that will collide with fixed obstacles. Second, prune the execution paths that violate the time constraint conditions, such as the paths that cannot be completed within the specified time. Third, prune the execution paths whose evaluation scores are lower than the pruning threshold.

[0170] Optimize the paths of the pruned decision tree, and calculate the comprehensive cost of each feasible execution path. The comprehensive cost is calculated by weighting the path length, the number of actions, the resource consumption, and the safety margin. For example, the path length weight can be set to 0.3, the action number weight to 0.2, the resource consumption weight to 0.3, and the safety margin weight to 0.2. Select the execution path with the minimum comprehensive cost as the optimal execution path.

[0171] During the task execution process, obtain real-time environmental change information. When it is detected that the environmental constraint conditions change, perform online adjustment on the optimal execution path. For example, when a new moving obstacle is detected, re-perform decision tree pruning and path optimization to ensure that the execution path continuously meets the dynamic environmental constraint conditions.

[0172] The beneficial effects of this method include:

[0173] 1. By constructing a multi-layer decision tree and performing pruning optimization, the computational complexity of task planning can be effectively reduced, and the planning efficiency can be improved.

[0174] 2. Combining the environmental constraint model and the heuristic evaluation function for path optimization can obtain the optimal execution path that meets various constraint conditions, improving the reliability and efficiency of task execution.

[0175] 3. By adjusting the execution path in real time, it is possible to adapt to the dynamically changing environment and enhance the robustness and adaptability of the system.

[0176] In an alternative embodiment, according to the task priority and time constraint, calculating the task switching cost of each node in the decision tree, and generating the optimal task execution path includes:

[0177] Obtaining the task priority information and time constraint information, dividing the task priority into multiple levels and assigning different weight coefficients, and converting the time constraint information into time window parameters and deadline parameters;

[0178] Constructing a task switching cost calculation model, the task switching cost calculation model includes state transition cost, resource reconfiguration cost, and time delay cost, where the state transition cost is calculated based on the state difference between two task nodes, the resource reconfiguration cost is calculated based on the change in resource occupancy, and the time delay cost is calculated based on the matching degree between the task execution time and the time window;

[0179] Calculating the task switching cost between adjacent nodes in the decision tree. First, calculating the state transition matrix between nodes and calculating the state transition cost based on the state transition matrix. Secondly, calculating the change in resource configuration between nodes to obtain the resource reconfiguration cost, and then calculating the overlap degree between the task execution time and the time window to obtain the time delay cost;

[0180] Establishing a path optimization objective function, the path optimization objective function comprehensively considers the total task switching cost, task priority weight, and time constraint satisfaction degree, and uses the dynamic programming algorithm to traverse the decision tree to calculate the cumulative cost from the root node to each leaf node;

[0181] According to the calculation result of the path optimization objective function, selecting the path with the minimum cumulative cost as the candidate optimal path, and performing time constraint verification on the candidate optimal path to ensure that all task nodes on the path meet their respective deadline constraints;

[0182] During the task execution process, real-time monitoring of the task status is carried out. When it is detected that the task priority changes or the time constraint is adjusted, the task switching cost of the affected nodes is recalculated, and the optimal execution path is updated to achieve dynamic optimization adjustment of the path.

[0183] The specific implementation method is as follows:

[0184] First, obtain the task priority information and time constraint information. Task priorities are usually divided into three levels: high, medium, and low, with weight coefficients of 1.5, 1.0, and 0.5 assigned respectively. The time constraint information includes time window parameters and deadline parameters. For example, the time window for task A is from 8:00 to 10:00, and the deadline is 11:00.

[0185] Then, construct a task switching cost calculation model. The state transition cost can be calculated by comparing the differences in state parameters between two task nodes, such as memory usage rate, CPU occupancy rate, etc. The resource reconfiguration cost can be calculated based on the change range of resource occupancy. For example, the memory allocation changes from 2GB to 4GB. The time delay cost is based on the overlap degree between the task execution time and the time window. The higher the overlap degree, the lower the cost.

[0186] Next, calculate the task switching cost between adjacent nodes in the decision tree. Taking the example of task A switching to task B, first calculate the state transition matrix. Assume the state parameters include memory usage rate and CPU occupancy rate, and the state transition matrix from A to B is [[0.8, 0.2], [0.3, 0.7]]. Based on this matrix, the state transition cost is calculated as 0.6. Secondly, calculate the resource reconfiguration cost. For example, the memory increases from 2GB to 4GB, and the cost is 0.5. Then calculate the time delay cost. Assume the overlap degree between the execution time of task B and the time window is 80%, and the cost is 0.2. The total switching cost from A to B is comprehensively obtained as 1.3.

[0187] Then, establish a path optimization objective function. This function comprehensively considers the total task switching cost, task priority weights, and time constraint satisfaction. Use the dynamic programming algorithm to traverse the decision tree and calculate the cumulative cost from the root node to each leaf node. For example, the cumulative cost of path 1 is 5.8, path 2 is 6.2, and path 3 is 5.5.

[0188] According to the calculation results of the objective function, select path 3 with the minimum cumulative cost as the candidate optimal path. Conduct a time constraint verification on this path to check whether all task nodes on the path meet their respective deadline constraints. Assume there are 3 task nodes on path 3, and their execution completion times are 9:30, 10:45, and 11:50 respectively, all of which meet the deadline constraints. Therefore, path 3 is determined as the optimal execution path.

[0189] During the task execution process, monitor the task status in real time. When it is detected that the priority of task C changes from low to high, recalculate the task switching cost of the nodes related to C. After the update, it is found that the cumulative cost of path 2 drops to 5.3, which is lower than the original optimal path 3. After passing the time constraint verification, it is confirmed that path 2 meets the requirements, and it is updated as the new optimal execution path, realizing the dynamic optimization adjustment of the path.

[0190] The beneficial effects of this method include:

[0191] 1) By constructing an accurate task switching cost model and comprehensively considering various factors such as state migration, resource reconfiguration, and time delay, it is possible to more accurately evaluate the actual overhead of task switching, thereby generating a better task execution path.

[0192] 2) The decision tree is traversed using a dynamic programming algorithm, avoiding the high computational complexity of exhausting all paths and improving the efficiency of optimal path search. At the same time, through time constraint verification, the executability of the generated path is ensured.

[0193] 3) The dynamic optimization adjustment of the task execution path is realized, and the optimal execution path can be updated in a timely manner according to real-time situations such as task priority changes or time constraint adjustments, improving the adaptability and robustness of the system to environmental changes.

[0194] In an optional implementation manner, when it is detected that the environment has changed, the online optimization of the control instruction sequence is performed based on a deep reinforcement learning model, and the generation of new control instructions adapted to the environmental change includes:

[0195] Construct a deep reinforcement learning model, which includes a state space, an action space, and a reward function. The state space includes environmental state information and task execution state information. The action space includes a set of optional control instructions. The reward function is designed based on the execution effect of the control instruction and environmental adaptability;

[0196] Obtain real-time environmental data through an environmental perception module, parse the environmental data into obstacle information, light information, and temperature information, construct an environmental feature vector, and combine the environmental feature vector with the task execution state information to form a state input;

[0197] Adopt a dual network structure, including a policy network and a value network. The policy network is responsible for generating a control instruction sequence, and the value network is responsible for evaluating the execution effect of the control instruction. The two networks use a convolutional layer with shared parameters to extract state features;

[0198] Based on an experience replay mechanism, store historical interaction data, including state transition sequences, control instruction sequences, and reward sequences. Randomly sample training data batches from the experience pool, update the parameters of the value network using a temporal difference algorithm, and optimize the parameters of the policy network using a policy gradient algorithm;

[0199] Set an environmental change detection threshold. When the change amplitude of the environmental feature vector exceeds the detection threshold, trigger online optimization, input the current state into the deep reinforcement learning model, and generate multiple candidate control instruction sequences through the policy network;

[0200] Perform simulation evaluation on the candidate control instruction sequences, calculate the expected reward values of each sequence in the new environment, select the control instruction sequence with the highest expected reward value as the optimization result, and decompose the optimized control instruction sequence into individual control instructions for step-by-step execution.

[0201] The specific implementation method is as follows:

[0202] First, construct a deep reinforcement learning model. This model consists of three key elements: the state space, the action space, and the reward function. The state space is composed of environmental state information and task execution state information. The environmental state information includes obstacle positions, light intensities, temperatures, etc., and the task execution state information includes the current position, speed, direction, etc. The action space contains various optional control instructions, such as forward, backward, turning, accelerating, decelerating, etc. The reward function is designed based on the execution effect of the control instruction and environmental adaptability. For example, a positive reward for reaching the target point and a negative penalty for colliding with an obstacle can be set.

[0203] Next, obtain real-time environmental data through the environmental perception module. This module contains various sensing devices such as cameras, lidars, and temperature sensors. Parse and extract features from the obtained raw data to obtain obstacle information (such as position, size, shape, etc.), light information (such as brightness, color temperature, etc.), temperature information, etc. Combine this information into an environmental feature vector, which can be represented as [x1, x2,..., xn], where xi represents a certain environmental feature. Then combine the environmental feature vector with the task execution state information (such as the current position coordinates, speed, direction angle, etc.) to form a complete state input.

[0204] Then, adopt a dual-network structure, including a policy network and a value network. The policy network is responsible for generating control instruction sequences, and the value network is responsible for evaluating the execution effect of the control instructions. The two networks use a convolutional layer with shared parameters to extract state features. For example, a 3-layer convolutional network can be used, with convolutional kernel sizes of 5x5, 3x3, 3x3, strides of 2, 1, 1, and channel numbers of 32, 64, 64 respectively. After the convolutional layer, a fully connected layer is connected. The policy network outputs the action probability distribution, and the value network outputs the state value estimation.

[0205] Store historical interaction data based on the experience replay mechanism. Set a fixed-size experience pool, such as a capacity of 10000. After each interaction with the environment, store the state transition sequence, control instruction sequence, and reward sequence in the experience pool. When the experience pool is full, update it using the first-in-first-out strategy. During training, randomly sample data of batch size (such as 64) from the experience pool for learning. Use the temporal difference algorithm to update the value network parameters, with the goal of minimizing the mean square error between the predicted value and the actual value. Use the policy gradient algorithm to optimize the policy network parameters, with the goal of maximizing the expected cumulative reward.

[0206] Set an environmental change detection threshold, for example, trigger online optimization when the Euclidean distance change of the environmental feature vector exceeds 0.5. Input the current state into the deep reinforcement learning model, and generate multiple (such as 10) candidate control instruction sequences through the policy network. Each sequence contains control instructions for the next N steps (such as 10 steps).

[0207] Perform simulation evaluation on the candidate control instruction sequences. Simulate the execution of each sequence in the environmental model, and calculate the cumulative reward as the expected reward value. Select the control instruction sequence with the highest expected reward value as the optimization result. Decompose this sequence into individual control instructions and execute them step by step in order. Continuously monitor environmental changes during the execution process, and repeat the optimization process if the threshold is triggered again.

[0208] The beneficial effects of this method include:

[0209] 1) Achieve the adaptive optimization of the control instruction sequence through the deep reinforcement learning model, improving the adaptability of the system to environmental changes. 2) Adopt a dual network structure and an experience replay mechanism to improve the learning efficiency and model performance. 3) Set an environmental change detection threshold and an online optimization mechanism to achieve a rapid response to environmental changes, ensuring the real-time and effectiveness of control instructions.

[0210] Figure 2 It is a schematic structural diagram of the multi-mode task scheduling and intelligent execution system of the intelligent vehicle in the embodiment of the present invention, as Figure 2 shown, the system includes:

[0211] The first unit is used to obtain the historical task data of the intelligent vehicle, and the historical task data includes task type, task path, task duration, environmental parameters, and task completion status; perform clustering analysis on the historical task data based on the deep learning algorithm, classify tasks with similar characteristics into the same task mode, and construct a task model library containing multiple preset task modes; receive a new task instruction, and the task instruction includes a task target location, a task execution time, a task priority, and environmental constraint conditions; perform feature matching between the task instruction and the preset task modes in the task model library to determine the execution mode of the current task;

[0212] The second unit is used to construct a task execution decision tree based on the execution mode, and each node of the decision tree includes an action sequence, resource consumption, and time constraints; use a heuristic algorithm to prune and optimize the decision tree to remove execution paths that do not meet the environmental constraint conditions; calculate the task switching cost of each node of the decision tree according to the task priority and time constraints, and generate an optimal task execution path; convert the optimal task execution path into a control instruction sequence of the intelligent vehicle, and the control instruction sequence includes a motion trajectory, speed planning, and execution timing;

[0213] A third unit is used for the intelligent vehicle to start executing tasks according to the control instruction sequence, and to collect environmental data in real time during the task execution process. The environmental data includes obstacle information, light intensity, temperature, and road condition status. When it is detected that the environment has changed, the control instruction sequence is optimized online based on a deep reinforcement learning model to generate new control instructions adapted to the environmental change. The state data, environmental data, and optimization results generated during the task execution process are fed back to the task model library for updating and improving the preset task mode, so as to improve the execution efficiency of subsequent tasks.

[0214] In the third aspect of the embodiments of the present invention,

[0215] There is provided an electronic device, including:

[0216] A processor;

[0217] A memory for storing instructions executable by the processor;

[0218] Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.

[0219] In the fourth aspect of the embodiments of the present invention,

[0220] There is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.

[0221] The present invention can be a method, an apparatus, a system, and / or a computer program product. The computer program product may include a computer-readable storage medium, on which computer-readable program instructions for executing various aspects of the present invention are loaded.

[0222] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A multi-mode task scheduling and intelligent execution method for an intelligent vehicle, characterized in that, Including: Obtain the historical task data of the intelligent vehicle, where the historical task data includes task type, task path, task duration, environmental parameters, and task completion status; Based on deep learning algorithms, perform clustering analysis on the historical task data, classify tasks with similar characteristics into the same task pattern, and construct a task model library containing multiple preset task patterns; Receive a new task instruction, where the task instruction includes a task target location, a task execution time, a task priority, and environmental constraint conditions; Match the features of the task instruction with the preset task patterns in the task model library to determine the execution mode of the current task; Construct a task execution decision tree based on the execution mode, where each node of the decision tree includes an action sequence, resource consumption, and time constraints; Use a heuristic algorithm to prune and optimize the decision tree, removing execution paths that do not meet the environmental constraint conditions; According to the task priority and time constraints, calculate the task switching cost of each node of the decision tree to generate an optimal task execution path; Convert the optimal task execution path into a control instruction sequence for the intelligent vehicle, where the control instruction sequence includes a motion trajectory, speed planning, and execution timing; The intelligent vehicle starts to execute the task according to the control instruction sequence, and in the process of task execution, it collects environmental data in real time, where the environmental data includes obstacle information, light intensity, temperature, and road condition status; When it is detected that the environment has changed, based on a deep reinforcement learning model, perform online optimization on the control instruction sequence to generate new control instructions that adapt to the environmental changes; Feed back the status data, environmental data, and optimization results generated during the task execution to the task model library for updating and improving the preset task patterns to improve the execution efficiency of subsequent tasks.

2. The method according to claim 1, wherein Based on deep learning algorithms, performing clustering analysis on the historical task data, and classifying tasks with similar characteristics into the same task pattern, and constructing a task model library containing multiple preset task patterns includes: Obtain the historical task data of the intelligent vehicle, and construct a task feature vector through a multi-dimensional feature extraction method, where the multi-dimensional feature extraction method includes: performing one-hot encoding on the task type to convert it into a numerical feature, extracting the spatial features of the task path to obtain the path length, the number of turns, and the key node distribution parameters, performing time series analysis on the task duration to obtain the time distribution features, performing standardization processing on the environmental parameters to obtain the numerical representations of light intensity, temperature, and humidity, and converting the task completion status into a task completion rate, a task on-time rate, and a task quality score; Use a deep autoencoder network to perform dimensionality reduction processing on the task feature vector, where the number of input layer nodes of the deep autoencoder network is the same as the dimension of the task feature vector, compress the high-dimensional features to a low-dimensional latent space through multiple hidden layers, introduce a task correlation loss function during the network training process to ensure that the distances of similar tasks in the latent space are relatively close, and use a density peak clustering algorithm in the low-dimensional latent space to adaptively determine the number of clustering centers to form task clusters; Feature extraction and pattern summarization are performed on each of the task clusters. The central vector of the task cluster is calculated as the typical feature representation of this type of task. The path patterns, time windows, and resource requirements of the samples within the task cluster are extracted to form a standardized task pattern description. An evaluation index system including execution efficiency, resource utilization rate, and adaptability is established, and the task pattern description and the evaluation index system are stored in the task model library; When a new task is received, feature extraction is performed on the new task to obtain a new task feature vector. The similarity between the new task feature vector and the central vectors of each task pattern is calculated, and the task pattern with the highest similarity is selected as the reference pattern. The reference pattern is adjusted in combination with the current environmental constraint conditions to generate an execution plan, and the execution result is fed back to the task model library for incremental update.

3. The method according to claim 1, wherein Feature matching is performed between the task instruction and the preset task patterns in the task model library to determine the execution mode of the current task, including: Receive a task instruction and construct a feature vector. Perform spatial feature encoding on the task target location information to obtain target point coordinates, regional attributes, and spatial constraint conditions. Perform temporal feature analysis on the task execution time to obtain time windows, duration, and time key point information. Convert the task priority into a priority matrix, and perform parameterization processing on the environmental constraint conditions to obtain environmental complexity, dynamic obstacle distribution, and weather condition parameters. Form a task instruction feature vector through feature normalization processing; Establish a hierarchical feature matching framework, perform multi-dimensional matching between the task instruction feature vector and the preset task patterns. Calculate the matching degree of the task type dimension through cosine similarity, calculate the similarity of the path features using the dynamic time warping algorithm, analyze the matching degree of the execution time using the temporal pattern matching algorithm, and determine the matching degree of the environmental constraints through the calculation of environmental parameter similarity; Build a feature importance evaluation model based on historical task execution data, calculate the contribution of each dimension feature to the task execution effect, and build an adaptive learning model in combination with the current task characteristics and environmental conditions to dynamically adjust and online optimize the weight coefficients of each dimension feature; Calculate the weighted matching scores of each preset task pattern. The weighted matching score is the weighted sum of the matching results of each dimension feature and the corresponding weight coefficients. Establish a mode switching cost model to evaluate the resource consumption of state switching, and build a decision optimization model in combination with task priority and time constraints; When there are multiple candidate execution modes, use the fuzzy decision method to comprehensively evaluate the weighted matching score, mode switching cost, and task priority, select the optimal execution mode, and collect execution status data in real time during task execution to evaluate the pattern matching accuracy. Feed the matching result and evaluation data back to the feature importance evaluation model for update and optimization.

4. The method according to claim 1, characterized in that Build a task execution decision tree based on the execution mode. Each node of the decision tree contains an action sequence, resource consumption, and time constraints; Use a heuristic algorithm to prune and optimize the decision tree, and remove the execution paths that do not meet the environmental constraint conditions, including: Construct a multi - layer decision tree according to the task execution mode. The root node of the decision tree stores the task start state information, the intermediate nodes store the action sequence information, resource consumption information, and time constraint information, and the leaf nodes store the target state information. The connection edges between the nodes represent the state transition relationship; Obtain environmental perception data, including static obstacle position information, dynamic obstacle trajectory information, environmental light information, and temperature information, establish an environmental constraint model, and decompose the environmental constraint model into spatial constraint conditions and time constraint conditions; Construct a heuristic evaluation function. The heuristic evaluation function includes a path length evaluation term, an action coherence evaluation term, a resource consumption evaluation term, and a safety margin evaluation term, and evaluate and score each execution path in the decision tree; Set a pruning threshold, and perform hierarchical pruning on the decision tree based on the environmental constraint model and the heuristic evaluation function. First, prune the execution paths that violate the spatial constraint conditions, second, prune the execution paths that violate the time constraint conditions, and third, prune the execution paths with evaluation scores lower than the pruning threshold; Optimize the paths of the pruned decision tree, calculate the comprehensive cost of each feasible execution path. The comprehensive cost is calculated by weighted combination of path length, number of actions, resource consumption, and safety margin, and select the execution path with the minimum comprehensive cost as the optimal execution path; During the task execution process, obtain the environmental change information in real - time. When it is detected that the environmental constraint conditions change, perform online adjustment on the optimal execution path, and re - perform decision tree pruning and path optimization to ensure that the execution path continuously meets the dynamic environmental constraint conditions.

5. The method according to claim 1, characterized in that, According to the task priority and time constraint, calculate the task switching cost of each node in the decision tree. Generating the optimal task execution path includes: Obtain the task priority information and time constraint information, divide the task priority into multiple levels and assign different weight coefficients, and convert the time constraint information into time window parameters and deadline parameters; Construct a task switching cost calculation model. The task switching cost calculation model includes a state transition cost, a resource re - configuration cost, and a time delay cost. Among them, the state transition cost is calculated based on the state difference between two task nodes, the resource re - configuration cost is calculated based on the change in resource occupancy, and the time delay cost is calculated based on the matching degree between the task execution time and the time window; Calculate the task switching cost between adjacent nodes in the decision tree. First, calculate the state transition matrix between the nodes, calculate the state transition cost based on the state transition matrix, second, calculate the change in resource configuration between the nodes to obtain the resource re - configuration cost, and then calculate the overlap degree between the task execution time and the time window to obtain the time delay cost; Establish a path optimization objective function. The path optimization objective function comprehensively considers the total task switching cost, task priority weight, and time constraint satisfaction degree, and uses the dynamic programming algorithm to traverse the decision tree and calculate the cumulative cost from the root node to each leaf node; According to the calculation result of the path optimization objective function, select the path with the minimum cumulative cost as the candidate optimal path, and verify the time constraint of the candidate optimal path to ensure that all task nodes on the path meet their respective deadline constraints; During the task execution process, monitor the task status in real time. When a task priority change or time constraint adjustment is detected, recalculate the task switching cost of the affected nodes, update the optimal execution path, and achieve dynamic optimization and adjustment of the path.

6. The method according to claim 1, characterized in that When it is detected that the environment has changed, online optimization of the control instruction sequence is performed based on the deep reinforcement learning model to generate new control instructions adapted to the environmental change, including: Construct a deep reinforcement learning model, which includes a state space, an action space, and a reward function. The state space includes environmental state information and task execution state information. The action space includes a set of optional control instructions. The reward function is designed based on the execution effect of the control instructions and environmental adaptability; Obtain real-time environmental data through the environmental perception module, parse the environmental data into obstacle information, light information, and temperature information, construct an environmental feature vector, and combine the environmental feature vector with the task execution state information to form a state input; Adopt a dual network structure, including a policy network and a value network. The policy network is responsible for generating a control instruction sequence, and the value network is responsible for evaluating the execution effect of the control instructions. The two networks use a convolutional layer with shared parameters to extract state features; Store historical interaction data based on the experience replay mechanism, including state transition sequences, control instruction sequences, and reward sequences. Randomly sample training data batches from the experience pool, update the value network parameters using the temporal difference algorithm, and optimize the policy network parameters using the policy gradient algorithm; Set an environmental change detection threshold. When the change amplitude of the environmental feature vector exceeds the detection threshold, trigger online optimization. Input the current state into the deep reinforcement learning model, and generate multiple candidate control instruction sequences through the policy network; Perform simulation evaluation on the candidate control instruction sequences, calculate the expected reward value of each sequence in the new environment, select the control instruction sequence with the highest expected reward value as the optimization result, and decompose the optimized control instruction sequence into individual control instructions for step-by-step execution.

7. A multi-mode task scheduling and intelligent execution system for an intelligent vehicle, which is used to implement the method described in any one of the foregoing claims 1-6, characterized in that, Including: The first unit is used to obtain the historical task data of the intelligent vehicle. The historical task data includes task type, task path, task duration, environmental parameters, and task completion status; Perform clustering analysis on the historical task data based on the deep learning algorithm, classify tasks with similar characteristics into the same task pattern, and construct a task model library containing multiple preset task patterns; receive a new task instruction, which includes a task target location, a task execution time, a task priority, and environmental constraint conditions; match the features of the task instruction with the preset task patterns in the task model library to determine the execution mode of the current task; A second unit, configured to build a task execution decision tree based on the execution mode, where each node of the decision tree includes an action sequence, resource consumption, and time constraints; pruning and optimizing the decision tree by using a heuristic algorithm to remove execution paths that do not meet the environmental constraint conditions; calculating the task switching cost of each node of the decision tree according to the task priority and time constraints to generate an optimal task execution path; converting the optimal task execution path into a control instruction sequence for the intelligent vehicle, where the control instruction sequence includes a motion trajectory, speed planning, and execution timing; A third unit, configured to enable the intelligent vehicle to start executing a task according to the control instruction sequence, and to collect environmental data in real time during the task execution, where the environmental data includes obstacle information, light intensity, temperature, and road condition status; when it is detected that the environment changes, online optimizing the control instruction sequence based on a deep reinforcement learning model to generate new control instructions adapted to the environmental changes; feeding back the state data, environmental data, and optimization results generated during the task execution to the task model library for updating and improving the preset task mode to improve the execution efficiency of subsequent tasks.

8. An electronic device, characterized in that, Comprising: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Multi-task cooperative processing method and system of industrial vehicle-mounted computer

    CN120448145A

  • Action execution optimization method and device, equipment and storage medium

    CN120725054A

  • Edge cloud resource management collaborative scheduling method based on artificial intelligence

    CN120872528A

  • Intelligent cleaning control method for sleeper mold

    CN121305136A