An unmanned vehicle target search method, device, medium and product

CN118410855BActive Publication Date: 2026-09-25SHANGHAI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410269618.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-08
Publication Date
2026-09-25
Estimated Expiration
2044-03-08

AI Technical Summary

Benefits of technology

本发明提供了一种无人车目标搜索方法、装置、介质及产品,方法包括:获取无人车目标搜索任务;根据功能将所述无人车目标搜索任务分解为若干个基础任务;为每个基础任务构建交互式仿真环境;所述交互式仿真环境支持无人车执行决策动作以及接收相应的观测数据反馈;从所述交互式仿真环境中获取图像和雷达观测数据,并将所述图像和雷达观测数据转化为特征向量,得到无人车观测特征向量;利用任务预测网络根据所述无人车观测特征向量预测当前正在执行的任务,得到任务预测结果;将所述无人车观测特征向量输入到门控循环单元网络中,获得观测时序特征向量;利用多层感知器分别学习每个基础任务,得到若干个策略网络;将所述观测时序特征向量输入至每个策略网络中,得到一组动作;将所述任务预测结果与所述动作做内积,得到当前最佳动作;在所述交互式仿真环境中执行所述当前最佳动作,并获取观测数据、环境奖励和任务完成状态;计算策略网络和任务预测网络的损失,并更新网络参数,直至所述无人车完成所有基础任务。本发明通过构建多个相关的基础任务,使无人车在这些基础任务中进行联合训练,学习到可重用的策略表示,然后将学习到的策略表示应用到目标下游任务中,作为初始策略,跳过复杂目标任务的从零开始学习过程,同时,也可将基础任务策略直接用于目标任务运行时的策略优化,从而提高强化学习算法在复杂控制和决策任务上的性能。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118410855B_ABST
    Figure CN118410855B_ABST
Patent Text Reader

Abstract

The application discloses an unmanned vehicle target search method and device, a medium and a product, relates to the technical field of reinforcement learning and multi-task learning, and decomposes an unmanned vehicle target search task into a plurality of basic tasks according to functions; an interactive simulation environment is constructed, observation data is acquired and is converted into an unmanned vehicle observation feature vector, a task being currently executed is predicted; the observation feature vector is input into a gated recurrent unit network to obtain an observation time sequence feature vector; the observation time sequence feature vector is input into each policy network to obtain a group of actions; the task prediction result and the actions are subjected to inner product to obtain an optimal action and execute the optimal action, acquire observation data, an environment reward and a task completion state, calculate a loss and update network parameters. The application enables the unmanned vehicle to be jointly trained in the basic tasks, learns a reusable policy representation, skips a zero-start learning process of a complex target task, and thus improves the performance of a reinforcement learning algorithm on a complex control and decision task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of reinforcement learning and multi-task learning technology, and in particular to a method, apparatus, medium and product for target search of unmanned vehicles. Background Technology

[0002] Reinforcement learning is an important branch of machine learning that learns to take actions to maximize long-term cumulative rewards in order to complete a specified task through continuous interaction and trial and error with the environment. However, many complex decision-making and control tasks, such as precise grasping and manipulation of robots, autonomous driving navigation in complex urban environments, and game AI defeating top human players in strategy games, still pose significant challenges to current reinforcement learning algorithms.

[0003] Multi-task learning aims to improve learning efficiency and generalization ability by learning multiple related tasks simultaneously. It has been proven to be extremely effective in fields such as computer vision and natural language processing. Introducing the idea of ​​multi-task learning into the field of reinforcement learning and developing systems that can learn from multiple related basic tasks to improve the success rate and stability of unmanned systems on target tasks is of great promise.

[0004] Current research has begun to explore hierarchical training mechanisms in game AI and robot control. This involves pre-training in basic environments and simple tasks, then transferring the training to more complex target scenarios and tasks to accelerate learning. However, these efforts either rely on manually designed task curves or only consider a single pre-training task. How to automatically construct multi-task environments suitable for transfer learning and design training mechanisms remains to be further explored. Furthermore, in applications such as industrial robots and autonomous vehicles, in addition to the complexity of control and decision-making, the generalization ability and stability of algorithms under different environments must also be considered. Therefore, there is an urgent need in this field for a technical solution that can improve the performance of reinforcement learning algorithms on complex control and decision-making tasks. Summary of the Invention

[0005] The purpose of this invention is to provide an unmanned vehicle target search method, device, medium, and product. By constructing multiple related basic tasks, the unmanned vehicle can be jointly trained in these basic tasks to learn reusable policy representations. Then, the learned policy representations are applied to downstream target tasks as initial policies, skipping the learning process from scratch for complex target tasks. At the same time, the basic task policies can also be directly used for policy optimization during the target task runtime, thereby improving the performance of reinforcement learning algorithms on complex control and decision-making tasks.

[0006] To achieve the above objectives, the present invention provides the following solution: In a first aspect, the present invention provides a target search method for unmanned vehicles, the method comprising: Obtain the target search task for the autonomous vehicle; The unmanned vehicle target search task is broken down into several basic tasks based on its functions; An interactive simulation environment is constructed for each basic task; the interactive simulation environment supports the autonomous vehicle in executing decision-making actions and receiving corresponding observation data feedback. Image and radar observation data are acquired from the interactive simulation environment, and the image and radar observation data are converted into feature vectors to obtain unmanned vehicle observation feature vectors; The task prediction network is used to predict the currently executing task based on the observation feature vector of the unmanned vehicle, and the task prediction result is obtained. The unmanned vehicle observation feature vector is input into a gated recurrent unit network to obtain the observation time series feature vector; By using a multilayer perceptron to learn each basic task separately, several policy networks are obtained; The observed temporal feature vectors are input into each policy network to obtain a set of actions; The optimal action is obtained by taking the inner product of the task prediction result and the action. Execute the current optimal action in the interactive simulation environment and acquire observation data, environmental rewards, and task completion status; The loss of the policy network and the task prediction network is calculated, and the network parameters are updated until the autonomous vehicle completes all basic tasks.

[0007] Optional, several basic tasks include: entering the search area, avoiding dynamic obstacles, searching for the target, and leaving the search area.

[0008] Optionally, the step of building an interactive simulation environment for each basic task specifically includes: An entry channel for entering the search area is constructed in the virtual environment for the "enter the search area" task; when the autonomous vehicle correctly enters the target search area through the entry channel, it receives a reward of 1; otherwise, it receives a reward of 0. For the "Avoid Dynamic Obstacles" task, randomly moving obstacles are generated in the target search area; when the autonomous vehicle collides with an obstacle, it receives a reward of -1, and when the autonomous vehicle successfully crosses the target search area, it receives a reward of 1. Generate a target object in the virtual environment for the "search target" task; if the autonomous vehicle finds the target, it receives a reward of 1, otherwise it receives a reward of 0. Construct an exit path from the target search area for the "Leave the search area" task; the autonomous vehicle receives a reward of 1 after safely leaving the target search area, otherwise it receives a reward of 0.

[0009] Optionally, after building an interactive simulation environment for each basic task, the following may also be included: One-Hot coding is performed on each basic task.

[0010] Optionally, after taking the inner product of the task prediction result and the action to obtain the current best action, the method further includes: The observed time-series feature vector is concatenated with the One-Hot encoding of the basic task to obtain the concatenation result; The splicing result is input into the evaluation network to obtain the value estimate of the current state.

[0011] Optionally, the losses of the policy network and the task prediction network are calculated, and the network parameters are updated until the autonomous vehicle completes all basic tasks, specifically including: The loss of the policy network, evaluation network, and task prediction network is calculated, and the network parameters are updated until the autonomous vehicle completes all basic tasks. The losses of the computational policy network, evaluation network, and task prediction network specifically include: The cross-entropy between the task prediction result and the One-Hot encoding is calculated to obtain the loss of the task prediction network; The first step is calculated using the near-end strategy optimization algorithm. The loss of a policy network; The loss value of the evaluation network is calculated using the mean squared error loss.

[0012] Optionally, a task prediction network is used to predict the currently executing task based on the autonomous vehicle's observation feature vector, and the prediction result is obtained, specifically including: A three-layer multilayer perceptron is used to predict the currently executing task based on the observation feature vector of the autonomous vehicle. The multilayer perceptron consists of three fully connected layers, and its output is a task distribution prediction vector. ReLU activation is used between each layer, and a Softmax function is used after the output layer for activation. The formula is as follows: ; i and j represent indices in the vector, n represents the dimension of the vector, x represents the vector output by the hidden layer, and e represents the base of the natural logarithm. and .

[0013] In a second aspect, the present invention provides a computer device comprising: a memory, a processor, and computer instructions stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the unmanned vehicle target search method described in any of the preceding claims.

[0014] In a third aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the unmanned vehicle target search method described in any of the preceding claims.

[0015] In a fourth aspect, the present invention provides a computer program product comprising a computer program that, when executed by a processor, implements the steps of the unmanned vehicle target search method described in any of the preceding claims.

[0016] According to specific embodiments provided by the present invention, the present invention discloses the following technical effects: This invention provides a method, apparatus, medium, and product for unmanned vehicle target search. The method includes: acquiring an unmanned vehicle target search task; decomposing the unmanned vehicle target search task into several basic tasks according to functions; constructing an interactive simulation environment for each basic task; the interactive simulation environment supporting the unmanned vehicle to execute decision-making actions and receive corresponding observation data feedback; acquiring image and radar observation data from the interactive simulation environment, and converting the image and radar observation data into feature vectors to obtain unmanned vehicle observation feature vectors; using a task prediction network to predict the currently executed task based on the unmanned vehicle observation feature vectors to obtain a task prediction result; inputting the unmanned vehicle observation feature vectors into a gated recurrent unit network to obtain an observation time-series feature vector; using a multilayer perceptron to learn each basic task to obtain several policy networks; inputting the observation time-series feature vectors into each policy network to obtain a set of actions; performing an inner product of the task prediction result and the actions to obtain the current optimal action; executing the current optimal action in the interactive simulation environment and acquiring observation data, environmental rewards, and task completion status; calculating the losses of the policy networks and the task prediction network, and updating the network parameters until the unmanned vehicle completes all basic tasks. This invention constructs multiple related basic tasks, enabling autonomous vehicles to perform joint training on these tasks and learn reusable policy representations. The learned policy representations are then applied to downstream target tasks as initial policies, skipping the learning process from scratch for complex target tasks. At the same time, the basic task policies can also be directly used for policy optimization during the target task runtime, thereby improving the performance of reinforcement learning algorithms on complex control and decision-making tasks. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a flowchart of a target search method for unmanned vehicles provided in Embodiment 1 of the present invention.

[0019] Figure 2 The complex target search task provided in Embodiment 1 of the present invention is divided into four basic tasks according to its function.

[0020] Figure 3 The diagram illustrates the steps involved in dividing the complex Adroit Hammer task provided in Embodiment 1 of this invention into two basic tasks. Figure 3 (a) provides a complete overview of the complex task, including all environmental elements such as the robotic arm, hammer, and planks with nails. Figure 3 (b) is the first basic task after being broken down into steps. Its core function is to grab the hammer and put it in place. Figure 3 (c) is the second basic task after being broken down into steps. Its core function is to swing a hammer and knock nails.

[0021] Figure 4 This is a schematic diagram illustrating an example of reusing basic knowledge for unmanned vehicle target search tasks provided in Embodiment 1 of the present invention.

[0022] Figure 5 This is a diagram of the internal structure of a computer device. Detailed Implementation

[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0024] Multi-task learning aims to improve learning efficiency and generalization ability by learning multiple related tasks simultaneously. It has been proven to be extremely effective in fields such as computer vision and natural language processing. Introducing the idea of ​​multi-task learning into the field of reinforcement learning and developing systems that can learn from multiple related basic tasks to improve the success rate and stability of unmanned systems on target tasks is of great promise.

[0025] Current research has begun to explore hierarchical training mechanisms in game AI and robot control. This involves pre-training in basic environments and simple tasks, then transferring the training to more complex target scenarios and tasks to accelerate learning. However, these efforts either rely on manually designed task curves or only consider a single pre-training task. How to automatically construct multi-task environments suitable for transfer learning and design training mechanisms remains to be further explored. Furthermore, in applications such as industrial robots and autonomous vehicles, in addition to the complexity of control and decision-making, the generalization ability and stability of the algorithm in different environments need to be considered. Utilizing multi-task learning for reinforcement transfer can not only enhance generalization ability but also improve the algorithm's adaptation speed to new environments.

[0026] In summary, developing a reinforcement learning algorithm that can be trained on relevant basic tasks and reuse the knowledge learned from these basic tasks to help unmanned systems achieve better results on complex target tasks has significant theoretical value and practical application prospects.

[0027] The purpose of this invention is to provide an unmanned vehicle target search method, device, medium, and product. By constructing multiple related basic tasks, the unmanned vehicle can be jointly trained in these basic tasks to learn reusable policy representations. Then, the learned policy representations are applied to downstream target tasks as initial policies, skipping the learning process from scratch for complex target tasks. At the same time, the basic task policies can also be directly used for policy optimization during the target task runtime, thereby improving the performance of reinforcement learning algorithms on complex control and decision-making tasks.

[0028] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0029] Example 1 like Figure 1As shown, this embodiment provides a target search method for unmanned vehicles. The unmanned vehicle needs to start from a starting position, traverse to the target search area, locate the target object, and then safely leave through an evacuation passage. During this process, the unmanned vehicle needs to avoid randomly moving obstacles. This method first divides the target search task of the unmanned vehicle into four basic tasks: "entering the search area," "avoiding dynamic obstacles," "searching for the target," and "leaving the search area." Then, these basic tasks are simulated and modeled in Unity3D simulation software. Multiple policy networks are used for joint learning on these basic tasks. A task prediction network is also used to predict the corresponding task based on the current state and select the best action, thereby improving the success rate of completing the complex task. The main technical features are: 1) decomposing the complex task into multiple basic tasks; 2) performing joint learning on the basic tasks, including predicting the most relevant task, selecting the optimal action, and evaluating the current state; 3) continuously optimizing the inference process to ultimately improve the success rate of completing the complex task. The specific steps include: S1. Obtain the target search task for the unmanned vehicle.

[0030] S2. Based on the function, the unmanned vehicle target search task is decomposed into several basic tasks.

[0031] The autonomous vehicle target search task is broken down into four basic tasks based on its fundamental functions. The specific process is as follows: Target search task division: According to the division criteria of this embodiment, it is decomposed into four basic tasks based on their "function": "entering the search area," "avoiding dynamic obstacles," "searching for the target," and "leaving the search area." The task set after division is represented as follows: .

[0032] S3. Construct an interactive simulation environment for each basic task; the interactive simulation environment supports the unmanned vehicle in executing decision-making actions and receiving corresponding observation data feedback. An interactive simulation environment is built for each basic task using Unity3D simulation software. The autonomous vehicle can then execute decision-making actions and receive corresponding observation data feedback within these environments. The specific process is as follows: S31. Basic Task Environment Modeling: Using Unity3D simulation software, the autonomous vehicle, physical environment, and interaction logic for each basic task are constructed to closely resemble the real-world environment in appearance and function. Specifically, the "Enter Search Area" task will construct a channel in the virtual environment to enter the search area, and the autonomous vehicle needs to follow the channel to enter the target search area; the "Avoid Dynamic Obstacles" task will randomly generate randomly moving obstacles in the target search area, and the autonomous vehicle needs to pass through the target search area without collision; the "Search Target" task will generate a target object in the virtual environment, and the autonomous vehicle will explore the environment to find this target object; the "Leave Search Area" task will construct a channel to leave the target search area, and the autonomous vehicle needs to leave the target search area through this channel.

[0033] S32. Basic Task Reward Settings: Use the Unity ML-Agents Toolkit software to set the reward function for the basic task environment. Specifically, in the "Enter Search Area" task, the autonomous vehicle will receive a reward of 1 if it correctly enters the target search area, and a reward of 0 otherwise; in the "Avoid Dynamic Obstacles" task, the autonomous vehicle will receive a reward of -1 if it collides with an obstacle, and a reward of 1 if it successfully crosses the target search area; in the "Search Target" task, the autonomous vehicle will receive a reward of 1 if it finds the target, and a reward of 0 otherwise; in the "Leave Search Area" task, the autonomous vehicle will receive a reward of 1 if it safely leaves the target search area, and a reward of 0 otherwise.

[0034] After building an interactive simulation environment for each basic task, the process also includes One-Hot coding for each basic task.

[0035] Each basic task is one-hot encoded as an independent representation of the task. The specific process is as follows: One-Hot coding is applied to each basic task: Each basic task is assigned a unique One-Hot code, which is used to classify the tasks. The corresponding One-Hot encoding is denoted as Specifically, the task code for "entering the search area" is as follows: The task code for "avoiding dynamic obstacles" is as follows: The "search target" task is coded as follows: "Leave the search area" is coded as .

[0036] S4. Obtain image and radar observation data from the interactive simulation environment, and convert the image and radar observation data into feature vectors to obtain unmanned vehicle observation feature vectors.

[0037] Image and radar observation data are acquired from the simulation environment constructed in step S3 and transformed into feature vectors through a feature encoding network. The specific process is as follows: S41. Acquire unmanned vehicle observation data: Acquire unmanned vehicle observation data using Unity ML-Agents Toolkit software, denoted as... ,in Represents all possible observation data. Indicates task exist In actual observations at any given moment, the dimensions and types of observation information from all basic tasks are consistent.

[0038] S42. Obtain the observation feature vector of the unmanned vehicle: Observation data It can be an image, a vector, or a combination of both. In the basic task of "unmanned vehicle target search task", the observation data consists of (84,84,3)-dimensional image data and (400,)-dimensional single-line radar data. Its feature encoding network is denoted as... ,in The network parameters consist of a 3-layer 2D convolutional network encoding image data and a 3-layer 1D convolutional network encoding radar data. A fully connected network follows both to reduce the dimensionality to (64,) dimensions. The (64,)-dimensional image features and (64,)-dimensional radar features are then concatenated to obtain a (128,)-dimensional feature vector, represented as: (Followed by a fully connected network for dimensionality reduction to (64,) dimensions, 64*2 = 128 dimensions, which are then combined with (64,) dimensional image features and (64,) dimensional radar features.) For image data, (84, 84, 3) indicates that the image has a height of 84 pixels, a width of 84 pixels, and 3 channels (usually RGB color channels). This representation is very common in computer vision and can accurately describe the structure and color information of an image. For single-line radar data, (400,) indicates that the data is a one-dimensional array with 400 elements. This representation is common when processing single-line radar data, where each element may represent a different feature or observation.

[0039] S5. Using the task prediction network, predict the currently executing task based on the unmanned vehicle's observation feature vector to obtain the task prediction result; The autonomous vehicle's observed feature vectors are input into a task prediction network to predict the currently executing task, which is then used for subsequent action selection. The specific process is as follows: Autonomous Vehicle Basic Task Distribution Prediction: Task Prediction Network ,in The network parameters are defined as follows: It is a 3-layer multilayer perceptron. Its input layer is a (128,)-dimensional fully connected network that processes the feature vectors observed by the autonomous vehicle. The hidden layers are also (128,)-dimensional fully connected networks used to learn higher-order features from the data. The output layer is a (4,)-dimensional fully connected network that generates the selection probabilities for each basic task. ReLU activation is used between each layer, with the following formula: (1) Finally, the softmax function is used for activation after the output layer, with the following formula: (2) Where i and j represent indices in the vector, n represents the dimension of the vector, x represents the vector output by the hidden layer, and e represents the base of the natural logarithm. and .

[0040] The task prediction result is a constant vector, such as The prediction results are expressed as This is used for selecting the best action later.

[0041] S6. Input the unmanned vehicle observation feature vector into the gated recurrent unit network to obtain the observation time series feature vector.

[0042] The feature vector observed by the autonomous vehicle is input into a gated recurrent unit network to give it temporal characteristics. The specific process is as follows: Unmanned vehicle observation time-series feature vector extraction: The gated recurrent unit used in this embodiment has a (128,)-dimensional hidden vector as its input. , and (128,)-dimensional observation data feature vector The output is the observation time-series feature vector (128,). The gated recurrent unit includes a reset gate and an update gate to control the retention of current input information and the forgetting of previously memorized information. The output of the update gate... The calculation formula is: (3) in, and All of these are updates to the weights. For the task exist The feature vector of the observed data at time t, For the task exist The hidden vector at time step, This is the sigmoid activation function. Reset the gate's output. The calculation formula is: (4) in, and All of these are reset weights. For the task exist The feature vector of the observed data at time t, For the task exist The hidden vector at each time step. The hidden vector output by the gating unit. The calculation formula is: (5) in, To update the gate output, For the task exist Hidden vector at time step, candidate hidden state The calculation formula is: (6) in and All are candidate weights. For the task exist The feature vector of the observed data at time t, To reset the gate's output, For the task exist The hidden vector at time step 1, where tanh is the hyperbolic tangent activation function. The Hadamard product is used. Finally, the temporal feature vectors are observed. The calculation formula is: (7) Its dimension is (128,), where tanh is the hyperbolic tangent activation function. It is the hidden vector output by the gating unit.

[0043] S7. Use a multilayer perceptron to learn each basic task separately to obtain several policy networks.

[0044] S8. Input the observed time-series feature vector into each policy network to obtain a set of actions.

[0045] S9. Take the inner product of the task prediction result and the action to obtain the current best action.

[0046] The observation time series feature vector obtained in step S6 The policy network for all basic tasks is input to generate a set of possible actions. The inner product of these actions and the task prediction results obtained in step S5 is calculated, and the result is taken as the optimal action. The specific process is as follows: A1. Autonomous Vehicle Policy Network: The policy network group is a set of multilayer perceptions, represented as... ,in The network parameters consist of four policy networks, which learn four basic tasks: "entering the search area", "avoiding dynamic obstacles", "searching for the target", and "leaving the search area".

[0047] A2. Calculate the possible actions of the autonomous vehicle: Calculate the observation time-series feature vector obtained in step S6. The input is fed into each policy network to obtain a set of possible actions. , recorded as The continuous action with dimensions (4,2) represents the magnitude of the autonomous vehicle's "forward" and "turn" in the four basic tasks.

[0048] A3. Calculate the optimal maneuver for the autonomous vehicle: Use the task prediction results... With possible action groups Calculate the inner product to obtain the current optimal action. The formula is: (8) in Indicates the first There are 1 task, where t represents the current time. This indicates the basic task currently being executed. Assume the output action is: Therefore, the final calculation result is This action will be considered the best action. Executed in the basic tasks.

[0049] After taking the inner product of the task prediction result and the action to obtain the current best action, the method further includes: splicing observation time series feature vectors One-Hot coding for basic tasks The value of the current state is estimated and fed into the evaluation network to guide the policy network and task prediction network in better completing the target search task. The specific process is as follows: Vector concatenation: combining time-series feature vectors With One-Hot vector By concatenating the vectors, we obtain the following new vector: ,in The vector concatenation operation is used as input to the evaluation network. In the basic task of "unmanned vehicle target search task", the dimension of the new vector is (132,).

[0050] Estimating Environmental State Value: Evaluating Networks ,in The network parameters are defined by a multilayer perceptron, whose input is a (132,)-dimensional vector and whose output is a (1,)-dimensional evaluation vector. This can be represented as: (8) This vector is for the task. ,exist The state assessment at each moment is used in subsequent calculations to guide the policy network and task prediction network in completing the target search task.

[0051] S10. Execute the current best action in the interactive simulation environment and obtain observation data, environmental rewards, and task completion status.

[0052] The task constructed in step S3 Execute the selected optimal action in the simulation environment and collect the next time step data. Observational data Environmental rewards and task completion status Steps S4 to S10 are performed to iterate through all basic tasks, repeating this process 512 times. Each iteration records the observation data, optimal action, environmental reward, the next time-instance observation, and the task prediction result. The specific process is as follows: Autonomous vehicle action execution: The optimal actions are executed in the corresponding virtual simulation environment using Unity ML-Agents Toolkit software. For example, the actions output by the observations corresponding to the basic "target search" task will also be executed in the "target search" simulation environment, and the next observation data after the action execution will be returned. Environmental awards and task completion status .

[0053] Autonomous vehicle training data collection: Steps S4 to S10 are performed to traverse all basic tasks, enabling the autonomous vehicle policy network to learn from all basic tasks. This process is repeated 521 times, and the observation data, best action, environmental reward, next-time observation, and task prediction results are recorded in the experience pool for subsequent network updates.

[0054] S11. Calculate the loss of the policy network and the task prediction network, and update the network parameters until the autonomous vehicle completes all basic tasks.

[0055] The losses of the policy network, value network, and task prediction network are calculated, and the network parameters are updated using the backpropagation algorithm. The data collection and training process is repeated until the autonomous vehicle can stably complete all basic tasks. The specific process is as follows: Calculate the task prediction network loss: Calculate the cross-entropy between the task prediction result and the One-Hot label, use this as the loss value for the task prediction network, and set the coefficients. This is used to control the weight of cross-entropy loss in the overall loss, and the coefficient is... The initial value is 1, and it is gradually decreased to approach 0 as the number of training iterations increases. This allows the network to autonomously adjust its task based on rewards, even when different base tasks have the same observations, thus improving the flexibility of the task prediction network. Task prediction network loss value The calculation formula is: (7) in, This represents the parameters of the task prediction network. To control the weight of cross-entropy loss in the overall loss, Indicates a time step. Indicates the number of tasks. Indicates at time step Medium task The true label, Indicates at time step Medium task Task prediction vector, This indicates that for all time steps and tasks Summation, Let represent the natural logarithm function. Calculate the policy network loss value: Calculate the th... Loss value of each policy network Its formula is: (8) in This represents the parameters of the i-th policy network. Indicates a time step. Indicates time step Perform the desired operation. This indicates a truncation operation on the value, restricting it to a range. Inside, For gradient clipping hyperparameters, This indicates taking the smaller value. The importance weight is calculated using the following formula: (9) in, Indicates the first The parameters of each strategy Indicates the state Below, according to the policy network Get action The probability, Indicates the state Below, based on the old strategy network Get action The probability of that. Generalized dominance function. The calculation formula is: (10) in, Indicates a time step. Indicates the next time step. To reward the discount exceeding the parameters, For the generalized dominance function hyperparameter, in the basic tasks divided into "unmanned vehicle target search task" , , The timing difference error is calculated using the following formula: (11) in, Indicates a time step. Represents the parameters of the value network. Indicates in Momentary environmental rewards This indicates that the reward discount exceeds the parameter. Indicates at time step In the middle, value function networks are used The obtained state value estimate, Indicates at time step In the middle, value function networks are used The obtained state value estimate. The overall loss of the policy network group is the sum of the losses of all policy networks, expressed as: The calculation formula is as follows: (12) in, This represents the parameters of the entire policy network group. Indicates the first The parameters of the policy network represent the... The loss of a policy network, This indicates the number of policy networks in the policy network group. This indicates a summation calculation. Calculate the evaluation network loss value: Calculate the loss value of the evaluation network using the mean squared error loss. The calculation formula is as follows: (13) in, Represents the parameters of the value network. Indicates the number of tasks. This indicates a summation calculation. This represents the state value estimate obtained using the evaluation network φ. Indicates the first Time step in each task The observation, Indicates the first Time step in each task The true value of the state, Indicates the first Time step in each task The advantage function. Overall network loss. This is the sum of the losses from each part, calculated as follows: (14) in, , This represents the parameters of the task prediction network. The parameters represent the evaluation network. This represents the task prediction loss value. This represents the overall loss value of the policy network. This represents the evaluation of the network loss value. Backpropagation is performed based on the overall loss value to update... Network parameters. Repeat data collection and network parameter updates until the autonomous vehicle can stably complete all basic tasks.

[0056] After training, the autonomous vehicle can autonomously select the most suitable strategy for the current moment through the task prediction network, and can stably complete the target search task.

[0057] Experimental Description and Results: To verify the effectiveness of the proposed method, two virtual simulation environments, "Unmanned Vehicle Target Search" and "AdroitHammer," were constructed. The tasks and their divisions are as follows: Figure 2 and Figure 3 (a) Figure 3 (b) Figure 3 As shown in (c), the experiment compared the proposed method with traditional reinforcement learning methods such as PPO, MTRL, and EAR. The results are shown in Table 1. The proposed method significantly outperforms the comparison methods in terms of success rate and stability on complex tasks. Specifically, the PPO algorithm, trained directly on complex tasks, struggles to make progress; the EAR algorithm is prone to interference between different simple tasks; and while the MTRL algorithm can learn simple tasks, it cannot effectively distinguish between tasks, thus making it difficult to complete complex tasks. In contrast, the method of this invention can clearly distinguish between different tasks, such as... Figure 4 As shown, this method effectively utilizes knowledge acquired from basic tasks, thereby significantly improving performance in complex task environments. This result strongly demonstrates the advantages of this method in handling complex tasks.

[0058] Table 1 Comparison Results

[0059] Example 2 A computer device, which may be a database, may have an internal structure diagram as shown below. Figure 5 As shown, the computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores pending transactions. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements the unmanned vehicle target search method in Embodiment 1.

[0060] Example 3 A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the unmanned vehicle target search method in Embodiment 1.

[0061] Example 4 A computer program product includes a computer program that, when executed by a processor, implements the steps of the unmanned vehicle target search method in Embodiment 1.

[0062] It should be noted that the object information (including but not limited to object device information, object personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this invention are all information and data authorized by the object or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0063] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided by this invention can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided by this invention may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided by this invention may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0064] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0065] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A target search method for unmanned vehicles, characterized in that, The method includes: S1. Obtain the target search task for the unmanned vehicle; S2. Based on the function, the unmanned vehicle target search task is decomposed into four basic tasks: "entering the search area", "avoiding dynamic obstacles", "searching for the target" and "leaving the search area". S3. Construct an interactive simulation environment for each basic task; the interactive simulation environment supports the unmanned vehicle in executing decision-making actions and receiving corresponding observation data feedback. S4. Obtain image and radar observation data from the interactive simulation environment, and convert the image and radar observation data into feature vectors to obtain unmanned vehicle observation feature vectors; S5. Using the task prediction network, predict the currently executing task based on the unmanned vehicle's observation feature vector to obtain the task prediction result; S6. Input the unmanned vehicle observation feature vector into the gated recurrent unit network to obtain the observation time series feature vector; S7. Use a multilayer perceptron to learn each basic task separately to obtain several policy networks; S8. Input the observed time-series feature vector into each policy network to obtain a set of actions; S9. Take the inner product of the task prediction result and the set of actions to obtain the current best action; S10. Execute the current best action in the interactive simulation environment and obtain observation data, environmental rewards and task completion status; then traverse all basic tasks according to steps S4 to S10, so that the autonomous vehicle policy network can learn in all basic tasks, and repeat 512 times, and record the observation data, best action, environmental rewards, next time observation and task prediction results in the experience pool for subsequent network updates. S11. Calculate the loss of the policy network and the task prediction network, and update the network parameters until the autonomous vehicle completes all basic tasks.

2. The unmanned vehicle target search method according to claim 1, characterized in that, The construction of an interactive simulation environment for each basic task specifically includes: For the "enter the search area" task, an entry channel to the search area is constructed in the virtual environment; when the autonomous vehicle correctly enters the target search area through the entry channel, it receives a reward of 1; otherwise, it receives a reward of 0. For the "Avoid Dynamic Obstacles" task, randomly moving obstacles are generated in the target search area; when the autonomous vehicle collides with an obstacle, it receives a reward of -1, and when the autonomous vehicle successfully crosses the target search area, it receives a reward of 1. Generate a target object in the virtual environment for the "search target" task; if the autonomous vehicle finds the target, it receives a reward of 1, otherwise it receives a reward of 0. Construct an exit path from the target search area for the "Leave the search area" task; the autonomous vehicle receives a reward of 1 after safely leaving the target search area, otherwise it receives a reward of 0.

3. The unmanned vehicle target search method according to claim 1, characterized in that, After building an interactive simulation environment for each basic task, it also includes: One-Hot coding is performed on each basic task.

4. The unmanned vehicle target search method according to claim 3, characterized in that, After taking the inner product of the task prediction result and the action to obtain the current best action, the method further includes: The observed time-series feature vector is concatenated with the One-Hot encoding of the basic task to obtain the concatenation result; The splicing result is input into the evaluation network to obtain the value estimate of the current state.

5. The unmanned vehicle target search method according to claim 4, characterized in that, Calculate the loss of the policy network and the task prediction network, and update the network parameters until the autonomous vehicle completes all basic tasks, specifically including: The loss of the policy network, evaluation network, and task prediction network is calculated, and the network parameters are updated until the autonomous vehicle completes all basic tasks. The losses of the computational policy network, evaluation network, and task prediction network specifically include: The cross-entropy between the task prediction result and the One-Hot encoding is calculated to obtain the loss of the task prediction network; The first step is calculated using the near-end strategy optimization algorithm. The loss of a policy network; The loss value of the evaluation network is calculated using the mean squared error loss.

6. The unmanned vehicle target search method according to claim 1, characterized in that, The task prediction network is used to predict the currently executing task based on the observation feature vector of the autonomous vehicle, and the prediction result is obtained, specifically including: A three-layer multilayer perceptron is used to predict the currently executing task based on the observation feature vector of the autonomous vehicle. The multilayer perceptron consists of three fully connected layers, and its output is a task distribution prediction vector. ReLU activation is used between each layer, and a Softmax function is used after the output layer for activation. The formula is as follows: ; i and j represent indices in the vector, n represents the dimension of the vector, x represents the vector output by the hidden layer, and e represents the base of the natural logarithm. and This represents the operation of the exponential function.

7. A computer device, comprising: The memory and processor contain computer instructions stored in the memory and executable on the processor, characterized in that the processor executes the computer instructions to implement the steps of the unmanned vehicle target search method according to any one of claims 1-6.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the steps of the unmanned vehicle target search method according to any one of claims 1-6.

9. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the steps of the unmanned vehicle target search method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Unmanned chariot target search strategy generation method

    CN115423214A

  • Integration-based cooperative multi-agent deep reinforcement learning method

    CN116468107A