Target data association method based on reinforcement learning
Through the target data association method based on reinforcement learning, the correlation probability and reward value are calculated using new object detection, ring wavegate screening and Bayesian network, the problem of traditional data association algorithms being affected by the system model is solved, and higher accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510380324.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-25
AI Technical Summary
In traditional target tracking technology, data association algorithms are susceptible to unknown or inaccurate target system models, resulting in difficult to guarantee the accuracy of correlation results and poor tracking performance.
Using a target data association method based on reinforcement learning, a new target detection, annular wavegate screening, long-term and short-term memory networks and Bayesian networks are used to calculate the correlation probability and reward value, and an adjustment mechanism that automatically corrects point trace association errors is designed to decouple the target tracking and filtering process.
It realizes learning target motion information in a small amount of prior information environment, improves the accuracy and robustness of data association, has a wide range of application and high application value.
Smart Images

Figure CN120372439A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a target data association method based on reinforcement learning, which is particularly suitable for dealing with target data association problems in the field of information fusion and belongs to the target data association technology. Background Art
[0002] The principle of data association technology is to determine the correct tracks and trajectories by establishing the relationship between radar measurement data at a certain moment and measurement data at other moments. If the association result deviates greatly from the actual situation, it may directly lead to a large error in the filtering estimation, thus affecting the tracking accuracy. For traditional target tracking technology, filtering and association are inseparable, promoting and restricting each other. Therefore, most data association algorithms include a filtering process, such as the global nearest neighbor association filtering algorithm, the probabilistic data association filtering algorithm, the joint probabilistic data association filtering algorithm, and so on. The basic idea of these algorithms is to first predict the position of the target at the next moment based on the state of the target at the current moment, then find the most suitable track according to the predicted position according to a certain rule, and finally combine the associated track to update the target state to obtain the filtering estimation. Although some practical problems of target tracking can be solved through this idea, it also has a natural defect, that is, the state prediction needs to rely on the system model to achieve, and the association result is directly affected by the predicted position. If the system model is unknown or inaccurate at this time, the accuracy of the association result cannot be guaranteed. In fact, in the actual application process, the motion state of the target is controlled by people, and the system model of the target is very difficult to predict, so the accuracy of the association result is very difficult to guarantee, and the tracking performance is poor.
[0003] In recent years, the development of artificial intelligence technology has advanced by leaps and bounds, and most data analysis and processing methods have begun to use artificial intelligence means. Especially the reinforcement learning technology, which has the ability to learn autonomously from the environment. Applied to the data association field, if the reinforcement learning technology can learn the motion information of the target from the environment, the data association can be decoupled from the tracking process and no longer affected by the filtering process. However, how to design a target track association architecture based on reinforcement learning, which can learn the system model from an environment with a small amount of prior information and achieve accurate association between the target and the track is an actual problem that needs to be solved urgently. Summary of the Invention
[0004] A target data association method based on reinforcement learning of the present invention can effectively solve the problems that classical data association algorithms are easily affected by factors such as the target system model and strong clutter by establishing a new data association network architecture based on the reinforcement learning framework.
[0005] The target data association method based on reinforcement learning of the present invention is characterized in that it includes the following steps:
[0006] Step 1: Based on the state at the current moment, perform new target detection on the sensor detection area;
[0007] Step 2: According to the velocity threshold of the target, design a circular wave gate screening mechanism to process the measurement set obtained at the next sampling moment, and screen out the possible selected traces for each target;
[0008] Step 3: Input each target and the corresponding trace screening result into the long short-term memory network to obtain the association probability between the target and the trace, and form an association probability matrix;
[0009] Step 4: Based on the association probability, select the trace associated with each target to obtain the association probability;
[0010] Step 5: Combine each target and the corresponding trace, input them into the reward function to obtain the reward value;
[0011] Step 6: Based on the target state, corresponding trace and reward value of this process, enter the learning session;
[0012] Step 7: Further process the data and enter the next loop process.
[0013] Preferably, the specific steps of Step 1 are as follows:
[0014] Step 1.1: Select the measurement sets obtained at adjacent 7 sampling moments for processing;
[0015] Step 1.2: If there are existing targets at this time, remove the target-related traces in these 7 measurement sets; otherwise, do not perform any processing;
[0016] Step 1.3: Process the measurement set obtained in the previous step, and extract all possible target track information according to the velocity threshold of the target;
[0017] Step 1.4: If there is target track information, based on the criteria that "one target generates only one trace" and "one trace is only related to one target", group the obtained target track information and enter Step 1.5; otherwise, directly enter Step 2;
[0018] Step 1.5: Calculate the variance value of the 7 traces in each track of each group, and select the track information with the smallest variance value as the new target;
[0019] Preferably, the specific steps of Step 4 are as follows:
[0020] Step 4.1: If it is in the training state at this time, enter Step 4.2; otherwise, if it is in the test state, enter Step 4.3;
[0021] Step 4.2: Each target randomly selects a track from the tracks screened in Step 2 and outputs the association probability;
[0022] Step 4.3: Referring to the association probability matrix obtained in Step 3, each target selects the track with the highest association probability therefrom and outputs the corresponding association probability;
[0023] Preferably, the specific steps of Step 5 are as follows:
[0024] Step 5.1: Select a Bayesian network and set it to a five-classification mode;
[0025] Step 5.2: Input the state into the network in Step 5.1, output the probabilities of different classifications, and select the classification number with the largest probability value;
[0026] Step 5.3: Using the classification number obtained in Step 5.2 as the order number, perform prediction by the least squares method of the corresponding order number to obtain the predicted value of the target position;
[0027] Step 5.4: Based on the predicted value obtained in Step 5.3 and the track position, calculate the reward value through the Bayesian recursion function;
[0028] Preferably, the specific steps of Step 6 are as follows:
[0029] Step 6.1: If it is in the training state at this time, go to Step 6.2; otherwise, if it is in the test state, go to Step 6.3;
[0030] Step 6.2: Feed back the state, track, and reward value of each target to the long short-term memory network for training and learning;
[0031] Step 6.3: Save the state, track, and reward value of each target;
[0032] Preferably, the specific steps of Step 7 are as follows:
[0033] Step 7.1: For the training state, collect the measurement set obtained by sampling in the next cycle and directly go back to Step 1;
[0034] Step 7.2: For the test state, check whether there is a target association interruption. If so, first perform adaptive adjustment and then go to Step 1; otherwise, directly go to Step 1;
[0035] Preferably, Step 7.2 specifically includes the following sub-steps:
[0036] Step 7.2.1: If entering the adaptive adjustment link, input the state of the target and all the tracks screened in Step 2 into the reward function to obtain a set of reward values;
[0037] Step 7.2.2: Compare the reward value of the selected trace with that of other traces.
[0038] Step 7.2.3: If the reward value of the selected trace is the largest, it is considered that the selected trace is correct, and the adaptive adjustment process ends, then enter Step 1. Otherwise, enter Step 7.2.4.
[0039] Step 7.2.4: If the reward value of the selected trace is not the largest, then select all the traces whose reward values are greater than this reward value.
[0040] Step 7.2.5: Continue to screen out the traces that have not been selected by other targets from the selected traces to form a set, and sort them in descending order of the reward value.
[0041] Step 7.2.6: If the trace set is empty, enter Step 1. Otherwise, select the first trace to replace the trace originally selected by the target, and then enter Step 1.
[0042] A target data association method based on reinforcement learning according to the present invention specifically includes the following technical measures: First, based on the criteria of "one target generates only one trace" and "one trace is only related to one target", all data in adjacent 7 sampling intervals are traversed for new target detection. Then, according to the velocity threshold, a circular wave gate screening mechanism is set. Next, the association probability between the target and the trace is obtained through a long short-term memory network, and then the trace is selected. Then, in combination with the Bayesian network and the Bayesian recursion function, the reward value of the selected trace is calculated. Finally, an adjustment mechanism that can automatically correct the trace association error is designed to solve the problem of weak transfer ability of reinforcement learning. A target data association method based on reinforcement learning proposed by the present invention can first use some measurement data for training and then test based on this, and has the advantages of wide application range, strong robustness, high application value, etc. The proposed data association method can be directly used to solve the corresponding practical problems. Brief Description of the Drawings
[0043] Figure 1 is the overall block diagram of a target data association method based on reinforcement learning according to the present invention;
[0044] Figure 2 is the circular wave gate screening mechanism diagram of a target data association method based on reinforcement learning according to the present invention. Detailed Embodiment
[0045] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments.
[0046] Embodiment 1
[0047] A target data association method based on reinforcement learning in this embodiment is shown in the attached Figure 1-2 , including the following steps:
[0048] Step 1: Based on the current state, new targets are detected in the sensor detection area;
[0049] Step 1.1: Select the measurement set obtained at 7 adjacent sampling moments for processing;
[0050] Step 1.2: If the target already exists at this time, the target-related points in the 7 measurement sets are removed; otherwise, no processing is performed;
[0051] Step 1.3: Process the measurement set obtained in the previous step and extract all possible target track information according to the target speed threshold;
[0052] Step 1.4: If there is target track information, then based on the principle of "one target only generates one point track" and "one point track is only related to one target", the obtained target track information is grouped and the process goes to step 1.5; otherwise, the process goes directly to step 2;
[0053] Step 1.5: Calculate the variance of the 7 points in each track in each group, and select the track information with the smallest variance as the new target;
[0054] Step 2: According to the speed threshold of the target, a circular wave gate screening mechanism is designed to process the measurement set obtained at the next sampling time and screen out the possible points for each target;
[0055] Step 3: Input each target and the corresponding point trace screening results into the long short-term memory network to obtain the association probability between the target and the point trace, and form an association probability matrix;
[0056] Step 4: Select the point trace associated with each target based on the association probability to obtain the association probability;
[0057] Step 4.1: If it is in the training state, go to step 4.2; otherwise, if it is in the testing state, go to step 4.3;
[0058] Step 4.2: Each target randomly selects a point from the points selected in step 2 and outputs the association probability;
[0059] Step 4.3: Compare the association probability matrix obtained in step 3, select the point trace with the largest association probability for each target, and output the corresponding association probability;
[0060] Step 5: Combine each target and the corresponding point trace, input the reward function, and get the reward value;
[0061] Step 5.1: Select a Bayesian network and set it to a five-classification mode;
[0062] Step 5.2: Input the status into the network in Step 5.1, output the probabilities of different classifications, and select the classification number with the largest probability value;
[0063] Step 5.3: Using the classification number obtained in Step 5.2 as the order, perform prediction using the least squares method of the corresponding order to obtain the predicted value of the target position;
[0064] Step 5.4: Based on the predicted value and the track position obtained in Step 5.3, calculate the reward value through the Bayesian recursive function;
[0065] Step 6: Based on the target status, corresponding tracks, and reward values of this process, enter the learning session;
[0066] Step 6.1: If it is in the training state at this time, enter Step 6.2. Otherwise, if it is in the test state, enter Step 6.3;
[0067] Step 6.2: Feed back the status, tracks, and reward values of each target to the long short-term memory network for training and learning;
[0068] Step 6.3: Save the status, tracks, and reward values of each target;
[0069] Step 7: Further process the data and enter the next loop process;
[0070] Step 7.1: For the training state, collect the measurement set obtained from the next cycle sampling and directly enter Step 1 again;
[0071] Step 7.2: For the test state, check if there is a target association interruption. If so, first perform adaptive adjustment and enter Step 7.3. Otherwise, directly enter Step 1;
[0072] Step 7.3: Input the status of the target and all the tracks selected from Step 2 into the reward function to obtain a set of reward values;
[0073] Step 7.4: Compare the reward value of the selected track with the reward values of other tracks;
[0074] Step 7.5: If the reward value of the selected track is the largest, it is considered that the selected track is correct, and the adaptive adjustment session ends. Enter Step 1. Otherwise, enter Step 7.6;
[0075] Step 7.6: If the reward value of the selected track is not the largest, select all the tracks with reward values greater than this reward value;
[0076] Step 7.7: Continue to select the points that have not been selected by other targets from the selected points to form a set, and arrange them in order from large to small according to the reward value;
[0077] Step 7.8: If the point set is empty, go to step 1. Otherwise, select the first point to replace the originally selected point of the target, and then go to step 1.
[0078] Example 2
[0079] To better illustrate the present invention, the following uses the measurement data of a certain type of radar as a specific embodiment to describe the steps of the present invention in detail:
[0080] Step 11: Based on the current state, new target detection is performed in the sensor detection area;
[0081] Step 11.1: Select a set of measurements obtained at 7 consecutive sampling moments {Z t-6 ,Z t-5 ,...,Z t-1 ,Z t} for processing;
[0082] Step 11.2: If the target already exists Then the target related points in these 7 measurement sets are removed; otherwise, no processing is done;
[0083] Step 11.3: Process the measurement set obtained in the previous step and extract all possible target track information based on the target speed threshold.
[0084] Step 11.4: If there is target track information, then based on the principle of "one target only generates one point track" and "one point track is only related to one target", the obtained target track information is grouped and the process goes to step 11.5; otherwise, the process goes directly to step 12;
[0085] Step 11.5: Calculate each track in each group The variance of the 7 points Select the track information with the smallest variance value as the new target;
[0086] Step 12: According to the speed threshold of the target, a ring wave gate screening mechanism is designed to process the measurement set obtained at the next sampling time and screen out the possible points for each target. If the sampling interval of the sensor is T_sample, the measurement data is z, the maximum speed v_max and the minimum speed v_min of the target movement, the ring wave gate screening formula is set to screen out the points that meet the speed threshold, that is,
[0087]
[0088] Step 13: Input each target and the corresponding point track screening result into the long short-term memory network to obtain the association probability between the target and the point track, and form an association probability matrix
[0089] Step 14: Select the point tracks associated with each target based on the association probability to obtain the association probability;
[0090] Step 14.1: If it is in the training state at this time, go to Step 14.2; otherwise, if it is in the test state, go to Step 14.3;
[0091] Step 14.2: Each target randomly selects a point track from the point tracks screened in Step 12 and outputs the association probability;
[0092] Step 14.3: Refer to the association probability matrix P obtained in Step 13 t , and each target selects the point track with the maximum association probability from it and outputs the corresponding association probability;
[0093] Step 15: Combine each target and the corresponding point track, input the reward function, and obtain the reward value;
[0094] Step 15.1: Select the Bayesian network and set it to the five-classification mode;
[0095] Step 15.2: Input the state into the network in Step 15.1, output the probabilities of different classifications, and select the classification number g with the maximum probability value;
[0096] Step 15.3: Use the least squares method of the corresponding order with the classification number g obtained in Step 15.2 for prediction to obtain the predicted value of the target position;
[0097] Step 15.4: Calculate the reward value through the Bayesian recursion function based on the predicted value obtained in Step 15.3 and the point track position That is
[0098]
[0099] Among them, represents the i-th point track selected from the measurement set at time t; K t represents the clutter intensity at time t, that is num t is the number of point tracks in the measurement set, and TS is the area of the radar detection area; R is the measurement covariance matrix, which is determined by the radar measurement error; P_D is the detection probability.
[0100] Step 16: Based on the target state, the corresponding point track, and the reward value of this process, enter the learning link;
[0101] Step 16.1: If it is in the training state at this time, go to Step 16.2; conversely, if it is in the test state, go to Step 16.3;
[0102] Step 16.2: Feed back the state, track, and reward value of each target to the long short-term memory network for training and learning;
[0103] Step 16.3: Save the state, track, and reward value of each target;
[0104] Step 17: Further process the data and enter the next loop;
[0105] Step 17.1: For the training state, collect the measurement set obtained by sampling in the next cycle and directly go back to Step 11;
[0106] Step 17.2: For the test state, check if there is a target association interruption. If so, first perform adaptive adjustment and go to Step 17.3; conversely, directly go to Step 11;
[0107] Step 17.3: Input the state of the target and all the tracks filtered from Step 12 into the reward function to obtain a set of reward values
[0108] Step 17.4: Compare the reward value of the selected track with the reward values of other tracks;
[0109] Step 17.5: If the reward value of the selected track is the largest, it is considered that the selected track is correct, and the adaptive adjustment process ends. Go to Step 11; conversely, go to Step 17.6;
[0110] Step 17.6: If the reward value of the selected track is not the largest, select all the tracks with a reward value greater than this reward value;
[0111] Step 17.7: Continue to filter out the tracks that have not been selected by other targets from the selected tracks to form a set, and sort them in descending order of the reward value;
[0112] Step 17.8: If the track set is empty, go to Step 11; conversely, select the first track to replace the track originally selected by the target, and then go to Step 11.
[0113] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.
Claims
1. A target data association method based on reinforcement learning, characterized in that The following steps are involved: Step 1: Based on the current state, new targets are detected in the sensor detection area; Step 2: According to the speed threshold of the target, a circular wave gate screening mechanism is designed to process the measurement set obtained at the next sampling time and screen out the possible points for each target; Step 3: Input each target and the corresponding point trace screening results into the long short-term memory network to obtain the association probability between the target and the point trace, and form an association probability matrix; Step 4: Select the point trace associated with each target based on the association probability to obtain the association probability; Step 5: Combine each target and the corresponding point trace, input the reward function, and get the reward value; Step 6: Based on the target state, corresponding point traces and reward values of this process, enter the learning phase; Step 7: Further process the data and enter the next cycle.
2. The method for target data association based on reinforcement learning according to claim 1, wherein The specific steps of step 1 are: Step 1.1: Select the measurement set obtained at 7 adjacent sampling moments for processing; Step 1.2: If the target already exists at this time, the target-related points in the 7 measurement sets are removed; otherwise, no processing is performed; Step 1.3: Process the measurement set obtained in the previous step and extract all possible target track information according to the target speed threshold; Step 1.4: If there is target track information, then based on the principle of "one target only generates one point track" and "one point track is only related to one target", the obtained target track information is grouped and the process goes to step 1.5; otherwise, the process goes directly to step 2; Step 1.5: Calculate the variance of the 7 points in each track in each group, and select the track information with the smallest variance as the new target.
3. The object data association method based on reinforcement learning according to claim 1, wherein The specific steps of step 4 are: Step 4.1: If it is in the training state, go to step 4.2; otherwise, if it is in the testing state, go to step 4.3; Step 4.2: Each target randomly selects a point from the points selected in step 2 and outputs the association probability; Step 4.3: Compare the association probability matrix obtained in step 3, select the point trace with the largest association probability for each target, and output the corresponding association probability.
4. A target data association method based on reinforcement learning according to claim 1, characterized in that The specific steps of step 5 are: Step 5.1: Select the Bayesian network and set it to five-class classification mode; Step 5.2: Input the state into the network of step 5.1, output the probabilities of different classifications, and select the classification number with the largest probability value; Step 5.3: Take the number of classifications obtained in step 5.2 as the order, use the least square method of the corresponding order to make predictions, and obtain the predicted value of the target position; Step 5.4: Based on the predicted value and point position obtained in step 5.3, the reward value is calculated using the Bayesian recursive function.
5. A target data association method based on reinforcement learning according to claim 1, characterized in that The specific steps of step 6 are: Step 6.1: If it is in the training state, go to step 6.2; otherwise, if it is in the testing state, go to step 6.3; Step 6.2: Feed the state, trace and reward value of each target together to the long short-term memory network for training and learning; Step 6.3: Save the state, traces, and reward values for each target.
6. The method for target data association based on reinforcement learning according to claim 1, characterized in that The specific steps of step 7 are: Step 7.1: For the training state, collect the measurement set obtained by sampling in the next cycle and directly enter Step 1 again; Step 7.2: For the test state, check whether there is a target association interruption. If this is the case, first perform adaptive adjustment and then enter Step 1. Otherwise, directly enter Step 1.
7. A target data association method based on reinforcement learning according to claim 6, characterized in that The specific steps of Step 7.2 include the following sub-steps: Step 7.2.1: If entering the adaptive adjustment link, input the state of the target and all the traces selected from Step 2 into the reward function to obtain a set of reward values; Step 7.2.2: Compare the reward value of the selected trace with the reward values of other traces; Step 7.2.3: If the reward value of the selected trace is the largest, it is considered that the selected trace is correct, the adaptive adjustment link ends, and enter Step 1. Otherwise, enter Step 7.2.4; Step 7.2.4: If the reward value of the selected trace is not the largest, select all the traces with a larger reward value than this reward value; Step 7.2.5: Continue to screen the traces that have not been selected by other targets from the selected traces to form a set, and arrange them in descending order of reward value; Step 7.2.6: If the trace set is empty, enter Step 1. Otherwise, select the first trace to replace the trace originally selected by the target, and then enter Step 1.