Multi-target data interconnection method based on Actor-Critic
The ‘dual window’ sliding learning mechanism is designed through the Actor-Critic method, combined with Gaussian neural network and Bayesian recursive network, and the accuracy of multi-objective data interconnection in complex environments is solved, effective learning and prediction of target motion trends is achieved, and the stability and accuracy of data interconnection is improved.
Patent Information
- Application Number
- CN202510380135.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-25
AI Technical Summary
In complex environments, traditional multi-target data interconnection algorithms are difficult to effectively deal with problems such as strong clutter interference, correlation interruption and target loss, resulting in inaccurate interconnection results.
Using the multi-objective data interconnection method based on Actor-Critic, the "dual-window" sliding learning mechanism is designed, combined with a random Gaussian neural network and a Bayesian recursive network, and the correlation probability and correlation coefficient are calculated using principal component analysis technology to achieve learning and prediction of target motion trends.
It reduces the impact of environmental clutter and sensor errors on the interconnection results, can handle strong target mobility and new targets, improves the accuracy and stability of the interconnection, and has a wide range of applications.
Smart Images

Figure CN120372347A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a multi-object data interconnection method based on Actor-Critic, which is particularly applicable to the problems of target tracking and data interconnection in the field of information fusion, and belongs to the fields of information fusion and data processing. Background Art
[0002] The data interconnection problem is a basic problem in the field of radar data processing, and its difficulty lies in establishing the relationship between radar measurement data at adjacent sampling intervals. Generally, the data interconnection process is closely related to the target tracking process. First, with the predicted value of the determined track as the center, eligible tracks are interconnected according to a specific criterion, and then filtering is performed using the tracks to achieve the purpose of tracking. There are many factors affecting the interconnection result in the data interconnection process. For example, due to external interference, a lot of clutter is mixed in the radar measurement data, the detection error of the sensor itself causes inaccurate data positions, and the change of the target motion model increases the difficulty of interconnection. In order to eliminate the influence of these factors and improve the accuracy of the interconnection result, for decades, scholars have conducted in-depth research, and many target data interconnection algorithms have emerged. The key to these traditional algorithms is the establishment of the system model, that is, using the system state equation and the measurement equation to describe the dynamic characteristics of target motion. Whether the interconnection result is effective and reliable depends on whether the established system model matches the changes of the actual system. Once the true motion of the target is inconsistent with the system model, incorrect data interconnection may occur. Therefore, in traditional algorithms, assuming that the system model is known is a prerequisite for achieving effective data interconnection. However, in the real environment, the motion modes of air targets are diverse and highly maneuverable, and it is difficult to predict its motion model in advance. It is precisely because of this prerequisite that the practicality and stability of traditional algorithms are greatly reduced.
[0003] After the wide application of artificial intelligence technology, many new multi-target data interconnection methods have emerged. The core of these methods is still to determine the system model of the target. Generally speaking, they can be divided into two categories. The first category is the supervised learning method, which sets some fixed system models and matches the optimal motion trend for the target by training a large amount of data. However, this method has two natural defects. First, the types of the set system models are limited and cannot cover all situations. Second, there may be multiple system models in a track, and it is difficult to determine the labels of the training data. The second category is the unsupervised learning method, which applies the learning experience of simple scenarios to real scenarios and finds the system model of the target during the learning process. The problem with this method is that there are significant differences between simple scenarios and real scenarios, and direct application is prone to errors, which may even mislead the learning direction and make it very difficult to re-learn, affecting the learning efficiency and the accuracy of the interconnection results. Therefore, how to reasonably apply artificial intelligence technology in a complex environment with scarce prior information to learn the system model that most conforms to the true motion trend of the target and achieve the precise interconnection of multiple targets is an urgent problem to be solved. Summary of the Invention
[0004] A multi-target data interconnection method based on Actor-Critic of the present invention aims to overcome problems such as strong clutter interference, association interruption, and target loss encountered by the target during tracking and data interconnection in a complex environment.
[0005] The multi-target data interconnection method based on Actor-Critic of the present invention is characterized in that it includes the following steps:
[0006] Step 1: Model the target environment and design a "double-window" sliding learning mechanism;
[0007] Step 2: Detect the new target state from the "state data";
[0008] Step 3: Input the states of all targets and the measurement data obtained by sampling at the next moment into the Actor network to generate an association probability matrix;
[0009] Step 4: Under the condition of adhering to the uniqueness principle of the dot plot and the target, select the dot plot with the largest possible association probability value for each target and output the corresponding association probability;
[0010] Step 5: Make a judgment according to the current actual situation;
[0011] Step 6: Output the reward value through the Critic network and feedback it to the Actor network to guide the motion trend of the learning target;
[0012] Step 7: Based on the learning results of the "learning window", associate the traces at the next moment, determine the next "learning window" and "status window", return to Step 2, and enter the next learning process;
[0013] Step 8: Repeat the process until there is no measurement data in the next sampling interval;
[0014] Preferably, the specific steps of Step 1 are as follows:
[0015] Step 1.1: Set the measurement data obtained by the sensor in 10 consecutive sampling intervals as "learning data". 10 is the "learning window", representing the length of the learning process;
[0016] Step 1.2: Select the measurement data obtained by the sensor in the first 5 sampling intervals from the "learning data" as "status data". 5 is the "status window", representing the length of the target track whose motion trend needs to be learned;
[0017] Step 1.3: Through continuous trial and error, learn the motion trend of the target within the "status window", and assist in associating the traces originating from the target from the data obtained by the sensor in the 11th sampling interval.
[0018] Preferably, the specific steps of Step 2 are as follows:
[0019] Step 2.1: If there are targets at the current moment, the trace data of these targets need to be deleted from the "status data";
[0020] Step 2.2: Set thresholds according to the maximum and minimum speeds of the target motion, process the "status data", and screen out all eligible trace combinations;
[0021] Step 2.3: If there are trace combinations, enter Step 2.4; otherwise, enter Step 3;
[0022] Step 2.4: Use the cosine theorem formula to quantify the motion trends of all trace combinations;
[0023] Step 2.5: Classify all trace combinations according to the uniqueness principle of the trace and the target, and determine the number of new targets;
[0024] Step 2.6: Calculate the variance of the motion trends of each type of trace combination, and find the trace combination with the smallest variance value as the state of the target;
[0025] Preferably, the specific steps of Step 3 are as follows:
[0026] Step 3.1: According to the detection range of the sensor, preprocess the states of all targets, that is, divide the state data of all targets by the detection range to limit the state data value within [-1, 1];
[0027] Step 3.2: Input the status data processed in the previous step into a random Gaussian neural network to obtain a status prediction value;
[0028] Step 3.3: Input the status prediction value and the track together into a Bayesian recursive network, and output the association probability between the target and the track;
[0029] Step 3.4: The association probabilities between all targets and all tracks form an association probability matrix;
[0030] Preferably, the specific steps of step 5 are as follows:
[0031] Step 5.1: If there is a target association interruption, directly return to step 2 and re-learn;
[0032] Step 5.2: If not all the measurement data within the "learning window" have been tested, directly return to step 3 and continue to test the measurement data at the next sampling moment;
[0033] Step 5.3: If all the measurement data within the "learning window" have been tested and there is no phenomenon of association interruption, it indicates that for the measurement data within the "status window", the track information of the target may originate from the real track of the target at this time, and step 6 can be entered for further operation;
[0034] Preferably, the specific steps of step 6 are as follows:
[0035] Step 6.1: According to the track associated with the target, the new status of the target at the next moment can be obtained;
[0036] Step 6.2: Use the principal component analysis technique to reduce both the new status data and the previous status data from two dimensions to one dimension;
[0037] Step 6.3: Process the two status data obtained in step 6.2 and calculate their Pearson correlation coefficient;
[0038] Step 6.4: Multiply the Pearson correlation coefficient by the association probability to obtain a reward value.
[0039] A multi-target data interconnection method based on Actor-Critic of the present invention specifically includes the following technical measures: First, model the target environment and design a "double-window" sliding learning mechanism. Then, preprocess the "state data" to ensure real-time detection of possible new targets. Next, add a Bayesian recursive network to the end of the random Gaussian neural network, enabling the Actor to predict the interconnection probability between the measurement and its possible various source targets, and the target selects the traces based on the interconnection probability. Finally, after reducing the dimension of the adjacent state matrices using the principal component analysis technique in the Critic, calculate the Pearson correlation coefficient between the two to evaluate the rewards and punishments selected by the Actor. The multi-target data interconnection method based on Actor-Critic proposed by the present invention can not only reduce the influence of environmental clutter and sensor detection errors on the interconnection results, but also effectively handle the strong maneuverability and sudden appearance of new targets that may occur during the interconnection process. At the same time, it can also get rid of the dependence of traditional target tracking and data interconnection algorithms on the system model. Through testing with a large number of sample data, the test results are good, and it has the advantages of a wide application range and high practical application value. The proposed target data interconnection method can be directly applied to corresponding practical problems. Description of the Drawings
[0040] Figure 1 is the overall block diagram of a multi-target data interconnection method based on Actor-Critic of the present invention;
[0041] Figure 2 is the single-action selection process of a multi-target data interconnection method based on Actor-Critic of the present invention. Detailed Embodiment
[0042] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.
[0043] Embodiment 1
[0044] A target tracking and data interconnection method based on reinforcement learning in this embodiment refers to the attached Figure 1-2 and includes the following steps:
[0045] Step 1: Model the target environment and design a "double-window" sliding learning mechanism;
[0046] Step 1.1: Set the measurement data obtained by the sensor at 10 consecutive sampling intervals as the "learning data", and 10 is the "learning window", indicating the length of the learning process;
[0047] Step 1.2: Select the measurement data obtained by the first 5 sampling interval sensors from the "learning data" as the "state data". Here, 5 is the "state window", representing the target track length for which the motion trend needs to be learned.
[0048] Step 1.3: Through continuous trial and error, learn the motion trend of the target within the "state window", and assist in associating the tracks originating from the target from the data obtained by the 11th sampling interval sensor.
[0049] Step 2: Detect new target states from the "state data".
[0050] Step 2.1: If there are existing targets at the current moment, the track data of these targets need to be deleted from the "state data".
[0051] Step 2.2: Set thresholds based on the maximum and minimum speeds of target motion, process the "state data", and screen out all eligible track combinations.
[0052] Step 2.3: If there are track combinations, go to Step 2.4; otherwise, go to Step 3.
[0053] Step 2.4: Use the cosine theorem formula to quantify the motion trends of all track combinations.
[0054] Step 2.5: Classify all track combinations according to the uniqueness principle of tracks and targets, and determine the number of new targets.
[0055] Step 2.6: Calculate the variance of the motion trends of each class of track combinations, and find the track combination with the smallest variance value as the state of the target.
[0056] Step 3: Input the states of all targets and the measurement data obtained by sampling at the next moment into the Actor network to generate an association probability matrix.
[0057] Step 3.1: According to the detection range of the sensor, preprocess the states of all targets, that is, divide the state data of all targets by the detection range to limit the state data values within [-1, 1].
[0058] Step 3.2: Input the state data processed in the previous step into a random Gaussian neural network to obtain state prediction values.
[0059] Step 3.3: Input the state prediction values and tracks into a Bayesian recursive network to output the association probability between the target and the track.
[0060] Step 3.4: The association probabilities between all targets and all tracks form an association probability matrix.
[0061] Step 4: Under the condition of adhering to the principle of uniqueness of the track and the target, select the track with the largest possible association probability value for each target, and output the corresponding association probability;
[0062] Step 5: Make a judgment according to the current actual situation;
[0063] Step 5.1: If there is a target association interruption, directly return to Step 2 and re-learn;
[0064] Step 5.2: If not all the measurement data within the "learning window" have been tested, directly return to Step 3 and continue to test the measurement data at the next sampling moment;
[0065] Step 5.3: If all the measurement data within the "learning window" have been tested and there is no phenomenon of association interruption, it indicates that for the measurement data within the "state window", the track information of the target may originate from the true track of the target at this time, and Step 6 can be entered for continued operation;
[0066] Step 6: Output the reward value through the Critic network, feedback to the Actor network, and guide the movement trend of the learning target;
[0067] Step 6.1: According to the track associated with the target, the new state of the target at the next moment can be obtained;
[0068] Step 6.2: Use the principal component analysis technique to reduce both the new state data and the previous state data from two dimensions to one dimension;
[0069] Step 6.3: Process the two state data obtained in Step 6.2 and calculate their Pearson correlation coefficient;
[0070] Step 6.4: Multiply the Pearson correlation coefficient by the association probability to obtain the reward value;
[0071] Step 7: Based on the learning results of the "learning window", associate the tracks at the next moment, determine the next "learning window" and "state window", return to Step 2, and enter the next learning process;
[0072] Step 8: Repeat in a loop until there is no measurement data in the next sampling interval.
[0073] Embodiment 2
[0074] To better illustrate the present invention, the following takes the measurement data of a certain type of radar as a specific embodiment to elaborate on the steps of the present invention in detail:
[0075] Step 11: Model the target environment and design a "double-window" sliding learning mechanism;
[0076] Step 11.1: Set the measurement data {Z t-9 , Z t-8 ,..., Z t-1 , Z t} obtained by the sensor at 10 consecutive sampling intervals as "learning data", where 10 is the "learning window", representing the length of the learning process;
[0077] Step 11.2: Select the measurement data obtained by the sensor at the first 5 sampling intervals from the "learning data" as "state data", where 5 is the "state window", representing the length of the target track for which the motion trend needs to be learned;
[0078] Step 11.3: Through continuous trial and error, learn the motion trend of the target within the "state window", and assist in associating the tracks originating from the target in the data obtained by the 11th sampling interval sensor;
[0079] Step 12: Detect new target states from the "state data"
[0080] Step 12.1: If there are targets at the current moment, the track data of these targets need to be deleted from the "state data";
[0081] Step 12.2: Set thresholds according to the maximum and minimum speeds of the target motion, and process the "state data" to screen out all eligible track combinations. Given the sampling interval T_sample of the sensor, the measurement data z, the maximum speed v_max and the minimum speed v_min of the target motion, screen out the track combinations that meet the speed threshold, that is
[0082]
[0083] Step 12.3: If there are track combinations, go to Step 12.4; otherwise, go to Step 13;
[0084] Step 12.4: Use the cosine theorem formula to quantify the motion trend of all track combinations, that is
[0085]
[0086] Step 12.5: Classify all track combinations according to the uniqueness principle of the track and the target, and determine the number of new targets;
[0087] Step 12.6: Calculate the variance Variance = var(f) of the motion trend of each class of track combinations, and find the track combination with the smallest variance value as the state of the target;
[0088] Step 13: Input the states of all targets and the measurement data obtained by sampling at the next moment into the Actor network to generate an association probability matrix
[0089] Step 13.1: According to the detection range of the sensor, preprocess the states of all targets, that is, divide the state data of all targets by the detection range to limit the state data value within [-1, 1].
[0090] Step 13.2: Input the state data processed in the previous step into the random Gaussian neural network to obtain the state prediction value.
[0091] Step 13.3: Input the state prediction value and the tracklet together into the Bayesian recursive network to output the association probability between the target and the tracklet.
[0092] Step 13.4: The association probabilities between all targets and all tracklets form the association probability matrix.
[0093] Step 14: Under the condition of adhering to the uniqueness principle of tracklets and targets, select the tracklet with the largest possible association probability value for each target and output the corresponding association probability.
[0094] Step 15: Make a judgment according to the current actual situation.
[0095] Step 15.1: If there is a target association interruption, directly return to Step 12 to re-learn.
[0096] Step 15.2: If not all the measurement data within the "learning window" have been tested, directly return to Step 13 to continue testing the measurement data at the next sampling moment.
[0097] Step 15.3: If all the measurement data within the "learning window" have been tested and there is no phenomenon of association interruption, it indicates that for the measurement data within the "state window", the tracklet information of the target may originate from the true tracklet of the target at this time, and Step 16 can be entered for further operation.
[0098] Step 16: Output the reward value through the Critic network, feedback to the Actor network, and guide the movement trend of the learning target.
[0099] Step 16.1: According to the tracklet associated with the target, the new state of the target at the next moment can be obtained
[0100] Step 16.2: Use the principal component analysis technique to reduce both the new state data and the previous state data from two dimensions to one dimension.
[0101] Step 16.3: Process the two state data obtained in Step 16.2 and calculate their Pearson correlation coefficient, that is
[0102]
[0103] Because the value range of is [-1, 1], so by setting the value range of
[0104] Step 16.4: Multiply the Pearson correlation coefficient by the association probability to obtain the reward value
[0105] Step 17: Based on the learning results of the "learning window", associate the traces at the next moment, and determine the next "learning window" and "state window", then return to Step 12 to enter the next learning process;
[0106] Step 18: Repeat the process until there is no measurement data in the next sampling interval.
[0107] As described above, only the preferred specific embodiments of the present invention are provided, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes, shall be covered by the protection scope of the present invention.
Claims
1. A multi-objective data interconnection method based on Actor-Critic, characterized in that It includes the following steps: Step 1: Model the target environment and design a "double-window" sliding learning mechanism; Step 2: Detect new target states from the "state data"; Step 3: Input the states of all targets and the measurement data obtained by sampling at the next moment into the Actor network to generate an association probability matrix; Step 4: Under the condition of adhering to the principle of uniqueness of tracks and targets, select the track with the largest possible association probability value for each target and output the corresponding association probability; Step 5: Make a judgment according to the current actual situation; Step 6: Output the reward value through the Critic network, feedback to the Actor network, and guide the movement trend of the learning target; Step 7: Based on the learning results of the "learning window", associate the tracks at the next moment, determine the next "learning window" and "state window", return to Step 2, and enter the next learning process; Step 8: Repeat until there is no measurement data in the next sampling interval.
2. The multi-objective data interconnection method based on Actor-Critic according to claim 1, wherein The specific steps of Step 1 are as follows: Step 1.1: Set the measurement data obtained by the sensor in 10 consecutive sampling intervals as the "learning data". 10 is the "learning window", indicating the length of the learning process; Step 1.2: Select the measurement data obtained by the sensor in the first 5 sampling intervals from the "learning data" as the "state data". 5 is the "state window", indicating the length of the target track whose movement trend needs to be learned; Step 1.3: Through continuous trial and error, learn the movement trend of the target within the "state window" to assist in associating the tracks originating from the target from the data obtained by the sensor at the 11th sampling interval.
3. A multi-objective data interconnection method based on Actor-Critic according to claim 1, characterized in that The specific steps of Step 2 are as follows: Step 2.1: If there are targets at the current moment, the track data of these targets need to be deleted from the "state data"; Step 2.2: Set thresholds according to the maximum and minimum speeds of target movement, process the "state data", and screen out all eligible track combinations; Step 2.3: If there are track combinations, enter Step 2.4, otherwise, enter Step 3; Step 2.4: Use the cosine theorem formula to quantify the movement trends of all track combinations; Step 2.5: Classify all track combinations according to the principle of uniqueness of tracks and targets to determine the number of new targets; Step 2.6: Calculate the variance of the movement trends of each type of track combination, and find the track combination with the smallest variance value as the state of the target.
4. A multi-objective data interconnection method based on Actor-Critic according to claim 1, characterized in that The specific steps of Step 3 are as follows: Step 3.1: According to the detection range of the sensor, preprocess the states of all targets, that is, divide the state data of all targets by the detection range to limit the state data value within [-1, 1]; Step 3.2: Input the state data processed in the previous step into a random Gaussian neural network to obtain the state prediction value; Step 3.3: Input the state prediction value and the track into the Bayesian recursive network to output the association probability between the target and the track; Step 3.4: The association probabilities between all targets and all tracks form an association probability matrix.
5. A multi-objective data interconnection method based on Actor-Critic according to claim 1, characterized in that The specific steps of Step 5 are as follows: Step 5.1: If there is a target association interruption, directly return to Step 2 and re-learn; Step 5.2: If not all the measurement data within the "learning window" have been tested, directly return to Step 3 to continue testing the measurement data at the next sampling moment; Step 5.3: If all the measurement data within the "learning window" have been tested and there is no phenomenon of association interruption, it indicates that for the measurement data within the "status window", the target track information at this time may originate from the true track of the target, and Step 6 can be entered for further operation.
6. A multi-objective data interconnection method based on Actor-Critic according to claim 1, characterized in that The specific steps of Step 6 are as follows: Step 6.1: Based on the tracks associated with the target, the new state of the target at the next moment can be obtained; Step 6.2: Using the principal component analysis technique, both the new state data and the previous state data are reduced from two dimensions to one dimension; Step 6.3: Process the two state data obtained in Step 6.2 and calculate their Pearson correlation coefficient; Step 6.4: Multiply the Pearson correlation coefficient by the association probability to obtain the reward value.