VTS radar station multi-station cooperation method, device and equipment and storage medium
By constructing a collaborative decision-making matrix and using a reinforcement learning model to optimize the real-time coverage capability weights, the problem of weight calculation lag in traditional VTS radar site collaboration methods is solved, reasonable resource allocation and improved target tracking accuracy are achieved, and the robustness of the system in complex environments is enhanced.
Patent Information
- Application Number
- CN202511302805.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-12
- Publication Date
- 2025-10-14
- Estimated Expiration
- 2045-09-12
AI Technical Summary
In the traditional VTS radar station multi-site coordination method, weight calculation relies on static or semi-static data such as the number of historical failures and the target loss rate in the overlapping area. This makes it impossible to accurately reflect the real-time operating status and capabilities of the radar site, resulting in resource waste and reduced tracking accuracy, especially in dynamically changing water traffic environments. The system robustness is reduced.
By building a collaborative decision-making matrix, obtaining historical performance data and real-time observation data of each radar site, using the reinforcement learning model to optimize the real-time coverage capability weight, combined with the reward function to guide learning, dynamically adjust the task allocation strategy, and optimize the coverage capability weight.
It achieves rational allocation and efficient utilization of resources, improves target tracking accuracy and system robustness, and effectively copes with complex water traffic environments.
Smart Images

Figure CN120779347A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of radar collaboration, and in particular to a VTS radar station multi-site collaboration method, device, equipment and storage medium. Background Art
[0002] A key step in traditional VTS radar multi-site coordination methods is the construction of a collaborative decision-making matrix. This matrix guides task allocation among radar sites and ensures efficient and accurate operation of the entire system. However, existing collaborative decision-making matrix construction methods have significant flaws, particularly in the calculation of real-time coverage capability weights.
[0003] Traditional weight calculations often rely on static or semi-static data such as historical failure counts and target loss rates in overlapping areas. This results in weight assignments that fail to accurately reflect the real-time operational status and capabilities of radar sites. For example, a radar site that has recently undergone maintenance and is currently performing well may have its actual coverage capability underestimated due to its historical failure history, leading to marginalization in task allocation, resulting in wasted resources and a mismatch between demand and performance. Furthermore, traditional methods struggle to cope with the dynamic and changing traffic environment of waterways.
[0004] When radar sites malfunction or performance degrades, the system cannot adjust its task allocation strategy in a timely manner due to lags in weight calculation, resulting in reduced tracking accuracy and system robustness. In complex and changing traffic environments, this lag can lead to global coordination failures, seriously impacting the operational effectiveness of the VTS system. Summary of the Invention
[0005] The purpose of this application is to overcome the defects in the above-mentioned prior art and provide a VTS radar station multi-site collaboration method, device, equipment and storage medium.
[0006] This application provides a VTS radar station multi-site collaboration method, including:
[0007] Constructing a collaborative decision-making matrix, wherein the collaborative decision-making matrix includes the real-time coverage capability weights of each radar site;
[0008] generating a site task allocation instruction set based on the collaborative decision matrix;
[0009] The collaborative decision-making matrix construction includes: obtaining historical performance data and real-time observation data of each radar site; optimizing the real-time coverage capability weight through a reinforcement learning model based on the historical performance data and the real-time observation data; wherein the reinforcement learning model guides learning through a reward function, and the reward function is: , 、 、 is the control coefficient, G is the tracking continuity, Z is the resource consumption, The actual coverage capability of the site is evaluated based on real observation data. The current estimated weight output by the reinforcement learning model; the optimized real-time coverage capability weight is output to the collaborative decision-making matrix.
[0010] Optionally, the reinforcement learning model guides learning through a reward function, including:
[0011] Get real-time tracking continuity indicators;
[0012] Get current resource consumption indicators;
[0013] Calculate the absolute error between the actual coverage capacity weight and the estimated coverage capacity weight;
[0014] A reinforcement learning reward value is generated by the reward function based on the real-time tracking continuity indicator, the current resource consumption indicator and the absolute error.
[0015] Optionally, obtaining real-time observation data of each radar site includes:
[0016] Collect real-time signal-to-noise ratio of radar sites;
[0017] Collect the real-time effective detection area of the radar site;
[0018] Collect real-time communication bandwidth utilization of radar sites;
[0019] A real-time observation data set is constructed based on the real-time signal-to-noise ratio, the real-time effective detection area, and the real-time communication bandwidth utilization.
[0020] Optionally, obtaining historical performance data of each radar site includes:
[0021] Count the number of site failures within the sliding time window;
[0022] Calculate the target loss rate of the overlapping area within the sliding time window;
[0023] A historical performance data set is constructed based on the number of site failures and the target loss rate of the overlapping area.
[0024] Optionally, optimizing the real-time coverage capability weight by using a reinforcement learning model includes:
[0025] Feed historical performance data and real-time observation data into the reinforcement learning model;
[0026] Evaluate the action value through the reward function;
[0027] Update the action policy to minimize the weight prediction error;
[0028] Output the optimized real-time coverage capability weight.
[0029] Optionally, generating a site task allocation instruction set based on the collaborative decision matrix includes:
[0030] Sort the site priorities based on real-time coverage capability weights;
[0031] Calculate the optimal takeover site set based on the location of the failed site;
[0032] Generate a task allocation instruction set including a takeover strategy.
[0033] Optionally, obtaining the real-time tracking continuity indicator includes:
[0034] Count the number of target tracks successfully spliced within the preset time window;
[0035] Calculate the total number of targets that should be tracked;
[0036] Track continuity values via ratio calculations:
[0037] Tracking continuity = number of successfully spliced tracks / total number of targets to be tracked.
[0038] The present application also provides a VTS radar station multi-site collaboration device, comprising:
[0039] A construction module is used to construct a collaborative decision matrix, wherein the collaborative decision matrix includes the real-time coverage capability weights of each radar site;
[0040] an allocation module, generating a site task allocation instruction set based on the collaborative decision matrix;
[0041] The collaborative decision-making matrix construction includes: obtaining historical performance data and real-time observation data of each radar site; optimizing the real-time coverage capability weight through a reinforcement learning model based on the historical performance data and the real-time observation data; wherein the reinforcement learning model guides learning through a reward function, and the reward function is: , 、 、 is the control coefficient, G is the tracking continuity, Z is the resource consumption, The actual coverage capability of the site is evaluated based on real observation data. The current estimated weight output by the reinforcement learning model; the optimized real-time coverage capability weight is output to the collaborative decision-making matrix.
[0042] The present application also provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the above method is implemented.
[0043] The present application also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed in a computer, the computer is caused to execute the above method.
[0044] The beneficial effects of this application are:
[0045] The present application provides a multi-site collaboration method for VTS radar stations, comprising: constructing a collaborative decision matrix, the collaborative decision matrix including the real-time coverage capability weight of each radar site; generating a site task allocation instruction set based on the collaborative decision matrix; wherein constructing the collaborative decision matrix includes: obtaining historical performance data and real-time observation data of each radar site; and optimizing the real-time coverage capability weight through a reinforcement learning model based on the historical performance data and the real-time observation data; wherein the reinforcement learning model guides learning through a reward function, and the reward function is: , 、 、 is the control coefficient, G is the tracking continuity, Z is the resource consumption, The actual coverage capability of the site is evaluated based on real observation data. The current estimated weight output by the reinforcement learning model is output; the optimized real-time coverage capability weight is output to the collaborative decision-making matrix. This application uses a reinforcement learning model to optimize the real-time coverage capability weight of the VTS radar station collaborative decision-making matrix, achieving reasonable resource allocation and efficient utilization, improving target tracking accuracy and system robustness, and effectively coping with complex water traffic environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 This is a schematic diagram of the multi-site collaboration process of the VTS radar station in this application;
[0047] Figure 2 This is a schematic diagram of the multi-source data preprocessing process in this application;
[0048] Figure 3 This is a schematic diagram of the timing of dynamic weight adjustment of the collaborative decision-making layer in this application;
[0049] Figure 4 This is a schematic diagram of the dual-mode weight switching state machine in this application. DETAILED DESCRIPTION
[0050] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the accompanying drawings, it is understood that various forms of implementing the present disclosure are not limited by the embodiments set forth herein. Rather, the embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.
[0051] Please refer to Figure 1 As shown, the present application provides a VTS radar station multi-site collaboration method, including:
[0052] S101: Construct a collaborative decision matrix, where the collaborative decision matrix includes the real-time coverage capability weight of each radar site.
[0053] Constructing a collaborative decision matrix is a core component of the multi-site collaboration approach for VTS radar stations. It aims to dynamically evaluate each site's coverage capability to optimize task allocation. The collaborative decision matrix is a data structure used to store and update key radar site status parameters, including real-time coverage capability weights, quantified values for detection overlap between adjacent sites, and site fault status indicators.
[0054] The real-time coverage capability weight indicates the detection efficiency of the station for water targets at the current moment. The higher the value, the stronger the coverage capability.
[0055] Constructing this matrix requires integrating historical performance data and real-time observation data, and optimizing weight values through a reinforcement learning model.
[0056] During the construction of the collaborative decision-making layer, historical performance data for each radar site is first obtained. This historical performance data includes the site's operational records within a specific time window. Specifically, the system uses a sliding window strategy to truncate the historical data, focusing on the site's performance within the last hour. This processing mechanism first establishes a time-indexed cache structure, registering the time of each site's failure and the corresponding target loss records within the overlapping area in real time. The system periodically slides the window boundaries, retaining only data with timestamps falling within the current window range and automatically discarding any data outside the time window.
[0057] For site failures , count the timestamps of each failure occurring in the current time window, and recalculate the number in each round of weight update to ensure that it reflects the most recent stability status.
[0058] For the target loss rate in the overlapping area ( ), the system compares the theoretical number of targets that should be detected in the overlapping area with the actual number of successfully tracked targets, and calculates the loss rate through the difference ratio, which is calculated as follows:
[0059]
[0060] in, is the number of targets that are actually continuously detected by multiple stations and whose trajectories are spliced within the time window. is the theoretical number of targets that should be successfully tracked.
[0061] Both are verified by the fusion engine for consistency of the target trajectory. This approach ensures that historical data only reflects the current system state, avoiding outdated information interference.
[0062] At the same time, the system obtains real-time observation data of each radar site.
[0063] Real-time observation data includes the instantaneous performance indicators of the site. The specific operation is: the system collects the real-time signal-to-noise ratio of the radar site , which is obtained by measuring the target reflection intensity of the radar system, reflecting the signal quality; collects the real-time effective detection area , indicating the current water area covered by the site; collects the real-time communication bandwidth utilization , reflecting the data transmission capacity of the site. These data are collected in real time through sensors and network monitoring modules, and are constructed into a real-time observation data set. Real-time observation data is used to dynamically update the site state to avoid historical bias.
[0064] Please refer to Figure 2 , further, each radar site first collects the target original point cloud data in the water area covered by the radar system in real time through a high-resolution radar system. These point cloud data reflect the reflection characteristics of the target in the radar beam, including spatial coordinates, reflection intensity, and timestamp information. Due to the influence of environmental noise, electromagnetic interference and the performance of the device itself, the original point cloud data usually contains a large number of non-target points, scattered points and artifacts, so noise filtering algorithm needs to be used for purification.
[0065] The filtering process uses a spatial anomaly rejection method based on density clustering to remove isolated points and low-energy echo points by evaluating the point density and intensity threshold of the local area in the point cloud, thereby improving the clarity of the target outline. After denoising, the next step is to perform coordinate unification processing. The point cloud data obtained by different sites is usually in their own local coordinate system, in order to realize subsequent cross-site trajectory matching, all point cloud data must be converted to a unified georeferenced coordinate system. The coordinate conversion process is based on the pre-calibrated geographic position parameters and attitude angle information of the site, and is converted through a rigid body transformation model, the formula is:
[0066]
[0067] wherein, represents the position vector of the target in the unified coordinate system, is the original local coordinate position, is the rotation matrix, is the translation vector, both of which are obtained from the site's geographic attitude calibration.
[0068] After coordinate unification is completed, the motion trajectory of the point cloud data in the time series is extracted, and the continuous observation segments of the target are converted into standardized target trajectory segments through a trajectory fitting method based on a sliding time window. Each segment contains a unified format of spatiotemporal position vector, velocity vector and confidence score, providing a highly consistent data foundation for subsequent distributed data fusion and trajectory matching.
[0069] Please refer to Figure 3 As shown, further, in the distributed data fusion stage, the standardized target trajectory segments that have been pre-processed and formed by each radar site are input into the regional data fusion engine in real time. The engine has an efficient trajectory matching and fusion algorithm core. First, based on the continuity of the target in spatial position and timestamp, the system performs spatiotemporal correlation analysis on the target trajectory across sites, and identifies the trajectory segments belonging to the same target but observed by different sites by comparing the target's velocity vector, trajectory trend and spatial distance between adjacent frames. For multiple sets of matched observation values, the fusion engine uses an adaptive weighted fusion algorithm to generate a unified target trajectory point. This fusion process ensures that the fusion result maintains accuracy while taking into account the communication performance status of each site by assigning observation weights to different sites. Specifically, the fusion value ( ) is calculated as:
[0070]
[0071] in, is the measurement value of the target at the i-th station, is the fusion weight of the observation. Weight ( The determination of takes into account two factors: measurement accuracy and communication bandwidth conditions, and its expression is:
[0072]
[0073] in, It represents the measurement variance of the i-th station, reflecting the accuracy of its observation data. The smaller the value, the smaller the error and the higher the credibility. Indicates the current communication bandwidth utilization of the site. It is the maximum bandwidth utilization in the system and is used for normalization.
[0074] This weighting design prioritizes measurement reliability while dynamically adjusting the influence of different sites in the fusion process, preventing sites with poor communication capabilities from excessively impacting fusion accuracy. In this way, the fusion engine constructs a unified set of target trajectories across the entire region, encompassing observation information from all radar sites. This ensures robust data consistency in complex network conditions, laying a solid foundation for subsequent collaborative decision-making and task allocation.
[0075] After obtaining historical performance data and real-time observation data, the system optimizes the real-time coverage capability weight through a reinforcement learning model.
[0076] Please refer to Figure 4 As shown in the figure, in collaborative task scheduling, to cope with the dynamic changes in radar site status, the system has designed a dual-mode weight calculation switching mechanism to ensure a balance between system stability and responsiveness. In normal mode, the system uses a combination of historical statistical data and real-time observations, using Kalman filtering and sliding window processing to calculate the signal-to-noise ratio and effective detection area estimate for each site, thereby producing a more stable weighting result suitable for most stable operation scenarios.
[0077] However, when the system detects a sudden change in the status of a site, such as a sharp drop in the continuous signal-to-noise ratio, a sudden decrease in the detection area, or frequent failures in a short period of time, the system will immediately trigger the switching mechanism and enter a temporary real-time emergency mode. In this mode, the weight is no longer based on historical information or filtered prediction values, but is directly calculated based on the current real-time observation data. ), the formula is:
[0078]
[0079] in, is the instantaneous signal-to-noise ratio measured at the current site, is the current actual detection area, Indicates the maximum value of the real-time signal-to-noise ratio among all sites, is the sum of the current detection areas of all stations, and It is the basic weight coefficient used in emergency mode, emphasizing real-time response priority.
[0080] When the mutation indicator continues to decline or remains fluctuating, the system maintains real-time mode operation to quickly adjust the task allocation structure. Once the indicator returns to stability and meets the stability threshold, the system will automatically switch back to normal mode and resume the integrated use of historical trends and status prediction results, thereby taking into account both stability and sensitivity, and improving the accuracy and resilience of the system's task allocation in complex dynamic environments.
[0081] The reinforcement learning model is an algorithmic framework that uses historical station performance and real-time observations as state inputs. It integrates trajectory tracking continuity, resource efficiency, and prediction accuracy to continuously adjust weight estimation strategies through an interactive process. The core mechanism is the design of a reward function to guide model learning.
[0082] The reward function is defined as ( ):
[0083]
[0084] in, 、 、 is the control coefficient, which controls the impact of the three indicators on the reward respectively; G is tracking continuity, which measures the system's ability to maintain the integrity of the target trajectory; Z is resource consumption, which represents the comprehensive indicator of communication and computing resources consumed by the station per unit time; The actual coverage capability of the site is evaluated based on real observation data; The current estimated weights output by the reinforcement learning model.
[0085] The temporal difference method is used to update the action value function. By positively reinforcing excellent strategies and negatively punishing deviant behaviors, the model is forced to learn the high-dimensional mapping relationship between variables such as signal-to-noise ratio, detection area, fault records and coverage capability.
[0086] The optimization process involves inputting historical performance data and real-time observation data into a reinforcement learning model; evaluating the value of actions using a reward function to generate reward values; updating action strategies to minimize weight prediction errors; and finally outputting optimized real-time coverage capability weights to the collaborative decision-making matrix. These weights more closely reflect the actual performance of the site, improving environmental adaptability.
[0087] The specific implementation of the reinforcement learning model guided by the reward function includes obtaining the real-time tracking continuity indicator, obtaining the current resource consumption indicator, and calculating the absolute error between the actual coverage capacity weight and the estimated coverage capacity weight. The real-time tracking continuity indicator is obtained by counting the number of target tracks successfully spliced within the preset time window, calculating the total number of targets to be tracked, and calculating the tracking continuity value through the ratio:
[0088]
[0089] Where D is the tracking continuity, is the number of successfully spliced trajectories, The total number of targets that should be tracked. This metric reflects the system's ability to maintain target continuity.
[0090] The current resource consumption index is obtained by measuring the resource usage of the site in scanning and data transmission through the monitoring module. The absolute error is calculated as , directly based on the observed values. Based on these indicators, the reward function generates reinforcement learning reward values to drive model optimization.
[0091] S102: Generate a site task allocation instruction set based on the collaborative decision matrix.
[0092] The execution phase of the method involves generating a set of station assignment instructions, which are used to dynamically schedule radar station assignments. This set includes instructions such as a dynamic takeover strategy for the coverage area of a failed station, ensuring continuous water monitoring. Specifically, the system prioritizes stations based on their real-time coverage capability weights in the collaborative decision-making matrix, assigning tasks to stations with higher weights. The system also calculates the optimal set of takeover stations based on the location of the failed station. Finally, a set of assignment instructions is generated, including the takeover strategy.
[0093] After the system completes the collaborative decision-making matrix construction and generates the task allocation instruction set, it enters the adaptive target tracking and output phase. This phase, based on the instruction set, dynamically schedules healthy radar stations to enhance their monitoring capabilities over the areas previously covered by adjacent faulty stations. The system first spatially reallocates scanning tasks based on trajectory continuity requirements, coverage redundancy, and current station load capacity. By increasing scanning frequency and adjusting beam pointing strategies, it fills in the gaps in perception left by the faulty station. For example, the default scanning frame rate for healthy stations is increased from 4 revolutions per minute in normal mode to 8 revolutions per minute in emergency mode. This increases the refresh rate by increasing the scanning cycle, shortening the target update interval, reducing track gaps caused by takeover, and enhancing target trajectory continuity. During this process, the fusion module continuously receives standardized target trajectory segments uploaded by each station and updates the target status in real time based on the previously constructed global unified trajectory structure. To avoid information delays or data conflicts, the system employs a timestamp synchronization mechanism and a data weighting strategy to ensure consistency in timing and accuracy between target information reported by multiple stations. The final fusion output includes the target's globally unique identifier, historical trajectory sequence, current motion status, and predicted trend. This unified package is then pushed to the VTS central display and control terminal, enabling global perception and visualization of dynamic targets within the waters, supporting higher-level functions such as command and dispatch, early warning linkage, and situation analysis. This fault-tolerant and self-healing mechanism not only ensures the system maintains trajectory continuity in the event of node failure, but also enhances its ability to respond to emergencies in highly dynamic environments, improving the robustness and operational effectiveness of the overall monitoring system.
[0094] The site prioritization operation is: read the real-time coverage capability weight in the matrix , sites are arranged in descending order, and sites with high weights are considered healthy and efficient and are prioritized for key area monitoring.
[0095] The calculation operation of the optimal takeover site set is: when the system detects a faulty site, a weighted minimization strategy is introduced to select the optimal subset from the candidate set of takeover sites.
[0096] This strategy uses the inverse ratio of distance cost and weight as a comprehensive evaluation indicator, and the objective function is:
[0097]
[0098] in, represents the spatial distance from site j to fault site k, It is the maximum effective takeover distance set by the system. is the weight of site j. This function reflects that sites with close distance and high weight are selected as takeover nodes first. The system traverses all sites that meet ≤ The site, screening minimizes the set of objective functions { } and generates specific takeover instructions. After the instruction set is updated, schedule an enhanced scan of the health site.
[0099] During the adaptive target tracking and output phase, the system dispatches healthy radar sites based on the instruction set to enhance the monitoring capability of the original coverage area of the faulty site. The system redistributes scanning tasks within the spatial range according to the trajectory continuity requirements, coverage redundancy, and current site load capacity, and fills the perception gaps by increasing the scanning frequency and adjusting the beam pointing strategy. For example, the scanning frame rate is increased from 4 revolutions per minute to 8 revolutions per minute, and the scanning cycle is encrypted. The fusion module receives standardized target trajectory segments, updates the target status based on the global unified trajectory structure, and adopts a timestamp synchronization mechanism and a data weight control strategy to ensure consistency. The fusion results are finally output to the VTS central display and control terminal to support ship traffic management.
[0100] During the fault takeover process, in order to alleviate the problem of velocity vector mutation during the cross-site inter-takeover process, the system uses the Doppler frequency shift compensation algorithm. This algorithm calculates the target's expected velocity value within the cross-site time window based on the target's radial velocity change trend in the historical continuous frames, and corrects the velocity discontinuity points in the fused trajectory caused by observation errors. The correction method is based on weighted interpolation. Assume that the expected velocity is The actual measured value is , by generating a corrected velocity ( ):
[0101]
[0102] Among them, λ is the correction coefficient, which is dynamically adjusted between 0 and 1 to balance real-time performance and continuity.
[0103] This mechanism effectively mitigates tracking jumps caused by heterogeneous radar distribution, improving track fusion smoothness and overall system stability during the failover period. The entire process is completed automatically within the fusion engine, requiring no human intervention, enabling rapid response and precise compensation in the faulty site area.
[0104] The present application also provides a VTS radar station multi-site collaboration device, comprising:
[0105] The constructing module constructs a cooperative decision matrix, and the cooperative decision matrix comprises real-time coverage capability weights of each radar station;
[0106] The distributing module generates a station task distribution instruction set based on the cooperative decision matrix.
[0107] The constructing the cooperative decision matrix comprises: obtaining historical performance data and real-time observation data of each radar station; and optimizing the real-time coverage capability weights based on the historical performance data and the real-time observation data through a reinforcement learning model; wherein the reinforcement learning model is guided to learn through a reward function, and the reward function is: , , , is a control coefficient, G is tracking continuity, Z is resource consumption, is actual coverage capability of a station evaluated based on real observation data, is a current estimated weight output by the reinforcement learning model; and the optimized real-time coverage capability weights are output to the cooperative decision matrix.
[0108] The application further provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method.
[0109] The application further provides a computer readable storage medium, which stores a computer program, and the computer program, when executed in a computer, causes the computer to execute the method.
[0110] The above description of the embodiments is for facilitating those skilled in the art to understand and apply the application. Those skilled in the art can easily make various modifications to the above embodiments, and apply the general principles described herein to other embodiments without creative labor. Therefore, the application is not limited to the above embodiments, and the improvements and modifications made to the application by those skilled in the art according to the disclosure of the application should be within the protection scope of the application.
Claims
1. A VTS radar station multi-site collaboration method, characterized in that: include: Constructing a collaborative decision-making matrix, wherein the collaborative decision-making matrix includes the real-time coverage capability weights of each radar site; generating a site task allocation instruction set based on the collaborative decision matrix; The collaborative decision-making matrix construction includes: obtaining historical performance data and real-time observation data of each radar site; optimizing the real-time coverage capability weight through a reinforcement learning model based on the historical performance data and the real-time observation data; wherein the reinforcement learning model guides learning through a reward function, and the reward function is: , 、 、 is the control coefficient, G is the tracking continuity, Z is the resource consumption, The actual coverage capability of the site is evaluated based on real observation data. The current estimated weight output by the reinforcement learning model; the optimized real-time coverage capability weight is output to the collaborative decision-making matrix.
2. The method according to claim 1, characterized in that The reinforcement learning model guides learning through a reward function, including: Get real-time tracking continuity indicators; Get current resource consumption indicators; Calculate the absolute error between the actual coverage capacity weight and the estimated coverage capacity weight; A reinforcement learning reward value is generated by the reward function based on the real-time tracking continuity indicator, the current resource consumption indicator and the absolute error.
3. The method according to claim 1, characterized in that The acquisition of real-time observation data from each radar site includes: Collect real-time signal-to-noise ratio of radar sites; Collect the real-time effective detection area of the radar site; Collect real-time communication bandwidth utilization of radar sites; A real-time observation data set is constructed based on the real-time signal-to-noise ratio, the real-time effective detection area, and the real-time communication bandwidth utilization.
4. The method according to claim 1, wherein The acquisition of historical performance data of each radar site includes: Count the number of site failures within the sliding time window; Calculate the target loss rate of the overlapping area within the sliding time window; A historical performance data set is constructed based on the number of site failures and the target loss rate of the overlapping area.
5. The method according to claim 1, wherein The optimization of the real-time coverage capability weight by the reinforcement learning model includes: Feed historical performance data and real-time observation data into the reinforcement learning model; Evaluate the action value through the reward function; Update the action policy to minimize the weight prediction error; Output the optimized real-time coverage capability weight.
6. The method according to claim 1, characterized in that Generating a site task allocation instruction set based on the collaborative decision matrix includes: Sort the site priorities based on real-time coverage capability weights; Calculate the optimal takeover site set based on the location of the failed site; Generate a task allocation instruction set including a takeover strategy.
7. The method according to claim 2, characterized in that The obtaining of real-time tracking continuity indicators includes: Count the number of target tracks successfully spliced within the preset time window; Calculate the total number of targets that should be tracked; Track continuity values via ratio calculations: Tracking continuity = number of successfully spliced tracks / total number of targets to be tracked.
8. A VTS radar station multi-site collaboration device, characterized in that: include: A construction module is used to construct a collaborative decision matrix, wherein the collaborative decision matrix includes the real-time coverage capability weights of each radar site; an allocation module, generating a site task allocation instruction set based on the collaborative decision matrix; The collaborative decision-making matrix construction includes: obtaining historical performance data and real-time observation data of each radar site; optimizing the real-time coverage capability weight through a reinforcement learning model based on the historical performance data and the real-time observation data; wherein the reinforcement learning model guides learning through a reward function, and the reward function is: , 、 、 is the control coefficient, G is the tracking continuity, Z is the resource consumption, The actual coverage capability of the site is evaluated based on real observation data. The current estimated weight output by the reinforcement learning model; the optimized real-time coverage capability weight is output to the collaborative decision-making matrix.
9. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that A computer program is stored thereon, and when the computer program is executed in a computer, the computer is caused to execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Modal layered enhanced multi-agent cooperative control method and related device
CN120508137A
USV formation path-following method based on deep reinforcement learning
US20220004191A1
Methods and systems for controlling multi-UAV cooperative tracking of ground-moving targets
US20250224735A1
Collaborative task offloading and service caching method based on graph attention multi-agent reinforcement learning
WO2025050608A1
Cited By
Radar resource dynamic allocation method and system for multi-task reinforcement learning
CN120993332A
A method and system for dynamic allocation of radar resources based on multi-task reinforcement learning
CN120993332B