A method, apparatus, equipment, and storage medium for multi-site coordination of VTS radar stations.
By constructing a collaborative decision matrix and using a reinforcement learning model to optimize the weights of real-time coverage capability, the problem of lagging weight calculation in the traditional VTS radar site collaboration method is solved, achieving reasonable resource allocation and improved target tracking accuracy, and enhancing the system's robustness in complex environments.
Patent Information
- Application Number
- CN202511302805.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-12
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-09-12
AI Technical Summary
In traditional VTS radar station multi-site collaboration methods, weight calculation relies on static or semi-static data such as historical failure counts and target loss rates in overlapping areas. This results in an inability to accurately reflect the real-time operating status and capabilities of radar stations, leading to resource waste and decreased tracking accuracy. In particular, the system's robustness is reduced in dynamically changing water traffic environments.
By constructing a collaborative decision matrix, historical performance data and real-time observation data of each radar station are obtained. The weights of real-time coverage capability are optimized using a reinforcement learning model. Combined with reward function-guided learning, a set of station task allocation instructions is generated to ensure the accuracy and flexibility of task allocation.
It achieves rational allocation and efficient utilization of resources, improves target tracking accuracy and system robustness, and can effectively cope with the dynamic changes in complex water traffic environment.
Smart Images

Figure CN120779347B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of radar coordination, and in particular relates to a method, apparatus, equipment and storage medium for multi-site coordination of VTS radar stations. Background Technology
[0002] In traditional VTS radar station multi-site coordination methods, a key step is constructing a collaborative decision matrix. This matrix guides task allocation among radar stations, ensuring the efficient and accurate operation of the entire system. However, existing collaborative decision matrix construction methods have significant shortcomings, particularly in the calculation of real-time coverage capability weights.
[0003] Traditional weighting calculations often rely on static or semi-static data such as historical failure counts and target loss rates in overlapping areas. This results in weight allocation failing to accurately reflect the real-time operational status and capabilities of radar sites. For example, a recently maintained radar site with good current performance may have its actual coverage capability underestimated by the system due to its historical failure records, leading to its marginalization in task allocation, wasted resources, and a mismatch between requirements and needs. Furthermore, traditional methods are inadequate when dealing with dynamically changing water traffic environments.
[0004] When radar stations malfunction or experience performance degradation, the system cannot adjust its task allocation strategy in a timely manner due to the lag in weight calculation, leading to decreased tracking accuracy and reduced system robustness. In complex and ever-changing traffic environments, this lag can trigger global coordination failures, severely impacting the operational efficiency of the VTS system. Summary of the Invention
[0005] The purpose of this application is to overcome the deficiencies in the prior art and provide a method, apparatus, device and storage medium for multi-site collaboration of VTS radar stations.
[0006] This application provides a multi-site coordination method for VTS radar stations, including:
[0007] Construct a collaborative decision matrix, which includes the real-time coverage capability weights of each radar station;
[0008] A set of site task allocation instructions is generated based on the collaborative decision matrix;
[0009] The construction of the collaborative decision matrix includes: acquiring historical performance data and real-time observation data of each radar station; optimizing the real-time coverage capability weights based on the historical performance data and the real-time observation data using a reinforcement learning model; wherein the reinforcement learning model is guided by a reward function, and the reward function is: , , , Here, G represents the control coefficient, G represents the tracking continuity, and Z represents resource consumption. To assess the actual coverage capability of stations based on real observation data, To reinforce the current estimated weights output by the learning model, the optimized real-time coverage capability weights are output to the collaborative decision matrix.
[0010] Optionally, the reinforcement learning model guides learning through a reward function, including:
[0011] Obtain real-time tracking continuity metrics;
[0012] Get current resource consumption metrics;
[0013] Calculate the absolute error between the actual coverage weight and the estimated coverage weight;
[0014] Based on the real-time tracking continuity indicator, the current resource consumption indicator, and the absolute error, a reinforcement learning reward value is generated through the reward function.
[0015] Optionally, acquiring real-time observation data from each radar station includes:
[0016] Real-time signal-to-noise ratio of radar stations;
[0017] Collect the real-time effective detection area of the radar station;
[0018] Real-time communication bandwidth utilization of radar sites;
[0019] A real-time observation dataset is constructed based on the real-time signal-to-noise ratio, the real-time effective detection area, and the real-time communication bandwidth utilization.
[0020] Optionally, acquiring historical performance data for each radar station includes:
[0021] Count the number of site failures within the sliding time window;
[0022] Calculate the target loss rate in the overlapping area within the sliding time window;
[0023] A historical performance dataset is constructed based on the number of site failures and the target loss rate in the overlapping area.
[0024] Optionally, optimizing the real-time coverage capability weights through a reinforcement learning model includes:
[0025] Input historical performance data and real-time observation data into the reinforcement learning model;
[0026] The value of an action is evaluated using the reward function.
[0027] Update the action policy to minimize weight prediction error;
[0028] Output the optimized real-time coverage capability weights.
[0029] Optionally, generating the site task allocation instruction set based on the collaborative decision matrix includes:
[0030] Site priority is sorted by real-time coverage capability weight;
[0031] Calculate the optimal set of takeover sites based on the location of the faulty site;
[0032] Generate a task assignment instruction set that includes the takeover strategy.
[0033] Optionally, obtaining the real-time tracking continuity indicator includes:
[0034] Count the number of target trajectories successfully spliced within a preset time window;
[0035] Calculate the total number of targets to be tracked;
[0036] The value of continuous tracking is calculated using the ratio:
[0037] Tracking continuity = Number of successfully stitched trajectories / Total number of targets to be tracked.
[0038] This application also provides a VTS radar station multi-site coordination device, including:
[0039] The module constructs a collaborative decision matrix, which includes the real-time coverage capability weights of each radar station.
[0040] The allocation module generates a set of site task allocation instructions based on the collaborative decision matrix;
[0041] The construction of the collaborative decision matrix includes: acquiring historical performance data and real-time observation data of each radar station; optimizing the real-time coverage capability weights based on the historical performance data and the real-time observation data using a reinforcement learning model; wherein the reinforcement learning model is guided by a reward function, and the reward function is: , , , Here, G represents the control coefficient, G represents the tracking continuity, and Z represents resource consumption. To assess the actual coverage capability of stations based on real observation data, To reinforce the current estimated weights output by the learning model, the optimized real-time coverage capability weights are output to the collaborative decision matrix.
[0042] This application also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the above-described method.
[0043] This application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to perform the above-described method.
[0044] The beneficial effects of this application are:
[0045] This application provides a multi-site collaboration method for VTS radar stations, comprising: constructing a collaborative decision matrix, wherein the collaborative decision matrix includes the real-time coverage capability weights of each radar station; generating a site task allocation instruction set based on the collaborative decision matrix; wherein, constructing the collaborative decision matrix includes: acquiring historical performance data and real-time observation data of each radar station; optimizing the real-time coverage capability weights based on the historical performance data and the real-time observation data through a reinforcement learning model; wherein, the reinforcement learning model is guided by a reward function, wherein the reward function is: , , , Here, G represents the control coefficient, G represents the tracking continuity, and Z represents resource consumption. To assess the actual coverage capability of stations based on real observation data, To reinforce the current estimated weights output by the reinforcement learning model, the optimized real-time coverage capability weights are output to the collaborative decision matrix. This application optimizes the real-time coverage capability weights of the VTS radar station collaborative decision matrix using a reinforcement learning model, achieving rational allocation and efficient utilization of resources, improving target tracking accuracy and system robustness, and effectively coping with complex water traffic environments. Attached Figure Description
[0046] Figure 1 This is a schematic diagram of the multi-site collaborative process of the VTS radar station in this application;
[0047] Figure 2 This is a schematic diagram of the multi-source data preprocessing process in this application;
[0048] Figure 3 This is a time-series diagram of the dynamic weight adjustment of the collaborative decision-making layer in this application;
[0049] Figure 4 This is a schematic diagram of the dual-mode weight switching state machine in this application. Detailed Implementation
[0050] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it is to be understood that various forms of implementation of the present disclosure are intended and should not be limited to the embodiments set forth herein. Rather, the embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0051] Please refer to Figure 1 As shown, this application provides a multi-site cooperation method for VTS radar stations, including:
[0052] S101. Construct a collaborative decision matrix, which includes the real-time coverage capability weights of each radar station.
[0053] Constructing a collaborative decision matrix is a core component of the VTS radar station multi-site collaborative method, aiming to dynamically evaluate the coverage capabilities of each site to optimize task allocation. The collaborative decision matrix is a data structure used to store and update key status parameters of radar sites, including real-time coverage capability weights, quantized values of detection overlap areas between adjacent sites, and site fault status indicators.
[0054] The real-time coverage capability weight represents the station's detection efficiency of water targets at the current moment; the higher the value, the stronger the coverage capability.
[0055] Constructing this matrix requires integrating historical performance data and real-time observation data, and optimizing the weight values through a reinforcement learning model.
[0056] During the construction of the collaborative decision-making layer, the historical performance data of each radar station is first acquired. This historical performance data includes the station's operational records within a specific time window. Specifically, the system employs a sliding window strategy to truncate the historical data, focusing on the station's performance within the most recent hour. This processing mechanism first establishes a time index cache structure, registering in real-time the time of each station's failure and the corresponding target loss records within the overlapping area. The system periodically slides the window boundaries, retaining only data whose timestamps fall within the current window range, automatically discarding any data outside the time window.
[0057] Number of site failures The timestamp of each failure within the current time window is counted, and the number is recalculated in each round of weight update to ensure that the recent stability status is reflected.
[0058] For the target loss rate in the overlapping area ( The system compares the theoretically expected number of targets to be detected with the actual number of successfully tracked targets within the overlapping area, and calculates the loss rate using the difference ratio. The calculation method is as follows:
[0059]
[0060] in, This represents the actual number of targets that were continuously detected by multiple stations and whose trajectories were stitched together within the time window. This represents the number of targets that should theoretically be successfully tracked.
[0061] The consistency of the target trajectory is verified through a fusion engine. This method ensures that historical data only reflects the current system state and avoids interference from outdated information.
[0062] At the same time, the system acquires real-time observation data from each radar station.
[0063] Real-time observation data includes the instantaneous performance indicators of the stations. Specifically, the system collects the real-time signal-to-noise ratio of the radar stations. This value is obtained by measuring the target reflection intensity through the radar system, reflecting the signal quality; the real-time effective detection area is also collected. This indicates the current water area covered by the station; it also collects real-time communication bandwidth utilization. This data reflects the site's data transmission capabilities. It is collected in real-time by sensors and network monitoring modules and compiled into a real-time observation dataset. This real-time observation data is used to dynamically update the site's status, avoiding historical biases.
[0064] Please refer to Figure 2 As shown, each radar station first acquires raw point cloud data of targets within its coverage area in real time using a high-resolution radar system. This point cloud data reflects the reflection characteristics of the targets in the radar beam, including spatial coordinates, reflection intensity, and timestamp information. Due to the influence of environmental noise, electromagnetic interference, and the performance of the equipment itself, the raw point cloud data usually contains a large number of non-target points, scattering points, and artifacts. Therefore, noise filtering algorithms are required to purify it.
[0065] The filtering process employs a density-based spatial anomaly removal method. This method evaluates the point density and intensity thresholds of local regions in the point cloud to remove isolated points and low-energy echo points, thereby improving the clarity of the target contour. After denoising, the next step is coordinate unification. Point cloud data acquired from different sites are typically in their respective local coordinate systems. To achieve subsequent cross-site trajectory matching, all point cloud data must be transformed to a unified geographic reference coordinate system. The coordinate transformation process is based on the pre-calibrated geographic location parameters and attitude angle information of the sites, and is performed using a rigid body transformation model. The formula is:
[0066]
[0067] in, This represents the position vector of the target in a unified coordinate system. This is the original local coordinate position. For rotation matrix, The translation vectors are both obtained from the site's geographic orientation calibration.
[0068] After coordinate unification, motion trajectories are extracted from point cloud data within the time series. A trajectory fitting method based on sliding time windows is used to convert continuous observation segments of the target into standardized target trajectory segments. Each segment contains a unified spatiotemporal position vector, velocity vector, and confidence score, providing a highly consistent data foundation for subsequent distributed data fusion and trajectory matching.
[0069] Please refer to Figure 3 As shown, further, in the distributed data fusion stage, the standardized target trajectory segments preprocessed and formed by each radar station are input in real time to the regional-level data fusion engine, which possesses a core of efficient trajectory matching and fusion algorithms. First, based on the continuity of the target in spatial location and timestamp, the system performs spatiotemporal correlation analysis on the target trajectories across stations. By comparing the target's velocity vector, trajectory trend, and spatial distance between adjacent frames, it identifies trajectory segments belonging to the same target but observed by different stations. For multiple matched sets of observations, the fusion engine uses an adaptive weighted fusion algorithm to generate unified target trajectory points. This fusion process ensures that the fusion result maintains accuracy while taking into account the communication performance of each station by allocating observation weights to different stations. Specifically, the fused value ( The calculation form of ) is:
[0070]
[0071] in, It is the measurement value of the target from the i-th station. This represents the fusion weight for this observation. Weight ( The determination of takes into account two factors: measurement accuracy and communication bandwidth, and its expression is:
[0072]
[0073] in, This represents the measurement variance of the i-th station, reflecting the accuracy of its observation data. The smaller the value, the smaller the error and the higher the reliability. This indicates the current communication bandwidth utilization rate of the site. This represents the maximum bandwidth utilization rate in the system, used for normalization.
[0074] This weighted design ensures that while prioritizing measurement reliability, the influence of different stations in the fusion process is dynamically adjusted, preventing stations with poor communication capabilities from excessively affecting fusion accuracy. In this way, the fusion engine constructs a unified target trajectory set across the entire domain, encompassing observation information from all radar stations, and guarantees robust consistency of data under complex network conditions, laying a solid foundation for subsequent collaborative decision-making and task allocation.
[0075] After acquiring historical performance data and real-time observation data, the system optimizes the weights of real-time coverage capability through a reinforcement learning model.
[0076] Please refer to Figure 4 As shown, in collaborative task scheduling, to cope with the dynamic changes in radar site status, the system is designed with a dual-mode weight calculation switching mechanism to ensure a balance between system stability and response agility. In the normal mode, the system comprehensively uses historical statistical data and real-time observations, and calculates the signal-to-noise ratio and effective detection area estimate for each site through Kalman filtering and sliding window processing mechanisms, thereby obtaining a more stable weight result, which is suitable for most stable operating scenarios.
[0077] However, when the system detects a sudden change in the state of a site, such as a sharp drop in continuous signal-to-noise ratio, a sudden reduction in detected area, or frequent failures in a short period of time, the system will immediately trigger a switching mechanism and enter a temporary real-time emergency mode. In this mode, the weights no longer rely on historical information or filtered prediction values, but are directly calculated based on the current real-time observation data to determine the coverage capability weights. The formula is:
[0078]
[0079] in, The instantaneous signal-to-noise ratio measured at the current site. This represents the current actual detected area. This represents the maximum real-time signal-to-noise ratio across all stations. The sum of the current detected areas of all stations. and It is the basic weighting coefficient used in emergency mode, emphasizing the priority of real-time response.
[0080] When the mutation index continues to decline or fluctuates, the system maintains real-time operation to quickly adjust the task allocation structure. Once the index returns to stability and meets the stability threshold, the system will automatically switch back to normal mode and resume the integrated use of historical trends and state prediction results, thereby balancing stability and sensitivity and improving the accuracy and resilience of task allocation in complex dynamic environments.
[0081] A reinforcement learning model is an algorithmic framework that uses historical site performance and real-time observations as state inputs, combining dimensions such as trajectory tracking continuity, resource utilization efficiency, and prediction accuracy, and continuously adjusts the weight estimation strategy through an interactive process. The core mechanism is the design of a reward function to guide the model's learning.
[0082] The reward function is defined as ( ):
[0083]
[0084] in, , , The control coefficients control the impact of the three indicators on the reward; G represents tracking continuity, which measures the system's ability to maintain the integrity of the target trajectory; Z represents resource consumption, which represents the combined indicator of communication and computing resources consumed by the station per unit time. This is an assessment of the actual coverage capability of the stations based on real observation data. The current estimated weights are used to reinforce the output of the learning model.
[0085] The action value function is updated using a temporal difference method. By positively reinforcing excellent policies and negatively penalizing biased behaviors, the model is forced to learn a high-dimensional mapping relationship between variables such as signal-to-noise ratio, detection area, and fault records and coverage capability.
[0086] The optimization process includes: inputting historical performance data and real-time observation data into the reinforcement learning model; evaluating the value of actions through a reward function and generating reward values; updating the action policy to minimize weight prediction errors; and finally outputting the optimized real-time coverage capability weights into the collaborative decision matrix. These weights more closely reflect the actual performance of the site, improving environmental adaptability.
[0087] The reinforcement learning model, guided by a reward function, involves acquiring real-time tracking continuity metrics, acquiring current resource consumption metrics, and calculating the absolute error between the actual coverage weights and the estimated coverage weights. The acquisition of real-time tracking continuity metrics involves: counting the number of successfully stitched target trajectories within a preset time window, calculating the total number of targets to be tracked, and then calculating the tracking continuity value using the ratio.
[0088]
[0089] Where D represents tracking continuity, The number of successfully spliced trajectories. This refers to the total number of targets to be tracked. This metric reflects the system's ability to maintain target continuity.
[0090] The current resource consumption metrics are obtained by measuring the site's resource usage during scanning and data transmission using the monitoring module. The absolute error is calculated as follows: It is based directly on observations. Based on these metrics, the reward function generates reinforcement learning reward values, driving model optimization.
[0091] S102. Generate a set of site task allocation instructions based on the collaborative decision matrix.
[0092] The generation of the site task allocation instruction set is the execution phase of the method, used for dynamically scheduling radar site tasks. The instruction set includes instructions such as dynamic takeover strategies for the coverage areas of faulty sites, ensuring the continuity of waterway monitoring. Specifically, the system prioritizes sites according to their real-time coverage capability weights in the collaborative decision matrix, allocating tasks to sites with higher weights first; it calculates the optimal set of sites to take over based on the location of the faulty site; and finally, it generates a task allocation instruction set containing the takeover strategy.
[0093] After the system completes the construction of the collaborative decision matrix and generates the task allocation instruction set, it enters the adaptive target tracking and output stage. This stage uses the instruction set as the execution basis, dynamically scheduling radar stations currently in a healthy state to enhance their monitoring capabilities over the original coverage areas of nearby faulty stations. The system first reallocates scanning tasks within the spatial range based on trajectory continuity requirements, coverage redundancy, and the current station load capacity. This is achieved by increasing the scanning frequency and adjusting beam pointing strategies to fill the sensing gaps caused by faulty stations. For example, the default scanning frame rate for healthy stations is increased from 4 revolutions per minute in normal mode to 8 revolutions per minute in emergency mode. The refresh rate is increased by encrypting the scanning cycle, shortening the target update interval, reducing trajectory gap time caused by takeover, and enhancing target trajectory continuity. During this process, the fusion module continuously receives standardized target trajectory fragments uploaded by each station and updates the target status in real time based on the previously constructed unified trajectory structure. To avoid information delays or data conflicts, the system employs a timestamp synchronization mechanism and a data weighting adjustment strategy to ensure consistency in timing and accuracy of the same target information reported by multiple stations. The final fusion output includes the target's globally unique identifier, historical trajectory sequence, current motion status, and predicted trend. This unified encapsulation is then pushed to the VTS central display terminal, enabling comprehensive perception and visualization of dynamic targets within the water area. This supports upper-level functions such as command and dispatch, early warning linkage, and situational analysis. This mechanism possesses fault tolerance and self-healing capabilities, ensuring trajectory continuity even in the event of node failure. It also enhances the system's responsiveness to emergencies in highly dynamic environments, improving the overall robustness and operational effectiveness of the monitoring system.
[0094] The site priority sorting operation is as follows: read the real-time coverage capability weights from the matrix. Sites are sorted in descending order, with higher-weighted sites considered healthy and efficient, and prioritized for monitoring key areas.
[0095] The calculation operation of the optimal takeover site set is as follows: when the system detects a faulty site, a weighted minimization strategy is introduced to select the optimal subset from the candidate takeover sites.
[0096] This strategy uses the inverse relationship between distance cost and weight as a comprehensive evaluation index, and the objective function is:
[0097]
[0098] in, This represents the spatial distance from station j to the faulty station k. This is the maximum effective take-off distance set by the system. Let j be the weight of site j. This function reflects the principle that sites with closer proximity and higher weights are preferentially selected as takeover nodes. The system iterates through all sites that meet the following conditions. ≤ The site, filter the set that minimizes the objective function { The system then generates specific takeover instructions. After the instruction set is updated, an enhanced scan of the health sites is scheduled.
[0099] During the adaptive target tracking and output phase, the system, based on the instruction set, schedules healthy radar stations to enhance monitoring capabilities over the original coverage areas of faulty stations. The system reallocates scanning tasks spatially based on trajectory continuity requirements, coverage redundancy, and current station load capacity, filling perception gaps by increasing scanning frequency and adjusting beam pointing strategies. For example, the scanning frame rate is increased from 4 revolutions per minute to 8 revolutions per minute, and the scanning cycle is encrypted. The fusion module receives standardized target trajectory segments, updates the target status based on a unified trajectory structure across the entire domain, and employs a timestamp synchronization mechanism and data weight adjustment strategy to ensure consistency. Finally, the fusion result is output to the VTS central display and control terminal to support ship traffic management.
[0100] During fault takeover, to mitigate the abrupt velocity vector changes during cross-site indirect management, the system employs a Doppler frequency shift compensation algorithm. This algorithm calculates the expected velocity value of the target within the cross-site time window based on the radial velocity variation trend in historical continuous frames, and corrects velocity discontinuities in the fused trajectory caused by observation errors. The correction method is based on weighted interpolation, with the expected velocity set as... The actual measured value is By generating a corrected velocity ( ):
[0101]
[0102] Wherein, λ is a correction coefficient, which is dynamically adjusted between 0 and 1 to balance real-time performance and continuity.
[0103] This mechanism effectively mitigates tracking jumps caused by the heterogeneity of radar distribution, improving the smoothness of trajectory fusion and the overall stability of the system during fault takeover periods. The entire process is completed automatically within the fusion engine without manual intervention, enabling rapid response and accurate compensation for faulty site areas.
[0104] This application also provides a VTS radar station multi-site coordination device, including:
[0105] The module constructs a collaborative decision matrix, which includes the real-time coverage capability weights of each radar station.
[0106] The allocation module generates a set of site task allocation instructions based on the collaborative decision matrix;
[0107] The construction of the collaborative decision matrix includes: acquiring historical performance data and real-time observation data of each radar station; optimizing the real-time coverage capability weights based on the historical performance data and the real-time observation data using a reinforcement learning model; wherein the reinforcement learning model is guided by a reward function, and the reward function is: , , , Here, G represents the control coefficient, G represents the tracking continuity, and Z represents resource consumption. To assess the actual coverage capability of stations based on real observation data, To reinforce the current estimated weights output by the learning model, the optimized real-time coverage capability weights are output to the collaborative decision matrix.
[0108] This application also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the above-described method.
[0109] This application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to perform the above-described method.
[0110] The above description of the embodiments is provided to enable those skilled in the art to understand and apply this application. Those skilled in the art will readily make various modifications to the above embodiments and apply the general principles described herein to other embodiments without inventive effort. Therefore, this application is not limited to the above embodiments, and any improvements and modifications made to this application based on the disclosure thereof should be within the scope of protection of this application.
Claims
1. A multi-site coordination method for VTS radar stations, characterized in that, include: Construct a collaborative decision matrix, which includes the real-time coverage capability weights of each radar station; A set of site task allocation instructions is generated based on the collaborative decision matrix; The construction of the collaborative decision matrix includes: acquiring historical performance data and real-time observation data of each radar station; optimizing the real-time coverage capability weights based on the historical performance data and the real-time observation data using a reinforcement learning model; wherein the reinforcement learning model is guided by a reward function, and the reward function is: , , , Here, G represents the control coefficient, G represents the tracking continuity, and Z represents resource consumption. To assess the actual coverage capability of stations based on real observation data, To reinforce the current estimated weights output by the learning model, the optimized real-time coverage capability weights are output to the collaborative decision matrix.
2. The method according to claim 1, characterized in that, The reinforcement learning model is guided by a reward function, including: Obtain real-time tracking continuity metrics; Get current resource consumption metrics; Calculate the absolute error between the actual coverage weight and the estimated coverage weight; Based on the real-time tracking continuity indicator, the current resource consumption indicator, and the absolute error, a reinforcement learning reward value is generated through the reward function.
3. The method according to claim 1, characterized in that, The acquisition of real-time observation data from each radar station includes: Real-time signal-to-noise ratio of radar stations; Collect the real-time effective detection area of the radar station; Real-time communication bandwidth utilization of radar sites; A real-time observation dataset is constructed based on the real-time signal-to-noise ratio, the real-time effective detection area, and the real-time communication bandwidth utilization.
4. The method according to claim 1, characterized in that, The acquisition of historical performance data for each radar station includes: Count the number of site failures within the sliding time window; Calculate the target loss rate in the overlapping area within the sliding time window; A historical performance dataset is constructed based on the number of site failures and the target loss rate in the overlapping area.
5. The method according to claim 1, characterized in that, The optimization of real-time coverage capability weights through a reinforcement learning model includes: Input historical performance data and real-time observation data into the reinforcement learning model; The value of an action is evaluated using the reward function. Update the action policy to minimize weight prediction error; Output the optimized real-time coverage capability weights.
6. The method according to claim 1, characterized in that, The generation of the site task allocation instruction set based on the collaborative decision matrix includes: Site priority is sorted by real-time coverage capability weight; Calculate the optimal set of takeover sites based on the location of the faulty site; Generate a task assignment instruction set that includes the takeover strategy.
7. The method according to claim 2, characterized in that, The acquisition of real-time tracking continuity indicators includes: Count the number of target trajectories successfully spliced within a preset time window; Calculate the total number of targets to be tracked; The value of continuous tracking is calculated using the ratio: Tracking continuity = Number of successfully stitched trajectories / Total number of targets to be tracked.
8. A VTS radar station multi-site coordination device, characterized in that, include: The module constructs a collaborative decision matrix, which includes the real-time coverage capability weights of each radar station. The allocation module generates a set of site task allocation instructions based on the collaborative decision matrix; The construction of the collaborative decision matrix includes: acquiring historical performance data and real-time observation data of each radar station; optimizing the real-time coverage capability weights based on the historical performance data and the real-time observation data using a reinforcement learning model; wherein the reinforcement learning model is guided by a reward function, and the reward function is: , , , Here, G represents the control coefficient, G represents the tracking continuity, and Z represents resource consumption. To assess the actual coverage capability of stations based on real observation data, To reinforce the current estimated weights output by the learning model, the optimized real-time coverage capability weights are output to the collaborative decision matrix.
9. An electronic device, characterized in that, It includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed in a computer, causes the computer to perform the method described in any one of claims 1-7.
Citation Information
Patent Citations
Modal layered enhanced multi-agent cooperative control method and related device
CN120508137A
USV formation path-following method based on deep reinforcement learning
US20220004191A1