Dynamic path planning method, system and device based on ground wave radar
By using a shore-based and shipborne radar collaborative sensing network, reinforcement learning, and ocean numerical assimilation technology, the problems of static paths, limited coverage, insufficient coordination, and slow emergency response in ocean observation have been solved, enabling efficient, safe, and adaptive observation path planning in dynamic ocean environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZHUHAI OCEAN CENTER OF THE MINISTRY OF NATURAL RESOURCES (ZHUHAI OCEAN FORECAST STATION OF THE MINISTRY OF NATURAL RESOURCES)
- Filing Date
- 2026-02-24
- Publication Date
- 2026-05-12
AI Technical Summary
Existing marine observation technologies suffer from problems such as static observation paths, inability to adapt to dynamic marine environments, limited coverage of single-point sampling, insufficient coordination between radar and mobile platforms, slow emergency replanning response, and uneven distribution of observation data value.
By constructing a collaborative sensing network of shore-based and shipborne dual-layer radars, combining reinforcement learning and a lightweight emergency replanning mechanism, and introducing marine numerical assimilation to form a closed-loop optimization, efficient observation path planning in dynamic marine environments can be achieved.
It achieves maximum observation coverage and value, strong adaptability to dynamic environments, high resource utilization efficiency, system intelligence and self-evolution, deep integration of multi-source information, and improved timeliness and security of path planning.
Smart Images

Figure CN121702412B_ABST
Abstract
Description
Technical Field
[0001] This application pertains to the field of data application of ground wave radar detection signals, and particularly relates to a dynamic path planning method, system, and device based on ground wave radar. Background Technology
[0002] With the deepening of the national strategy of building a maritime power, the demand for real-time, comprehensive, and accurate observation data in fields such as marine dynamic element monitoring, emergency maritime response, and marine resource exploration is becoming increasingly urgent. Marine dynamic observation technology has become a core foundation supporting marine scientific research and the protection of maritime rights. Currently, marine observation technology has formed a composite system of "shore-based fixed observation + mobile platform sampling." Shore-based radar can achieve wide-area coverage, while mobile platforms such as unmanned survey vessels and underwater gliders can complete fine-grained point sampling. Reinforcement learning algorithms are also gradually being applied to mobile platform path planning. However, existing technologies still have many bottlenecks.
[0003] Current marine observation technologies face several major challenges: static observation paths that are ill-suited to dynamic marine environments; limited coverage of single-point sampling; insufficient coordination between radar and mobile platforms; slow response to emergency replanning; and uneven distribution of the value of observation data.
[0004] Therefore, there is a lack of radar-mobile platform collaborative ocean dynamic observation path planning schemes that can provide technical support for efficient observation in dynamic ocean environments.
[0005] The foregoing statements are for informational purposes only and are not intended to provide background information in connection with this application. Unless otherwise stated herein, the content described in this section is not prior art to the rest of this application. Summary of the Invention
[0006] This invention proposes a dynamic path planning method, system, and device based on ground wave radar. By constructing a shore-based and shipborne dual-layer radar collaborative sensing network, combining reinforcement learning and a lightweight emergency replanning mechanism, and introducing ocean numerical assimilation to form a closed-loop optimization, this invention ultimately achieves efficient, safe, and adaptive planning of observation paths in dynamic ocean environments.
[0007] According to a first aspect of the embodiments of this application, a dynamic path planning method based on ground wave radar is provided, comprising the following steps:
[0008] Wide-area dynamic data of the observation sea area are obtained through shore-based high-frequency ground wave radar, and short-range environmental information around the mobile platform is obtained through shipborne navigation radar.
[0009] Preprocessing and fusing wide-area dynamic data and near-range environmental information to construct a dynamic state vector containing multi-dimensional environmental features and uncertainty representations;
[0010] The dynamic state vector is input into a path planning model based on reinforcement learning. Through optimization using a multi-objective reward function, the initial heading and speed of the mobile platform are output, forming the initial observation path.
[0011] In some embodiments of this application, after forming the initial observation path, the method further includes:
[0012] Control the mobile platform to execute the initial observation path and continuously acquire real-time radar observation data and platform status data;
[0013] Based on real-time radar observation data and platform status data, the preset hierarchical replanning trigger conditions are verified in real time.
[0014] If any replanning trigger condition is triggered, the lightweight reinforcement learning model is invoked, an emergency priority weighting term is introduced into the multi-objective reward function, and an emergency replanning path is generated and issued.
[0015] In some embodiments of this application, after generating and issuing the emergency rerouting path, the method further includes:
[0016] Step S7: Collaborative optimization and closed-loop feedback: Control the mobile platform to execute emergency replanning of the path, and feed back the path information of the mobile platform to the shore-based high-frequency ground wave radar to adjust its detection beam direction. At the same time, input the radar observation data into the ocean numerical assimilation module to correct the ocean forecast model, and feed back the corrected forecast data to generate a new path plan.
[0017] In some embodiments of this application, the dynamic state vector is:
[0018] ;
[0019] in, The data is wide-area observation grid data for shore-based radar, with dimensions of 51×51×3, corresponding to the flow field, wave height and target density channels respectively.
[0020] The data is grid data for short-range observation by shipborne radar, with dimensions of 51×51×2, corresponding to obstacle distribution and signal-to-noise ratio channels, respectively.
[0021] This represents the electromagnetic interference intensity and frequency band vector. Provides real-time Cartesian coordinates for mobile platforms.
[0022] In some embodiments of this application, the multi-objective reward function includes a composite reward function comprising radar observation value reward, obstacle avoidance reward, observation effectiveness reward, and platform energy consumption reward. The expression for the multi-objective reward function is as follows:
[0023] ;
[0024] Where α, β, γ, and δ are adaptive weighting coefficients that are dynamically adjusted according to the environment type;
[0025] Rewards for radar observation of ocean dynamic elements, and flow field gradients Wave height gradient Positive correlation;
[0026] The obstacle avoidance reward is dynamically assigned based on the distance between the path and the obstacle grid.
[0027] The reward for radar observation effectiveness is determined based on the signal-to-noise ratio (SNR).
[0028] It is a platform energy consumption reward, which is negatively correlated with the platform's actual power.
[0029] In some embodiments of this application, the replanning triggering conditions include at least one of wide-area environmental abrupt change, near-range risk approach, and radar observation failure;
[0030] The triggering condition for wide-area environmental abrupt change is: when the offset distance of the center of the flow field gradient region monitored in real time by the shore-based radar relative to the historical benchmark exceeds the first threshold, or when the abrupt change value of the wave height data exceeds the second threshold.
[0031] The triggering conditions for short-range risk are: when the shipborne radar detects in real time that the distance between the moving obstacle and the platform is less than the third threshold, or when the short-range signal-to-noise ratio drops sharply to below the fourth threshold within a preset time.
[0032] The triggering condition for radar observation failure is: when the intensity of electromagnetic interference exceeds the system's anti-interference threshold and the radar echo signal is continuously lost.
[0033] In some embodiments of this application, the path information of the mobile platform is fed back to the shore-based high-frequency ground wave radar to adjust its detection beam pointing. Specifically, the beam pointing angle is calculated using the following formula. :
[0034] ;
[0035] in,( , ) represents the location of the radar station. , () is the target area center of the platform.
[0036] According to a second aspect of the embodiments of this application, a dynamic path planning system based on ground wave radar is provided, comprising:
[0037] The collaborative sensing data acquisition module is used to acquire wide-area dynamic data of the observation sea area through shore-based high-frequency ground wave radar and to acquire short-range environmental information around the mobile platform through shipborne navigation radar.
[0038] The dynamic state model construction module is used to preprocess and fuse wide-area dynamic data and short-range environmental information to construct a dynamic state vector containing multi-dimensional environmental features and uncertainty representations.
[0039] The initial observation path generation module is used to input the dynamic state vector into the reinforcement learning-based path planning model, optimize it through a multi-objective reward function, and output the initial heading and speed of the mobile platform to form the initial observation path.
[0040] According to a third aspect of the embodiments of this application, a dynamic path planning device based on ground wave radar is provided, comprising: a storage unit for storing executable instructions; and a processing unit for connecting to the storage unit to execute the executable instructions to complete the dynamic path planning method based on ground wave radar.
[0041] According to a fourth aspect of the embodiments of this application, a computer-readable storage medium is provided having a computer program stored thereon; the computer program is executed by a processor to implement a dynamic path planning method based on ground wave radar.
[0042] The dynamic path planning method, system, and apparatus based on ground-wave radar of this application include: acquiring wide-area dynamic data of the observation sea area through shore-based high-frequency ground-wave radar and acquiring near-range environmental information of the mobile platform's surroundings through shipborne navigation radar; preprocessing and fusing the wide-area dynamic data and near-range environmental information to construct a dynamic state vector containing multi-dimensional environmental features and uncertainty representations; inputting the dynamic state vector into a path planning model based on reinforcement learning, optimizing it through a multi-objective reward function, and outputting the initial heading and speed of the mobile platform to form the initial observation path. This application achieves radar-mobile platform collaborative dynamic marine observation path planning by constructing a shore-based and shipborne dual-layer radar collaborative sensing network, providing technical support for efficient observation in dynamic marine environments.
[0043] Meanwhile, this application combines reinforcement learning with a lightweight emergency replanning mechanism and introduces ocean numerical assimilation to form a closed-loop optimization, ultimately achieving efficient, safe and adaptive planning of observation paths in dynamic ocean environments.
[0044] Compared with the prior art, the present invention has the following significant advantages:
[0045] 1) Observation coverage and value maximization: By combining the wide-area scanning of shore-based radar with short-range blind spot filling of shipborne radar, and combined with a reward function guided by the gradient of ocean dynamic elements, the platform is guided to prioritize the coverage of observation areas with high scientific value, which significantly improves the data benefits of a single observation mission.
[0046] 2) Strong adaptability to dynamic environments: The proposed hierarchical replanning triggering mechanism and lightweight emergency response model enable the system to respond to dynamic risks such as sudden changes in ocean currents, sudden obstacles, and strong electromagnetic interference in millisecond time, ensuring the continuity of observation missions and the safety of the platform.
[0047] 3) High resource utilization efficiency: Through the radar-platform collaborative interface, the radar beam can be tracked and focused on the target area, avoiding the energy waste of wide-area scanning and improving the utilization efficiency of radar resources.
[0048] 4) System intelligence and self-evolution: The introduction of an ocean numerical assimilation module continuously uses real-time observation data to improve the environmental model and feeds the improved model back to the path planning, forming a closed-loop system of continuous learning, which enables the path planning strategy to be continuously optimized as the task is executed.
[0049] 5) Deep fusion of multi-source information: The dynamic state modeling module not only integrates multi-source data, but also innovatively adds uncertainty representations such as error covariance and confidence weight, providing more reliable and richer environmental information input for reinforcement learning decision-making and improving the robustness of decision-making. Attached Figure Description
[0050] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0051] Figure 1 The diagram illustrates the steps of a dynamic path planning method based on ground wave radar according to an embodiment of this application.
[0052] Figure 2 The diagram illustrates the steps of another dynamic path planning method based on ground wave radar according to an embodiment of this application;
[0053] Figure 3 The diagram illustrates the steps of another dynamic path planning method based on ground wave radar according to an embodiment of this application;
[0054] Figure 4 The flowchart of the dynamic path planning and replanning method according to an embodiment of this application is shown in the figure;
[0055] Figure 5The diagram shows a schematic representation of a dynamic path planning system based on ground wave radar according to an embodiment of this application.
[0056] Figure 6 The diagram shows a system deployment principle according to an embodiment of this application;
[0057] Figure 7 The diagram shows the system module composition and data flow according to an embodiment of this application;
[0058] Figure 8 The diagram illustrates the principle of generating an initial path using the DQN algorithm according to an embodiment of this application.
[0059] Figure 9 The diagram shows a structural schematic of a ground wave radar-based dynamic path planning device 400 according to an embodiment of this application. Detailed Implementation
[0060] Regarding this application, the radar-mobile platform collaborative ocean dynamic observation path planning integrates the wide-area detection capability of shore-based ground wave radar and the short-range perception capability of shipborne radar. Combined with the reinforcement learning-driven dynamic perception and strategy replanning mechanism, it can make up for the shortcomings of poor adaptability of static paths and small single-point sampling coverage in traditional ocean observation, and improve the real-time performance, robustness and data value of observation.
[0061] Current ocean observation technologies face the following main problems:
[0062] 1. Static observation paths are not suitable for dynamic marine environments: Traditional path-dependent static forecast fields are difficult to cope with dynamic changes such as sudden changes in flow fields and sudden obstacles.
[0063] 2. Limited coverage of single-point sampling: Traditional mobile platforms have a small sampling range, making it difficult to achieve wide-area coverage and accurate sampling in high-value areas.
[0064] 3. Insufficient coordination between radar and mobile platform: Radar and mobile platform mostly operate independently, lacking two-way information linkage, resulting in low equipment utilization.
[0065] 4. Slow response time for emergency replanning: Traditional path replanning algorithms have long response times, making it difficult to meet real-time requirements.
[0066] 5. Uneven value of observation data: Traditional sampling is prone to missing high-value areas (such as flow field gradient areas and wave height change areas).
[0067] Specifically, on the one hand, traditional mobile platform path planning relies heavily on static ocean forecast fields or historical data, which cannot respond in real time to dynamic changes such as sudden changes in ocean current fields, sudden obstacles (such as floating objects and fishing boats), and electromagnetic interference. This can lead to the failure of preset paths, interruption of observation missions, and even platform safety risks. For example, some patented technologies propose a ship energy-saving path planning method based on high-frequency ground wave radar ocean current data, but it only considers ocean current factors for energy-saving purposes and lacks a comprehensive consideration of short-range risks, observation effectiveness, and dynamic replanning.
[0068] On the other hand, existing radar observations and mobile platform operations are often independent and lack coordination. As a wide-area sensing device, radar beam scanning is usually fixed or performed according to a preset mode, making it impossible to focus on high-value observation areas that the mobile platform is about to reach, resulting in low utilization of observation resources. At the same time, the data collected by the mobile platform has not been effectively fed back in real time and used to revise the marine environment model, thus failing to form a closed loop of "perception-decision-modeling".
[0069] Furthermore, while there are studies applying reinforcement learning to path planning, in complex and dynamic environments such as ocean observation, how to construct a state representation that integrates multi-source radar data (wide-area and short-range), design a multi-objective optimization function that simultaneously considers observational value, platform safety, and energy consumption, and achieve millisecond-level emergency replanning remain pressing technical challenges. For example, some patents disclose a vehicle wading path planning system that employs multi-sensor fusion and state estimation, but its application scenario, sensor architecture, and planning objectives are fundamentally different from those of dynamic ocean observation, and its methods cannot be directly applied to solve the problems addressed in this invention.
[0070] Therefore, there is an urgent need for a marine mobile observation path planning system and method that can achieve deep collaboration between radar and mobile platforms, possess real-time dynamic environment perception and intelligent replanning capabilities, and form a closed-loop feedback of observation data.
[0071] Against this backdrop, this invention proposes a development concept for a radar-mobile platform collaborative marine dynamic observation path planning system. This concept aims to integrate shore-based and shipborne dual-layer radar detection technologies, reinforcement learning dynamic decision-making technology, and marine numerical assimilation technology to construct a closed-loop observation system of "wide-area perception - intelligent planning - real-time replanning - data iteration," providing technical support for efficient observation in dynamic marine environments.
[0072] The concept for this invention is based on the following considerations:
[0073] 1. Radar detection and mobile platform observation have a basis for synergy and complementarity: shore-based high-frequency ground wave radar can realize wide-area dynamic monitoring of ocean current field and wave height within a range of 200km, shipborne navigation radar can complete short-range obstacle perception within 2km around the platform, and mobile platform can perform fine sampling of high-value areas identified by radar. The combination of the three can overcome the limitations of single equipment in terms of "wide-area but not accurate, and accurate but not full-area", and achieve the unity of observation range and accuracy.
[0074] 2. Reinforcement learning provides algorithmic support for dynamic path planning: Deep reinforcement learning (DQN) algorithms have verified their sequential decision-making advantages in the field of path planning. They can learn the optimal strategy through continuous interaction with the marine environment and adapt to changes in dynamic elements such as ocean currents and obstacles. Lightweight model modification can further meet the low-latency decision-making requirements of emergency scenarios and provide a feasible solution for dynamic replanning.
[0075] 3. Ocean numerical assimilation technology can enable two-way empowerment of observation and forecast: observation data from radar and mobile platforms can be assimilated into ocean numerical forecast models in real time to correct forecast errors, while the assimilated accurate forecast field can feed back into path planning, improve the targeting of observation tasks, and form an iterative optimization closed loop of "observation-forecast-planning".
[0076] This will solve the following technical problems:
[0077] 1) Solving the problem of poor adaptability of traditional observation paths: Traditional mobile platform observations mostly rely on static forecast fields to generate fixed paths, which cannot cope with dynamic scenarios such as sudden changes in flow fields and sudden obstacles. This system can achieve adaptive adjustment of paths through real-time radar perception and reinforcement learning replanning, thereby improving the completion rate of observation tasks.
[0078] 2) Solving the problem of uneven value of observation data: Traditional sampling has the drawbacks of "blindly covering and missing high-value areas". This system takes the gradient area of dynamic elements identified by radar as the core guide and combines it with a multi-objective reward function, which can significantly increase the sampling ratio of high-value areas and enhance the scientific research and application value of observation data.
[0079] 3) Solving the problem of insufficient coordination between radar and mobile platform: In the existing technology, radar and mobile platform mostly operate independently. The radar cannot focus on the observation area of the platform, and the platform cannot avoid the interference / risk areas monitored by the radar. The coordination interface of this system can realize two-way information linkage, improve the overall utilization rate of equipment and the safety of observation.
[0080] 4) Solving the problem of observation strategies not being able to continuously evolve: Traditional path planning algorithms have fixed parameters and cannot accumulate and optimize strategies with observation tasks. This system stores dynamic scene samples in an experience playback pool and uses the fusion and inversion results to provide feedback optimization, which can realize the self-iteration of strategies and improve the observation efficiency under complex sea conditions in the long term.
[0081] Compared to existing technologies, it has the following technological advantages:
[0082] 1. Strong wide-area dynamic observation coverage: shore-based high-frequency ground wave radar achieves wide-area coverage, while shipborne radar completes short-range blind spot filling. The two work together to break through the range limitation of single-point sampling of traditional mobile platforms. Dynamic state modeling technology integrates wide-area and short-range radar observation data, which can map large-scale changes in marine dynamic environment and small-scale sudden risk conditions in real time. Compared with the traditional static forecast field driven observation scheme, it realizes the perception of the dynamic environment of the whole area without blind spots and significantly expands the observation coverage.
[0083] 2. High timeliness of path planning decisions: Initial path generation relies on a multi-branch neural network to quickly extract core features from radar observations, ensuring decision delays are kept low. Emergency replanning employs a pruned and optimized lightweight DQN model, which, by simplifying the number of parameters and reusing pre-trained parameters, can quickly output a new, adapted path. Compared to the long response time of traditional path planning, timeliness is significantly improved, accurately matching the high-frequency refresh requirements of short-range radar data.
[0084] 3. Excellent Adaptability to Observational Task Value: The reward function uses the gradients of dynamic elements such as current field and wave height observed by radar as core weights, guiding the mobile platform to prioritize high-value observation areas such as areas with abrupt changes in ocean dynamic elements. The multi-source fusion process verifies the actual value of the observation area through feature matching and inversely optimizes the reward function weights. Compared to traditional single-element-oriented observation schemes, this significantly increases the sampling ratio of high-value areas, thus significantly enhancing the supporting role of observational data in marine scientific research.
[0085] 4. High platform operational safety: The system verifies operational risks in real time using quantitative conditions such as obstacle distance and signal-to-noise ratio thresholds. Once a replanning process is triggered, the emergency reward mechanism will proactively increase obstacle avoidance weight, forcibly guiding the platform to prioritize avoiding nearby obstacles such as fishing boats and ice floes. Simultaneously, radar interference warnings can be linked to dynamic path adjustments to avoid entering areas where observation has failed, thus preventing safety hazards. Compared to solutions without a replanning mechanism, this significantly reduces the risk of platform collisions and substantially improves operational stability under extreme sea conditions.
[0086] 5. High utilization rate of radar observation resources: The mobile platform's path information is fed back to the shore-based radar system in real time, enabling the radar beam to accurately point to the target observation area and reducing ineffective energy consumption during wide-area scanning; the data preprocessing stage uses the CFAR algorithm to remove sea clutter interference, increasing the proportion of effective echoes and avoiding the waste of radar resources on noise signal processing. Compared with non-coordinated radar observation modes, this significantly improves radar resource utilization and effectively reduces radar energy consumption per batch of observation missions.
[0087] 6. High accuracy of environmental state perception: During dynamic state modeling, error covariance and confidence weights are added to radar observation data to accurately quantify the reliability of the observation data; in the data assimilation stage, a three-dimensional variational algorithm is used to fuse radar observation data and platform sampling data, effectively correcting the inherent errors of the ocean forecasting model. Compared with traditional single-data-source state perception schemes, this significantly improves the accuracy of flow field inversion and significantly reduces the target location error.
[0088] 7. Excellent multi-source data collaboration and compatibility: A spatiotemporal calibration function unifies the timestamps and coordinate systems of radar wide-area data, platform sampling data, and assimilated forecast data, efficiently resolving the heterogeneity problem of multi-source data. The multi-branch neural network can simultaneously process grid-type radar data and vector-type platform location data, compatible with various data types. Compared to traditional single-data-driven models, this significantly expands the data compatibility dimension and substantially improves the fusion efficiency of multi-source information.
[0089] 8. Strong strategy iteration and self-optimization capability: The experience replay pool continuously stores path decision samples under sudden environmental changes, providing sufficient data support for DQN network parameter optimization; if there is a deviation between the fused inversion results and the state model, the reward function weights will be adjusted in reverse, driving continuous iteration of the path planning strategy. Compared with planning algorithms with fixed parameters, the system can continuously evolve as the observation mission progresses, gradually enhancing its ability to adapt to complex marine environments, and continuously improving the overall benefits of path planning.
[0090] 9. Broad cross-scenario task adaptability: For different application scenarios such as strong interference, dense targets, and emergency observation, the reward function can adaptively adjust the weights of each dimension (e.g., prioritizing the weight of observation effectiveness in strong interference scenarios); the lightweight replanning model can output differentiated path solutions such as obstacle avoidance and relocation based on the type of emergency. Compared to traditional single-scenario adaptable planning schemes, it can cover a variety of marine observation tasks such as nearshore monitoring, offshore emergency response, and target tracking, significantly expanding the scope of scenario adaptability.
[0091] To make the technical solutions and advantages of the embodiments of this application clearer, the exemplary embodiments of this application will be described in further detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not an exhaustive list of all embodiments. It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of this application can be combined with each other.
[0092] Example 1
[0093] Figure 1 The diagram illustrates the steps of a dynamic path planning method based on ground wave radar according to an embodiment of this application.
[0094] like Figure 1 As shown, a dynamic path planning method based on ground wave radar is provided, including the following steps:
[0095] S1: Acquire collaborative sensing data: acquire wide-area dynamic data of the observed sea area through shore-based high-frequency ground wave radar, and acquire short-range environmental information around the mobile platform through shipborne navigation radar.
[0096] S2: Constructing a dynamic state model: Preprocessing and fusing wide-area dynamic data and short-range environmental information to construct a dynamic state vector containing multi-dimensional environmental features and uncertainty representations;
[0097] S3: Generate initial observation path: Input the dynamic state vector into the reinforcement learning-based path planning model, optimize it through a multi-objective reward function, and output the initial heading and speed of the mobile platform to form the initial observation path.
[0098] Figure 2 The diagram illustrates the steps of another dynamic path planning method based on ground wave radar according to an embodiment of this application.
[0099] like Figure 2 As shown, in other preferred embodiments, after the initial observation path is formed in S3, the following steps are also included:
[0100] S4: Execution Path and Monitoring Environment: Controls the mobile platform to execute the initial observation path and continuously acquires real-time radar observation data and platform status data;
[0101] S5: Determine if replanning is triggered: Based on real-time radar observation data and platform status data, verify the preset hierarchical replanning trigger conditions in real time; the replanning trigger conditions include at least one of wide-area environmental change, near-range risk approach, and radar observation failure.
[0102] S6: Execute emergency replanning: If any replanning trigger condition is triggered, the lightweight reinforcement learning model is invoked to generate and distribute the emergency replanning path.
[0103] Figure 3 The diagram illustrates the steps of another dynamic path planning method based on ground wave radar according to an embodiment of this application.
[0104] like Figure 3 As shown, in other preferred embodiments, after forming the initial observation path, the following steps are also included:
[0105] S7: Collaborative Optimization and Closed-Loop Feedback: The mobile platform is controlled to execute emergency replanning of its path, and the path information of the mobile platform is fed back to the shore-based high-frequency ground wave radar to adjust its detection beam direction. At the same time, the radar observation data is input into the ocean numerical assimilation module to correct the ocean forecast model, and the corrected forecast data is fed back to generate a new path plan. This forms a closed-loop iteration of observation-planning-assimilation.
[0106] In specific implementation, the dynamic state vector in S2 is:
[0107] ;
[0108] in, The data is wide-area observation grid data for shore-based radar, with dimensions of 51×51×3, corresponding to the flow field, wave height and target density channels respectively.
[0109] The data is grid data for short-range observation by shipborne radar, with dimensions of 51×51×2, corresponding to obstacle distribution and signal-to-noise ratio channels, respectively.
[0110] This represents the electromagnetic interference intensity and frequency band vector. Provides real-time Cartesian coordinates for mobile platforms.
[0111] In practical implementation, the uncertainty characteristics in S2 include:
[0112] An error covariance matrix is introduced into the flow field inversion data. The error covariance matrix is a 51×51 diagonal matrix, and the diagonal elements are the variances of the flow field data of each grid.
[0113] Add a confidence weight ω to the target location data, calculated as follows:
[0114] ;
[0115] Where SNR is the current signal-to-noise ratio, and SNRmax is the system's maximum signal-to-noise ratio. The echo signal coherence coefficient, This is the historical error attenuation factor.
[0116] In another preferred embodiment, this application proposes a radar data confidence fusion method based on Bayesian inference. This method not only uses the signal-to-noise ratio but also introduces multi-dimensional indicators such as Doppler spectral width, echo coherence, and historical error distribution to construct a dynamic confidence model as follows:
[0117] ;
[0118] in, For Doppler spectral width.
[0119] This significantly enhances the dimensions of data reliability assessment compared to simple signal-to-noise ratio threshold judgment.
[0120] In specific implementation, in S3, the multi-objective reward function includes a composite reward function of radar observation value reward, obstacle avoidance reward, observation effectiveness reward, and platform energy consumption reward. The expression for the multi-objective reward function R is:
[0121] ;
[0122] Rewards for radar observation of ocean dynamic elements, and flow field gradients Wave height gradient Positive correlation;
[0123] The obstacle avoidance reward is dynamically assigned based on the distance between the path and the obstacle grid.
[0124] The reward for radar observation effectiveness is determined based on the signal-to-noise ratio (SNR).
[0125] It is a platform energy consumption reward, which is negatively correlated with the platform's actual power.
[0126] Where α, β, γ, and δ are adaptive weighting coefficients that are dynamically adjusted according to the environment type;
[0127] In optimal implementation, a reward weight adjustment network based on meta-learning is introduced to automatically learn the weights of each target according to the current environment type (e.g., strong interference, dense targets, emergency scenarios). For example, the calculation formula is as follows:
[0128] ;
[0129] in, It is a strong environmental disturbance factor; This is an environmental emergency scenario mode.
[0130] In S5, the replanning trigger conditions include at least one of the following: wide-area environmental abrupt change, near-range risk approach, and radar observation failure.
[0131] The triggering condition for wide-area environmental abrupt change is: when the offset distance of the center of the flow field gradient region monitored in real time by the shore-based radar relative to the historical benchmark exceeds the first threshold, or when the abrupt change value of the wave height data exceeds the second threshold.
[0132] The triggering conditions for short-range risk are: when the shipborne radar detects in real time that the distance between the moving obstacle and the platform is less than the third threshold, or when the short-range signal-to-noise ratio drops sharply to below the fourth threshold within a preset time.
[0133] The triggering condition for radar observation failure is: when the intensity of electromagnetic interference exceeds the system's anti-interference threshold and the radar echo signal is continuously lost.
[0134] In S7, the mobile platform's path information is fed back to the shore-based high-frequency ground wave radar to adjust its detection beam pointing. Specifically, the beam pointing angle is calculated using the following formula. :
[0135] ;
[0136] in,( , ) represents the location of the radar station. , () is the target area center of the platform.
[0137] Figure 4 The flowchart of a dynamic path planning and replanning method according to an embodiment of this application is shown.
[0138] like Figure 4 As shown, the process begins with "Start Observation Task".
[0139] First, collaborative radar data acquisition is performed, simultaneously acquiring shore-based wide-area and shipborne short-range data.
[0140] The process then proceeds to data preprocessing and dynamic state modeling, where the raw data is cleaned, fused, and dynamic state vectors with uncertainty representations are constructed.
[0141] Based on this, reinforcement learning is triggered to generate the initial path. The system calls the trained multi-branch neural network and multi-objective reward function to calculate the initial optimal path that takes into account observation value, safety and energy consumption.
[0142] Next, the mobile platform begins to execute the initial path. During this process, the system does not execute passively but continuously monitors in real time, including: the platform's own status (position, speed, attitude); changes in the wide-area environment (flow field, wave height) transmitted back by dual radars; the situation of near-range obstacles and radar signal-to-noise ratio (SNR); and the level of electromagnetic interference.
[0143] Next, the process enters the core judgment stage: Does the replanning condition meet? This judgment box contains three types of parallel triggering conditions:
[0144] Abrupt changes in the wide-area environment: for example, the center of the flow field gradient region shifts by more than 5 km, or the wave height increases dramatically by more than 1.5 m in a short period of time.
[0145] Near-range risks are imminent: for example, a newly appearing obstacle is less than 2km away from the platform, or the radar SNR drops from good to poor within 3 seconds.
[0146] Radar observation failure: For example, strong electromagnetic interference causes a continuous loss of radar echoes.
[0147] Once any condition is met, the process will switch to the "yes" branch.
[0148] If the determination is "yes", then the lightweight model is invoked for emergency replanning. The system will switch to a pruned and quantized lightweight reinforcement learning model, which quickly (e.g., within milliseconds) calculates a new emergency path based on the latest environmental state. The new path is constrained in terms of heading adjustment range and speed range to ensure platform maneuver safety.
[0149] Then, regardless of whether it has been replanned, the system will feed back the (new) path information to the shore-based radar to guide its beam to focus on the target area; at the same time, it will send the information of the interference zone sensed by the radar to the platform to assist in subsequent avoidance.
[0150] Next, ocean numerical assimilation and model updates are performed: the system inputs all observation data (radar + platform sampling) collected in this mission into the assimilation module to correct the ocean numerical prediction model and improve the accuracy of future environmental predictions.
[0151] Finally, the feedback loop: After step S7, the process is explicitly pointed to either step S2 or S3 by a dashed arrow. This indicates that the updated and more accurate ocean forecast field will be injected as new background knowledge into the next round of state modeling or path planning. This mechanism enables the system's environmental awareness and decision-making capabilities to continuously iterate and optimize as the task is executed, forming a powerful self-learning and adaptive closed loop.
[0152] Ultimately, the process leads to "task completion or continuous observation". For a single task, the process ends upon reaching the endpoint; for long-term observation tasks, this process will continue to run in a loop.
[0153] In summary, the dynamic path planning method based on ground-wave radar proposed in this application includes acquiring wide-area dynamic data of the observation sea area through shore-based high-frequency ground-wave radar and acquiring near-range environmental information around the mobile platform through shipborne navigation radar. The wide-area dynamic data and near-range environmental information are preprocessed and fused to construct a dynamic state vector containing multi-dimensional environmental features and uncertainty representations. This dynamic state vector is then input into a reinforcement learning-based path planning model, and through multi-objective reward function optimization, the initial heading and speed of the mobile platform are output, forming the initial observation path. This application, by constructing a shore-based and shipborne dual-layer radar collaborative sensing network, realizes radar-mobile platform collaborative dynamic marine observation path planning, providing technical support for efficient observation in dynamic marine environments.
[0154] Meanwhile, this application combines reinforcement learning with a lightweight emergency replanning mechanism and introduces ocean numerical assimilation to form a closed-loop optimization, ultimately achieving efficient, safe and adaptive planning of observation paths in dynamic ocean environments.
[0155] Example 2
[0156] This embodiment provides a dynamic path planning system based on ground wave radar. For details not disclosed in the dynamic path planning system based on ground wave radar in this embodiment, please refer to the specific implementation of the dynamic path planning scheme based on ground wave radar in other embodiments.
[0157] Figure 5 The diagram shows a schematic representation of a dynamic path planning system based on ground wave radar according to an embodiment of this application.
[0158] like Figure 5 As shown, the dynamic path planning system based on ground wave radar includes:
[0159] The collaborative sensing data acquisition module 10 is used to acquire wide-area dynamic data of the observation sea area through shore-based high-frequency ground wave radar and acquire short-range environmental information around the mobile platform through shipborne navigation radar.
[0160] The dynamic state model construction module 20 is used to preprocess and fuse wide-area dynamic data and short-range environmental information to construct a dynamic state vector containing multi-dimensional environmental features and uncertainty representations.
[0161] The initial observation path generation module 30 is used to input the dynamic state vector into the path planning model based on reinforcement learning, optimize it through a multi-objective reward function, and output the initial heading and speed of the mobile platform to form the initial observation path.
[0162] This application also provides another radar-mobile platform collaborative ocean dynamic observation path planning system, which mainly consists of a shore-based high-frequency ground wave radar 1, a shipborne navigation radar 2, a radar data preprocessing module 3, a dynamic state modeling module 4, a reinforcement learning path planning module 5, a replanning module, a radar-mobile platform collaborative interface, and an ocean numerical assimilation module.
[0163] Figure 6 The diagram shows a system deployment principle according to an embodiment of this application.
[0164] Figure 7 The diagram shows the system module composition and data flow according to an embodiment of this application.
[0165] The system adopts a layered design, forming a closed-loop workflow of perception → processing → decision-making → execution → feedback.
[0166] like Figure 6 , Figure 7 As shown, firstly, the shore-based high-frequency ground wave radar 1 and the shipborne navigation radar 2 are jointly deployed in the target observation sea area. The shore-based high-frequency ground wave radar 1 is responsible for collecting wide-area dynamic data such as ocean surface current field, wave height, and distribution of sea targets.
[0167] The shipborne navigation radar 2 focuses on short-range environmental information such as the location of moving obstacles and the signal-to-noise ratio of radar echoes within the area surrounding the platform.
[0168] The radar data preprocessing module 3 is installed in both the shore-based radar station and the mobile platform. Due to the complex conditions of the marine observation environment, such as strong sea clutter and electromagnetic interference, it needs to be designed as a dedicated processing unit with clutter suppression, spatiotemporal registration and multi-source data fusion capabilities. In addition, the shipborne module needs to meet the requirements of waterproof and anti-turbulence protection for marine operations.
[0169] Radar data preprocessing module 3 performs multi-stage processing on the acquired raw radar echo data, employing anti-interference and high-precision processing algorithms to ensure that the data can directly support path planning decisions. For data from shore-based high-frequency ground wave radar 1, the module uses existing technology, constant false alarm rate (CFAR), to remove sea clutter and electromagnetic interference.
[0170] Its detection and judgment formula is:
[0171] ;
[0172] in Radar echo signal amplitude, An adaptive threshold;
[0173] Subsequently, core elements were extracted using an ocean current field inversion algorithm. The formula for inverting the radial velocity of the surface flow field is as follows:
[0174] ;
[0175] In the formula, For radar operating wavelength, For Doppler frequency shift, This refers to the radar carrier frequency.
[0176] For data from shipborne navigation radar 2, the module uses a polar coordinate-rectangular coordinate transformation algorithm to convert the original polar coordinate data (azimuth angle) into rectangular coordinate data. ,distance Convert the data to Cartesian coordinate grid using the following formula:
[0177] ;
[0178] Simultaneously, signal-to-noise ratio (SNR) analysis technology is used to screen for high-risk obstacle areas at close range. The SNR calculation formula is:
[0179] ;
[0180] in, For effective target echo power, For noise power, when The area was identified as a low signal-to-noise region and marked as a potential observation failure zone.
[0181] Next, by constructing a dynamic state modeling module 4, preprocessed radar data, electromagnetic interference monitoring data, and real-time location information of the mobile platform are integrated to construct a multi-dimensional dynamic state vector. Its expression is:
[0182] ;
[0183] In the formula, For shore-based radar wide-area observation grid data (dimension: The three channels correspond to flow field, wave height, and target density, respectively. This data is collected by the shore-based radar through the ocean area within the scanning radius, with a grid resolution of 100m, and then filtered and denoised. The flow field channel records the horizontal velocity and direction components of the seawater in each grid, the wave height channel reflects the average effective wave height in the area, and the target density channel counts the number of targets such as ships and buoys in the grid.
[0184] The data is for short-range observation grid data of shipborne radar (dimension 51×51×2, with two channels corresponding to obstacle distribution and signal-to-noise ratio, respectively). The shipborne radar covers an area with a radius of 500m centered on the mobile platform, with a grid resolution of 100m. The obstacle distribution channel marks the presence status of obstacles such as reefs and small floating objects in the grid (0 for none, 1 for present). The signal-to-noise ratio channel represents the power ratio of the radar echo signal to the background noise in the grid.
[0185] This is a vector of electromagnetic interference intensity and frequency band (dimension 1×2), where the first element is the equivalent power intensity of electromagnetic interference in the current environment (in dBm), and the second element is the frequency band range where the interference signal is concentrated (in MHz). The real-time Cartesian coordinates (dimension 1×2) of the mobile platform were recorded with the pre-set ocean reference origin as the reference, and the horizontal and vertical coordinates of the platform were recorded (unit: m).
[0186] Simultaneously, uncertainty characterization is added to radar observation data, and an error covariance matrix is introduced into the flow field inversion data. The matrix is a 51×51 diagonal matrix, and the elements on the diagonal correspond to the variance of each flow field grid data. The variance value is determined by the flow velocity measurement accuracy of the shore-based radar and the error propagation model of the data fusion algorithm.
[0187] Add a confidence weight ω to the target location data. The confidence score is calculated as follows:
[0188] ;
[0189] in, The maximum measurable signal-to-noise ratio of the radar system (determined by the radar hardware performance, usually taken as 40dB) is the maximum measurable signal-to-noise ratio of the radar system. When the actual signal-to-noise ratio within the grid is higher, the confidence weight of the corresponding target position data is closer to 1. Conversely, the weight is reduced to weaken the impact of low-quality data.
[0190] Through the integration and uncertainty characterization of the above multi-source data, the dynamic state modeling module 4 can accurately depict the complete scene of the dynamic marine environment, which includes not only the wide-area marine hydrology and target distribution characteristics, but also the short-range obstacles and electromagnetic environment around the mobile platform. At the same time, it distinguishes data quality through confidence weights, providing highly reliable and multi-dimensional environmental inputs for the subsequent path planning of the mobile platform.
[0191] Reference Figure 6 , Figure 7 Next, the reinforcement learning path planning module 5 receives the dynamic state vector. The core features are extracted using a multi-branch neural network. This network includes a wide-area radar data branch, a short-range radar data branch, and a fully connected branch for platform location.
[0192] The wide-area radar data branch employs a spatial convolutional layer with a kernel size of 8×4×4, used for processing... The flow field, wave height, and target density channels are used to extract features. After convolution, batch normalization and ReLU activation function are used to enhance the nonlinear expression of the features.
[0193] The short-range radar data branch uses a spatial convolutional layer with a kernel size of 8×3×3, and is matched with... The obstacle distribution and signal-to-noise ratio channels are compressed into spatial dimensions after convolution through a max-pooling layer.
[0194] The platform location fully connected branch will The real-time coordinates are input to two fully connected layers and mapped to feature vectors with the same dimension as the radar branches. The feature vectors output from each branch are concatenated and fused before being input to a decision layer consisting of three fully connected layers, completing the mapping from features to the action space. Simultaneously, a multi-target reward function is constructed. Its expression is:
[0195] ;
[0196] In the formula, The reward for radar-observed ocean dynamic elements is calculated as follows: ,in The flow field gradient (characterizing the spatial rate of change of flow velocity between adjacent grid cells) is the flow field gradient. Wave height gradient (reflecting the degree of regional variation in wave height). , These are the weighting coefficients;
[0197] Obstacle avoidance reward: When the planned path is... When the distance between the marked obstacle grid cells is greater than or equal to 20, the obstacle is considered to have been successfully avoided. = 10; When the path overlaps with the obstacle grid, a collision is determined, and at this time... =-50;
[0198] Rewards for radar observation effectiveness: based on The signal-to-noise ratio (SNR) is determined when the SNR is... At 5 o'clock, the radar echo data is reliable. = 8; when SNR At 3 o'clock, the echo is severely affected by noise, resulting in low data reliability. = -20;
[0199] The platform's energy consumption reward is calculated using the following formula: , This represents the platform's actual power (positively correlated with the current speed). This represents the platform's maximum power (corresponding to energy consumption at maximum speed).
[0200] α, β, γ, and δ are weighting coefficients that satisfy α + β + γ + δ = 1: γ = 0.4 is used in strong interference scenarios to enhance the weight of observation effectiveness; α = 0.4 is used in dense target scenarios to prioritize the adaptation to the influence of the marine dynamic environment.
[0201] DQN algorithm generates initial path
[0202] Figure 8 The diagram illustrates the principle of generating an initial path using the DQN algorithm according to an embodiment of this application.
[0203] like Figure 8 As shown, the module uses the Discrete Quantization (DQN) algorithm in the discrete action space to generate the initial path, and its objective function is:
[0204] ;
[0205] in, ;
[0206] These are the current Q-network parameters. The target Q-network parameters are updated synchronously with the current Q-network parameters every 100 training steps. The experience replay pool (capacity set to 10,000, storing past state-action-reward-next state samples, randomly sampled during training to reduce data correlation). This is the discount factor. Ultimately, it passes... The action output by the network is valued, and the action with the highest value is selected as the course for the mobile platform. With speed This forms an initial optimal observation path that adapts to the marine environment and platform constraints.
[0207] Next, the mobile observation platform executes the observation task according to the initial path command. The platform control system adjusts the thrusters and rudder according to the path parameters. Its kinematic model is as follows:
[0208] ;
[0209] in This is the angular velocity of the heading.
[0210] At the same time, the platform will display its real-time location (x, y) and remaining battery power. Data such as sensor operating status are transmitted back to the system via the shipborne communication module. The shore-based and shipborne radars maintain continuous detection and periodically synchronize the latest environmental data to the system, ensuring that the system can simultaneously grasp the platform's operating status and changes in the marine environment.
[0211] Next, the system processes the returned radar observation data. Millisecond-level real-time threshold verification is performed against mobile platform status data. Three hierarchical replanning trigger conditions are constructed. If any condition is met, the policy replanning process is initiated. If no condition is triggered, the original path is maintained for real-time tracking and fine-tuning.
[0212] (1) Wide-area environmental mutation triggering:
[0213] The system continuously compares real-time data collected by shore-based radar with historical baseline data. When a shift in the flow field gradient region is detected, the system calculates the shift distance. :
[0214] ;
[0215] in,( , The coordinates of the center of the original flow field gradient region are obtained by fitting the grid set with the most significant flow field changes from historical data; , () represents the center coordinates of the current flow field gradient region.
[0216] If Δd > 5km, it is determined that a large-scale change has occurred in the flow field environment, which may cause the speed matching of the original path to fail. At the same time, the change value Δh of the wave height data (the difference between the current average wave height and the average wave height in the past 10 minutes) is monitored. If Δh > 1.5m, it is determined that the wave height environment exceeds the wave resistance threshold of the original path, and wide-area environment replanning is triggered.
[0217] (2) Short-range risk triggering:
[0218] Shipborne radar scans at a frequency of 10 Hz The coverage area is short-range, and when moving obstacles (such as small fishing boats or floating objects) are detected, the distance is considered. Calculation formula:
[0219] ;
[0220] in,( , The real-time grid coordinates of the obstacle are used for target identification algorithms based on radar echoes; , () represents the real-time rectangular coordinates of the mobile platform.
[0221] The distance between the obstacle and the platform is calculated. If the distance dobs is less than 2km, the obstacle is determined to have entered the platform's short-range warning range, and the obstacle avoidance margin of the original path is insufficient. At the same time, the signal-to-noise ratio (SNR) of the radar echo is monitored in real time. If the SNR drops sharply to below 3dB within 3 seconds, the reliability of the short-range observation data is determined to be lost, and it cannot support obstacle avoidance of the original path. At this time, short-range risk-based replanning is triggered.
[0222] (3) Radar observation failure trigger:
[0223] The system continuously collects data. When the electromagnetic interference data is detected to exceed the anti-interference threshold of the radar system and the radar echo signal is completely lost, it is determined that the radar observation function has temporarily failed and the environmental data dependent on the original path cannot be updated. At this time, a radar observation failure replanning is triggered. If the radar echo returns to normal after the interference is removed, the replanning process is terminated and the execution of the original path is restored.
[0224] For the triggered replanning scenario, the system prioritizes calling the lightweight DQN model to perform emergency path calculation: This model performs multi-dimensional pruning optimization on the basis of the original DQN network, which not only reduces the number of convolutional layers, but also greatly reduces the number of parameters through weight quantization (and channel pruning). At the same time, it adopts a hardware-accelerated inference strategy to ensure the rapid completion of emergency path generation and verification, meeting the low latency requirements of replanning scenario.
[0225] The reward function for replanning the path adds an emergency priority weighting term to the original multi-objective reward function. The emergency reward function is as follows:
[0226] ;
[0227] in, The emergency priority reward is dynamically assigned based on the risk level of the triggering scenario: in obstacle avoidance emergency scenarios, the distance between the obstacle and the platform is relatively short and the risk of collision is high; in observation failure emergency scenarios, the environmental perception capability is limited and the path uncertainty is increased; η is the emergency weight, which is used to balance the influence ratio of the original reward function and the emergency priority.
[0228] To accommodate the maneuverability constraints of the mobile platform, the action output of the replanning path is limited to a safe range: heading adjustment. (To avoid excessive turning that could cause platform instability), speed range (Taking into account both mobility and energy efficiency).
[0229] The generated new path will be accompanied by a path confidence indicator (calculated from the reliability of the current environmental data) and sent to the navigation control system of the mobile platform in real time via a low-latency communication link. After receiving the new path, the platform will complete a smooth switch between the original path and the new path to ensure the stability of the navigation process.
[0230] The embodiments of this application use a radar-mobile platform collaborative interface to achieve bidirectional communication and linkage.
[0231] On one hand, the mobile platform feeds back the new path information to the shore-based radar, which then adjusts the beam pointing angle according to the area the platform is about to enter. The beam pointing adjustment formula is:
[0232] ;
[0233] in,( , ) represents the location of the radar station. , The mobile platform uses the target area center as its objective, thereby improving the observation accuracy of the target area. On the other hand, the shore-based radar pushes real-time electromagnetic interference early warning information to the mobile platform, assisting it in avoiding areas where observations fail. Simultaneously, the system inputs the radar observation data from this round into the ocean numerical assimilation module, employing a three-dimensional variational assimilation algorithm to correct the ocean numerical prediction model. Its objective function is:
[0234] ;
[0235] In the formula, For the analysis field, As background scene, The background error covariance matrix, For radar observation data, For the observation operator, The observation error covariance matrix is used. The assimilated dynamic forecast field serves as a new state input, feeding back into reinforcement learning path planning module 5, forming a closed-loop iteration of "observation-planning-correction".
[0236] Finally, as the core data processing node of the marine observation system, the shore-based base station will continuously receive observation data from shore-based radar, shipborne radar, real-time data transmitted back from mobile platforms, and historical forecast data assimilated by the marine forecasting system. These data constitute the basic input for multi-source data fusion.
[0237] (1) Time synchronization and spatial alignment: First, time synchronization is performed, assuming... This is a radar observation dataset, where each data item... All include the original data collection timestamp. ; For the platform, sample the dataset, each data item Includes original sampling timestamps .
[0238] Using time mapping function The original timestamp of the radar data Calibration The platform's raw timestamps Calibration This ensures that the timestamp error after calibration is controlled within 1 second.
[0239] After time synchronization is completed, spatial alignment is performed. The original location of the radar data. It is a rectangular coordinate system (unit: meters) based on the ocean reference origin, and the two coordinate systems are completely different.
[0240] At this point, construct the space mapping function. First, select at least 3 known common control points of "radar grid-rectangular coordinates". Calculate the rotation, translation, and scaling parameters of the coordinate transformation through affine transformation. Then, apply these parameters to the grid identifier of all radar data to convert it into rectangular coordinates consistent with the platform data. At the same time, the rectangular coordinates of the platform data will be "rasterized" based on the resolution of the radar grid to ensure that each platform data point corresponds to a unique radar grid.
[0241] Ultimately, the raw location of the radar data Calibrated as The original location of platform data Calibrated as This enables precise matching of spatial locations.
[0242] (2) Feature extraction and matching: After time synchronization and spatial alignment are completed, multi-source data enter the feature extraction and matching stage.
[0243] The core objective of this stage is to extract relevant features from different types of data and establish mapping relationships between features to provide a logical link for subsequent fusion.
[0244] For radar observation datasets Its core information is the spatial distribution of the marine dynamic environment, therefore, a feature set needs to be extracted. It includes three dimensions: first, the flow field intensity, which is the modulus of the seawater flow velocity in each grid, reflecting the hydrodynamics of the area; second, the wave height amplitude, which is the maximum effective wave height in each grid, reflecting the degree of sea surface fluctuation; and third, the target density, which is the number of targets such as ships and buoys in each grid, reflecting the navigation status of the area.
[0245] These features are calculated by performing spectral analysis and threshold segmentation on the raw radar echo signal, and are ultimately represented as a 5-dimensional vector for each radar data item.
[0246] For platform sampling dataset Its core information is the real-time status of the marine hydrological environment, therefore, a feature set needs to be extracted. It includes four dimensions: first, water temperature, i.e., the temperature of the seawater at the sampling point; second, salinity, i.e., the salinity of the seawater at the sampling point; third, water pressure, i.e., the pressure of the seawater at the sampling point, reflecting the sampling depth; and fourth, turbidity, i.e., the transparency of the seawater at the sampling point, reflecting the content of suspended particles in the water.
[0247] These features are directly collected by the sensors on the platform and processed by filtering and noise reduction, ultimately representing the features of each platform data item in the form of a 4-dimensional vector.
[0248] Next, feature matching is performed, using cosine similarity as a metric. The core logic is to calculate the cosine value of the angle between two feature vectors in high-dimensional space. The closer the value is to 1, the more similar the features are.
[0249] For any eigenvector in the radar feature set , and any feature vector in the platform feature set The similarity calculation formula is:
[0250] ;
[0251] The numerator is the dot product of the two vectors, and the denominator is the product of the magnitudes of the two vectors. To ensure the reliability of the matching, a matching threshold is set. When the calculated similarity When it is determined that the radar characteristics and platform characteristics belong to the observation results under the same marine environmental conditions, the corresponding radar data items are... With platform data items The association is a fused data record; when the similarity is less than 0.8, it is marked as "data to be verified" and will be further verified in conjunction with the forecast data.
[0252] In practice, feature matching does not involve traversing all feature vectors one by one. Instead, it combines the completed time synchronization and spatial alignment results: first, radar and platform data items with consistent timestamps are selected, then items with the same spatial location are selected, and finally, the similarity of the selected feature vectors is calculated. After feature matching, the originally independent radar and platform datasets are transformed into a fused dataset containing the "radar observation - platform sampling" correlation. Each data item simultaneously possesses characteristics of both the dynamic environment and the hydrological environment, providing multi-dimensional input for subsequent marine environmental state inversion.
[0253] (3) Fusion Inversion: After feature matching is completed, the fusion inversion stage is entered. This stage adopts a hybrid algorithm combining Kalman filtering and deep learning. It utilizes both the time-series prediction capability of Kalman filtering and the nonlinear fitting capability of deep learning to finally output a high-precision marine environmental status.
[0254] Kalman filtering is a recursive estimation algorithm based on the minimum mean square error criterion, suitable for state estimation of dynamic systems. Its core principle is to continuously refine the state estimate through an iterative "prediction-update" process. First, the state vector of the marine environment is defined. It includes eight dimensions of state variables, such as flow velocity, wave height, water temperature, and salinity, as... State estimate at time 1.
[0255] The first step in Kalman filtering is state prediction: using the state estimate from the previous time step. Through the state transition matrix Predict the current state and the state transition matrix. It is based on the ocean dynamics model and reflects the changes of state variables such as flow field and wave height over time. For example, the changes in flow field are affected by Coriolis force and pressure gradient force. These relationships are quantified into elements in the matrix.
[0256] Meanwhile, the mobile platform's control inputs, such as speed and heading, are also... It will be through the control matrix Effects on state prediction, such as the flow of surrounding water caused by the platform's navigation, will be incorporated into the prediction process.
[0257] Therefore, the expression for state prediction is: .
[0258] The second step is covariance prediction: covariance matrix The covariance matrix at the previous time step reflects the uncertainty of the state estimate. Through the state transition matrix Transform by the transpose, and simultaneously add the process noise covariance matrix. The process noise mainly comes from random fluctuations in the marine environment, and its matrix elements are obtained based on variance statistics of historical data. Therefore, the expression for covariance prediction is: .
[0259] The third step is to calculate the Kalman gain: Kalman gain It is a core parameter for weighing the predicted and observed values, and its calculation requires the use of the covariance matrix. Observation matrix Covariance matrix of observation noise Observation matrix Used to map state vectors to observation vectors, observation noise covariance matrix It reflects the error level of the observed data; therefore, the expression for the Kalman gain is: .
[0260] The fourth step is state update: using the observation data at the current moment. (i.e., the fused dataset after feature matching), combined with Kalman gain The previous state predictions were corrected to obtain more accurate state estimates, expressed as follows: ,in It is the observation residual, which reflects the difference between the predicted value and the actual observation.
[0261] The fifth step is covariance update: The covariance matrix is corrected using Kalman gain to obtain the covariance matrix at the current time step, expressed as follows: ,in It is an identity matrix.
[0262] After iterative processing using Kalman filtering, the obtained state estimate is... While it already possesses high accuracy, the linear assumption of Kalman filtering introduces certain errors due to the complex nonlinear system of the marine environment. Therefore, it is necessary to introduce a deep learning model for further optimization.
[0263] The multilayer perceptron (MLP) model is used here. It is a deep learning model with a simple structure but strong fitting ability. It consists of an input layer, a hidden layer and an output layer. The dimension of the input layer is the same as the dimension of the state vector of the Kalman filter. The hidden layer is set to 3 layers, each containing 64 neurons. The dimension of the output layer is the same as the dimension of the state vector and is used to output the optimized state estimate.
[0264] The input to the MLP model is the preliminary state estimate from the Kalman filter. Compared with the original observation data The concatenated vectors, the model expression is as follows: ,in It is the set of parameters of the model, including the weights and biases between neurons in each layer.
[0265] The model is trained using supervised learning, with a training dataset... middle, It is the concatenated vector of the Kalman filter output and the original observation. It is based on real-world marine environmental conditions obtained from on-site measurements. The training objective is to minimize the loss function. ,in This is the mean squared error loss function, which calculates the mean squared difference between the model output and the true label. During training, the stochastic gradient descent algorithm is used to update the model parameters. .
[0266] Example 3
[0267] This embodiment provides a dynamic path planning device based on ground wave radar. For details not disclosed in the dynamic path planning device based on ground wave radar in this embodiment, please refer to the specific implementation of the dynamic path planning method or system based on ground wave radar in other embodiments.
[0268] Figure 9 The diagram shows a structural schematic of a ground wave radar-based dynamic path planning device 400 according to an embodiment of this application.
[0269] like Figure 9 As shown, the ground wave radar-based dynamic path planning device 400 includes: a storage unit 402 for storing executable instructions; and a processing unit 401 for connecting to the storage unit 402 to execute the executable instructions to complete the ground wave radar-based dynamic path planning method.
[0270] Those skilled in the art will understand that the illustration Figure 9 This is merely an example of a ground wave radar-based dynamic path planning device 400 and does not constitute a limitation on the ground wave radar-based dynamic path planning device 400. It may include more or fewer components than shown, or combine certain components, or different components. For example, the ground wave radar-based dynamic path planning device 400 may also include input / output devices, network access devices, buses, etc.
[0271] The processing unit 401 (Central Processing Unit, CPU) can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor, or the processing unit 401 can be any conventional processor. The processing unit 401 is the control center of the ground wave radar-based dynamic path planning device 400, connecting all parts of the device using various interfaces and lines.
[0272] Storage unit 402 can be used to store computer-readable instructions. Processing unit 401 implements various functions of the ground wave radar-based dynamic path planning device 400 by running or executing the computer-readable instructions or modules stored in storage unit 402 and calling the data stored in storage unit 402. Storage unit 402 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the ground wave radar-based dynamic path planning device 400, etc. In addition, storage unit 402 may include hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, read-only memory (ROM), random access memory (RAM), or other non-volatile / volatile storage devices.
[0273] If the integrated module of the ground wave radar-based dynamic path planning device 400 is implemented as a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can also be implemented by instructing related hardware through computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium, and when executed by a processor, they can implement the steps of the various method embodiments described above.
[0274] Example 5
[0275] This embodiment provides a computer-readable storage medium on which a computer program is stored; the computer program is executed by a processor to implement the dynamic path planning method based on ground wave radar in other embodiments.
[0276] This application utilizes a ground-wave radar-based dynamic path planning device and storage medium. Wide-area dynamic data of the observation sea area is acquired via a shore-based high-frequency ground-wave radar, while near-range environmental information surrounding the mobile platform is acquired via a shipborne navigation radar. The wide-area dynamic data and near-range environmental information are preprocessed and fused to construct a dynamic state vector containing multi-dimensional environmental features and uncertainty representations. This dynamic state vector is then input into a reinforcement learning-based path planning model. Through multi-objective reward function optimization, the initial heading and speed of the mobile platform are output, forming the initial observation path. This application, by constructing a shore-based and shipborne dual-layer radar collaborative sensing network, realizes radar-mobile platform collaborative dynamic marine observation path planning, providing technical support for efficient observation in dynamic marine environments.
[0277] Meanwhile, this application combines reinforcement learning with a lightweight emergency replanning mechanism and introduces ocean numerical assimilation to form a closed-loop optimization, ultimately achieving efficient, safe and adaptive planning of observation paths in dynamic ocean environments.
[0278] Those skilled in the art will understand that the terminology used in this invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The singular forms “a,” “the,” and “the” as used in this invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.
[0279] It should be understood that although the terms first, second, third, etc., may be used in this invention to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first information may also be referred to as second information without departing from the scope of this invention, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."
[0280] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.
[0281] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A dynamic path planning method based on ground wave radar, characterized in that, Includes the following steps: Wide-area dynamic data of the observation sea area are obtained through shore-based high-frequency ground wave radar, and short-range environmental information around the mobile platform is obtained through shipborne navigation radar. Based on the wide-area dynamic data and near-range environmental information, preprocessing and fusion are performed to construct a dynamic state vector containing multi-dimensional environmental features and uncertainty representations; The dynamic state vector is input into a path planning model based on reinforcement learning. Through multi-objective reward function optimization, the initial heading and speed of the mobile platform are output to form the initial observation path. After forming the initial observation path, the process also includes: The mobile platform is controlled to execute the initial observation path and continuously acquire real-time radar observation data and platform status data; Based on the real-time radar observation data and platform status data, the preset hierarchical replanning trigger conditions are verified in real time. If any of the aforementioned replanning trigger conditions are triggered, the lightweight reinforcement learning model is invoked, an emergency priority weighting term is introduced into the multi-objective reward function, and an emergency replanning path is generated and issued. After generating and issuing the emergency rerouting path, the process also includes: The mobile platform is controlled to execute the emergency replanning path, and the path information of the mobile platform is fed back to the shore-based high-frequency ground wave radar to adjust its detection beam direction; Simultaneously, the radar observation data is input into the ocean numerical assimilation module to correct the ocean forecast model, and the corrected forecast data is fed back to generate a new path plan.
2. The dynamic path planning method based on ground wave radar according to claim 1, characterized in that, The dynamic state vector is: ; in, The data is wide-area observation grid data for shore-based radar, with dimensions of 51×51×3, corresponding to the flow field, wave height and target density channels respectively. The data is grid data for short-range observation by shipborne radar, with dimensions of 51×51×2, corresponding to obstacle distribution and signal-to-noise ratio channels, respectively. This represents the electromagnetic interference intensity and frequency band vector. Provides real-time Cartesian coordinates for mobile platforms.
3. The dynamic path planning method based on ground wave radar according to claim 1, characterized in that, The multi-objective reward function includes a composite reward function of radar observation value reward, obstacle avoidance reward, observation effectiveness reward, and platform energy consumption reward. The expression of the multi-objective reward function is as follows: ; Where α, β, γ, and δ are adaptive weighting coefficients that are dynamically adjusted according to the environment type; The reward for radar observation of ocean dynamic elements is positively correlated with the flow field gradient and wave height gradient; The obstacle avoidance reward is dynamically assigned based on the distance between the path and the obstacle grid. The reward for radar observation effectiveness is determined based on the signal-to-noise ratio (SNR). It is a platform energy consumption reward, which is negatively correlated with the platform's actual power.
4. The dynamic path planning method based on ground wave radar according to claim 1, characterized in that, The replanning triggering conditions include at least one of wide-area environmental abrupt change, near-range risk approach, and radar observation failure. The triggering condition for the wide-area environmental change is: when the offset distance of the center of the flow field gradient region monitored in real time by the shore-based radar relative to the historical benchmark exceeds the first threshold, or when the change value of the wave height data exceeds the second threshold. The triggering condition for the near-range risk is: when the shipborne radar detects in real time that the distance between the moving obstacle and the platform is less than the third threshold, or when the near-range signal-to-noise ratio drops sharply to below the fourth threshold within a preset time. The triggering condition for radar observation failure is: when the electromagnetic interference intensity exceeds the system's anti-interference threshold and the radar echo signal is continuously lost.
5. The dynamic path planning method based on ground wave radar according to claim 1, characterized in that, The path information of the mobile platform is fed back to the shore-based high-frequency ground wave radar to adjust its detection beam pointing. Specifically, the beam pointing angle is calculated using the following formula. : ; in, Location of the radar site. The target area center of the platform.
6. A dynamic path planning system based on ground wave radar, characterized in that, include: The collaborative sensing data acquisition module is used to acquire wide-area dynamic data of the observation sea area through shore-based high-frequency ground wave radar and to acquire short-range environmental information around the mobile platform through shipborne navigation radar. The dynamic state model construction module is used to preprocess and fuse the wide-area dynamic data and near-range environmental information to construct a dynamic state vector containing multi-dimensional environmental features and uncertainty representations. The initial observation path generation module is used to input the dynamic state vector into the path planning model based on reinforcement learning, optimize it through a multi-objective reward function, and output the initial heading and speed of the mobile platform to form the initial observation path. After forming the initial observation path, the process also includes: The mobile platform is controlled to execute the initial observation path and continuously acquire real-time radar observation data and platform status data; Based on the real-time radar observation data and platform status data, the preset hierarchical replanning trigger conditions are verified in real time. If any of the aforementioned replanning trigger conditions are triggered, the lightweight reinforcement learning model is invoked, an emergency priority weighting term is introduced into the multi-objective reward function, and an emergency replanning path is generated and issued. After generating and issuing the emergency rerouting path, the process also includes: The mobile platform is controlled to execute the emergency replanning path, and the path information of the mobile platform is fed back to the shore-based high-frequency ground wave radar to adjust its detection beam direction; Simultaneously, the radar observation data is input into the ocean numerical assimilation module to correct the ocean forecast model, and the corrected forecast data is fed back to generate a new path plan.
7. A dynamic path planning device based on ground wave radar, characterized in that, include: Storage unit, used to store executable instructions; as well as A processing unit is configured to be connected to a memory to execute executable instructions to perform the method as described in any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, It stores a computer program thereon; the computer program is executed by a processor to implement the method as described in any one of claims 1-5.