Large-area sea area Beidou base station control network adaptive layout method

By adaptively adjusting the operating strategy of the Beidou base station control network through deep reinforcement learning intelligent agents, the problem of adaptive deployment in the marine environment was solved, efficient and stable base station layout optimization was achieved, and the learning efficiency and data utilization of the model were improved.

CN120630729AActive Publication Date: 2025-09-12THE 2ND ENG CO LTD MBEC +2
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511113984.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-09-12
Estimated Expiration
2045-08-11

AI Technical Summary

Technical Problem

Existing technologies in Beidou base station control networks over large sea areas have problems such as poor adaptability to dynamic environments, limited model generalization capabilities, and low efficiency in multi-objective optimization, making it difficult to achieve adaptive deployment.

Method used

Through the deep reinforcement learning intelligent agent, the operating strategy of the control network is adaptively adjusted. The state evaluation neural network, deep feature extraction network and deep Q network are used to generate adaptive system configuration and adjustment strategies. Combined with the composite reward function and unsupervised cluster analysis, the automatic classification and identification of the current operating mode of the system is achieved.

Benefits of technology

It achieves adaptive deployment in complex sea environments, improves learning efficiency and data utilization, ensures stable convergence of the model and high-performance decision-making, can quickly respond to environmental changes, and provide high-quality base station layout solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120630729A_ABST
    Figure CN120630729A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of machine learning, in particular to a large-area sea area Beidou base station control network adaptive layout method, which deeply integrates machine learning and reinforcement learning technologies and comprises the following steps: firstly, evaluating system stability by utilizing supervised learning, and identifying a control network operation mode through a variational auto-encoder and unsupervised clustering; a deep reinforcement learning decision-making agent is constructed, and the agent takes a system state and environment characteristics as input and is trained under the guidance of a composite reward function of adaptive risk perception. The deep Q network structure DQN is adopted to construct the decision neural network, the fitting state-action value function is iteratively optimized, the intelligent agent can learn and generate the optimal node layout, topology reconstruction and parameter adjustment instruction, and efficient, autonomous and intelligent operation of the control network in the complex marine environment is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of machine learning technology, and in particular to a method for adaptively deploying a Beidou base station control network in a large sea area. Background Art

[0002] In the deployment of Beidou base station control networks in maritime areas, existing neural network-based machine learning methods face several key technical bottlenecks. Regarding adaptability to dynamic environments, traditional static training paradigms struggle to meet real-time response requirements, while mainstream time series modeling methods struggle to strike a balance between computational efficiency and long-term dependencies. Small-sample learning scenarios face challenges such as data scarcity and feature nonlinearity, limiting model generalization. Multi-objective optimization faces issues such as fixed objective weights and inefficient optimization in high-dimensional solution spaces. These technical bottlenecks collectively reflect the adaptability challenges faced by neural network models in complex maritime environments.

[0003] To this end, innovative breakthroughs are needed in algorithm architecture and training paradigm to achieve adaptive deployment of the control network. Summary of the Invention

[0004] This invention aims to address the poor survivability and low intelligence of Beidou base station control networks in large ocean areas in harsh environments. By using a deep reinforcement learning agent to adaptively adjust the control network's operating strategy, this system builds an intelligent, efficient, and highly resilient ocean positioning solution.

[0005] A method for adaptively deploying a Beidou base station control network in a large sea area includes the following steps: Taking multi-dimensional system operating state parameters and environmental feature vectors as input, the state evaluation neural network is trained through supervised learning methods to fit and generate a system stability index that can characterize the operating state of the system in a specific environment. A deep feature extraction network model is used to characterize and learn the underlying features of the collected inter-node data channels, generating a core state vector that reflects the health of the system. This core state vector is then subjected to unsupervised cluster analysis to automatically classify and identify the current operating mode of the system. Equipped with a deep reinforcement learning decision-making intelligent agent, the state space of the intelligent agent includes the system stability index, environmental feature vector, system node resource occupancy rate and data transmission path efficiency index; the action space of the intelligent agent includes adjustment instructions for system node layout, connection topology and operating parameters; the intelligent agent is trained through an adaptive risk-aware composite reward function, and the decision-making neural network is iteratively optimized to generate an adaptive system configuration and adjustment strategy.

[0006] Preferably, the system operation status parameters select at least three types of key indicators: the first type is the quality factor of the node-to-node interaction signal reflecting the clarity of the data flow, the second type is the data interaction delay index reflecting the processing efficiency, and the third type is the data block error rate reflecting the transmission fidelity; the environmental characteristic vector includes wave level data, salt spray concentration index, and refractive index profile parameters of the atmospheric duct effect.

[0007] Preferably, the step of using a deep feature extraction network model to perform feature extraction is specifically as follows: selecting a variational autoencoder as the deep feature extraction network model, the variational autoencoder being symmetrically composed of an encoder and a decoder; inputting the collected underlying features of the data channel into the encoder, compressing the input data through multi-layer nonlinear transformations, mapping it to a low-dimensional latent space, and outputting parameters for defining a local probability distribution in the latent space, wherein the parameters are specifically the mean vector and variance vector of the Gaussian distribution; applying reparameterization techniques to perform random sampling from the Gaussian distribution, and determining the sampling results as the core state vector.

[0008] Preferably, the unsupervised cluster analysis includes: based on the set core point, boundary radius and minimum sample number parameters, using a density-based noise application space clustering algorithm to divide the core state vector into different dense areas, and establish a mapping rule between the dense areas and the physical operation mode of the system; mapping high-density sample clusters into continuous system macroscopic operation states, while discrete low-density samples are identified as isolated incidental events.

[0009] Preferably, the action space of the deep reinforcement learning decision-making agent contains adjustment instructions specifically as follows: at least three different levels of operation strategies, the first level is the adjustment of the physical interaction parameters of the nodes, dynamically adjusting the output signal energy level of a node; the second level is the reselection of the logical path of the data flow. When it is monitored that the performance of a key data transmission path is about to deteriorate, the agent can take the initiative to make a decision to switch the data flow carried by the path to one or more predefined redundant data paths; the third level is the resource scheduling of the system topology. During the peak period of system load or when a key node fails, the agent can issue instructions to remotely activate the backup node connection to reconstruct the system connection topology.

[0010] Preferably, the composite reward function includes: a positive core performance reward item, which is positively correlated with the improvement or maintenance level of the system stability index; a positive service quality reward item, which is triggered when the performance indicator of the key data transmission path is maintained above a preset threshold; and a negative resource consumption penalty item, which is related to the system resources or energy consumption generated by executing the adjustment instruction; by configuring different weight coefficients for the core performance reward item, service quality reward item and resource consumption penalty item, a composite reward function for balancing multiple objectives is constructed.

[0011] Preferably, the generation of adaptive system configuration and adjustment strategy includes: the decision neural network can dynamically adjust the decision priority according to the environmental risk level reflected by the real-time values ​​of the system stability index and the data transmission path efficiency index. When the indicator deteriorates due to increased environmental risk, the adjustment instruction that can maximize the sum of core performance and service quality rewards is given priority; conversely, when the indicator is in good condition and the environmental risk is low, the instruction that can minimize the resource consumption penalty is given priority to generate an adaptive system configuration and adjustment strategy.

[0012] Preferably, the implementation and optimization of the decision neural network specifically include: using a deep Q network structure DQN to construct the decision neural network to fit the state-action value function of the deep reinforcement learning decision agent; during the training process, the agent generates experience data through interaction with the environment, and samples the experience data through an experience replay mechanism and establishes a target network with parameter update delay for calculating the target Q value, and the calculation of the target Q value incorporates a composite reward function; by minimizing the temporal difference error between the Q value predicted by the decision neural network and the target Q value calculated by the target network as the loss function, and using the gradient descent algorithm to update the weight parameters of the decision neural network until the network model converges.

[0013] Compared with the prior art, the beneficial effects of the present invention are embodied in: 1. This method, employing the core concept of deep Q-networks within a vast and complex environment, transforms the complex deployment problem into a sequential decision-making problem. Through repeated interactions with the environment, the model learns an optimal state-action-value function. Through self-learning and exploration, it autonomously finds a near-optimal base station deployment strategy. Rather than relying on fixed rules or formulas, this method dynamically adjusts based on real-time feedback during the deployment process, achieving truly adaptive deployment and ultimately obtaining a high-quality layout solution.

[0014] 2. By storing interaction experiences in a data pool and randomly sampling them, the temporal correlation between training data is broken, allowing data to be reused, greatly improving learning efficiency and sample value. This method significantly shortens the time required for model training and makes better use of the data generated by each simulation deployment. This avoids the problem of low learning efficiency caused by strong data correlation, allowing the model to learn from historical experience more quickly.

[0015] 3. This method effectively addresses the problem of model parameter fluctuations or non-convergence caused by constantly changing learning objectives during training. By establishing a target network with delayed parameter updates, separating the network used to calculate the target Q value from the decision network being updated and delaying the target network's update, this provides a relatively fixed "bull's eye" for training, avoiding the instability of "chasing a moving target" training and ensuring reliable convergence of the algorithm. This ensures a smoother and more reliable training process, ultimately resulting in a stable, high-performance decision model. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 This is a flowchart of the steps of a method for adaptively deploying a Beidou base station control network in a large sea area proposed in the present invention; Figure 2 This is a flowchart of the deep reinforcement learning decision-making agent action space adjustment instructions proposed in the embodiment of the present invention; Figure 3 This is a flowchart of the steps of the decision neural network proposed in the embodiment of the present invention. DETAILED DESCRIPTION

[0017] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0018] See also Figures 1 to 3 The present invention provides a method for adaptively deploying a Beidou base station control network in a large sea area. The technical solution is as follows: Taking multi-dimensional system operating state parameters and environmental feature vectors as input, the state evaluation neural network is trained through supervised learning methods to fit and generate a system stability index that can characterize the operating state of the system in a specific environment. A deep feature extraction network model is used to characterize and learn the underlying features of the collected inter-node data channels, generating a core state vector that reflects the health of the system. This core state vector is then subjected to unsupervised cluster analysis to automatically classify and identify the current operating mode of the system. Equipped with a deep reinforcement learning decision-making intelligent agent, the state space of the intelligent agent includes the system stability index, environmental feature vector, system node resource occupancy rate and data transmission path efficiency index; the action space of the intelligent agent includes adjustment instructions for system node layout, connection topology and operating parameters; the intelligent agent is trained through an adaptive risk-aware composite reward function, and the decision-making neural network is iteratively optimized to generate an adaptive system configuration and adjustment strategy.

[0019] Example 1: This implementation scenario uses a deep-sea region as an example. This region exhibits typical complexity, including frequent extreme weather events, complex hydrogeology, a volatile atmosphere, high-salinity sea breezes, limited communications, and high maintenance costs. In such an environment, severe weather or equipment failure can lead to widespread disruption of navigation services, severely impacting offshore operations. Therefore, the core goal of this implementation approach is to create a networked agent capable of real-time perception, intelligent diagnosis, autonomous decision-making, and adaptive execution.

[0020] The BeiDou base station control network in this implementation scenario consists of several floating or semi-submersible base station platforms distributed in a deep-sea area. Each base station platform integrates BeiDou signal receiving, processing and forwarding modules, high-precision atomic clocks, communication modules, various environmental sensors, internal status monitoring units, and energy supply systems. All base stations conduct two-way data exchange and command issuance with the central intelligent control platform located on land through highly reliable satellite communication links or long-distance microwave links. Figure 1 , which is a flow chart of the steps of a method for adaptive deployment of Beidou base station control network in a large sea area proposed in the present invention application.

[0021] This method integrates five major functions: data fusion, state assessment, pattern recognition, decision optimization, and instruction generation. Each step is closely connected, combining the control network itself with environmental characteristics for analysis. After deep reinforcement learning, the intelligent agent makes decisions and provides adjustment plans, forming a closed-loop perception from analysis to decision-making to execution.

[0022] Furthermore, data collection at the base station end is continuously collected and encrypted for transmission at each base station node. Specifically, network performance data includes the signal quality factor (SQF) of the interaction signal between nodes, the data interaction delay index (DID), and the data block error rate (DBER); environmental characteristic data includes wave level data, salt fog concentration index, and the refractive index profile parameters of the atmospheric waveguide effect; node resource status data includes the base station's CPU memory usage, power supply voltage and current, internal temperature, storage space, and equipment operating time; data channel underlying characteristics include instantaneous channel gain, phase noise, multipath effect characteristics, frequency selective fading mode, background noise power spectrum density, and demodulation error vector amplitude high-dimensional raw signal data; The intelligent control platform receives and integrates various data uploaded by base stations, distributing it to various AI modules for processing. These modules work together to generate the optimal adjustment strategy, which is then transmitted back to the relevant base station nodes via secure communication links to execute control network configuration and parameter adjustments. This approach of receiving and issuing commands at any time allows for more intuitive perception of the current state of the control network and the implementation of the most appropriate adjustment actions.

[0023] Furthermore, input parameters and data acquisition include the inter-node signal quality factor (SQF), which reflects the clarity of the data stream. Each base station's communication module incorporates a high-performance digital signal processor that monitors the signal strength RSSI, signal-to-noise ratio (SNR), and bit error rate (BER) between other base stations or satellite links in real time, and calculates a comprehensive quality factor percentage. The SQF decreases significantly when wind and waves in a particular sea area cause the base station platform to shake, or when atmospheric ducting causes signal fading.

[0024] The Data Interaction Delay (DID) metric, which reflects processing efficiency, is measured by measuring the round-trip time and end-to-end latency of data packets between base stations. DID increases when communication links are congested, the base station CPU is overloaded, or the number of routing hops increases.

[0025] The Block Error Rate (DBER), which reflects transmission fidelity, uses the cyclic redundancy check (CRC) or forward error correction (FEC) report of data packets to calculate the proportion of error bits in received data blocks. DBER can increase significantly in the presence of strong interference, severe multipath effects, or channel fading.

[0026] These system operating parameters are collected in real time by the base station's network monitoring agent at a frequency of once per second. By taking data of different dimensions in real time and packaging it and uploading it to the system for analysis and processing, the operating status of each module of the base station is fully controlled. The operating status is transmitted to the system for full-angle monitoring, ensuring that humans can perceive and record the real-time parameters of the control network.

[0027] Furthermore, the environmental feature vector includes wave level data, which measures wave height and period in real time and automatically categorizes the wave level from 0 to 9 according to international sea condition standards. During a typhoon, the wave level can rapidly soar from 3-4 to 7-9.

[0028] Salt fog concentration index: In summer, the temperature is high and the humidity is high, and the salt fog concentration is high all year round, which seriously corrodes the equipment and also affects the wave transmission performance of the RF window.

[0029] Refractive index profile parameters of atmospheric ducting: Based on historical meteorological data and numerical forecast models, we calculate a profile curve of atmospheric refractive index as it varies with altitude. Key parameters such as duct height, duct strength, and duct top are extracted from this data. When strong evaporation ducts are present, signals in certain frequency bands propagate far near the sea surface, while signals at slightly higher altitudes experience severe attenuation.

[0030] These parameters directly affect the stability and fading characteristics of beyond-line-of-sight communication links. Sea conditions are complex and change extremely rapidly. By monitoring these three types of environmental characteristic vectors, the current sea environment can be understood remotely. Future sea changes can be further predicted based on the current environmental characteristics of the sea area. The risk assessment of sea environment changes can be graded based on the prediction results, providing more forward-looking suggestions for parameter regulation and avoiding uncontrollable emergencies caused by bad weather.

[0031] Furthermore, at the land-based control center, a large amount of historical operational data is pre-collected, covering various sea conditions, meteorological conditions, and system operating modes. A team of experts manually annotates each set of data with a stability score for the overall system stability. This data is used to train a multi-layer perceptron (MLP) neural network. The training process allows the state assessment neural network to learn how to accurately predict system stability based on six input dimensions: SQF, DID, DBER, wave level, salt spray concentration, and waveguide parameters. Once training is complete, this state assessment neural network can receive current data in real time and immediately generate a quantitative system stability index (SSI). For example, when a typhoon approaches, wave levels suddenly increase, accompanied by a deterioration in DID and DBER. The SSI will quickly drop from its usual 85 points to 40 points, clearly indicating that the system has entered an unstable state.

[0032] Furthermore, the underlying characteristics of the inter-node data channel captured in real time are extracted, including but not limited to: instantaneous channel gain, which reflects the real-time amplification or attenuation of signal propagation; phase noise, which measures the degree of random jitter in the signal phase and affects data demodulation accuracy; multipath effect characteristics. In complex oceans, signals form multiple paths to the receiver through sea surface reflection, atmospheric refraction, and other factors. VAE analyzes the time delay, intensity, and phase differences of these multipath signals; frequency selective fading, which refers to the fact that signals in certain frequency bands experience more severe fading under specific channel conditions; background noise power spectral density, which measures the intensity distribution of background noise and directly affects the signal-to-noise ratio; and demodulation error vector magnitude, which measures the deviation between the actual modulated signal and the ideal modulated signal and directly reflects the signal transmission quality. Together, these characteristics reflect the smoothness of the inter-node data channel. These raw characteristics are high-sampling-rate, high-dimensional time series data, such as collecting 100 signal snapshots per second, each containing different underlying physical parameters.

[0033] To extract comprehensive indicators to reflect the smoothness of data channels between nodes and improve computational efficiency, a variational autoencoder is selected as the deep feature extraction network. It consists of two main parts: an encoder and a decoder.

[0034] Specifically, the encoder takes as input high-dimensional raw channel feature data, for example, a 100-dimensional vector composed of ten signal snapshots. Through multiple layers of complex nonlinear transformations, this massive data set is compressed and mapped into a low-dimensional latent space. During this process, the encoder does not directly output a fixed point, but rather outputs parameters that define a "local probability distribution" in this latent space. For a Gaussian distribution, this output consists of a series of mean and variance vectors, representing the center position and spread of the distribution, respectively.

[0035] Specifically, the decoder reconstructs the original input data from the sampling results in the latent space. The decoder also consists of multiple layers of nonlinear transformations, which is roughly symmetrical with the encoder structure.

[0036] Furthermore, VAE does not directly sample the mean and variance during learning, but instead uses a reparameterization technique. This technique allows the system to indirectly sample by introducing a random, pre-distributed noise during the learning process. This noise is combined with the mean and variance of the encoder output to produce a specific vector extracted from the latent space probability distribution. This core state vector CSV, which is a sampled low-dimensional vector that contains the most essential features of the original channel data, is mapped into a 16-dimensional vector, each dimension of which represents a clear abstract feature of the channel health.

[0037] By integrating predicted environmental trends, real-time base station status, and historical fault data as inputs to the risk assessment neural network, comprehensive awareness and proactive early warning of Beidou base station control network risks are achieved, rather than just passive response. Its training data combines expert experience, preset rules, and real-world historical records to ensure the accuracy and reliability of risk labels, making the quantitative risk levels output by the neural network more realistic and instructive. This not only improves the accuracy and foresight of risk assessments, but also provides a high-quality foundation for subsequent deep reinforcement learning agents to make timely and effective adaptive deployment and risk avoidance decisions, significantly enhancing network resilience and management efficiency.

[0038] Furthermore, the density-based noise application spatial clustering algorithm DBSCAN is used for unsupervised cluster analysis. The advantage of DBSCAN is that it can identify data clusters of any shape and can effectively distinguish "dense" data points from "sparse" abnormal data points. Specifically, two key parameters of DBSCAN are pre-set: the boundary radius is defined as the "neighborhood" around a data point. For example, in the 16-dimensional space of CSVs, a specific distance value of 0.5 is set, indicating that if the distance between two CSVs is less than 0.5, they are considered to be adjacent. The minimum number of samples is defined as the minimum number of data points in the neighborhood required for a point to be considered a "core point." For example, a setting of 5 means that a CSV must have at least 5 other CSVs within a radius of 0.5 to be considered a core point.

[0039] Furthermore, the massive core state vectors (CSVs) generated by the VAE during system operation are fed into the DBSCAN algorithm. Based on a set boundary radius and minimum sample count, DBSCAN automatically identifies high-density regions in the CSV space and divides the CSVs within these regions into distinct clusters.

[0040] Specifically, each identified high-density CSV cluster is mapped to a persistent macroscopic system operating state. Normal operation corresponds to a dense cluster with a large number of CSVs, indicating a healthy channel with stable performance. Atmospheric waveguide interference corresponds to another CSV cluster and is highly correlated with abnormal changes in refractive index profile parameters, leading to specific changes in signal propagation characteristics. Localized channel degradation patterns, caused by base station antenna icing, salt spray deposition, or localized hardware aging, manifest as a collective deterioration of specific channel parameters. Mild electromagnetic interference patterns, associated with nearby ship radars or illegal signal sources, manifest as specific patterns of background noise.

[0041] Discrete low-density samples are identified by DBSCAN as isolated CSV points that do not belong to any high-density clusters. These points are identified as isolated incidents or transient anomalies, such as a brief strong pulse interference, a momentary communication link interruption, or a transient jump in sensor readings.

[0042] Furthermore, while the system is running, each newly generated CSV is compared in real time with the learned clusters. If it falls into a known cluster, the system automatically identifies the current operating mode. If it is identified as a noise point by DBSCAN, an anomaly alarm is triggered, indicating that an unknown or transient anomaly has occurred.

[0043] Furthermore, the agent's state space, a snapshot of the agent's perceived environment, provides a comprehensive basis for decision-making. It comprises the following real-time data: the System Stability Index (SSI), which reflects the overall health of the system; the Environmental Eigenvector, which reflects the degree of challenge in the external environment; the System Node Resource Utilization, which reflects the operational load and hardware health of each base station; and the Data Transmission Path Performance Indicator, which reflects the quality of network service.

[0044] These multi-dimensional data are integrated into a unified, high-dimensional vector as input for each decision of the intelligent agent. Using different data dimensions as comprehensive indicators saves computing resources and presents them as a more clear numerical value.

[0045] The agent's action space stores a set of executable discrete operation instructions. To achieve refined and multi-level network adjustments, the agent's action space includes the following three different levels of operation strategies, covering a wide range of adjustments from the physical layer to the network layer: Specifically, the first level involves adjusting the physical interaction parameters of nodes. When the signal detected in a specific direction is severely attenuated by atmospheric waveguides, the agent can instruct base station X to increase its transmit power by 3 decibels to penetrate the attenuated area. When sea conditions are good and the signal is sufficient, base station Y can be instructed to reduce its transmit power to save energy and minimize potential interference with other communications. To address pointing deviations caused by the floating platform's swaying in the waves, the agent can instruct base station Z to fine-tune the antenna's azimuth or elevation to ensure the beam is precisely aligned with the target node or satellite. Second level: When the performance of a critical data transmission path is detected to be deteriorating, a typhoon is predicted to affect the path, or latency is detected to be consistently exceeding preset thresholds or packet loss is increasing, the agent can proactively switch the navigation data flow or control command flow carried by that path to one or more predefined, redundant data paths with better performance. This is achieved by modifying the routing table on the network router or the flow table rules in the software-defined network (SDN) controller, ensuring the continuity and reliability of critical data transmission. When multiple paths are available, the agent can intelligently distribute data traffic across them based on their real-time load and bandwidth, avoiding single points of congestion and improving overall network throughput.

[0046] The third level: During peak system load periods, when critical nodes fail, such as when a primary base station fails due to wind and waves or aging equipment, or when the environment in a specific sea area deteriorates dramatically, the agent can issue commands to remotely activate pre-deployed backup base station nodes in a low-power standby state. If a typhoon approaches and a primary base station is expected to temporarily lose connection, the agent can pre-activate a nearby backup base station, proactively establish a new communication link, and quickly reconfigure the system's connectivity topology to ensure network coverage and redundancy. It also shuts down unnecessary nodes to conserve energy. Based on predicted traffic volume and environmental conditions, the agent can adjust the base station's periodic sleep and wake-up strategies to further optimize energy consumption.

[0047] Reference Figure 2 This is a flowchart of the action space adjustment instructions for the deep reinforcement learning decision-making agent proposed in an embodiment of the present invention. The decision-making actions define three different layers of action strategies, greatly enhancing the deep reinforcement learning agent's ability to fine-tune and multi-dimensionally control the Beidou base station control network. This comprehensive, three-level, hierarchical action space gives the agent flexible and powerful adaptive deployment and management capabilities, effectively navigating complex and dynamic maritime environments.

[0048] Furthermore, the adaptive risk-aware compound reward function is a guiding signal for the agent to learn good behavior. To balance multiple objectives such as network stability, service quality, and resource consumption, this method constructs a comprehensive compound reward function. This reward function is a weighted combination of multiple terms: A positive core performance reward: When the System Stability Index (SSI) improves compared to the previous level or remains at a high level, the agent will receive a positive reward. This reward encourages the agent to take actions that can maintain or improve the overall health of the network.

[0049] A positive quality of service reward: When the performance indicators of key data transmission paths, such as average latency below 100 milliseconds and packet loss rate below 0.5%, remain above the preset quality threshold, the agent will also receive a positive reward. This reward ensures that the agent optimizes overall network performance without affecting the transmission quality of critical services.

[0050] A negative resource consumption penalty: Any adjustment command executed by the agent consumes corresponding system resources or energy. Increasing transmit power and activating backup nodes consume more energy; frequent routing increases computational and communication overhead. This penalty is proportional to the amount of resources consumed, encouraging the agent to choose more efficient and energy-efficient strategies.

[0051] These three reward components are weighted and summed according to preset weight coefficients to construct the final composite reward value. These weight coefficients are adjustable. For example, in the early stages of training, the weights can be evenly distributed; after actual deployment, they can be fine-tuned based on operational strategy preferences.

[0052] A single reward function considers only one dimension and cannot achieve balanced optimization of multiple objectives. By constructing a composite reward function, we achieve a comprehensive consideration of both positive and negative rewards and penalties. Positive core performance rewards and service quality rewards incentivize the agent to continuously improve network performance and user experience. Meanwhile, negative resource consumption penalties guide the agent to achieve performance goals while also considering energy efficiency and operating costs. By assigning adjustable weight coefficients to each reward item, this method can flexibly adjust the optimization focus based on actual business needs or environmental changes, making the strategies learned by the agent more in line with actual needs, achieving refined control of intelligent decision-making and optimal resource allocation.

[0053] Furthermore, a deep Q network (DQN) structure is used to construct a decision neural network. It is a feedforward neural network with multiple hidden layers. The input layer receives the current complete state vector, and the output layer is the same size as the action space. Each output neuron represents the Q value of an action, that is, the long-term cumulative reward expected after executing the action. This reward is updated through the Q value back propagation to update the network parameters to achieve long-term positive feedback, organically combining the gradient optimization of deep learning with the value iteration of reinforcement learning to achieve efficient and stable strategy learning. Figure 3 , which is a flowchart of the steps of the decision neural network proposed in the embodiment of the application of the present invention.

[0054] During training, every experience the agent generates while interacting with the simulated environment—including its current state, actions taken, rewards received, and the next state—is stored in an experience replay buffer. During training, the system randomly draws a small batch of data from this buffer for learning, rather than following the actual chronological order of occurrence. This effectively breaks down data dependencies and improves training stability and efficiency.

[0055] Furthermore, to address the issue of unstable Q-value calculation targets during training, the system establishes an additional target network with the same structure as the main decision neural network. The parameters of the target network are periodically copied from the main decision network, but remain fixed between copies. This provides a more stable reference point for calculating the target Q-value, thereby stabilizing the training process.

[0056] The goal of training is to make the Q-values ​​predicted by the decision neural network as close as possible to the true Q-values ​​calculated by the target network. The difference between the two is used as the loss function, and the system attempts to minimize this difference. This difference is measured by calculating the temporal difference error between the two.

[0057] To minimize the loss function, the system uses a gradient descent algorithm to iteratively update the "weight parameters" within the decision neural network. By continuously adjusting these parameters, the predicted Q-values ​​become increasingly accurate, enabling the agent to select the optimal action. Training continues until the network model reaches convergence, meaning that the predicted Q-values ​​stabilize and no longer fluctuate significantly.

[0058] The Deep Q-Network (DQN) enables intelligent agents to evaluate the long-term rewards of different actions by fitting a state-action value function. The experience replay mechanism effectively breaks the temporal correlation between training samples, improving training stability and data utilization. The design of the target network further stabilizes the Q-value update process and alleviates oscillation issues during training. By minimizing the temporal difference error and iteratively updating network weights using the gradient descent algorithm, the decision neural network can efficiently converge and learn the optimal decision strategy. This provides accurate and reliable adaptive configuration and adjustment instructions for the Beidou base station control network, which is a key technical guarantee for achieving intelligent decision-making.

[0059] Furthermore, the process of adaptive system configuration and adjustment strategy to generate risk perception is as follows: Strategy in High-Risk Mode: When SSI performance degrades to below 60 due to extreme conditions, such as typhoons or strong waveguides, or when the performance of critical data transmission paths continues to deteriorate, the agent identifies the situation as high-risk. In this mode, the decision neural network prioritizes adjustment instructions that maximize the sum of core performance rewards and quality of service rewards. This means that even if the instruction incurs a high resource consumption penalty, the agent will execute it without hesitation to ensure the uninterrupted transmission of the network's core functions and critical services. In the compound reward function, the weights assigned to core performance rewards and quality of service rewards are dynamically increased, while the weight of resource consumption penalties is relatively reduced.

[0060] Low-risk mode strategy: Conversely, when the SSI is greater than 80, indicating good conditions, low environmental risk, calm sea conditions, low salt spray concentration, no significant waveguide effects, and all key data transmission path performance indicators are healthy, the agent identifies the low-risk mode. In this mode, the decision neural network prioritizes instructions that minimize resource consumption penalties. It tends to adopt more energy-efficient and conservative strategies, such as moderately reducing the transmit power of certain nodes to conserve energy or shutting down temporarily unneeded backup nodes. This maximizes resource utilization while maintaining good performance. In this case, the weight of the resource consumption penalty in the composite reward function increases dynamically, while the weight of the core performance reward and service quality reward decreases.

[0061] Through this dynamic, risk-aware weight adjustment mechanism, the intelligent agent can demonstrate a high degree of resilience and adaptability, ensuring core functions under severe challenges, and striving to optimize efficiency in a stable environment, thus realizing the true intelligence and adaptability of the Beidou base station control network. In addition, the decision neural network can dynamically adjust the decision priority and generate a more environmentally adaptable system configuration and adjustment strategy. This embodiment first analyzes the operating status and environmental characteristics of the control network to determine the current risk level, and then implements priority adjustment based on the adaptive decision-making of the risk level. The intelligent judgment of the system reduces manual dependence, allowing the Beidou base station control network to intelligently balance performance and efficiency in different environments, and to make adjustments at any time according to the environment, reasonably allocate resources, and avoid unnecessary waste of resources. While considering the intelligence of the system, it also considers economic benefits.

[0062] Example 2: Fifteen Beidou base station nodes are deployed in a distributed mesh structure in a deep-sea area. Five base stations (numbered BS01-BS05) located in the typhoon's path are key areas affected by the typhoon. The central intelligent control platform is located on Hainan Island.

[0063] The scenario event is an approaching strong typhoon, which is divided into three stages to realize the response of the intelligent agent.

[0064] Phase 1: Typhoon Approach Warning and Early Stage, responding to intelligent perception and preventive adjustments. At this time, the typhoon center was approximately 72 hours away from the BS01-BS05 area. The pressure sensor detected a continuous drop in air pressure, the wind speed sensor began to detect increasing gusts, and the wave buoy indicated a gradual increase in wave severity from Category 2 light waves to Category 4 moderate waves. The atmospheric refractive index profile parameters showed that the evaporation duct effect began to strengthen, indicating that the signal propagation characteristics near the sea surface would change.

[0065] Specifically, at the input of the SSI evaluation module, although SQF, DID, and DBER have not deteriorated significantly at this time, changes in the environmental feature vector have begun to cause the SSI score to drop slightly from 88 to 80, indicating that the system stability is beginning to be threatened.

[0066] VAE continues to process the underlying channel features, and the CSV clustering model still mainly identifies the “normal operation mode” cluster, but occasionally there are CSVs falling near the boundaries of “unstable propagation mode” or “channel fluctuation mode”.

[0067] Furthermore, the intelligent control platform receives the updated SSI and environmental characteristics, identifying the system's transition from "low risk" to "medium risk." Based on the trained decision neural network, the DRL agent predicts that the system's state will deteriorate further in the future. Under the adaptive risk-aware mechanism, the reward function now begins to weight "core performance" and "service quality."

[0068] The intelligent agent makes decisions and issues preventive adjustment instructions: it instructs BS01-BS05 and 2-3 surrounding nodes to increase the transmission power by 1-2 decibels to enhance the stability of the signal link and respond to potential channel fading. It proactively identifies several key data transmission paths that will be most severely affected along the typhoon path, and pre-switches some non-critical data streams on these paths, such as device logs and secondary monitoring data, to other redundant paths to free up more bandwidth on the primary path and prepare for subsequent critical data transmission.

[0069] Instruct two backup base stations (such as BS06 and BS07) located outside the typhoon path but able to establish connections with BS01-BS05 to enter the pre-activation state to preheat their communication modules so that they can go online quickly when needed.

[0070] Phase 2: The typhoon's core impact area maintains resilience in the high-risk mode. The typhoon's center passes directly over BS02 and BS03. Waves surge to Category 9, salt spray concentrations rise sharply, and the atmospheric ducting effect is temporarily weakened or becomes extremely unstable due to strong wind shear.

[0071] The SQF dropped dramatically from 90 to 40, the DID soared from 50ms to 300ms, and the DBER increased dramatically from 0.1% to over 5%. The SSI score instantly plummeted from 80 to below 20, putting the system in extremely high-risk mode. The internal temperature or power supply voltage of the BS02 base station experienced abnormal fluctuations, indicating that the equipment was under significant stress.

[0072] A large number of CSVs generated by VAE fall into severe channel fading patterns or high noise interference pattern clusters, and a large number of CSVs are identified by DBSCAN as isolated sporadic events, that is, they are recovered after a short loss of connection.

[0073] Furthermore, the intelligent control platform identified that the system was in an extremely high-risk mode. The adaptive risk perception mechanism was activated to the maximum extent possible, with the reward function weighted entirely toward core performance and service quality, and the penalty weight for resource consumption was minimized. The DRL agent quickly assessed the current state and decided to implement high-intensity, safeguarding adjustments: Instruct BS01-BS05 to increase the transmission power of all available links to the maximum limit and try to maintain the communication connection. Data flow logical path reselection: When the intelligent agent detected a temporary interruption of the critical microwave link between BS02 and BS03 due to high winds and high waves, it immediately decided to force all Beidou differential correction data flowing through that link to be switched to a backup satellite backhaul link forwarded via BS01 and BS04. The intelligent agent also decided to maximize compression of non-critical data streams, even suspending their transmission, to ensure that critical navigation and positioning data could be transmitted with minimal delay.

[0074] Instruct the pre-activated BS06 and BS07 to go fully online and attempt to establish new redundant links with all affected base stations to form new communication paths and build a temporary, more robust network topology.

[0075] Phase 3: Typhoon Passage and Recovery, Intelligent Optimization and Energy Saving. At this point, the typhoon center has been away from the BS01-BS05 area for approximately 24 hours. Wave severity gradually decreases to Category 5, salt spray concentration begins to decrease, and the atmospheric ducting effect stabilizes.

[0076] Indicators such as SQF, DID, and DBER gradually improved, and the SSI score steadily climbed from around 20 to over 75, shifting the system from high risk to medium-low risk. The CSV clustering model began to identify more "normal operating mode" clusters, and incidents decreased.

[0077] Furthermore, the intelligent control platform recognizes that system risk has decreased. The adaptive risk perception mechanism begins to prioritize resource consumption, and the reward function weight gradually returns to a balanced state.

[0078] The DRL agent evaluates the current state and makes decisions to implement optimization and energy-saving adjustments: Instruct BS01-BS05 and related nodes to gradually adjust the transmit power back to normal or slightly below normal levels to save energy.

[0079] The agent reassessed the load and performance of all links and discovered that the previously interrupted link between BS02 and BS03 had been restored. It then decided to gradually switch critical data flows, which had been routed during the typhoon, back to their original, more efficient or direct paths. It also optimized traffic distribution to ensure network efficiency.

[0080] The intelligent agent detected that network redundancy and performance had returned to a healthy level, and decided to gradually downgrade backup base stations such as BS06 and BS07, which were urgently activated during the typhoon, to low-power standby mode, or to completely hibernate if traffic volume permits, to maximize energy savings.

[0081] When making decisions, the intelligent agent will review the temporary links and topologies established during the typhoon to respond to emergencies and evaluate the long-term benefits. For temporary links that are no longer needed or inefficient, the intelligent agent will issue instructions to shut them down or downgrade them to restore to a leaner and more energy-efficient topology.

[0082] This embodiment, through the highly collaborative and adaptive mechanisms of artificial intelligence technology, achieves performance improvement, risk avoidance, resilience maintenance, and energy efficiency optimization of the Beidou base station control network in complex marine environments. This improves network availability, significantly shortens service interruptions, and reduces fault recovery time from hours to minutes, significantly reducing operation and maintenance costs and improving service quality. Through this dynamic, phased adaptive response, the Beidou base station control network can demonstrate unprecedented resilience, reliability, and intelligence in large and complex marine environments, especially in extreme weather conditions, ensuring the continuity of Beidou services and effectively balancing performance and resource consumption.

[0083] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A method for adaptively deploying a BeiDou base station control network in a large sea area, characterized in that: The following steps are involved: Taking multi-dimensional system operating state parameters and environmental feature vectors as input, the state evaluation neural network is trained through supervised learning methods to fit and generate a system stability index that can characterize the operating state of the system in a specific environment. A deep feature extraction network model is used to characterize and learn the underlying features of the collected inter-node data channels, generating a core state vector that reflects the health of the system. This core state vector is then subjected to unsupervised cluster analysis to automatically classify and identify the current operating mode of the system. Equipped with a deep reinforcement learning decision-making intelligent agent, the state space of the intelligent agent includes the system stability index, environmental feature vector, system node resource occupancy rate and data transmission path efficiency index; the action space of the intelligent agent includes adjustment instructions for system node layout, connection topology and operating parameters; the intelligent agent is trained through an adaptive risk-aware composite reward function, and the decision-making neural network is iteratively optimized to generate an adaptive system configuration and adjustment strategy.

2. The method for adaptively deploying a BeiDou base station control network in a large sea area according to claim 1, characterized in that: At least three types of key indicators are selected as the system operation status parameters: the first type is the quality factor of the node-to-node interaction signal reflecting the clarity of the data flow, the second type is the data interaction delay indicator reflecting the processing efficiency, and the third type is the data block error rate reflecting the transmission fidelity; the environmental feature vector includes wave level data, salt spray concentration index, and refractive index profile parameters of the atmospheric duct effect.

3. The method for adaptively deploying a BeiDou base station control network in a large sea area according to claim 1, characterized in that: The steps of extracting features using the deep feature extraction network model are specifically as follows: A variational autoencoder is selected as the deep feature extraction network model, wherein the variational autoencoder is symmetrically composed of an encoder and a decoder; Input the collected underlying features of the data channel into the encoder, compress the input data through multiple layers of nonlinear transformations, map it to a low-dimensional latent space, and output parameters for defining a local probability distribution in the latent space, wherein the parameters are specifically the mean vector and variance vector of the Gaussian distribution; A reparameterization technique is applied to randomly sample from the Gaussian distribution, and the sampling result is determined as the core state vector.

4. The method for adaptively deploying a BeiDou base station control network in a large sea area according to claim 1, characterized in that: The unsupervised cluster analysis includes: Based on the set core point, boundary radius and minimum sample number parameters, a density-based noise application spatial clustering algorithm is used to divide the core state vector into different dense areas, and a mapping rule between the dense areas and the physical operation mode of the system is established; High-density sample clusters are mapped to the persistent macroscopic operating state of the system, while discrete low-density samples are identified as isolated incidental events.

5. The method for adaptively deploying a BeiDou base station control network in a large sea area according to claim 1, characterized in that: The action space of the deep reinforcement learning decision agent contains the following adjustment instructions: It includes at least three different levels of operation strategies. The first level is the adjustment of node physical interaction parameters, dynamically adjusting the output signal energy level of a node; the second level is the reselection of data flow logical paths. When it is monitored that the performance of a key data transmission path is about to deteriorate, the intelligent agent can take the initiative to make a decision to switch the data flow carried by the path to a predefined redundant data path; the third level is the resource scheduling of the system topology structure. During the peak period of system load or when a key node fails, the intelligent agent can issue instructions to remotely activate the backup node connection and reconstruct the system connection topology.

6. The method for adaptively deploying a BeiDou base station control network in a large sea area according to claim 1, characterized in that: The compound reward function includes: A positive core performance reward item that is positively correlated with the improvement or maintenance level of the system stability index; A positive service quality reward is triggered when the performance indicator of the key data transmission path remains above the preset threshold; and a negative resource consumption penalty term, which is related to the amount of system resources or energy consumed by executing the adjustment instruction; By configuring different weight coefficients for the core performance reward item, the service quality reward item and the resource consumption penalty item, a composite reward function for balancing multiple objectives is constructed.

7. The method for adaptively deploying a BeiDou base station control network in a large sea area according to claim 1, characterized in that: The generating of the adaptive system configuration and adjustment strategy includes: The decision neural network can dynamically adjust the decision priority based on the environmental risk level reflected by the real-time values ​​of the system stability index and the performance index of the data transmission path. When the performance index deteriorates due to the increase of environmental risk, the decision neural network will give priority to the adjustment instructions that can maximize the sum of the core performance and service quality rewards. On the contrary, when the performance indicator is in good condition and the environmental risk is low, instructions that can minimize resource consumption penalties are preferentially selected to generate an adaptive system configuration and adjustment strategy.

8. The method for adaptively deploying a BeiDou base station control network in a large sea area according to claim 1, characterized in that: The implementation and optimization of the decision neural network specifically include: The decision neural network is constructed using a deep Q-network structure (DQN) to fit the state-action value function of the deep reinforcement learning decision agent. During training, the agent generates experience data through interaction with the environment, samples the experience data through an experience replay mechanism, and establishes a target network with parameter update delay to calculate the target Q value. The calculation of the target Q value is integrated with the composite reward function. The loss function is determined by minimizing the temporal difference error between the Q value predicted by the decision neural network and the target Q value calculated by the target network, and the weight parameters of the decision neural network are updated using the gradient descent algorithm until the network model converges.

Citation Information

Patent Citations

  • Unmanned ship risk adaptive navigation algorithm based on distributed reinforcement learning

    CN118747519A

  • Multi-agent deep reinforcement learning path planning method based on improved A*heuristic

    CN118759846A

  • Coal mine dust diffusion simulation and control method based on artificial intelligence

    CN119623238A

  • DRL-based control logic design method for continuous microfluidic biochips

    US20230401367A1