Satellite differential data broadcasting method, device, equipment, medium and program product

CN122764286APending Publication Date: 2026-09-15CHINA MOBILE SHANGHAI ICT CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610805846.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-05
Publication Date
2026-09-15

AI Technical Summary

Technical Problem

[0004]本发明的目的在于提供一种卫星差分数据的播发方法、装置、设备、介质及程序产品,用以解决现有RTK差分数据播发技术存在的固定播发模式低效、云带宽成本和运营成本高的问题

Benefits of technology

本发明实施例中,通过获取多模态数据;其中,所述多模态数据包括以下的一项或多项:地理环境数据、用户场景数据、终端特性与状态数据、电离层指数分布数据;对所述多模态数据进行特征提取与融合处理,获得融合特征向量;通过Actor网络的一个或多个全连接层,对所述融合特征向量进行非线性变换,得到策略隐向量;对所述策略隐向量进行分层动作决策,生成播发策略,所述播发策略包括播发频率、压缩等级、星座、频点权重以及高级改正数的启用或禁用;按照所述播发策略,进行卫星差分数据的播发,这样,采用本发明的方法能够根据地理环境数据、用户场景数据、终端特性与状态数据和/或电离层指数分布数据,动态、智能地调整播发策略,从而实现资源优化、提升定位性能并降低云带宽成本和运营成本。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122764286A_ABST
    Figure CN122764286A_ABST
Patent Text Reader

Abstract

The application provides a satellite differential data broadcasting method, device, equipment, medium and program product. The method comprises: acquiring multi-modal data; wherein the multi-modal data comprises one or more of the following: geographic environment data, user scene data, terminal characteristic and state data, ionospheric index distribution data; performing feature extraction and fusion processing on the multi-modal data to obtain a fusion feature vector; performing nonlinear transformation on the fusion feature vector through one or more fully connected layers of an Actor network to obtain a policy hidden vector; performing hierarchical action decision on the policy hidden vector to generate a broadcasting strategy, wherein the broadcasting strategy comprises a broadcasting frequency, a compression level, a constellation, a frequency point weight, and the enabling or disabling of high-level corrections; and broadcasting satellite differential data according to the broadcasting strategy. The application can dynamically and intelligently adjust the broadcasting strategy, thereby optimizing resources, improving positioning performance, and reducing cloud bandwidth costs and operating costs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data broadcasting technology, and in particular to a method, apparatus, equipment, medium, and program product for broadcasting satellite differential data. Background Technology

[0002] Real-Time Kinematic (RTK) is a core technology for high-precision satellite navigation and positioning. Its basic principle is that the base station broadcasts differential correction information of its observations to the mobile station (user terminal) via a data link. The mobile station uses this information to perform real-time joint calculations with its own observations, thereby obtaining positioning accuracy at the centimeter or even millimeter level. Therefore, the differential data broadcasting scheme is a crucial link in achieving high-precision positioning in an RTK system. The content of RTK differential data broadcast depends mainly on the differential data format and standard used. A common one is the internationally recognized Global Navigation Satellite System (GNSS) differential data standard protocol—RTCM SC-104. Its main broadcast content includes observation data (such as pseudorange, carrier phase, etc.), base station coordinate information (broadcasting the precise coordinates of the base station), ionospheric and tropospheric correction parameters (used to improve the calculation accuracy of medium- and long baselines), and satellite orbit corrections (state-space characterization information).

[0003] Traditional RTK networks use a fixed 1Hz broadcast frequency, failing to consider that static terminals in industries like surveying and mapping do not require high-frequency broadcasting. Furthermore, broadcasting the full amount of data to all terminals results in significant bandwidth waste, with much of the data remaining unused due to hardware limitations. This leads to a bandwidth waste rate of up to 60%. When there are millions of online users, the annual cloud internet bandwidth rental costs are substantial. In other words, existing RTK differential data broadcasting technology suffers from inefficiencies in its fixed broadcasting mode and high cloud bandwidth and operational costs. Summary of the Invention

[0004] The purpose of this invention is to provide a method, apparatus, equipment, medium, and program product for broadcasting satellite differential data, in order to solve the problems of inefficiency in fixed broadcasting mode, high cloud bandwidth cost, and high operating cost in existing RTK differential data broadcasting technology.

[0005] To achieve the above objectives, in a first aspect, embodiments of the present invention provide a method for broadcasting satellite differential data, comprising: Acquire multimodal data; wherein the multimodal data includes one or more of the following: geographic environment data, user scenario data, terminal characteristics and status data, and ionospheric index distribution data; The multimodal data is subjected to feature extraction and fusion processing to obtain a fused feature vector; The policy latent vector is obtained by performing a nonlinear transformation on the fused feature vector through one or more fully connected layers of the Actor network. A hierarchical action decision is made on the strategy latent vector to generate a broadcast strategy, which includes enabling or disabling broadcast frequency, compression level, constellation, frequency point weight, and advanced correction numbers. Satellite differential data is broadcast according to the broadcasting strategy described above.

[0006] In some embodiments, the step of performing feature extraction and fusion processing on the multimodal data to obtain a fused feature vector includes: The geographic environment data is used to extract features through a geographic convolutional neural network to obtain a first feature vector. The second feature vector is obtained by extracting features from the ionospheric exponential distribution data using a time-series coding neural network. The third feature vector is obtained by extracting features from the user scenario data and the terminal characteristics and status data through a structured data neural network. Based on the first feature vector, the second feature vector, and / or the third feature vector, feature fusion processing is performed to obtain the fused feature vector.

[0007] In some embodiments, the step of extracting features from the user scenario data and the terminal characteristics and status data using a structured data neural network to obtain a third feature vector includes: The user scenario data is encoded using the encoding unit of a structured data neural network to obtain a low-dimensional and sparse binary vector. The terminal characteristics and status data are numerically normalized using the normalization unit of the structured data neural network to obtain a hybrid vector. The low-dimensional and sparse binary vector and the hybrid vector are input into the fully connected layer of the structured data neural network to obtain the third feature vector.

[0008] In some embodiments, the feature fusion process based on the first feature vector, the second feature vector, and / or the third feature vector to obtain the fused feature vector includes: A linear transformation is performed on the first feature vector to obtain a first query matrix, a first key matrix, and a first value matrix; a first attention weight is calculated based on the first query matrix and the first key matrix; and a weighted fusion is performed on the first value matrix based on the first attention weight to obtain a first output result. A linear transformation is performed on the second feature vector to obtain a second query matrix, a second key matrix, and a second value matrix; based on the second query matrix and the second key matrix, a second attention weight is calculated; and based on the second attention weight, the second value matrix is ​​weighted and fused to obtain a second output result. A linear transformation is performed on the third feature vector to obtain a third query matrix, a third key matrix, and a third value matrix; based on the third query matrix and the third key matrix, a third attention weight is calculated; and based on the third attention weight, the third value matrix is ​​weighted and fused to obtain a third output result. The first output result, the second output result, and / or the third output result are concatenated and then subjected to a linear transformation to obtain the fused feature vector.

[0009] In some embodiments, the step of performing hierarchical action decision-making on the policy latent vector to generate a broadcasting policy includes: Gaussian distribution sampling is performed on the latent vectors of the strategy to obtain the broadcast frequency and compression level; Gumbel-Softmax sampling is performed on the latent vector of the strategy to obtain constellation and frequency point weights; The latent vectors of the policy are sampled using a Bernoulli distribution and processed with an activation function to determine whether advanced corrections are enabled or disabled.

[0010] In some embodiments, the method further includes: The broadcast strategy is processed by action masking to obtain compliant actions; The compliant actions are verified for budget constraints to obtain budget-safe actions; Based on the budgeted safety actions, the final safety actions are obtained through pessimistic Q-value initialization and safety constraint processing of the safety action set, and the final safety actions are determined as the final broadcast strategy. The broadcasting of satellite differential data according to the broadcasting strategy includes: Satellite differential data will be broadcast according to the final broadcast strategy described above.

[0011] In some embodiments, the step of performing action masking processing on the broadcast strategy to obtain compliant actions includes: Based on the terminal capabilities, the broadcasting strategy is masked to obtain a broadcasting strategy that matches the terminal capabilities. Based on business rules, the broadcasting strategy that matches the terminal capabilities is masked to obtain the compliant action.

[0012] In some embodiments, the step of performing budget constraint verification on the compliance action to obtain a budget-safe action includes: The resource consumption prediction model is used to obtain the predicted resources consumed by the compliance action. Get the currently consumed resources; If the sum of the predicted resources consumed by the compliance action and the currently consumed resources is greater than a preset threshold, the compliance action is corrected to obtain the budgeted security action, wherein the resources consumed by the budgeted security action are less than or equal to the difference between the predicted resources consumed by the compliance action and the currently consumed resources. If the sum of the predicted resources consumed by the compliance action and the currently consumed resources is less than or equal to the preset threshold, the compliance action is determined to be the budget security action.

[0013] In some embodiments, obtaining the final safety action based on the budgeted safety action, through pessimistic Q-value initialization and safety constraint processing of the safety action set, includes: The Q-value output of the Crtic network is initialized to a constant that is much lower than the normal reward range; Determine whether the budgeted safety action belongs to the set of safety actions; If it is, then the budget security action is determined to be the final security action, and the Q value is updated after the final security action is executed; If it does not belong to the set of safety actions, then select the safety alternative action that is most similar to the budgeted safety action from the set of safety actions, and determine the safety alternative action as the final safety action.

[0014] In some embodiments, the method further includes: Obtain multi-dimensional feedback data after the broadcast strategy is executed; Based on the multi-dimensional feedback data, calculate the corresponding reward value for each; The total reward value is calculated by weighting the values ​​based on their respective reward values ​​using a discount factor. Based on the total reward value, update the Q value of the Critical network; The policy parameters of the Actor network are updated based on the Q-value gradient, wherein the Q-value gradient is obtained based on the updated Q-value and the Q-value before the update.

[0015] Secondly, embodiments of the present invention also provide a satellite differential data broadcasting device, comprising: The first acquisition module is used to acquire multimodal data; wherein, the multimodal data includes one or more of the following: geographic environment data, user scenario data, terminal characteristics and status data, and ionospheric index distribution data; The feature extraction and fusion module is used to perform feature extraction and fusion processing on the multimodal data to obtain a fused feature vector; The first processing module is used to perform a nonlinear transformation on the fused feature vector through one or more fully connected layers of the Actor network to obtain the policy latent vector; The broadcast strategy generation module is used to make hierarchical action decisions on the strategy latent vector and generate a broadcast strategy. The broadcast strategy includes enabling or disabling broadcast frequency, compression level, constellation, frequency point weight, and advanced correction numbers. The data broadcasting module is used to broadcast satellite differential data according to the broadcasting strategy.

[0016] Thirdly, embodiments of the present invention also provide a satellite differential data broadcasting device, including a processor and a transceiver, wherein the transceiver receives and transmits data under the control of the processor, and the processor is used to perform the following operations: Acquire multimodal data; wherein the multimodal data includes one or more of the following: geographic environment data, user scenario data, terminal characteristics and status data, and ionospheric index distribution data; The multimodal data is subjected to feature extraction and fusion processing to obtain a fused feature vector; The policy latent vector is obtained by performing a nonlinear transformation on the fused feature vector through one or more fully connected layers of the Actor network. A hierarchical action decision is made on the strategy latent vector to generate a broadcast strategy, which includes enabling or disabling broadcast frequency, compression level, constellation, frequency point weight, and advanced correction numbers. Satellite differential data is broadcast according to the broadcasting strategy described above.

[0017] Fourthly, embodiments of the present invention also provide a satellite differential data broadcasting device, including a memory, a processor, and a program stored in the memory and executable on the processor; when the processor executes the program, it implements the satellite differential data broadcasting method as described in the first aspect.

[0018] Fifthly, embodiments of the present invention also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the satellite differential data broadcasting method as described in the first aspect.

[0019] In a sixth aspect, embodiments of the present invention also provide a computer program product, including computer instructions, which, when executed by a processor, implement the steps in the satellite differential data broadcasting method as described in the first aspect.

[0020] The above-described technical solution of the present invention has at least the following beneficial effects: In this embodiment of the invention, multimodal data is acquired, including one or more of the following: geographic environment data, user scenario data, terminal characteristic and status data, and ionospheric index distribution data. Feature extraction and fusion processing are performed on the multimodal data to obtain a fused feature vector. A nonlinear transformation is applied to the fused feature vector through one or more fully connected layers of an Actor network to obtain a policy latent vector. Hierarchical action decisions are made on the policy latent vector to generate a broadcast strategy, which includes enabling or disabling broadcast frequency, compression level, constellation, frequency point weights, and advanced corrections. Satellite differential data is broadcast according to the broadcast strategy. Thus, the method of this invention can dynamically and intelligently adjust the broadcast strategy based on geographic environment data, user scenario data, terminal characteristic and status data, and / or ionospheric index distribution data, thereby achieving resource optimization, improving positioning performance, and reducing cloud bandwidth and operating costs. Attached Figure Description

[0021] Figure 1 A schematic diagram illustrating the existing methods for broadcasting satellite differential data; Figure 2 A flowchart illustrating the satellite differential data broadcasting method provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the heterogeneous feature fusion encoder provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of the hierarchical action decision-maker provided in an embodiment of the present invention; Figure 5 This is a schematic diagram of the structure of the safety constraint exploration module provided in an embodiment of the present invention; Figure 6 This is a schematic diagram illustrating the structure of the tiered reward calculator provided in an embodiment of the present invention; Figure 7 A schematic diagram illustrating the instant reward calculation provided in an embodiment of the present invention; Figure 8 A schematic diagram illustrating the short-term reward calculation provided in an embodiment of the present invention; Figure 9 A schematic diagram illustrating the long-term reward calculation provided in an embodiment of the present invention; Figure 10 This diagram illustrates a block diagram of a satellite differential data broadcasting system based on reinforcement learning, as provided in an embodiment of this application. Figure 11 A schematic diagram of the module for broadcasting satellite differential data provided in an embodiment of the present invention; Figure 12 This is a schematic diagram of the hardware structure of the satellite differential data broadcasting device provided in an embodiment of the present invention. Detailed Implementation

[0022] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0023] To facilitate understanding of the present invention, the relevant content involved in the present invention will be introduced first.

[0024] See Figure 1 The existing RTK differential data broadcasting method is as follows: the base station uploads differential data to the NTRIP Caster (data stream server) via the Internet, and the mobile station obtains a continuous, real-time data stream from the Caster using mobile networks such as GPRS / 3G / 4G / 5G. The data stream is continuous and real-time, relying on the TCP / IP connection between the mobile station and the server, and is typically broadcast once per second.

[0025] Current RTK differential networks generally adopt a static broadcast architecture, whose core characteristics can be summarized as follows: 1. Fixed broadcast frequency: The typical configuration is 1Hz, which may be reduced to 0.5Hz in some scenarios. Once the terminal selects the configuration, it cannot be changed during service.

[0026] 2. Templated broadcast content: The current approach involves configuring broadcast content to different mounting points. Terminals select one mounting point to obtain the corresponding service. For example, the "five-star, sixteen-frequency" configuration includes observation data from five navigation satellite systems (GPS, BeiDou, GLONASS, Galileo, and Quasi-Zenith Satellite System, QZSS) and sixteen frequency points.

[0027]

[0028] The pain points of existing technologies are as follows: 1. Fixed broadcast mode is inefficient Traditional RTK networks use a fixed 1Hz broadcast frequency, failing to consider that static terminals in industries such as surveying and mapping do not require high-frequency broadcasting. Furthermore, broadcasting all data to all terminals results in significant bandwidth waste, as much of the data is unusable due to hardware limitations. This leads to a bandwidth waste rate of up to 60%. When dealing with millions of online users, the annual cloud internet bandwidth rental costs are enormous.

[0029] 2. Poor environmental adaptability Once the terminal is configured with its broadcast mount point, the broadcast content is fixed. During ionospheric storms, it cannot dynamically enhance the broadcast with high-precision ionospheric corrections (such as regional precise ionospheric grids or gradient vectors), causing the system to be unable to effectively compensate for rapid positioning deviations caused by drastic changes in electron content, thus reducing positioning accuracy. In scenarios with significant elevation differences, such as hilly or mountainous areas, there is significant asymmetry (anisotropy) in the tropospheric path delay between the base station and the rover. The fixed broadcast mode cannot dynamically enhance the broadcast with tropospheric horizontal gradient corrections, resulting in an inability to effectively compensate for this path difference, causing a decrease in positioning accuracy, especially in the elevation direction.

[0030] 3. Terminal resource mismatch Because the broadcast content is fixed, low-performance terminals (such as IoT devices) are forced to receive redundant data, resulting in a surge in processing latency and increased terminal power consumption. For some IoT devices in the field where it is inconvenient to replace batteries, this accelerates battery wear.

[0031] 4. Server-side management is difficult. Existing RTK service systems employ a static broadcast strategy. To meet the diverse needs of different user groups or application scenarios, service providers are forced to operate by tailoring broadcast content on demand, creating new mount points, and gradually abandoning old mount points. This manually driven and configuration-intensive management model not only results in slow response times and high operating costs, but also leads to increasingly bloated and fragmented server-side mount point resources. Management difficulty and maintenance costs rise sharply with business growth.

[0032] As mentioned above, existing RTK differential data broadcasting technology suffers from inherent drawbacks such as inefficiency due to fixed broadcasting modes, poor environmental adaptability, mismatch of terminal resources, and high management difficulty for service providers. The inventors discovered that the root cause lies in the fact that existing broadcasting systems are open-loop, static systems that cannot perceive changes in terminal status and dynamic fluctuations in the environment, thus failing to make intelligent decisions.

[0033] To address the aforementioned technical problems, this invention provides a method, apparatus, device, medium, and program product for broadcasting satellite differential data. The method and apparatus are based on the same concept, and since the principles by which they solve the problems are similar, their implementations can be mutually referenced; repeated details will not be elaborated further.

[0034] like Figure 2 The diagram shown is a flowchart illustrating a satellite differential data broadcasting method provided in an embodiment of the present invention. This method may include: Step 201: Obtain multimodal data; wherein the multimodal data includes one or more of the following: geographic environment data, user scenario data, terminal characteristics and status data, and ionospheric index distribution data; The geographic environment data is collected in the form of a 10-meter resolution digital elevation model (DEM) and building outline raster maps, with dimensions of 64×64×2. This preserves the characteristics associated with spatial occlusion and multipath effects. The geographic environment data describes the physical terrain and building layout around the terminal and is crucial for determining satellite signal occlusion and multipath effects.

[0035] User scenario data describes the types of terminal applications (such as static mapping, drones, autonomous driving, and agricultural IoT), which have very different requirements for positioning accuracy, latency, and power consumption.

[0036] Terminal characteristics and status data is a hybrid vector that includes terminal hardware capabilities (such as supported constellations and frequencies, represented by binary masks) and real-time dynamic information (such as battery percentage, current positioning mode, convergence time, and residual RMS).

[0037] Ionospheric index distribution data reflects the main propagation errors caused by current space weather to satellite signals. It can be collected in the form of 5×5 regional TEC grid and gradient vector to capture the error impact of space weather on satellite signals and provide a complete environmental and terminal profile for subsequent decision-making.

[0038] Step 202: Perform feature extraction and fusion processing on the multimodal data to obtain a fused feature vector; It should be noted that in RTK high-precision positioning systems, the environmental state information driving intelligent decision-making is inherently multimodal, heterogeneous, and structurally complex. Traditional feature stitching methods cannot effectively capture the deep correlations between different modalities, leading to insufficient information utilization. Specific challenges are as follows: 1. Significant differences in data structures: Geographic environment: Represented as spatially distributed raster data (DEM + building outline, 64×64×2), with local correlation and translation invariance.

[0039] Ionospheric index: It is represented by time-varying sequence data (TEC grid and gradient, 30-dimensional 5×5) and has strong time-series dependence.

[0040] User scenarios and terminal status are represented as structured vector data, including categorical and continuous numerical values.

[0041] 2. Complex cross-modal relationships: There are nonlinear interactions between information from different modes. For example, the geographical features of an "urban canyon" can amplify the impact of the "ionospheric gradient" on positioning accuracy, and this coupling relationship is dynamic and complex.

[0042] 3. Dynamic changes in feature contribution: The importance of various features to the final decision is not constant in different scenarios. For example, in open areas, ionospheric activity is the main source of error; while in urban canyons, the multipath effect dominates.

[0043] To address the aforementioned problems and challenges, this invention utilizes a hierarchical heterogeneous feature fusion encoder (i.e., a multimodal feature fusion module based on the Transformer architecture) to extract and fuse features from multimodal data, obtaining a fused feature vector. Specifically, multimodal data can be processed using a divide-and-conquer approach, where dedicated neural network branches first process the data types they are best at (multimodal data includes multiple data types), and finally, a fusion mechanism integrates global information. This enables the learning of complex cross-modal relationships.

[0044] Step 203: The fused feature vector is nonlinearly transformed through one or more fully connected layers of the Actor network to obtain the policy latent vector; Here, the Actor network is the brain of the system, responsible for generating specific broadcasting strategies. One or more fully connected layers of the Actor network (e.g., from 128 dimensions to 64 dimensions) serve as the basic part of the Actor network, performing nonlinear transformations on the fused feature vectors and ultimately outputting a policy latent vector.

[0045] Step 204: Perform hierarchical action decision-making on the strategy latent vector to generate a broadcast strategy. The broadcast strategy includes enabling or disabling broadcast frequency, compression level, constellation, frequency point weight, and advanced correction numbers. It should be noted that in an adaptive RTK differential data broadcasting system, the action parameters that the agent needs to decide on have inherent heterogeneity in terms of data type: Continuous actions: such as broadcast frequency (1.0Hz ~ 10.0Hz) and data compression level (0.3 ~ 1.0). These actions require fine adjustment within a continuous numerical range.

[0046] Discrete actions: such as constellation and frequency weighting (e.g., prioritizing broadcasting GPS L1 / L2, BDS B1 / B2, or Galileo E1 / E5a). These actions are either / or category choices.

[0047] Boolean actions: such as enabling or disabling advanced corrections (e.g., whether to broadcast tropospheric gradient corrections, whether to enable multipath suppression parameters). These are binary on / off decisions.

[0048] Training instability: The discreteness of discrete actions and Boolean actions can disrupt the gradient flow during backpropagation.

[0049] Inefficient exploration: The agent cannot explore effectively in a way that is consistent with the nature of the data type (for example, for constellation selection, it needs to "jump" between different discrete options rather than "fine-tune" on continuous values).

[0050] Policy performance is limited: It cannot accurately express hard decision logic such as "in a specific scenario, a certain correction should be absolutely enabled".

[0051] To address the aforementioned problems and challenges, this invention integrates a hierarchical action decision-maker at the end of the Actor network to perform hierarchical action decisions on the policy latent vectors and generate broadcast strategies. In other words, it decodes the policy latent vectors into actions at different time scales that conform to the characteristics of RTK broadcasting.

[0052] For example, generating fast-changing layer continuous motion, that is, generating motion parameters that require high-frequency adjustment (adjustable at each step), including broadcast frequency (1.0Hz ~ 10.0Hz) and compression level (0.3 ~ 1.0).

[0053] For example, generating slow-varying layer discrete actions, i.e. generating constellation and frequency point weights, can lead to unstable terminal solutions if such parameters change frequently. However, the above decision processing can be used to periodically update the system based on long-term statistical performance, thus ensuring the stability of the system.

[0054] For example, it generates trigger layer Boolean actions, namely enabling or disabling advanced corrections. Here, advanced correction data can include tropospheric gradient corrections, multipath suppression parameters, ionospheric gradient vectors, and inter-frequency bias. The trigger layer connects to a rule daemon that activates the broadcasting of these high-overhead, high-efficiency parameters only when the network output "uncertainty" or a specific indicator (such as the ionospheric gradient) is triggered, thus achieving "on-demand enhancement."

[0055] Through the aforementioned hierarchical action decision-making, the system can learn the optimal strategy for collaboratively optimizing all broadcast parameters end-to-end, significantly improving the system's adaptability and overall performance in complex scenarios.

[0056] Step 205: Broadcast the satellite differential data according to the broadcasting strategy.

[0057] Specifically, the corresponding observation values ​​are selected according to constellation and frequency weights, advanced corrections are activated as needed, and the above satellite differential data are output according to broadcast frequency and compression level, and then broadcast to the mobile station via the Internet to finally complete the data distribution.

[0058] The satellite differential data broadcasting method of this invention acquires multimodal data, which includes one or more of the following: geographic environment data, user scenario data, terminal characteristic and status data, and ionospheric index distribution data. Feature extraction and fusion processing are performed on the multimodal data to obtain a fused feature vector. A nonlinear transformation is applied to the fused feature vector through one or more fully connected layers of an Actor network to obtain a policy latent vector. Layered action decisions are made on the policy latent vector to generate a broadcasting strategy, which includes enabling or disabling broadcasting frequency, compression level, constellation, frequency point weights, and advanced corrections. Satellite differential data is broadcast according to the broadcasting strategy. Thus, the method of this invention can dynamically and intelligently adjust the broadcasting strategy based on geographic environment data, user scenario data, terminal characteristic and status data, and / or ionospheric index distribution data, thereby achieving resource optimization, improving positioning performance, and reducing cloud bandwidth and operating costs.

[0059] See Figure 3 As an optional implementation, step 202, which involves feature extraction and fusion processing of the multimodal data to obtain a fused feature vector, may specifically include: Step 2021: Extract features from the geographic environment data using a geographic convolutional neural network to obtain a first feature vector; As mentioned above, a hierarchical heterogeneous feature fusion encoder is used to extract and fuse features from multimodal data to obtain a fused feature vector. Specifically, the heterogeneous feature fusion encoder can include three parallel branches: a geographic convolutional neural network (geo-convolutional branch), a temporal coding neural network (temporal coding branch), and a structured data neural network (structured data branch).

[0060] The geographic convolutional neural network is a lightweight convolutional neural network that processes rasterized geographic environmental data (DEM + building raster) to obtain the first feature vector. It should be noted that the convolutional kernel of the geographic convolutional neural network can automatically extract key spatial occlusion features such as "urban canyons" and "open areas", transforming the original 64x64x2 image into a compact feature vector (i.e., the first feature vector).

[0061] See Figure 3For example, a geographic convolutional neural network includes a convolutional layer with a kernel size of 5×5, 16 output channels, and a stride of 2, connected to a ReLU activation function; followed by a convolutional layer with a kernel size of 3×3, 32 output channels, and a stride of 1, also connected to a ReLU activation function; then a max pooling layer with a window size of 2×2; followed by a convolutional layer with a kernel size of 3×3, 64 output channels, and a stride of 1, also connected to a ReLU activation function; and finally, a global average pooling layer. The geographic environment data processed by this geographic convolutional neural network outputs a 64-dimensional feature vector.

[0062] Step 2022: Extract features from the ionospheric index distribution data using a time-series coding neural network to obtain a second feature vector; Optionally, the temporal coding neural network can be a Long Short-Term Memory (LSTM) network or a Gated Recurrent Unit (GRU) encoder. The temporal coding neural network processes ionospheric TEC data and its historical sequences. It can learn short-term variation patterns of ionospheric perturbations, enabling predictive perception of ionospheric errors rather than simply reactive responses.

[0063] See Figure 3 For example, a temporal coding neural network includes an LSTM / GRU layer (128 units), followed by taking the hidden state of the last time step (i.e., extracting key information from the output of the LSTM / GRU layer); finally, a fully connected layer is connected (mapping the extracted key information into a 64-dimensional feature vector). The ionospheric exponential distribution data processed by this temporal coding neural network outputs a 64-dimensional feature vector.

[0064] Step 2023: Extract features from the user scenario data and the terminal characteristics and status data using a structured data neural network to obtain a third feature vector; See Figure 3 As an optional implementation, a third feature vector is obtained by extracting features from the user scenario data and the terminal characteristics and status data using a structured data neural network. This feature vector may specifically include: The user scenario data is encoded using the encoding unit of a structured data neural network to obtain a low-dimensional and sparse binary vector. Here, the user scenario data is encoded using One-Hot encoding to generate a low-dimensional and sparse binary vector that clearly represents the top-level business requirements.

[0065] The terminal characteristics and status data are numerically normalized using the normalization unit of the structured data neural network to obtain a hybrid vector. Here, the terminal characteristic and status data (these continuous or binary values) are normalized to ensure numerical stability. The hybrid vector defines the terminal's capability boundaries and current health.

[0066] The low-dimensional and sparse binary vector and the hybrid vector are input into the fully connected layer of the structured data neural network to obtain the third feature vector.

[0067] Here, the structured data neural network uses fully connected layers to process structured vector data such as user scenario data, terminal characteristics and state data, and can learn the mapping relationship between them and the optimal policy.

[0068] The following table illustrates the processing methods for different modal data and their encoding functions.

[0069]

[0070]

[0071] Step 2024: Perform feature fusion processing based on the first feature vector, the second feature vector, and / or the third feature vector to obtain the fused feature vector.

[0072] See Figure 3 As an optional implementation, feature fusion processing is performed based on the first feature vector, the second feature vector, and / or the third feature vector to obtain the fused feature vector, which may specifically include: A linear transformation is performed on the first feature vector to obtain a first query matrix, a first key matrix, and a first value matrix; a first attention weight is calculated based on the first query matrix and the first key matrix; and a weighted fusion is performed on the first value matrix based on the first attention weight to obtain a first output result. A linear transformation is performed on the second feature vector to obtain a second query matrix, a second key matrix, and a second value matrix; based on the second query matrix and the second key matrix, a second attention weight is calculated; and based on the second attention weight, the second value matrix is ​​weighted and fused to obtain a second output result. A linear transformation is performed on the third feature vector to obtain a third query matrix, a third key matrix, and a third value matrix; based on the third query matrix and the third key matrix, a third attention weight is calculated; and based on the third attention weight, the third value matrix is ​​weighted and fused to obtain a third output result. It should be noted that by performing linear transformations on the feature vectors of each of the above branches (geographic convolutional neural network, temporal coding neural network, and structured data neural network) to generate the corresponding query matrix Q, key matrix K, and value matrix V, the features of each modality can be projected into a common semantic space.

[0073] Next, using a scaled dot product attention mechanism, the attention weights of each modality's features relative to other modalities are calculated. Specifically, these attention weights are calculated based on the query matrix Q and the key matrix K. These attention weights quantify the correlation strength between geographic and ionospheric information within the current context. Here, multiple attention heads are computed in parallel to capture cross-relationships in different subspaces, enhancing the model's expressive power.

[0074] Then, the calculated attention weights are used to weight and fuse the respective value matrices V to calculate the output results (i.e., the first output result, the second output result, and / or the second output result).

[0075] The first output result, the second output result, and / or the third output result are concatenated and then subjected to a linear transformation to obtain the fused feature vector.

[0076] Here, after processing by the hierarchical heterogeneous feature fusion encoder described above, a 128-dimensional fused feature vector is finally output. This fused feature vector serves as a shared input for the Actor network and the Critic network, providing a unified and information-rich state basis for subsequent decisions.

[0077] It's important to understand that the Actor network and Critic network are the two core components of the Actor-Critic architecture in reinforcement learning, responsible for decision-making and evaluation, respectively. The Actor network is responsible for selecting actions based on the current environment state, outputting either a probability distribution of the action or a specific action value. It acts as the "executor," continuously optimizing the policy through policy gradient updates to maximize long-term rewards. The Critic network is responsible for evaluating the quality of the actions selected by the Actor network, outputting a value estimate of the current state or state-action pairs. It acts as the "judge," guiding the Actor network's update direction and improving learning stability through temporal difference (TD) error. When the Actor network executes an action, the environment returns a reward and a new state; the Critic network calculates the value based on the new state, deriving the TD error (the difference between the predicted and actual values). This error serves as a feedback signal, used to adjust the Actor network's policy (e.g., increasing the probability of good actions and reducing the selection of bad actions); simultaneously, the Critic network itself updates its parameters based on the error, improving evaluation accuracy.

[0078] This invention introduces a multimodal feature fusion module based on the Transformer architecture (i.e., a hierarchical heterogeneous feature fusion encoder). This module achieves deep cross-inference of multi-source information such as geographical environment, ionospheric activity, and terminal status through a self-attention mechanism. Unlike traditional simple concatenation, this module can dynamically learn the nonlinear correlations between features in different contexts, achieving intelligent weight allocation based on global context. For example, when the system detects the simultaneous occurrence of "urban canyon" and "drastic ionospheric gradient change," it automatically strengthens the interaction weights of these two types of features, generating a fusion state representation pointing to "enabling a strong correction strategy."

[0079] This relational reasoning capability enables the system to understand complex application scenarios, rather than mechanically responding to a single metric, thereby achieving a more refined balance between bandwidth, accuracy, and energy consumption while ensuring service quality. Context-aware fusion: Unlike fixed-weight fusion methods, attention mechanisms can dynamically adjust the contribution of each modality feature. For example, when the system identifies an urban canyon, it will automatically increase the weight of "geographical features" and "multipath suppression parameters" in the decision-making process.

[0080] Relational reasoning ability: The encoder can learn complex cross-modal relationships, such as "in hilly terrain, the north-south ionospheric gradient has a more significant impact on positioning error".

[0081] Lossless information condensation: The original data is processed through a dedicated branch, avoiding information loss in the early fusion. Then, selective condensation is performed through an attention mechanism. The final output state vector provides SAC decision-making with input that has extremely high information density and is rich in semantics.

[0082] In some embodiments, step 204 above, which involves performing hierarchical action decision-making on the policy latent vector to generate a broadcasting policy, includes: Gaussian distribution sampling is performed on the latent vectors of the strategy to obtain the broadcast frequency and compression level; It should be noted that the aforementioned Actor network, with one or more fully connected layers (e.g., 128-dimensional → 64-dimensional) serving as its foundation, performs a nonlinear transformation on the fused feature vector, ultimately outputting a policy latent vector. To address the core challenge of heterogeneous action parameters in RTK differential data broadcasting strategies, this invention integrates a hierarchical action decision-maker at the end of the Actor network to perform hierarchical action decisions on the policy latent vector, generating a broadcasting strategy. For details, see [link to details]. Figure 4 By integrating a hybrid action space output head at the end of the Actor network, unified and parallel policy modeling and gradient calculation can be achieved for three types of actions (including continuous actions, discrete actions, and Boolean actions).

[0083] The hybrid action space output head includes a continuous action head, a discrete action head, and a Boolean action head. Here, the continuous action head is used to sample the policy latent vector using a Gaussian distribution to obtain the broadcast frequency and compression level. Specifically, a fully connected layer maps the policy latent vector to the mean (μ) and log-standard deviation (logσ) of each continuous action (i.e., Gaussian distribution parameterization); then, a re-parameterization technique is used to extract the Gaussian distribution N(μ, exp(logσ)). 2 The sampled values ​​are obtained from the broadcast frequency (e.g., 1.5Hz) and compression level. This technique ensures that the gradient can be backpropagated through random nodes. It should be noted that the sampled values ​​can be compressed to [-1, 1] using the tanh activation function, and then linearly mapped to the actual action range [1, 10]Hz.

[0084] Gumbel-Softmax sampling is performed on the latent vector of the strategy to obtain constellation and frequency point weights; Here, the policy latent vector can be sampled using Gumbel-Softmax through the discrete action head to obtain the constellation and frequency weights. Specifically, a fully connected layer outputs the logits (unnormalized log probabilities) of all possible options for each discrete action; then, the Gumbel-Softmax reparameterization technique is used for sampling to obtain the constellation and frequency weights. The principle is: Gumbel noise is added to the logits, and then the probability distribution is calculated using the Softmax function. By introducing a temperature coefficient (τ), the smoothness of the distribution is controlled: in the early stage of training, τ is larger, and the distribution is smooth to encourage exploration; in the later stage of training, τ approaches 0, and the distribution approaches one-hot to form deterministic decisions.

[0085] It should be noted that the Gumbel-Softmax reparameterization technique sampling provides a differentiable approximation, which allows gradients to be computed for actions sampled from discrete distributions, thus supporting end-to-end training.

[0086] The latent vectors of the policy are sampled using a Bernoulli distribution and processed with an activation function to determine whether advanced corrections are enabled or disabled.

[0087] Here, the activation or deactivation of the advanced correction can be determined by sampling the policy latent vector using a Bernoulli distribution and applying an activation function through a Boolean action head. Specifically, a fully connected layer outputs a Logit value for each Boolean action; then, it is processed using Bernoulli distribution sampling and a Sigmoid activation function. Specifically, the Logit is treated as a probability parameter p of a Bernoulli distribution, and mapped to the [0, 1] interval using the Sigmoid function. By comparing the Sigmoid(logit) with a threshold sampled from a uniform distribution U(0,1), the output is determined to be True (enabled) or False (disabled). This process can also achieve gradient calculation through reparameterization techniques.

[0088] It should be noted that although the sampling methods for the three action heads differ, their policy gradients can all be uniformly calculated within the SAC framework during training. Policy gradient estimation follows a variation of the following formula:

[0089] For continuous actions, It is the logarithmic probability of a Gaussian distribution.

[0090] For discrete actions and Boolean actions is the log probability of the Gumbel-Softmax or Bernoulli distribution. Q(s,a) is the Critic network's evaluation of the state-action pair, and b(s) is the baseline function (in SAC, it is undertaken by the value function V(s) or another Critic).

[0091] To address the core challenge of heterogeneous action parameters in RTK differential data broadcasting strategies, this invention provides a reinforcement learning decision module for hybrid action spaces. This module designs parallel multi-type action output heads at the end of the Actor network, innovatively integrating Gaussian distribution, Gumbel-Softmax reparameterization, and Bernoulli distribution modeling. This enables unified policy representation and differentiable sampling for continuous (broadcasting frequency, compression level), discrete (constellation and frequency weights), and Boolean (advanced correction number switching) actions. This design solves the gradient breakage and inefficient exploration problems inherent in traditional methods when handling hybrid action spaces, allowing the agent to learn the optimal policy for co-optimizing all broadcasting parameters end-to-end, significantly improving the system's adaptability and overall performance in complex scenarios.

[0092] In online RTK services, the exploratory behavior of reinforcement learning agents must be strictly limited to avoid catastrophic consequences. Traditional optimistic exploration fundamentally contradicts the high reliability requirements of real-world systems. The challenge lies in: Catastrophic actions: The agent may discover strategies that lead to service disruption or resource exhaustion.

[0093] Training instability: Individual negative experiences may over-impact the strategy, leading to training oscillations.

[0094] The trade-off between safety and exploration: How to fully explore to find the optimal strategy while absolutely ensuring that the system operates within the safety boundaries.

[0095] To address the aforementioned challenges, after processing the policy hidden variables through the hierarchical action decision-maker and generating the broadcast strategy, it is necessary to limit the corresponding action parameters (such as broadcast frequency, compression level, constellation and frequency weights, etc.) in the broadcast strategy to a safe range. This can be achieved through the following optional methods.

[0096] As an optional implementation, the method of the present invention further includes: 1) Perform action masking on the broadcast strategy to obtain compliant actions; Here, the present invention designs a safe exploration constraint module to limit the corresponding action parameters in the broadcast strategy within a safe range. See also Figure 5 Specifically, it includes three levels of constraints: Level 1: hardware physical constraints through action masks; Level 2: resource boundary protection through budget constraint verification; and Level 3: service security boundary protection through pessimistic Q-value constraints and security action sets.

[0097] This step belongs to the first-level action masking process. As an optional implementation, the broadcast strategy is processed by action masking to obtain compliant actions, including: Based on the terminal capabilities, the broadcasting strategy is masked to obtain a broadcasting strategy that matches the terminal capabilities. Here, discrete actions (constellation and frequency weights) / Boolean actions (advanced correction number switches) and terminal capability vectors are input. Then, based on the terminal capabilities, the discrete / Boolean actions in the broadcast strategy are masked. Specifically, the probabilities of actions not supported by the terminal hardware are set to zero, generating a binary mask matrix, i.e., a_masked = a_original ⊙ M_capability. For example, if the terminal is a single-frequency receiver, all L2 / L5 frequency-related actions are masked.

[0098] Based on business rules, the broadcasting strategy that matches the terminal capabilities is masked to obtain the compliant action.

[0099] Next, based on business rules, broadcast strategies matching terminal capabilities, such as continuous actions (broadcast frequency, compression level) and user scenario tags, are masked to obtain compliant actions. Specifically, action boundaries are set for different business scenarios; for example, mapping users: frequency ∈ [0.5, 1] ​​Hz, autonomous driving users: frequency ∈ [1, 5] Hz. The specific implementation is as follows: a_bounded=clamp(a_masked,a_min(scenario),a_max(scenario)).

[0100] It should be noted that this level is based on prior knowledge of the system, establishing inviolable physical and logical constraints.

[0101] 2) Perform budget constraint verification on the compliant actions to obtain budget-safe actions; As an optional implementation, the budget constraint verification of the compliant action to obtain a budget-safe action includes: The resource consumption prediction model is used to obtain the predicted resources consumed by the compliance action. This implementation corresponds to level two: verifying resource boundaries through budget constraints. The resource consumption prediction model is a pre-built lightweight neural network C(a)=[bandwidth,computation,power] to accurately predict the resource cost of executing action a (i.e., the compliant action), that is, to predict the resources consumed by the compliant action.

[0102] Get the currently consumed resources; Here, real-time budget monitoring can be used to dynamically track the currently consumed resources, B_current(t)=[bandwidth_used,computation_used,…].

[0103] If the sum of the predicted resources consumed by the compliance action and the currently consumed resources is greater than a preset threshold, the compliance action is corrected to obtain the budgeted security action, wherein the resources consumed by the budgeted security action are less than or equal to the difference between the predicted resources consumed by the compliance action and the currently consumed resources. Here, if C(a) + B_current(t) > B_threshold (preset threshold), then the compliant action is corrected. Specifically, gradient descent optimization is initiated, which means finding a corrected action that minimizes the difference between it and the original action to be executed (i.e., the compliant action a_input) (minimizing the squared difference). In other words, a_corrected is the budget safety action. There needs to be a constraint here: the resources consumed by the budget safety action are less than or equal to the difference between the predicted resources consumed by the compliant action and the currently consumed resources, i.e., stC(a) ≤ B_threshold - B_current.

[0104] It should be noted that in practice, a fast heuristic algorithm is used, such as degrading action components by priority.

[0105] If the sum of the predicted resources consumed by the compliance action and the currently consumed resources is less than or equal to the preset threshold, the compliance action is determined to be the budget security action.

[0106] Level 2 ensures that system resource usage does not exceed the operating budget, representing an economic-level security constraint.

[0107] 3) Based on the budgeted safety actions, the final safety actions are obtained through pessimistic Q-value initialization and safety constraint processing of the safety action set, and the final safety actions are determined as the final broadcast strategy; As an optional implementation, based on the budgeted safety actions, the final safety actions are obtained through pessimistic Q-value initialization and safety constraint processing of the safety action set, including: The Q-value output of the Crtic network is initialized to a constant that is much lower than the normal reward range; This implementation corresponds to level three: pessimistic Q-value constraints and a set of safe actions, in order to protect the service security boundary.

[0108] It should be noted that before training begins, the Q-value output of the Critical network is initialized to a constant Q_low that is much lower than the normal reward range. For example, Q_low = -10 is set, which is much lower than the normal reward range [-1, +1].

[0109] Psychological analogy: If an intelligent agent is made to believe that "all unknown things are dangerous", the effect is that when exploring unknown states-action pairs, the agent will expect very low rewards, and therefore will tend to quickly return to the known safe zone, or only try small, incremental explorations.

[0110] Determine whether the budgeted safety action belongs to the set of safety actions; Here, a set of safe actions, A_safe(s), can be predefined. It is a set of state-related actions that have been verified not to cause a serious degradation in system performance. The state mentioned here refers to multimodal data.

[0111] Initialization of the safety set: Extract state-action pairs from historical log data, and filter out records whose key indicators such as positioning accuracy and system load are within the normal range after execution to form the initial safety set.

[0112] Online updates for the security set: During online operation, the effect of each actual action 'a' is monitored in real time.

[0113] Define the safety metrics: S(a) = f(positioning accuracy, convergence time, resource utilization, ...).

[0114] If S(a) > S_threshold for a period of time, then a is considered safe in the current state s and is added to A_safe(s).

[0115] If an action that was originally in the safe set causes a performance violation, it is removed from A_safe(s).

[0116] For each action a_proposed output by the Actor network, it is first filtered by the first two levels (level one and level two), and then it is checked whether the filtered action a_filtered (i.e. budget safe action) belongs to the safe action set A_safe(s) under the current state s.

[0117] If it is, then the budget security action is determined to be the final security action, and the Q value is updated after the final security action is executed; If a_filtered (i.e., budget-safe action) ∈ A_safe(s) (set of safe actions): allow the execution of that budget-safe action (as the final safe action) and update the Q-value with the actual reward obtained. Since the initial Q-value is pessimistic, the actual reward (even a small positive reward) will bring a strong positive surprise, quickly boosting the Q-value of the state-action pair.

[0118] If it does not belong to the set of safety actions, then select the safety alternative action that is most similar to the budgeted safety action from the set of safety actions, and determine the safety alternative action as the final safety action.

[0119] If a_filtered (i.e., a budget-safe action) A_safe(s) (safe action set): Refuse to execute the original action (i.e. refuse to execute the budgeted safe action), and instead select a "safe alternative action" a_safe from A_safe(s) that is most similar to a_filtered (i.e. the budgeted safe action) to execute.

[0120] Here, the exploration incentive is to provide a small additional reward for exploratory actions that successfully add new actions to the safe set, in order to encourage agents to expand the safe set.

[0121] Accordingly, step 205 above, broadcasting satellite differential data according to the broadcasting strategy, includes: broadcasting satellite differential data according to the final broadcasting strategy.

[0122] To address the security risks associated with reinforcement learning exploration in RTK online services, this invention proposes a dual security constraint mechanism based on pessimistic Q-value initialization and a dynamic set of safe actions. This mechanism initializes the Critic network with pessimistic value estimation, guiding the agent to exercise caution in the unknown domain. Simultaneously, it constructs and dynamically maintains a state-dependent set of safe actions, strictly limiting the agent's exploration behavior within verified safety boundaries.

[0123] This design realizes an intelligent security exploration paradigm of "initially conservative, validated and advanced, and rule-breaking and fallback," which fundamentally avoids the emergence of catastrophic strategies and provides a solid security foundation for the practical deployment of reinforcement learning in mission-critical systems.

[0124] Safe and incremental exploration: Intelligent agents will not conduct blind and high-risk large-scale exploration, but will "test" the boundaries of the safe set and steadily expand the safe area.

[0125] Training stability: Pessimistic initialization avoids the destructive impact of negative data generated by random strategies on the model in the early stages of training.

[0126] Feasibility of online learning: Even when conducting online learning in a real production environment, this mechanism ensures that the quality of service will not fall below the safety baseline.

[0127] Theoretical guarantee: This scheme conforms to the theoretical framework of pessimistic reinforcement learning and has a good convergence guarantee.

[0128] It should be noted that in RTK adaptive broadcasting systems, the optimization objective is inherently multi-timescale and multi-dimensional. A single reward signal cannot accurately express this complex optimization requirement. 1. Time scale conflict: Immediate effect: Reducing the broadcast frequency will immediately save bandwidth, but may slightly affect instantaneous positioning accuracy.

[0129] Short-term effects: Enabling advanced corrections will immediately increase bandwidth overhead, but may improve positioning accuracy and fixation rate within minutes.

[0130] Long-term effects: Continuously providing high-quality service can improve user stickiness and satisfaction, but the short-term costs are high.

[0131] 2. Conflict in target dimensions: Bandwidth cost (operator benefits) vs. positioning accuracy (user benefits) vs. terminal power consumption (user experience).

[0132] Traditional single scalar rewards cannot distinguish the effects of different time scales and dimensions, leading agents to learn short-sighted, locally optimal strategies, such as indiscriminately reducing bandwidth while ignoring long-term service quality. To address this problem, this invention designs a hierarchical reward calculator. Instead of calculating a single reward value within a decision cycle, it calculates a set of reward components evaluating the policy's effectiveness at different time scales, and finally synthesizes the final training signal through a weighted mechanism. Specifically, as an optional implementation, the method of this invention further includes: Obtain multi-dimensional feedback data after the broadcast strategy is executed; See Figure 6 Multi-dimensional feedback data after the broadcast strategy is executed can be obtained through environmental interaction. Here, environmental interaction includes RTK timing engine, network bandwidth monitoring, and terminal status feedback.

[0133] Based on the multi-dimensional feedback data, calculate the corresponding reward value for each; Here, a tiered reward calculator can be used based on multi-dimensional feedback data (see...). Figure 6 Each tiered reward system calculates its corresponding reward value. The tiered rewards specifically include immediate rewards, short-term rewards, and long-term rewards.

[0134] The immediate reward is calculated right after the action is executed, primarily reflecting the instantaneous impact of the strategy on system resources. For details, see [link to details]. Figure 7 The instant reward R_instant includes broadband efficiency reward, terminal power consumption reward, and computational overhead reward.

[0135] The bandwidth efficiency bonus is calculated as R_bandwidth = -λ_bw * (Data_volume / Data_budget), where λ_bw represents the bandwidth cost weighting coefficient, Data_volume represents the amount of data generated in this broadcast, and Data_budget represents the preset bandwidth budget. The purpose of the bandwidth efficiency bonus is to encourage data compression and reduce unnecessary broadcast frequencies.

[0136] The terminal power consumption reward R_power = -λ_power * (P_estimated / P_max); where λ_power represents the terminal power consumption weighting coefficient; P_estimated represents the terminal processing power consumption estimated based on actions (such as frequency and data volume); and P_max represents the terminal's maximum power consumption threshold. The purpose of the terminal power consumption reward is to protect the terminal's battery power, which is especially crucial for IoT devices.

[0137] The computational overhead reward R_cpu = -λ_cpu * (Server_Load / Load_max); where λ_cpu represents the computational overhead weighting coefficient; Server_Load represents the server load value; and Load_max represents the server's maximum load value. The purpose of the computational overhead reward is to prevent server overload and ensure the overall stability of the system.

[0138] Subsequently, based on the broadband efficiency reward, terminal power consumption reward, and computational overhead reward, an instant reward is synthesized to obtain the instant reward R_instant.

[0139] The short-term reward is obtained several time steps (approximately 1-5 minutes) after the action is executed, and its effectiveness depends on observing the output of the localization engine. For details, see [link to documentation]. Figure 8 The short-term reward R_short includes positioning accuracy reward, convergence time reward, and fixation rate reward.

[0140] in: The positioning accuracy reward R_accuracy = λ_accuracy * exp(-position_error / error_threshold); where λ_accuracy represents the positioning accuracy weighting coefficient; position_error represents the positioning error reported by the terminal; and error_threshold represents the positioning error threshold. The positioning accuracy reward is the core reward, incentivizing high-precision positioning in an exponential manner.

[0141] The convergence time reward R_convergence = -λ_converge * (T_converge / T_max); where λ_converge represents the convergence time weight coefficient; T_converge represents the time required to re-fix the solution after the lock is lost; and T_max represents the maximum convergence time. The purpose of the convergence time reward is to encourage the strategy to quickly help the terminal recover high-precision positioning.

[0142] The fixed rate reward R_fix = λ_fix * (Fix_Rate); where λ_fix represents the fixed rate weighting coefficient; and Fix_Rate represents the proportion of fixed solutions over a period of time. The purpose of the fixed rate reward is to improve the reliability and continuity of positioning.

[0143] Subsequently, based on the positioning accuracy reward, convergence time reward, and fixation rate reward, a short-term reward is synthesized to obtain the short-term reward R_short.

[0144] The long-term reward is calculated periodically (e.g., hourly / daily) to reflect the macro and long-term effects of the strategy. See details below. Figure 9The long-term reward R_long includes service stability rewards, user satisfaction rewards, and system health rewards.

[0145] The service stability reward R_stability = λ_stability * (Uptime_ratio - Downtime_penalty); where λ_stability represents the service stability weight coefficient, Uptime_ratio represents the service uptime rate, and Downtime_penalty represents the downtime penalty. The purpose of the service stability reward is to encourage the provision of continuously available services.

[0146] The user satisfaction reward R_satisfaction = λ_satisfaction * (User_retention_rate + Positive_feedback); where λ_satisfaction represents the user satisfaction weight coefficient, User_retention_rate represents the user retention rate, and Positive_feedback represents positive feedback. The purpose of the user satisfaction reward is to integrate business goals into the optimization process and encourage improvements in user experience.

[0147] The system health reward R_health = λ_health * (1 - Load_imbalance_metric); where λ_health represents the system health weighting coefficient, and Load_imbalance_metric represents the degree of uneven load distribution. The purpose of the system health reward is to encourage load balancing and prevent a single server or link from becoming a bottleneck.

[0148] Subsequently, based on service stability rewards, user satisfaction rewards, and system health rewards, long-term rewards are synthesized to obtain the long-term reward R_long.

[0149] The total reward value is calculated by weighting the values ​​based on their respective reward values ​​using a discount factor. Specifically, based on the immediate reward R_instant, the short-term reward R_short, and the long-term reward R_long, the total reward value R_total(t) is calculated using the following formula by weighting the discount factors: R_total(t) = R_instant(t) + γ_short * R_short(t + Δt) + γ_long * R_long(t + T); where γ_short and γ_long both represent discount factors used to balance the present value of rewards at different time scales.

[0150] It should be noted that the weights [λ_bw, λ_power,…] can be dynamically adjusted through the attention mechanism, for example, the weight of λ_bw can be automatically increased when the network is congested.

[0151] Based on the total reward value, update the Q value of the Critical network; It should be noted that there are two Crtic networks. Each Crtic network includes continuous states and actions, followed by a fully connected layer of 256, then a ReLU activation function, and finally a fully connected layer of 128.

[0152] The policy parameters of the Actor network are updated based on the Q-value gradient, wherein the Q-value gradient is obtained based on the updated Q-value and the Q-value before the update.

[0153] It should be noted that the Critic network using the SAC algorithm learns a value function V(s), which automatically backpropagates and assigns credit to the previously obtained short-term and long-term rewards, assigning them to a series of states and actions that led to this result. This allows the agent to understand that "although enabling advanced corrections increases immediate costs, it brings superior short-term positioning performance, thereby improving long-term user satisfaction."

[0154] To address the inherent time-scale conflicts and short-sighted decision-making problems in multi-objective optimization of RTK broadcasting, this system proposes a hierarchical reward mechanism. This mechanism designs reward functions at three time scales: immediate (resource cost), short-term (location performance), and long-term (service value), and employs discounted weighted synthesis and credit allocation to provide the reinforcement learning agent with a comprehensive and time-aware optimization objective. This hierarchical reward mechanism balances immediate costs and long-term benefits, learning an optimal broadcasting strategy that is both cost-effective and ensures service quality and user satisfaction, fundamentally solving the challenge of multi-objective collaborative optimization. It should be noted that the reward feedback sequence is: action execution (t-1), immediate reward (t0), short-term reward (t+n), and long-term reward (t+T).

[0155] This design enables the agent to understand the causal chain that sacrifices instantaneous bandwidth for improved positioning accuracy, ultimately increasing long-term user satisfaction, thereby learning the optimal broadcasting strategy that is both cost-effective and ensures service quality and system sustainability. Overcoming short-sightedness: The agent will not sacrifice positioning accuracy for immediate bandwidth savings because it can foresee subsequent short-term penalties.

[0156] Balancing multiple objectives: By adjusting the weighting coefficients, the system can flexibly set preferences among objectives such as bandwidth, accuracy, and energy consumption.

[0157] More robust strategies: Strategies that consider long-term stability will avoid "risky" behaviors that may have good short-term performance but could lead to system instability.

[0158] It aligns with business logic: the reward structure is directly linked to real business operation metrics (cost, user satisfaction), making the learned strategies more commercially valuable.

[0159] The method of this invention can be applied to a satellite differential data broadcasting system based on reinforcement learning. A schematic diagram of the system framework can be found in [reference needed]. Figure 10 .

[0160] The system includes a state awareness module, an SAC intelligent decision center, and a strategy execution module.

[0161] Among them, the state perception module is responsible for collecting multi-dimensional information (i.e., the aforementioned multimodal data) from inside and outside the system, as the basis for the agent's decision-making.

[0162] The SAC Intelligent Decision Center is the brain of the system, containing a built-in reinforcement learning agent, including a heterogeneous feature fusion encoder, an Actor network, and a Critic network. Specifically, the heterogeneous feature fusion encoder extracts and fuses features from multimodal data to obtain a fused feature vector. Based on this fused feature vector, the Actor network makes hierarchical action decisions, generating a broadcast strategy (i.e., outputting action parameters such as broadcast frequency, compression level, constellation, frequency weights, and enabling or disabling advanced corrections). The Critic network evaluates the quality of the Actor network's output actions, outputting a value estimate of the current state or state-action pairs.

[0163] The policy execution module is responsible for receiving instructions from the agent and controlling broadcast servers such as NTRIP Caster to execute specific broadcast actions. These actions may include one or more of the following: broadcast frequency, compression level, constellation and frequency weights, tropospheric gradient correction, multipath suppression parameters, ionospheric gradient vector, and inter-frequency bias.

[0164] This invention utilizes the Soft Actor-Critic reinforcement learning algorithm as the core of the system's intelligent decision-making, realizing a dynamic RTK differential data broadcasting method with self-learning and evolutionary capabilities. By establishing a reward mechanism with multi-dimensional optimization objectives of "bandwidth overhead, positioning accuracy, terminal power consumption, and service stability," the system can: Sensing complex terminal states, geographical environments, and space weather; Determine the optimal broadcast strategy (including frequency, content, compression rate, etc.). Execute and collect feedback to continuously optimize yourself.

[0165] Ultimately, the system dynamically and accurately balances the core contradictions between cloud bandwidth costs, user positioning accuracy, and terminal battery life, fundamentally solving the resource mismatch and inefficiency problems caused by the existing fixed broadcasting mode.

[0166] This invention constructs an intelligent broadcasting decision engine based on reinforcement learning. By deeply perceiving and fusing multi-dimensional dynamic environmental states (including user scenarios, geographical environment, ionospheric index distribution, terminal characteristics, and real-time status), it achieves autonomous optimization and dynamic adjustment of differential data broadcasting strategies. Compared to traditional fixed-parameter broadcasting modes, this solution achieves breakthroughs in the following aspects: Significantly improved bandwidth efficiency: By intelligently adjusting the broadcast frequency, content, and compression level, redundant data transmission is effectively eliminated, achieving bandwidth cost savings of 30%-70%.

[0167] Enhanced terminal battery life: The adaptive broadcasting strategy based on terminal status significantly reduces data processing load, bringing 20%-40% energy consumption optimization to restricted terminals such as IoT devices, and effectively extending device working time.

[0168] Improved positioning performance and reliability: In complex environments (such as urban canyons and ionospheric disturbances), it can dynamically enhance the broadcast of correction data, improve positioning accuracy and fixation rate, and ensure service quality.

[0169] The system's intelligence level has been significantly improved: it possesses self-learning and continuous evolution capabilities, fundamentally solving the inherent problems of poor environmental adaptability and rigid resource allocation in traditional systems.

[0170] like Figure 11 As shown, an embodiment of the present invention provides a satellite differential data broadcasting device, the device comprising: The first acquisition module 1101 is used to acquire multimodal data; wherein, the multimodal data includes one or more of the following: geographic environment data, user scenario data, terminal characteristics and status data, and ionospheric index distribution data; The feature extraction and fusion module 1102 is used to perform feature extraction and fusion processing on the multimodal data to obtain a fused feature vector; The first processing module 1103 is used to perform a nonlinear transformation on the fused feature vector through one or more fully connected layers of the Actor network to obtain a policy latent vector; The broadcast strategy generation module 1104 is used to make hierarchical action decisions on the strategy latent vector and generate a broadcast strategy. The broadcast strategy includes enabling or disabling broadcast frequency, compression level, constellation, frequency point weight, and advanced correction numbers. The data broadcasting module 1105 is used to broadcast satellite differential data according to the broadcasting strategy.

[0171] In some embodiments, the feature extraction and fusion module 1102 includes: The first feature extraction unit is used to extract features from the geographic environment data through a geographic convolutional neural network to obtain a first feature vector; The second feature extraction unit is used to extract features from the ionospheric exponential distribution data through a time-series coding neural network to obtain a second feature vector. The third feature extraction unit is used to extract features from the user scenario data and the terminal characteristics and status data through a structured data neural network to obtain a third feature vector; The feature fusion unit is used to perform feature fusion processing based on the first feature vector, the second feature vector, and / or the third feature vector to obtain the fused feature vector.

[0172] In some embodiments, the third feature extraction unit is specifically used for: The user scenario data is encoded using the encoding unit of a structured data neural network to obtain a low-dimensional and sparse binary vector. The terminal characteristics and status data are numerically normalized using the normalization unit of the structured data neural network to obtain a hybrid vector. The low-dimensional and sparse binary vector and the hybrid vector are input into the fully connected layer of the structured data neural network to obtain the third feature vector.

[0173] In some embodiments, the feature fusion unit is specifically used for: A linear transformation is performed on the first feature vector to obtain a first query matrix, a first key matrix, and a first value matrix; a first attention weight is calculated based on the first query matrix and the first key matrix; and a weighted fusion is performed on the first value matrix based on the first attention weight to obtain a first output result. A linear transformation is performed on the second feature vector to obtain a second query matrix, a second key matrix, and a second value matrix; based on the second query matrix and the second key matrix, a second attention weight is calculated; and based on the second attention weight, the second value matrix is ​​weighted and fused to obtain a second output result. A linear transformation is performed on the third feature vector to obtain a third query matrix, a third key matrix, and a third value matrix; based on the third query matrix and the third key matrix, a third attention weight is calculated; and based on the third attention weight, the third value matrix is ​​weighted and fused to obtain a third output result. The first output result, the second output result, and / or the third output result are concatenated and then subjected to a linear transformation to obtain the fused feature vector.

[0174] In some embodiments, the broadcast strategy generation module 1104 includes: The first strategy generation unit is used to perform Gaussian distribution sampling on the strategy latent vector to obtain the broadcast frequency and compression level. The second policy generation unit is used to perform Gumbel-Softmax sampling on the policy latent vector to obtain constellation and frequency point weights. The third policy generation unit is used to perform Bernoulli distribution sampling and activation function processing on the policy latent vector to determine whether the advanced correction number is enabled or disabled.

[0175] In some embodiments, the apparatus of the present invention further includes: The second processing module is used to perform action masking processing on the broadcast strategy to obtain compliant actions; The third processing module is used to perform budget constraint verification on the compliant actions to obtain budget-safe actions. The fourth processing module is used to obtain the final safety action based on the budgeted safety action by initializing the pessimistic Q value and processing the safety constraint of the safety action set, and to determine the final safety action as the final broadcast strategy. Accordingly, the data broadcasting module 1105 includes: The data broadcasting unit is used to broadcast satellite differential data according to the final broadcasting strategy.

[0176] In some embodiments, the second processing module includes: The first processing unit is used to perform masking processing on the broadcasting strategy based on the terminal capabilities to obtain a broadcasting strategy that matches the terminal capabilities. The second processing unit is used to perform masking processing on the broadcasting strategy that matches the terminal capabilities based on business rules, so as to obtain the compliance action.

[0177] In some embodiments, the third processing module includes: The third processing unit is used to obtain the predicted resources consumed by the compliance action using a resource consumption prediction model. The acquisition unit is used to acquire the currently consumed resources; An action correction unit is used to correct the compliant action when the sum of the predicted resources consumed by the compliant action and the currently consumed resources is greater than a preset threshold, thereby obtaining the budgeted safety action, wherein the resources consumed by the budgeted safety action are less than or equal to the difference between the predicted resources consumed by the compliant action and the currently consumed resources. The fourth processing unit is configured to determine the compliance action as the budget security action if the sum of the predicted resources consumed by the compliance action and the currently consumed resources is less than or equal to the preset threshold.

[0178] In some embodiments, the fourth processing module includes: The fifth processing unit is used to initialize the Q-value output of the Crtic network to a constant that is much lower than the normal reward range; The judgment unit is used to determine whether the budget safety action belongs to the set of safety actions; The sixth processing unit is configured to determine the budget safety action as the final safety action when the budget safety action belongs to the set of safety actions, and update the Q value after executing the final safety action; The seventh processing unit is used to select the most similar safety alternative action to the budgeted safety action from the safety action set when the budgeted safety action does not belong to the safety action set, and to determine the safety alternative action as the final safety action.

[0179] In some embodiments, the apparatus of the present invention includes: The second acquisition module is used to acquire multi-dimensional feedback data after the broadcasting strategy is executed; The first calculation module is used to calculate the corresponding reward value based on the multi-dimensional feedback data; The second calculation module is used to calculate the total reward value by weighting the corresponding reward value using a discount factor. The fifth processing module is used to update the Q value of the Crtic network based on the total reward value; The sixth processing module is used to update the policy parameters of the Actor network based on the Q-value gradient, wherein the Q-value gradient is obtained based on the updated Q-value and the Q-value before the update.

[0180] The satellite differential data broadcasting device of this invention acquires multimodal data, wherein the multimodal data includes one or more of the following: geographic environment data, user scenario data, terminal characteristic and status data, and ionospheric index distribution data; performs feature extraction and fusion processing on the multimodal data to obtain a fused feature vector; performs a nonlinear transformation on the fused feature vector through one or more fully connected layers of an Actor network to obtain a policy latent vector; performs hierarchical action decision-making on the policy latent vector to generate a broadcasting strategy, wherein the broadcasting strategy includes enabling or disabling broadcasting frequency, compression level, constellation, frequency point weight, and advanced correction numbers; and broadcasts satellite differential data according to the broadcasting strategy. Thus, the method of this invention can dynamically and intelligently adjust the broadcasting strategy based on geographic environment data, user scenario data, terminal characteristic and status data, and / or ionospheric index distribution data, thereby achieving resource optimization, improving positioning performance, and reducing cloud bandwidth costs and operating costs.

[0181] To better achieve the above objectives, such as Figure 12 As shown, this embodiment of the invention also provides a satellite differential data broadcasting device, including a processor 1200 and a transceiver 1210. The transceiver 1210 receives and transmits data under the control of the processor 1200, and the processor 1200 is used to execute the following processes: Acquire multimodal data; wherein the multimodal data includes one or more of the following: geographic environment data, user scenario data, terminal characteristics and status data, and ionospheric index distribution data; The multimodal data is subjected to feature extraction and fusion processing to obtain a fused feature vector; The policy latent vector is obtained by performing a nonlinear transformation on the fused feature vector through one or more fully connected layers of the Actor network. A hierarchical action decision is made on the strategy latent vector to generate a broadcast strategy, which includes enabling or disabling broadcast frequency, compression level, constellation, frequency point weight, and advanced correction numbers. Satellite differential data is broadcast according to the broadcasting strategy described above.

[0182] In some embodiments, the processor 1200 is further configured to: The geographic environment data is used to extract features through a geographic convolutional neural network to obtain a first feature vector. The second feature vector is obtained by extracting features from the ionospheric exponential distribution data using a time-series coding neural network. The third feature vector is obtained by extracting features from the user scenario data and the terminal characteristics and status data through a structured data neural network. Based on the first feature vector, the second feature vector, and / or the third feature vector, feature fusion processing is performed to obtain the fused feature vector.

[0183] In some embodiments, the processor 1200 is further configured to: The user scenario data is encoded using the encoding unit of a structured data neural network to obtain a low-dimensional and sparse binary vector. The terminal characteristics and status data are numerically normalized using the normalization unit of the structured data neural network to obtain a hybrid vector. The low-dimensional and sparse binary vector and the hybrid vector are input into the fully connected layer of the structured data neural network to obtain the third feature vector.

[0184] In some embodiments, the processor 1200 is further configured to: A linear transformation is performed on the first feature vector to obtain a first query matrix, a first key matrix, and a first value matrix; a first attention weight is calculated based on the first query matrix and the first key matrix; and a weighted fusion is performed on the first value matrix based on the first attention weight to obtain a first output result. A linear transformation is performed on the second feature vector to obtain a second query matrix, a second key matrix, and a second value matrix; based on the second query matrix and the second key matrix, a second attention weight is calculated; and based on the second attention weight, the second value matrix is ​​weighted and fused to obtain a second output result. A linear transformation is performed on the third feature vector to obtain a third query matrix, a third key matrix, and a third value matrix; based on the third query matrix and the third key matrix, a third attention weight is calculated; and based on the third attention weight, the third value matrix is ​​weighted and fused to obtain a third output result. The first output result, the second output result, and / or the third output result are concatenated and then subjected to a linear transformation to obtain the fused feature vector.

[0185] In some embodiments, the processor 1200 is further configured to: Gaussian distribution sampling is performed on the latent vectors of the strategy to obtain the broadcast frequency and compression level; Gumbel-Softmax sampling is performed on the latent vector of the strategy to obtain constellation and frequency point weights; The latent vectors of the policy are sampled using a Bernoulli distribution and processed with an activation function to determine whether advanced corrections are enabled or disabled.

[0186] In some embodiments, the processor 1200 is further configured to: The broadcast strategy is processed by action masking to obtain compliant actions; The compliant actions are verified for budget constraints to obtain budget-safe actions; Based on the budgeted safety actions, the final safety actions are obtained through pessimistic Q-value initialization and safety constraint processing of the safety action set, and the final safety actions are determined as the final broadcast strategy. Satellite differential data will be broadcast according to the final broadcast strategy described above.

[0187] In some embodiments, the processor 1200 is further configured to: Based on the terminal capabilities, the broadcasting strategy is masked to obtain a broadcasting strategy that matches the terminal capabilities. Based on business rules, the broadcasting strategy that matches the terminal capabilities is masked to obtain the compliant action.

[0188] In some embodiments, the processor 1200 is further configured to: The resource consumption prediction model is used to obtain the predicted resources consumed by the compliance action. Get the currently consumed resources; If the sum of the predicted resources consumed by the compliance action and the currently consumed resources is greater than a preset threshold, the compliance action is corrected to obtain the budgeted security action, wherein the resources consumed by the budgeted security action are less than or equal to the difference between the predicted resources consumed by the compliance action and the currently consumed resources. If the sum of the predicted resources consumed by the compliance action and the currently consumed resources is less than or equal to the preset threshold, the compliance action is determined to be the budget security action.

[0189] In some embodiments, the processor 1200 is further configured to: The Q-value output of the Crtic network is initialized to a constant that is much lower than the normal reward range; Determine whether the budgeted safety action belongs to the set of safety actions; If it is, then the budget security action is determined to be the final security action, and the Q value is updated after the final security action is executed; If it does not belong to the set of safety actions, then select the safety alternative action that is most similar to the budgeted safety action from the set of safety actions, and determine the safety alternative action as the final safety action.

[0190] In some embodiments, the processor 1200 is further configured to: Obtain multi-dimensional feedback data after the broadcast strategy is executed; Based on the multi-dimensional feedback data, calculate the corresponding reward value for each; The total reward value is calculated by weighting the values ​​based on their respective reward values ​​using a discount factor. Based on the total reward value, update the Q value of the Critical network; The policy parameters of the Actor network are updated based on the Q-value gradient, wherein the Q-value gradient is obtained based on the updated Q-value and the Q-value before the update.

[0191] The satellite differential data broadcasting device of this invention acquires multimodal data, which includes one or more of the following: geographic environment data, user scenario data, terminal characteristic and status data, and ionospheric index distribution data. It performs feature extraction and fusion processing on the multimodal data to obtain a fused feature vector. Through one or more fully connected layers of an Actor network, it performs a nonlinear transformation on the fused feature vector to obtain a policy latent vector. It then performs hierarchical action decision-making on the policy latent vector to generate a broadcasting strategy, which includes enabling or disabling broadcasting frequency, compression level, constellation, frequency point weights, and advanced corrections. The satellite differential data is broadcast according to the broadcasting strategy. Thus, the method of this invention can dynamically and intelligently adjust the broadcasting strategy based on geographic environment data, user scenario data, terminal characteristic and status data, and / or ionospheric index distribution data, thereby achieving resource optimization, improving positioning performance, and reducing cloud bandwidth and operating costs.

[0192] This invention also provides a satellite differential data broadcasting device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the various processes in the above-described satellite differential data broadcasting method embodiments and achieves the same technical effect. To avoid repetition, it will not be described again here.

[0193] This invention also provides a computer-readable storage medium storing a computer program. When executed by a processor, this program implements the various processes described above in the satellite differential data broadcasting method embodiment, achieving the same technical effect. To avoid repetition, it will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.

[0194] This invention also provides a computer program product, including computer instructions, which, when executed by a processor, implement the above-described functionality. Figure 2 The steps in the satellite differential data broadcasting method shown are illustrated.

[0195] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-readable storage media (including, but not limited to, disk storage and optical storage) containing computer-usable program code.

[0196] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 A device for one or more processes and / or the functions specified in one or more boxes.

[0197] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce a paper article including an instruction means, the instruction means being implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0198] These computer program instructions can also be loaded onto a computer or other programmable data processing equipment, causing the computer or other programmable equipment to perform a series of operational steps to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0199] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A method of broadcasting satellite differential data, characterized by, include: Acquire multimodal data; wherein the multimodal data includes one or more of the following: geographic environment data, user scenario data, terminal characteristics and status data, and ionospheric index distribution data; The multimodal data is subjected to feature extraction and fusion processing to obtain a fused feature vector; The policy latent vector is obtained by performing a nonlinear transformation on the fused feature vector through one or more fully connected layers of the Actor network. A hierarchical action decision is made on the strategy latent vector to generate a broadcast strategy, which includes enabling or disabling broadcast frequency, compression level, constellation, frequency point weight, and advanced correction numbers. Satellite differential data is broadcast according to the broadcasting strategy described above.

2. The method of claim 1, wherein, The step of performing feature extraction and fusion processing on the multimodal data to obtain a fused feature vector includes: The geographic environment data is used to extract features through a geographic convolutional neural network to obtain a first feature vector. The second feature vector is obtained by extracting features from the ionospheric exponential distribution data using a time-series coding neural network. The third feature vector is obtained by extracting features from the user scenario data and the terminal characteristics and status data through a structured data neural network. Based on the first feature vector, the second feature vector, and / or the third feature vector, feature fusion processing is performed to obtain the fused feature vector.

3. The method of claim 2, wherein, The step of extracting features from the user scenario data and the terminal characteristics and status data using a structured data neural network to obtain a third feature vector includes: The user scenario data is encoded using the encoding unit of a structured data neural network to obtain a low-dimensional and sparse binary vector. The terminal characteristics and status data are numerically normalized using the normalization unit of the structured data neural network to obtain a hybrid vector. The low-dimensional and sparse binary vector and the hybrid vector are input into the fully connected layer of the structured data neural network to obtain the third feature vector.

4. The method of claim 2, wherein, The step of performing feature fusion processing based on the first feature vector, the second feature vector, and / or the third feature vector to obtain the fused feature vector includes: A linear transformation is performed on the first feature vector to obtain a first query matrix, a first key matrix, and a first value matrix; a first attention weight is calculated based on the first query matrix and the first key matrix; and a weighted fusion is performed on the first value matrix based on the first attention weight to obtain a first output result. A linear transformation is performed on the second feature vector to obtain a second query matrix, a second key matrix, and a second value matrix; based on the second query matrix and the second key matrix, a second attention weight is calculated; and based on the second attention weight, the second value matrix is ​​weighted and fused to obtain a second output result. A linear transformation is performed on the third feature vector to obtain a third query matrix, a third key matrix, and a third value matrix; based on the third query matrix and the third key matrix, a third attention weight is calculated; and based on the third attention weight, the third value matrix is ​​weighted and fused to obtain a third output result. The first output result, the second output result, and / or the third output result are concatenated and then subjected to a linear transformation to obtain the fused feature vector.

5. The method according to claim 1, characterized in that, The step of performing hierarchical action decision-making on the policy latent vector to generate a broadcasting policy includes: Gaussian distribution sampling is performed on the latent vectors of the strategy to obtain the broadcast frequency and compression level; Gumbel-Softmax sampling is performed on the latent vector of the strategy to obtain constellation and frequency point weights; The latent vectors of the policy are sampled using a Bernoulli distribution and processed with an activation function to determine whether advanced corrections are enabled or disabled.

6. The method according to claim 5, characterized in that, The method further includes: The broadcast strategy is processed by action masking to obtain compliant actions; The compliant actions are verified for budget constraints to obtain budget-safe actions; Based on the budgeted safety actions, the final safety actions are obtained through pessimistic Q-value initialization and safety constraint processing of the safety action set, and the final safety actions are determined as the final broadcast strategy. The broadcasting of satellite differential data according to the broadcasting strategy includes: Satellite differential data will be broadcast according to the final broadcast strategy described above.

7. The method according to claim 6, characterized in that, The step of performing action masking on the broadcast strategy to obtain compliant actions includes: Based on the terminal capabilities, the broadcasting strategy is masked to obtain a broadcasting strategy that matches the terminal capabilities. Based on business rules, the broadcasting strategy that matches the terminal capabilities is masked to obtain the compliant action.

8. The method according to claim 6, characterized in that, The process of verifying the compliance actions against budget constraints to obtain budget-safe actions includes: The resource consumption prediction model is used to obtain the predicted resources consumed by the compliance action. Get the currently consumed resources; If the sum of the predicted resources consumed by the compliance action and the currently consumed resources is greater than a preset threshold, the compliance action is corrected to obtain the budgeted security action, wherein the resources consumed by the budgeted security action are less than or equal to the difference between the predicted resources consumed by the compliance action and the currently consumed resources. If the sum of the predicted resources consumed by the compliance action and the currently consumed resources is less than or equal to the preset threshold, the compliance action is determined to be the budget security action.

9. The method according to claim 6, characterized in that, The process of obtaining the final safety action based on the budgeted safety action, through pessimistic Q-value initialization and safety constraint processing of the safety action set, includes: The Q-value output of the Crtic network is initialized to a constant that is much lower than the normal reward range; Determine whether the budgeted safety action belongs to the set of safety actions; If it is, then the budget security action is determined to be the final security action, and the Q value is updated after the final security action is executed; If it does not belong to the set of safety actions, then select the safety alternative action that is most similar to the budgeted safety action from the set of safety actions, and determine the safety alternative action as the final safety action.

10. The method according to claim 9, characterized in that, The method further includes: Obtain multi-dimensional feedback data after the broadcast strategy is executed; Based on the multi-dimensional feedback data, calculate the corresponding reward value for each; The total reward value is calculated by weighting the values ​​based on their respective reward values ​​using a discount factor. Based on the total reward value, update the Q value of the Critical network; The policy parameters of the Actor network are updated based on the Q-value gradient, wherein the Q-value gradient is obtained based on the updated Q-value and the Q-value before the update.

11. A satellite differential data broadcasting device, characterized in that, include: The first acquisition module is used to acquire multimodal data; wherein, the multimodal data includes one or more of the following: geographic environment data, user scenario data, terminal characteristics and status data, and ionospheric index distribution data; The feature extraction and fusion module is used to perform feature extraction and fusion processing on the multimodal data to obtain a fused feature vector; The first processing module is used to perform a nonlinear transformation on the fused feature vector through one or more fully connected layers of the Actor network to obtain the policy latent vector; The broadcast strategy generation module is used to make hierarchical action decisions on the strategy latent vector and generate a broadcast strategy. The broadcast strategy includes enabling or disabling broadcast frequency, compression level, constellation, frequency point weight, and advanced correction numbers. The data broadcasting module is used to broadcast satellite differential data according to the broadcasting strategy.

12. A satellite differential data broadcasting device, comprising a processor and a transceiver, wherein the transceiver receives and transmits data under the control of the processor, characterized in that, The processor is used to perform the following operations: Acquire multimodal data; wherein the multimodal data includes one or more of the following: geographic environment data, user scenario data, terminal characteristics and status data, and ionospheric index distribution data; The multimodal data is subjected to feature extraction and fusion processing to obtain a fused feature vector; The policy latent vector is obtained by performing a nonlinear transformation on the fused feature vector through one or more fully connected layers of the Actor network. A hierarchical action decision is made on the strategy latent vector to generate a broadcast strategy, which includes enabling or disabling broadcast frequency, compression level, constellation, frequency point weight, and advanced correction numbers. Satellite differential data is broadcast according to the broadcasting strategy described above.

13. A satellite differential data broadcasting device, comprising a memory, a processor, and a computer program stored in the memory and running thereon, characterized in that, When the processor executes the program, it implements the satellite differential data broadcasting method as described in any one of claims 1 to 10.

14. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the steps of the satellite differential data broadcasting method as described in any one of claims 1 to 10.

15. A computer program product, characterized in that, It includes computer instructions that, when executed by a processor, implement the steps in the satellite differential data broadcasting method as described in any one of claims 1 to 10.