Ionospheric disturbance detection and gnss positioning response method and system
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-16
- Publication Date
- 2026-08-11
AI Technical Summary
上述技术流程的ROTI阈值法仅利用了统计量的标量信息,对局部斑块状闪烁、带状漂移扰动和渐变式增强等复杂扰动模式的识别能力不足
[0007]从以上技术方案可以看出,本申请具有以下优点:同步构建TEC格网图、ROTI格网图形成双通道图像,以N帧时序序列作为输入,视觉编码器同步提取两类图像的空间纹理特征与时序演化特征,完整捕获扰动区域空间斑块形态、TEC 梯度分布、扰动随时间推移的演化趋势,避免单一标量指标的信息局限,能够准确识别各类局部、动态、非均匀电离层扰动,降低扰动漏检、误检概率,有效避免了因电离层相位抖动误判导致的PPP定位精度退化。通过行动头输出的四维连续动作向量,实时更新周跳检测阈值、ROTI异常权重、TEC梯度权重和时间平滑权重,在强扰动场景下自动抬升周跳检测阈值以规避电离层相位抖动误判,并同步强化电离层扰动项约束权重,在电离层平静期恢复基础参数以保障定位收敛速度,实现了定位处理流程根据电离层状态变化的实时自适应调整,有效抑制了电离层扰动带来的定位误差,保障了高精度定位服务的连续可靠运行。
Smart Images

Figure CN122546249A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of satellite navigation and positioning technology, specifically to an ionospheric disturbance detection and GNSS positioning response method and system. Background Technology
[0002] In the field of satellite navigation and positioning (GNSS), ionospheric disturbance is one of the main factors affecting the continuity and reliability of high-precision positioning services. When GNSS satellite signals pass through the ionosphere, their propagation speed and direction change, resulting in signal delay and phase flicker, which in turn affects the positioning accuracy of the receiving terminal.
[0003] Currently, ionospheric disturbance monitoring and GNSS positioning response are mainly implemented based on the following technical process: First, the oblique total electron content (STEC) is calculated using GNSS dual-frequency observations. Through differential code bias (DCB) correction and the ionospheric puncture point (IPP) mapping function, the oblique STEC is converted to the vertical total electron content (VTEC). Second, an ionospheric disturbance index is constructed based on the rate of change of total electron content (ROT) and its standard deviation (ROTI). A fixed numerical threshold (typically ROTI > 0.5 TECU / min) is used to determine whether a disturbance has occurred. Finally, in GNSS precise positioning processing, a preset fixed cycle slip detection threshold (typically 3-5 times the standard deviation of carrier phase noise) is used for data quality control. This threshold remains unchanged during ionospheric calm periods and scintillation periods during storms; that is, when a disturbance occurs, the positioning processing flow does not adaptively adjust according to changes in the ionospheric state. The ROTI thresholding method described above only utilizes scalar information from statistics, and its ability to identify complex disturbance patterns such as localized patchy flicker, banded drift disturbances, and gradual enhancement is insufficient. When ionospheric disturbances cause frequent and severe fluctuations in the carrier phase, the traditional ROTI thresholding method struggles to distinguish between true cycle slips and phase jitter caused by the ionosphere, resulting in a degradation of PPP positioning accuracy to tens of meters or even complete interruption. Summary of the Invention
[0004] To address the aforementioned issues, this invention provides an ionospheric disturbance detection and GNSS positioning response method and system. By using a VLA model, it achieves synchronous perception of the spatial texture and temporal evolution of ionospheric disturbances and adaptively adjusts GNSS positioning parameters to ensure the continuous and reliable operation of high-precision positioning services under ionospheric disturbance conditions.
[0005] In a first aspect, the technical solution of the present invention provides an ionospheric disturbance detection and GNSS positioning response method, comprising the following steps: Acquire observation data from the Global Navigation Satellite System, and construct TEC and ROTI maps based on this data; The TEC map and ROTI map of the current time step and the previous N-1 time steps are combined to form an N-frame dual-channel input sequence, where N≥2; the N-frame dual-channel input sequence is input into a pre-trained visual-language-action model, which includes a visual encoder, a language inference module and an action head; The visual encoder extracts spatial texture features and temporal evolution features from the TEC map and ROTI map to generate visual features; the language inference module performs cross-modal semantic inference on the visual features to generate fused features; and the action head outputs ionospheric disturbance warning level and GNSS positioning parameter adjustment suggestions based on the fused features. Based on the GNSS positioning parameter adjustment suggestions, the GNSS positioning algorithm is triggered to switch parameters.
[0006] Secondly, the technical solution of the present invention provides an ionospheric disturbance detection and GNSS positioning response system, comprising: The data preprocessing unit is used to acquire observation data from the Global Navigation Satellite System and construct TEC and ROTI maps based on this observation data. The model inference unit is used to construct an N-frame dual-channel input sequence from the TEC map and ROTI map of the current time and the previous N-1 time steps, where N≥2; and input the N-frame dual-channel input sequence into a pre-trained visual-language-action model, which includes a visual encoder, a language inference module, and an action head; wherein, the visual encoder extracts spatial texture features and temporal evolution features from the TEC map and ROTI map to generate visual features; the language inference module performs cross-modal semantic inference on the visual features to generate fused features; and the action head outputs ionospheric disturbance warning level and GNSS positioning parameter adjustment suggestions based on the fused features; The parameter adjustment unit is used to trigger the GNSS positioning algorithm to switch parameters based on GNSS positioning parameter adjustment suggestions.
[0007] As can be seen from the above technical solutions, this application has the following advantages: It simultaneously constructs TEC grid maps and ROTI grid maps to form dual-channel images. With N frames of temporal sequence as input, the visual encoder simultaneously extracts the spatial texture features and temporal evolution features of the two types of images, fully capturing the spatial patch morphology of the perturbation area, TEC gradient distribution, and the evolution trend of the perturbation over time. It avoids the information limitations of a single scalar index, can accurately identify various local, dynamic, and non-uniform ionospheric perturbations, reduces the probability of missed or false detections of perturbations, and effectively avoids the degradation of PPP positioning accuracy caused by misjudgment of ionospheric phase jitter. By using the four-dimensional continuous motion vector output by the action head, the cycle slip detection threshold, ROTI anomaly weight, TEC gradient weight, and time smoothing weight are updated in real time. In strong disturbance scenarios, the cycle slip detection threshold is automatically raised to avoid misjudgment of ionospheric phase jitter, and the constraint weight of the ionospheric disturbance term is strengthened at the same time. During the ionospheric calm period, the basic parameters are restored to ensure the positioning convergence speed. This realizes the real-time adaptive adjustment of the positioning processing flow according to the changes in the ionospheric state, effectively suppresses the positioning error caused by ionospheric disturbance, and ensures the continuous and reliable operation of high-precision positioning services. Attached Figure Description
[0008] To more clearly illustrate the technical solution of this application, the accompanying drawings used in the description will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0009] Figure 1 This is a schematic diagram of a method for ionospheric disturbance detection and GNSS positioning response provided in an embodiment of the present invention.
[0010] Figure 2 This is a schematic diagram of the visual-language-action model structure.
[0011] Figure 3 This is a schematic block diagram of an ionospheric disturbance detection and GNSS positioning response system provided in an embodiment of the present invention. Detailed Implementation
[0012] To make the purpose, features, and advantages of this application more apparent and understandable, specific embodiments and accompanying drawings will be used to clearly and completely describe the technical solution protected by this application. Obviously, the embodiments described below are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0013] Unless otherwise defined, all technical and scientific terms used in this application have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used in this application and in the specification of this invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention.
[0014] Figure 1 This is a schematic flowchart of an ionospheric disturbance detection and GNSS positioning response method provided in an embodiment of the present invention. Figure 1 The executing entity can be an ionospheric disturbance detection and GNSS positioning response system. The ionospheric disturbance detection and GNSS positioning response method provided in this embodiment of the invention is executed by a computer device; correspondingly, the ionospheric disturbance detection and GNSS positioning response system runs on the computer device. Depending on different requirements, the order of the steps in this flowchart can be changed, and some steps can be omitted.
[0015] like Figure 1 As shown, the method includes the following steps.
[0016] S1: Acquire observation data from the Global Navigation Satellite System and construct TEC and ROTI maps based on this data.
[0017] S2, construct an N-frame dual-channel input sequence from the TEC map of the current time step and the ROTI map of the N-1 time steps before it, where N≥2; input the N-frame dual-channel input sequence into a pre-trained visual-language-action model, which includes a visual encoder, a language inference module and an action head.
[0018] Specifically, the visual encoder extracts spatial texture features and temporal evolution features from the TEC map and ROTI map to generate visual features; the language inference module performs cross-modal semantic inference on the visual features to generate fused features; and the action head outputs ionospheric disturbance warning level and GNSS positioning parameter adjustment suggestions based on the fused features.
[0019] S3, based on the GNSS positioning parameter adjustment suggestion, triggers the GNSS positioning algorithm to switch parameters.
[0020] As a refinement and extension of the specific implementation of the above embodiments, in order to fully explain the specific implementation process of this embodiment, the following will provide possible embodiments to describe the specific implementation of the above steps in a non-limiting manner.
[0021] In this embodiment, step S1 involves acquiring observation data from the Global Navigation Satellite System and constructing a TEC map and a ROTI map based on this observation data. This step specifically includes the following steps.
[0022] S1.1 Perform cycle slip detection and repair, gross error removal and cutoff elevation angle screening on the acquired GNSS dual-frequency pseudorange observations and carrier phase observations to obtain preprocessed dual-frequency pseudorange observations and carrier phase observations.
[0023] Cycle slip detection and repair is used to detect and repair integer cycle count jumps in carrier phase observations caused by signal loss of lock, so as to ensure the continuity of carrier phase observations; gross error removal is used to remove abnormal observations with large errors to avoid contaminating subsequent TEC estimation; cutoff elevation angle screening is used to remove observations from low elevation angle satellites to reduce the impact of tropospheric delay and multipath effects.
[0024] S1.2, obtain satellite ephemeris, differential code offset and receiver coordinates, and calculate satellite position based on satellite ephemeris.
[0025] Satellite ephemeris is used to determine the satellite's position coordinates in space; differential code bias (DCB) includes pseudorange hardware delay bias at the satellite and receiver ends, and is used to eliminate systematic bias in pseudorange observations caused by hardware delay; receiver coordinates are used to determine the receiver's position.
[0026] S1.3. Using the preprocessed dual-frequency pseudorange observations and carrier phase observations, a dual-frequency geometrically independent combined observation equation is constructed. Based on the satellite position, differential code deviation, and receiver coordinates, the oblique total electron content in the line-of-sight direction from the satellite to the receiver is estimated.
[0027] The basic principle of dual-frequency geometrically independent combination is that the propagation delay of a GNSS signal when it traverses the ionosphere is inversely proportional to the square of the signal frequency. By combining pseudorange and carrier phase observations at two frequencies, frequency-independent geometric distance, satellite clock bias, and receiver clock bias can be eliminated, thereby separating the combined observations containing ionospheric delay information and estimating the oblique total electron content along the signal propagation path.
[0028] S1.4, using a mapping function under the assumption of a thin ionospheric layer, based on the satellite elevation angle determined by the receiver coordinates and satellite position, the oblique total electron content is converted into the vertical total electron content in the receiver zenith direction.
[0029] The thin ionospheric layer hypothesis considers the ionosphere as an infinitely thin spherical shell located at a certain altitude above the Earth. The intersection of the signal propagation path and this thin layer is called the ionospheric puncture point. The mapping function uses the satellite elevation angle at the puncture point as the independent variable and projects the STEC along the inclined path to the zenith direction to eliminate the influence of different satellite elevation angles on the TEC observation values, making the observation values of different receivers and different satellites comparable.
[0030] S1.5, the discrete vertical total electron content scatter points are interpolated using the inverse distance weighted interpolation method according to the preset spatial grid to construct a spatial gridded TEC map.
[0031] Because the ionospheric puncture points corresponding to the line-of-sight of each receiver and satellite differ, the spatial distribution of VTEC observations is a non-uniform, discrete scatter, which cannot be directly used as image input for subsequent visual encoders. The basic principle of inverse distance weighted interpolation is as follows: for each grid point, the inverse distance from the known surrounding points to that grid point is used as the weight, and the VTEC values of all known points are weighted and averaged, with closer points assigned higher weights. This interpolation method normalizes the discrete scatter points into image data on a uniform spatial grid, resulting in the TEC map.
[0032] S1.6 Calculate the TEC rate of change and its standard deviation based on the TEC time series to obtain the ROTI value. Use the same gridding method as the TEC plot to interpolate the ROTI scatter points to the spatial grid to obtain the ROTI plot.
[0033] The TEC rate of change reflects the speed at which TEC changes between adjacent epochs, while its standard deviation (ROTI) characterizes the degree of fluctuation in the TEC rate of change within a certain time window. The ROTI values are also distributed as discrete scatter points at each puncture point location, and are gridded into an ROTI map using the same inverse distance weighted interpolation method as in step S1.5.
[0034] In this embodiment, step S2 involves constructing an N-frame dual-channel input sequence from the TEC map of the current time and the N-1 time steps before it, along with the ROTI map. Specifically, this includes the following steps S2.1 to S2.4.
[0035] S2.1 Extract the TEC map and ROTI map of the current time (time t) and the N-1 times before it to form a four-frame time-series image sequence.
[0036] Since ionospheric perturbation is a dynamic process that evolves over time, a single frame image can only reflect the spatial distribution at a certain moment and cannot provide information on whether the perturbation is strengthening, weakening, or maintaining its evolutionary trend. This embodiment selects four frames, including the current moment and the N-1 moments before it, to capture the short-term evolutionary trend of the perturbation through multi-frame temporal information. By controlling the length of the input sequence, the computational burden on the model is avoided by using too many historical frames.
[0037] S2.2, the TEC map and ROTI map in each frame are treated as two independent channels and spliced along the channel dimension to form a single-frame input tensor of size 2×H×W, where H=W is the spatial grid size.
[0038] The TEC map reflects the spatial distribution background of ionospheric electron content, presenting large-scale ionospheric structural features; the ROTI map reflects the fluctuation degree of the TEC rate of change, highlighting the active regions of the ionospheric irregular structure and small-scale perturbation features. By stitching the two as two independent channels, the subsequent visual encoder can simultaneously utilize the background distribution information of TEC and the perturbation intensity information of ROTI during the same feature extraction process, achieving complementarity and fusion of the two types of information.
[0039] S2.3 Stack the N single-frame input tensors along the time dimension to obtain a multi-frame dual-channel input tensor with dimensions T×C×H×W, where T is the number of time frames and C is the number of channels.
[0040] By stacking the time dimension, four time-series images are organized into a unified four-dimensional tensor structure, in which the time dimension and the channel dimension are independent of each other, enabling the model to simultaneously perceive "which regions have perturbations" (spatial dimension), "when the perturbations occur" (time dimension), and "what type of perturbation they exhibit" (channel dimension).
[0041] S2.4 performs z-score normalization on the TEC channel and ROTI channel in the multi-frame dual-channel input tensor respectively.
[0042] Because TEC and ROTI have different physical dimensions and numerical ranges, if directly input into the model, features with larger numerical ranges will dominate gradient updates, affecting the stability and convergence speed of model training. This embodiment performs z-score normalization on both channels separately, ensuring that the mean of each channel is 0 and the standard deviation is 1, thus eliminating the impact of the difference in dimensions on model training. The normalization formula is:
[0043] Where X represents the original channel data. This is the mean of the channel. This represents the standard deviation of the channel. During the training phase, the mean and standard deviation are statistically obtained from the training set; during the inference phase, the mean and standard deviation parameters saved during the training phase are used for normalization.
[0044] For example, the TEC map and ROTI map of the current time step and the three time steps before it are combined to form a four-frame dual-channel input sequence. In each frame, the TEC map and ROTI map are concatenated as two independent channels along the channel dimension to form a single-frame input tensor of size [2×H×W]. The dimension of the input tensor after stacking the four frames is [T×C×H×W], where T=4 is the number of time frames, C=2 is the number of channels, H=W=64 is the spatial grid size, and the TEC channel and ROTI channel are z-score normalized respectively.
[0045] The multi-frame dual-channel input tensor simultaneously contains information in the following four dimensions: the number of time frames T=4, indicating that the ionospheric temporal evolution information of the current time and the previous 3 time moments has been captured; the number of channels C=2, indicating that two types of information, namely TEC background distribution and ROTI perturbation intensity, are input simultaneously; and the spatial grid H×W=64×64, indicating that the location, shape, range and texture structure of the perturbation region are preserved in the spatial dimension.
[0046] In this embodiment, an N-frame dual-channel input sequence is fed into a pre-trained Vision-Language-Action Model (VLA model), which then outputs ionospheric disturbance detection and GNSS positioning parameter adjustment suggestions. Figure 2 As shown, the VLA model includes a visual encoder, a language inference module, and an action head.
[0047] The VLA model is an end-to-end decision-making architecture. The visual encoder extracts spatial texture features and temporal evolution features from the TEC map and ROTI map to generate visual features. The language inference module performs cross-modal semantic inference on the visual features to generate fused features. The action head outputs ionospheric disturbance warning level and GNSS positioning parameter adjustment suggestions based on the fused features.
[0048] In this embodiment, the visual encoder includes a patch embedding layer, a position encoding layer, a visual Transformer encoding body, and a temporal aggregation module. The visual encoder extracts spatial texture features and temporal evolution features from the TEC map and the ROTI map, specifically including the following steps S2.5 to S2.8.
[0049] S2.5, the Patch embedding layer uses two-dimensional convolution to divide the TEC map and ROTI map of each frame in the multi-frame dual-channel input tensor into non-overlapping image blocks, and projects each image block to the visual token space to obtain the initial token of each image block in each frame.
[0050] Specifically, for each frame's two-channel input tensor of size 2×H×W (H=W=64), the Patch embedding layer performs a convolution operation on it using a two-dimensional convolution kernel of size 8×8 and stride 8. This convolution operation uniformly divides each frame image into N=(H / 8)×(W / 8)=8×8=64 non-overlapping image patches in spatial dimension, with each patch containing 8×8×2=128 pixel values (including TEC and ROTI channels). Subsequently, each image patch is mapped to a vector image through the projection operation of the convolution kernel. Each image patch is represented by a 128-dimensional visual token, which encodes the spatial distribution information of TEC and ROTI within that local region. After the above processing, each frame of the image is represented by N=64 visual tokens, and each token has a dimension of... .
[0051] S2.6, the location encoding layer adds spatial location information and temporal order information to the initial token of each image block to obtain the encoded visual token sequence.
[0052] The location encoding layer includes spatial location encoding and temporal location encoding. Spatial location encoding identifies the position of each image patch in the spatial grid, such as the i-th row and j-th column, enabling the model to perceive the spatial distribution and geometry of the perturbation region. Temporal location encoding identifies the order of each frame in the entire temporal sequence, such as the current frame, the previous frame, the second frame before that, and the third frame before that, enabling the model to perceive the temporal evolution stage of the perturbation. The spatial and temporal location encodings are superimposed on the corresponding image patch tokens, so that each token carries information about both "which spatial location it comes from" and "which temporal frame it comes from". After merging the four frames, a total of T×N=4×64=256 tokens, a sequence of encoded visual tokens with dimensions [B×256×128] is obtained, where B is the batch size.
[0053] S2.7, the visual Transformer encoding subject receives the encoded visual token sequence, extracts the spatial texture features of the TEC map and ROTI map through a multi-head self-attention mechanism and a feedforward multilayer perceptron, and obtains the spatial feature tokens of each spatial image patch in each time frame. The spatial texture features include the TEC gradient structure, the ROTI abnormal patch morphology and the spatial distribution pattern of the perturbation region.
[0054] The main body of the visual Transformer encoding consists of The structure consists of stacked Transformer blocks, each containing a multi-head self-attention mechanism (number of heads). This model employs a multi-head self-attention mechanism and a feedforward multilayer perceptron (MLP), utilizing layer normalization and residual connections to accelerate training and improve model performance. In the MLP, each token aggregates information from all other tokens through the interaction of query, key, and value, enabling the features of each image patch to fuse global contextual information and effectively capture spatial associations and long-range dependencies between perturbed regions. The feedforward MLP then performs independent nonlinear transformations on each token, enhancing the expressive power of the features. The model uses stacked Transformer blocks to extract spatial texture features of ionospheric perturbations layer by layer. These features include the TEC gradient structure (reflecting the drastic spatial variation in ionospheric electron content), ROTI anomalous patch morphology (reflecting the shape, size, and boundary features of the perturbation region), and the spatial distribution pattern of the perturbation region (reflecting the spatial arrangement and distribution pattern of the perturbation region). After processing by the visual Transformer encoding body, the output feature dimension remains unchanged. That is, each spatial image patch corresponds to one in each time frame. A spatial feature token, which encodes the spatial texture information of the local region.
[0055] S2.8 The temporal aggregation module receives the spatial feature tokens of each spatial image block in each time frame, performs a temporal Transformer attention operation on the token sequence of each spatial image block in N time frames, captures the temporal evolution trend of ionospheric perturbation, and performs average pooling on the time dimension to compress the N frames of temporal information into a single feature representation, thereby obtaining a visual feature that integrates spatial texture features and temporal evolution features.
[0056] After processing by S2.7, each spatial image patch has a spatial feature token across T=4 time frames, and these tokens are independent of each other in spatial dimension. Ionospheric perturbation is a dynamic evolutionary process, and static spatial texture features from a single time frame are insufficient to determine whether the perturbation is strengthening, weakening, or dissipating. To this end, the temporal aggregation module processes each spatial image patch independently: for the i-th spatial image patch, it takes the feature tokens from its T=4 time frames and arranges them in chronological order into a temporal token sequence with a dimension of [4×128]; then, it performs a temporal Transformer attention operation on this sequence, calculating the attention weights between each time frame in the time dimension, so that the model can focus on the moments that play a key role in judging the perturbation state and capture the dynamic change pattern of the perturbation over time, i.e., the "temporal evolution features", such as whether the perturbation region gradually expands or shrinks, whether the ROTI anomaly intensity continuously increases or gradually decreases, and other evolutionary trends; finally, it performs average pooling on the weighted features of the T=4 time frames along the time dimension, that is, it calculates the element-wise average of the feature vectors of the same spatial image patch in the 4 time frames, compressing the temporal information into one. A single feature vector of dimension 1.
[0057] After average pooling, the time dimension is eliminated, and the output dimension becomes... This single feature vector contains both the spatial texture feature information of the spatial image patch (from the output of the S2.7 visual Transformer encoding subject) and the temporal evolution information of the spatial image patch over T=4 time frames (from the weighted processing of temporal attention and the comprehensive expression after average pooling), thus achieving deep fusion of spatial texture features and temporal evolution features.
[0058] In this embodiment, the visual-language-action model further includes a visual-language dimension projection layer, which is used to map the output of the visual encoder to the input space of the language inference module, specifically including the following steps S2.9 and S2.10.
[0059] S2.9, the visual-language dimension projection layer maps the visual features output by the visual encoder from the visual space dimension to the language reasoning space dimension of the language reasoning module.
[0060] Specifically, as shown in step S2.8, the visual feature dimension output by the visual encoder is... Where N=64 is the number of image patches This represents the visual token dimension. The working dimension of the language reasoning module is... Since the two dimensions are inconsistent, direct input is not possible. Therefore, this embodiment introduces a learnable linear projection layer. Each visual token is linearly mapped from a 128-dimensional visual space to a 192-dimensional language reasoning space, i.e.:
[0061] in, The visual features output by the visual encoder. The projected weight matrix is a learnable matrix. This is the projected sequence of visual features. Through this projection operation, the visual features are transformed into a feature space that matches the language inference module, enabling subsequent cross-modal attention calculations to be performed on a unified dimension. The projected feature dimension is... .
[0062] S2.10, learnable cue tokens and action tokens are concatenated onto the projected visual feature sequence to generate the input sequence for the language reasoning module.
[0063] Specifically, before the projected N=64 visual tokens, concatenation is performed. Each token has one learnable prompt token and one learnable action token. Both the prompt token and the action token are dimensional. The learnable vectors are initialized through random initialization or a specific pre-training strategy and are optimized along with other parameters during model training.
[0064] The length of the input sequence for the concatenated language reasoning module is The dimensions are [B×89×192]. Among them, the front... The first position is the cue token, the middle N=64 positions are the projected visual feature tokens, and the last position is the action token.
[0065] The cue tokens are used to encode the task semantics for ionospheric disturbance state assessment and positioning parameter adjustment recommendations. Specifically, the cue tokens encode the task semantics of "determining the ionospheric disturbance level based on the TEC and ROTI map sequences and providing parameter adjustment recommendations," guiding the language inference module to focus on decision elements related to GNSS positioning response, i.e., determining the ionospheric disturbance level based on the TEC and ROTI map sequences and providing parameter adjustment recommendations. During cross-modal attention computation, these cue tokens, as task-related contextual information, interact with visual feature tokens, enabling the language inference module to selectively focus on and understand visual information under the semantic guidance of "what task needs to be completed." Specifically, during training, the cue tokens learn semantic representations related to ionospheric disturbance detection and GNSS positioning response tasks, and during the inference phase, they serve as task guidance signals, causing the output of the language inference module to converge towards the task objective.
[0066] Action tokens are used to aggregate information related to action decisions. In the multi-head self-attention mechanism of the language reasoning module, the action token, as the last position in the sequence, can access information from all cue tokens and visual feature tokens in the sequence through the attention mechanism, aggregating and encoding the task semantics and visual features scattered throughout the sequence into its own hidden state. Therefore, the output of the language reasoning module is the hidden state of the last action token position. As input to the subsequent action head, this hidden state contains both visual feature information and task semantic information, providing information-rich feature representations for the action head to output warning levels and positioning parameter adjustment suggestions.
[0067] In this embodiment, the language reasoning module receives the concatenated input sequence and, through... Layered Transformerblock stacks are used for cross-modal fusion inference.
[0068] Each Transformer block contains a multi-head self-attention mechanism and a feedforward multilayer perceptron (MLP), and employs layer normalization and residual connections. In the multi-head self-attention mechanism, each token in the input sequence is computed through the interaction of query, key, and value, aggregating information from all tokens in the sequence.
[0069] through After processing by the Transformer block, the output sequence maintains a dimension of [B×89×192], where the hidden state of the last action token position is... This is extracted and used as input for subsequent action heads. This hidden state simultaneously contains visual feature information (from the visual feature token), task semantic information (from the cue token), and converged decision-related information (from the attention calculation of the action token itself).
[0070] In this embodiment, the language reasoning module is trained using a LoRA (Low-Rank Adaptation) fine-tuning strategy to adapt to the scarcity of ionospheric perturbation samples during storms.
[0071] During the training of the language reasoning module, low-rank decomposition increments are applied to the original pre-trained weight matrices of the query linear layer, key linear layer, and value linear layer, respectively. Let the original pre-trained weight matrix be... ,in These are the input and output dimensions of the linear layer, respectively, in this embodiment. In LoRA, the fine-tuned weight matrix is represented as:
[0072] in, , Let r be the rank of the low-rank decomposition, and in this embodiment, r = 8. Since r is much smaller than... The number of parameters in A and B is much smaller than that in the original weight matrix W. During training, the original pre-trained weight matrix W is frozen, and only the parameters in the low-rank matrices A and B are trained. The number of trainable parameters is less than 5 million, which effectively avoids the overfitting problem under the condition of scarce samples during peak training.
[0073] During forward propagation, the input x is computed through this layer as follows:
[0074] That is, the output is the result of the original pre-trained weights. Results with the low-rank increment part The results are obtained by adding them together. Through LoRA fine-tuning, this embodiment retains the original general capabilities of the pre-trained language model while achieving targeted optimization for ionospheric disturbance detection and GNSS positioning response tasks with only a small number of trainable parameters.
[0075] The specific process of the language inference module performing cross-modal semantic inference on visual features and generating fused features includes: the language inference module receives the concatenated input sequence (including cue tokens, visual feature tokens, and action tokens) and processes it through stacked Transformer blocks. In each layer, a multi-head self-attention mechanism enables each token in the sequence to interact with all other tokens. Spatial texture features and temporal evolution features in visual tokens, task semantics encoded in cue tokens, and the convergence function of action tokens achieve cross-modal fusion in this process. Each layer abstracts progressively, ultimately forming a fused feature at the action token location that includes both visual perception information and task semantic understanding. Specifically: in each Transformer block, each token in the input sequence (including cue tokens, visual feature tokens, and action tokens) calculates its attention weight with all other tokens through a query-key-value mechanism, and weighted aggregates the information of each token. This allows the spatial texture and temporal evolution information in visual feature tokens to be transmitted to action tokens, and the task semantics in cue tokens to guide the attention direction of visual features; after... Through layer-by-layer processing of the Transformer block, the model gradually completes the mapping from low-level visual features to high-level task semantics, ultimately forming fused features at the action token location. The action token is located at the end of the sequence. Through a self-attention mechanism, it can access all cue tokens and visual feature tokens, thus converging visual information and task semantics scattered throughout the sequence into its hidden state. middle.
[0076] In this embodiment, the action token output by the language reasoning module has a hidden state. The data is input to the action head, which outputs an ionospheric disturbance warning level and GNSS positioning parameter adjustment recommendations. The action head consists of two parallel branches: a classification output layer and a regression output layer.
[0077] The classification output layer will hide the state. The data is mapped to four-dimensional classification logits, and the probability distribution of each warning level is obtained through the softmax function. The category corresponding to the highest probability is taken as the predicted ionospheric disturbance warning level.
[0078] The ionospheric disturbance warning level is divided into four levels: Level 0 is the normal state, indicating that the TEC and ROTI changes are weak and the ionosphere is in a relatively calm state; Level 1 is a slight disturbance, indicating that the TEC or ROTI in a local area has increased, but has not yet formed a significant spatial structure; Level 2 is a moderate disturbance, indicating that the disturbance area has expanded or the temporal changes are more obvious, and the irregular structure of the ionosphere tends to be more active; Level 3 is a strong disturbance, indicating that the ROTI and TEC gradients are significantly enhanced, the ionospheric scintillation activity is strong, and it poses a significant threat to the quality of GNSS signals.
[0079] Specifically, the classification output layer passes through a fully connected linear layer. (dimension) Mapped to a four-dimensional logits vector :
[0080] in, This is the classification weight matrix. This is the classification bias vector. Then, the logits are converted into a probability distribution using the softmax function.
[0081] in Let be the probability of the i-th warning level. The final predicted warning level is: .
[0082] The regression output layer will contain the hidden state. It is mapped to a four-dimensional continuous motion vector and output as a GNSS positioning parameter adjustment suggestion.
[0083] The regression output layer is passed through a fully connected linear layer. Mapped to a four-dimensional continuous action vector:
[0084] in, For the regression weight matrix, This is the regression bias vector; When expanded, it appears as follows:
[0085] in, This represents the predicted value of the increase in the cycle slip detection threshold, which is used to adjust the cycle slip detection threshold in the GNSS positioning algorithm to cope with carrier phase jitter caused by ionospheric scintillation. This represents the predicted value of the ROTI anomaly weight, which is used to adjust the constraint strength of the ROTI anomaly term in the localization solution. When ionospheric disturbances are enhanced, increasing this weight can improve the localization algorithm's sensitivity to ROTI anomaly patches. This represents the predicted value of the TEC gradient weight, which is used to adjust the constraint strength of the TEC gradient term in the localization solution. When ionospheric disturbances are enhanced, increasing this weight can enhance the localization algorithm's sensitivity to the TEC spatial gradient. The predicted value represents the time smoothing weight, which is used to adjust the strength of time smoothing between epochs. When ionospheric disturbances are enhanced, reducing this weight can reduce the blurring of the true disturbance boundary caused by the smoothing effect, allowing the positioning algorithm to respond more quickly to rapid changes in the ionosphere. When the ionosphere is in a calm period, this weight is restored to improve the positioning convergence speed.
[0086] The regression output layer outputs a four-dimensional continuous motion vector to the GNSS positioning solution module, which then calculates the position based on the input vector. Update the cycle slip detection threshold, expressed as:
[0087] in, The baseline cycle slip detection threshold coefficient is used. This represents the standard deviation of carrier phase noise.
[0088] In this embodiment, the model outputs a four-dimensional continuous action vector. The data is transmitted in real time to the GNSS precise point positioning (PPS) module. The PPS module reads the motion vector at each epoch and updates the cycle slip detection threshold, ROTI anomaly weights, TEC gradient weights, and time smoothing weights based on each component, thus dynamically adjusting the positioning parameters. When the model determines a strong disturbance, the cycle slip detection threshold is automatically increased (by raising the threshold). To avoid misjudging ionospheric phase jitter as cycle slip, and at the same time increase the ROTI anomaly weight (increase the weight of ROTI anomalies). ) and TEC gradient weights (increase) ), reduce time smoothing weight (reduce) This enhances the positioning algorithm's responsiveness to disturbance boundaries and rapid changes. When the model determines a calm state, the baseline cycle slip detection threshold and default weights are restored to ensure positioning convergence speed and solution efficiency. Through dynamic adjustment, GNSS positioning solution parameters can be adaptively switched in real time according to the ionospheric disturbance state, effectively suppressing positioning errors introduced by ionospheric disturbances and ensuring the continuous and reliable operation of high-precision positioning services.
[0089] In this embodiment, the vision-language-action model is trained end-to-end using a joint loss function. Training data includes TEC and ROTI map sequences and their corresponding warning level labels and behavioral parameter labels. When training samples contain manually labeled data, these supervised labels are used directly for training; when training samples lack manually labeled data, weakly supervised labels are constructed based on the physical statistics of the TEC and ROTI maps. During training, the model optimizes both the warning level classification accuracy and the behavioral parameter regression accuracy by minimizing the joint loss function.
[0090] For training samples lacking manual annotation, a comprehensive perturbation score is first constructed based on the physical statistics of the TEC and ROTI plots. Then based on the comprehensive disturbance score The system compares data with multiple preset thresholds to generate a four-level warning label. Finally, based on the warning level Generate a vector of behavioral parameter labels.
[0091] Overall perturbation score The average intensity of the ROTI map at the current time is obtained by weighted summation of three components. It reflects the overall activity level of the irregular structure of the ionosphere; the average absolute change of the TEC diagram at the current moment relative to the TEC diagram at the previous moment. This reflects the degree of drastic change in the ionosphere over time; the spatial standard deviation of the TEC plot at the current moment. This reflects the degree of non-uniformity of the ionosphere in spatial dimensions. It is represented as:
[0092] in, The calculation formula is:
[0093] The calculation formula is:
[0094] The calculation formula is:
[0095] in, This represents the average value of the TEC plot at the current moment.
[0096] Let be the weighting coefficient, satisfying For example, .
[0097] According to the comprehensive disturbance analysis The number is compared with multiple preset thresholds to generate a four-level warning label, as exemplified by:
[0098] Level 0 represents normal disturbance, Level 1 represents slight disturbance, Level 2 represents moderate disturbance, and Level 3 represents strong disturbance. These thresholds are set based on the statistical distribution of low-latitude ionospheric disturbance events and are matched with the values of the weighting coefficients.
[0099] According to the warning level Generate behavioral parameter label vectors , is represented as:
[0100]
[0101]
[0102]
[0103] in, Labels indicating the increase in the cycle slip detection threshold. Labels representing ROTI outlier weights. Labels representing TEC gradient weights Labels representing time smoothing weights, This is a preset constant.
[0104] For example, The values are 1, 0.25, 0.2, 0.5, 0.2, 0.18, and 0.1. Among them, As the baseline value, This represents the step coefficient.
[0105] The model training employs a joint loss function to simultaneously optimize the classification accuracy of the warning level and the regression accuracy of the behavioral parameters.
[0106] in, The cross-entropy loss for classifying warning levels is used to measure the predicted warning level. Compared with the actual warning level Differences between ; represents the Smooth L1 loss for behavioral parameter regression, used to measure the predicted four-dimensional continuous action vector. Labels of real behavior parameters The difference between them. Smooth L1 loss combines the advantages of L1 loss's robustness to outliers and L2 loss's smooth differentiability near zero, providing stable gradient updates during model training.
[0107] For example, The classification loss weight is higher than the regression loss weight to ensure the accuracy of the warning level classification during the early stages of model training. It should be noted that the above loss weight coefficients can be adjusted according to the specific application scenario.
[0108] The optimizer uses AdamW, and the learning rate is... The weight decay coefficient λ = 0.01. The AdamW optimizer decouples weight decay from gradient updates, building upon the standard Adam optimizer, thus more effectively suppressing overfitting. During training, a cosine annealing learning rate scheduling strategy is employed, gradually reducing the learning rate from its initial value to near zero, which helps the model converge to the optimal solution with fine-grained precision in the later stages of training.
[0109] Due to the scarcity of ionospheric perturbation samples during storms, full-parameter fine-tuning can easily lead to overfitting. In this embodiment, a LoRA (Low-Rank Adaptation) fine-tuning strategy is adopted in the query linear layer, key linear layer, and value linear layer of the language inference module, as detailed above, and will not be repeated here.
[0110] The foregoing has described in detail an embodiment of an ionospheric disturbance detection and GNSS positioning response method. Based on the ionospheric disturbance detection and GNSS positioning response method described in the above embodiment, this invention also provides an ionospheric disturbance detection and GNSS positioning response system corresponding to the method.
[0111] Figure 3 This is a schematic block diagram of an ionospheric disturbance detection and GNSS positioning response system provided in an embodiment of the present invention. In this embodiment, the ionospheric disturbance detection and GNSS positioning response system performs its functions. The module referred to in this invention is a series of computer program segments that can be executed by at least one processor and perform a fixed function, and is stored in memory.
[0112] The data preprocessing unit is used to acquire observation data from the Global Navigation Satellite System and construct TEC and ROTI maps based on this observation data.
[0113] The model inference unit is used to construct an N-frame dual-channel input sequence from the TEC map and ROTI map of the current time and the previous N-1 time steps, where N≥2; and input the N-frame dual-channel input sequence into a pre-trained vision-language-action model, which includes a visual encoder, a language inference module, and an action head; wherein, the visual encoder extracts spatial texture features and temporal evolution features from the TEC map and ROTI map to generate visual features; the language inference module performs cross-modal semantic inference on the visual features to generate fused features; and the action head outputs ionospheric disturbance warning level and GNSS positioning parameter adjustment suggestions based on the fused features.
[0114] The parameter adjustment unit is used to trigger the GNSS positioning algorithm to switch parameters based on GNSS positioning parameter adjustment suggestions.
[0115] The ionospheric disturbance detection and GNSS positioning response system of this embodiment is used to implement the aforementioned ionospheric disturbance detection and GNSS positioning response method. Therefore, the specific implementation of this system can be found in the embodiment section of the ionospheric disturbance detection and GNSS positioning response method above. Thus, its specific implementation can be referred to the description of the corresponding embodiments, and will not be elaborated here.
[0116] Furthermore, since the ionospheric disturbance detection and GNSS positioning response system of this embodiment is used to implement the aforementioned ionospheric disturbance detection and GNSS positioning response method, its function corresponds to the function of the above method, and will not be described again here.
[0117] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method of ionospheric disturbance detection and GNSS positioning response, characterized in that, Includes the following steps: Acquire observation data from the Global Navigation Satellite System, and construct TEC and ROTI maps based on this data; The TEC map and ROTI map of the current time step and the previous N-1 time steps are combined to form an N-frame dual-channel input sequence, where N≥2; the N-frame dual-channel input sequence is input into a pre-trained visual-language-action model, which includes a visual encoder, a language inference module and an action head; The visual encoder extracts spatial texture features and temporal evolution features from the TEC map and ROTI map to generate visual features; the language inference module performs cross-modal semantic inference on the visual features to generate fused features; and the action head outputs ionospheric disturbance warning level and GNSS positioning parameter adjustment suggestions based on the fused features. Based on the GNSS positioning parameter adjustment recommendations, the GNSS positioning algorithm is triggered to switch parameters.
2. The ionospheric disturbance detection and GNSS positioning response method according to claim 1, characterized in that, Based on this observational data, TEC and ROTI maps were constructed, specifically including: The acquired GNSS dual-frequency pseudorange and carrier phase observations are subjected to cycle slip detection and repair, gross error removal and cutoff elevation angle screening to obtain preprocessed dual-frequency pseudorange and carrier phase observations. Obtain satellite ephemeris, differential code offset, and receiver coordinates, and calculate satellite position based on satellite ephemeris; Using preprocessed dual-frequency pseudorange and carrier phase observations, a dual-frequency geometrically independent combined observation equation is constructed. Based on the satellite position, differential code deviation, and receiver coordinates, the oblique total electron content in the line-of-sight direction from the satellite to the receiver is estimated. Using a mapping function under the assumption of a thin ionosphere, and based on the satellite elevation angle determined by the receiver coordinates and satellite position, the oblique total electron content is converted into the vertical total electron content in the receiver zenith direction; The discrete vertical total electron content scatter points are interpolated using the inverse distance weighted interpolation method according to a preset spatial grid to construct a spatial gridded TEC map; The TEC rate of change and its standard deviation are calculated based on the TEC time series to obtain the ROTI value. The ROTI scatter points are then interpolated to a spatial grid using the same gridding method as the TEC plot to obtain the ROTI plot.
3. The ionospheric disturbance detection and GNSS positioning response method according to claim 1, characterized in that, Before inputting the N-frame dual-channel input sequence into the pre-trained vision-language-action model, the process also includes tensor construction and normalization of the N-frame dual-channel input sequence, specifically including: Each frame's TEC map and ROTI map in N frames are treated as two independent channels and spliced along the channel dimension to form a single-frame input tensor. Stack N single-frame input tensors along the time dimension to obtain a multi-frame dual-channel input tensor. The TEC and ROTI channels in the multi-frame dual-channel input tensor are normalized respectively.
4. The ionospheric disturbance detection and GNSS positioning response method according to claim 3, characterized in that, The visual encoder includes a patch embedding layer, a positional encoding layer, a visual Transformer encoding body, and a temporal aggregation module; The Patch embedding layer uses two-dimensional convolution to divide the TEC map and ROTI map of each frame in the multi-frame dual-channel input tensor into non-overlapping image blocks, and projects each image block into the visual token space to obtain the initial token of each image block in each frame. The location encoding layer adds spatial location information and temporal order information to the initial token of each image block to obtain the encoded visual token sequence; The visual Transformer encoding subject receives the encoded visual token sequence and extracts the spatial texture features of the TEC map and ROTI map through a multi-head self-attention mechanism and a feedforward multilayer perceptron to obtain the spatial feature tokens of each spatial image patch in each time frame. The spatial texture features include the TEC gradient structure, the ROTI abnormal patch morphology and the spatial distribution pattern of the perturbation region. The temporal aggregation module receives spatial feature tokens of each spatial image block in each time frame, performs temporal Transformer attention operation on the token sequence of each spatial image block in N time frames, captures the temporal evolution trend of ionospheric perturbation, and performs average pooling on the time dimension to compress the N frames of temporal information into a single feature representation, thus obtaining a visual feature that integrates spatial texture features and temporal evolution features.
5. The ionospheric disturbance detection and GNSS positioning response method according to claim 1, characterized in that, The visual-language-action model also includes a visual-language dimension projection layer; The visual-language dimension projection layer maps the visual features output by the visual encoder from the visual space dimension to the language reasoning space dimension of the language reasoning module, and concatenates learnable cue tokens and action tokens on the projected visual feature sequence to generate the input sequence of the language reasoning module. The prompt token is used to encode the task semantics for determining the ionospheric disturbance state and generating suggestions for adjusting positioning parameters.
6. The ionospheric disturbance detection and GNSS positioning response method according to claim 1, characterized in that, The language reasoning module is trained using a LoRA low-rank adaptive fine-tuning strategy, including: The original pre-training weight matrix of the query linear layer, the key linear layer and the value linear layer of the language inference module The low-rank decomposition increment is respectively applied to the upper , to obtain the fine-tuned weight matrix ; wherein, , , The rank of the low-rank decomposition is The input dimension and the output dimension of the linear layer in the language inference module are respectively Freezing original pre-trained weight matrices during training process Training only parameters in low-rank matrices A and B.
7. The ionospheric disturbance detection and GNSS positioning response method according to claim 1, characterized in that, The action head includes a classification output layer and a regression output layer; The classification output layer outputs ionospheric disturbance warning levels, which include four levels: normal, slight disturbance, moderate disturbance, and strong disturbance. The regression output layer outputs GNSS positioning parameter adjustment suggestions, which are four-dimensional continuous action vectors, including cycle slip detection threshold increase, ROTI anomaly weight, TEC gradient weight, and time smoothing weight.
8. The ionospheric disturbance detection and GNSS positioning response method according to claim 7, characterized in that, The regression output layer calculates the four-dimensional continuous action vector as follows: Set The action token hidden state output by the language reasoning module is mapped by a fully connected linear layer to a four-dimensional continuous action vector: wherein, is a regression weight matrix, is a bias vector; The four-dimensional continuous action vector is represented as follows: wherein, a prediction of the cycle slip detection threshold step, a prediction of the ROTI anomaly weight, a prediction of the TEC gradient weight, a prediction of the time smoothing weight; The regression output layer outputs a four-dimensional continuous action vector to a GNSS positioning solution module, which determines the GNSS position of the vehicle based on the action vector and the GNSS measurements. The cycle slip detection threshold is updated and is expressed as: wherein is a reference cycle slip detection threshold coefficient, is a carrier phase noise standard deviation.
9. The ionospheric disturbance detection and GNSS positioning response method according to claim 1, characterized in that, The vision-language-action model is trained using a joint loss function, which is a weighted sum of the cross-entropy loss for warning level classification and the Smooth L1 loss for behavior parameter regression. Before training, when the training samples lack manual annotation, weakly supervised labels are constructed based on the physical statistics of the TEC and ROTI maps, including: Constructing a composite disturbance score is represented as: wherein, is the average intensity of the ROTI map at the current time, is the average absolute change of the TEC map at the current time relative to the TEC map at the previous time, is the spatial standard deviation of the TEC map at the current time, is a weight coefficient, and ; According to the comprehensive disturbance score A four-level warning label is generated by comparing with preset thresholds , respectively corresponding to normal, slight disturbance, moderate disturbance and strong disturbance; According to the warning level Generate behavioral parameter label vectors , is represented as: in, Labels indicating the increase in the cycle slip detection threshold. Labels representing ROTI outlier weights. Labels representing TEC gradient weights Labels representing time smoothing weights, This is a preset constant.
10. An ionospheric disturbance detection and GNSS positioning response system, characterized in that, include: The data preprocessing unit is used to acquire observation data from the Global Navigation Satellite System and construct TEC and ROTI maps based on this observation data. The model inference unit is used to construct an N-frame dual-channel input sequence from the TEC map and ROTI map of the current time and the previous N-1 time steps, where N≥2; and input the N-frame dual-channel input sequence into a pre-trained visual-language-action model, which includes a visual encoder, a language inference module, and an action head; wherein, the visual encoder extracts spatial texture features and temporal evolution features from the TEC map and ROTI map to generate visual features; the language inference module performs cross-modal semantic inference on the visual features to generate fused features; and the action head outputs ionospheric disturbance warning level and GNSS positioning parameter adjustment suggestions based on the fused features; The parameter adjustment unit is used to trigger the GNSS positioning algorithm to switch parameters based on the GNSS positioning parameter adjustment suggestions.