Satellite observation signal scene recognition method based on deep learning

Through intelligent algorithms based on deep learning, the Beidou-3 satellite observation data and the attention mechanism deep learning model are used to accurately identify signal interference characteristics in complex environments, solving the problems of insufficient recognition capabilities and high computing overhead in complex environments in existing technologies, and achieving higher adaptability and accuracy.

CN120086673APending Publication Date: 2025-06-03NANJING NORTH OPTICAL ELECTRONICS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411968646.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify and distinguish different signal interference characteristics such as high reflection, occlusion and multipath effect in complex environments. The calculation overhead is large, the real-time is insufficient, and the adaptability is limited, so it has failed to fully utilize the potential advantages of the Beidou system in complex environments.

Method used

Design an intelligent algorithm based on deep learning, using the observation data of the Beidou-3 satellite, and through the attention mechanism deep learning model, accurately identify environmental scenes and optimize signal processing strategies.

Benefits of technology

It significantly improves the adaptability and accuracy of the Beidou-3 satellite positioning system in diverse environments, improves the accuracy and efficiency of scene recognition, simplifies the feature extraction and classification steps, and enhances the generalization ability and robustness of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086673A_ABST
    Figure CN120086673A_ABST
Patent Text Reader

Abstract

The invention discloses a scene environment automatic identification and classification method adopting a satellite navigation observation signal, which realizes scene classification and identification at a satellite navigation receiver by combining feature extraction of the observation signal with a deep learning technology, and improves the environmental adaptability of a satellite navigation system. The method comprises the following steps: acquiring satellite navigation observation data in different scenes; acquiring and marking the observation data; and dividing a data set. And a deep learning algorithm based on an attention mechanism is adopted, and the network model is evaluated by adopting accuracy. And analyzing the real-time observation data through a neural network to obtain the category of the current scene. Compared with the prior art, the method provided by the invention only adopts the Beidou No.3 satellite observation data, does not depend on other sensor sources, utilizes a deep learning algorithm, can process observation data of more dimensions, accurately identifies various complex scene types, and significantly improves the adaptability and accuracy of the Beidou No.3 satellite positioning system in various environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of satellite navigation detection, and particularly relates to a method for automatically identifying scenarios based on satellite navigation observation signals. Background Art

[0002] With the rapid development of satellite navigation technology, the problems of the influence of satellite signals in different environments have gradually emerged, and scenario recognition technology has become a key factor in improving navigation and positioning accuracy. Especially for the Beidou-3 satellite system, how to identify and process signal interference in complex environments and optimize signal processing strategies is an important research direction at present.

[0003] In the prior art, various technical paths have been proposed for scenario recognition and signal processing in complex environments. For example, Blanco-Delgado and Nunes (2010) proposed a multi-constellation GNSS satellite selection method based on convex geometry, which improved the signal processing efficiency in a multi-satellite system by optimizing the selection of satellite subsets. However, this method mainly focuses on geometric optimization and fails to effectively cope with the interference problems caused by signal reflection, occlusion, and multipath effects in complex environments. On the other hand, Mohanty and Gao (2023) combined graph neural networks with Kalman filtering to enhance signal modeling and filtering capabilities through deep learning, and improved the positioning accuracy in complex scenarios to a certain extent. However, such methods rely heavily on high-quality training data and have high model computational overhead, making it difficult to meet real-time requirements.

[0004] In addition, some studies have attempted to perform scenario recognition through signal quality indicators (such as signal strength, carrier-to-noise ratio C / N0, etc.). Although such methods are relatively simple to implement, they are difficult to effectively cope with non-linear interference characteristics in complex scenarios, such as high-rise building reflections, tree occlusions, and multipath effects, because they can only capture linear features.

[0005] Therefore, the existing scenario recognition technologies still have obvious deficiencies in dealing with complex environments: First, the ability to distinguish scenarios is insufficient, and it is difficult to accurately identify and distinguish different signal interference characteristics such as high reflection, occlusion, and multipath effects; second, the computational overhead is large and the real-time performance is insufficient. Some methods rely on complex models and large amounts of data training, making it difficult to meet the real-time processing requirements of practical applications; third, the adaptability is limited, and the existing technologies are not well applied in multi-constellation or Beidou-3 satellite systems, and the potential advantages of the Beidou system in complex environments have not been fully utilized.

[0006] To address these issues, through the automatic recognition of environmental scenarios, the Beidou-3 system is expected to dynamically adjust signal processing strategies according to different environmental conditions to improve the system's environmental adaptability and positioning accuracy in complex scenarios. For example, in an urban canyon, after identifying areas with dense buildings, multi-path effects can be preferentially processed. However, how to efficiently and accurately identify environmental features in complex scenarios, especially the scenario adaptability problem for Beidou-3 satellite observation signals, remains a technical challenge that needs to be solved urgently. This poses new challenges and opportunities for further optimizing the navigation services of the Beidou-3 system in complex environments. Summary of the Invention

[0007] The core objective of the present invention is to design and implement an innovative intelligent algorithm that specifically identifies the current environmental scenario based on the observation data of Beidou-3 satellites. By deeply analyzing and processing Beidou-3 satellite signals, this algorithm can accurately determine the specific environment where the Beidou-3 satellite receiver is located, thereby significantly enhancing the adaptability and accuracy of the satellite navigation system in various environments, and optimizing its performance and reliability in different application scenarios.

[0008] The technical solution for achieving the objective of the present invention is as follows:

[0009] A method for identifying satellite observation signal scenarios based on deep learning, and the specific process is as follows:

[0010] Determination of satellite scenario categories;

[0011] Collection and annotation of satellite navigation observation data;

[0012] Preprocessing and data partitioning of the observation data: Remove the observation data in the satellite search stage, select corresponding features from the observation data as inputs, and partition the processed data into a training set and a test set;

[0013] Adopt an algorithm based on the attention mechanism to construct a deep learning model, use cross-entropy as the loss function, adjust appropriate hyperparameters, and load the training data set to train the model;

[0014] Evaluate the model based on the accuracy rate, and adjust the data set and model parameters according to the evaluation results;

[0015] Deploy the trained neural network model in the Beidou-3 satellite signal receiving system, and use the real-time obtained satellite observation data for scenario recognition.

[0016] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0017] (1) The present invention only uses the observation data of Beidou-3 satellites, without relying on other sensor sources. By using a deep learning algorithm based on the attention mechanism, it can process more-dimensional observation data and accurately identify various complex scene types, significantly improving the adaptability and accuracy of the Beidou-3 satellite positioning system in diverse environments.

[0018] (2) Data collection and annotation are carried out in accordance with the principle of equal proportion, ensuring the representativeness of various scenes and the balance of data, thereby enhancing the generalization ability and robustness of the model.

[0019] (3) The present invention introduces a multi-head attention mechanism, enabling the model to concurrently focus on multiple features of the observation data in different subspaces, effectively capturing key information and complex patterns in different scenarios, thereby further improving the accuracy and efficiency of scene recognition.

[0020] (4) By designing an end-to-end deep learning framework, the cumbersome feature extraction and classification steps in traditional scene recognition are simplified. The original observation data is directly used for automatic feature learning and classification, enhancing the automation degree and real-time processing ability of the scene recognition process. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 The network structure diagram of the examples given in the present invention;

[0022] Figure 2 The flowchart of the examples given in the present invention; DETAILED DESCRIPTION OF THE INVENTION

[0023] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments. According to the specific description of the following embodiments, the purpose, technical solution and advantages of the present application can be made more clear and understood. It should be understood that the embodiments for implementing the present invention are not limited to the specific embodiments described below. All other embodiments obtained by those of ordinary skill in the art based on the embodiments provided in the present application without creative efforts shall fall within the scope of protection of the present application.

[0024] In the implementation of the present invention, first, satellite observation data is collected by an in-vehicle Beidou-3 satellite receiver in different environments, including the number of visible satellites, satellite elevation angle, etc. Then, the collected data is cleaned and missing values are processed. Subsequently, the data is divided into a training set and a test set, and the weights and parameters of the neural network model based on the attention mechanism are initialized. The neural network model receives the processed training data and performs forward propagation and backward propagation to calculate the output according to the current weights and input data. The softmax function is used to convert the output into probability values, and then the cross-entropy loss function is used to evaluate the model performance. If the error of the model exceeds the preset threshold, the model parameters and the data set are adjusted and then retrained. The optimized neural network model is used to predict the test set data, and finally the scene classification result is output.

[0025] A method for satellite observation signal scene recognition based on deep learning disclosed by the present invention specifically comprises the following implementation steps:

[0026] Step 1: Determination of satellite scene categories: Considering the influences of signal occlusion and multipath effects, the scene categories are determined as urban canyon, urban block, urban square, on viaduct, under viaduct, open area, and tree-lined area; a total of 7 scene categories.

[0027] The Beidou-3 satellite signal receiver is fixed on the upper part of the vehicle platform, and different driving routes are planned to ensure coverage of all predetermined scene categories, including urban canyon, urban block, urban square, on viaduct, under viaduct, open area, and tree-lined area, and driving is ensured at different time periods to collect diverse data, and the data sample size of each scene type should be approximately the same. During driving, the vehicle-mounted device continuously collects the original observation data of the Beidou-3 satellite signal.

[0028] Step 2: Acquisition and annotation of satellite navigation observation data: First, Beidou-3 satellite signal data is collected from different locations and at different times to ensure coverage of the aforementioned seven different scene categories, and the data acquisition of each category should follow the equal proportion principle, that is, the data sample size of each scene type should be approximately the same. Subsequently, these data are exhaustively scene-annotated to indicate their corresponding scene categories.

[0029] Step 3: Preprocessing of observation data and data division: The observation data in the satellite search stage is removed, corresponding features are selected from the observation data as inputs, and the processed data is divided into a training set and a test set.

[0030] For the collected data, the observation data in the satellite search stage needs to be removed first. j represents the satellite number, and the satellite code N is adopted in this method sat , the satellite elevation angle θ sat,j , the Doppler frequency shift f Doppler,j, Carrier-to-noise ratio C / N 0 The statistical data of 0 includes the mean of the last 10 samples of the carrier-to-noise ratio and the variance of the last 10 samples Pseudorange residual ΔP, altitude h of satellite j sat,j , A total of 8 features form the input vector of the model

[0031] Satellite code N sat The satellites are numbered in the form of 5-bit binary codes. For example, the 10th satellite code is

[01010] .

[0032] Among them, the pseudorange residual ΔP is calculated by the following formula: ΔP = P obs -P calc , where P obs is the observed value of the pseudorange, and P calc is the theoretical pseudorange calculated based on the known positions of the satellite and the receiver

[0033] Select the above features as the input X of the model j , where j represents the satellite number

[0034]

[0035] In the formula, X j is the output of the satellite numbered j, j ∈ {1, 2,..., n}, where n is the number of BeiDou-3 satellites;

[0036] A single sample consists of the outputs X of satellites with different numbers j to form a complete input matrix X, denoted as X = [X 1 , X 2 , …, X n , where n is the number of BeiDou-3 satellites. For satellites that are not observed, their feature vectors need to be filled with 0

[0037] Divide the processed data above into a training set and a test set

[0038] In addition, its corresponding output label is Y. Here, 01 coding is used to represent 7 different scenarios. Each scenario is represented by a 7-bit binary vector, where one bit is 1 and the rest are 0. Each bit represents a scenario. 1 to 7 correspond to urban canyon, urban block, urban square, on viaduct, under viaduct, open area, and tree-lined area respectively. The label Y can be expressed as:

[0039] Y = [y 1 , y 2 , y 3 , y 4 , y 5 , y 6, y 7

[0040] where y 1 to y 7 respectively represent the label status of the corresponding scenarios. When a certain scenario is recognized, the value at the corresponding label position is 1, and the values at the remaining positions are 0.

[0041] Step 4: Use an algorithm based on the attention mechanism to construct a deep learning model, use cross-entropy as the loss function, adjust appropriate hyperparameters, and load the training dataset to train the model.

[0042] Step 4.1: Construct a deep learning model;

[0043] (1) Fully connected embedding layer, transfer the low-dimensional vector to the high-dimensional space

[0044] First, map the low-dimensional input feature vector to the high-dimensional space through a fully connected layer to enhance the model's ability to capture complex patterns.

[0045] The weight matrix W of the first fully connected layer 1 has a dimension of 32 * 11, and the weight matrix W of the second fully connected layer 2 has a dimension of 64 * 32. The input vector X i is linearly transformed through the weight matrix W, where W is a general term for the weight matrix and is specifically represented by W 1 and W 2 in different layers. Then, add the bias vector b, that is, WX + b. The ReLU function is used as the activation function between each layer, where ReLU(x) = max(0, x), and max(0, x) means setting all negative outputs to 0 while keeping all positive outputs unchanged, and x generally refers to the numerical value.

[0046] The classifier based on the Transformer (deep learning model) architecture has the following core components:

[0047] (2) Position encoding: Since the Transformer architecture itself does not have the ability to process the sequential relationship in sequence data, the present invention adds a position encoding module to this architecture to endow the model with the ability to process sequence order.

[0048] (3) Encoder layer: This layer is composed of multiple encoder units stacked. Each unit contains a self-attention mechanism and a feed-forward neural network. The key of the attention mechanism lies in calculating the weight distribution of the input vector, and its formula is expressed as:

[0049]

[0050] ​Among them, Q, K, and V are the Query, Key, and Value matrices respectively, representing different feature representations. d k represents the dimension of the input vector. In this process, the model calculates the attention weights of each element to all other elements, thereby capturing the global dependencies of the entire sequence. The self-attention mechanism enables the model to consider the information of all other elements in the sequence when processing each element in the sequence.

[0051] (4) Multi-Head Attention Mechanism: The attention mechanism is divided into multiple "heads", and then the outputs of these heads are concatenated. It allows the model to learn information in parallel in different subspaces. Each head learns different features in the sequence. The attention calculation formula for each head is:

[0052]

[0053] In the formula, Head m represents the attention calculation result of the m-th head, where m ∈ {1,.., h}, and h is the number of heads of the multi-head attention, represent the query, key, and value weight matrices of the m-th head respectively. After the attention calculation of independent heads, the model concatenates the outputs of different heads to form a comprehensive feature representation. This process can be expressed as:

[0054] MultiHead(Q, K, V) = Concat(Head 1 , Head 2 , …, Head h )W O

[0055] Among them, MultiHead represents the multi-head attention mechanism, which captures the features in different subspaces of the input sequence by parallelly calculating the results of multiple attention heads. Concat represents concatenating the outputs of all heads. h is the total number of heads, and W O is the weight matrix of the output layer.

[0056] (5) Output Head: The output head located at the top of the Transformer consists of one or more fully connected layers, and its function is to convert the output of the encoder layer into the final classification prediction result.

[0057] (6) Classification Layer: At the end of the entire architecture, a fully connected layer with a softmax activation function is set up to output the prediction probabilities of the model for each category. The softmax activation function formula is as follows:

[0058]

[0059] where softmax(z) k is the output of the softmax function, representing the predicted probability of the k-th class, where e is the base of the natural logarithm, and z k represents the original predicted value of the k-th element of the vector z, which is the output layer of the model, and z c is the original predicted value for class c, where c is the scene class, and C represents the total number of scene classes.

[0060] Step 4.2, Model Training: The training of the deep learning network model includes forward propagation, backward propagation, and performance evaluation using the cross-entropy loss function.

[0061] In the training phase, a preprocessed dataset is used to train the model. Configure the Transformer model (deep learning model) and set the hyperparameters: the number of encoder layers is 3, the number of heads in the multi-head attention is set to 8, the number of training epochs is set to 300, and an early stopping mechanism is added (where the model training includes an early stopping mechanism to prevent the model from overfitting on the training data), and the batch-size is set to 128.

[0062] Use cross-entropy as the loss function L:

[0063]

[0064] where c is the scene class, C is the total number of scene classes, and y c is the true vector label, and p c is the probability that the model predicts the sample belongs to the class.

[0065] Step 5: Evaluate the model based on accuracy, and adjust the dataset and model parameters according to the evaluation results;

[0066] The accuracy used to evaluate the model is:

[0067]

[0068] In the formula, Accuracy is the accuracy, TP is the true positive class, that is, the number of samples that accurately identify the scene, TN is the true negative class, representing the number of samples that the model correctly predicts as the negative class, FP is the false positive class, representing the number of samples that the model incorrectly predicts the negative class as the positive class, and FN is the false negative class, representing the number of samples that the model incorrectly predicts the positive class as the negative class.

[0069] In the model testing phase, the collected data is used to evaluate the scene recognition accuracy of the model. By comparing the predicted results of the model with the actual scenes, the accuracy of the model in different scenes can be calculated, a confusion matrix can be generated, and the model can be optimized according to the results of the performance evaluation.

[0070] Specifically, adjust the structure of the model, such as increasing or decreasing the number of encoder layers, adjusting the number of heads in the multi-head attention mechanism, or adjusting hyperparameters during training, such as the learning rate, batch size, etc. Retrain the model and evaluate it on the validation set using the new optimized parameters. If the performance of the model improves, deploy it in practical applications. Otherwise, continue with parameter adjustment and model optimization.

[0071] Step 6: Deploy the trained neural network model in the Beidou-3 satellite signal receiving system and use the real-time obtained satellite observation data for scene recognition.

[0072] During the actual environment testing phase, the specific steps include:

[0073] Deploy the trained model to the in-vehicle Beidou-3 satellite signal receiver. The receiver is installed on the top of the vehicle and fixed above the roof to ensure that it can receive signals from satellites without obstruction. During the actual operation, continuously monitor the performance of the model. Record the recognition accuracy of the model in various scenarios and promptly detect possible problems. Test the model under different road conditions and environmental scenarios, including urban canyons, urban blocks, urban squares, overpasses, under overpasses, open areas, and tree-lined areas to evaluate the performance of the model in various actual situations. During the testing process, collect the prediction results of the model and the corresponding actual scenario data for further detailed analysis and comparison of the scene recognition accuracy.

[0074] Through the above steps, the present invention constructs an in-vehicle scene recognition system based on Beidou-3 satellite signals. The system can effectively process and analyze satellite signal data through an advanced neural network model and attention mechanism, achieve accurate recognition of the vehicle's environment, and provide more accurate and reliable information for the vehicle navigation system.

Claims

1. A method for satellite observation signal scene recognition based on deep learning, characterized in that: Determination of satellite scene categories; Collection and annotation of satellite navigation observation data; Preprocessing and data division of observation data: remove the observation data in the satellite search phase, select corresponding features from the observation data as input, and divide the processed data into training set and test set; Adopt an algorithm based on the attention mechanism to build a deep learning model, use cross entropy as the loss function, adjust appropriate hyperparameters, and load the training data set to train the model; Evaluate the model based on accuracy, and adjust the data set and model parameters according to the evaluation results; The trained neural network model is deployed in the BeiDou-3 satellite signal receiving system, and scene recognition is performed using real-time satellite observation data.

2. The method for satellite observation signal scene recognition based on deep learning according to claim 1, characterized in that: In determining the satellite scene categories, the scene categories are determined as urban canyons, urban blocks, urban squares, on viaducts, under viaducts, open areas, and shaded areas.

3. The method for satellite observation signal scene recognition based on deep learning according to claim 1, characterized in that: In the collection and annotation of satellite navigation observation data, Beidou-3 satellite signal data are collected from different locations and at different times to ensure that all different scene categories are covered, and the data collection of each category should follow the principle of equal proportion, that is, the data sample size of each scene type should be the same; these data are scene labeled to indicate their corresponding scene categories.

4. The method for satellite observation signal scene recognition based on deep learning according to claim 1, characterized in that: The input vector definition of the deep learning model is: Where, X j is the output vector of the satellite numbered j, i represents the satellite number, N sat is the satellite code, θ sat,j is the satellite altitude angle, f Doppler,j is the Doppler shift, is the mean of the last 10 samples of the carrier-to-noise ratio, is the variance of the last 10 samples of the carrier-to-noise ratio, ΔP is the pseudorange residual, and h sat,j is the altitude of satellite j.

5. The method for satellite observation signal scene recognition based on deep learning according to claim 1, characterized in that: Building a deep learning model specifically includes: Fully connected embedding layer, converting low-dimensional vectors into high-dimensional space; Positional encoding module, which gives the model the ability to process sequence order; Encoder layer, each encoder layer contains a self-attention mechanism and a feed-forward neural network; Multi-head attention mechanism, split the attention mechanism into eight heads and concatenate the outputs of these heads; The output head, which consists of a fully connected layer, converts the output of the encoder layer into classification prediction results; The classification layer sets a fully connected layer with a softmax activation function to output the model's predicted probability for each category.

6. The method for satellite observation signal scene recognition based on deep learning according to claim 5, characterized in that: In the fully connected embedding layer, the input vector is linearly transformed by the weight matrix, and then the bias vector b is added, that is, the transformation matrix is: WX+b, Where W is a general term for the weight matrix, which is represented by the weight matrix W1 of the first fully connected layer and the weight matrix W2 of the second fully connected layer in different layers, X is the complete output of a single sample, and b is the bias vector.

7. The method for satellite observation signal scene recognition based on deep learning according to claim 5, characterized in that: At the encoder layer, the attention mechanism formula is expressed as: Among them, Attention is the attention mechanism function, Q, K, V are query, key and value matrices respectively, d k Represents the dimension of the vector and the softmax activation function; in this process, the model calculates the attention weight of each element to all other elements, thereby capturing the global dependencies of the entire sequence.

8. The method for satellite observation signal scene recognition based on deep learning according to claim 7, characterized in that: In the multi-head attention mechanism, the attention calculation formula for each head is: In the formula, Head m is the attention calculation result of the mth head, where m∈{1,..,h}, h is the total number of heads of multi-head attention, Represent the query, key, and value weight matrices of the mth head respectively; After the attention calculation of the independent heads, the model connects the outputs of different heads to form a comprehensive feature representation. This process is expressed as: MultiHead(Q,K,V)=Concat(Head1,Head2,…,Head h )W O Among them, MultiHead represents the multi-head attention mechanism, which captures the features of different subspaces in the input sequence by computing the results of multiple attention heads in parallel, Concat represents connecting the outputs of all heads, and W O is the weight matrix of the output layer.

9. The method for satellite observation signal scene recognition based on deep learning according to claim 5, characterized in that: Classification layer; fully connected layer with softmax activation function, used to output the model's predicted probability for each category; the softmax activation function formula is as follows: Where softmax(z) k is the output of the softmax function, representing the predicted probability of the kth category, where e is the base of the natural logarithm, z k represents the original prediction value of the model output layer of the kth element of the vector z, c is the scene category, and z c is the original prediction value of category c, and C is the total number of scene categories.

10. The method for satellite observation signal scene recognition based on deep learning according to claim 5, characterized in that: Use cross entropy as the loss function L: Where c is the scene category, C is the total number of scene categories, and y c is the true vector label, p c is the probability that the model predicts that the sample belongs to the category.