Multimodal feature fusion multipath signal intelligent identification method in complex and changeable environment

By introducing multimodal feature fusion technology into the multipath signal recognition method, and using a deep learning framework to extract and fusion the satellite signals, the problem of low multipath signal recognition accuracy in complex environments is solved, and higher recognition accuracy and robustness are achieved.

CN120147803AActive Publication Date: 2025-06-13GUANGDONG UNIV OF TECH
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510233928.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-13
Estimated Expiration
2045-02-28

Smart Images

  • Figure CN120147803A_ABST
    Figure CN120147803A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal feature fusion multipath signal intelligent identification method in a complex and changeable environment. The method comprises the following steps: acquiring original satellite data and carrying out cutting preprocessing; performing image processing on the original satellite data, generating two-dimensional image data and three-dimensional image data, and integrating the two-dimensional image data and the three-dimensional image data to obtain an image data set; constructing a deep learning framework of multi-modal feature fusion, wherein the framework comprises a feature extraction module and a feature fusion module; training the deep learning framework by using an image data set and satellite observation values at corresponding moments; and converting a real-time GNSS baseband signal into image data, and inputting the image data into the trained model to complete signal classification. According to the invention, high-precision signal identification can be realized under the condition of low signal-to-noise ratio or high overlapping of signal paths. The method can be widely applied to the field of signal identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of signal recognition, and particularly to a multi-path signal intelligent recognition method with multi-modal feature fusion in a complex and changeable environment. Background Art

[0002] As the core of modern positioning and navigation technology, the Global Navigation Satellite System (GNSS) has played an important role in fields such as national defense, transportation, agriculture, communication, and disaster relief. With the rapid development of emerging technologies such as intelligent driving and unmanned aerial vehicles, the demand for high-precision GNSS positioning has been further improved. However, the multi-path effect in complex environments remains one of the main obstacles restricting GNSS positioning accuracy.

[0003] The multi-path effect can cause an offset in the arrival time estimation of satellite signals, resulting in errors in pseudo-range measurement and carrier phase measurement, thereby directly affecting the positioning accuracy of the receiver. For critical application scenarios, such as an autonomous driving vehicle navigating in an urban canyon or a drone performing tasks in a forest environment, traditional GNSS receiving technologies are difficult to meet their requirements for centimeter-level or even millimeter-level positioning accuracy. Studying how to effectively distinguish pure direct signals from direct satellite signals contaminated by multi-paths has become a key technical path to solve the multi-path problem.

[0004] Most of the existing GNSS multi-path signal recognition methods are implemented based on deep learning models. Since they do not pay attention to the satellite observation value modality and do not compare the relationships between image features in multiple dimensions, the classification in a complex environment with a lot of noise is easily greatly affected, resulting in limited classification accuracy in a complex and changeable environment. Summary of the Invention

[0005] In view of this, in order to solve the technical problem that the existing multi-path signal recognition methods do not introduce multi-modal features for complex environments, and thus have low classification accuracy, the present invention proposes a multi-path signal intelligent recognition method with multi-modal feature fusion in a complex and changeable environment. The method includes the following steps:

[0006] Obtain the original satellite data and perform cropping preprocessing;

[0007] Perform image processing on the original satellite data, generate two-dimensional image data and three-dimensional image data, and integrate them to obtain an image dataset;

[0008] Construct a deep learning framework with multi-modal feature fusion, which includes a feature extraction module and a feature fusion module;

[0009] Train the deep learning framework using the image dataset and the satellite observation values at corresponding times;

[0010] Convert the real-time GNSS baseband signal into image data and input it into the trained model to complete signal classification.

[0011] In some embodiments, the step of acquiring the original satellite data and performing cropping preprocessing is specifically as follows:

[0012] Receive the data file, crop a large file into small files of one second according to the size, and name the files with the determined time when the satellite sends the signal.

[0013] In some embodiments, the step of performing image processing on the original satellite data, generating two-dimensional image data and three-dimensional image data, and integrating to obtain an image dataset, representing the signal in the form of an image, specifically includes:

[0014] Perform calculation on the original satellite data, and obtain the required features from the original data, including the correlation intensity, carrier characteristics, and the number of chips, and generate three-dimensional image data based on this combination;

[0015] Process the original satellite data, convert the correlation intensity into the strength of color for representation, find the number of chips at the maximum peak of the correlation intensity, and intercept a certain range of the number of chips centered on the peak to generate two-dimensional image data;

[0016] Label the generated three-dimensional image data and two-dimensional image data, and assign labels to obtain an image dataset.

[0017] In some embodiments, the step of training the deep learning framework using the image dataset and the satellite observations at the corresponding time to generate an identification model specifically includes:

[0018] Input the constructed multi-dimensional satellite signal image dataset (including three-dimensional image data and two-dimensional image data) and the satellite observation value modality at the corresponding time into the deep learning framework;

[0019] Based on the feature extraction module, perform feature extraction on the three-dimensional image data and the two-dimensional image data respectively to obtain two-dimensional image features and three-dimensional image features. Among them, the feature extraction module includes a convolutional neural network, a convolutional attention network, and a feedforward neural network;

[0020] Based on the feature fusion module, process the two-dimensional image features, the three-dimensional image features, and the satellite observation value modality at the corresponding time. Among them, the feature fusion module includes a multi-head attention network, a multi-dimensional image feature fusion network, a self-attention block, a cross-attention block, and a fully connected neural network;

[0021] Learn the differences in the features of the pure direct signal and the satellite signal interfered by multipath in the three-dimensional and two-dimensional dimensions, and adjust the model parameters in combination with the loss function.

[0022] Based on the above solution, the present invention provides an intelligent multi-path signal recognition method with multi-modal feature fusion in a complex and variable environment. The GNSS baseband signal is converted into two-dimensional and three-dimensional images, and image processing technology is used for the recognition and classification of multi-path signals. Depending on multi-dimensional image features, signal classification can be more effectively carried out through a deep learning model. In addition, through the multi-modal feature fusion technology, the present invention has stronger adaptability and robustness in complex and dynamic environments, and can still maintain a high recognition accuracy under the conditions of low signal-to-noise ratio or highly overlapping signal paths, solving the problem that traditional methods fail in complex environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 is a flowchart of the steps of an intelligent multi-path signal recognition method with multi-modal feature fusion in a complex and variable environment according to the present invention;

[0024] Figure 2 is a schematic diagram of the construction process of the image data set according to the present invention;

[0025] Figure 3 is a schematic diagram of the steps of generating two-dimensional image data in a specific embodiment of the present invention;

[0026] Figure 4 is a schematic three-dimensional image diagram of a pure direct signal according to the present invention;

[0027] Figure 5 is a schematic two-dimensional image diagram of a pure direct signal according to the present invention;

[0028] Figure 6 is a schematic diagram of the structure of a deep learning framework with multi-modal feature fusion according to the present invention;

[0029] Figure 7 is a schematic diagram of the data flow of the feature extraction module in a specific embodiment of the present invention;

[0030] Figure 8 is a schematic diagram of the structure of a convolutional neural network in a specific embodiment of the present invention;

[0031] Figure 9 is a schematic diagram of the structure of a convolutional attention network in a specific embodiment of the present invention;

[0032] Figure 10 is a schematic diagram of the data flow of the feature fusion module in a specific embodiment of the present invention;

[0033] Figure 11 is a schematic diagram of the structure of a multi-dimensional image feature fusion network in a specific embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0034] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts belong to the scope of protection of the present application.

[0035] It should be noted that for the convenience of description, only the parts related to the relevant invention are shown in the accompanying drawings. Without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.

[0036] It should be understood that the "system", "device", "unit" and / or "module" used in the present application are a way to distinguish different components, elements, parts, portions or assemblies at different levels. However, if other words can achieve the same purpose, the word can be replaced by other expressions.

[0037] As shown in the present application and the claims, unless the context clearly indicates an exception, words such as "a", "an", "one" and / or "the" are not specifically singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of the clearly identified steps and elements, and these steps and elements do not constitute an exclusive list. The method or device may also include other steps or elements. An element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, commodity or device including the element.

[0038] In the description of the embodiments of the present application, "a plurality" means two or more than two. The following terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features.

[0039] In addition, flowcharts are used in the present application to illustrate the operations performed by the system according to the embodiments of the present application. It should be understood that the previous or subsequent operations do not necessarily need to be executed precisely in sequence. On the contrary, they can be executed in reverse order or simultaneously. At the same time, other operations can also be added to these processes, or one or several steps of operations can be removed from these processes.

[0040] Refer to Figure 1 , which is a schematic flowchart of an optional example of the multi-path signal intelligent recognition method for multi-modal feature fusion in a complex and changeable environment proposed by the present invention. This method can be applied to computer devices. The recognition method proposed in this embodiment may include but is not limited to the following steps:

[0041] Step S1: Obtain the original satellite data and construct a multi-dimensional satellite signal image dataset;

[0042] Step S2: Construct a deep learning framework for multi-modal feature fusion based on a feature extraction module and a feature fusion module;

[0043] Step S3: Train the deep learning framework based on the multi-dimensional satellite signal image dataset and the satellite observation value modality at the corresponding moment to obtain an identification model.

[0044] In some feasible embodiments, Step S1 further includes:

[0045] Preprocess the original satellite data.

[0046] For the satellite signal data files received by the labsat receiver in complex and changeable environments, such as urban complex overpass environments, urban canyons, urban forests, and urban tunnels, etc., obtain the original satellite data files in the ls3w suffix format.

[0047] Since this part of the files usually contain original satellite data files with a long time (more than thirty minutes) and are large in size, preliminary cropping processing needs to be performed through a python script. According to the satellite receiver parameter file paired with the original satellite data file, the python script will calculate the size of the data to be cropped per second based on parameters such as the sampling rate, acquisition bit depth, and number of channels of the original satellite data file collected by the receiver. The formula is:

[0048]

[0049] Where S is the sampling rate of the satellite signal data collected by the receiver, in Hz; D is the bit depth of the receiver, that is, the number of bits of each data, the number of bits for encoding, and the bit depth of general receivers is 2 bits; C is the number of channels, in bytes, so divide by 8 here for unit conversion, and the processed S in the formula is the size to be cropped per second.

[0050] Input the result into the python script. The code will crop a large file into small files of one second according to the size and name the files with the determined time when the satellite sends the signal, that is, the time format accurate to seconds. For example, for the data collected at 3:11:11 pm on October 21st, the original satellite file data for this second is named 1021031111, which is used as the index to query the label data matching the satellite data according to the file name for convenient access. Through this processing, a series of original satellite small data files in seconds are obtained for the next step of processing.

[0051] In some feasible embodiments, the construction process of the multi-dimensional satellite signal image dataset in step S1 refers to Figure 2 , and specifically includes:

[0052] First, perform preprocessing operations on signal data such as collecting, demodulating, filtering, and sampling the satellite baseband signal.

[0053] S1.1. Generation of three-dimensional image data.

[0054] Extract the correlation strength between signals, and its formula is as follows:

[0055]

[0056] Among them, S r (t) is the baseband signal after being processed into a small data file and received, S ref (t - γ) is the pseudo-random code reference signal generated locally, γ is the time delay, and T is the integration time. By sliding the time delay parameter γ, calculate the strength of the correlation function R(γ) to obtain the correlation distribution of the signal.

[0057] After that, it is necessary to calculate the carrier characteristics of the correlation distribution. The frequency and phase characteristics of the carrier can help identify the signal propagation path. When extracting the carrier characteristics, use the Fourier transform to calculate the signal spectrum, and its formula is:

[0058]

[0059] Through spectrum analysis, extract the carrier frequency f and phase information e -j2πft of the baseband signal. Combining with the time-domain characteristics x(t), the carrier characteristics can be used as the information of one dimension of the signal.

[0060] Furthermore, it is necessary to determine the number of chips. GNSS signals are usually modulated by pseudo-random codes, and the number of chips directly reflects the periodic structure of the signal in the time domain. Using the periodic characteristics of the correlation function, the number of chips N chip can be accurately extracted:

[0061]

[0062] Among them, T signal is the time of signal sampling, and T chip is the period of a single chip.

[0063] Take the above three extracted characteristics as the three-dimensional characteristics of the signal, and generate image data: the X-axis is the number of chips N chip, the Y-axis is the carrier frequency f, and the Z-axis is the correlation intensity R(γ). Through the peak shape diagram of the baseband signal formed among the three, the basic signal shape of the approximate signal during this period can be basically understood. By observing the signal peak shape gap between the pure direct signal and the direct signal interfered by multipath, the differences in the characteristics manifested in the peak shape between the two can be understood. Through this, the unknown true signal dataset can be tagged (divided into pure direct signals and direct signals interfered by multipath), laying a foundation for more in-depth classification processing later.

[0064] S1.2, Generation of two-dimensional image data, the process refers to Figure 3 .

[0065] In the first step, search for the peak with the maximum correlation intensity in the processed satellite signal sequence, and determine its chip count and carrier frequency. In the second step, convert the three-dimensional correlation intensity (Z-axis) to be represented by color intensity, that is, convert it into a two-dimensional image. In the third step, for the chip count of the peak with the maximum correlation intensity, adjust the search range to be centered on the chip count of the peak with the maximum correlation intensity, with the upper and lower limits being 1.5 chips. In the fourth step, perform relevant processing on the image to obtain a two-dimensional signal data image.

[0066] It should be noted that this process directly generates a two-dimensional image from satellite signal data, rather than generating a two-dimensional image from a three-dimensional image, so as to ensure the uniqueness, independence, and contrast of the final features of the image data in two dimensions.

[0067] Convert the three-dimensional correlation intensity to a two-dimensional image by representing the correlation intensity as color intensity. The formula is:

[0068]

[0069] Among them, I(x,y) represents the color value of the pixel point in the two-dimensional image, R max , R min are the minimum and maximum values of the correlation intensity, used for normalization processing.

[0070] For the peak R max with the maximum correlation intensity, adjust the search range of the chip count N chip , and limit it to the floating range centered on the peak. Specifically, it is expressed as:

[0071] N chip ∈[N chip,max - 1.5, N chip,max + 1.5]#(6)

[0072] At the same time, the carrier frequency f is adjusted accordingly, and the upper and lower limits of the frequency range corresponding to the peak value are taken. This can highlight the main signal characteristics and at the same time remove irrelevant noise interference.

[0073] S1.3. Assign labels.

[0074] After obtaining the two-dimensional signal data images, the next step is to label these image data for subsequent use in training the model. The purpose of data labeling is to classify the images into two categories: pure direct signals and direct signals interfered by multipath, and assign a label to each image.

[0075] By observing the characteristics in the two-dimensional images, significant differences in the peak shape and distribution between pure direct signals and multipath interference signals can be found:

[0076] Pure direct signal: When shown in a three-dimensional image, it usually shows a single peak shape, with the correlation intensity concentrated and symmetric, the peak value high and sharp. Refer to Figure 4 , for the three-dimensional image of the pure direct signal; when shown in a two-dimensional image, it is a relatively concentrated circular red pattern. Refer to Figure 5 , for the two-dimensional image of the pure direct signal.

[0077] Multipath interference signal: When shown in a three-dimensional image, it usually shows multiple peak shapes, with the correlation intensity distribution relatively dispersed, the peak value low and wide, and due to the multipath interference signal, various phenomena such as Doppler frequency shift will occur due to reflection, refraction, etc.; when shown in a two-dimensional image, it cannot remain consistent with the pure direct signal. This is an important point for identifying pure direct signals and satellite signals interfered by multipath.

[0078] Based on this difference, both the three-dimensional images and two-dimensional images obtained from the collected satellite signal data are assigned labels for subsequent input into the classification model as classification labels for classification. The higher the accuracy of labeling, the higher the classification accuracy obtained by the final classification model.

[0079] In some feasible embodiments, the deep learning framework in step S2, the specific model framework is represented as Figure 6 , and it specifically includes:

[0080] S2.1. Build a multi-dimensional image feature extraction module based on the convolutional attention network mechanism.

[0081] For the prepared two-dimensional image data and three-dimensional image data, they will be separately input into the feature extraction module based on the convolutional attention network mechanism to extract and optimize the key features of each dimension. First, the image data will pass through a convolutional neural network (CNN) with residual connections. This network gradually learns and extracts image features through multiple two-dimensional convolutional layers and pooling layers. Then, the extracted feature vectors will be input into the Convolutional Block Attention Module (CBAM) to further enhance and optimize the performance of important features using the attention mechanism. Finally, these optimized features will be input into a fully connected neural network with residual connections for deeper processing and understanding. Figure 7 It shows the data processing process of this feature extraction module and the specific structure of the multi-dimensional image feature extraction module.

[0082] Convolutional Neural Network (CNN), the structure is as Figure 8 shown. The network first receives the RGB image vector (128×128) input from the image encoder, and uses 64 convolutional kernels of size 5×5 (Conv2D) in convolutional layer 1 to perform preliminary feature extraction on the image. Subsequently, a 2×2 max pooling layer (PoolingLayer) is used to reduce the dimension of the feature map, halving its size. The calculation formula of the convolutional layer is:

[0083] W = Relu(conv2(H, S) + R) #(7)

[0084] where H is the input feature map, S is the convolutional kernel (also known as the filter), and R is the bias. After the convolution operation through the conv function, a Relu activation function is used to introduce non-linearity.

[0085] Next, 128 3×3 convolutional kernels and 256 3×3 convolutional kernels are used to perform deeper feature extraction on the feature map. These convolutional layers gradually learn more complex image features. In the final stage, the extracted features will be input into the fully connected layer (FullyConnected Layer) and flattened into a one-dimensional vector form, that is, arranging each pixel value of the feature map into a row to obtain the final one-dimensional image feature as the output result of the network.

[0086] Convolutional Block Attention Module (CBAM). After the convolutional neural network extracts the image features, they will be input into the convolutional attention network for further optimization of the image features. The specific situation is as Figure 9 shown.

[0087] Channel Attention Module (CAM). The design of the channel attention module aims to capture the global importance information in the input feature map. First, the feature map is processed by average pooling and max pooling respectively. Average pooling calculates the global average value of the feature map to obtain the overall information, while max pooling extracts the global maximum value to capture the significant regions, and finally generates two feature maps of size 1×1×C. These features are then further processed by a multi-layer perceptron (MLP) with shared weights.

[0088] The shared MLP contains two layers: the first layer uses a hidden unit number of C / r (where r is the reduction rate) and the activation function is ReLU, and the number of neurons in the second layer is C. By sharing weights, not only the number of model parameters is significantly reduced, but also the generalization ability of the model is improved. Then, the two processed feature maps are added element-wise, and the channel attention weights are generated through the sigmoid activation function. These weights are used to perform a weighting operation on the original input feature map, significantly enhancing the sensitivity and attention of the model to key features. The formula is expressed as:

[0089]

[0090] where σ represents the sigmoid function, W 0 ∈RC / r×C, W 1 ∈RC×C / r, and W 0 and W 1 are two parameters in the MLP and are both shared.

[0091] Spatial attention module (SAM). The spatial attention module aims to capture the spatial distribution information of the feature map so that the model can more precisely focus on the important regions in the input image. Its input is the feature map optimized by the channel attention module. First, two feature maps are generated by global average pooling and global max pooling respectively, which describe the importance of the features from different perspectives. Subsequently, they are concatenated in the channel dimension to form a new feature representation.

[0092] Next, the concatenated feature map undergoes a convolution operation with a size of 7×7 to reduce the number of channels to 1, generating a spatial attention feature map. This convolution operation can capture the local and global spatial relationships in the feature map and enhance the representation ability. Finally, the convolution result is normalized through the sigmoid activation function to generate the spatial attention weights. After the input feature map and this weight are multiplied element-wise, the generated output feature pays more attention to the regions with significant spatial information in the input image. The formula is expressed as:

[0093]

[0094] where σ represents the sigmoid function, and f 7×7 represents that the size of the filter is 7×7.

[0095] Feed Forward Neural Network (FFN). The optimized feature vectors extracted by the attention module are input into a feed-forward network composed of fully connected layers for further processing and understanding. The feed-forward neural network can achieve efficient information expression and classification by learning the non-linear relationships of the feature vectors. Its core structure includes two fully connected layers, which are used for feature transformation and output generation respectively. The first layer realizes a linear transformation through a fully connected operation, and the weight matrix W 1 and the bias vector b 1 are used to map the input features to the hidden layer. The ReLU activation function is used to introduce non-linearity and help the network capture complex feature patterns. Subsequently, the second layer maps the hidden layer features to the final output features through another fully connected operation. The formula is as follows:

[0096] FFN(x) = Relu(xW 1 +b 1 )W 2 +b 2 #(10)

[0097] where x is the input vector, W 1 and W 2 are weight matrices, b 1 and b 2 are bias vectors. The introduction of the ReLU activation function not only increases the expression ability of the network but also avoids the gradient vanishing problem, making the training more stable and efficient. Through this structure, the feed-forward neural network can fully learn the complex relationships between features, is suitable for various task scenarios such as classification and regression, and provides a powerful feature representation ability for the subsequent decision-making module.

[0098] S2.2. Build an attention network module based on multi-modal feature fusion

[0099] Aiming at the problem that multi-modal features cannot be well fused, a multi-modal feature fusion module with a specific structure as Figure 10 shown is invented. This module mainly uses the multi-head attention mechanism and the multi-dimensional feature fusion network (MDFN) to coordinate and cooperate with each other to jointly complete the extraction operation of multi-modal features. At the same time, it performs feature fusion on the obtained multi-dimensional image features, then enhances the association between different features through the cross-attention layer, and finally inputs them into the fully connected neural network to output the final fusion feature vector of the multi-modal feature fusion and complete the extraction operation.

[0100] Multi-head attention network (Multi-headattention). In the design of this module, the multi-head attention network (Multi-head Attention) is the core component of feature extraction, especially suitable for processing multi-dimensional image modal features and satellite observation value modal features. Its role is to extract fine-grained global information from different attention subspaces, so as to better capture the relationship between multi-modal features. The input multi-modal features (such as two-dimensional image features, three-dimensional image features, and at the same time satellite observation value features (signal-to-noise ratio features)) are respectively projected into the query (Query), key (Key), and value (Value) vector spaces. The specific formulas are:

[0101] Q = W Q ×X#(11)

[0102] K = W K ×X#(12)

[0103] Q = W V ×X#(13)

[0104] Here, W Q 、W K 、W V are learned weight matrices, and X represents the input features.

[0105] Through the multi-head mechanism, the attention weights in multiple subspaces are calculated respectively, and each subspace is implemented through scaled dot-product attention:

[0106]

[0107] Among them, is a scaling factor, the purpose of which is to balance the gradient size and ensure training stability. Each attention head is calculated independently in different subspaces to capture the detailed correlation of multi-dimensional image modal features and satellite observation modal features. The formula of the multi-head mechanism is:

[0108] MultiHead(Q,K,V)=Concat(head 1 , head 2 ,……,head h )W o #(15)

[0109] Concat here means concatenating the outputs of multiple headers. o It is a transformation matrix used to fuse the results of multiple attention heads. Multi-head attention can simultaneously process the complex correlations between 2D and 3D image features and mine potential high-order dependencies. Since the attention mechanism covers all input features, the model can model the global information of the image and avoid focusing only on local areas.

[0110] Multi-dimensional image feature fusion network (MDFN). In order to solve the problem that multi-dimensional features are difficult to fuse and the correlation between two-dimensional and three-dimensional features is difficult to capture, this module designs a multi-dimensional image feature fusion network. This network can fuse two-dimensional image features with three-dimensional image features, and significantly enhance the correlation and fusion effect of different dimensional features by automatically adjusting the fusion weights between multi-dimensional features. The specific structure of this network is as follows Figure 11 The multi-dimensional image feature fusion network is shown in Figure 1. Its main modules include two-dimensional feature embedding, three-dimensional feature embedding, multi-dimensional feature fusion and attention mechanism, which are used to achieve adaptive optimization and weighted fusion of global and local information.

[0111] The network targets the features of two-dimensional images. The network embeds the features through global average pooling (GAP) to extract the global information of the two-dimensional image. The specific calculation formula is:

[0112]

[0113] Among them: X∈RH×S×R is the image feature, H is the height, S is the width, and R is the number of channels. Through global average pooling, the overall information of the two-dimensional image is embedded into a compact global feature representation to avoid information loss.

[0114] For the extracted 3D image features, the network designed a fully connected neural network (FCNN) with a residual structure for embedding. The formula is as follows:

[0115] f(y) = w 1 y + w 2 (ReLU(w 1 y))#(17)

[0116] where: y ∈ Rd is the multi-dimensional image feature; w 1 ∈ Rd*c and w 2 ∈ Rc*c are weight matrices for aligning the feature dimensions. The addition of the residual structure improves the stability of 3D feature embedding and can non-linearly enhance the 3D features to extract high-order spatial information.

[0117] After the multi-dimensional image features are embedded, a fusion feature function is used for feature fusion, and the fusion result is:

[0118] θ = g(x) + f(y)#(18)

[0119] Meanwhile, to further improve the discrimination ability of the fusion features, this module introduces an attention mechanism to dynamically adjust the weights of the global information using the Sigmoid function. The specific formula is as follows:

[0120] S = H(β, ε) = sigmoid(Q 1 (ReLU(Q 2 β)))#(19)

[0121] where H(β, ε) is the activation operation of the sigmoid function, weight matrices Q1 ∈ RC×C / r, Q2 ∈ RC / r×C; the reduction ratio r represents the proportion of the reduction in the number of channels between two fully connected layers, and r is taken as 16. After reducing the dimension of the input features through two fully connected layers and activating, the weight S is obtained, which is used to represent the attention coefficients of the model for image (2D or 3D) features in different dimensions. The weight S is applied to the input feature x and the multi-dimensional feature y respectively, and an element-wise multiplication operation is performed in the channel direction. The result is:

[0122] L = f j (f c (S, x), f c (S, y))#(20)

[0123] where L ∈ RH×S×R, fj is the addition of corresponding channel elements, and fc is the multiplication operation of corresponding channel elements.

[0124] Through the design of feature embedding and fusion functions, two-dimensional and three-dimensional image features can achieve efficient information interaction at the global and local levels. The attention mechanism dynamically adjusts the feature weights of different dimensions, effectively avoiding information loss in feature fusion. The residual structure and global pooling mechanism ensure the stable performance of the module in complex tasks and can adapt to different types of image features.

[0125] Self-Attention Block. The self-attention block is the core unit of the multi-head attention mechanism, which is used to process image features of the same dimension and establish global dependencies within the features. For multi-dimensional image features, this module uses self-attention blocks when processing two-dimensional and three-dimensional image features respectively.

[0126] By calculating the attention weights between image features of the same dimension, the self-attention block can capture the correlations within the features. The weight calculation formula is the same as formula (14). This mechanism ensures that each feature point can dynamically perceive the importance of other feature points, thereby strengthening the global representation. Multiple heads are used to capture the detailed information in different subspaces of the image features, improving the richness of feature expression.

[0127] For example, the texture information of two-dimensional images and the depth information of three-dimensional images can be separately extracted through self-attention blocks, ultimately achieving more fine-grained feature modeling. Deep enhancement is performed on two-dimensional image features to extract details such as texture and color. Spatial structure information is extracted from three-dimensional image features to enhance the representation ability of three-dimensional features.

[0128] Cross Attention Block. It aims to fuse image features from different dimensions, such as the correlations between two-dimensional and three-dimensional image features. This module realizes the deep fusion of multi-dimensional features by constructing cross-dimensional correlations. The features in one dimension are adjusted according to the features in another dimension to enhance the information exchange and complementarity between the two dimensions. The formula is expressed as:

[0129]

[0130] where x and y respectively represent the feature maps of two different-dimensional images, and W x and W y are weight matrices for dimensionality reduction, and d kis the scaling factor of the feature dimension, and the softmax is used to normalize the attention weights. To ensure the bidirectionality of the information flow, the model also performs reverse calculations, that is, using the three-dimensional image features as queries and the two-dimensional image features as keys and values, thereby further strengthening the feature interaction. The features output by the cross-attention block fuse the global context information of image features of different dimensions, which not only improves the distinctiveness of the features but also provides a more robust input for the subsequent fully connected layer.

[0131] Multimodal feature fusion loss (MFFL). To fully exploit the potential of the multi-feature fusion network (MDFN) and ensure its effective adaptation to the image classification task, this scheme proposes a loss function that combines the characteristics of multi-dimensional feature fusion, adaptive weight optimization, and the requirements of the classification task. Its form is:

[0132]

[0133] where is the standard cross-entropy loss for classification supervision; is the multi-dimensional image feature fusion regularization loss for enhancing the compactness and discriminability of the fused features; is the consistency loss of the attention mechanism for optimizing the attention distribution weights. α and β are adjustment hyperparameters that control the contributions of different loss terms.

[0134] The cross-entropy loss is the core loss function in the classification task, which is used to measure the difference between the predicted distribution and the true label distribution. Its formula is:

[0135]

[0136] where N is the number of samples, C is the number of classification categories, y ij is the true label distribution of sample i, is the predicted probability distribution of sample i. Through the cross-entropy loss, it is ensured that the network can effectively learn the class discriminant information of the input images.

[0137] To further optimize the representation ability of the fused feature (θ), a multi-dimensional feature fusion regularization loss is designed to constrain the two-dimensional feature g(x) and the three-dimensional feature f(y) to maintain specific compactness and diversity during the fusion process. Its formula is:

[0138]

[0139] where is l 2The norm is used to minimize the difference between two-dimensional and three-dimensional features; KL(g(x)|||f(x)) is the KL divergence, which is used to measure the similarity between two feature distributions; and are the balancing weights that control the contributions of the two regularization terms. By minimizing the feature difference and aligning the distributions, the representation ability after multi-dimensional feature fusion is enhanced, and the classification performance of the model is improved.

[0140] To optimize the learning effect of the attention mechanism, a regularization loss based on the consistency of attention weights is proposed. Its formula is:

[0141]

[0142] where S x and S y respectively represent the attention weights calculated for two-dimensional and three-dimensional features, is the norm of l 2 which is used to constrain the consistency of attention allocation. It ensures that the model can dynamically adjust the weights of image features in different dimensions during the process of fusing features, while avoiding the instability of attention allocation.

[0143] Finally, to ensure the balance of the contributions of different loss terms to the training process, we introduce a dynamic weight adjustment strategy to adaptively adjust the values of α and β according to the convergence of the loss during the training process:

[0144]

[0145] where t is the current training iteration number, and ∈ is a smoothing factor to avoid a zero denominator.

[0146] In some feasible embodiments, step S3 specifically includes:

[0147] For the obtained three-dimensional and two-dimensional satellite signal coherent image modal data and the satellite observation value modal (the satellite observation value used here is the signal-to-noise ratio) at the corresponding moment, this patent constructs a deep learning framework for multi-modal feature fusion. By building an image feature extraction module based on the convolutional attention mechanism, the image features of two-dimensional satellite signal images and three-dimensional satellite signal images are extracted; these features are paired with the satellite observation value modal (signal-to-noise ratio) at the corresponding moment and input into the deep learning framework for multi-modal feature fusion containing a multi-head attention network. By comparing the differences in features between three-dimensional satellite signal images and two-dimensional satellite signal images, the differences in the features of the pure direct signal and the satellite signal interfered by multipath are learned in two dimensions of three-dimensional and two-dimensional, and at the same time, learning and training are carried out by cooperating with the satellite observation value, that is, the signal-to-noise ratio, which can better explain the current satellite data features.

[0148] In summary, the present invention provides a multi-path signal intelligent recognition method for multi-modal feature fusion in a complex and variable environment. The research aims to intelligently recognize pure direct signals and satellite signals interfered by multi-path, reduce the phenomenon of impure satellite signals caused by multi-path effects in complex and variable scenarios, which in turn leads to positioning deviation, and improve the accuracy and robustness of multi-path signal recognition in complex and variable scenarios. By capturing the original satellite signals collected by the receiver, the original satellite signal dataset and the satellite observation value dataset at the corresponding time are obtained. Then, relying on the matlab platform, the satellite signal dataset is transformed to obtain a three-dimensional color intensity image of the frequency domain, the number of chips, and the correlation intensity. The satellite signal data is further processed to obtain a two-dimensional image at the corresponding time. The multi-dimensional image modality and the satellite observation value modality are simultaneously input into the deep learning framework for multi-modal feature fusion for training and intelligent recognition.

[0149] In some feasible embodiments, it further includes:

[0150] Converting the real-time GNSS baseband signal into the image data to be measured;

[0151] Inputting the image data to be measured into the recognition model to complete signal classification.

[0152] Based on the above solution

[0153] A multi-path signal intelligent recognition device for multi-modal feature fusion in a complex and variable environment:

[0154] At least one processor;

[0155] At least one memory for storing at least one program;

[0156] When the at least one program is executed by the at least one processor, the at least one processor implements the above-mentioned multi-path signal intelligent recognition method for multi-modal feature fusion in a complex and variable environment.

[0157] The content in the above method embodiments is applicable to the device embodiments of the present invention. The functions specifically implemented by the device embodiments of the present invention are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0158] A storage medium, in which processor-executable instructions are stored, and the processor-executable instructions are used to implement the above-mentioned multi-path signal intelligent recognition method for multi-modal feature fusion in a complex and variable environment when executed by the processor.

[0159] The content in the above method embodiments is applicable to this storage medium embodiment. The functions specifically implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.

[0160] The above is a specific description of the preferred embodiments of the present invention. However, the present invention is not limited to the described embodiments. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.

Claims

1. A multipath signal intelligent recognition method based on multimodal feature fusion in a complex and changing environment, characterized in that: The following steps are involved: Obtain raw satellite data and construct a multi-dimensional satellite signal image dataset; Construct a deep learning framework for multimodal feature fusion based on feature extraction module and feature fusion module; Based on the multi-dimensional satellite signal image data set and the satellite observation value modality at the corresponding time, the deep learning framework is trained to obtain a recognition model.

2. According to claim 1, the multipath signal intelligent identification method based on multimodal feature fusion under complex and changeable environment further comprises: Convert real-time GNSS baseband signals into image data to be tested; The image data to be tested is input into the recognition model to complete signal classification.

3. According to claim 2, a multipath signal intelligent identification method based on multimodal feature fusion under complex and changeable environment, characterized in that: The step of obtaining the original satellite data and constructing a multi-dimensional satellite signal image data set specifically includes: Capture raw satellite data; Solving the original satellite data, extracting correlation strength, carrier frequency and number of chips as three-dimensional features, and generating three-dimensional image data; Processing the original satellite data, converting the correlation intensity into color intensity for representation, searching with the number of chips with the maximum peak value of the correlation intensity as the center, and outputting two-dimensional image data; Labels are assigned to the three-dimensional image data and the two-dimensional image data to obtain a multi-dimensional satellite image data set.

4. According to claim 3, a multipath signal intelligent identification method based on multimodal feature fusion in a complex and changeable environment is characterized in that: The formula is expressed as follows: Among them, S r (t) is the baseband signal of the original satellite data received, S ref (t-γ) is the pseudo-random code reference signal generated locally, γ is the time delay, and T is the integration time.

5. According to claim 3, a multipath signal intelligent identification method based on multimodal feature fusion under complex and changeable environment, characterized in that: The step of processing the original satellite data, converting the correlation intensity into color intensity for representation, searching with the number of chips with the maximum peak value of the correlation intensity as the center, and outputting the two-dimensional image data specifically includes: Determining the correlation strength, the number of chips and the carrier frequency of the original number of satellites; The strength of the correlation is represented by the intensity of the color; With respect to the maximum peak value of the correlation intensity, the search range of the number of chips is adjusted to be limited to an upper and lower floating range centered on the peak value, and two-dimensional image data is output.

6. According to claim 5, a multipath signal intelligent identification method based on multimodal feature fusion in a complex and changeable environment is characterized in that: The formula for expressing the correlation strength by color intensity is as follows: Among them, I(x,y) is the color value of the pixel (x,y) in the two-dimensional image obtained after normalization, and R max is the maximum value of the correlation strength, R min is the minimum value of the correlation strength, Represents the original signal intensity value of the three-dimensional correlation intensity at the position (x, y).

7. According to claim 3, a multipath signal intelligent identification method based on multimodal feature fusion in a complex and changeable environment is characterized in that: The step of training the deep learning framework based on the multi-dimensional satellite signal image data set and the satellite observation value modality at the corresponding time to obtain the recognition model specifically includes: Inputting the multi-dimensional satellite signal image data set and the satellite observation value modality at the corresponding time into the deep learning framework; Based on the feature extraction module, feature extraction is performed on the three-dimensional image data and the two-dimensional image data to obtain two-dimensional image features and three-dimensional image features; Based on the feature fusion module, the satellite observation value modalities of the two-dimensional image features and the three-dimensional image features at corresponding moments are processed to learn the difference between the features of a pure direct signal and a satellite signal interfered by multipath; The model parameters are adjusted in combination with the loss to obtain the recognition model.

8. According to claim 7, a multipath signal intelligent identification method based on multimodal feature fusion in a complex and changeable environment is characterized in that: The feature extraction module includes a convolutional neural network, a convolutional attention network and a feedforward neural network.

9. According to claim 7, a multipath signal intelligent identification method based on multimodal feature fusion in a complex and changeable environment is characterized in that: The feature fusion module includes a multi-head attention network, a multi-dimensional image feature fusion network, a self-attention block, a cross-attention block and a fully connected neural network.

10. The multipath signal intelligent identification method of multi-modal feature fusion under complex and changeable environment according to claim 9, characterized in that: The loss function expression of the feature fusion module is as follows: in, is the standard cross entropy loss, is the multi-dimensional image feature fusion regularization loss, α and β are adjustment hyperparameters, is the consistency loss of the attention mechanism.

Citation Information

Patent Citations

  • High-precision fast acquisition method for Beidou satellite weak signal based on blind separation technology

    CN109188473A

  • GNSS deception jamming detection method and system in signal acquisition phase

    CN109782304A

  • Signal modulation identification method based on feature fusion

    CN114881092A

  • Wireless indoor positioning method and system based on multi-modal fusion and deep learning

    CN118118855A

  • Image modal multipath signal suppression method based on satellite baseband signal correlation peak

    CN118501906A