Short-wave radiation source time difference positioning method based on transfer learning and deep neural network
By generating simulation data and preprocessing of actual data, combining transfer learning and deep neural networks, a multi-scale convolution feature extraction module is designed to solve the positioning error problems caused by ionosphere time degeneration and noise interference in short-wave communication, and high-precision and robust short-wave radiation source positioning are achieved.
Patent Information
- Application Number
- CN202510334471.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-29
AI Technical Summary
In short-wave communication, due to large positioning errors caused by ionosphere time degeneration and noise interference, the positioning accuracy of neural networks is limited when the actual data volume is insufficient.
By generating simulation data and actual data for preprocessing, using transfer learning and deep neural networks for training and improvement, multi-scale one-dimensional convolution and hollow convolution feature extraction modules are designed, and feature fusion and prediction modules are combined to achieve high-precision positioning.
It improves the accuracy and adaptability of the time difference positioning of short-wave radiation sources, is suitable for a variety of communication scenarios, reduces the dependence on data volume, and improves the accuracy and robustness of positioning.
Smart Images

Figure CN120385973A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of radio monitoring and positioning, and more particularly to a new time difference positioning estimation method under short-wave skywave propagation conditions, and specifically to a short-wave radiation source time difference positioning method based on transfer learning and deep neural network. Background Art
[0002] Short wave refers to electromagnetic waves with frequencies ranging from 3 to 30 MHz. Radio communication using the short-wave frequency band is called short-wave communication. Short-wave communication can achieve long-distance transmission by relying solely on the ionosphere as a medium without using relay stations, and is widely used in military, civil aviation, meteorology, broadcasting and other fields. In recent years, with the continuous development of new short-wave communication technologies, a new system of modern short-wave communication systems has gradually taken shape. However, due to the scarcity of the short-wave frequency band as a spectrum resource, the phenomenon of radio interference in short-wave communication is increasing day by day, seriously affecting the communication quality and posing unprecedented challenges to short-wave monitoring and positioning.
[0003] Traditional short-wave signal monitoring and positioning technologies rely on national short-wave monitoring networks and are combined with related interferometers or spatial spectrum direction-finding and other technical means. However, such systems usually require large and medium-sized antenna arrays and multi-channel signal acquisition equipment, and the difficulty of site selection for building stations has become a global problem. Therefore, the sustainable development of this system faces many challenges. To address these challenges, short-wave time difference positioning technology has gradually become the focus of research. This technology synchronously collects target signals through a distributed single-antenna system, estimates the position of the emission source by using the measured time difference parameters, and then realizes positioning. This method not only has positioning advantages such as low cost, high precision and strong equipment mobility, but also has strong adaptability to complex electromagnetic environments, effectively overcoming the limitations of traditional short-wave positioning technologies. However, due to the time-varying nature of the virtual height of the ionosphere, it is difficult to accurately predict its height, which in turn leads to deviations in the prediction of the propagation distance and distance difference, resulting in positioning errors. At the same time, the propagation process may also be affected by random or burst noise. These factors together greatly restrict the accuracy of short-wave time difference positioning. In recent years, neural networks have been widely used in time difference positioning problems due to their significant advantages in solving nonlinear problems. The working principle of this method draws on the neurons of the human brain, and continuously adjusts the network weights and biases by learning complex features and patterns from a large amount of raw data to improve the prediction accuracy. However, this method has been rarely studied in the field of short-wave radiation source positioning. The main reason is that the training of neural networks usually relies on a large amount of data, while the number of available short-wave radiation source samples in practice is limited, resulting in insufficient available data. Summary of the Invention
[0004] According to the above technical problem of large positioning error in the existing positioning method, a short-wave radiation source time difference positioning method based on transfer learning and deep neural network is provided. The present invention needs to generate simulation data and collect actual data, and perform the same preprocessing on the two types of data. Then, the preprocessed simulation data is used to train the designed deep neural network, and high-precision positioning under simulation conditions is achieved through testing. Then, transfer learning technology is used to fine-tune the network parameters and improve the network according to the actual data. The preprocessed actual data is used to train the improved network to achieve parameter fine-tuning, and high-precision positioning is achieved through testing, bringing a new solution for the time difference positioning of short-wave radiation sources in practice.
[0005] The technical means adopted by the present invention are as follows:
[0006] A short-wave radiation source time difference positioning method based on transfer learning and deep neural network, comprising:
[0007] S1. Establish a short-wave time difference positioning system and use the ionosphere model to calculate relevant parameters to generate simulation data required for deep neural network training, and collect actual data through monitoring stations;
[0008] S2. Perform the same preprocessing operations on the simulation data and the actual data;
[0009] S3. Design a deep neural network, and use the preprocessed simulation data to train the deep neural network to achieve high-precision positioning under simulation conditions;
[0010] S4. Use transfer learning technology to transfer the network and parameters trained with simulation data, and improve the transferred deep neural network according to the actual situation. Then, use the preprocessed actual data to retrain the improved deep neural network to adapt to the actual data law and achieve high-precision positioning.
[0011] Further, step S1 specifically includes:
[0012] S11. Determine the propagation path of the short-wave signal, the positions of the short-wave radiation source and the receiving station;
[0013] S12. Select the International Reference Ionosphere IRI-2016 model to determine the ionospheric reflection height of the short-wave signal;
[0014] S13. Calculate the propagation distance and distance difference of the short-wave signal according to the positions of the short-wave radiation source and the receiving station, and the ionospheric reflection height.
[0015] Further, step S2 specifically includes:
[0016] S21. Normalize the data, where the noisy distance difference is normalized by the maximum value, the time is normalized by the maximum value, and the frequency is linearly normalized;
[0017] S22. Measure the noisy distance difference multiple times and take the average.
[0018] Further, in step S3, the designed deep neural network includes:
[0019] Feature extraction part: Design a multi-scale one-dimensional ordinary convolution feature extraction module and a multi-scale one-dimensional dilated convolution feature extraction module to extract local features of adjacent positions and non-adjacent positions respectively;
[0020] Feature fusion part: Design a feature fusion module to integrate information from different feature sources, including features from the feature extraction part, self-features of time, frequency, and distance difference vectors, and interaction features between time, frequency, and distance difference vectors;
[0021] Prediction part: Design a prediction module to perform localization prediction using the features from the feature fusion part.
[0022] Further, in the feature extraction part:
[0023] The multi-scale one-dimensional ordinary convolution feature extraction module is used to capture adjacent position features, including shallow feature extraction and deep feature extraction. In shallow feature extraction, multi-scale one-dimensional ordinary convolution is used, the size range of the convolution kernel is from 2 to M - 2, the stride is 1, then layer normalization and PReLU activation function are applied, and global average pooling and global max pooling operations are performed; after shallow feature extraction, deep feature extraction is carried out. Except for the convolution with a convolution kernel size of M - 2, the remaining convolutions are connected to one-dimensional ordinary convolutions with a convolution kernel size of 2 and a stride of 1, then layer normalization and PReLU activation function are applied, and global average pooling and global max pooling operations are performed;
[0024] The multi-scale one-dimensional dilated convolution feature extraction module is used to capture non-adjacent position features. Multi-scale one-dimensional dilated convolution is used, the size of the convolution kernel is fixed at 2, the stride is 1, then layer normalization and PReLU activation function are applied, and global average pooling and global max pooling operations are performed, and the dilation rate ranges from 2 to M - 2.
[0025] Further, in the feature fusion part, feature fusion is performed on the preprocessed and replicated time and frequency and the preprocessed distance difference vector through element-wise multiplication and element-wise addition to obtain the fused feature of the three; then, through the Concat splicing method, the self-features of the preprocessed time, frequency, and distance difference vector, the fused feature of the three, and the features of the feature extraction part are fused.
[0026] Further, in the prediction part, a fully connected neural network with a continuous up and down dimensional fully connected layer is used to predict the longitude and latitude coordinates of the short-wave radiation source.
[0027] Further, step S4 specifically includes:
[0028] S41. Freeze the weights and biases of the feature extraction and feature fusion parts in the simulation model to retain their relevant change trend features;
[0029] S42. Add a Dropout layer in the middle of the fully connected layer in the prediction part to avoid overfitting;
[0030] S43. Retrain the prediction part to adapt to the laws of actual data.
[0031] Compared with the prior art, the present invention has the following advantages:
[0032] 1. A short-wave radiation source time difference positioning method based on transfer learning and deep neural network provided by the present invention uses simulation data to train a network model, performs model transfer through transfer learning technology, and improves the transferred network to quickly adapt to actual data, making up for the problem of insufficient actual data volume.
[0033] 2. A short-wave radiation source time difference positioning method based on transfer learning and deep neural network provided by the present invention can effectively learn the complex features of short-wave signal propagation by designing a deep neural network, and further improve the adaptability and positioning accuracy of the model to actual data in combination with transfer learning technology.
[0034] 3. A short-wave radiation source time difference positioning method based on transfer learning and deep neural network provided by the present invention is applicable to a variety of short-wave communication scenarios and has good adaptability to different modulation methods and signal environments.
[0035] For the above reasons, the present invention can be widely promoted in the fields of radio monitoring and positioning, etc. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0037] Figure 1 It is a schematic diagram of the skywave propagation of short-wave signals provided by an embodiment of the present invention;
[0038] Figure 2 It is a detailed schematic diagram of any short-wave radiation source propagating to any receiving station provided by an embodiment of the present invention;
[0039] Figure 3 It is the overall flowchart of the method provided by an embodiment of the present invention;
[0040] Figure 4 It is a schematic diagram of a multi-scale one-dimensional ordinary convolution feature extraction module provided by an embodiment of the present invention;
[0041] Figure 5 It is a schematic diagram of a multi-scale one-dimensional dilated convolution feature extraction module provided by an embodiment of the present invention;
[0042] Figure 6 It is a process diagram of the feature fusion part provided by an embodiment of the present invention;
[0043] Figure 7 It is a process diagram of the prediction part provided by an embodiment of the present invention;
[0044] Figure 8 It is a process diagram of the improved prediction part provided by an embodiment of the present invention;
[0045] Figure 9 It is a schematic diagram of the receiving station location provided by an embodiment of the present invention;
[0046] Figure 10 It is the positioning average error of different numbers of receiving stations at different signal-to-noise ratios provided by an embodiment of the present invention;
[0047] Figure 11 It is a schematic diagram of the actual short-wave signal position provided by an embodiment of the present invention. Detailed implementation manners
[0048] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0049] It should be noted that the terms "first", "second", etc. in the description, claims and above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0050] The present invention provides a short-wave radiation source time difference positioning method based on transfer learning and deep neural network, including:
[0051] S1. Establish a short-wave time difference positioning system and use the ionospheric model to calculate relevant parameters to generate simulation data required for deep neural network training, and collect actual data through monitoring stations;
[0052] Specifically, as a preferred embodiment of the present invention, step S1 specifically includes:
[0053] S11. Determine the propagation path of the short-wave signal, the positions of the short-wave radiation source and the receiving station;
[0054] S12. Select the International Reference Ionosphere IRI-2016 model to determine the ionospheric reflection height of the short-wave signal;
[0055] S13. Calculate the propagation distance and distance difference of the short-wave signal according to the positions of the short-wave radiation source and the receiving station, and the ionospheric reflection height.
[0056] In this embodiment, the short-wave signal propagates in the form of sky wave and reaches the receiving station after being reflected by the ionosphere. Within a range of 4000 kilometers, it can be considered that the short-wave signal is reflected by the ionosphere once. If the target is located within the country, a single-hop model can be used for positioning, and its sky wave propagation path schematic diagram is as Figure 1 shown. Wherein, T j represents the short-wave radiation source, j = 1, 2, 3... J, J represents the number of short-wave radiation sources, and R m represents the receiving station, m = 1, 2, 3... M, and M represents the number of receiving stations.
[0057] In addition, in short-wave time difference positioning, in addition to the longitude and latitude of the short-wave radiation source, the virtual height of the ionosphere is also unknown. Therefore, assuming that the ionospheric height is consistent, at least four receiving stations are required for positioning, but in reality, the ionospheric height is not consistent, so the number of receiving stations M is greater than 4.
[0058] Meanwhile, to establish the short-wave time difference positioning equation, Figure 2 a detailed schematic diagram of the propagation of any short-wave radiation source to any receiving station is given. Figure 2 In [reference], let the longitude and latitude coordinate vector of the short-wave radiation source be represented as u j = [x j , y j T , and the longitude and latitude coordinate vector of the receiving station be represented as s m = [λ m , α m T . In addition, O represents the center of the sphere and R represents the radius of the earth. Then the great circle distance d from the j-th short-wave radiation source to the m-th receiving station j,m can be calculated using the Haversine formula in the WGS-84 coordinate system, but first, the conversion between degrees and radians is required. The calculation formula is:
[0059]
[0060] where and are the radians corresponding to x j , y j , λ m and α m respectively, and π is approximately equal to 3.14.
[0061] Then calculate the radian difference of the longitude from the j-th short-wave radiation source to the m-th receiving station and the radian difference of the latitude from the j-th short-wave radiation source to the m-th receiving station From this, the great circle distance d from the j-th short-wave radiation source to the m-th receiving station can be obtained j,m , and its calculation formula is:
[0062]
[0063] where a1 and a2 are intermediate variables in the Haversine formula for calculating d j,m . According to d j,m , the following can be obtained:
[0064]
[0065] where θ j,m is half of the central angle corresponding to d j,m , is the chord length on the sphere between the j-th short-wave radiation source and the m-th receiving station, is the distance from the midpoint of the chord length on the sphere between the j-th short-wave radiation source and the m-th receiving station to the corresponding reflection point Zj,m height. Observing the above formula, it can be seen that the ionospheric reflection height h from the j-th shortwave radiation source to the m-th receiving station j,m is the only unknown parameter. During the simulation process, h can be obtained by selecting an appropriate ionospheric model j,m .
[0066] Compared with the Chapman model, QP model, NeQuick model, and Bent model, the IRI model is an empirical model constructed based on a large amount of measured data and can more accurately reflect the real ionospheric environment. Therefore, the most commonly used International Reference Ionospheric IRI-2016 (International Reference Ionospheric-2016) model is selected to obtain h j,m . Therefore, the propagation distance and distance difference between the j-th shortwave radiation source to the m-th receiving station and the reference receiving station s1 can be expressed as D j,m and ΔD j,m,1 , respectively, and their calculation formulas are as follows:
[0067]
[0068] Due to the inevitable interference of noise during the propagation process, the shortwave time difference positioning equation can be redefined as:
[0069]
[0070] where n j,m,1 represents the noise between the j-th shortwave radiation source to the m-th receiving station and the reference receiving station s1, represents the noisy distance difference between the j-th shortwave radiation source to the m-th receiving station and the reference receiving station s1. Also, because h j,m is affected by the corresponding frequency f j and time t j , and thus will affect , so not only the noisy distance difference corresponding to the shortwave radiation and the longitude and latitude coordinate vector u j of each shortwave radiation source in reality are used as the simulation data set, but also their f j and t j need to be taken into account. And the longitude and latitude coordinate vector u j of each shortwave radiation source, time t j , frequency f j and noisy distance difference and other data are provided by each monitoring station. Among them, the noisy distance difference is obtained by multiplying the time delay difference by the speed of light c to complete the establishment of the actual data set.
[0071] S2. Perform the same preprocessing operations on the simulation data and the actual data;
[0072] Specifically, as a preferred embodiment of the present invention, step S2 specifically includes:
[0073] S21. Normalize the data, where the noisy distance difference uses maximum normalization, the time uses maximum normalization, and the frequency uses linear normalization;
[0074] S22. Measure the noisy distance difference multiple times and take the average value to reduce the influence of noise on the distance difference.
[0075] In this embodiment, to improve the model prediction accuracy, the model input will be preprocessed. First, to avoid the influence of the noisy distance difference time t j and frequency f j on the model due to different dimensions, the data will be normalized. This operation can accelerate the model convergence speed, prevent gradient disappearance and gradient explosion. However, according to the characteristics of the data itself, a suitable normalization method needs to be selected. Therefore, for the noisy distance difference maximum normalization is adopted, but since noise will affect the distance difference, the average value is taken at the same time to reduce the influence of noise on the distance difference.
[0076] Specifically, the distance difference of each short-wave radiation source is measured N times separately under noise conditions, and the maximum noisy distance difference measured is denoted as while the noisy distance difference measured at the v-th measurement between the j-th short-wave radiation source and the m-th receiving station and the reference receiving station s1 is denoted as After maximum normalization, it is obtained:
[0077]
[0078] Then, the distance difference between the j-th short-wave radiation source and the m-th receiving station and the reference receiving station s1 is measured N times, and after maximum normalization and taking the average value, the noisy distance difference can be expressed as:
[0079]
[0080] In addition, the time t j is also normalized by the maximum value to obtain:
[0081]
[0082] where t max is the time maximum value, that is, 24.
[0083] For the frequency f j in terms of, because the frequency f jThe range is between 3 - 30 MHz, so linear normalization is performed on it, and its calculation formula is:
[0084]
[0085] Among them, f max and f min are the maximum and minimum values of the frequency f j respectively, that is, 30 and 3.
[0086] S3. Design a deep neural network, and use the preprocessed simulation data to train the deep neural network to achieve high-precision positioning under simulation conditions;
[0087] In this embodiment, as Figure 3 shown, the preprocessed distance difference vector and the corresponding and are used as model inputs, and u j is used as the model output. It can be seen from Figure 3 that the designed deep neural network is divided into three parts, where the first part is the feature extraction part, the second part is the feature fusion part, and the third part is the prediction part.
[0088] Next, some formulas required for deep neural network calculation will be introduced. Without specific explanation, the input is all X = [x1,..., x i ,..., x I T , with a length of I. The calculation formula for one-dimensional ordinary convolution is:
[0089]
[0090] Among them, Y i is the new feature value obtained by weighted summation of the i-th position to the i + K - 1-th position of X. w i and b i are the weights and biases of the i-th position to the i + K - 1-th position in the convolution kernel K respectively.
[0091] The calculation formula for the output length I′ after X is convolved by one-dimensional ordinary convolution is:
[0092]
[0093] Among them, P represents the padding size, and S represents the stride.
[0094] For one-dimensional dilated convolution, its calculation process and the calculation of the output length are similar to those of one-dimensional ordinary convolution. The difference lies in the size of the dilation rate d. The d of one-dimensional ordinary convolution is 1, while the d of one-dimensional dilated convolution is d > 1. Therefore, the calculation formula for one-dimensional dilated convolution is:
[0095]
[0096] Among them, Y i ′ is the new eigenvalue obtained by performing weighted summation on the i-th position to the (i+(K-1)d-1)-th position of X.
[0097] The calculation formula for the output length I″ after X passes through one-dimensional dilated convolution is:
[0098]
[0099] In addition, the calculation formula for layer normalization (LN) is:
[0100]
[0101] Among them, is the x at the i-th position in X i output after passing through LN, u is the mean value of X, σ is the standard deviation of X, and γ and β are learnable parameters, representing scaling and bias respectively.
[0102] The calculation formula for the activation function PReLU is:
[0103]
[0104] Among them, α is a parameter learned by the network, controlling the slope of the negative value part.
[0105] In addition, the calculation formulas for global max pooling (GMP) and global average pooling (GlobalAverage Pooling, GAP) are respectively:
[0106]
[0107] Among them, GMP selects the maximum value in X, and GAP calculates the mean value of X.
[0108] The output of the fully connected layer can be expressed as:
[0109] Y1 = XW + B
[0110] Among them, W and B represent the weight vector and the bias vector respectively.
[0111] Finally, element-wise multiplication, element-wise addition, and Concat splicing are also used. Suppose there are vectors X == [x1,...,x i ,...,x I T and Z = [z1,...,z i ,...,z I T After the above three feature fusions, the outputs are Y2, Y3, and Y4 respectively, and their calculation formulas are as follows:
[0112]
[0113] In specific implementation, as a preferred implementation manner of the present invention, in step S3, the designed deep neural network includes:
[0114] The feature extraction part designs a multi-scale one-dimensional ordinary convolution feature extraction module and a multi-scale one-dimensional dilated convolution feature extraction module. The multi-scale one-dimensional ordinary convolution feature extraction module is used to capture adjacent position features, including shallow feature extraction and deep feature extraction. Among them, during shallow feature extraction, multi-scale one-dimensional ordinary convolution is used, the size range of the convolution kernel is from 2 to M - 2, the stride is 1, then layer normalization and PReLU activation function are applied, and global average pooling and global max pooling operations are performed; after shallow feature extraction, deep feature extraction is carried out. Except for the convolution with a convolution kernel size of M - 2, the remaining convolutions are connected to one-dimensional ordinary convolutions with a convolution kernel size of 2 and a stride of 1, then layer normalization and PReLU activation function are applied, and global average pooling and global max pooling operations are performed; the multi-scale one-dimensional dilated convolution feature extraction module is used to capture non-adjacent position features, adopts multi-scale one-dimensional dilated convolution, the size of the convolution kernel is fixed at 2, the stride is 1, then layer normalization and PReLU activation function are applied, and global average pooling and global max pooling operations are performed, and the dilation rate ranges from 2 to M - 2. More specifically, as follows:
[0115] In this embodiment, in the feature extraction part, mainly for the preprocessed distance difference vector local feature extraction is carried out, and a multi-scale one-dimensional ordinary convolution feature extraction module is designed as Figure 4 shown, and a multi-scale one-dimensional dilated convolution feature extraction module is designed as Figure 5 shown.
[0116] In the multi-scale one-dimensional ordinary convolution feature extraction module, first use multi-scale one-dimensional ordinary convolution for local feature extraction, and the size of the convolution kernel K ranges from 2 to M - 2, aiming to capture feature information of different scales and help the network capture the local change relationship between data under different receptive fields. At the same time, in order to better retain the local feature information of the data and reduce information loss, the stride S is set to 1. However, due to the limitation of the geographical location of the receiving station, the preprocessed distance difference vector It is relatively short. To avoid interference from irrelevant features to the model, the padding method uses valid padding, that is, P = 0. In addition, after multi-scale one-dimensional ordinary convolution, LN is connected to stabilize the training process, accelerate the convergence speed, and improve the overall performance of the model. And because the data in it has positive and negative values, the activation function selects PReLU, which introduces a learnable parameter on the basis of ReLU, enabling the model to respond more flexibly to negative inputs. Therefore, the input after passing through multi-scale one-dimensional ordinary convolution, LN, and PReLU is calculated as follows:
[0117]
[0118] where Conv K (·) is multi-scale one-dimensional ordinary convolution, LN(·) is layer normalization, and PRelu(·) is the PReLU activation function. After passing through GAP and GMP, the purpose is to select global comprehensive features and global key features from the local features extracted at each scale, capturing the comprehensive information and important information between data while reducing the parameter calculation amount. The calculation formula is as follows:
[0119]
[0120] where GAP(·) is global average pooling and GMP(·) is global max pooling. However, only performing multi-scale feature extraction once on is equivalent to only performing shallow feature extraction once, and may not be able to capture the deeper feature relationships between data, thereby restricting the model effect. Therefore, in addition to the one-dimensional ordinary convolution with K = M - 2, other multi-scale one-dimensional ordinary convolutions are connected with a convolution with K = 2, LN, and PReLU after non-linearity and normalization to continue capturing more complex local features and enhancing non-linear representation, so as to achieve deeper feature learning. After the one-dimensional ordinary convolution with K = M - 2 extracts local features, the feature vector becomes too short. When using one-dimensional ordinary convolution for deep feature extraction later, it will cause the feature values of subsequent GAP and MAP to be the same, restricting the model effect. Therefore, it is no longer connected with a convolution with K = 2 later. Then, by performing GAP and MAP operations on the deep features again to capture the global comprehensive features and global key features in the deep features. The calculation process is as follows:
[0121]
[0122] Next, the shallow features and deep features at each scale and the outputs of their corresponding GAP and MAP are concatenated by Concat, and the calculation process is as follows:
[0123]
[0124] Note that in the above formula, only performs and concatenation. After the above operations, all the feature information at each scale can be obtained. Finally, the outputs at each scale are concatenated by Concat to obtain the final output denoted as output1, achieving the comprehensive capture of the adjacent local feature relationships. The formula is as follows:
[0125]
[0126] In the multi-scale one-dimensional dilated convolution feature extraction module, the overall design is similar to that of the multi-scale one-dimensional ordinary convolution feature extraction module, but there are also some differences. Specifically, multi-scale one-dimensional dilated convolution is adopted. The size of the convolution kernel K is 2, the stride S is set to 1, the padding method is valid padding (P = 0), and the dilation rate d ranges from 2 to M - 2. Subsequently, LN, PReLU, GAP, and GMP are used to improve the model performance and capture the global comprehensive features and global key features. The calculation process is as follows:
[0127]
[0128] where DConv d (·) is the one-dimensional dilated convolution with multiple dilation rates. Next, the features at each scale after LN and PReLU and the outputs of their corresponding GAP and MAP are concatenated by Concat. The calculation process is as follows:
[0129]
[0130] In the above formula, Concat(·) is the Concat concatenation. Finally, the outputs at each scale are concatenated by Concat to obtain the final output denoted as output2, achieving the comprehensive capture of the non-adjacent local feature relationships. The formula is as follows:
[0131]
[0132] In summary, using the multi-scale one-dimensional dilated convolution module and the multi-scale one-dimensional ordinary convolution module for local feature extraction, the local features of are obtained from multiple dimensions, and the captured feature information is more comprehensive.
[0133] Feature Fusion Part: Design a feature fusion module to integrate information from different feature sources, including the features from the feature extraction part, the self - features of time, frequency, and distance - difference vectors, as well as the interaction features between time, frequency, and distance - difference vectors. In the said feature fusion part, feature fusion is performed on the pre - processed and replicated time and frequency, and the pre - processed distance - difference vector through element - wise multiplication and element - wise addition to obtain the fused feature of the three; then, through the Concat splicing method, the self - features of the pre - processed time, frequency, and distance - difference vectors, the fused feature of the three, and the features from the feature extraction part are fused.
[0134] In this embodiment, in the feature fusion part, as Figure 6 shown, integrate information from different feature sources, enabling the model to more comprehensively capture the interaction relationship among the pre - processed time pre - processed frequency and pre - processed distance - difference vector when processing the input data. Specifically, since and will affect the value and the lengths of the three are inconsistent, first perform a replication operation to expand the dimensions of and to make them consistent with the dimension, and then capture the complex relationship among the three through element - wise multiplication and element - wise addition. The calculation formula is as follows:
[0135]
[0136] Among them, and are the results of element - wise addition and element - wise multiplication respectively, and are the results of and after the replication operation respectively. Then, concatenate and output1 and output2 through Concat to obtain the final output denoted as output3. The calculation formula is as follows:
[0137]
[0138] The purpose of doing this is to transfer the original features, fused features, and features from the feature extraction part to the subsequent layers together, ensuring that the network can subsequently extract richer expressions from the original features, fused features, and features from the feature extraction part, and enhancing the model's ability to model complex relationships.
[0139] In the prediction part, a prediction module is designed to perform positioning prediction using the features from the feature fusion part. In this prediction part, a fully connected neural network with continuously rising and falling dimensionality fully connected layers is adopted to predict the longitude and latitude coordinates of the short-wave radiation source.
[0140] In this embodiment, in the prediction part, a prediction model is designed as Figure 7 shown. FC represents the fully connected layer. The number of neurons in each fully connected layer from left to right is 256, 128, 256, 128, and 2 respectively. This design with continuously rising and falling dimensionality fully connected layers can extract important information more effectively, enhance the flexibility and expressive ability of the model. Compared with a simple fully connected layer with one-time rising and falling dimensionality, it can better capture the complex relationships between data, continuously and gradually abstract and compress information, thereby improving the overall performance of the model and reducing information loss, significantly improving the prediction effect. Its calculation formula is as follows:
[0141]
[0142] Among them, for the sake of simplifying the formula, PRelu(·) is changed to f(·), and w1, w2, w3, w4 and b1, b2, b3, b4 are the weights and biases of each fully connected layer from left to right except for the last fully connected layer, is the final output of the model.
[0143] S4. Use transfer learning technology to transfer the network and parameters trained with simulation data, and improve the transferred deep neural network according to the actual situation, and then retrain the improved deep neural network with the preprocessed actual data to adapt to the actual data law and achieve high-precision positioning.
[0144] Specifically in implementation, as a preferred implementation manner of the present invention, step S4 specifically includes:
[0145] S41. Freeze the weights and biases of the feature extraction and feature fusion parts in the simulation model to retain their relevant changing trend features;
[0146] S42. Add a Dropout layer in the middle of the fully connected layer in the prediction part to avoid overfitting;
[0147] S43. Retrain the prediction part to adapt to the law of the actual data.
[0148] In this embodiment, due to the high similarity between the actual task and the simulation task, model-based transfer learning is adopted, that is, the overall structure of the model is retained, and only the parameters are fine-tuned according to the actual data. However, the model needs to be improved to better adapt to the actual data. Specifically, considering that the change trends of the actual data and the simulation data are generally the same, in transfer learning, the weights and biases of the feature extraction and feature fusion parts in the simulation model are frozen to retain their relevant change trend features. In the prediction part, due to the extremely reduced amount of data, direct training may lead to overfitting, affecting the prediction results. Therefore, regularization technology will be adopted, and a dropout layer will be added in the middle of the fully connected layer. The improved prediction part model is as shown in Figure 8 shown. If the input is X, the output after passing through the dropout layer can be expressed as:
[0149] Dropout(X) = X * bernoulli(r)
[0150] where bernoulli(r) represents independent sampling from a Bernoulli distribution with a probability parameter of r. That is, the probability that each element is set to 0 is r, and the probability of remaining unchanged is 1 - r. r is the dropout rate, representing the proportion of neurons to be discarded. In actual situations, it is more complex. Even if the time frequencies are the same, there may be differences in the numerical values of the simulation data and the actual data. Therefore, the prediction part needs to be retrained to adapt to the laws of the actual data, and the final estimation result is obtained through testing. The calculation process is as follows:
[0151]
[0152] where Dropout(·) is the dropout layer.
[0153] Embodiment
[0154] To verify the feasibility of the invention, the following operations are carried out for testing.
[0155] Two points need to be noted for this invention. First, to ensure the positioning result in the actual situation, the positioning result in the simulation situation should be guaranteed first. Second, during the experiment, it is necessary to ensure that the longitude and latitude coordinates of the receiving station and the generation range of the short-wave radiation source are the same in both the simulation situation and the actual situation. For the simulation data, an AM modulated signal is generated, with a bandwidth of 9 kHz and a transmission power of 500 kW, within the area from 90°E to 122°E and from 18°N to 42°N. The number of receiving stations is set to 7, and its position schematic diagram is as shown in Figure 9. The relevant parameter settings of the IRI-2016 model required to generate the simulation data are given in Table 1 below. In this way, the simulation data can be generated, and the signal received by R1 is used as a reference.
[0156] Table 1 IRI-2016 model parameter settings
[0157]
[0158] The above operations were carried out in MATLAB R2020a. At the same time, in order to save computer resources, it is necessary to determine a reasonable setting for the number of reference sources in the simulation experiment. The simulation experiment will be carried out on a Windows 11 system computer equipped with an NVIDIA GeForce RTX 3060 graphics card, and a software simulation platform based on the TensorFlow framework will be built using the Python 3.8 programming language. The Adam optimizer is used during the training process, and the learning rate is set to 0.0001. A total of 100 epochs are trained. To sum up, when N = 20, the average positioning errors of the simulation data under the condition of SNR = 30 trained with different numbers of reference sources and tested with 100 test sources are shown in Table 2.
[0159] Table 2 Average positioning errors with different numbers of reference sources
[0160]
[0161] It can be seen that when the amount of training data reaches about 50,000, increasing the amount of data further does not significantly improve the positioning effect. Therefore, the limit positioning error of the model can reach about 15 km. Therefore, in subsequent simulation experiments, the number of reference sources is set to 50,000, and the number of test sources is set to 100. Without explicitly stating the SNR value, the default SNR = 30.
[0162] As shown in Table 3, under different signal-to-noise ratios, the average positioning error of 100 test sources gradually decreases as the signal-to-noise ratio increases, but the decreasing amplitude gradually slows down. At the same time, Figure 10 shows the positioning errors with different numbers of receiving stations under different signal-to-noise ratios. It can be seen from this that when the number of receiving stations is small, the positioning error is large, but as the number of receiving stations increases, the error gradually decreases and tends to an asymptotic limit. In order to verify the effectiveness of the feature extraction module and the feature fusion module in the proposed model, ablation experiments will be carried out as shown in Table 4, and the experimental results prove the effectiveness of each designed module.
[0163] Table 3 Average positioning errors under different signal-to-noise ratios
[0164]
[0165] Table 4 Ablation experiments on feature extraction and feature fusion
[0166]
[0167] To obtain actual data, a shortwave geolocation system can be established. Specifically, seven sets of shortwave receivers with GPS synchronization and omnidirectional antennas are deployed in seven different cities in China. The positions of the receivers are the same as those of the simulated receiving stations, as Figure 9 shown. The target signal is synchronously captured by the receivers and uploaded to the fusion center for processing in the form of baseband I / Q samples through a high-speed wired network. Subsequently, positioning calculations are performed using the processed I / Q data. At the same time, to ensure positioning accuracy, the sampling rates of all receivers should be kept consistent and greater than twice the signal bandwidth. Taking the signal received by R1 as a reference, a total of 10 shortwave data are collected, and the position schematic diagram is as Figure 1 shown. Among them, the modulation mode of the signals of T1, T2, T3, T4, T7, T9, and T 10 is AM, the modulation mode of the signal of T5 is Morse, the modulation mode of the signal of T6 is FSK, and the modulation mode of the signal of T8 is DRM. Due to the small amount of data, it is necessary to fine-tune and improve the migrated model, including freezing the weights and biases of the feature extraction part and the feature fusion part of the simulation model, and adding a Dropout layer in the prediction module. In addition, the learning rate needs to be set to 0.000001, and the epoch is set to 1000. Secondly, considering the importance of the feature extraction part and the feature fusion part to the prediction results, and the batch size batchsize = 1 during the training process, the dropout rate r of the Dropout layer is not set in the conventional range of 0.2 - 0.5, but is set to 0.0001 to minimize the loss of key information, adapt to the actual data law, and avoid the negative impact of overfitting on the model performance. Finally, it is necessary to ensure that the hardware and software facilities of the actual experiment and the simulation experiment are the same. During the actual experiment, 8 training data are randomly selected, and 2 test data are randomly selected.
[0168] As can be seen from Table 5, the UTC time, frequency f of the test data, and the minimum and maximum errors of the actual test data of the present invention are shown. It can be seen that the minimum error can reach 2.13 km, and the maximum error is only 58.36 km, which proves that the present invention can still achieve high positioning accuracy in actual situations. At the same time, Table 6 shows the comparison of the average errors of the actual data before and after transfer learning. The results show that the average positioning error without transfer learning is relatively large, verifying the effectiveness of transfer learning.
[0169] Table 5 Minimum and maximum errors of actual test data
[0170]
[0171] Table 6 Error comparison without model transfer
[0172]
[0173] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A short-wave radiation source time difference location method based on transfer learning and deep neural network, characterized in that Including: S1. Establish a short-wave time-difference positioning system, calculate relevant parameters using the ionospheric model to generate simulation data required for deep neural network training, and collect actual data through monitoring stations; S2. Perform the same preprocessing operations on the simulation data and actual data; S3. Design a deep neural network, use the preprocessed simulation data to train the deep neural network to achieve high-precision positioning under simulation conditions; S4. Use transfer learning technology to transfer the network and parameters trained with simulation data, improve the transferred deep neural network according to the actual situation, and then retrain the improved deep neural network with the preprocessed actual data to adapt to the actual data law and achieve high-precision positioning.
2. The shortwave radiation source time difference positioning method based on transfer learning and deep neural network according to claim 1, characterized in that Step S1 specifically includes: S11. Determine the propagation path of the short-wave signal, the positions of the short-wave radiation source and the receiving station; S12. Select the International Reference Ionosphere IRI-2016 model to determine the ionospheric reflection height of the short-wave signal; S13. Calculate the propagation distance and distance difference of the short-wave signal according to the positions of the short-wave radiation source and the receiving station, and the ionospheric reflection height.
3. A shortwave radiation source time difference positioning method based on transfer learning and deep neural network according to claim 1, characterized in that, Step S2 specifically includes: S21. Normalize the data. Among them, the noisy distance difference uses maximum normalization, the time uses maximum normalization, and the frequency uses linear normalization; S22. Measure the noisy distance difference multiple times and take the average value.
4. A shortwave radiation source time difference positioning method based on transfer learning and deep neural network according to claim 1, characterized in that In step S3, the designed deep neural network includes: Feature extraction part: Design a multi-scale one-dimensional ordinary convolution feature extraction module and a multi-scale one-dimensional dilated convolution feature extraction module to extract local features of adjacent positions and non-adjacent positions respectively; Feature fusion part: Design a feature fusion module to integrate information from different feature sources, including the features of the feature extraction part, the self-features of the time, frequency, and distance difference vectors, and the interaction features between the time, frequency, and distance difference vectors; Prediction part: Design a prediction module to perform positioning prediction using the features of the feature fusion part.
5. A shortwave radiation source time difference positioning method based on transfer learning and deep neural network according to claim 4, characterized in that, In the feature extraction part: The multi-scale one-dimensional ordinary convolution feature extraction module is used to capture adjacent position features, including shallow feature extraction and deep feature extraction. Among them, during shallow feature extraction, multi-scale one-dimensional ordinary convolution is used, the size range of the convolution kernel is from 2 to M - 2, the stride is 1, then layer normalization and PReLU activation function are applied, and global average pooling and global max pooling operations are performed; after shallow feature extraction, deep feature extraction is carried out. Except for the convolution with a convolution kernel size of M - 2, the rest of the convolutions are connected to one-dimensional ordinary convolutions with a convolution kernel size of 2 and a stride of 1, then layer normalization and PReLU activation function are applied, and global average pooling and global max pooling operations are performed; The multi-scale one-dimensional dilated convolution feature extraction module is used to capture non-adjacent position features. Multi-scale one-dimensional dilated convolution is used, the size of the convolution kernel is fixed at 2, the stride is 1, then layer normalization and PReLU activation function are applied, and global average pooling and global max pooling operations are performed, and the dilation rate ranges from 2 to M - 2.
6. A short-wave radiation source time difference positioning method based on transfer learning and deep neural network according to claim 4, characterized in that In the feature fusion part, feature fusion is performed on the time and frequency after preprocessing and replication operations and the preprocessed distance difference vector through element-wise multiplication and element-wise addition to obtain the fused feature of the three; Then, through the way of Concat splicing, the self-features including the preprocessed time, frequency and distance difference vector, the fused feature of the three and the features of the feature extraction part are fused.
7. A shortwave radiation source time difference positioning method based on transfer learning and deep neural network according to claim 4, characterized in that In the prediction part, a fully connected neural network with a fully connected layer of continuous dimension raising and lowering is used to predict the longitude and latitude coordinates of the short-wave radiation source.
8. A shortwave radiation source time difference positioning method based on transfer learning and deep neural network according to claim 1, characterized in that, Step S4 specifically includes: S41. Freeze the weights and biases of the feature extraction and feature fusion parts in the simulation model to retain their relevant change trend features; S42. Add a Dropout layer in the middle of the fully connected layer in the prediction part to avoid overfitting; S43. Retrain the prediction part to adapt to the law of actual data.