Unmanned aerial vehicle positioning method and device based on multi-modal fusion, and medium

Through the multimodal signal feature extraction and dimensionality reduction processing methods, different modal signals are fused to output drone positioning information, solving the accuracy and stability problems of traditional positioning methods in complex environments, and achieving more efficient anti-interference and positioning continuity.

CN119919499AInactive Publication Date: 2025-05-02INSPUR GENERSOFT CO LTD

Patent Information

Application Number
CN202510405753.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-05-02
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional drone positioning methods have accuracy and stability problems in complex environments and are susceptible to noise and environmental interference.

Method used

The drone positioning method based on multimodal fusion is adopted, and feature extraction and dimensionality reduction are performed through preset multiple modal signals (such as radar, photoelectric, audio, radio), and the signal characteristics are fused to output the positioning information of the drone.

Benefits of technology

It realizes the accuracy and stability of drone positioning in complex environments, enhances anti-interference ability through the complementarity of multi-source data, and maintains positioning continuity in specific environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119919499A_ABST
    Figure CN119919499A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned aerial vehicle positioning method and device based on multi-mode fusion and a medium, and relates to the field of unmanned aerial vehicle positioning, and the method comprises the steps: collecting corresponding mode signals based on a plurality of modes corresponding to a preset unmanned aerial vehicle; performing feature extraction on the modal signal through a pre-trained neural network to obtain a signal feature; dimensionality reduction processing is carried out on at least part of the modal signals to obtain dimensionality reduction features; performing fusion according to the signal features and the dimension reduction features to obtain fusion features; and outputting positioning information of the unmanned aerial vehicle according to the fusion features. By fusing multi-source heterogeneous data, the limitation of physical characteristics of a single sensor is broken through. Performance degradation of a single sensor in a specific scene is effectively compensated, and multi-dimensional information complementarity enhancement is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of drone positioning, and specifically to a drone positioning method, device and medium based on multi-source information. Background Art

[0002] With the advancement of science and communication technology, drones have gradually come into people's view.

[0003] Traditional drone positioning methods include radar, photoelectric, audio, and radio. However, these dimensions also have corresponding shortcomings: for example, the application of radar detection technology is limited in densely populated urban areas or complex background environments. Although photoelectric detection technology is simple and effective and can monitor small targets, the detection results are often not accurate enough. Audio detection technology is relatively simple in actual deployment and can simultaneously realize the detection and positioning functions of drones, but its accuracy and stability are easily affected by the surrounding noise environment, thereby reducing its effectiveness. Radio frequency detection technology can also effectively realize the detection and positioning of drones, but it is easily affected by environmental factors. Summary of the invention

[0004] In order to solve the above problems, this application proposes a UAV positioning method based on multimodal fusion, including: Based on the preset multiple modes corresponding to the UAV, corresponding modal signals are collected respectively; Extracting features of the modal signal through a pre-trained neural network to obtain signal features; and performing dimensionality reduction processing on at least part of the modal signal to obtain dimensionality reduction features; Fusing the signal feature and the dimension reduction feature to obtain a fusion feature; According to the fusion features, the positioning information of the UAV is output.

[0005] In one example, feature extraction is performed on the modal signal to obtain signal features, specifically including: When the modal signal is at least one of a radar signal, an audio signal, and a radio frequency signal, performing wavelet transform on the modal signal through a pre-trained neural network, and decomposing the modal signal through a wavelet basis function to obtain an approximate coefficient and a detail coefficient; Recursively decomposing the approximate coefficients, and performing noise suppression on the high-frequency parts of the detail coefficients obtained by decomposition at each layer that exceed a preset threshold; Reconstructing the approximate coefficients and the detail coefficients by inverse wavelet transform; Based on the reconstructed signal, decompose it according to the time series; According to the approximate coefficient, detail coefficient and time series decomposition result corresponding to each layer, the corresponding signal characteristics are obtained.

[0006] In one example, performing dimensionality reduction processing on at least part of the modal signal to obtain dimensionality reduction features specifically includes: When the modal signal is at least one of a radar signal, an audio signal, and a radio frequency signal, performing a wavelet scattering transform on the modal signal, and extracting corresponding approximate coefficients and detail coefficients through wavelet filtering; Recursively decomposing the approximate coefficients, and performing a modular operation on detail coefficients obtained by decomposition at each layer to obtain scattering coefficients; According to the approximate coefficient and detail coefficient corresponding to each layer and the scattering path composed of the scattering coefficient, the corresponding scattering transformation feature is obtained; The scattering transformation features are used as input data and arranged in time windows to form a high-dimensional feature matrix; Standardizing the high-dimensional feature matrix, determining the corresponding covariance matrix, and performing eigenvalue decomposition on the covariance matrix to obtain eigenvectors and eigenvalues; Based on the size of the eigenvalue, the eigenvector is selected to obtain a plurality of principal component directions; According to the selected multiple principal component directions, the high-dimensional feature matrix is ​​projected to obtain dimensionality reduction features.

[0007] In one example, feature extraction is performed on the modal signal to obtain signal features, specifically including: When the modal signal is a photoelectric signal, determining a visible light image and an infrared image contained in the photoelectric signal; Performing spatial and temporal alignment on the visible light image and the infrared image, and performing image preprocessing; The visible light image and the infrared image are respectively subjected to feature extraction through branches in a pre-trained neural network to obtain signal features.

[0008] In one example, performing spatiotemporal alignment on the visible light image and the infrared image specifically includes: Establishing a coordinate mapping relationship between a visible light acquisition device and an infrared acquisition device, and mapping the visible light image coordinates of the visible light image and the infrared image coordinates of the infrared image into the same spatial coordinate system; Performing a spatial constraint check based on the coordinate distance between the visible light image coordinates and the infrared image coordinates; and performing a temporal constraint check based on the motion trajectory and motion parameters corresponding to the visible light image coordinates and the infrared image coordinates; If the spatial constraint check and / or the temporal constraint check fails, determining an abnormality classification based on the number of failed frames; If the anomaly is classified as a temporary deviation, for the spatial constraint check, a local affine transformation matrix is ​​calculated based on the visible light image coordinates and the infrared image coordinates, and the visible light image coordinates and the infrared image coordinates are corrected; for the temporal constraint check, the motion trajectory is corrected based on trajectory filtering prediction; If the anomaly is classified as a systematic deviation, the coordinate mapping relationship is regenerated through dynamic calibration.

[0009] In one example, the signal feature and the dimension reduction feature are fused to obtain a fusion feature, which specifically includes: For a single modal signal, the features contained in the set are concatenated to obtain the modal features corresponding to the modal signal; Based on the current scene, determine a first weight corresponding to each modal signal, and based on the first weight, fuse each modal feature to obtain a fused feature; Determining positioning information of the UAV based on the fusion feature output; Determine the contribution of each modal feature in the fused feature to the positioning information, and adjust the first weight according to the contribution, so as to continue to fuse each modal feature according to the adjusted first weight to obtain the fused feature.

[0010] In one example, determining the contribution of each modal feature in the fusion feature to the positioning information and adjusting the first weight according to the contribution specifically includes: For each modal feature, a corresponding ring buffer is established, and the performance impact data in the most recent multiple frames is recorded through the ring buffer; Determine, based on the performance impact data corresponding to each frame and whether the positioning information when the modal feature is included, the contribution of the frame to the positioning information; Determine the total contribution corresponding to the modal feature based on the second weight corresponding to each frame; wherein the second weight is an exponential decay weight; Normalizing the total contribution to obtain a weight adjustment value corresponding to the first weight; The first weight is adjusted according to a preset upper limit of the weight adjustment and the weight adjustment value.

[0011] In one example, based on the performance impact data corresponding to each frame and whether the positioning information when the modal feature is included, determining the contribution of the frame to the positioning information specifically includes: Acquire first positioning information when the modal feature is included, and second positioning information when the modal feature is not included, and determine a positioning difference between the first positioning information and the second positioning information; Try to obtain the real coordinate information of the drone; If the real coordinate information is successfully obtained, determining a corresponding positioning effect according to the positioning information and the real coordinate information; Determining, according to the positioning difference and the positioning effect, a contribution of the frame to the positioning information; If the real coordinate information is not successfully obtained, the contribution of the frame to the positioning information is determined according to the positioning difference through a plurality of pre-set quantitative indicators.

[0012] On the other hand, the present application also proposes a UAV positioning device based on multimodal fusion, comprising: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the drone positioning method based on multimodal fusion as described in any of the above examples.

[0013] On the other hand, the present application also proposes a non-volatile computer storage medium storing computer executable instructions, wherein the computer executable instructions are configured to be: a drone positioning method based on multimodal fusion as described in any of the above examples.

[0014] The UAV positioning method based on multimodal fusion proposed in this application can bring the following beneficial effects: 1. By integrating heterogeneous data from multiple sources such as radar, optoelectronics, audio, and radio, the physical characteristics limitations of a single sensor have been broken through. Different modal signals form a multi-dimensional cross-validation in the space-spectrum-time domain in a complex loop, effectively compensating for the performance degradation of a single sensor in a specific scenario and achieving enhanced multi-dimensional information complementarity.

[0015] 2. By fusing multimodal features to form a composite discrimination basis, single-channel noise interference can be eliminated through cross-modal consistency testing to achieve noise suppression, thereby improving anti-interference capabilities, maintaining positioning continuity in specific environments, and achieving environmental adaptability.

[0016] 3. Achieve performance balance through a dual-path feature processing architecture. The dimension reduction branch retains the main separable features and reduces computational complexity. The full-feature branch maintains detailed information through the network structure. Feature-level fusion reduces invalid feature interactions compared to data-level fusion, thereby optimizing computational efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings: Figure 1 Schematic diagram of the process of the UAV positioning method based on multimodal fusion in the embodiment of the present application; Figure 2 This is a schematic diagram of a UAV positioning method based on multimodal fusion in one scenario in an embodiment of the present application; Figure 3 This is a schematic diagram of extracting signal features in one scenario in an embodiment of the present application; Figure 4 A schematic diagram of wavelet scattering transformation in one case in an embodiment of the present application; Figure 5 This is a schematic diagram of a fully connected layer in one scenario in an embodiment of the present application; Figure 6 This is a schematic diagram of a sigmoid function image in one case in an embodiment of the present application; Figure 7 This is a schematic diagram of a UAV positioning device based on multimodal fusion in an embodiment of the present application. DETAILED DESCRIPTION

[0018] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the present application will be clearly and completely described below in combination with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.

[0019] The technical solutions provided by various embodiments of the present application are described in detail below in conjunction with the accompanying drawings.

[0020] like Figure 1 As shown, the embodiment of the present application provides a drone positioning method based on multimodal fusion, including: S101: Based on the preset multiple modes corresponding to the UAV, corresponding modal signals are collected respectively.

[0021] like Figure 2 As shown, the pre-set multiple modes may include: radar, audio, radio, optoelectronics, etc., and the modal signals obtained therefrom may include: radar signals, audio signals, radio frequency signals, optoelectronic signals, etc.

[0022] Based on different environments, actual needs, and hardware equipment capabilities, only some of the multiple modes can be selected to collect the corresponding modal signals. The environment can be determined by the number of surrounding buildings, altitude, and number of high-altitude obstacles. The actual needs can include monitoring of drones launched by the user and countermeasures against illegal drones. The hardware equipment can include radars, microphone arrays, radio detection systems, image acquisition equipment, etc.

[0023] S102: extracting features from the modal signal using a pre-trained neural network to obtain signal features; and performing dimensionality reduction processing on at least part of the modal signal to obtain dimensionality reduction features.

[0024] The neural network may be a convolutional neural network (CNN). For different modal signals, separate neural networks may be trained to extract corresponding signal features, or signal features of different modal signals may be extracted through different branches in one neural network.

[0025] Dimensionality reduction can be achieved through principal component analysis (PCA).

[0026] Different modal signals may be processed in different ways. Figure 2 As shown, for radar signals, audio signals, and radio frequency signals, signal features can be extracted and dimension reduction can be performed to obtain reduced dimension features. For optoelectronic signals, signal features can be extracted from multiple aspects.

[0027] Specifically, Figure 3 As shown, when the modal signal is a radar signal, an audio signal, or a radio frequency signal, the modal signal is subjected to wavelet transform through a pre-trained neural network, and the modal signal is decomposed through a wavelet basis function to obtain approximate coefficients and detail coefficients.

[0028] Among them, the wavelet basis function can select the Haar wavelet basis, first perform a first-level decomposition, retain the low frequency through a low-pass filter, generate approximate coefficients (corresponding to low-frequency signals), which reflect the overall outline of the signal (for example, the uniform flight signal of a drone); extract the high frequency through a high-pass filter, and generate detail coefficients (corresponding to high-frequency signals), which capture transient changes (for example, motor noise or environmental interference).

[0029] After the first level of decomposition, the multi-layer recursive decomposition is continued to be performed, and the approximate coefficients are recursively decomposed. In the decomposition process of each level, the approximate coefficients of the previous level are decomposed into new approximate coefficients and detail coefficients.

[0030] At this time, for the detail coefficients obtained by each layer decomposition, noise suppression is performed on the high-frequency part that exceeds the preset threshold. Among them, noise suppression can be performed by setting a corresponding threshold and directly truncating the part that exceeds the threshold. Of course, the threshold can also be dynamically adjusted according to the signal characteristics to avoid excessive weakening of effective high-frequency information.

[0031] At this time, the approximate coefficients and detail coefficients corresponding to each layer are extracted as part of the signal features.

[0032] Through inverse wavelet transform, the approximate coefficients and detail coefficients are reconstructed. Only the processed approximate coefficients and the detail coefficients after noise reduction are retained, and the signal is reconstructed through inverse wavelet transform. Due to the sparsity of wavelet transform, most of the noise coefficients have been set to zero. The reconstructed signal can retain the main information and significantly reduce the noise.

[0033] At this time, based on the reconstructed signal, decomposition is performed according to the time series. For example, the reconstructed long signal is divided into short time windows, and decomposed and reconstructed frame by frame to reduce the single calculation load.

[0034] At this time, in addition to the approximate coefficients and detail coefficients corresponding to each layer, the final time series decomposition results can also be used as corresponding signal features.

[0035] Wavelet transform can accurately separate and weaken the noise components in the signal through multi-scale decomposition, which is crucial for accurately extracting valuable information from signals mixed with a lot of noise. As an efficient and fast algorithm, wavelet transform can significantly shorten the signal processing time cycle while maintaining processing accuracy, making real-time signal processing possible.

[0036] Compared with the traditional Fourier transform, the wavelet transform has shown superiority in processing non-stationary signals and nonlinear signals. Although the Fourier transform is good at processing stationary signals, its effect is often greatly reduced when facing time-varying signals or signals with significant nonlinear characteristics. The wavelet transform can adaptively adjust the analysis window to better capture the local characteristics of the signal, thereby showing higher processing efficiency and accuracy when processing such complex signals.

[0037] The collected signal is decomposed into wavelet coefficients at different scale levels using wavelet transform. These coefficients not only reflect the characteristics of the signal in different frequency domains, but also contain rich time domain information, thus forming a wavelet coefficient time-frequency diagram. In order to mine more representative features from the time-frequency diagram, a convolutional neural network (CNN) is introduced. With its powerful feature learning ability and robustness, CNN can extract useful feature information from the wavelet coefficient time-frequency diagram. By constructing a special CNN framework, accurate extraction of time-frequency diagram features is achieved.

[0038] In addition, if Figure 4 As shown, when the modal signal is a radar signal, an audio signal, or a radio frequency signal, a wavelet scattering transform is performed on the modal signal through a pre-trained neural network, and corresponding approximate coefficients and detail coefficients are extracted through wavelet filtering.

[0039] Similar to wavelet transform, the modal signal is decomposed into approximate coefficients and detail coefficients in the first layer. At this time, a modular operation is performed on the detail coefficients to obtain the scattering coefficients. The modular operation refers to finding the absolute value of the detail coefficients, thereby eliminating phase sensitivity and obtaining the scattering coefficients.

[0040] Then, similar to the wavelet transform, the approximate coefficients are recursively decomposed, and the detail coefficients obtained by each layer of decomposition are subjected to a modular operation to obtain the scattering coefficients.

[0041] At this time, all the scattering coefficients are combined, and the approximate coefficients corresponding to the last layer (since it is not further decomposed, the approximate coefficients are selected for the last layer) are combined to obtain the corresponding scattering path.

[0042] At this time, the approximate coefficient, detail coefficient, and scattering path corresponding to each layer are used as the corresponding scattering transformation features.

[0043] The scattering transform features are used as input data and arranged in time windows to form a high-dimensional feature matrix. For example, it can be an N×M-dimensional high-dimensional feature matrix, where N is the number of samples and M is the feature dimension. If 100 coefficients are generated per layer, the total dimension is 300 for a 3-layer decomposition.

[0044] The high-dimensional feature matrix is ​​standardized, including returning its mean to zero and normalizing its variance, to eliminate the influence of dimensional differences on PCA.

[0045] Determine the corresponding covariance matrix to characterize the correlation between features. The covariance matrix is ​​directly calculated from the original high-dimensional feature matrix.

[0046] Perform eigenvalue decomposition on the covariance matrix to obtain eigenvectors and eigenvalues. The eigenvectors represent the directions of the principal components, while the eigenvalues ​​represent the variances in each direction, reflecting the importance of the principal component directions.

[0047] Based on the size of the eigenvalue, the eigenvector is selected to obtain multiple principal component directions. Usually, the eigenvalues ​​are sorted from large to small, and the first k principal components are selected (for example, the cumulative variance contribution rate exceeds 95%). These directions retain the most significant change patterns in the data.

[0048] According to the selected multiple principal component directions, the high-dimensional feature matrix is ​​projected to obtain the reduced-dimensional features. For example, after the original 100-dimensional high-dimensional feature matrix is ​​reduced by PCA, the final 10-dimensional principal component covers 95% of the original information.

[0049] The wavelet scattering transform, with its unique ability of multi-level refinement, can perform layer-by-layer decomposition operations at different scale levels, which enables it to gradually and deeply extract increasingly refined feature information from the signal. Specifically, the wavelet scattering transform effectively captures the subtle changes and structural features of the signal at different scales by introducing nonlinear scattering operations on the basis of wavelet transform, thus providing a deep understanding of the intrinsic characteristics and structure of the signal.

[0050] In order to further optimize the data processing process, reduce the dimension of the feature space, and thus reduce the computational burden and improve processing efficiency, principal component analysis (PCA) data dimensionality reduction was used. The PCA method projects the original high-dimensional feature space onto the low-dimensional principal component space through linear transformation, while retaining the key information and structure in the original data as much as possible. This processing step not only simplifies the complexity of data processing, but also obtains a more refined and representative feature set. The combination of wavelet scattering transform and principal component analysis method not only improves the accuracy and depth of signal feature extraction, but also effectively simplifies the data processing process through dimensionality reduction, providing strong technical support and theoretical guarantee for practical applications in the field of signal processing and analysis.

[0051] S103: Fusing the signal feature and the dimension reduction feature to obtain a fusion feature.

[0052] Specifically, during fusion, for a single modal signal, the features contained in the signal are concatenated to obtain the modal features corresponding to the modal signal. First, for each modal signal, the internal features are fused. The fusion method can be selected as direct concatenation.

[0053] For multiple modal signals, simple splicing is no longer used. Instead, a corresponding weight is set for each modal signal, and the fusion is performed in a weighted summation manner through the weight.

[0054] Based on the current scene, a first weight corresponding to each modal signal is determined, and based on the first weight, each modal feature is fused to obtain a fused feature.

[0055] The current scene can be determined by the number of surrounding buildings, altitude, number of high-altitude obstacles, etc. For example, the more surrounding buildings and high-altitude obstacles there are, the more complex the current environment is, and the lower the first weight corresponding to the visible light signal is. Alternatively, the higher the noise volume in the surrounding environment, the lower the first weight corresponding to the audio signal. Alternatively, the more complex the surrounding radio environment is, the lower the first weight corresponding to the radio signal is. Of course, the first weight can also be initially set manually, or a corresponding default value can be set.

[0056] After the UAV is positioned according to the fusion feature, the positioning information of the UAV based on the fusion feature output is determined.

[0057] At this time, the first weight is dynamically adjusted according to the actual environment to achieve adaptive optimization and increase the adaptability of the fusion feature to the environment.

[0058] During adjustment, the contribution of each modal feature in the fused feature to the positioning information is determined, and the first weight is adjusted according to the contribution, so as to continue to fuse each modal feature according to the adjusted first weight to obtain the fused feature.

[0059] Specifically, for each modal feature, a corresponding ring buffer is established, and the performance impact data in the most recent multiple frames is recorded through the ring buffer.

[0060] Ring Buffer, also known as Circular Buffer, is a data structure that can efficiently process continuous data streams. It implements cyclic writing and reading of data through a fixed-size circular array and two moving pointers (head and tail), which is suitable for scenarios such as drone positioning that require real-time processing, memory reuse, or decoupling of production and consumption speeds.

[0061] The performance impact data may include whether the feature is used this time, the first weight when used, the final positioning effect, etc. Among them, if the drone is a controllable drone, accurate positioning information can be transmitted back by the positioning device carried by the drone itself, and the positioning effect can be determined based on the gap between the positioning information and the output positioning information. If the drone is an illegal drone, it is often necessary to counter the drone based on the output positioning information of the drone. At this time, the positioning effect can be determined by the counter-measure effect on the drone. If the illegal drone is not affected or is less affected within a certain period of time, the positioning effect is considered to be poor. If the flight trajectory of the illegal drone shows obvious shaking and instability, it is considered to be affected to a certain extent and the positioning effect is good. If the direct countermeasure is successful, the positioning effect is considered to be excellent, and the weight at this time is recorded, which can continue to be applied as the initial first weight for the next countermeasure.

[0062] At this time, based on the performance impact data corresponding to each frame and whether the positioning information when the modal feature is included, the contribution of the frame to the positioning information is determined. Among them, when the positioning effect is good, the contribution can be determined for each modal signal by removing the positioning information after the modal signal is removed and the positioning information after the modal signal is included. The larger the difference, the greater its role in the good positioning effect of this round, so the higher its contribution. When the positioning effect is poor, the difference in the positioning information of whether the modal signal is included can also be determined. At this time, on the contrary, the larger the difference, the lower the contribution.

[0063] Specifically, the first positioning information including the modal feature and the second positioning information not including the modal feature are obtained, and the positioning difference between the first positioning information and the second positioning information is determined. Both the first positioning information and the second positioning information can be obtained by changing the input of the neural network, and the positioning difference can be obtained by the Euclidean distance between the coordinates.

[0064] At this time, try to obtain the real coordinate information of the drone. For controllable drones, the real coordinate information can be fed back through the positioning module set in it, but for illegal drones, it is difficult to directly obtain the real coordinate information.

[0065] If the real coordinate information is successfully obtained, the corresponding positioning effect is determined according to the positioning information and the real coordinate information. The smaller the positioning difference between the positioning information and the real coordinate information, the better the positioning effect, which can be obtained by normalization.

[0066] According to the positioning difference and the positioning effect, the contribution of the frame to the positioning information is determined. When the positioning effect is good (higher than a certain preset value), the larger the positioning difference between the first positioning information and the second positioning information, the more serious the positioning degradation after removing the modal signal, indicating that its contribution is higher. When the positioning effect is poor (lower than a certain preset value), the larger the positioning difference between the first positioning information and the second positioning information, the more improved the positioning after removing the modal signal, indicating that its contribution is lower.

[0067] If the real coordinate information is not successfully obtained, the contribution of the frame to the positioning information is determined according to the positioning difference through multiple pre-set quantitative indicators.

[0068] Among them, the quantitative indicators can be determined by consistency score, motion rationality index, environmental matching, historical similarity, etc. Among them, the consistency score CS=1−maximum deviation of multi-source positioning / preset threshold. The maximum deviation of multi-source positioning is the maximum spatial difference between the positioning information obtained by positioning different modal signals separately. The closer the score is to 1, the higher the multi-sensor consistency and the better the positioning effect. Motion rationality index MVI=number of frames that conform to physical laws / total number of frames. Whether it conforms to physical laws can be determined by the spatial constraint verification mentioned above. The higher the index, the more consistent the trajectory is with the dynamic characteristics of the drone, and the better the positioning effect. For environmental matching, GIS (geographic information system) and DEM (digital elevation model) verification are used to count the frequency of conflicts between positioning results and the environment. The fewer conflicts, the higher the score, indicating that the positioning effect is better. Historical similarity HS=DTW similarity with the benchmark trajectory / historical maximum similarity. The higher the similarity, the more reliable the positioning result, indicating that the positioning effect is better. Among them, the full name of DTW in English is Dynamic Time Warping, and the Chinese name is dynamic time warping. It is an algorithm used to measure the similarity of two time series.

[0069] In actual work, some or all of the quantitative indicators can be selected based on needs, and each indicator can be assigned a corresponding weight, and the weighted sum can be taken after normalization to evaluate the positioning effect.

[0070] Based on this, the contribution of the modal feature in each frame can be obtained. At this time, the total contribution of the modal feature is determined based on the second weight corresponding to each frame; wherein the second weight is an exponential decay weight. In the second weight, the closer to the current time point, the higher the corresponding second weight, thereby taking the contribution at the current moment into consideration more, and obtaining the total contribution corresponding to each modal signal.

[0071] The total contribution is normalized to obtain a weight adjustment value corresponding to the first weight, for example, normalization is performed by a Sigmoid function or linear scaling. Assuming that the normalized total contribution of a modal signal is 40%, its weight adjustment value may be 0.4.

[0072] At this time, the first weight is adjusted according to the preset weight adjustment upper limit and weight adjustment value. To avoid drastic fluctuations in weight, the weight change rate is limited, for example, the weight change of a single frame does not exceed ±5%. If the difference between the weight adjustment value and the current first weight does not exceed the weight adjustment upper limit, the first weight is adjusted to the weight adjustment value, otherwise, it is adjusted to the corresponding weight adjustment upper limit.

[0073] S104: Outputting positioning information of the UAV according to the fusion features.

[0074] like Figure 5 As shown in the figure, after all the features are obtained, linear fusion is performed and a prediction model is established. In this model, a fully connected layer and a sigmoid layer are added. In order to ensure the accuracy and reliability of the model, the fully connected layer can collect the extracted features together and give a probability. Among them, the positioning information can include: longitude, latitude, altitude, safety, etc.

[0075] The feature matrix of the extracted fusion features is transformed into a 1×n matrix, and then the output result is obtained through the fully connected layer.

[0076] like Figure 5 As shown in the figure, x1, x2, and x3 are the inputs of the fully connected layer, which are transformed from feature matrices into 1×n matrices. a1, a2, and a3 are the outputs of the fully connected layer. The calculation process of a1, a2, and a3 is shown in Formula 1 to Formula 3: Formula 1; Formula 2: Formula 3; Its matrix form is shown in Formula 4: Formula 4.

[0077] in, is the weight of the i-th row and j-th column, and b1~b3 are the corresponding bias coefficients. The weight W and bias b are updated through training, and the correct weight W and bias b are finally determined according to the corresponding labels of the data set to complete the training.

[0078] like Figure 6 As shown, a sigmoid function image is provided, and the Sigmoid function expression is shown in Formula 5: Formula five.

[0079] 1. By integrating heterogeneous data from multiple sources such as radar, optoelectronics, audio, and radio, the physical characteristics limitations of a single sensor have been broken through. Different modal signals form a multi-dimensional cross-validation in the space-spectrum-time domain in a complex loop, effectively compensating for the performance degradation of a single sensor in a specific scenario and achieving enhanced multi-dimensional information complementarity.

[0080] 2. By fusing multimodal features to form a composite discrimination basis, single-channel noise interference can be eliminated through cross-modal consistency testing to achieve noise suppression, thereby improving anti-interference capabilities, maintaining positioning continuity in specific environments, and achieving environmental adaptability.

[0081] 3. Achieve performance balance through a dual-path feature processing architecture. The dimension reduction branch retains the main separable features and reduces computational complexity. The full-feature branch maintains detailed information through the network structure. Feature-level fusion reduces invalid feature interactions compared to data-level fusion, thereby optimizing computational efficiency.

[0082] In one embodiment, the modal signal may also be a photoelectric signal. Figure 2 As shown, when the modal signal is a photoelectric signal, it is only necessary to extract the signal features without performing dimensionality reduction feature processing.

[0083] Specifically, for photoelectric signals, signal features can be extracted from two dimensions.

[0084] Determine the visible light image and infrared image contained in the photoelectric signal. The visible light image and infrared image can be collected by corresponding cameras respectively.

[0085] At this time, the visible light image and the infrared image are aligned in time and space, and image preprocessing is performed.

[0086] Among them, spatiotemporal alignment can be achieved through hardware synchronization and / or software registration to ensure that the two images at the same time cover the same field of view. Image preprocessing can include resolution normalization, data enhancement, etc. Among them, resolution normalization can upsample infrared images to the resolution of visible light images (for example, 640×480) to retain the interpolation accuracy of thermal radiation values. Data enhancement can be used to add illumination changes to visible light images (for example, to simulate cloudy days and nights) and to add thermal noise to infrared images (for example, to simulate ambient temperature fluctuations) during the model training phase to enhance the generalization ability of the model.

[0087] In the neural network, there are two branches. The visible light branch can use the pre-trained ResNet-50, and modify its input layer to a 3-channel RGB mode to adapt to high-resolution texture extraction. The infrared branch can use a lightweight MobileNetV2, and modify its input layer to a 1-channel grayscale mode to adapt to single-channel thermal radiation data processing.

[0088] Through the branches in the pre-trained neural network, feature extraction is performed on the visible light image and the infrared image respectively to obtain signal features.

[0089] For visible light features, the input is an aligned RGB image, and the output is a 1024-dimensional feature vector extracted from the penultimate layer of ResNet-50, representing color, edges, and shapes (e.g., the outline of a drone propeller).

[0090] For infrared features, the input is the aligned thermal radiation grayscale image. The output is a 512-dimensional feature vector extracted from the global average pooling layer of MobileNetV2, which represents the temperature distribution and abnormal hot spots (for example, the heating area of ​​the motor).

[0091] Furthermore, when performing spatiotemporal alignment, the coordinate mapping relationship between the visible light acquisition device and the infrared acquisition device is first established. The coordinate mapping relationship can be established through a calibration plate or feature point matching (for example, SIFT algorithm). For example, the drone is controlled to hover in advance, the calibration plate image is synchronously photographed, and the affine transformation matrix is ​​calculated.

[0092] Map the visible light image coordinates of the visible light image and the infrared image coordinates of the infrared image in the same spatial coordinate system. The spatial coordinate system can be newly created, or one of the coordinate systems corresponding to the visible light image coordinates and the infrared image coordinates can be selected as the mapping spatial coordinate system. For example, map the target coordinates in the infrared image to the visible light coordinate system.

[0093] Based on the coordinate distance between the visible light image coordinates and the infrared image coordinates, a spatial constraint check is performed; and based on the motion trajectory and motion parameters corresponding to the visible light image coordinates and the infrared image coordinates, a temporal constraint check is performed.

[0094] Among them, the spatial constraint check is mainly used to detect whether the position deviation of the visible light and infrared target in the same frame exceeds the limit. The temporal constraint check is mainly used to verify whether the motion trajectory of the target in multiple consecutive frames conforms to the physical laws.

[0095] For spatial constraint verification, in the same mapped spatial coordinate system, the image coordinate distance between the positioning results for the same target in the same frame is calculated. If the distance exceeds the preset value (for example, more than 10 pixels), it is considered to have failed.

[0096] For time constraint verification, the instantaneous velocity and acceleration corresponding to each frame are determined in the most recent frames. The instantaneous velocity can be obtained based on the distance between the previous and next frames and the interval time of each frame, while the acceleration can be obtained from the adjacent instantaneous velocities.

[0097] If the acceleration exceeds the preset value (for example, set to 30 meters per second squared), or the distance within several consecutive frames exceeds the preset value (for example, the position jump exceeds 50 meters within 3 consecutive frames), it is considered that the time constraint check has not passed.

[0098] If the spatial constraint check and / or the temporal constraint check fails, the abnormal classification is determined based on the number of frames that failed. Generally speaking, if the number of failed frames is small, only a single frame or less than 10 frames failed, it can be considered a temporary deviation, which may be caused by camera shaking, temporary occlusion, etc. If the number of failed frames is large, more than 10 frames have failed in a row, it can be considered a systematic deviation, which may be caused by calibration parameter failure, hardware synchronization abnormality, etc.

[0099] If the anomaly is classified as a temporary deviation, it can be quickly corrected and the alignment restored through lightweight dynamic adjustments without interrupting the detection process.

[0100] For spatial constraint verification, the local affine transformation matrix is ​​calculated based on the visible light image coordinates and the infrared image coordinates, and the visible light image coordinates and the infrared image coordinates are corrected. For example, in the visible light image and the infrared image of the current frame, the feature points of the target area are extracted through the scale-invariant feature transform (SIFT) algorithm, and then the local affine transformation matrix is ​​calculated through the random sample consensus algorithm (RANSAC) algorithm, so as to fine-tune the infrared image coordinates to the visible light image coordinates.

[0101] For time constraint verification, the motion trajectory is corrected based on trajectory filter prediction. For example, Kalman filter prediction is used to predict the coordinates of the target in the following frames based on the coordinates of the target in the previous frames (referring to the frames that have passed the time constraint verification), thereby correcting the motion trajectory.

[0102] If the anomaly is classified as a systematic deviation, the hardware capabilities of the system can be determined first. If there is no anomaly in the hardware, the coordinate mapping relationship can be regenerated through dynamic calibration. Dynamic calibration refers to controlling the takeoff of a controllable drone carrying an LED array and hovering, using the LED array as a reference target. The visible light image acquisition device and the infrared image acquisition device simultaneously capture the reference target, calculate a new affine transformation matrix based on the position therein, and regenerate the coordinate mapping relationship.

[0103] In traditional solutions, visible light images are highly dependent on good lighting conditions. Once it encounters night, rainy days or insufficient light, the image quality will be significantly reduced, making it difficult to clearly identify the target, thereby increasing the risk of false detection and missed detection. In addition, natural factors such as reflections, shadows and obstructions also make the accuracy of visible light images in anti-UAV detection very challenging, limiting the effectiveness and reliability of visible light images in anti-UAV detection.

[0104] Considering that infrared image detection technology can overcome the limitations of lighting conditions to a certain extent. Infrared images are based on the thermal radiation characteristics of objects and are highly sensitive to targets with fever or abnormal temperatures. Similarly, the use of infrared image detection technology alone also has its inherent limitations. When the target temperature is close to the ambient temperature, the thermal anomaly signal on the infrared image will become weak, resulting in poor detection results. In addition, the spatial resolution of infrared images is relatively low, and it is difficult to provide detailed information comparable to visible light images, which to a certain extent limits its scope of application in anti-UAV detection.

[0105] Based on this, in order to overcome the limitations of a single image modality and improve the accuracy and efficiency of anti-UAV detection, the advantages of visible light images and infrared images are combined, and the key features of the two images are extracted through CNN feature extraction. In the feature extraction stage, the CNN model can automatically learn and extract the key feature information of the two images, such as color, texture, shape, and thermal radiation characteristics, to achieve effective integration of complementary information, thereby improving the robustness and accuracy of target detection.

[0106] These feature information are effectively integrated to form a more comprehensive and accurate target detection model. By integrating the high-resolution texture information of visible light images with the temperature anomaly detection capability of infrared images, accurate and efficient identification of anti-UAV detection is achieved.

[0107] like Figure 7 As shown, the embodiment of the present application also proposes a drone positioning device based on multimodal fusion, including: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the drone positioning method based on multimodal fusion as described in any of the above embodiments.

[0108] The embodiment of the present application also proposes a non-volatile computer storage medium storing computer executable instructions, wherein the computer executable instructions are configured to be: the drone positioning method based on multimodal fusion described in any of the above embodiments.

[0109] Each embodiment in this application is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device and medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.

[0110] The devices and media provided in the embodiments of the present application correspond one-to-one to the methods. Therefore, the devices and media also have similar beneficial technical effects as the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be repeated here.

[0111] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.

Claims

1. A UAV positioning method based on multimodal fusion, characterized in that: include: Based on the preset multiple modes corresponding to the UAV, corresponding modal signals are collected respectively; Extracting features of the modal signal through a pre-trained neural network to obtain signal features; and performing dimensionality reduction processing on at least part of the modal signal to obtain dimensionality reduction features; Fusing the signal feature and the dimension reduction feature to obtain a fusion feature; According to the fusion features, the positioning information of the UAV is output.

2. The method for positioning a UAV based on multimodal fusion according to claim 1, characterized in that: Feature extraction is performed on the modal signal to obtain signal features, specifically including: When the modal signal is at least one of a radar signal, an audio signal, and a radio frequency signal, performing wavelet transform on the modal signal through a pre-trained neural network, and decomposing the modal signal through a wavelet basis function to obtain an approximate coefficient and a detail coefficient; Recursively decomposing the approximate coefficients, and performing noise suppression on the high-frequency parts of the detail coefficients obtained by decomposition at each layer that exceed a preset threshold; Reconstructing the approximate coefficients and the detail coefficients by inverse wavelet transform; Based on the reconstructed signal, decompose it according to the time series; According to the approximate coefficient, detail coefficient and time series decomposition result corresponding to each layer, the corresponding signal characteristics are obtained.

3. The UAV positioning method based on multimodal fusion according to claim 1, characterized in that: Performing dimensionality reduction processing on at least part of the modal signal to obtain dimensionality reduction features, specifically including: When the modal signal is at least one of a radar signal, an audio signal, and a radio frequency signal, performing a wavelet scattering transform on the modal signal, and extracting corresponding approximate coefficients and detail coefficients through wavelet filtering; Recursively decomposing the approximate coefficients, and performing a modular operation on detail coefficients obtained by decomposition at each layer to obtain scattering coefficients; According to the approximate coefficient and detail coefficient corresponding to each layer and the scattering path composed of the scattering coefficient, the corresponding scattering transformation feature is obtained; The scattering transformation features are used as input data and arranged in time windows to form a high-dimensional feature matrix; Standardizing the high-dimensional feature matrix, determining the corresponding covariance matrix, and performing eigenvalue decomposition on the covariance matrix to obtain eigenvectors and eigenvalues; Based on the size of the eigenvalue, the eigenvector is selected to obtain a plurality of principal component directions; According to the selected multiple principal component directions, the high-dimensional feature matrix is ​​projected to obtain dimensionality reduction features.

4. The method for positioning an unmanned aerial vehicle based on multimodal fusion according to claim 1, characterized in that: Feature extraction is performed on the modal signal to obtain signal features, specifically including: When the modal signal is a photoelectric signal, determining a visible light image and an infrared image contained in the photoelectric signal; Performing spatial and temporal alignment on the visible light image and the infrared image, and performing image preprocessing; The visible light image and the infrared image are respectively subjected to feature extraction through branches in a pre-trained neural network to obtain signal features.

5. The method for positioning a UAV based on multimodal fusion according to claim 4, characterized in that: The visible light image and the infrared image are temporally and spatially aligned, specifically comprising: Establishing a coordinate mapping relationship between a visible light acquisition device and an infrared acquisition device, and mapping the visible light image coordinates of the visible light image and the infrared image coordinates of the infrared image into the same spatial coordinate system; Performing a spatial constraint check based on the coordinate distance between the visible light image coordinates and the infrared image coordinates; and performing a temporal constraint check based on the motion trajectory and motion parameters corresponding to the visible light image coordinates and the infrared image coordinates; If the spatial constraint check and / or the temporal constraint check fails, determining an abnormality classification based on the number of failed frames; If the anomaly is classified as a temporary deviation, for the spatial constraint check, a local affine transformation matrix is ​​calculated based on the visible light image coordinates and the infrared image coordinates, and the visible light image coordinates and the infrared image coordinates are corrected; for the temporal constraint check, the motion trajectory is corrected based on trajectory filtering prediction; If the anomaly is classified as a systematic deviation, the coordinate mapping relationship is regenerated through dynamic calibration.

6. The method for positioning an unmanned aerial vehicle based on multimodal fusion according to claim 1, characterized in that: According to the signal feature and the dimension reduction feature, a fusion feature is obtained by fusing, which specifically includes: For a single modal signal, the features contained in the set are concatenated to obtain the modal features corresponding to the modal signal; Based on the current scene, determine a first weight corresponding to each modal signal, and based on the first weight, fuse each modal feature to obtain a fused feature; Determining positioning information of the UAV based on the fusion feature output; Determine the contribution of each modal feature in the fused feature to the positioning information, and adjust the first weight according to the contribution, so as to continue to fuse each modal feature according to the adjusted first weight to obtain the fused feature.

7. The method for positioning a UAV based on multimodal fusion according to claim 6, characterized in that: Determining the contribution of each modal feature in the fusion feature to the positioning information, and adjusting the first weight according to the contribution, specifically includes: For each modal feature, a corresponding ring buffer is established, and the performance impact data in the most recent multiple frames is recorded through the ring buffer; Determine, based on the performance impact data corresponding to each frame and whether the positioning information when the modal feature is included, the contribution of the frame to the positioning information; Determine the total contribution corresponding to the modal feature based on the second weight corresponding to each frame; wherein the second weight is an exponential decay weight; Normalizing the total contribution to obtain a weight adjustment value corresponding to the first weight; The first weight is adjusted according to a preset upper limit of the weight adjustment and the weight adjustment value.

8. The method for positioning a UAV based on multimodal fusion according to claim 7, characterized in that: Based on the performance impact data corresponding to each frame and whether the positioning information when the modal feature is included, the contribution of the frame to the positioning information is determined, specifically including: Acquire first positioning information when the modal feature is included, and second positioning information when the modal feature is not included, and determine a positioning difference between the first positioning information and the second positioning information; Try to obtain the real coordinate information of the drone; If the real coordinate information is successfully obtained, determining a corresponding positioning effect according to the positioning information and the real coordinate information; Determining, according to the positioning difference and the positioning effect, a contribution of the frame to the positioning information; If the real coordinate information is not successfully obtained, the contribution of the frame to the positioning information is determined according to the positioning difference through a plurality of pre-set quantitative indicators.

9. A UAV positioning device based on multimodal fusion, characterized in that: include: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the drone positioning method based on multimodal fusion as described in any one of claims 1 to 8.

10. A non-volatile computer storage medium storing computer executable instructions, characterized in that: The computer executable instructions are set to: the drone positioning method based on multimodal fusion as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Multi-modal brain-computer interface data fusion method based on wavelet transform

    CN114145752A

  • UAV (unmanned aerial vehicle) cross-modal fusion detection method based on CFT-OfficientDet

    CN118485932A

  • Multi-modal fusion algorithm for electrocardiosignal anomaly detection

    CN118520279A

  • Elevator traction machine fault diagnosis algorithm based on multi-source information fusion

    CN119059386A

  • Partial discharge defect category intelligent identification method fusing optical and electrical characteristics

    CN119442096A

Cited By

  • Unmanned aerial vehicle multi-dimensional information fusion method and system based on acousto-optic-electric composite detection

    CN120257215A

  • A multi-dimensional information fusion method and system for UAV based on acoustic, optical and electrical composite detection

    CN120257215B