Multi-source big data intelligent analysis and fusion processing method based on ground-based detection equipment

Through a multi-source big data intelligent analysis and fusion processing method, the problem of multi-source data processing complexity of foundation detection equipment is solved, and the accurate feature extraction and fusion of data is realized, which improves the robustness of modeling and the reliability of prediction results.

CN120067980APending Publication Date: 2025-05-30CHINESE PEOPLES LIBERATION ARMY UNIT 63610
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510132989.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art is difficult to adapt to the complexity of data when processing multi-source data of foundation detection equipment, especially in data preprocessing, feature extraction, multimodal modeling and prediction reliability evaluation.

Method used

A multi-source big data intelligent analysis and fusion processing method based on foundation detection equipment is adopted, including data standardization, time and space alignment, missing value filling, manifold learning, high-order Gaussian process model and variational Bayesian deep network, to achieve efficient fusion and feature extraction of multi-source data.

Benefits of technology

By introducing a higher-order Gaussian process method of multimodal interactive modeling, the correlation and interaction relationship between the internal and external modes are comprehensively portrayed, the more accurate feature extraction and fusion of data is achieved, the robustness and expression ability of subsequent modeling are improved, and the reliability of prediction results is improved through uncertainty quantification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067980A_ABST
    Figure CN120067980A_ABST
Patent Text Reader

Abstract

The invention relates to the field of foundation detection and multi-source big data intelligent analysis, and discloses a multi-source big data intelligent analysis fusion processing method based on foundation detection equipment, and the method comprises the following steps: carrying out the standardization processing, sampling frequency unification, missing value filling, and completion time and space alignment of multi-source data collected by the foundation detection equipment; based on the preprocessed multi-source data, constructing a similarity weight matrix between the data points for measuring the correlation between the data points; a manifold learning algorithm is adopted to embed multi-source data into a low-dimensional manifold space, and low-dimensional embedding representation used for describing a data local geometric structure is constructed; the method is based on a high-order Gaussian process model. By introducing the high-order Gaussian process method of multi-modal interaction modeling, the internal correlation of the modals and the complex interaction relationship between the modals can be comprehensively described. The problem that the nonlinear relation of the multi-modal data cannot be captured is solved, and more accurate feature extraction and fusion of the data are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of ground-based detection and multi-source big data intelligent analysis, and specifically to a method for intelligent analysis and fusion processing of multi-source big data based on ground-based detection equipment. Background Technique

[0002] With the rapid development of ground-based detection technology, multi-source detection equipment has been widely used in fields such as space target tracking, surveillance and early warning, and range tests. These devices can collect rich data information from different dimensions, such as radar echo signals, angle measurement, range measurement, speed measurement, etc., forming typical multi-modal and multi-source big data. The characteristics of these data are large volume, complex structure, and strong correlation, containing a large amount of potential information. However, to fully exploit the value of these data and apply them to range test and tracking surveillance tasks, an efficient and intelligent fusion processing system is needed. However, from the perspective of existing technologies, there are still many bottlenecks in this field, especially in data preprocessing, feature extraction, multi-modal modeling, and prediction reliability evaluation, and a unified technical framework has not been formed yet.

[0003] The data collected by different detection devices have significant differences in format, time resolution, spatial distribution, and feature dimensions. Existing technologies often use simple alignment and normalization methods when processing these heterogeneous data. These methods usually rely on artificial rules or single algorithms and are difficult to adapt to the complexity of multi-source data. In addition, the data of ground-based detection equipment often contains missing values, especially when the equipment fails or the communication is abnormal, this situation is particularly serious. Traditional linear interpolation or fixed-value filling methods ignore the potential correlation and trend between data, resulting in poor accuracy of the filling results and ultimately affecting downstream analysis. For this reason, those skilled in the art have proposed a method for intelligent analysis and fusion processing of multi-source big data based on ground-based detection equipment to solve the above problems. Summary of the Invention

[0004] Aiming at the deficiencies of the existing technology, the present invention provides a method for intelligent analysis and fusion processing of multi-source big data based on ground-based detection equipment, which solves the problem that the existing technology often uses simple alignment and normalization methods when processing these heterogeneous data.

[0005] To achieve the above objectives, the present invention is realized through the following technical solutions: A method for intelligent analysis and fusion processing of multi-source big data based on ground-based detection equipment, including the following steps:

[0006] Perform standardization processing on the multi-source data collected by the ground-based detection equipment, unify the sampling frequency, fill in the missing values, and complete the time and space alignment;

[0007] Based on the preprocessed multi-source data, construct a similarity weight matrix between data points to measure the correlation between data points;

[0008] The manifold learning algorithm is used to embed multi-source data into a low-dimensional manifold space, and a low-dimensional embedding representation for describing the local geometric structure of the data is constructed;

[0009] Based on the high-order Gaussian process model, an intra-modal feature correlation kernel and an inter-modal interaction kernel are constructed to form a kernel function for jointly describing the features of multi-modal data;

[0010] A fusion framework based on variational Bayesian deep network is constructed. High-dimensional features of multi-source data are extracted through deep learning, the global representation of data fusion is optimized, and the uncertainty of the fusion result is quantified.

[0011] Preferably, the preprocessing includes the following:

[0012] Perform format standardization processing on data of different modalities;

[0013] Perform interpolation resampling on data with inconsistent sampling frequencies to unify the sampling frequency;

[0014] Fill in missing values using statistical methods;

[0015] Use the dynamic time warping algorithm to align the data on the time axis, and perform spatial alignment on the data of different detection devices through spatial interpolation methods.

[0016] Preferably, the calculation of the similarity weight includes the following steps:

[0017] Calculate the Euclidean distance between each pair of data points for the preprocessed multi-source data;

[0018] Calculate the similarity weight through an exponential function based on the distance value of the data points. The similarity weight is used to measure the correlation between data points, and the value range is from zero to one.

[0019] Preferably, the embedding of the manifold is realized through Laplacian eigenmaps, and its steps include:

[0020] Construct a Laplacian matrix based on the similarity weight matrix;

[0021] Obtain the embedding representation of the data in the low-dimensional manifold space by minimizing the distance change in the embedding space through the optimization objective;

[0022] Ensure that the embedding representation satisfies the unit orthogonality constraint condition.

[0023] Preferably, the intra-modal high-order interaction modeling is realized by defining an intra-modal correlation kernel function, which is calculated through the distance between data points and can represent the feature distribution and its local changes within the modality.

[0024] Preferably, the interaction modeling between modalities is achieved by defining an interaction kernel function between modalities, which is calculated from the feature mappings of different modality data and quantifies the feature interactions of multi-modal data based on interaction weights.

[0025] Preferably, the high-order Gaussian process performs feature modeling through a joint kernel function, which includes an intra-modal correlation kernel and an inter-modal interaction kernel, and the joint kernel function is used to generate a joint representation of multi-modal features.

[0026] Preferably, the fusion framework of deep learning is designed based on the multi-head attention mechanism, which captures the non-linear interaction relationships between different modality data through the multi-head attention mechanism and optimizes the deep network parameters through the variational Bayesian method.

[0027] Preferably, the uncertainty quantification of the fusion result is achieved through a variational Bayesian deep network, where the mean and variance of the fusion result are calculated by inferring the posterior distribution of deep learning, and the variance is used to characterize the degree of uncertainty.

[0028] A system for intelligent analysis and fusion processing of multi-source big data based on ground-based detection equipment, comprising:

[0029] A data acquisition module for obtaining multi-source data from multiple ground-based detection equipment;

[0030] A data preprocessing module for performing standardization, time-space alignment, and missing value filling on the acquired multi-source data;

[0031] A similarity weight calculation module for calculating the similarity weights between data points based on the preprocessed data;

[0032] A manifold embedding module for mapping multi-source data to a low-dimensional manifold space based on the calculated similarity weights;

[0033] A modality interaction modeling module for constructing an intra-modal correlation kernel function and an inter-modal interaction kernel function based on a high-order Gaussian process model;

[0034] A data fusion module for achieving high-dimensional feature fusion, uncertainty quantification, and result output of multi-source data through a variational Bayesian deep network.

[0035] The present invention provides a method for intelligent analysis and fusion processing of multi-source big data based on ground-based detection equipment.

[0036] It has the following beneficial effects:

[0037] 1. By introducing the high-order Gaussian process method for multi-modal interaction modeling, the present invention can comprehensively characterize the intra-modal correlation and the complex interaction relationship between modalities. Compared with the methods in the prior art that only model single-modal features or perform simple linear fusion, the problem of being unable to capture the non-linear relationship of multi-modal data is solved, and more accurate feature extraction and fusion of data are achieved.

[0038] 2. The present invention uses manifold embedding technology to map high-dimensional multi-source data to a low-dimensional manifold space, maintaining the local geometric structure and global distribution characteristics of the data. Traditional methods have limited processing capabilities for high-dimensional features and are difficult to effectively retain the feature distribution of complex spaces. This technical solution avoids information loss and improves the robustness and expressive ability of subsequent modeling.

[0039] 3. By adopting the variational Bayesian deep network, the present invention realizes the quantification of uncertainty in multi-source data fusion, significantly improving the reliability and interpretability of prediction results. Compared with the models in the prior art that ignore uncertainty assessment, the present invention can quantify the credibility of prediction results, making up for the defect that a single result cannot judge credibility, and providing more robust decision-making support for practical applications.

[0040] 4. In data preprocessing, the present invention realizes the standardization, time-space alignment, and missing value filling of multi-source data, greatly improving the consistency and integrity of the data. Traditional methods have insufficient processing capabilities for heterogeneous multi-source data, often resulting in feature loss or errors. This solution ensures the accuracy of data processing through multi-step optimization, laying a high-quality foundation for subsequent analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 is a schematic flow chart of the method of the present invention;

[0042] Figure 2 is a schematic diagram of the system architecture of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0043] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0044] Please refer to the attached Figure 1 , the embodiments of the present invention provide a method for intelligent analysis and fusion processing of multi-source big data based on ground-based detection equipment, including:

[0045] S1. Standardize the multi-source data collected by ground-based detection equipment, unify the sampling frequency, fill in missing values, and complete time and space alignment.

[0046] Specifically, step S1 is the starting step of the data fusion analysis process, which mainly preprocesses the multi-source big data collected by ground-based detection equipment. Since multi-source data usually comes from different types of detection equipment, including radar, optical, telemetry detection equipment, etc., there are significant differences in the format, sampling frequency, spatial distribution and time dimension of the data. Direct use may lead to a decrease in the accuracy of subsequent modeling results. Therefore, before starting data modeling, multi-source data needs to be standardized, aligned and repaired to ensure the consistency and availability of input data. This step provides a solid foundation for subsequent data similarity modeling and fusion analysis.

[0047] Data collected by different detection equipment may exist in different forms due to differences in sampling methods, equipment characteristics and storage formats. For example, radar data is usually stored in a two-dimensional matrix form, optical data is angle measurement data based on time series, and telemetry data is mainly in the form of telemetry frames. In order to facilitate subsequent processing, the present invention adopts a unified data structure to convert all input data into a standard matrix form.

[0048] Radar data is flattened into a matrix of row samples and column features, where each row corresponds to a sampling point and each column represents the relevant parameters of radar reflection. Optical data uses time segmentation to divide the sequence into windows of fixed length, and each window is stored as a row in the matrix. For telemetry data, a unified matrix is ​​generated by concatenating the point cloud coordinates with the corresponding parameter values. If the data provided by some devices is already in matrix form, it is read directly without additional processing.

[0049] The sampling frequencies of multi-source data may be different. For example, radar and optical equipment may collect data every 50 milliseconds, while telemetry equipment usually records data at the second level. The difference in sampling frequency will cause the multi-source data to not match in the time dimension. Generally, low-frequency data will be interpolated to the same time scale as high-frequency data.

[0050] The present invention uses linear interpolation to resample low-frequency data. Assume that the time series of low-frequency data is {t 1 ,t 2 ,…,t m}, the corresponding observation value is {x 1 ,x 2 ,…,x m}. For the target time point t′, the interpolation formula is as follows:

[0051]

[0052] Where: x′ is the interpolated observed value, i.e., the estimated value at the target time point t′; t′ is the target time point, satisfying t i ≤t′≤t i+1 ; t i is the i-th time point in the low-frequency time series; t i+1 is the (i + 1)-th time point in the low-frequency time series; x i is the observed value corresponding to the time point t i ; x i+1 corresponds to the observed value at the time point t i+1 . Through interpolation, the low-frequency data is mapped to the same time resolution as the high-frequency data.

[0053] The present invention can also adopt a high-order interpolation method (such as spline interpolation) to improve the interpolation accuracy, which is applicable to data scenarios with obvious non-linear changes.

[0054] During the actual detection process, some devices may cause missing sampling data due to environmental interference, hardware failures, or measurement limitations. The existence of missing values will affect the subsequent modeling results. Therefore, the present invention provides a filling method for different types of missing values.

[0055] For time series data, a linear interpolation method based on adjacent time points is used to complete the missing values. Assume that the value x k corresponding to the time point t k is missing, then the observed values of the two adjacent points t k-1 and t k+1 are used for estimation, and the calculation formula is:

[0056]

[0057] Where: x k is the estimated value at the target time point t k , i.e., the observed value after interpolation and completion; x k-1 is the observed value at the adjacent time point t k-1 ; x k+1 is the observed value at the adjacent time point t k+1 ; t k is the target time point, and the value x k corresponding to it is missing; t k-1 and t k+1 are the two adjacent time points of the target time point t k .

[0058] For spatially distributed data, a weighted interpolation-based method is used for repair. Assume that the value at a certain spatial point is missing, then according to the values of its n adjacent points {x 1 , x 2 , …, x n} Calculate the interpolation value x′, and the calculation formula is as follows:

[0059]

[0060] Where: x′ is the interpolation estimated value of the target space point; x i is the observed value of the i-th neighboring space point; w i is the distance weight of the i-th neighboring space point, and the value is the reciprocal of the distance between this point and the target point; n is the number of neighboring space points used for interpolation; p is the spatial coordinate of the target point; p i is the coordinate of the i-th neighboring space point; ||p - p i || is the Euclidean distance between the target point p and the neighboring point p i .

[0061] If there are too many missing values or systematic missing values (such as complete loss of data within a certain time period), the present invention can adopt a method combining data interpolation and prediction model for repair. For example, use historical data to train a time series prediction model (such as an ARIMA model) for interpolation prediction.

[0062] The distribution of multi-source data in the time and space dimensions may have significant differences. For example, the starting points of different devices on the time axis may be different, or the spatial data coordinates may be inconsistent due to different device positions. Therefore, the present invention provides technical solutions for time alignment and space alignment.

[0063] In time alignment, the present invention adopts the dynamic time warping (DTW) algorithm to minimize the time distance between two sequences by adjusting the alignment path of the time series. The core of the DTW algorithm lies in defining the cumulative distance matrix D( i , j):

[0064] D(i, j) = d(i, j) + min{D(i - 1, j), D(i, j - 1), D(i - 1, j - 1)}

[0065] Where: D(i, j) represents the minimum cumulative distance from the starting point (1, 1) to the point (i, j); d(i, j) is the local distance between two time points i and j in the time series; D(i - 1, j) is the minimum cumulative distance from the point (1, 1) to the point (i - 1, j); D(i, j - 1) is the minimum cumulative distance from the point (1, 1) to the point (i, j - 1); D(i - 1, j - 1) is the minimum cumulative distance from the point (1, 1) to the point (i - 1, j - 1).

[0066] In space alignment, the present invention maps data with different spatial distributions to a unified coordinate system according to the geographic coordinate information of the detection device. The bilinear interpolation method is used to map the spatial values of discrete points to a regular grid, thereby realizing the alignment of spatial data.

[0067] S2. Based on the preprocessed multi-source data, construct a similarity weight matrix between data points to measure the correlation between data points.

[0068] Specifically, step S2 is to further calculate the similarity weight matrix between data points after completing the data preprocessing in step S1, in order to achieve non-linear modeling and fusion analysis between multi-source data. Step S1 ensures the consistency and integrity of the input data by performing processing such as format standardization, sampling frequency unification, and time and space alignment on the multi-source data. On this basis, step S2 quantifies the degree of association between data points by constructing a weight matrix, so as to provide a basic data structure for subsequent manifold embedding and high-order modeling. The core of this step is to construct a set of mathematical models to describe similarity according to the geometric distance and distribution characteristics between data points.

[0069] The distance between different data points is the basic measure of similarity weight. To ensure the applicability of distance calculation, the present invention uses Euclidean distance as the distance measure between data points. Euclidean distance is a common geometric distance that can intuitively describe the relationship between points.

[0070] Suppose the data point set is where is a data vector containing d features. The Euclidean distance between any two points is defined as:

[0071]

[0072] where: d ij is the Euclidean distance between data points X i and X j ; X i , X j are any two points in the data point set, which are the feature vectors of i and j respectively; d is the dimension of the feature space, indicating that each data point contains d features; X ik , X jk respectively represent the values of data points X i and X j on the k-th feature; ||X i -X j || is the vector norm representation of the Euclidean distance between data points X i and X j .

[0073] If the importance of some features is higher than that of other features, weight coefficients w k can be introduced for each feature, so as to construct a weighted Euclidean distance:

[0074]

[0075] where: d ij is the weighted Euclidean distance between data points X i and X j ; w k is the weight of the k-th feature, representing the relative importance of this feature in distance calculation, satisfying w k > 0 and d is the dimension of the feature space; X ik , X jk respectively represent the values of data points X i and X j on the k-th feature. By adjusting the feature weights, the contribution of a specific dimension can be highlighted, thereby enhancing the applicability of similarity calculation.

[0076] To map the geometric distance between data points into a non-linear similarity metric, the present invention constructs a similarity weight matrix using a weight function based on a Gaussian kernel. The Gaussian kernel is a widely used kernel function that can effectively capture the local relationships between data points.

[0077] The similarity weight W ij is calculated as follows:

[0078]

[0079] where: W ij represents the similarity weight between point X i and point X j ; d ij is the Euclidean distance between the two points; σ > 0 is the kernel width parameter, used to control the distribution range of similarity.

[0080] A smaller σ value will make the weights more concentrated in the local area, while a larger σ value will enhance the similarity of distant points. Therefore, the selection of σ needs to be adjusted according to the distribution characteristics of the actual data. In one possible implementation, the kernel width parameter σ can be taken as the mean of the distances between all pairs of points, and the formula is:

[0081]

[0082] where: σ is the kernel width parameter, used to control the sensitivity of distance weighting or the kernel function to distant and near data points; n is the total number of data points; d ij is the distance between point X i and point X j .

[0083] To reduce storage and computational complexity, the present invention sparsifies the similarity weight matrix. The main idea of sparsification is to retain the weights between each data point and its nearest neighbor data points, and set other weights to zero.

[0084] After calculating the weight matrix W, for any data point X i , only retain the weights between its k nearest neighbor points, defined as:

[0085]

[0086] where: W ij represents the similarity weight between point X i and point X j ; d ij is the Euclidean distance between point X i and point X j ; σ is the kernel width parameter used to adjust the sensitivity of the weight to the distance; N k (X i ) is the set of k nearest neighbor points of point X i ; X j is any point in the data set. The sparsified weight matrix can still retain the main similarity relationships between data points while significantly reducing the storage requirements.

[0087] The present invention can also introduce an adaptive sparsification strategy according to the global distribution characteristics of the data, that is, dynamically adjust the number k of nearest neighbor points for different data points. For example, for data points in a dense area, the value of k can be reduced; while for data points in a sparse area, the value of k can be appropriately increased to improve the flexibility of similarity modeling.

[0088] The relationship between data points may not be accurately described by simple Euclidean distance or Gaussian kernel function. The present invention provides corresponding extension methods for these scenarios.

[0089] For data points with a time dimension, the similarity between time series can be calculated by combining dynamic time warping (DTW). Suppose two time series are X i ={x i1 ,x i2 ,…,x im} and X j ={x j1 ,x j2 ,…,x jn}, the DTW algorithm constructs a cumulative distance matrix D(i,j) to minimize the alignment path length between sequences, and the recurrence formula is:

[0090] D(i,j) = d(i,j) + min{D(i - 1,j), D(i,j - 1), D(i - 1,j - 1)}

[0091] Where: D(i,j) represents the minimum cumulative distance from the starting point (1,1) to the point (i,j); d(i,j) is the local distance between two time points i and j in the time series; D(i - 1,j) is the minimum cumulative distance from the point (1,1) to the point (i - 1,j); D(i,j - 1) is the minimum cumulative distance from the point (1,1) to the point (i,j - 1); D(i - 1,j - 1) is the minimum cumulative distance from the point (1,1) to the point (i - 1,j - 1).

[0092] S3. Use the manifold learning algorithm to embed multi-source data into a low-dimensional manifold space, and construct a low-dimensional embedding representation for describing the local geometric structure of the data.

[0093] Specifically, step S3 is a key step carried out after the similarity weight calculation in step S2 to further reveal the intrinsic geometric characteristics of multi-source data. The similarity weight matrix accurately describes the local correlation between data points, and the high-dimensional distribution characteristics of the data make it difficult to directly analyze. Therefore, in the present invention, through the manifold learning technology, the high-dimensional data is projected into a low-dimensional space, so as to realize the retention of the local geometric structure of the data and the reduction of the computational complexity. The manifold embedding not only provides a simplified feature space for subsequent modeling, but also enhances the representation ability of the data.

[0094] Specifically, the following technical means are adopted to realize the manifold embedding of the data: The manifold embedding needs to construct a matrix based on the similarity between data points that can reflect its local structure. The present invention uses the Laplacian matrix as the structured representation of the data, and its construction method is as follows:

[0095] First, according to the similarity weight matrix W calculated in step S2, define the diagonal matrix D:

[0096]

[0097] Where: n is the total number of data points; W ij is the similarity weight between X i and the point X j ; D ii represents the sum of weights related to the point X i , which is used to measure the global connectivity of the point X i . Then, construct the Laplacian matrix L through D and W:

[0098] L = D - W

[0099] Where: L is a symmetric positive definite matrix, which can effectively reflect the local adjacent relationship between data points.

[0100] To further optimize the alignment and processing performance of data, the present invention may adopt a normalized Laplacian matrix L norm , which is defined as follows:

[0101] L norm = I - D -1 / 2 WD -1 / 2

[0102] where: i is the identity matrix, with the same dimension as D and W; D -1 / 2 is the inverse square root matrix of the diagonal matrix D. The normalization process can reduce the influence of uneven data distribution on the embedding effect.

[0103] After completing the construction of the Laplacian matrix, the present invention projects the high-dimensional data into a low-dimensional space through Laplacian Eigenmaps to preserve its local geometric properties.

[0104] The objective of the present invention is to minimize the change in the local distance between the embedding points of the data in the low-dimensional space by optimizing the following objective function:

[0105]

[0106] where: represents the embedding representation of the data in the low-dimensional space, d' is the dimension of the low-dimensional space; W ij is the similarity weight between X i and the point X j ; ||Y i - Y j || 2 is the square of the Euclidean distance between X i and the point X j in the low-dimensional space.

[0107] To ensure the solvability of the optimization problem, the present invention further introduces an orthogonality constraint:

[0108]

[0109] where: I is the identity matrix, representing the orthogonality constraint; D is the diagonal matrix, representing the global connectivity of the points. This constraint condition ensures that different dimensions of the embedding representation are orthogonal to each other, thus avoiding information redundancy.

[0110] By solving the following generalized eigenvalue problem, the low-dimensional embedding representation can be efficiently obtained:

[0111] LY = λDY

[0112] Among them: L is the Laplacian matrix, representing the local adjacent relationship between data points; λ is the eigenvalue, representing the degree of deformation in a specific direction; Y is the eigenvector, representing the embedded representation in the low-dimensional space.

[0113] The embedding dimension d′ is an important parameter in manifold learning, and its selection directly affects the expressiveness and simplification degree of the embedding result. The selection of the embedding dimension in the present invention is based on the distribution of eigenvalues. By observing the decreasing rate of the generalized eigenvalue λ, the number of eigenvalues before a point with a large rate of change is selected as the embedding dimension.

[0114] The embedding dimension can be automatically determined by the following criterion:

[0115]

[0116] Among them: d′ is the dimension of the embedded low-dimensional space; λ k is the k-th eigenvalue; ∈>0 is a user-defined threshold, controlling the determination criterion for the change amount between eigenvalues. This method can adapt to the characteristics of different data distributions, thus ensuring the reliability of the embedding result.

[0117] In order to improve the robustness of the embedded representation, the present invention further post-processes the result after manifold embedding. For example, for high-noise data, a low-pass filter can be used to smooth the embedded representation Y to eliminate the influence of local discontinuous points.

[0118] In addition, for data containing time information, the present invention can correct the embedded representation by combining time smoothing techniques. For example, applying a moving average to the time series of each dimension in the embedded space, the formula is as follows:

[0119]

[0120] Among them: Y′ i is the smoothed embedded representation, the corrected value corresponding to the time point i; Y i+j is the original value of the embedded representation at the time point i + j; r is the radius of the moving window, representing the range of time points considered in the smoothing operation; 2r + 1 is the length of the moving window, representing the total number of time points included in the smoothing operation.

[0121] In some specific scenarios, such as when the data has complex non-linear relationships, the present invention can also use other manifold learning algorithms to replace Laplacian eigenmaps. For example, using Locally Linear Embedding (LLE) to preserve the linear local characteristics of the data, or adopting t-SNE (t-distributed Stochastic Neighbor Embedding) technology to better handle the clustering problem of high-dimensional dense data.

[0122] For the LLE algorithm, the optimization objective is as follows:

[0123]

[0124] where: Y i is the embedded representation of the data point X i in the low-dimensional space; Y j is the embedded representation of the neighborhood point X j in the low-dimensional space; N i is the set of neighborhood points of the data point X i ; W ij represents the similarity weight between the point X i and the point X j ; n is the total number of data points. Through this optimization objective, the linear retention ability of the manifold embedding can be further enhanced.

[0125] S4. Based on the high-order Gaussian process model, construct the feature correlation kernel within the modality and the interaction kernel between modalities to form a kernel function for jointly describing the features of multimodal data.

[0126] Specifically, after the data manifold embedding in step S3 is completed, the multi-source data has been projected into a low-dimensional space. At this time, although the geometric structure of the data is well preserved, the feature expression ability of each modality data and the interaction relationship between modalities have not been fully modeled. In order to further characterize the distribution law of the features within the modality and the correlation between modalities, the present invention introduces a high-order Gaussian process model in step S4, and constructs a joint feature expression through feature modeling within the modality and interaction modeling between modalities, laying a solid foundation for subsequent data fusion and uncertainty quantification.

[0127] Specifically, the following technical means are adopted to achieve the high-order interaction modeling within and between modalities: There are still local feature changes in the embedded manifold space for data points of different modalities. In order to characterize this change trend, this embodiment uses a Gaussian kernel function to model the features within the modality.

[0128] Assume that the data set contains N modalities, and the data of each modality is represented as where n i represents the number of data points in the i-th modality, and d′ represents the dimension of the embedded space. For any two data points X i,j and X i,k in the modality i, define its feature kernel function within the modality as:

[0129]

[0130] where: k s (·,·) is the kernel function within the modality, which is used to measure the similarity between data points within the modality; ||Xi,j -X i,k || is the Euclidean distance between the data point X i,j and X i,k ; l is the kernel width parameter, which is used to control the sensitivity of the kernel function and determines the decay range of similarity.

[0131] The kernel width l can be automatically set according to the average distance between the points within the modality to adapt to the characteristics of different modality data. For example:

[0132]

[0133] where: n i is the total number of data points in modality i; ||X i,j -X i,k || is the Euclidean distance between the data point X i,j and X i,k .

[0134] Through the above Gaussian kernel function, the characteristic distribution and its local correlation of the data within the modality can be effectively characterized.

[0135] To describe the association relationship between different modality data, in this embodiment, an interaction kernel function is used to model the features between modalities.

[0136] Assume that the feature representations of two modalities i and j are X i and X j , and the cross-modal interaction kernel function between a point X i,p in modality i and a point X j,q in modality j is defined as:

[0137]

[0138] where: X i and X j represent the feature representations of modalities i and j respectively; X i,p is the p-th feature point of modality i; X j,q is the q-th feature point of modality i; represents the cross-modal interaction kernel function between modalities i and j; φ i (X i,p ) and φ j (X j,q ) are the feature mapping functions of modalities i and j respectively.

[0139] The feature mapping function φ i (·) can be designed as a non-linear mapping according to the characteristics of the modality data, such as polynomial feature expansion or neural network embedding representation. As a possible implementation, the feature mapping function can be defined as:

[0140]

[0141] where: φ i (·) is a feature mapping function used to non - linearly map the original feature X i,p to a high - dimensional space to better capture the complex relationships between features; X i,p represents the feature of the p - th data point of modality i; c i is the center point of the feature mapping and is the Gaussian kernel center corresponding to modality i; σ i is the kernel width parameter in the feature mapping, which controls the range and sensitivity of the Gaussian kernel function; ||X i,p -c i || 2 is the square of the Euclidean distance between the feature point X i,p and the kernel center c i , measuring the similarity between the feature point and the kernel center.

[0142] Through the above - mentioned inter - modality interaction kernel function, the feature correlation relationships between different modalities can be effectively captured.

[0143] To uniformly describe the intra - modality features and the inter - modality interaction relationships, in this embodiment, a joint kernel function is used to perform high - order feature modeling on multi - modality data. The joint kernel function is defined as:

[0144]

[0145] where: k(X,X′) is the joint kernel function, used to describe the overall similarity between samples X and X′; k s (X i ,X′ i ) is the intra - modality kernel function of modality i, used to describe the similarity between data points within modality i; is the inter - modality interaction kernel function between modality i and modality j, used to characterize the data association between modalities; w ij represents the weight of the interaction between modality i and modality j, used to adjust the contribution ratio of the interaction between different modalities.

[0146] The weight parameter w ij can be automatically adjusted through the optimization objective to ensure the balanced feature contribution of multi - modality data.

[0147] After the joint kernel function is constructed, the present invention further expresses the joint features of multi - modality data based on a high - order Gaussian process model.

[0148] The high - order Gaussian process model assumes that the joint feature F of multi - modality data satisfies the following distribution:

[0149] F~GP(m(X),k(X,X′))

[0150] Where: F represents the joint features of multi-modal data; GP represents the Gaussian process distribution; m(X) is the mean function, usually set to the zero function; k(X, X′) is the joint kernel function, which describes the similarity and correlation between feature points X and X′.

[0151] Through the high-order Gaussian process model, it is possible to effectively capture the complex interaction relationships between modalities while maintaining the consistency of intra-modal features, providing a global feature representation for subsequent in-depth fusion analysis.

[0152] To further improve the flexibility of inter-modal interaction modeling, this embodiment can also introduce a sparse Gaussian process to reduce the computational complexity of the kernel matrix by selecting subset points. Specifically, the sparse Gaussian process constructs a low-rank approximate kernel matrix by selecting M representative points:

[0153]

[0154] Where: Z represents the set of M representative points selected from the data; K X,Z is the kernel matrix between representative points; K Z,X′ is the kernel matrix between representative points and new data points X′.

[0155] Through the sparse Gaussian process, it is possible to maintain the modeling accuracy while reducing the computational cost.

[0156] S5. Construct a fusion framework based on the variational Bayesian deep network, extract high-dimensional features of multi-source data through deep learning, optimize the global representation of data fusion, and quantify the uncertainty of the fusion result.

[0157] Specifically, after completing the intra-modal and inter-modal high-order interaction modeling in step S4, the joint feature expression of multi-modal data has been established. However, the feature representation of the high-order Gaussian process is still local and static, and does not fully utilize the global optimization ability of the deep learning framework. In addition, during the data fusion process, the uncertainty of the result is not quantified, which may affect the reliability of practical applications. Therefore, the present invention proposes a global fusion and optimization method based on the variational Bayesian deep network in step S5, aiming to further optimize the global expression of multi-modal joint features through deep learning while quantifying the uncertainty of the result.

[0158] The following technical implementation methods are specifically adopted to complete global fusion and optimization: To effectively extract the global features within and between modalities, this embodiment introduces a deep learning network structure based on the multi-head attention mechanism. The multi-head attention mechanism improves the expression ability of data fusion by capturing the non-linear interaction relationships between different modalities.

[0159] Assume that the joint feature representation of multi-modal data is X = {X 1, X 2 , …, X N}, where \(x\) i represents the feature of the \(i\)-th modality. For any modalities \(i\) and \(j\), the calculation process of its attention mechanism includes the following steps:

[0160] First, define the query matrix \(Q\), key matrix \(K\), and value matrix \(V\), which are generated from the input feature \(X\) through weight matrices \(W\) q , \(W\) k , \(W\) v :

[0161] \(Q = W\) q \(X\), \(K = W\) k \(X\), \(V = W\) v \(X\)

[0162] where: are the linear transformation weights of the query matrix, key matrix, and value matrix respectively; represent the query matrix, key matrix, and value matrix obtained through linear transformation respectively; \(d\) represents the dimension of the input feature; \(d\) k represents the dimension of the attention head.

[0163] Then, calculate the dot product of the query matrix and the key matrix, and normalize to obtain the attention weights:

[0164]

[0165] where: Calculate the dot product between the query matrix and the key matrix to obtain the attention scores, representing the correlation between features; is the scaling factor, used to avoid the problem of gradient disappearance caused by too large dot product values, especially in high-dimensional spaces; softmax(·) normalizes the attention scores to ensure that they are in the range of [0, 1] and satisfy the properties of probability distribution; the \(V\) value matrix is weighted and summed through the attention weights to generate the output related to the input feature.

[0166] Finally, concatenate the outputs of the multi-head attention and generate the final global fusion feature through linear transformation:

[0167] \(H = Concat(head\) 1 , \(head\) 2 , …, \(head\) h ) \(W\) o

[0168] where: \(Concat(·)\) represents concatenating the outputs of multiple attention heads together; The weights of the linear transformation in the fusion layer; H is the final global fusion feature, which contains the relationships between different modalities. Through the multi-head attention mechanism, the global feature relationships between different modalities can be fully extracted.

[0169] To further quantify the uncertainty of the deep learning fusion results, in this embodiment, a variational Bayesian optimization framework is introduced for the network parameters. Through variational inference, the posterior distribution of the network parameters is optimized, enabling the network to not only generate fusion prediction values but also provide an estimate of the credibility of the prediction results.

[0170] Assume that the parameters of the deep learning network are

[0171]

[0172] Where: θ represents the parameters of the deep learning model; p(θ) is the prior distribution, representing the assumption about the model parameters θ before observing the training data D; q(θ) is the variational distribution, used to approximate the posterior distribution p(θ∣D); β is a regulating factor, used to balance the data reconstruction error (data fitting) and the regularization term (constraint of prior information); KL(q(θ)||P(θ)) is the Kullback-Leibler (KL) divergence, measuring the difference between the variational distribution q(θ) and the prior distribution p(θ); is the data reconstruction error, measuring the difference between the model prediction value and the true value Y i gap.

[0173] Specifically, the data reconstruction error is defined as:

[0174]

[0175] Where: n is the number of samples in the training data; Y i is the true label of the i-th sample; is the predicted output of the model.

[0176] The prior distribution p(θ) can be set as a Gaussian distribution while the variational distribution q(θ) is approximated in a parametric form as where μ and σ are learnable parameters.

[0177] After completing the network optimization, the present invention further quantifies the uncertainty of the results through the posterior distribution of the parameters. The uncertainty mainly includes the following two parts:

[0178] One part is the predicted mean of the model, used to represent the best prediction value of the model:

[0179]

[0180] Where: denotes the predicted mean of the model, i.e., the expectation of the output value of the model under the input data x; f(x,θ) represents the output of the neural network, i.e., the prediction result of the model under the input x and parameters θ; x is the input data point, which is the input feature of the model; θ represents the parameters (weights, biases, etc.) of the neural network; q(θ) represents the variational distribution, which is an approximation of the posterior distribution p(θ∣D) of the model parameters; E q(θ) [·] denotes taking the expectation with respect to the distribution q(θ), i.e., calculating the mean of the samples from the parameter distribution.

[0181] Another part is the uncertainty variance, which is used to quantify the credibility of the model prediction:

[0182]

[0183] Where: represents the uncertainty of the model prediction result, quantifying the credibility of the model prediction on the input x; the prediction result f(x,θ) and the predicted mean of the squared error, representing the degree of deviation of a single prediction from the mean; E q(θ) denotes taking the expectation with respect to the variational distribution q(θ).

[0184] In this embodiment, the accuracy of uncertainty estimation can be further improved by sampling the parameters θ~q(θ) multiple times and calculating the mean and variance of the output results.

[0185] In some specific scenarios, such as applications with high real-time computing requirements, this embodiment can reduce the computational complexity through the sparse Bayesian method. Specifically, by selecting important parameters for Bayesian inference and using fixed values for other parameters, the computational amount of the posterior distribution is reduced.

[0186] In addition, for high-dimensional data, the present invention can preprocess the input features by combining dimensionality reduction techniques (such as principal component analysis) to further optimize the computational efficiency of the network.

[0187] A system for intelligent analysis and fusion processing of multi-source big data based on ground-based detection equipment described below can be correspondingly referred to the method for intelligent analysis and fusion processing of multi-source big data based on ground-based detection equipment described above.

[0188] Please refer to the attached Figure 2 , this embodiment of the present invention also provides a system for intelligent analysis and fusion processing of multi-source big data based on ground-based detection equipment, including:

[0189] A data acquisition module, configured to obtain multi-source data from multiple ground-based detection equipment;

[0190] A data preprocessing module for standardizing, time-space aligning, and filling missing values in the collected multi-source data;

[0191] A similarity weight calculation module for calculating the similarity weights between data points based on the preprocessed data;

[0192] A manifold embedding module for mapping multi-source data to a low-dimensional manifold space based on the calculated similarity weights;

[0193] A modal interaction modeling module for constructing an intra-modal correlation kernel function and an inter-modal interaction kernel function based on a high-order Gaussian process model;

[0194] A data fusion module for achieving high-dimensional feature fusion, uncertainty quantification, and result output of multi-source data through a variational Bayesian deep network.

[0195] Specifically, for the data acquisition module, the function description is as follows: The core function of the data acquisition module is to obtain multi-source data from multiple ground-based detection devices. This data can come from different types of detection devices, such as radar, optical, telemetry sensors, etc., covering multi-modal, multi-time and space dimension information. Features: It supports data access from heterogeneous devices and can collect various data types including structural radar data, optical data, telemetry parameters, etc. It adopts real-time or offline acquisition methods to ensure the integrity and timeliness of the data. It provides meta-information records of device data, including acquisition time, geographical location, and device type, etc.

[0196] A data preprocessing module, the function description is as follows: This module performs standardization, time-space alignment, and filling of missing values on the collected multi-source data to ensure the accuracy and consistency of subsequent analysis.

[0197] Preprocessing steps: Standardization processing: Normalize or standardize data from different sources to eliminate the influence of dimension differences on subsequent analysis.

[0198] Time-space alignment: Time alignment: Synchronize data from different devices according to the sampling timestamps to ensure consistency in the time dimension. Space alignment: Calibrate the geographical location data of detection devices to solve the problem of spatial inconsistency caused by different device positions. Filling of missing values: Use interpolation methods (such as linear interpolation, spline interpolation) or similarity-based methods to fill in missing data to ensure the integrity of the data.

[0199] A similarity weight calculation module, the function description is as follows: This module calculates the similarity weights between data points based on the preprocessed multi-source data to characterize the strength of the relationship between data.

[0200] Computational Features: Multidimensional Similarity: Combining the dimensions of time, space, and eigenvalue to comprehensively measure the similarity between data points. Weight Matrix Construction: Constructing a weighted adjacency matrix based on the similarity results to provide a basis for subsequent manifold embedding. Adaptive Parameter Adjustment: Dynamically adjusting the kernel width or distance metric method according to the data distribution characteristics to improve the computational accuracy.

[0201] Manifold Embedding Module, Functional Description: This module maps high-dimensional multi-source data to a low-dimensional manifold space based on similarity weights for subsequent feature analysis and processing.

[0202] Main Technologies: Manifold Learning Methods: Using methods such as Laplacian Eigenmaps or Locally Linear Embedding (LLE) to reveal the low-dimensional manifold structure of the data. Low-Dimensional Representation: Projecting the original high-dimensional data onto a low-dimensional space that can preserve local geometric properties for subsequent modeling and analysis. Noise Robustness: Considering the impact of noisy data on the manifold structure and improving the stability of low-dimensional embedding through weight sparsification or regularization methods.

[0203] Modal Interaction Modeling Module, Functional Description: This module constructs an intra-modal correlation kernel function and an inter-modal interaction kernel function respectively through a high-order Gaussian process model to capture the feature distribution within the modality and the complex associations between modalities.

[0204] Main Tasks: Intra-Modal Modeling: Using a Gaussian kernel function or a polynomial kernel function to model the features of a single modality and extract the intra-modal correlation features. Inter-Modal Interaction: Quantifying the feature interaction relationship between different modalities based on the interaction kernel function to reveal the dependency structure between modalities. Parameter Optimization: Using variational inference methods to optimize the parameters of the kernel function to improve the model's ability to describe the multi-modal feature relationship.

[0205] Data Fusion Module, Functional Description: The data fusion module integrates the high-dimensional features of multi-source data through a variational Bayesian deep network to achieve uncertainty quantification and result output. Main Functions: High-Dimensional Feature Fusion: Using a deep neural network to map the features of different modalities to a unified high-dimensional space to generate a joint feature representation.

[0206] Uncertainty Quantification: Based on the Bayesian framework, quantifying the uncertainty of the fusion result and providing a credibility assessment of the prediction result. The uncertainty information can assist in decision-making. For example, high-uncertainty regions can be further verified manually.

[0207] Result Output: Generating the final analysis results, including low-dimensional feature representations, prediction results, and their uncertainty scores. Supporting visual output, such as feature distribution maps, prediction confidence intervals, etc.

[0208] Although embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for intelligent analysis and fusion processing of multi-source big data based on ground-based detection equipment, characterized in that: The following steps are involved: Standardize the multi-source data collected by ground-based detection equipment, unify the sampling frequency, fill in missing values, and complete time and space alignment; Based on the preprocessed multi-source data, a similarity weight matrix between data points is constructed to measure the correlation between data points; The multi-source data is embedded into a low-dimensional manifold space using a manifold learning algorithm to construct a low-dimensional embedding representation for describing the local geometric structure of the data. Based on the high-order Gaussian process model, the feature correlation kernel within the modality and the interaction kernel between the modalities are constructed to form a kernel function for jointly describing the characteristics of multimodal data. A fusion framework based on variational Bayesian deep network is constructed to extract high-dimensional features of multi-source data through deep learning, optimize the global representation of data fusion, and quantify the uncertainty of the fusion results.

2. According to claim 1, a method for intelligent analysis and fusion processing of multi-source big data based on ground-based detection equipment is characterized in that: The preprocessing includes the following contents: Standardize the formats of data of different modalities; Perform interpolation and resampling processing on data with inconsistent sampling frequencies to unify the sampling frequency; Statistical methods were used to fill missing values; The dynamic time warping algorithm is used to align the data on the time axis, and the spatial interpolation method is used to spatially align the data from different detection devices.

3. According to claim 1, a method for intelligent analysis and fusion processing of multi-source big data based on ground-based detection equipment is characterized in that: The calculation of the similarity weight comprises the following steps: Calculate the Euclidean distance between each pair of data points for the preprocessed multi-source data; The similarity weight is calculated by an exponential function according to the distance value of the data points. The similarity weight is used to measure the correlation between the data points, and the value range is zero to one.

4. The method for intelligent analysis and fusion processing of multi-source big data based on ground-based detection equipment according to claim 1 is characterized in that: The embedding of the manifold is achieved by Laplace eigenmapping, and the steps include: Construct a Laplacian matrix based on the similarity weight matrix; By optimizing the objective to minimize the distance change in the embedding space, the embedded representation of the data in the low-dimensional manifold space is obtained; Make sure the embedded representation satisfies the unit orthogonality constraint.

5. The method for intelligent analysis and fusion processing of multi-source big data based on ground-based detection equipment according to claim 1 is characterized in that: The high-order interaction modeling within the modality is achieved by defining an intra-modal correlation kernel function, which is calculated by the distance between data points and can represent the feature distribution and its local changes within the modality.

6. The method for intelligent analysis and fusion processing of multi-source big data based on ground-based detection equipment according to claim 1 is characterized in that: The inter-modal interaction modeling is achieved by defining an inter-modal interaction kernel function, which is calculated from feature maps of different modal data and quantifies the feature interactions of multi-modal data based on interaction weights.

7. The method for intelligent analysis and fusion processing of multi-source big data based on ground-based detection equipment according to claim 1 is characterized in that: The high-order Gaussian process performs feature modeling through a joint kernel function, wherein the joint kernel function includes an intra-modal correlation kernel and an inter-modal interaction kernel, and the joint kernel function is used to generate a joint representation of multi-modal features.

8. The method for intelligent analysis and fusion processing of multi-source big data based on ground-based detection equipment according to claim 1 is characterized in that: The deep learning fusion framework is designed based on a multi-head attention mechanism, which captures the nonlinear interaction between different modal data and optimizes the deep network parameters through a variational Bayesian method.

9. The method for intelligent analysis and fusion processing of multi-source big data based on ground-based detection equipment according to claim 1 is characterized in that: The uncertainty quantification of the fusion result is achieved through a variational Bayesian deep network, in which the mean and variance of the fusion result are calculated by inferring the posterior distribution of deep learning, and the variance is used to characterize the degree of uncertainty.

10. A system for intelligent analysis and fusion processing of multi-source big data based on ground-based detection equipment, applied to a method for intelligent analysis and fusion processing of multi-source big data based on ground-based detection equipment as described in any one of claims 1 to 9, characterized in that: include: A data acquisition module, used to obtain multi-source data from multiple ground-based detection devices; The data preprocessing module is used to standardize, align time and space, and fill missing values ​​on the collected multi-source data; A similarity weight calculation module is used to calculate the similarity weights between data points based on the preprocessed data; A manifold embedding module for mapping multi-source data into a low-dimensional manifold space based on the calculated similarity weights; Modal interaction modeling module, which is used to construct intra-modal correlation kernel function and inter-modal interaction kernel function based on high-order Gaussian process model; The data fusion module is used to realize high-dimensional feature fusion, uncertainty quantification and result output of multi-source data through variational Bayesian deep network.

Citation Information

Cited By

  • Energy terminal multi-source data fusion analysis and self-adaptive regulation and control method and system

    CN120408532A

  • Wind power plant data reconstruction method based on graph attention mechanism and manifold auto-encoder

    CN120492826A

  • Multi-source heterogeneous data fusion method and system based on edge calculation

    CN120541795A

  • River monitoring data processing method and device and electronic equipment

    CN120653902A

  • Multi-source heterogeneous data fusion processing and key feature extraction method and system

    CN121881247A