Data aggregation method based on multi-modal features
Through the ball tree algorithm and local linear embedding technology, combined with the value frequency and probability density to calculate marginal probability, the problems of information redundancy and low computing efficiency in multimodal data fusion are solved, and efficient and accurate data aggregation is achieved.
Patent Information
- Application Number
- CN202510576972.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-05-06
AI Technical Summary
In the multimodal data fusion, the existing technology has problems such as redundancy in information, low computing efficiency and decision-making deviations, and the weight allocation is disconnected from the intrinsic distribution characteristics of the data.
The spherical tree algorithm is used to locate the neighborhood points of high-dimensional data, the medium-dimensional data dynamically adjust the neighborhood range, the low-dimensional data calculates the neighborhood through Euclidean distance, and maps it to the low-dimensional space through local linear embedding, and calculates the marginal probability based on the value frequency and probability density to determine the data aggregation weight.
The utilization efficiency and accuracy of multimodal features are improved, the data processing efficiency and feature utilization effect in different dimensions are balanced, and the dynamic weight allocation strengthens the complementarity between multimodal features, reduces redundant information, and provides efficient cross-modal information fusion.
Smart Images

Figure CN120372555A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of resource management, and particularly to a data aggregation method based on multi-modal features. Background Art
[0002] With the rapid development of the Internet of Things and intelligent sensing technologies, multi-modal data fusion has become a core technology in fields such as smart cities, industrial inspection, and medical diagnosis. Traditional data aggregation methods mostly adopt early fusion or late fusion architectures: the former integrates multi-source data through simple feature splicing, but ignores the dimensional differences and semantic gaps between different modalities, resulting in information redundancy and low computational efficiency; the latter relies on independent models to process data of each modality, although it alleviates dimensional conflicts, it is difficult to capture cross-modal correlation features and is prone to decision-making biases in dynamic scenarios; although existing improvement schemes partially solve the above problems, their computational complexity increases exponentially, and the weight allocation process is disconnected from the inherent distribution characteristics of the data. Summary of the Invention
[0003] The purpose of the present invention is to provide a data aggregation method based on multi-modal features.
[0004] To achieve the above purpose, the present invention is implemented according to the following technical solutions: The first aspect of the present invention provides a data aggregation method based on multi-modal features, including: Collect multi-modal data, and preprocess the multi-modal data to obtain dimensional data, where the dimensional data includes high-dimensional, medium-dimensional, and low-dimensional data; Locate the nearest neighbor points of the data points for the high-dimensional data through the ball tree algorithm, dynamically adjust the neighborhood range according to the distribution density of the medium-dimensional data, calculate the neighborhood for the low-dimensional data through the Euclidean distance, and map the dimensional data to a low-dimensional space through local linear embedding based on the neighborhood; Traverse the dimensional data in the low-dimensional space, count the value frequencies of discrete data, obtain the probability density of continuous data, and combine the value frequencies and probability densities to obtain marginal probabilities; Use the marginal probabilities to determine the weights of the dimensional data in data aggregation, and splice the dimensional data features after dimensionality reduction according to the weights to obtain the aggregated comprehensive features.
[0005] As a further method, the method of collecting multi-modal data and preprocessing the multi-modal data to obtain dimensional data includes: Collect multi-modal data, clean and denoise the data of different modalities, calculate the means of the data of different modalities to fill in missing values, and correct outliers using the interquartile range method; Divide according to the modality of multimodal data to obtain text data, numerical data, audio data, image data, and video data. Obtain the word vector dimension for text data through a word vector model, obtain the numerical dimension according to the numerical columns of numerical data, extract the spectral features and time-domain features of audio data and merge them into audio dimension information, determine the basic dimension for image data through the resolution and the number of color channels, and obtain the video spatio-temporal dimension information based on the image data and the frame rate duration of key frames of video data. Obtain the number of dimensions of each modality dimension information, count all modality dimension numbers, sort them from high to low, divide the top 20% into the high-dimensional interval, the middle 60% into the medium-dimensional interval, and the bottom 20% into the low-dimensional interval. Put the modality data of different dimensions into the dimension interval to obtain the dimension dataset; Obtain the maximum-minimum difference, mean, and standard deviation of different modality data in the dimension dataset, use the maximum-minimum difference, mean, and standard deviation to determine the scaling threshold, and scale the dimension data based on the scaling threshold.
[0006] As a further method, the method of using the maximum-minimum difference, mean, and standard deviation to determine the scaling threshold and scaling the corresponding modality data based on the scaling threshold includes: Obtain the maximum-minimum difference, mean, and standard deviation of different modality data in the dimension dataset, calculate the information entropy of the maximum-minimum difference, mean, and standard deviation for data feature description, obtain the dynamic threshold weight coefficients of the maximum-minimum difference, mean, and standard deviation after normalizing the information entropy, calculate the data difference degree based on the historical threshold, the mean of the current data window, and the maximum-minimum difference, and obtain the scaling threshold through the difference degree. The formula for the scaling threshold is: , where is the current newly input data point, n is the total number of current data points, is the cumulative mean of the first n - 1 data points, is the cumulative variance of the first n - 1 data points, M is the maximum value within the data window, m is the minimum value within the data window, last is the historical threshold, , , are the weight coefficients of the maximum-minimum difference, mean, and standard deviation for the threshold, is the attenuation coefficient; Scale the corresponding modality data based on the scaling threshold.
[0007] As a further method, the method of locating the nearest neighbor points of data points for high-dimensional data through the ball tree algorithm includes: Based on a high-dimensional dataset, regarding high-dimensional data points as spatial points, recursively partitioning the data space, calculating the mean of the data subset of the current node to determine the center point, taking the center point as the center of the hypersphere, and taking the maximum distance from each point to the center point as the radius to generate a ball tree node containing the center, radius, and child nodes. Repeat the partitioning until the leaf node data volume threshold or the recursive depth limit is met to complete the construction of the ball tree; Perform a nearest neighbor search on the target high-dimensional data point based on the ball tree structure. Start traversing from the root node of the ball tree, calculate the distance between the target point and the center of the hypersphere of the current node. If the distance exceeds the radius of the hypersphere, prune this branch; if not, continue to search the child nodes. After reaching the leaf node, calculate the distances between the target point and all data points within the leaf node, and record the initial nearest neighbor points; After obtaining the initial nearest neighbor points, optimize the nearest neighbor set by backtracking the tree structure. Backtrack along the tree from the leaf node. For each node passed through, calculate the minimum distance between the target point and the surface of the hypersphere of the node. If this distance is less than the distance of the current farthest nearest neighbor point, expand the node, recalculate the distances between the data points within the node and the target point, and update the nearest neighbor point set until the entire tree search is completed, finally determining the nearest neighbor points of the high-dimensional data points and obtaining the high-dimensional neighborhood set.
[0008] As a further method, the method of dynamically adjusting the neighborhood range according to the distribution density of the middle-dimensional data and calculating the neighborhood for low-dimensional data by Euclidean distance includes: Based on each data point in the middle-dimensional dataset, calculate its contribution to the density of the target point through the Gaussian kernel function. Obtain the bandwidth parameter using Silverman's rule for the standard deviation of the middle-dimensional dataset. Divide the sum of the contributions of all data points by the total number of data points and the bandwidth parameter to obtain the density estimate value of each data point. Sort the density estimate values in ascending order, take the 75th percentile value as the high density threshold, and the 25th percentile value as the low density threshold; Calculate the proportion of data points higher than the high density threshold , if , the reduction factor takes , if , take the default value 0.8; calculate the proportion of data points lower than the low density threshold , if , the expansion factor takes , if , take the default value 1.2; The preset neighborhood radius is twice the average of the Euclidean distances between pairwise middle-dimensional data points. If the density estimate value of the data point region is higher than the high density threshold, multiply the preset neighborhood radius by the reduction factor to narrow the range. If it is lower than the low density threshold, multiply the preset neighborhood radius by the expansion factor to expand the range, and obtain the middle-dimensional neighborhood set after dynamic adjustment; Calculate the Euclidean distance between any two data points in the low-dimensional data. Calculate the interquartile range based on the Euclidean distance. Add the lower quartile to 1.5 times the interquartile range to obtain the distance threshold. With the target point as the center, include the data points smaller than the distance threshold in the low-dimensional neighborhood set.
[0009] As a further method, the method of mapping dimensional data to a low-dimensional space based on neighborhoods through local linear embedding includes: Perform local linear embedding separately based on the neighborhood sets of different-dimensional data. For each data point in the neighborhood set, use other data points within its neighborhood for linear reconstruction, construct a reconstruction error function, and add constraint conditions to the error function. , where is the reconstruction weight of the neighborhood point for the data point. Solve the error function to obtain the optimal reconstruction weight; Use the optimal reconstruction weight to construct the low-dimensional embedding objective function for different neighborhood sets. The formula of the objective function is: , where Z is the low-dimensional embedding coordinate matrix, is the low-dimensional coordinate of the i-th data point, M is the total number of data points, K is the neighborhood size, is the optimal reconstruction weight of neighborhood point j for data point i, is the low-dimensional coordinate of neighborhood point j for data point i; Convert the objective function into the form of a Gram matrix, calculate the eigenvalues and eigenvectors of the Gram matrix, retain the eigenvectors corresponding to the smallest 2 - 3 non-zero eigenvalues, obtain the mapping coordinates of the three-dimensional data sets in the low-dimensional space, map the dimensional data to the low-dimensional space, and splice the three types of low-dimensional features in descending order to form a unified low-dimensional feature vector.
[0010] As a further method, the method of obtaining the marginal probability by combining the value frequency and probability density includes: Traverse the discrete data of the low-dimensional feature vector in the low-dimensional space, count the number of occurrences of each discrete value, and calculate its value frequency; Traverse all data points for the continuous data to calculate the kernel function values, and perform a weighted sum of all kernel function values to obtain the probability density estimate value. The formula for the probability density estimate value is: , where is the probability density estimate value, N is the total number of sample data points, d is the dimension of the data, h is the key hyperparameter for kernel density estimation, x is the target data point, is the n-th sample object; Integrate the probability density estimation results corresponding to all data points to form a probability density function of continuous data, obtain the probability density according to the probability density function, and normalize the probability density; Calculate the information entropy of discrete data and continuous data. Divide the information entropy of discrete data by the sum of the information entropy of discrete data and continuous data to obtain the information weight. Multiply the information weight by the value frequency of discrete data, multiply the density probability of continuous data by the inverse weight of the information weight, and add the two results to obtain the marginal probability.
[0011] As a further method, the method of using the marginal probability to determine the weight of dimensional data in data aggregation and splicing the dimensional data features after dimensionality reduction according to the weight to obtain the aggregated comprehensive features includes: Divide the dimensional data in the low-dimensional space into three levels according to the marginal probability. Retain the features of the high-probability layer, perform element-wise multiplication on the features of the medium-probability layer and the high-probability layer, and filter the low-probability layer through the gating formula. The formula is: , where is the low-dimensional feature, is the low-dimensional marginal probability, is the high-dimensional marginal probability, is the medium-dimensional marginal probability, is the normalization factor, is the total sum of feature dimensions, is the low-dimensional feature after gating screening, and G is the gating signal; Piecewise splice the high-dimensional, medium-dimensional, and low-dimensional features after hierarchical processing in the order of high, medium, and low levels. After splicing, perform probability normalization on each level block internally to obtain the aggregated comprehensive features.
[0012] The second aspect of the present invention provides a data aggregation system based on multi-modal features, including: A data acquisition module for collecting multi-modal data and preprocessing the multi-modal data to obtain dimensional data, where the dimensional data includes high-dimensional, medium-dimensional, and low-dimensional data; A data conversion module for locating the nearest neighbor points of data points for high-dimensional data through the ball tree algorithm, dynamically adjusting the neighborhood range according to the distribution density of medium-dimensional data, calculating the neighborhood for low-dimensional data through the Euclidean distance, and mapping the dimensional data to the low-dimensional space based on the neighborhood through local linear embedding; A data processing module for traversing the dimensional data in the low-dimensional space, counting the value frequency of discrete data, obtaining the probability density of continuous data, and combining the value frequency and probability density to obtain the marginal probability; A data aggregation module, which is used to determine the weight of dimensional data in data aggregation using marginal probability, and splice the dimensional data features after dimensionality reduction according to the weight to obtain the aggregated comprehensive features.
[0013] Compared with the prior art, the embodiments of the present invention have at least the following advantages or beneficial effects: (1) The present invention implements differential processing for different dimensional data characteristics. High-dimensional data uses the ball tree algorithm to locate neighbor points to improve the search efficiency. Medium-dimensional data dynamically adjusts the neighborhood range according to the distribution density. Low-dimensional data calculates the neighborhood through the Euclidean distance. This hierarchical strategy avoids the drawbacks of the traditional "one-size-fits-all" approach, balances the processing efficiency of different dimensional data and the feature utilization effect, and significantly improves the overall feature utilization efficiency. (2) The present invention calculates the marginal probability through the value frequency, probability density, and information entropy to determine the aggregation weight. It statistically counts the value frequency for discrete data, estimates the probability density for continuous data and normalizes it, and then allocates the weight based on the information entropy ratio, highlighting high-information features and suppressing redundant information. The dynamic weight allocation mechanism strengthens the complementarity between multimodal features and effectively improves the aggregation accuracy. (3) The present invention uses local linear embedding to map high-dimensional data to a low-dimensional space, reducing the dimension while preserving the local structural relationship of the data, reducing the subsequent calculation complexity. Then, the high-, medium-, and low-dimensional mapped features are spliced hierarchically to achieve deep cross-modal information fusion, which not only reduces redundancy but also retains the essential features and cross-modal associations of the data, providing an efficient input for subsequent analysis tasks. Description of the Drawings
[0014] Figure 1 It is a flowchart of the steps of a data aggregation method based on multimodal features in an embodiment of the present invention. Detailed Embodiments
[0015] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0016] Refer to Figure 1 As shown, the present invention provides a data aggregation method based on multimodal features, including: Step A: Collect multimodal data and preprocess the multimodal data to obtain dimensional data, where the dimensional data includes high-dimensional, medium-dimensional, and low-dimensional data.
[0017] In the actual evaluation, multi-modal monitoring data of the lithium battery coater is collected, including numerical data such as doctor blade temperature, coating pressure, and motor speed; text data such as equipment logs; image data captured by an industrial camera of the doctor blade contact point; video data of the coating process; and audio data of abnormal noises from the doctor blade. The above data is preprocessed. For numerical data such as temperature and pressure, outliers are cleaned and missing values are filled. Text data filters out unstructured characters and generates 768-dimensional word vectors through the BERT model. Image data with a resolution of 1920×1080 is reduced to 200 dimensions by ResNet50 and normalized according to the ImageNet standard. 5 key frames are extracted from the video data, and each frame outputs 2048-dimensional features by ResNet50. After frame-by-frame splicing, the dimension is reduced to 512 dimensions by PCA. The audio data extracts 128-dimensional spectral features after Fourier transform. Finally, they are sorted by the number of dimensions: 512 dimensions for video and 768 dimensions for text are high-dimensional, 200 dimensions for image and 128 dimensions for audio are medium-dimensional, and 3 dimensions for numerical values are low-dimensional; In the actual evaluation, video modal feature dimension data is collected. The maximum-minimum difference of the statistical data window is 0.1. The current newly input data point is determined to be 0.25, the total number of data points is 200, the cumulative mean of the first n - 1 data points is 0.15, and the cumulative variance is 0.08. The dynamic scaling threshold is calculated to obtain 0.89915. The 512-dimensional feature value of the video modal is 0.75 and is scaled to obtain 0.834.
[0018] In step B, the high-dimensional data locates the nearest neighbor points of the data points through the ball tree algorithm, dynamically adjusts the neighborhood range according to the distribution density of the medium-dimensional data, calculates the neighborhood for the low-dimensional data through the Euclidean distance, and maps the dimensional data to a low-dimensional space through local linear embedding based on the neighborhood.
[0019] It should be noted that due to the dimensionality characteristics differences of multi-modal data, traditional nearest neighbor search for high-dimensional data is computationally complex due to the curse of dimensionality. The ball tree algorithm optimizes the search process through spatial recursive partitioning and hypersphere pruning to ensure the efficiency of high-dimensional data neighborhood location; the distribution of medium-dimensional data often has non-uniformity, and dynamically adjusting the neighborhood range can adaptively expand or shrink the search radius based on density estimation to avoid feature extraction biases caused by a fixed neighborhood; the low-dimensional data structure is simple, and the Euclidean distance can be directly used to quickly calculate the neighborhood, taking into account both efficiency and accuracy. Inputting the three neighborhood results into local linear embedding, through reconstruction weight optimization and objective function solution, it can not only retain the local linear structure features of each dimensional data but also achieve dimensionality reduction and unification of cross-dimensional features, providing a concise and effective low-dimensional feature representation for subsequent multi-modal data fusion analysis.
[0020] In actual evaluation, 1000 high-dimensional data are taken from the high-dimensional dataset as spatial points. The recursive termination condition is set that the data volume of the leaf node is less than or equal to 30. Take a subset of 200 data of a node, calculate the mean value of each dimension, determine the center of the hypersphere, calculate the distances from all points within the subset to the center, take the maximum distance as the radius of the hypersphere, divide the data into child nodes according to the distance, and repeat this process. After 3-layer recursion, a ball tree structure including the center, radius, and child nodes is generated. Select the target high-dimensional data point, start traversing the ball tree from the root node, calculate the Euclidean distance between the data point and the center of the hypersphere of the current node. After reaching the leaf node, calculate the Euclidean distance from the data point to all data points within the leaf node, record the 5 points with the smallest distances as the initial nearest neighbor points, trace back along the tree from the leaf node, calculate the minimum distance between the data point and the surface of the hypersphere of each node passed, when the minimum distance is less than the distance of the farthest nearest neighbor point, expand the node, recalculate the distances between the data points within the node and the target high-dimensional data point, and update the nearest neighbor point set, and finally determine the high-dimensional neighborhood set including 10 nearest neighbor points.
[0021] In actual evaluation, 800 middle-dimensional data are obtained. For each middle-dimensional data point, take a middle-dimensional data point and another data point, their Euclidean distance is 0.2, substitute it into the Gaussian kernel function to obtain its contribution to the density of the target point as 0.328, and the bandwidth parameter is obtained by the Silverman rule through the variance of the middle-dimensional data as 0.134. The density estimate value of a data point is calculated as 2.239. After traversing and calculating all data points, the density values are sorted in ascending order. Take the 75% quantile value as the high-density threshold and the 25% quantile value as the low-density threshold. Calculate the proportion of data points higher than the high threshold as 0.22. , the shrinkage coefficient is taken as 0.956; The preset neighborhood radius is 2 times the average of the Euclidean distances between data points. If the regional density of the data point is higher than the high threshold, adjust the radius to 1.147 to generate the middle-dimensional neighborhood set; Take 500 low-dimensional data, calculate the pairwise Euclidean distances between data points to obtain a distance set, use the distance set to calculate the interquartile range, determine the distance threshold as 1.3, take a target numerical point as the center, and include the data points with distances less than 1.3 into the low-dimensional neighborhood set, and finally obtain the low-dimensional neighborhood set including 20 nearest neighbor points.
[0022] In actual evaluation, for the high-, middle-, and low-dimensional neighborhood sets, for each data point , use the neighborhood of 5 points to perform linear reconstruction, construct an error function , add a constraint , force the sum of the reconstruction weights of the neighborhood points to be 1, and solve the optimal reconstruction weights by the Lagrange multiplier method. For a high-dimensional data, the optimal reconstruction weight vector is [0.18, 0.22, 0.20, 0.19, 0.21]. The objective function is constructed and transformed into the form of Gram matrix. A Gram matrix of order 1000 is calculated, and its eigenvalues and eigenvectors are solved. After mapping the 1280-dimensional high-dimensional data, 2D coordinates are obtained. For the medium-dimensional data of 328 dimensions, the mapping is For the low-dimensional data of 3 dimensions, the mapping is They are concatenated in the order of high, medium and low to form a 6D unified low-dimensional feature vector [0.32, -0.15, 0.08, 0.21, -0.09, 0.12], providing a standardized input for the subsequent fault diagnosis model.
[0023] Step C traverses the dimensional data in the low-dimensional space, counts the value frequencies of discrete data, obtains the probability density of continuous data, and combines the value frequencies and probability densities to obtain the marginal probability.
[0024] It should be noted that actual data often mixes discrete and continuous features. A single analysis method cannot fully describe the data distribution characteristics. By integrating and statistically analyzing discrete frequencies step by step, the category distribution information is retained, the continuous kernel density is calculated to capture the numerical distribution law, and then the weights are determined based on information entropy for weighting. When the proportion of discrete information entropy is higher, its corresponding frequency dominates in the calculation of marginal probability, so as to ensure that high-value information dominates the result. The fused marginal probability adapts to the characteristics of mixed data and more accurately depicts the probability essence of the data, providing a more reliable probability basis for subsequent tasks such as feature aggregation and fault diagnosis.
[0025] In actual evaluation, 1000 low-dimensional feature vector data of a lithium battery coater are collected. Through traversal and statistics, 580 are normal, 220 have abnormal doctor blades, 150 have abnormal coating pressures, and 50 have abnormal motor speeds. The extracted value frequencies are 0.58 for normal, 0.22 for abnormal doctor blades, 0.15 for abnormal coating pressures, and 0.05 for abnormal motor speeds. The mapping value of the coating speed in the low-dimensional feature vector is continuous data, with a dimension of 1, a total of 1000 samples, and a kernel density estimation hyperparameter of 0.3. All data points are traversed to calculate the kernel function values, and the weighted sum of all kernel function values is 620 to obtain the probability density estimation value. Taking the target data point 3.2 as an example, the probability density estimation value is 2.275. The probability density estimation values of all data points are integrated to construct a probability density function, and the probability density is obtained according to the probability density function. The probability density is normalized to obtain ; The information entropy of discrete data is calculated to be 1.276. The information entropy of a continuous data is calculated by integration to be 0.92. The information weight is obtained by dividing the information entropy of the discrete data by the sum of the information entropies of the discrete data and the continuous data, which is 0.42. Multiply the information weight by the value frequency of the discrete data, multiply the density probability of the continuous data by the inverse weight of the information weight, and add the two results to obtain the marginal probability of 0.2032.
[0026] In step D, the marginal probability is used to determine the weight of the dimensional data in data aggregation. The dimensional data features after dimensionality reduction are spliced according to the weight to obtain the aggregated comprehensive features.
[0027] In the actual evaluation, the marginal probability of the low-dimensional features of the lithium battery coater is 0.35 for the high-dimensional (video + text fusion features), 0.28 for the medium-dimensional (image + audio features), and 0.22 for the low-dimensional (numerical features). According to the marginal probability, the dimensional data in the low-dimensional space is divided into three levels. The high-probability layer selects the 2D features of the video texture feature values and text keyword probability values that are strongly related to the abnormal coating quality. The medium-probability layer selects the 2D features of the image edge density and audio spectrum energy values that reflect the operating state of the coater in the marginal probability. The low-probability layer selects the rotational speed fluctuation coefficient among the remaining marginal probability features. Keep the high-probability layer features, and perform element-wise multiplication of the medium-probability layer features and the high-probability layer features to obtain The low-probability layer is obtained by screening through the gating formula. According to the high, medium, and low hierarchical order, the high-dimensional, medium-dimensional, and low-dimensional features after hierarchical processing are block-spliced to obtain [0.75, 0.68, 0.4125, 0.3264, 0.1786]. After splicing, probability normalization is performed on each hierarchical block internally to obtain the aggregated comprehensive features [0.5245, 0.4755, 0.5583, 0.4417, 0.1786], providing high-value inputs for equipment status analysis and fault diagnosis model input.
[0028] In this embodiment, the method for collecting multi-modal data and preprocessing the multi-modal data to obtain dimensional data includes: Collect multi-modal data, clean and denoise the data of different modalities, calculate the mean values of the data of different modalities to fill in the missing values, and use the interquartile range method to correct the outliers. Divide according to the modality of multimodal data to obtain text data, numerical data, audio data, image data, and video data. Obtain the word vector dimension for text data through a word vector model, obtain the numerical dimension according to the numerical columns of numerical data, extract the spectral features and time-domain features of audio data and merge them into audio dimension information, determine the basic dimension for image data through the resolution and the number of color channels, and obtain the video spatio-temporal dimension information based on the image data and the frame rate duration of key frames of video data. Obtain the number of dimensions of each modality dimension information, count all modality dimensions and sort them from high to low, divide the top 20% into the high-dimensional interval, the middle 60% into the medium-dimensional interval, and the bottom 20% into the low-dimensional interval. Put the modality data of different dimensions into the dimension interval to obtain the dimension dataset; Obtain the maximum-minimum difference, mean, and standard deviation of different modality data in the dimension dataset, use the maximum-minimum difference, mean, and standard deviation to determine the scaling threshold, and scale the dimension data based on the scaling threshold.
[0029] In this embodiment, the method of using the maximum-minimum difference, mean, and standard deviation to determine the scaling threshold and scaling the corresponding modality data based on the scaling threshold includes: Obtain the maximum-minimum difference, mean, and standard deviation of different modality data in the dimension dataset, calculate the information entropy of the maximum-minimum difference, mean, and standard deviation for data feature description, obtain the dynamic threshold weight coefficients of the maximum-minimum difference, mean, and standard deviation after normalizing the information entropy, calculate the data difference degree based on the historical threshold, the current data window mean, and the maximum-minimum difference, and obtain the scaling threshold through the difference degree. The formula for the scaling threshold is: , where is the currently newly input data point, n is the total number of current data points, is the cumulative mean of the first n - 1 data points, is the cumulative variance of the first n - 1 data points, M is the maximum value within the data window, m is the minimum value within the data window, last is the historical threshold, , , are the weight coefficients of the maximum-minimum difference, mean, and standard deviation for the threshold, is the attenuation coefficient; Scale the corresponding modality data based on the scaling threshold.
[0030] In this embodiment, the method of locating the nearest neighbor points of data points for high-dimensional data through the ball tree algorithm includes: Based on the high-dimensional dataset, regarding high-dimensional data points as spatial points, recursively partitioning the data space, calculating the mean of the current node's data subset to determine the center point, taking the center point as the center of the hypersphere, and taking the maximum distance from each point to the center point as the radius, generating a ball tree node containing the center, radius, and child nodes, repeating the partitioning until the leaf node data volume threshold or the recursive depth limit is met, and completing the construction of the ball tree; Perform a nearest neighbor search on the target high-dimensional data point based on the ball tree structure. Start traversing from the root node of the ball tree, calculate the distance between the target point and the center of the current node's hypersphere. If the distance exceeds the radius of the hypersphere, prune this branch; if not, continue to search the child nodes. After reaching the leaf node, calculate the distances between the target point and all the data points within the leaf node, and record the initial nearest neighbor points; After obtaining the initial nearest neighbor points, optimize the nearest neighbor set by backtracking the tree structure. Backtrack along the tree from the leaf node. For each node passed through, calculate the minimum distance between the target point and the surface of the node's hypersphere. If this distance is less than the distance of the current farthest nearest neighbor point, expand the node, recalculate the distances between the data points within the node and the target point, and update the nearest neighbor point set until the entire tree search is completed, and finally determine the nearest neighbor points of the high-dimensional data point to obtain the high-dimensional neighborhood set.
[0031] In this embodiment, the method for dynamically adjusting the neighborhood range according to the distribution density of the middle-dimensional data and calculating the neighborhood for low-dimensional data through the Euclidean distance includes: Based on each data point in the middle-dimensional dataset, calculate its contribution to the density of the target point through the Gaussian kernel function. Obtain the bandwidth parameter using Silverman's rule for the standard deviation of the middle-dimensional dataset. Divide the sum of the contributions of all data points by the total number of data points and the bandwidth parameter to obtain the density estimate value of each data point. Sort the density estimate values in ascending order, take the 75th percentile value as the high-density threshold, and the 25th percentile value as the low-density threshold; Calculate the proportion of data points higher than the high-density threshold , if , the reduction factor takes , if , take the default value 0.8; calculate the proportion of data points lower than the low-density threshold , if , the expansion factor takes , if , take the default value 1.2; Preset the neighborhood radius as the average of twice the Euclidean distance between pairwise middle-dimensional data points. If the density estimate value of the data point region is higher than the high-density threshold, multiply the preset neighborhood radius by the reduction factor to narrow the range. If it is lower than the low-density threshold, multiply the preset neighborhood radius by the expansion factor to expand the range, and obtain the middle-dimensional neighborhood set after dynamic adjustment; Calculate the Euclidean distance between any two data points in the low-dimensional data. Calculate the interquartile range based on the Euclidean distance. Add the lower quartile to 1.5 times the interquartile range to obtain the distance threshold. With the target point as the center, include the data points with distances less than the distance threshold in the low-dimensional neighborhood set.
[0032] In this embodiment, the method of mapping dimensional data to a low-dimensional space through local linear embedding based on neighborhoods includes: Perform local linear embedding separately on the neighborhood sets of different dimensional data. For each data point in the neighborhood set, use other data points within its neighborhood for linear reconstruction, construct a reconstruction error function, and add constraint conditions to the error function , where is the reconstruction weight of the neighborhood point for the data point. Solve the error function to obtain the optimal reconstruction weight; Use the optimal reconstruction weight to construct the low-dimensional embedding objective function for different neighborhood sets. The formula of the objective function is: , where Z is the low-dimensional embedding coordinate matrix, is the low-dimensional coordinate of the i-th data point, M is the total number of data points, K is the neighborhood size, is the optimal reconstruction weight of the neighborhood point j for the data point i, is the low-dimensional coordinate of the neighborhood point j for the data point i; Convert the objective function into the form of a Gram matrix, calculate the eigenvalues and eigenvectors of the Gram matrix, retain the eigenvectors corresponding to the smallest 2 - 3 non-zero eigenvalues, obtain the mapping coordinates of the three-dimensional data sets in the low-dimensional space, map the dimensional data to the low-dimensional space, and splice the three types of low-dimensional features in descending order to form a unified low-dimensional feature vector.
[0033] In this embodiment, the method of obtaining the marginal probability by combining the value frequency and probability density includes: Traverse the discrete data of the low-dimensional feature vector in the low-dimensional space, count the number of occurrences of each discrete value, and calculate its value frequency; Traverse all data points for the continuous data to calculate the kernel function values, and perform a weighted sum of all kernel function values to obtain the probability density estimate value. The formula for the probability density estimate value is: , where is the probability density estimate value, N is the total number of sample data points, d is the dimension of the data, h is the key hyperparameter for kernel density estimation, x is the target data point, is the n-th sample object; Integrate the probability density estimation results corresponding to all data points to form the probability density function of continuous data, obtain the probability density according to the probability density function, and normalize the probability density; Calculate the information entropy of discrete data and continuous data. Divide the information entropy of discrete data by the sum of the information entropy of discrete data and continuous data to obtain the information weight. Multiply the information weight by the value frequency of discrete data, multiply the density probability of continuous data by the inverse weight of the information weight, and add the two results to obtain the marginal probability.
[0034] In this embodiment, the method of using the marginal probability to determine the weight of dimensional data in data aggregation and splicing the dimensional data features after dimensionality reduction according to the weight to obtain the aggregated comprehensive features includes: Divide the dimensional data in the low-dimensional space into three levels according to the marginal probability. Retain the features of the high-probability layer, perform element-wise multiplication of the features of the medium-probability layer and the high-probability layer, and screen the low-probability layer through the gating formula. The formula is: , where is the low-dimensional feature, is the low-dimensional marginal probability, is the high-dimensional marginal probability, is the medium-dimensional marginal probability, is the normalization factor, is the sum of feature dimensions, is the low-dimensional feature after gating screening, and G is the gating signal; Piecewise splice the high-dimensional, medium-dimensional, and low-dimensional features after hierarchical processing in the order of high, medium, and low levels. After splicing, perform probability normalization within each level block to obtain the aggregated comprehensive features.
[0035] The second aspect of the present invention also provides a data aggregation system based on multi-modal features, including: A data acquisition module for acquiring multi-modal data and preprocessing the multi-modal data to obtain dimensional data, where the dimensional data includes high-dimensional, medium-dimensional, and low-dimensional data; A data conversion module for locating the nearest neighbor points of data points for high-dimensional data through the ball tree algorithm, dynamically adjusting the neighborhood range according to the distribution density of medium-dimensional data, calculating the neighborhood for low-dimensional data through the Euclidean distance, and mapping the dimensional data to the low-dimensional space based on the neighborhood through local linear embedding; A data processing module for traversing the dimensional data in the low-dimensional space, counting the value frequency of discrete data, obtaining the probability density of continuous data, and combining the value frequency and probability density to obtain the marginal probability; A data aggregation module, which is used to determine the weight of dimension data in data aggregation using marginal probability, splice the dimension data features after dimensionality reduction according to the weight, and obtain the aggregated comprehensive features.
[0036] The above content is only an example and explanation of the structure of the present invention. Those skilled in the art of the present technology can make various modifications or supplements to the described specific embodiments or use similar methods for substitution. As long as they do not deviate from the structure of the invention or exceed the scope defined by this claim book, they should all fall within the protection scope of the present invention.
Claims
1. A data aggregation method based on multi-modal features, characterized in that Including the following steps: σ Collect multimodal data, preprocess the multimodal data to obtain dimensional data, and the dimensional data includes high-dimensional, medium-dimensional, and low-dimensional data; Locate the nearest neighbor points of the data points for the high-dimensional data through the ball tree algorithm, dynamically adjust the neighborhood range according to the distribution density of the medium-dimensional data, calculate the neighborhood for the low-dimensional data through the Euclidean distance, and map the dimensional data to the low-dimensional space through local linear embedding based on the neighborhood; Traverse the dimensional data in the low-dimensional space, count the value frequencies of discrete data, obtain the probability density of continuous data, and combine the value frequencies and probability densities to obtain the marginal probability; Use the marginal probability to determine the weights of the dimensional data in data aggregation, splice the dimensional data features after dimensionality reduction according to the weights, and obtain the aggregated comprehensive features.
2. The data aggregation method based on multimodal features according to claim 1, characterized in that The method of collecting multimodal data and preprocessing the multimodal data to obtain dimensional data includes: Collect multimodal data, clean and denoise the data of different modalities, calculate the mean values of the data of different modalities to fill in the missing values, and use the interquartile range method to correct the outliers; Divide according to the modalities of the multimodal data to obtain text data, numerical data, audio data, image data, and video data. Obtain the word vector dimension for the text data through the word vector model, obtain the numerical dimension according to the numerical columns of the numerical data, extract the spectral features and time-domain features of the audio data and combine them into audio dimension information, determine the basic dimension for the image data through the resolution and the number of color channels, obtain the video spatio-temporal dimension information based on the image data and the frame rate duration of the key frames of the video data, obtain the number of dimensions of the dimension information of each modality, count the number of dimensions of all modalities and sort them from high to low, divide the first 20% into the high-dimensional interval, the middle 60% into the medium-dimensional interval, and the last 20% into the low-dimensional interval, and put the modality data of different dimensions into the dimension interval to obtain the dimension dataset; Obtain the maximum-minimum difference, mean, and standard deviation of the modality data in the dimension dataset, use the maximum-minimum difference, mean, and standard deviation to determine the scaling threshold, and scale the dimensional data based on the scaling threshold.
3. The data aggregation method based on multi-modal features according to claim 2, wherein The method of using the maximum-minimum difference, mean, and standard deviation to determine the scaling threshold and scaling the corresponding modality data based on the scaling threshold includes: Obtain the maximum-minimum difference, mean, and standard deviation of the modality data in the dimension dataset, calculate the information entropy of the maximum-minimum difference, mean, and standard deviation for data feature description, obtain the dynamic threshold weight coefficients of the maximum-minimum difference, mean, and standard deviation after normalizing the information entropy, calculate the data difference degree based on the historical threshold, the current data window mean, and the maximum-minimum difference, and obtain the scaling threshold through the difference degree. The formula for the scaling threshold is: , where is the current newly input data point, n is the total number of current data points, is the cumulative mean of the previous n - 1 data points, is the cumulative variance of the previous n - 1 data points, M is the maximum value within the data window, m is the minimum value within the data window, and last is the historical threshold, , , are the weight coefficients of the maximum - minimum difference, mean, and standard deviation with respect to the threshold, is the attenuation coefficient; Scale the corresponding modality data based on the scaling threshold.
4. A data aggregation method based on multi-modal features according to claim 1, characterized in that, The method of locating the nearest neighbor points of the data points for the high-dimensional data through the ball tree algorithm includes: Based on the high-dimensional dataset, regarding high-dimensional data points as spatial points, recursively partitioning the data space, calculating the mean of the data subset of the current node to determine the center point, taking the center point as the center of the hypersphere, and taking the maximum distance from each point to the center point as the radius to generate a ball tree node containing the center, radius, and child nodes. Repeat the partitioning until the leaf node data volume threshold or the recursive depth limit is met to complete the construction of the ball tree; Perform a nearest neighbor search on the target high-dimensional data point based on the ball tree structure. Start traversing from the root node of the ball tree, calculate the distance between the target point and the center of the hypersphere of the current node. If the distance exceeds the radius of the hypersphere, prune this branch; if not, continue to search the child nodes. After reaching the leaf node, calculate the distances between the target point and all data points within the leaf node, and record the initial nearest neighbor points; After obtaining the initial nearest neighbor points, optimize the nearest neighbor set by backtracking the tree structure. Backtrack along the tree from the leaf node. For each node passed through, calculate the minimum distance between the target point and the surface of the hypersphere of the node. If this distance is less than the distance of the current farthest nearest neighbor point, expand the node, recalculate the distances between the data points within the node and the target point, and update the nearest neighbor point set until the entire tree search is completed. Finally, determine the nearest neighbor points of the high-dimensional data points to obtain the high-dimensional neighborhood set.
5. A multimodal feature-based data aggregation method according to claim 1, characterized in that, The method of dynamically adjusting the neighborhood range according to the distribution density of the middle-dimensional data and calculating the neighborhood for low-dimensional data by Euclidean distance includes: Based on each data point in the middle-dimensional dataset, calculate its contribution to the density of the target point through the Gaussian kernel function. Use the Silverman rule for the standard deviation of the middle-dimensional dataset to obtain the bandwidth parameter. Divide the sum of the contributions of all data points by the total number of data points and the bandwidth parameter to obtain the density estimate value of each data point. Sort the density estimate values in ascending order, take the 75th percentile value as the high-density threshold, and the 25th percentile value as the low-density threshold; Calculate the proportion of data points higher than the high density threshold , if , the reduction coefficient is taken as , if , take the default value 0.8; Calculate the proportion of data points lower than the low density threshold , if , the expansion coefficient is taken as , if , take the default value 1.2; Preset the neighborhood radius as the average of twice the Euclidean distances between pairwise middle-dimensional data points. If the density estimate value of the data point region is higher than the high-density threshold, multiply the preset neighborhood radius by a reduction factor to narrow the range. If it is lower than the low-density threshold, multiply the preset neighborhood radius by an expansion factor to expand the range, and obtain the middle-dimensional neighborhood set after dynamic adjustment; Calculate the Euclidean distance between any two data points for the low-dimensional data, calculate the interquartile range based on the Euclidean distance, add the lower quartile to 1.5 times the interquartile range to obtain the distance threshold. Taking the target point as the center, include the data points smaller than the distance threshold in the low-dimensional neighborhood set.
6. A multimodal feature-based data aggregation method according to claim 1, characterized in that The method of mapping dimensional data to a low-dimensional space based on the neighborhood through local linear embedding includes: Perform local linear embedding on the neighborhood sets based on data in different dimensions. For each data point in the neighborhood set, use other data points within its neighborhood for linear reconstruction, construct a reconstruction error function, and add constraint conditions to the error function , where is the reconstruction weight of the neighborhood point for the data point, and solve the error function to obtain the optimal reconstruction weight; Construct a low-dimensional embedding objective function for different neighborhood sets using the optimal reconstruction weights. The formula of the objective function is: , where Z is a low-dimensional embedding coordinate matrix, is the low-dimensional coordinate of the i-th data point, M is the total number of data points, K is the neighborhood size, is the optimal reconstruction weight of neighborhood point j for data point i, is the low-dimensional coordinate of neighborhood point j for data point i; Transform the objective function into the form of the Gram matrix, calculate the eigenvalues and eigenvectors of the Gram matrix, retain the eigenvectors corresponding to the smallest 2 - 3 non-zero eigenvalues, obtain the mapping coordinates of the three-dimensional datasets in the low-dimensional space, map the dimensional data to the low-dimensional space, and splice the three types of low-dimensional features in descending order to form a unified low-dimensional feature vector.
7. A multimodal feature-based data aggregation method according to claim 1, characterized in that The method of obtaining the marginal probability by combining the value frequency and the probability density includes: Traverse the discrete data of the low-dimensional feature vectors in the low-dimensional space, count the occurrence times of each discrete value, and calculate its value frequency; Traverse all data points for continuous data to calculate the kernel function values, and sum the weighted values of all kernel function values to obtain the probability density estimate value. The formula for the probability density estimate value is: , where is the probability density estimate value, N is the total number of sample data points, d is the dimension of the data, h is the key hyperparameter of kernel density estimation, x is the target data point, is the nth sample object; Integrate the probability density estimate value results corresponding to all data points to form the probability density function of the continuous data, obtain the probability density according to the probability density function, and normalize the probability density; Calculate the information entropy of the discrete data and the continuous data. Divide the information entropy of the discrete data by the sum of the information entropy of the discrete data and the continuous data to obtain the information weight. Multiply the information weight by the value frequency of the discrete data, multiply the density probability of the continuous data by the inverse weight of the information weight, and add the two results to obtain the marginal probability.
8. A multimodal feature-based data aggregation method according to claim 1, characterized in that, The method of using the marginal probability to determine the weight of the dimensional data in data aggregation and splicing the dimensional data features after dimensionality reduction according to the weight to obtain the aggregated comprehensive features includes: Divide the dimensional data in the low-dimensional space into three levels according to the marginal probability. Retain the features of the high-probability level, perform element-wise multiplication of the features of the medium-probability level and the high-probability level, and filter the low-probability level through the gating formula. The formula is: , wherein is a low-dimensional feature, is a low-dimensional marginal probability, is a high-dimensional marginal probability, is a medium-dimensional marginal probability, is a normalization factor, is the total sum of feature dimensions, is the low-dimensional feature after gated screening, and G is a gating signal; Piecewise splice the high-dimensional, medium-dimensional, and low-dimensional features after hierarchical processing in the order of high, medium, and low levels. After splicing, perform probability normalization on each level block internally to obtain the aggregated comprehensive features.
9. A data aggregation system based on multimodal features for performing a data aggregation method based on multimodal features according to any one of claims 1 to 8, characterized in that The system includes: A data acquisition module for acquiring multi-modal data and preprocessing the multi-modal data to obtain dimensional data, where the dimensional data includes high-dimensional, medium-dimensional, and low-dimensional data; A data conversion module for locating the nearest neighbor points of the data points for high-dimensional data through the ball tree algorithm, dynamically adjusting the neighborhood range according to the distribution density of the medium-dimensional data, calculating the neighborhood for low-dimensional data through the Euclidean distance, and mapping the dimensional data to the low-dimensional space based on the neighborhood through local linear embedding; A data processing module for traversing the dimensional data in the low-dimensional space, counting the value frequency of the discrete data, obtaining the probability density of the continuous data, and combining the value frequency and the probability density to obtain the marginal probability; A data aggregation module for using the marginal probability to determine the weight of the dimensional data in data aggregation and splicing the dimensional data features after dimensionality reduction according to the weight to obtain the aggregated comprehensive features.
Citation Information
Patent Citations
Multi-data-source semantic intelligent analysis and event scene restoration method and device
CN109635107A
Relation cluster database optimization method based on multi-modal learning
CN119046316A
Retrieval result determination method and device, electronic equipment and storage medium
CN119829793A
System and method for inducing sleep by transplanting mental states
US20190321583A1
Cited By
Multi-mode propeller fault diagnosis method and diagnosis system
CN121008566A
Multi-modal propeller fault diagnosis method and diagnosis system
CN121008566B
Intelligent data analysis method based on wastewater of chemical production line
CN121148536A