A data aggregation method based on multi-modal features

By using the ball tree algorithm and local linear embedding technology, the neighborhood range of multimodal data is dynamically adjusted. The weights are calculated by combining the value frequency and probability density. This solves the problems of information redundancy and computational complexity in multimodal data fusion, and achieves efficient cross-modal feature aggregation and accurate data fusion.

CN120372555BActive Publication Date: 2025-12-23CHINESE ACAD OF INSPECTION & QUARANTINE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510576972.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-12-23
Estimated Expiration
2045-05-06

AI Technical Summary

Technical Problem

Existing technologies suffer from information redundancy, low computational efficiency, and decision bias in multimodal data fusion, and have high computational complexity, making it difficult to effectively capture cross-modal correlation features.

Method used

The ball tree algorithm is used to locate the nearest neighbor of high-dimensional data, dynamically adjust the neighborhood range of mid-dimensional data, calculate the neighborhood of low-dimensional data through Euclidean distance, and map the data to the low-dimensional space using local linear embedding. The marginal probability is calculated by combining the value frequency and probability density, the data aggregation weight is determined, and feature splicing is performed.

Benefits of technology

It improves the efficiency of multimodal data processing, balances the feature utilization effect of data from different dimensions, enhances feature utilization efficiency and aggregation accuracy, reduces computational complexity, and achieves deep fusion of cross-modal information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372555B_ABST
    Figure CN120372555B_ABST
Patent Text Reader

Abstract

The application discloses a data aggregation method based on multi-modal features, comprising: collecting multi-modal data, pre-processing, extracting features according to modes and dividing into high, medium and low dimensional data, for high dimensional data, using ball tree algorithm to locate the near neighbor point; medium dimensional data is based on distribution density to dynamically adjust the neighborhood range; the low dimensional data is calculated by the Euclidean distance, and then is mapped to the low dimensional space by the aid of the local linear embedding, then traverses the low dimensional data, and the discrete data value frequency and the continuous data probability density are counted, the marginal probability is calculated by combining the information entropy, and the data aggregation weight is determined, finally, the features after dimension reduction are spliced in the order of high, medium and low levels, the probability normalization is carried out in each level block, and the aggregated comprehensive features are generated. The method realizes the aggregation of multi-modal features through multi-dimensional differentiated processing and weight calculation based on data distribution, and improves the feature complementarity and accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the field of resource management, and particularly relates to a data aggregation method based on multi-modal features. BACKGROUND

[0002] With the rapid development of the Internet of Things and intelligent sensing technology, multi-modal data fusion has become a core technology in the fields of smart cities, industrial detection, medical diagnosis, etc. Traditional data aggregation methods mostly adopt early fusion or late fusion architecture: the former integrates multi-source data through simple feature splicing, but ignores the dimensional differences and semantic gaps between different modalities, resulting in information redundancy and low computational efficiency; the latter relies on independent models to process each modality data, which alleviates the dimensional conflict, but is difficult to capture cross-modality correlation features, and is prone to decision bias in dynamic scenarios; existing improvement schemes partially solve the above problems, but their computational complexity grows exponentially, and the weight distribution process is disconnected from the intrinsic distribution characteristics of data. SUMMARY

[0003] The application aims to provide a data aggregation method based on multi-modal features.

[0004] To achieve the above-mentioned purpose, the application is implemented according to the following technical solutions:

[0005] The application provides a data aggregation method based on multi-modal features, comprising:

[0006] Collecting multi-modal data, pre-processing the multi-modal data to obtain dimensional data, the dimensional data including high-dimensional, medium-dimensional and low-dimensional data;

[0007] Locating the near neighbor points of the data points of the high-dimensional data through the ball tree algorithm, dynamically adjusting the neighborhood range according to the distribution density of the medium-dimensional data, calculating the neighborhood for the low-dimensional data through the Euclidean distance, and mapping the dimensional data to a low-dimensional space based on the neighborhood through local linear embedding;

[0008] Traversing the dimensional data in the low-dimensional space, counting the value frequency of the discrete data, obtaining the probability density of the continuous data, and combining the value frequency and the probability density to obtain the marginal probability;

[0009] Determining the weight of the dimensional data in data aggregation using the marginal probability, splicing the features of the reduced dimensional data according to the weight, and obtaining the aggregated comprehensive features.

[0010] As a further method, the method of collecting multi-modal data and pre-processing the multi-modal data to obtain dimensional data comprises:

[0011] The multi-modal data is collected, the data of different modalities is cleaned and denoised, the mean of the data of different modalities is calculated, the missing values are filled, and the quartile range method is used to correct the abnormal values.

[0012] According to the modalities of the multi-modal data, text data, numerical data, audio data, image data and video data are obtained, the word vector dimension of the text data is obtained through a word vector model, the numerical dimension is obtained according to the numerical column of the numerical data, the frequency spectrum feature and the time domain feature of the audio data are extracted and combined into audio dimension information, the basic dimension of the image data is determined through the resolution and the number of color channels, the video space-time dimension information is obtained based on the image data and the frame rate length of the key frames of the video data, the dimension number of each modality dimension information is obtained, all modality dimension numbers are sorted from high to low, the first 20% is divided into a high-dimensional interval, the middle 60% is a medium-dimensional interval, and the last 20% is a low-dimensional interval, and the modality data of different dimensions is put into the dimension interval to obtain a dimension data set.

[0013] The maximum difference, the mean and the standard deviation of the different modal data in the dimension data set are obtained, the scaling threshold is determined using the maximum difference, the mean and the standard deviation, and the dimension data is scaled based on the scaling threshold.

[0014] As a further method, the method of using the maximum difference, the mean and the standard deviation to determine the scaling threshold and scaling the corresponding modality data based on the scaling threshold comprises:

[0015] The maximum difference, the mean and the standard deviation of the different modal data in the dimension data set are obtained, the information entropy of the maximum difference, the mean and the standard deviation to data feature description is calculated, the dynamic threshold weight coefficient of the maximum difference, the mean and the standard deviation is obtained after normalization of the information entropy, the data difference degree is calculated based on the historical threshold, the current data window mean and the maximum difference, and the scaling threshold is obtained through the difference degree, and the scaling threshold calculation formula is:

[0016] ,

[0017] wherein is the current new input data point, n is the total number of the current data point, is the cumulative mean of the first n-1 data points, is the cumulative variance of the first n-1 data points, M is the maximum value in the data window, m is the minimum value in the data window, and last is the historical threshold, , , is the weight coefficient of the maximum minimum difference, the mean and the standard deviation to the threshold, is the decay coefficient;

[0018] The corresponding modality data is scaled based on the scaling threshold.

[0019] As a further method, the method for locating the near-neighbor points of the data points of the high-dimensional data by the ball tree algorithm comprises the following steps:

[0020] Based on the high-dimensional data set, the high-dimensional data points are regarded as space points, the data space is divided in a recursive manner, the mean value of the data subset of the current node is calculated to determine the center point, the center point is taken as the center of the hypersphere, the maximum distance of each point to the center point is taken as the radius, a ball tree node containing the center, the radius and the sub-node is generated, the division is repeated until the data amount threshold of the leaf node or the recursive depth limit is met, and the ball tree construction is completed;

[0021] Based on the ball tree structure, the near-neighbor search of the target high-dimensional data point is performed, the ball tree root node is traversed, the distance between the target point and the center of the hypersphere of the current node is calculated, if the distance exceeds the radius of the hypersphere, the branch is pruned; if not, the sub-node is continuously searched, after reaching the leaf node, the distance between the target point and all data points in the leaf node is calculated, and the initial near-neighbor points are recorded;

[0022] After the initial near-neighbor points are obtained, the near-neighbor set is optimized by backtracking the tree structure, the tree is backtracked from the leaf node, for each node passed, the minimum distance between the target point and the surface of the hypersphere of the node is calculated, if the distance is less than the distance of the farthest near-neighbor point, the node is expanded, the distance between the data points in the node and the target point is recalculated, the near-neighbor point set is updated, until the full tree search is completed, the near-neighbor points of the high-dimensional data point are finally determined, and the high-dimensional field set is obtained.

[0023] As a further method, the method for dynamically adjusting the neighborhood range according to the distribution density of the medium-dimensional data comprises the following steps:

[0024] Based on each data point of the medium-dimensional data set, the contribution of the data point to the density of the target point is calculated by using a Gaussian kernel function, the bandwidth parameter is obtained by using the Silverman rule for the standard deviation of the medium-dimensional data set, the total sum of the contributions of all data points is divided by the total number of data points and the bandwidth parameter to obtain the density estimation value of each data point, the density estimation values are arranged in ascending order, the 75% quantile value is taken as the high-density threshold, and the 25% quantile value is taken as the low-density threshold;

[0025] The proportion of the data points higher than the high-density threshold is calculated , if , the reduction coefficient is taken as , if , a default value 0.8 is taken; the proportion of the data points lower than the low-density threshold is calculated , if , the expansion coefficient is taken as , if , a default value 1.2 is taken;

[0026] The preset neighborhood radius is twice the average of the Euclidean distance between each pair of data points in the region, if the estimated value of the region density of the data points is higher than the high-density threshold, the preset neighborhood radius is multiplied by a reduction factor to reduce the range, if it is lower than the low-density threshold, the preset neighborhood radius is multiplied by an expansion factor to expand the range, and the medium-dimensional neighborhood set is obtained after dynamic adjustment;

[0027] The Euclidean distance between any two data points in the low-dimensional data is calculated, the interquartile range is calculated based on the Euclidean distance, the distance threshold is obtained by adding the lower quartile and 1.5 times the interquartile range, and the data points less than the distance threshold are included in the low-dimensional neighborhood set.

[0028] As a further method, the method for mapping the dimensional data to a low-dimensional space based on the neighborhood through local linear embedding, comprising:

[0029] The local linear embedding is performed based on the neighborhood set of different dimensional data, for each data point in the neighborhood set, the other data points in the neighborhood are used for linear reconstruction, a reconstruction error function is constructed, and a constraint condition is added to the error function , wherein is the reconstruction weight of the neighborhood point to the data point, and the error function is solved to obtain the optimal reconstruction weight;

[0030] The low-dimensional embedding target function of different neighborhood sets is constructed using the optimal reconstruction weight, and the target function formula is:

[0031] ,

[0032] wherein Z is a low-dimensional embedding coordinate matrix, is the low-dimensional coordinate of the i th data point, M is the total number of data points, and K is the neighborhood size, is the optimal reconstruction weight of the neighborhood point j to the data point i, is the low-dimensional coordinate of the neighborhood point j to the data point i;

[0033] The target function is converted into a Gram matrix form, the eigenvalues and eigenvectors of the Gram matrix are calculated, the eigenvectors corresponding to the smallest 2-3 non-zero eigenvalues are retained, the mapping coordinates of the three-dimensional data set in the low-dimensional space are obtained, the dimensional data is mapped to the low-dimensional space, and the three types of low-dimensional features are spliced in the order from high to low to form a unified low-dimensional feature vector.

[0034] As a further method, the method for obtaining the marginal probability by combining the value frequency and the probability density, comprising:

[0035] The discrete data of the low-dimensional feature vector in the low-dimensional space is traversed, the number of occurrences of each discrete value is counted, and the value frequency is calculated;

[0036] The kernel function value is calculated for all data points of continuous data, and a weighted sum of all kernel function values is obtained to obtain a probability density estimation value, and the probability density estimation value is calculated according to the following formula:

[0037] ,

[0038] wherein is the probability density estimation value, N is the total number of sample data points, d is the dimension of data, h is a key hyperparameter of kernel density estimation, x is a target data point, is the nth sample object;

[0039] The probability density estimation value corresponding to all data points is integrated to form a probability density function of continuous data, and the probability density is obtained according to the probability density function, and the probability density is normalized.

[0040] The information entropy of discrete data and continuous data is calculated, the information weight is obtained by dividing the information entropy of discrete data by the sum of the information entropy of discrete data and continuous data, the value frequency of discrete data is multiplied by the information weight, the density probability of continuous data is multiplied by the inverse weight of the information weight, and the marginal probability is obtained by adding the two results.

[0041] As a further method, the method for determining the weight of dimensional data in data aggregation using the marginal probability, and splicing the dimensional data features after dimension reduction according to the weight to obtain the aggregated comprehensive features, comprises:

[0042] According to the marginal probability, the dimensional data in the low-dimensional space is divided into three levels, the high-probability layer features are reserved, the medium-probability layer features are multiplied by the high-probability layer features element by element, and the low-probability layer is screened by a gating formula, and the formula is:

[0043] ,

[0044] wherein is the low-dimensional feature, is the low-dimensional marginal probability, is the high-dimensional marginal probability, is the medium-dimensional marginal probability, is a normalization factor, is the total sum of feature dimensions, is the low-dimensional feature after gating screening, and G is a gating signal.

[0045] The high-dimensional, medium-dimensional and low-dimensional features after the hierarchical processing are spliced in a block shape according to the order of high, medium and low levels, and the probability of each level block is normalized after splicing to obtain the aggregated comprehensive features.

[0046] The second aspect of the present application provides a data aggregation system based on multi-modal features, comprising:

[0047] a data acquisition module, configured to acquire multi-modal data, pre-process the multi-modal data to obtain dimensional data, the dimensional data comprising high-dimensional, medium-dimensional and low-dimensional data;

[0048] a data conversion module, configured to locate the near neighbor points of the data points of the high-dimensional data by using a ball tree algorithm, dynamically adjust the neighborhood range according to the distribution density of the medium-dimensional data, calculate the neighborhood for the low-dimensional data by using the Euclidean distance, and map the dimensional data to a low-dimensional space based on the neighborhood by using a locally linear embedding;

[0049] a data processing module, configured to traverse the dimensional data in the low-dimensional space, count the value frequency of discrete data, obtain the probability density of continuous data, and obtain the marginal probability by combining the value frequency and the probability density;

[0050] a data aggregation module, configured to determine the weight of the dimensional data in data aggregation using the marginal probability, splice the features of the dimensional data after dimension reduction according to the weight, and obtain the integrated features after aggregation.

[0051] Compared with the prior art, the embodiments of the present application have at least the following advantages or beneficial effects:

[0052] (1) The present application implements differential processing for different dimensional data characteristics, the high-dimensional data locates the near neighbor points by using a ball tree algorithm to improve the search efficiency, the medium-dimensional data dynamically adjusts the neighborhood range according to the distribution density, and the low-dimensional data calculates the neighborhood by using the Euclidean distance. This hierarchical strategy avoids the drawbacks of traditional “one-size-fits-all”, balances the processing efficiency and feature utilization effect of different dimensional data, and significantly improves the overall feature utilization efficiency.

[0053] (2) The present application calculates the marginal probability by using the value frequency, the probability density and the information entropy to determine the aggregation weight. The value frequency of discrete data is counted, the probability density of continuous data is estimated and normalized, and then the weight is proportionally distributed based on the information entropy, which highlights the high information content features and suppresses the redundant information. The dynamic weight distribution mechanism strengthens the complementarity between multi-modal features, effectively improves the aggregation accuracy.

[0054] (3) The present application uses a locally linear embedding to map the high-dimensional data to a low-dimensional space, which reduces the dimension while preserving the local structure of the data, reduces the subsequent calculation complexity, and then splices the mapping features of high, medium and low dimensions according to the hierarchical level to realize the deep fusion of cross-modal information, which not only reduces the redundancy but also preserves the essential features and cross-modal correlation of the data, providing efficient input for subsequent analysis tasks. BRIEF DESCRIPTION OF DRAWINGS

[0055] Figure 1A flow chart of steps of a data aggregation method based on multi-modal features in an embodiment of the present application. DETAILED DESCRIPTION

[0056] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative labor fall within the protection scope of the present application.

[0057] Referring to Figure 1 The present application provides a data aggregation method based on multi-modal features, comprising:

[0058] Step A: Collect multi-modal data, and pre-process the multi-modal data to obtain dimensional data, wherein the dimensional data includes high-dimensional, medium-dimensional and low-dimensional data.

[0059] In actual evaluation, multi-modal monitoring data of a lithium battery coating machine is collected, including numerical data such as scraper temperature, coating pressure and motor speed; text data such as device logs; image data of the scraper contact point captured by an industrial camera; video data of the coating process; and audio data of the scraper abnormal sound. The above data is pre-processed, the abnormal values of the temperature, pressure and other numerical data are cleaned and the missing values are filled, the unstructured characters in the text data are filtered through a BERT model to generate 768-dimensional word vectors, the 1920x1080 resolution image data is reduced to 200 dimensions by ResNet50 and normalized according to the ImageNet standard, the 5 key frames of the video data are extracted, each frame is output by ResNet50 to output 2048-dimensional features, and after frame-by-frame splicing, the data is reduced to 512 dimensions by PCA, and the 128-dimensional frequency spectrum features of the audio data are extracted by Fourier transform; finally, the dimensional numbers are sorted: the video 512-dimensional and the text 768-dimensional are high-dimensional, the image 200-dimensional and the audio 128-dimensional are medium-dimensional, and the numerical 3-dimensional is low-dimensional;

[0060] In actual evaluation, the video modal feature dimensional data is collected, the maximum difference of the data window is 0.1, the current new input data point is determined to be 0.25, the total number of data points is 200, the cumulative mean of the first n-1 data points is 0.15, the cumulative variance is 0.08, the dynamic scaling threshold is calculated to be 0.89915, and the scaling is performed on the 512-dimensional feature value of the video modal to obtain 0.834.

[0061] Step B: The near neighbor points of the data points are located by the ball tree algorithm for the high-dimensional data, the neighborhood range is dynamically adjusted according to the distribution density of the medium-dimensional data, the neighborhood is calculated by the Euclidean distance for the low-dimensional data, and the dimensional data is mapped to a low-dimensional space based on the neighborhood by local linear embedding.

[0062] It needs to be explained that for the dimensional characteristic difference of multi-modal data, high-dimensional data causes traditional neighbor search to be computationally complex due to the curse of dimensionality, the ball tree algorithm optimizes the search process through spatial recursive division and hypersphere pruning to ensure the efficiency of high-dimensional data neighborhood positioning; the distribution of medium-dimensional data often has unevenness, and dynamic adjustment of the neighborhood range can adaptively expand or reduce the search radius based on density estimation to avoid feature extraction deviation caused by fixed neighborhood; low-dimensional data structure is simple, and direct use of Euclidean distance can quickly calculate the neighborhood, balancing efficiency and accuracy, inputting three neighborhood results into local linear embedding, solving through reconstruction weight optimization and objective function, which can not only preserve the local linear structure characteristics of each dimension data, but also realize the dimension reduction of cross-dimensional features, providing simplified and effective low-dimensional feature representation for subsequent multi-modal data fusion analysis.

[0063] In actual evaluation, 1000 high-dimensional data points are taken as space points based on high-dimensional data set, the recursive termination condition is set as the data amount of leaf node being less than or equal to 30, a node data subset of 200 data is taken, the mean of each dimension is calculated, the hypersphere center is determined, the distance of all points in the subset to the center is calculated, the maximum distance is taken as the radius of the hypersphere, the data is divided into sub-nodes according to the distance, the process is repeated, and after 3 layers of recursion, the ball tree structure including the center, radius and sub-nodes is generated. Select the target high-dimensional data point, start traversing the ball tree from the root node, calculate the Euclidean distance between the data point and the hypersphere center of the current node, calculate the Euclidean distance between the data point and all data points in the leaf node after reaching the leaf node, record the 5 points with the smallest distance as the initial neighbor points, backtrack from the leaf node, calculate the minimum distance between the data point and the surface of each node, and expand the node when the minimum distance is less than the distance of the farthest neighbor point. Recalculate the distance between the data points in the node and the target high-dimensional data point, update the neighbor point set, and finally determine the high-dimensional region set containing 10 neighbor points.

[0064] In actual evaluation, 800 medium-dimensional data points are obtained, for each medium-dimensional data point, the Euclidean distance between one medium-dimensional data point and another data point is 0.2, the contribution of the Gaussian kernel function to the target point density is 0.328, the bandwidth parameter is obtained by Silverman rule through the variance of the medium-dimensional data as 0.134,

[0065] The density estimation value of a data point is calculated as 2.239, after traversing and calculating all data points, the density values are sorted in ascending order, the 75% quantile value is taken as the high-density threshold, the 25% quantile value is taken as the low-density threshold, and the proportion of data points higher than the high threshold is calculated as 0.22, because The reduction coefficient is taken as 0.956;

[0066] The preset neighborhood radius is 2 times the average of the Euclidean distance between each pair of data points. If the area density of the data points is higher than the high threshold, the radius is adjusted to 1.147, and the middle-dimensional neighborhood set is generated;

[0067] Take 500 low-dimensional data points, calculate the Euclidean distance between each pair of data points, obtain the distance set, calculate the interquartile range using the distance set, determine the distance threshold as 1.3, take a target numerical point as the center, and include the data points with a distance less than 1.3 in the low-dimensional neighborhood set, and finally obtain a low-dimensional neighborhood set containing 20 nearest neighbor points.

[0068] In actual evaluation, for high, middle and low dimensional neighborhood sets, respectively , the neighborhood of 5 points is linearly reconstructed to construct the error function , and the constraint is added to force the sum of the reconstruction weights of the neighborhood points to be 1, and the optimal reconstruction weight is solved by the Lagrange multiplier method , the optimal reconstruction weight vector of a high-dimensional data is [0.18, 0.22, 0.20, 0.19, 0.21], the target function is constructed, the target function is converted into Gram matrix form, the 1000-order Gram matrix is calculated, the eigenvalues and eigenvectors are solved, the high-dimensional data is mapped to 2-dimensional coordinates , the middle-dimensional data is mapped to , and the low-dimensional data is mapped to , and the high, middle and low dimensional data are spliced in order to form a 6-dimensional unified low-dimensional feature vector [0.32, -0.15, 0.08, 0.21, -0.09, 0.12], which provides a standardized input for subsequent fault diagnosis models.

[0069] Step C traverses the low-dimensional spatial dimension data, counts the value frequency of discrete data, obtains the probability density of continuous data, and obtains the marginal probability by combining the value frequency and the probability density.

[0070] It needs to be explained that actual data often mixes discrete and continuous features, and a single analysis method cannot completely describe the data distribution characteristics. By integrating the discrete frequency statistics to retain the category distribution information, calculating the continuous kernel density to capture the numerical distribution law, and then determining the weight based on the information entropy, the high-value information can be ensured to dominate the result when the discrete information entropy occupies a higher proportion, and the fused marginal probability can adapt to the characteristics of mixed data and more accurately depict the probability nature of the data, providing a more reliable probability basis for subsequent feature aggregation, fault diagnosis and other tasks.

[0071] In the actual evaluation, 1000 low-dimensional feature vector data of lithium battery coating machine is collected, and 580 normal, 220 scraper abnormal, 150 coating pressure abnormal, 50 motor speed abnormal are obtained by traversal statistics, the value frequency extraction is normal 0.58, scraper abnormal 0.22, coating pressure abnormal 0.15, motor speed abnormal 0.05;

[0072] The coating speed mapping value in the low-dimensional feature vector is continuous data, the dimension is 1, the total number of samples is 1000, the kernel density estimation hyperparameter is 0.3, the kernel function value is calculated by traversing all data points, the weighted sum of all kernel function values is obtained 620, the probability density estimation value is obtained, taking the target data point 3.2 as an example, the probability density estimation value is 2.275, the probability density estimation value of all data points is integrated to construct the probability density function, the probability density is obtained according to the probability density function, and the probability density is normalized to obtain ;

[0073] The information entropy of discrete data is 1.276, the information entropy of continuous data is calculated by integration 0.92, the information weight is obtained by using the information entropy of discrete data divided by the sum of information entropy of discrete data and continuous data 0.42, the value frequency of discrete data is multiplied by the information weight, the density probability of continuous data is multiplied by the inverse weight of information weight, and the marginal probability 0.2032 is obtained by adding the two results.

[0074] Step D uses the marginal probability to determine the weight of the dimensional data in data aggregation, and splices the dimensional data features after dimension reduction according to the weight to obtain the aggregated comprehensive features.

[0075] In the actual evaluation, the marginal probability of low-dimensional features of lithium battery coating machine is 0.35, the marginal probability of high-dimensional (video + text fusion features) is 0.28, the marginal probability of medium-dimensional (image + audio features) is 0.22, and the marginal probability of low-dimensional (numerical features) is 0.22, according to the marginal probability, the dimensional data in the low-dimensional space is divided into three levels, the high-probability layer selects 2-dimensional features of video texture feature value and text keyword probability value which are strongly related to coating quality abnormality , the medium-probability layer selects 2-dimensional features of image edge density and audio spectrum energy value which reflect the running state of the coating machine , and the low-probability layer selects the remaining marginal probability features of speed fluctuation coefficient , the high-probability layer features are reserved, the medium-probability layer features are multiplied with the high-probability layer features element by element to obtain , and the low-probability layer is obtained by screening through the gating formula , the high-dimensional, medium-dimensional and low-dimensional features after the hierarchical processing are spliced in a block manner in order of high, medium and low levels to obtain [0.75, 0.68, 0.4125, 0.3264, 0.1786], and after the splicing, probability normalization is performed on each level block to obtain the aggregated comprehensive features [0.5245, 0.4755, 0.5583, 0.4417, 0.1786], which provide high-value input for device state analysis and fault diagnosis model input.

[0076] In the embodiment, the method for collecting multi-modal data and pre-processing the multi-modal data to obtain dimensional data comprises:

[0077] The multi-modal data is collected, and the data of different modalities is cleaned and denoised, the mean value of the data of different modalities is calculated, the missing values are filled, and the quartile range method is used to correct the abnormal values.

[0078] According to the modalities of the multi-modal data, text data, numerical data, audio data, image data and video data are obtained, the word vector dimension is obtained through a word vector model for the text data, the numerical dimension is obtained according to the numerical columns of the numerical data, the frequency spectrum features and the time domain features of the audio data are extracted and merged into audio dimensional information, the basic dimension is determined through the resolution and the number of color channels for the image data, the video space-time dimensional information is obtained based on the image data of the key frames and the frame rate duration of the video data, the number of dimensions of each modality dimension information is obtained, all modality dimension numbers are sorted from high to low, the first 20% is divided into a high-dimensional interval, the middle 60% is a medium-dimensional interval, and the last 20% is a low-dimensional interval, and the modality data of different dimensions is put into the dimensional interval to obtain a dimensional data set.

[0079] The maximum difference, the mean value and the standard deviation of the different modality data in the dimensional data set are obtained, the scaling threshold is determined using the maximum difference, the mean value and the standard deviation, and the dimensional data is scaled based on the scaling threshold.

[0080] In the embodiment, the method for determining the scaling threshold using the maximum difference, the mean value and the standard deviation and scaling the corresponding modality data based on the scaling threshold comprises:

[0081] The maximum difference, the mean value and the standard deviation of the different modality data in the dimensional data set are obtained, the information entropy of the data feature description is calculated using the maximum difference, the mean value and the standard deviation, the dynamic threshold weight coefficient of the maximum difference, the mean value and the standard deviation is obtained after the information entropy is normalized, the data difference degree is calculated based on the historical threshold, the current data window mean value and the maximum difference, and the scaling threshold is obtained through the difference degree, and the scaling threshold calculation formula is:

[0082] ,

[0083] wherein is the current new input data point, n is the total number of current data points, is the cumulative mean of the previous n-1 data points, is the cumulative variance of the previous n-1 data points, M is the maximum value in the data window, m is the minimum value in the data window, last is the historical threshold value, , , is the maximum-minimum difference, the mean and the standard deviation of the threshold weight coefficient, is the decay coefficient;

[0084] The corresponding modal data is scaled based on the scaled threshold value.

[0085] In the embodiment, the method for locating the near neighbor points of the data points of the high-dimensional data by the ball tree algorithm comprises the following steps:

[0086] Based on the high-dimensional data set, the high-dimensional data points are regarded as space points, the data space is divided in a recursive manner, the mean of the current node data subset is calculated to determine the center point, the center point is taken as the center of the hypersphere, and the maximum distance of each point to the center point is taken as the radius, so as to generate a ball tree node containing the center, the radius and the subnode, and the division is repeated until the leaf node data amount threshold value or the recursive depth limit is met, and the ball tree construction is completed.

[0087] Based on the ball tree structure, the near neighbor search is performed on the target high-dimensional data point, the distance between the target point and the center of the hypersphere of the current node is calculated, if the distance exceeds the radius of the hypersphere, the branch is pruned, if the distance does not exceed the radius, the subnode is continuously searched, after reaching the leaf node, the distance between the target point and all data points in the leaf node is calculated, and the initial near neighbor points are recorded.

[0088] After the initial near neighbor points are obtained, the near neighbor set is optimized by backtracking the tree structure, the leaf node is backtracked along the tree, for each node passed, the minimum distance between the target point and the surface of the hypersphere of the node is calculated, if the distance is less than the distance of the farthest near neighbor point, the node is expanded, the distance between the data points in the node and the target point is recalculated, the near neighbor point set is updated, until the full tree search is completed, the near neighbor points of the high-dimensional data point are finally determined, and the high-dimensional field set is obtained.

[0089] In the embodiment, the method for dynamically adjusting the neighborhood range according to the distribution density of the middle-dimensional data, and the neighborhood of the low-dimensional data is calculated by the Euclidean distance, comprises the following steps:

[0090] Based on each data point of the middle-dimensional data set, the contribution of the data point to the target point density is calculated by using a Gaussian kernel function, the bandwidth parameter is obtained by using the Silverman rule for the standard deviation of the middle-dimensional data set, the density estimation value of each data point is obtained by dividing the sum of the contributions of all data points by the total number of data points and the bandwidth parameter, the density estimation values are arranged in ascending order, the 75th percentile value is taken as the high-density threshold value, and the 25th percentile value is taken as the low-density threshold value;

[0091] The proportion of data points higher than the high-density threshold value is calculated , if , the reduction coefficient is taken as , if , the default value 0.8 is taken; the proportion of data points lower than the low-density threshold value is calculated , if , the expansion coefficient is taken as , if , the default value 1.2 is taken;

[0092] The preset neighborhood radius is the average value of the Euclidean distance between each two data points in the middle-dimensional data set, if the density estimation value of the data point region is higher than the high-density threshold value, the preset neighborhood radius is reduced by multiplying the preset neighborhood radius by the reduction coefficient, if the density estimation value is lower than the low-density threshold value, the preset neighborhood radius is expanded by multiplying the preset neighborhood radius by the expansion coefficient, and the middle-dimensional neighborhood set is obtained by dynamically adjusting.

[0093] The Euclidean distance between any two data points in the low-dimensional data is calculated, the interquartile range is calculated based on the Euclidean distance, the distance threshold value is obtained by adding the lower quartile and 1.5 times the interquartile range, and the data points less than the distance threshold value are included in the low-dimensional neighborhood set.

[0094] In the embodiment, the method for mapping the dimensional data to the low-dimensional space based on the neighborhood by using the local linear embedding includes:

[0095] The local linear embedding is performed on the neighborhood sets of different dimensional data respectively, for each data point in the neighborhood set, the other data points in the neighborhood are used for linear reconstruction, a reconstruction error function is constructed, and a constraint condition is added to the error function, wherein is the reconstruction weight of the neighborhood point to the data point, and the optimal reconstruction weight is obtained by solving the error function;

[0096] The low-dimensional embedding target function of different neighborhood sets is constructed by using the optimal reconstruction weight, and the formula of the target function is:

[0097] ,

[0098] wherein Z is a low-dimensional embedding coordinate matrix, is the low-dimensional coordinate of the i th data point, M is the total number of data points, and K is the neighborhood size. is the optimal reconstruction weight of the neighborhood point j to the data point i, is the low-dimensional coordinate of the neighborhood point j to the data point i;

[0099] The target function is converted into a Gram matrix form, the eigenvalues and eigenvectors of the Gram matrix are calculated, the eigenvectors corresponding to the smallest 2-3 non-zero eigenvalues are retained, the mapping coordinates of the three-dimensional data set in the low-dimensional space are obtained, the dimensional data is mapped to the low-dimensional space, and the three types of low-dimensional features are spliced in the order from high to low to form a unified low-dimensional feature vector.

[0100] In the embodiment, the method for obtaining the marginal probability by combining the value frequency and the probability density comprises:

[0101] The discrete data of the low-dimensional feature vector in the low-dimensional space is traversed, the occurrence frequency of each discrete value is counted, and the value frequency is calculated;

[0102] The kernel function value is calculated by traversing all data points of the continuous data, and the probability density estimation value is obtained by weighted sum of all kernel function values, and the probability density estimation value calculation formula is:

[0103] ,

[0104] wherein is the probability density estimation value, N is the total number of sample data points, d is the dimension of the data, h is the key hyperparameter of the kernel density estimation, x is the target data point, is the nth sample object;

[0105] The probability density estimation value results corresponding to all data points are integrated to form the probability density function of the continuous data, the probability density is obtained according to the probability density function, and the probability density is normalized.

[0106] The information entropy of the discrete data and the continuous data is calculated, the information weight is obtained by using the information entropy of the discrete data divided by the total sum of the information entropy of the discrete data and the continuous data, the value frequency of the discrete data is multiplied by the information weight, the density probability of the continuous data is multiplied by the inverse weight of the information weight, and the marginal probability is obtained by adding the two results.

[0107] In the embodiment, the method for determining the weight of the dimensional data in data aggregation by using the marginal probability comprises:

[0108] According to the marginal probability, the dimensional data in the low-dimensional space is divided into three levels, the high-probability layer feature is retained, the medium-probability layer feature is multiplied by the high-probability layer feature element by element, and the low-probability layer is screened through a gating formula, and the formula is:

[0109] ,

[0110] wherein is a low-dimensional feature, is a low-dimensional marginal probability, is a high-dimensional marginal probability, is a medium-dimensional marginal probability, is a normalization factor, is a total sum of feature dimensions, is a low-dimensional feature after gating screening, and G is a gating signal;

[0111] The high-dimensional, medium-dimensional and low-dimensional features after the hierarchical processing are block spliced in order of high, medium and low levels, and the internal probability of each level block is normalized after splicing to obtain the aggregated comprehensive features.

[0112] The second aspect of the application also provides a data aggregation system based on multi-modal features, comprising:

[0113] a data acquisition module, configured to acquire multi-modal data, and to pre-process the multi-modal data to obtain dimensional data, wherein the dimensional data comprises high-dimensional, medium-dimensional and low-dimensional data;

[0114] a data conversion module, configured to locate the near neighbor points of data points of the high-dimensional data by using a ball tree algorithm, to dynamically adjust the neighborhood range according to the distribution density of the medium-dimensional data, to calculate the neighborhood of the low-dimensional data by using the Euclidean distance, and to map the dimensional data to a low-dimensional space by using local linear embedding based on the neighborhood;

[0115] a data processing module, configured to traverse the dimensional data in the low-dimensional space, to count the value frequency of discrete data, to obtain the probability density of continuous data, and to obtain the marginal probability by combining the value frequency and the probability density;

[0116] a data aggregation module, configured to determine the weight of the dimensional data in data aggregation by using the marginal probability, to splice the dimensional data features after dimension reduction according to the weight, and to obtain the aggregated comprehensive features.

[0117] The above content is merely an example and description of the structure of the application, and those skilled in the art can make various modifications or supplements or use similar ways to replace the described specific embodiments, as long as the modifications or supplements or replacements do not deviate from the structure of the application or exceed the scope defined by the claims, and should be within the protection scope of the application.

Claims

1. A data aggregation method based on multimodal features, characterized in that, Includes the following steps: Collect multimodal data, preprocess the multimodal data to obtain dimensional data, the dimensional data including high-dimensional, medium-dimensional and low-dimensional data; For high-dimensional data, the nearest neighbor of the data point is located by ball tree algorithm. The neighborhood range is dynamically adjusted according to the distribution density of medium-dimensional data. For low-dimensional data, the neighborhood is calculated by Euclidean distance. Based on the neighborhood, the dimensional data is mapped to the low-dimensional space by local linear embedding. The system iterates through the dimensional data in the low-dimensional space, counts the frequency of values ​​for discrete data, obtains the probability density of continuous data, and combines the frequency of values ​​and the probability density to obtain the marginal probability. Marginal probability is used to determine the weight of dimensional data in data aggregation. The features of the dimensional data after dimensionality reduction are then concatenated according to their weights to obtain the aggregated comprehensive features. The method for collecting multimodal data and preprocessing the multimodal data to obtain dimensional data includes: Collect multimodal data, clean and denoise the data of different modalities, calculate the mean of the data of different modalities to fill missing values, and use the interquartile range method to correct outliers. Based on the modality of the multimodal data, text data, numerical data, audio data, image data, and video data are obtained. For text data, word vector dimensions are obtained through a word vector model. Numerical dimensions are obtained based on the numerical columns of the numerical data. Spectral and temporal features of the audio data are extracted and merged into audio dimension information. For image data, the basic dimensions are determined by resolution and the number of color channels. Video spatiotemporal dimension information is obtained based on the image data and frame rate duration of keyframes in the video data. The number of dimensions for each modality is obtained. All modality dimensions are counted and sorted from high to low. The top 20% are divided into high-dimensional intervals, the middle 60% into medium-dimensional intervals, and the bottom 20% into low-dimensional intervals. Modality data of different dimensions are placed into dimension intervals to obtain a dimensional dataset. Obtain the difference between extreme values, mean, and standard deviation of different modal data in the dimensional dataset; use the difference between extreme values, mean, and standard deviation to determine the scaling threshold; and scale the dimensional data based on the scaling threshold. The method for determining a scaling threshold using the difference between extreme values, the mean, and the standard deviation, and scaling the corresponding modal data based on the scaling threshold, includes: Obtain the extreme value difference, mean, and standard deviation of different modal data in the dimensional dataset. Calculate the information entropy of the extreme value difference, mean, and standard deviation in relation to the data features. After normalizing the information entropy, obtain the dynamic threshold weight coefficients for the extreme value difference, mean, and standard deviation. Calculate the data variability based on historical thresholds, the current data window mean, and the extreme value difference. Obtain the scaling threshold through the variability. The scaling threshold calculation formula is as follows: in Let n be the number of newly input data points, and n be the total number of data points. This is the cumulative mean of the first n-1 data points. The cumulative variance of the first n-1 data points. The maximum value within the data window. The minimum value within the data window. For historical thresholds, The weighting coefficients for the maximum and minimum differences, mean, and standard deviation with respect to the threshold are: The attenuation coefficient; The corresponding modal data is scaled based on a scaling threshold; The method for mapping dimensional data to a low-dimensional space based on neighborhood through local linear embedding includes: Local linear embedding is performed on neighborhood sets based on data of different dimensions. For each data point in the neighborhood set, linear reconstruction is performed using other data points in its neighborhood, and a reconstruction error function is constructed. Constraints are added to the error function. in To determine the reconstruction weights of neighboring points to data points, solve the error function to obtain the optimal reconstruction weights. A low-dimensional embedding objective function is constructed using the optimal reconstruction weights for different neighborhood sets. The objective function formula is as follows: in For low-dimensional embedding coordinate matrices, These are the low-dimensional coordinates of the i-th data point. K is the total number of data points, and K is the size of the neighborhood. The optimal reconstruction weight of neighboring point j to data point i. The low-dimensional coordinates of neighboring point j relative to data point i; The objective function is transformed into Gram matrix form. The eigenvalues ​​and eigenvectors of the Gram matrix are calculated. The eigenvectors corresponding to the 2-3 smallest non-zero eigenvalues ​​are retained to obtain the mapping coordinates of the three-dimensional datasets in the low-dimensional space. The dimensional data is mapped to the low-dimensional space. The three types of low-dimensional features are concatenated in descending order to form a unified low-dimensional feature vector. The method for obtaining marginal probability by combining value frequency and probability density includes: Traverse the discrete data of low-dimensional feature vectors in low-dimensional space, count the occurrence frequency of each discrete value, and calculate its frequency. For continuous data, calculate the kernel function value by iterating through all data points, and then sum the weighted kernel function values ​​to obtain the probability density estimate. The formula for calculating the probability density estimate is as follows: in This is the probability density estimate. The total number of sample data points. For the dimensions of the data, The key hyperparameter for kernel density estimation, For the target data point, For the first One sample object; The probability density estimates of all data points are integrated to form the probability density function of continuous data. The probability density is obtained from the probability density function and then normalized. Calculate the information entropy of discrete data and continuous data. Divide the information entropy of discrete data by the sum of the information entropy of discrete data and continuous data to obtain the information weight. Multiply the information weight by the frequency of the discrete data value. Multiply the density probability of continuous data by the inverse weight of the information weight. Add the two results to obtain the marginal probability.

2. The data aggregation method based on multimodal features according to claim 1, characterized in that, The method for locating nearest neighbors of data points in high-dimensional data using a ball tree algorithm includes: Based on the high-dimensional dataset, the high-dimensional data points are regarded as spatial points. The data space is divided in a recursive manner. The mean of the data subset of the current node is calculated to determine the center point. The center point is used as the center of the hypersphere. The maximum distance from each point to the center point is used as the radius. A ball tree node containing the center, radius and child nodes is generated. The division is repeated until the threshold of leaf node data volume or the recursion depth limit is met to complete the construction of the ball tree. Based on the ball tree structure, a nearest neighbor search is performed on the target high-dimensional data point. Starting from the root node of the ball tree, the traverse is performed to calculate the distance between the target point and the center of the hypersphere of the current node. If the distance exceeds the radius of the hypersphere, the branch is pruned; if it does not exceed the radius, the search continues to the child nodes. After reaching the leaf node, the distance between the target point and all data points in the leaf node is calculated, and the initial nearest neighbor point is recorded. After obtaining the initial nearest neighbor points, the nearest neighbor set is optimized by backtracking the tree structure. Starting from the leaf nodes, backtracking along the tree, for each node visited, the minimum distance between the target point and the hypersphere surface of the node is calculated. If this distance is less than the distance of the current farthest nearest neighbor, the node is expanded, the distance between the data points in the node and the target point is recalculated, and the nearest neighbor set is updated until the full tree search is completed. Finally, the nearest neighbors of the high-dimensional data points are determined, and the high-dimensional neighborhood set is obtained.

3. The data aggregation method based on multimodal features according to claim 1, characterized in that, The method of dynamically adjusting the neighborhood range based on the distribution density of mid-dimensional data, and calculating the neighborhood using Euclidean distance for low-dimensional data, includes: Based on each data point in the middimensional dataset, its contribution to the density of the target point is calculated using the Gaussian kernel function. The bandwidth parameter is obtained by applying the Silverman rule to the standard deviation of the middimensional dataset. The sum of the contributions of all data points is divided by the total number of data points and the bandwidth parameter to obtain the density estimate of each data point. The density estimates are sorted in ascending order, and the 75th percentile value is taken as the high-order density threshold, and the 25th percentile value is taken as the low-order density threshold. Calculate the proportion of data points above the high bit density threshold. ,like The reduction factor is taken as ,like , Take the default value of 0.8; Calculate the proportion of data points below the low bit density threshold. ,like The expansion factor is taken as ,like The default value is 1.

2. The preset neighborhood radius is the average of the pairwise Euclidean distances between two-dimensional data points. If the estimated density of the data point region is higher than the high-dimensional density threshold, the preset neighborhood radius is multiplied by a shrinkage factor to reduce the range. If it is lower than the low-dimensional density threshold, the preset neighborhood radius is multiplied by an expansion factor to expand the range. After dynamic adjustment, a mid-dimensional neighborhood set is obtained. Calculate the Euclidean distance between any two data points in the low-dimensional data, calculate the interquartile range based on the Euclidean distance, add the lower quartile to 1.5 times the interquartile range to obtain the distance threshold, and include data points smaller than the distance threshold into the low-dimensional neighborhood set with the target point as the center.

4. The data aggregation method based on multimodal features according to claim 1, characterized in that, The method for determining the weights of dimensional data in data aggregation using marginal probability, and then concatenating the dimensionality-reduced dimensional data features according to their weights to obtain the aggregated comprehensive features includes: Based on marginal probabilities, the dimensional data in the low-dimensional space is divided into three levels. Features from the high-probability level are retained, while features from the medium-probability level are multiplied element-wise with those from the high-probability level. The low-probability level is then filtered using a gating formula, which is: in Low-dimensional features For low-dimensional marginal probabilities, For high-dimensional marginal probabilities, For the mid-dimensional marginal probability, As the normalization factor, The sum of feature dimensions, These are the low-dimensional features after gating. This is a gating signal; The high-dimensional, medium-dimensional, and low-dimensional features, after being processed by hierarchical layering, are pieced together in blocks. After the blocks are pieced together, the probability within each block is normalized to obtain the aggregated comprehensive features.

5. A data aggregation system based on multimodal features, used to execute the data aggregation method based on multimodal features according to any one of claims 1 to 4, characterized in that, The system includes: The data acquisition module is used to acquire multimodal data and preprocess the multimodal data to obtain dimensional data, which includes high-dimensional, medium-dimensional and low-dimensional data. The data transformation module is used to locate the nearest neighbor of data points in high-dimensional data using the ball tree algorithm, dynamically adjust the neighborhood range according to the distribution density of mid-dimensional data, calculate the neighborhood for low-dimensional data using Euclidean distance, and map the dimensional data to the low-dimensional space based on the neighborhood using local linear embedding. The data processing module is used to traverse dimensional data based on low-dimensional space, count the frequency of values ​​of discrete data, obtain the probability density of continuous data, and combine the value frequency and probability density to obtain the marginal probability. The data aggregation module is used to determine the weight of dimensional data in data aggregation using marginal probability, and to concatenate the features of the dimensionality-reduced dimensional data according to the weights to obtain the aggregated comprehensive features.

Citation Information

Patent Citations

  • Relation cluster database optimization method based on multi-modal learning

    CN119046316A

  • Retrieval result determination method and device, electronic equipment and storage medium

    CN119829793A