File processing method adaptive to structured and non-structured data
By identifying and classifying multiple data types, preprocessing and feature extraction, and using multimodal mutual information optimization and variational inference technology to dynamically adjust the fusion weight, the heterogeneity and uncertainty problems of multimodal data processing in traditional methods are solved, and efficient and robust multimodal data fusion is achieved.
Patent Information
- Application Number
- CN202510502855.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-04-22
AI Technical Summary
Traditional data processing methods are difficult to operate multimodal data uniformly, resulting in inconsistent information volume and difficult to quantify correlation between modals. The accuracy of weight allocation and feature expression during the fusion process is insufficient, which affects the credibility and robustness of the results.
By identifying and classifying structured, semi-structured and unstructured data, preprocessing and feature extraction are performed, the correlation between quantitative data modes is optimized based on multimodal mutual information, the fusion weight is dynamically adjusted, and the uncertainty of modeled data is used to infer variational in the modeled data.
It realizes efficient fusion of multimodal data, significantly improves the accuracy and robustness of data processing, solves the problem of insufficient correlation quantization and fusion robustness between modals, and provides a multimodal data fusion solution in complex application scenarios.
Smart Images

Figure CN120030193A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of file processing, and in particular to a file processing method adapted to structured and unstructured data. Background Art
[0002] With the rapid development of information technology, structured data and unstructured data are increasingly used in various fields. Structured data usually exists in the form of tables and databases, which have clear row-column relationships and data structures; while unstructured data includes images, text, audio, video and other forms, which lack a unified structured format. At the same time, semi-structured data is between the two and stores information in a hierarchical structure. In response to the processing needs of different types of data, traditional methods often face the following technical bottlenecks.
[0003] Structured, semi-structured and unstructured data differ significantly in storage format, feature expression and distribution form, which makes it difficult for traditional data processing tools to operate in a unified manner when processing multimodal data. There are problems such as inconsistency in the amount of information between multimodal data and difficulty in quantifying the correlation between modes, which makes the weight allocation and feature expression in the data fusion process inaccurate. Unstructured data such as images and text may contain noise or missing values, resulting in information loss or error accumulation in the multimodal fusion process, thereby reducing the credibility and robustness of the results. Traditional data processing methods are mostly targeted at a single data type or a specific field, lack universality, and have limited capabilities in dynamic weighted fusion and uncertainty modeling, making it difficult to meet the needs of complex application scenarios. Summary of the invention
[0004] In view of the deficiencies in the prior art, the present invention provides a file processing method adapted to structured and unstructured data, which solves the problems of heterogeneity, uncertainty and insufficient robustness of quantification and fusion of inter-modal correlation in multimodal data processing.
[0005] To achieve the above objectives, the present invention is implemented through the following technical solutions: A file processing method adapted to structured and unstructured data includes the following steps: S1. Identify the types of input files and classify them, including structured data, semi-structured data, and unstructured data; S2, preprocessing and feature extraction for different types of data; S3, quantify the correlation between different data modalities based on multimodal mutual information optimization; S4, verify the interaction consistency between data modalities and dynamically adjust the fusion weight; S5. Model and optimize data uncertainty using variational inference; S6. Output the processed unified data results and conduct performance evaluation.
[0006] Preferably, the step S1 specifically includes the following steps: S1.1. Preliminary identification of file types through file extensions and file header information; S1.2. Further analyze the file content and classify the table structure data containing clear row and column relationships as structured data; S1.3. For data files without a clear structure, identify their type through the file content pattern. If it is an image file, extract the grayscale distribution of the pixel matrix; S1.4. Record the recognition and classification results in a preset data type classification table, where structured data includes tabular data, semi-structured data includes hierarchical files, and unstructured data includes image data, audio data, video data, and natural language text data.
[0007] Preferably, the step S2 of preprocessing and extracting features of structured data includes: ; in, For time point The interpolation of and is the time point of the adjacent known value; (b) Identification and elimination of outliers by calculating the Z score ,like The data is considered abnormal. is the mean, is the standard deviation, is the threshold value.
[0008] Preferably, the step S2 of preprocessing and feature extraction of unstructured data includes: For image data, normalization and noise reduction are performed, and feature vectors are extracted using convolutional neural network (CNN). For text data, word segmentation, word frequency statistics and embedding models are used to convert text into vector representation. The formula for word frequency statistics is: ; Among them, TF Words In the documentation The word frequency in is the total number of documents, To contain words The number of documents, (c) For audio data, short-time Fourier transform is used to extract spectrum features. The formula is: ; in, For time and frequency The time-frequency characteristics of is a window function.
[0009] Preferably, the method based on multimodal mutual information optimization in step S3 includes: estimating the mutual information by optimizing the following variational lower bound ; in, is the conditional distribution, is a marginal distribution, which is parameterized using a neural network.
[0010] Preferably, the method for performing consistency verification in step S4 includes: calculating the interaction consistency score between modalities; ; in, Representing modality and The mutual information of and The modal and Information entropy.
[0011] Preferably, the dynamic weighting method based on consistency score in step S4 includes: Calculate the weight of each mode , ; The fusion result is: ; in, For the data of each modality, is the corresponding dynamic weighting coefficient.
[0012] Preferably, the uncertainty modeling based on variational inference in step S5 includes: The following evidence lower bound (ELBO) is used to determine the latent variables The posterior distribution of Make an approximation; ; in, is the variational distribution and KL is the KL divergence.
[0013] Preferably, the method based on uncertainty quantification in step S5 includes: ; Uncertainty values are used to weight data points with high noise or low confidence.
[0014] Preferably, the performance verification in step S6 includes: Mutual information gain evaluation, verifying the improvement of mutual information before and after data fusion; Perceptual performance verification, the accuracy of the output results is evaluated based on the IoU and mAP indicators.
[0015] The present invention provides a file processing method adapted to structured and unstructured data, which has the following beneficial effects: 1. The present invention achieves efficient fusion of multimodal data by classifying and extracting features from structured data, semi-structured data and unstructured data, and utilizing mutual information optimization and dynamic weighting strategies. Compared with traditional methods, the present invention can significantly improve the accuracy and robustness of data processing, and provides an efficient solution for multimodal data fusion in complex scenarios.
[0016] 2. The present invention introduces a mutual information optimization method based on a variational lower bound, and accurately quantifies the correlation between modes. Compared with the existing correlation analysis technology, mutual information optimization can better capture the complex correlation between modes and provide a solid theoretical basis for dynamic weighted fusion.
[0017] 3. The present invention adopts a dynamic weighting strategy to dynamically adjust the modal weights according to the interaction consistency scores between the modalities, so that the high-correlation modality contributes more to the fusion result, while the influence of the low-correlation modality is weakened. This method effectively solves the robustness problem caused by fixed weights in traditional methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is a stereogram of the present invention. DETAILED DESCRIPTION
[0019] The following will be combined with the drawings in the specification of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0020] Please see attached Figure 1 , an embodiment of the present invention provides a file processing method adapted to structured and unstructured data, comprising the following steps; S1. Identify the types of input files and classify them, including structured data, semi-structured data, and unstructured data; S2, preprocessing and feature extraction for different types of data; S3, quantify the correlation between different data modalities based on multimodal mutual information optimization; S4, verify the interaction consistency between data modalities and dynamically adjust the fusion weight; S5. Model and optimize data uncertainty using variational inference; S6. Output the processed unified data results and conduct performance evaluation; Specific; In this embodiment, the input file is first identified and classified by its type, specifically in the following steps; As an option, this embodiment performs a preliminary judgment based on the file extension and file header information. For some files with non-standard file extensions, a supplementary judgment can be further performed based on the file header information. Specifically, for files whose extensions have been determined, the file contents need to be further parsed. If the file has a clear row-column relationship, its row-column data structure can be extracted through a table parsing tool and classified as structured data. For semi-structured data, this embodiment classifies by parsing the hierarchical structure of the file. Files usually store information in the form of nested key-value pairs, and files use nested tags to represent hierarchical content. As an implementation method, these hierarchical structures can be traversed through the depth-first search algorithm (DFS) to record their nesting relationships and hierarchical depths; It should be emphasized that, for the complexity of hierarchical data, this embodiment further uses information entropy to quantify it. The specific formula is as follows: ; in, represents the entropy of hierarchical information, For Node The distribution probability of Representation Node The frequency of occurrence in the hierarchical structure. Information entropy can reflect the complexity of the data hierarchy, thus providing a reference for classification; For data files without a clear structure, such as pictures and audio, this embodiment classifies them through content pattern analysis. In an exemplary implementation, the grayscale histogram can be used to analyze the pixel distribution characteristics of image files, and the edge detection algorithm can be combined to further verify that they are image data. Text files can be analyzed through word segmentation and keyword extraction methods to determine their text attributes.
[0021] It should be noted that, in order to ensure the accuracy of classification, this embodiment records all recognition results in a unified data type classification table, which exemplarily includes the following: Structured data: such as table files and relational database export data; Semi-structured data: such as JSON files and XML files; Unstructured data: such as image data, audio files, video files, and free text.
[0022] In this embodiment, for the input multi-type files, their types are identified and classified through multi-level rules and algorithms. Specifically, the following contents are included: In this embodiment, step S2 mainly performs targeted preprocessing and feature extraction on structured data, semi-structured data, and unstructured data to achieve a unified representation of multimodal data. It should be noted that this step is intended to ensure that different modal data can be input into the subsequent fusion stage in the form of standardized high-dimensional features after specific processing, providing a basis for the final dynamic weighting and uncertainty modeling.
[0023] Specifically, structured data is normalized through outlier detection, interpolation and statistical feature extraction; semi-structured data is optimized through hierarchical structure analysis and information entropy calculation; unstructured data is combined with deep learning and signal processing methods to extract efficient feature representation; In this embodiment, the processing of structured data mainly involves data cleaning, outlier processing and statistical feature extraction; This embodiment first processes the missing values in the structured data. Specifically, the missing values can be filled using the mean filling method, interpolation method, or regression prediction method. For example, for time series data, the interpolation method calculates the value of the missing point using the following formula: ; in, For time point The interpolated value of and are the time points of adjacent known values; For outliers in the data, this embodiment uses the Z-score method to detect and remove them. The specific calculation formula is as follows: ; in, For data points The standardized score of and are the mean and standard deviation of the data respectively. (generally Take 3), then the data point can be judged as an outlier; It should be noted that, for the cleaned data, this embodiment further extracts its statistical features. Exemplarily, these features include mean, variance, skewness, kurtosis and correlation matrix, which are used to characterize the statistical characteristics of the data and the relationship between variables; Preprocessing and feature extraction of semi-structured data In this embodiment, for semi-structured data, feature representation is mainly achieved through hierarchical structure analysis and information entropy calculation.
[0024] Specifically, this embodiment uses a depth-first search algorithm to recursively parse the key-value pair nesting relationship of the JSON file and record the structural information of each layer of nodes. In a possible implementation, the tag nesting of the XML file can also be parsed into a tree structure by a similar method; In order to quantify the complexity of the hierarchical structure, this embodiment calculates the hierarchical information entropy , the formula is as follows; ; in, is the hierarchical information entropy, For the The type or label of the layer node, is the probability of occurrence of the node. It should be noted that the larger the hierarchical information entropy value is, the more complex the data structure is; This embodiment can also express the parsed hierarchical structure as a high-dimensional tensor for subsequent fusion processing; Preprocessing and feature extraction of unstructured data In this embodiment, image processing technology, natural language processing method and signal processing technology are used to extract features from unstructured data.
[0025] For image data, this embodiment uses preprocessing operations such as image normalization and noise reduction. As an option, a high-dimensional feature vector of the image can be extracted through a convolutional neural network.
[0026] For text data, this embodiment adopts word segmentation, stop word filtering and embedding model conversion methods. For example, the importance of each word can be counted using TF-IDF (term frequency-inverse document frequency), and the calculation formula is as follows: ; Among them, TF Words In the documentation The frequency of occurrence in is the total number of documents, To contain words The number of documents; For audio data, this embodiment uses short-time Fourier transform (STFT) to extract spectrum features. The specific formula is: ; in, For time and frequency The time-frequency characteristics of is a window function; The unstructured data after the above preprocessing are converted into standardized high-dimensional feature representations for the subsequent data fusion stage; The preprocessing of different types of data may need to be adjusted according to data distribution or business needs. For example: The cleaning strategy for structured data can be combined with specific business rules, such as fluctuation range constraints for financial data; Hierarchical information entropy calculation of semi-structured data can introduce label importance weights to highlight key fields; Feature extraction of unstructured data can select different embedding models or signal processing methods according to the application scenario; It should be noted that the above preprocessing and feature extraction process provides standardized input for subsequent data fusion, consistency verification and dynamic weighting, ensuring the robustness and consistency of multimodal data processing. Step S3 aims to quantify the correlation of different data modes through multimodal mutual information optimization method. It should be noted that multimodal data is usually difficult to directly fuse due to differences in modal characteristics, distribution forms and feature dimensions. This embodiment accurately measures the correlation of different modal data by calculating the mutual information between modalities, providing a theoretical basis for subsequent consistency verification and dynamic weighted fusion; In this embodiment, mutual information is used to describe the modal data and degree of association; In this embodiment, the joint distribution and marginal distribution , The mutual information is calculated by the logarithmic ratio of the two, and its mathematical expression is: ; It should be noted that in this formula is modal and The joint probability distribution of and The modal and The marginal probability distribution of , in a possible implementation, for high-dimensional feature data, it is difficult to directly calculate the above integral, therefore, this embodiment uses a variational inference method to optimize and solve the mutual information; In the present invention, by constructing an optimizable distribution ,right Approximate modeling is performed to optimize mutual information with the help of the lower bound. Specifically, the mathematical expression of the variational lower bound is: ; in: is the conditional distribution, used to approximate ; is the marginal distribution, given by and Derived are the parameters of the distribution parameterization model; It should be noted that the above lower bound optimization problem is solved by the gradient descent method, and the optimization objective is; ; In a possible implementation, this embodiment uses a deep neural network to and Perform parametric modeling and transform the input modal features into and Project to the joint space and learn the characteristics of its joint distribution; Normalized consistency score In this embodiment, in order to further measure the correlation between modal data In this embodiment, in order to further measure the correlation between modal data, the present invention calculates the consistency score between modalities. , the formula is as follows; ; It should be noted that in this formula; Representing modality Information entropy of Representing modality Information entropy of Used to normalize the mutual information value to ensure that the result is between 0 and 1. By calculating the consistency score, this embodiment can quantify the strength of the correlation between different modalities and provide a weight basis for subsequent dynamic weighted fusion; Multimodal Correlation Matrix In a possible implementation, this embodiment stores the mutual information values and consistency scores between all modalities as a multimodal correlation matrix, which is exemplarily represented as: ; in Representing modality and The correlation between them.
[0027] It should be noted that the generation of this matrix can provide a direct reference for weight allocation and consistency adjustment in the multimodal fusion stage.
[0028] In this embodiment, step S4 verifies the interactive consistency of different data modalities and quantifies the synergistic relationship between modalities, thereby providing a dynamic weighted basis for multimodal data fusion. It should be noted that the interactive consistency between data modalities is an important guarantee for the accuracy and robustness of multimodal fusion. This embodiment realizes the adaptive adjustment of the fusion strategy through the calculation of consistency scores and dynamic weight allocation; In this embodiment, in order to quantify the interactive consistency between modal data, a normalized consistency score calculation method based on mutual information is proposed; As an option, this embodiment uses the modal mutual information value calculated in step S3 , combined with modal information entropy and , calculate the consistency score The formula is: ; It should be noted that: Representing modality and The mutual information between and The modal and Information entropy of Normalization factor To convert the consistency score Restricted to the range [0,1]; Specifically, the closer the consistency score is to 1, the higher the degree of interaction between the modalities and the stronger the data consistency; when the score is close to 0, it means that there is a large deviation or noise influence between the modalities; Calculation of dynamic weights In this embodiment, in order to achieve dynamic adjustment of fusion weights, based on the consistency score Assign weights to each modality , the weight calculation formula is: ; in the formula; Indicates Data fusion weights of each modality; Representing modality and The consistency score of Denominator Used to normalize the weights so that the sum of the weights is 1; This embodiment dynamically adjusts the weights to make the high-consistency modality contribute more to the fusion result, while the influence of the low-consistency modality is weakened; Dynamic weighted calculation of fusion results In this embodiment, the fusion result By modal data and weight The weighted sum of is calculated, and its mathematical expression is: ; Specifically; Represents the fusion result of multimodal data; Indicates Data features of each modality; Represents the dynamic weight of the corresponding mode; The dynamic weighting formula can adaptively adjust the contribution of each modality to the fusion result according to the interaction consistency score between the modalities, thereby improving the robustness and accuracy of the fusion result; Storage of consistency matrix In one possible implementation, this embodiment stores the consistency scores between all modalities as a consistency matrix , its expression is; ; in Representing modality and It should be noted that the consistency matrix provides a global reference for subsequent modal screening and dynamic weighting. The modal with the highest correlation is selected for fusion according to the consistency score in the matrix, and the low-correlation modal is ignored. In this embodiment, step S5 uses the variational inference method to model and optimize the uncertainty of multimodal data, thereby improving the credibility of the fusion result. It should be noted that in the process of multimodal data processing, noise, missing values and differences between modalities may cause data uncertainty. If they are not modeled and optimized, they may have a negative impact on the reliability of the final result. The probability distribution model of the fusion result is used to quantify and optimize the uncertainty of the fusion result to further enhance the robustness of data processing; Modeling of latent variables In this embodiment, in order to describe the uncertainty of data, a hidden variable is introduced , and assuming It is multimodal data The potential representation of, its joint distribution can be expressed as; ; in the formula; Represents hidden variables The prior distribution of is usually assumed to be a standard normal distribution; Representing modality Given a hidden variable Likelihood distribution under conditions surface Shows the collection of all modal data; This embodiment uses a deep neural network to Perform parametric modeling to ensure it can adapt to complex modal distributions; Optimizing Variational Inference In this embodiment, in order to Approximate and introduce variational distribution , and achieve parameter learning by optimizing the evidence lower bound (ELBO). The specific optimization objectives are; ; It should be noted that; Item 1 is the log-likelihood term, indicating that the data Given a hidden variable The probability of generation under the conditions; Item 2 is the KL divergence, which is used to measure the variational distribution With the prior distribution The difference between As an option, this embodiment uses the reparameterization technique to For example, by It is expressed as; ; middle, and Hidden variables The mean and standard deviation parameters of Solve it through the gradient descent algorithm to learn The best parameters for Uncertainty Quantification In this embodiment, through the variational distribution The variance of quantifies the uncertainty of the fusion result. The specific formula is: ; It should be noted that: The latent variable representation representing the fusion result; Var is the variance of the variational distribution, which is used to measure the uncertainty of data fusion. When the uncertainty is large, this embodiment can reduce the weight of the corresponding data, thereby reducing its negative impact on the fusion result; This embodiment further optimizes the uncertainty quantification result, and the specific steps include: Adjust the weights of high uncertainty data points to reduce their contribution in the fusion process; Increase the weight of low-uncertainty data points to enhance the influence of high-confidence data; Specifically, the optimized weight calculation formula is: ; Among them, Uncertainty Representing modality The uncertainty quantification results of ,it should be explained that this optimization strategy can significantly improve the robustness of the fusion results, especially when the data quality is uneven; In this embodiment, step S6 outputs the processed multimodal fusion data and verifies the validity of the fusion result through a series of performance evaluation indicators. It should be noted that the unified data result not only includes the feature representation after fusion, but also has a description of uncertainty quantification, so as to provide reliable decision support in subsequent application scenarios. This embodiment ensures the integrity and practical applicability of the data processing solution by verifying the multi-dimensional performance of the fusion result. Output of data results In this embodiment, the processed unified data result The final feature representation including multimodal data fusion is calculated by the dynamic weighted fusion method in this embodiment , the formula is; ; in: Indicates Data features of each modality; Representing modality Dynamic weight of It should be noted that Represents the fused high-dimensional feature vector, which can be used for subsequent classification, regression or clustering tasks. As a possible implementation method, this embodiment further Output together with the uncertainty quantification results of each mode to enhance the interpretation of the data results. The uncertainty quantification results can be calculated by the following formula; ; in, Var Represents the variance of the latent variable distribution of the fusion result; Mutual Information Gain Evaluation In this embodiment, in order to verify the change in the amount of information before and after data fusion, the mutual information gain is calculated. Specifically, the mutual information gain is defined as: ; in: Represents the fusion result and the target variable The mutual information between Representing modality With the target variable The mutual information between It should be noted that a positive value of mutual information gain indicates that data fusion improves the amount of information, while a negative value may indicate that the fusion weights of some modalities need to be further optimized; In a possible implementation, this embodiment minimizes the following objective function to weight Make dynamic adjustments; ; The above optimization process can further improve the information validity of the fusion results; Perceptual performance verification In this embodiment, in order to further verify the practicality of the data processing solution, the perception performance evaluation index is used to measure the accuracy of the fusion result. This embodiment is verified based on the following common indicators: IoU (Intersection over Union): Applicable to target detection or segmentation tasks, the formula is; ; Among them, the numerator is the overlapping area of the predicted area and the true area, and the denominator is the union area of the two; mAP (Mean Average Precision): Applicable to multi-classification tasks, indicating the average precision score of all categories, the formula is; ; in, For Category The average precision of is the total number of categories; This embodiment comprehensively measures the performance of the fusion results in specific tasks through the calculation of the above indicators, thereby verifying the actual effect of the data processing solution; Summarize By classifying, preprocessing, extracting features, optimizing mutual information, modeling uncertainty, and dynamically merging multimodal data, the heterogeneity, uncertainty, and lack of consistency in multimodal data processing are solved. By introducing mutual information optimization and variational inference modeling, the present invention effectively improves the accuracy of correlation measurement between modalities, and optimizes the reliability of data fusion results through dynamic weighting strategies.
[0029] The technical effect of the present invention is that it can achieve efficient fusion and optimization of multi-type data, so that the data processing results can reach a high level in terms of information volume, robustness and practical application effect. In addition, through uncertainty quantification and performance evaluation, the present invention further ensures the credibility of the processing results and provides reliable technical support for a wide range of multimodal data application scenarios.
[0030] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A file processing method adapted to structured and unstructured data, characterized in that: The steps include: S1. Identify the types of input files and classify them, including structured data, semi-structured data, and unstructured data; S2, preprocessing and feature extraction for different types of data; S3, quantify the correlation between different data modalities based on multimodal mutual information optimization; S4, verify the interaction consistency between data modalities and dynamically adjust the fusion weight; S5. Model and optimize data uncertainty using variational inference; S6. Output the processed unified data results and conduct performance evaluation.
2. The file processing method for adapting structured and unstructured data according to claim 1, characterized in that: The step S1 specifically includes the following steps: S1.
1. Preliminary identification of file types through file extensions and file header information; S1.
2. Further analyze the file content and classify the table structure data containing clear row and column relationships as structured data; S1.
3. For data files without a clear structure, identify their type through the file content pattern. If it is an image file, extract the grayscale distribution of the pixel matrix; S1.
4. Record the recognition and classification results in a preset data type classification table, where structured data includes tabular data, semi-structured data includes hierarchical files, and unstructured data includes image data, audio data, video data, and natural language text data.
3. The file processing method for adapting structured and unstructured data according to claim 1, characterized in that: The step S2 of preprocessing and feature extraction of structured data includes: ; in, For time point The interpolated value of and is the time point of the adjacent known value; (b) Identification and elimination of outliers by calculating the Z score ,like The data is considered abnormal. is the mean, is the standard deviation, is the threshold value.
4. The file processing method for adapting structured and unstructured data according to claim 1, characterized in that: The step S2 of preprocessing and feature extraction of unstructured data includes: For image data, normalization and noise reduction are performed, and feature vectors are extracted using convolutional neural network (CNN). For text data, word segmentation, word frequency statistics and embedding models are used to convert text into vector representation. The formula for word frequency statistics is: ; Among them, TF Words In the documentation The word frequency in is the total number of documents, To contain words The number of documents, (c) For audio data, short-time Fourier transform is used to extract spectrum features. The formula is: ; in, For time and frequency The time-frequency characteristics of is a window function.
5. The file processing method for adapting structured and unstructured data according to claim 1, characterized in that: The method based on multimodal mutual information optimization in step S3 includes: estimating the mutual information by optimizing the following variational lower bound: ; in, is the conditional distribution, is a marginal distribution, which is parameterized using a neural network.
6. The file processing method for adapting structured and unstructured data according to claim 1, characterized in that: The method for performing consistency verification in step S4 includes: calculating the interaction consistency score between modalities; ; in, Representing modality and The mutual information of and The modal and Information entropy.
7. The file processing method for adapting structured and unstructured data according to claim 1, characterized in that: The dynamic weighting method based on consistency score in step S4 includes: Calculate the weight of each mode , ; The fusion result is: ; in, For the data of each modality, is the corresponding dynamic weighting coefficient.
8. The file processing method for adapting structured and unstructured data according to claim 1, characterized in that: The uncertainty modeling based on variational inference in step S5 includes: The following evidence lowers the ELBO bound on the latent variable The posterior distribution of Make an approximation; ; in, is the variational distribution and KL is the KL divergence.
9. The file processing method for adapting structured and unstructured data according to claim 1, characterized in that: The method based on uncertainty quantification in step S5 includes: ; Uncertainty values are used to weight data points with high noise or low confidence.
10. The file processing method for adapting structured and unstructured data according to claim 1, characterized in that: The performance verification in step S6 includes: Mutual information gain evaluation, verifying the improvement of mutual information before and after data fusion; Perceptual performance verification, the accuracy of the output results is evaluated based on the IoU and mAP indicators.
Citation Information
Patent Citations
Data processing method and device, equipment and storage medium
CN116432140A
Multi-modal knowledge graph method based on power grid dispatching
CN117171358A
Multi-modal standardized knowledge graph automatic generation method and system
CN119150971A
Monitoring method and system for data governance process
CN119202545A
High-precision surveying and mapping method based on multi-source data fusion
CN119203034A