Multimodal learning data fusion analysis method and system, and intelligent device
By identifying the application scenarios and data formats of multimodal data and selecting the optimal timing strategy and algorithm, the accuracy problem of semantic alignment algorithms in multimodal data fusion is solved, thereby improving the accuracy and efficiency of multimodal data fusion.
Patent Information
- Application Number
- CN202511431183.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-09
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2045-10-09
AI Technical Summary
Existing multimodal data fusion processes cannot accurately analyze multimodal data semantic alignment algorithms, which reduces the accuracy and quality of multimodal learning data fusion.
By collecting multimodal feature data, identifying application scenario types and data formats, analyzing semantic alignment algorithm features, selecting the optimal data fusion timing strategy, and matching the optimal fusion algorithm, accurate fusion of multimodal data can be achieved.
It improves the accuracy and quality of multimodal data fusion, achieves precise matching between scientific identification application scenarios and data formats, and enhances the efficiency and reliability of data fusion.
Smart Images

Figure CN120930069B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, specifically to a method, system, and intelligent device for multimodal learning data fusion and analysis. Background Technology
[0002] Multimodal data fusion refers to the integration and analysis of heterogeneous information from different sensors or data sources to extract more comprehensive and accurate decision-making information than a single modality. Data sources include images, text, audio, and video. The core of multimodal data fusion lies in addressing the differences, complementarities, and redundancies between modalities, and it is mainly divided into three fusion levels: data-level, feature-level, and decision-level. 1. Data-level fusion involves directly aligning and stitching the raw data at the original data level, such as aligning and fusing pixels of infrared and visible light images. 2. Feature-level fusion involves extracting high-order features from each modality and then fusing them, such as combining word vectors from text with CNN features from images. 3. Decision-level fusion involves summarizing the results after independent analysis of each modality, such as voting mechanisms or Bayesian inference. Existing multimodal data fusion processes cannot accurately analyze multimodal data semantic alignment algorithms based on the application scenarios and forms of multimodal data fusion, nor can they scientifically analyze multimodal data fusion algorithms based on the application scenarios and forms of multimodal data fusion, thus reducing the accuracy and quality of multimodal learning data fusion.
[0003] Chinese invention patent CN117828539B, published on May 24, 2024, discloses a data intelligent fusion analysis system and method. This system constructs a density among several data columns by using the similarity between any two data columns as the target distance; it analyzes and obtains the data quality coefficients of the data columns; if the data quality of a data column is lower than expected, it filters out outliers and replaces them to optimize the data column; it pre-constructs the fusion priority of each data column based on the similarity and data quality coefficients, and marks the corresponding data columns according to the fusion priority; it matches the corresponding fusion strategy from a pre-built data fusion knowledge graph based on the correspondence between data features and fusion strategies; and it uses the fusion strategy with the highest optimization coefficient as the target strategy. However, the above technical solution cannot intelligently select multimodal data fusion algorithms based on the application scenario of multimodal data fusion and the form of multimodal learning data, thus reducing the accuracy of multimodal learning data fusion. Summary of the Invention
[0004] (a) Technical problems to be solved
[0005] To address the issues raised by existing multimodal data fusion processes, which fail to accurately analyze multimodal data semantic alignment algorithms based on application scenarios and data formats, and thus reduce the accuracy and quality of multimodal learning data fusion, this paper aims to achieve the following objectives: scientifically identifying multimodal data fusion application scenarios, autonomously analyzing multimodal data fusion formats, scientifically analyzing multimodal data semantic alignment algorithms, intelligently selecting multimodal data fusion timing strategies, and intelligently matching multimodal data fusion algorithms.
[0006] (II) Technical Solution
[0007] This invention is achieved through the following technical solution: a multimodal learning data fusion and analysis method, the method comprising the following steps:
[0008] S1. Collect multimodal feature data and identify the application scenario type of multimodal data fusion to obtain multimodal data fusion application scenario information; acquire multimodal data attribute information and analyze the data form information of multimodal data fusion to obtain multimodal data fusion data form information;
[0009] S2. Based on the application scenario information and data format information of multimodal data, analyze the semantic alignment algorithm feature information of multimodal data fusion, obtain the semantic alignment algorithm information of multimodal data fusion and output it;
[0010] S3. Based on the application scenario information and data format information of multimodal data fusion, the optimal data fusion time strategy information for multimodal data fusion is selected to obtain the optimal multimodal data fusion time strategy information; based on the data fusion time strategy information of multimodal data, the feature information of the optimal data fusion algorithm required for multimodal data fusion is matched to obtain the optimal multimodal data fusion algorithm information.
[0011] Preferably, the steps for collecting multimodal feature data and identifying the application scenario type of multimodal data fusion to obtain multimodal data fusion application scenario information, and obtaining multimodal data attribute information and analyzing the data format information of multimodal data fusion to obtain multimodal data fusion data format information are as follows:
[0012] S11. Collect heterogeneous data from different data collection terminals online through the data management platform and generate a multimodal feature dataset. ,in Indicates the number of collections The multimodal feature data includes text data, image data, audio data, and video data specific to a particular domain.
[0013] S12. Based on the multimodal feature dataset and standard multimodal datasets for different application scenarios, perform multimodal data fusion to identify application scenario types and obtain multimodal data fusion application scenario information.
[0014] S13. The multimodal feature dataset is processed using the KD-tree nearest neighbor search algorithm. The multimodal feature data described in We search and collect data attribute information to obtain a multimodal data attribute information set. ,in This represents the multimodal feature data. The corresponding multimodal data attribute information includes data format information, data size information, data location information, and data opening method information; wherein the data format information includes TXT, DOC, HTML, JPEG, PNG, SVG, TIFF, MP3, FLAC, OGG, MP4, AVI, and MKV.
[0015] S14. Based on the multimodal data attribute information set and the standard multimodal data attribute datasets of different data formats, perform multimodal data fusion data format information analysis and processing to obtain multimodal data fusion data format information.
[0016] Preferably, the steps for performing multimodal data fusion application scenario type identification processing based on the multimodal feature dataset and standard multimodal datasets for different application scenarios to obtain multimodal data fusion application scenario information are as follows:
[0017] S121. Establish standard multimodal datasets for different application scenarios. ,in Indicates the first The application scenario types include medical, transportation, security, education, and entertainment standard multimodal data; the application scenario types include medical, transportation, security, education, and entertainment standard multimodal data; the standard heterogeneous data of the different application scenario types are set for the specific application scenario standard heterogeneous data of the different application scenario types.
[0018] S122. The XGBoost algorithm is used to process the multimodal feature dataset. The multimodal feature data described in Standard multimodal datasets for different application scenarios Standard multimodal data for different application scenarios described in the article Perform data matching to search for data matching with the multimodal feature dataset. Matching standard multimodal data for different application scenarios The corresponding application scenario type information is used to construct multimodal data fusion application scenario information through data identification.
[0019] Preferably, the steps for analyzing and processing the data format information of multimodal data fusion based on the multimodal data attribute information set and standard multimodal data attribute datasets of different data formats to obtain the multimodal data fusion data format information are as follows:
[0020] S141. Establish standard multimodal data attribute datasets in different data formats. ,in Indicates the first The data formats include text data, image data, audio data, and video data; the standard multimodal data attribute data for each data format corresponds to different data format types.
[0021] S142. The Rabin-Karp algorithm is used to process the multimodal data attribute information set. The multimodal data attribute information described in Different data formats of standard multimodal data attribute datasets The different data formats of standard multimodal data attribute data described in the article Perform text information matching to search for the multimodal data attribute information set. Matching all the different data formats of the standard multimodal data attribute data The corresponding data format type information is used to construct multimodal data fusion data format information through data identification. The multimodal data fusion data format information represents the text information of all data format types of the collected multimodal data.
[0022] Preferably, the steps for analyzing the semantic alignment algorithm features of multimodal data fusion based on the application scenario and data format information of the multimodal data, obtaining the semantic alignment algorithm information of multimodal data fusion, and outputting it are as follows:
[0023] S21. Based on the multimodal data fusion application scenario information, the multimodal data fusion data form information, and the multimodal data fusion semantic alignment algorithm feature information matrix, perform multimodal data fusion semantic alignment algorithm feature information analysis and processing to obtain multimodal data fusion semantic alignment algorithm information;
[0024] S22. The generated multimodal data is displayed and output using the semantic alignment algorithm information fused with the display screen.
[0025] Preferably, the steps for analyzing and processing the semantic alignment algorithm feature information of multimodal data fusion based on the multimodal data fusion application scenario information, the multimodal data fusion data format information, and the multimodal data fusion semantic alignment algorithm feature information matrix to obtain the multimodal data fusion semantic alignment algorithm information are as follows:
[0026] S211. Establish the feature information matrix of the multimodal data fusion semantic alignment algorithm. ,in Indicates the first The multimodal data feature combination type corresponds to the multimodal data fusion semantic alignment algorithm feature information; the multimodal data feature combination type represents the index data type mainly composed of application scenario information and data format information of multimodal data, which is used to search for the semantic alignment algorithm required for multimodal data fusion; the semantic alignment algorithms in the multimodal data fusion process include CLIP, ALIGN, WenLan, autoencoder, VITA model and TIP model; the multimodal data fusion semantic alignment algorithm feature information represents the name, version and source text information of the optimal semantic alignment algorithm for multimodal data fusion set for different multimodal data feature combination types;
[0027] S212. A bidirectional search algorithm is used to integrate the multimodal data fusion application scenario information, the multimodal data fusion data format information, and the feature information matrix of the multimodal data fusion semantic alignment algorithm. Feature information of the multimodal data fusion semantic alignment algorithm described in the article Perform keyword matching based on application scenarios and data formats to search for feature information of the multimodal data fusion semantic alignment algorithm that matches the application scenario information and data format information of the multimodal data fusion. Furthermore, a multimodal data fusion semantic alignment algorithm was constructed.
[0028] Preferably, the optimal data fusion time strategy information for multimodal data fusion is selected based on the application scenario information and data format information of multimodal data fusion, thus obtaining the optimal multimodal data fusion time strategy information; the operation steps for matching the optimal data fusion algorithm feature information required for multimodal data fusion based on the multimodal data fusion time strategy information to obtain the optimal multimodal data fusion algorithm information are as follows:
[0029] S31. Based on the multimodal data fusion application scenario information, the multimodal data fusion data format information, and the multimodal data fusion time strategy information matrix, perform the optimal data fusion time strategy screening process for multimodal data fusion to obtain the optimal multimodal data fusion time strategy information;
[0030] S311. Establish a multimodal data fusion time strategy information matrix. ,in Indicates the first Multimodal data fusion time strategy information corresponding to various multimodal data feature combination types. The multimodal data fusion time strategy information represents the optimal data fusion time stage scheme information based on the application scenario and data format of the data fusion. The multimodal data fusion time strategy information includes early data fusion, late data fusion, and mixed data fusion.
[0031] S312, combine the multimodal data fusion application scenario information, the multimodal data fusion data format information, and the multimodal data fusion time strategy information matrix. Multimodal data fusion timing strategy information described in the document Perform character matching based on application scenario and data format to search for multimodal data fusion time strategy information that matches the multimodal data fusion application scenario information and the multimodal data fusion data format information. And construct the optimal fusion time strategy information for multimodal data;
[0032] S32. Based on the optimal fusion time strategy information of the multimodal data and the feature information matrix of the multimodal data fusion algorithm, perform the matching processing of the feature information of the optimal data fusion algorithm required for multimodal data fusion to obtain the optimal multimodal data fusion algorithm information.
[0033] Preferably, the steps for matching the optimal data fusion algorithm feature information required for multimodal data fusion based on the optimal multimodal data fusion time strategy information and the feature information matrix of the multimodal data fusion algorithm to obtain the optimal multimodal data fusion algorithm information are as follows:
[0034] S321. Establish the feature information matrix of the multimodal data fusion algorithm. ,in This indicates the multimodal data fusion timing strategy information. The corresponding multimodal data fusion algorithm feature information; the name, version, and source text information of the optimal multimodal data fusion algorithm set for different multimodal data fusion time strategies; the multimodal data fusion algorithms include PCA, ICA, NMF, weight-based fusion methods, maximum value fusion, Progressive Fusion, and ALBEF;
[0035] S322, Combine the optimal fusion time strategy information of the multimodal data with the feature information of the multimodal data fusion algorithm. Perform text information matching of data fusion time strategy to search for feature information of the multimodal data fusion algorithm that matches the optimal fusion time strategy information of the multimodal data. And construct the optimal multimodal data fusion algorithm information.
[0036] A multimodal learning data fusion and analysis system is used to implement the multimodal learning data fusion and analysis method. The system includes a multimodal data fusion pre-analysis module, a multimodal data fusion semantic alignment scheme analysis module, and a multimodal data fusion scheme analysis module.
[0037] The multimodal data fusion pre-analysis module includes a multimodal data acquisition unit, a standard multimodal data storage unit for different application scenarios, a multimodal data fusion application scenario identification unit, a multimodal data attribute acquisition unit, a standard multimodal data attribute storage unit for different data formats, and a multimodal data fusion data format analysis unit.
[0038] The multimodal data acquisition unit collects multimodal feature data through a data management platform; the standard multimodal data storage unit for different application scenarios stores standard multimodal data attribute data in different data formats; the multimodal data fusion application scenario identification unit performs multimodal data fusion application scenario type identification processing based on the multimodal feature data and standard multimodal data for different application scenarios to obtain multimodal data fusion application scenario information; the multimodal data attribute acquisition unit collects multimodal data attribute information based on the multimodal feature data; the standard multimodal data attribute storage unit for different data formats stores standard multimodal data attribute data in different data formats; the multimodal data fusion data format analysis unit performs multimodal data fusion data format information analysis processing based on the multimodal data attribute information and standard multimodal data attribute data in different data formats to obtain multimodal data fusion data format information.
[0039] The multimodal data fusion semantic alignment scheme analysis module includes a multimodal data fusion semantic alignment algorithm text information storage unit, a multimodal data fusion semantic alignment algorithm analysis unit, and a multimodal data fusion semantic alignment algorithm output unit;
[0040] The multimodal data fusion semantic alignment algorithm text information storage unit is used to store the feature information of the multimodal data fusion semantic alignment algorithm; the multimodal data fusion semantic alignment algorithm analysis unit performs semantic alignment algorithm feature information analysis and processing based on the multimodal data fusion application scenario information, the multimodal data fusion data format information, and the multimodal data fusion semantic alignment algorithm feature information to obtain multimodal data fusion semantic alignment algorithm information; the multimodal data fusion semantic alignment algorithm output unit displays and outputs data based on the multimodal data fusion semantic alignment algorithm information and in conjunction with the display screen.
[0041] The multimodal data fusion scheme analysis module includes a multimodal data fusion time strategy information storage unit, a multimodal data fusion time strategy selection unit, a multimodal data fusion algorithm text information storage unit, and a multimodal data fusion algorithm matching unit;
[0042] The multimodal data fusion timing strategy information storage unit is used to store multimodal data fusion timing strategy information; the multimodal data fusion timing strategy selection unit performs optimal data fusion timing strategy filtering based on the multimodal data fusion application scenario information, the multimodal data fusion data format information, and the multimodal data fusion timing strategy information to obtain optimal multimodal data fusion timing strategy information; the multimodal data fusion algorithm text information storage unit is used to store multimodal data fusion algorithm feature information; the multimodal data fusion algorithm matching unit performs optimal data fusion algorithm feature information matching based on the optimal multimodal data fusion timing strategy information and the multimodal data fusion algorithm feature information to obtain optimal multimodal data fusion algorithm information.
[0043] A multimodal learning data fusion and analysis intelligent device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the multimodal learning data fusion and analysis method.
[0044] (III) Beneficial Effects
[0045] This invention provides a method, system, and intelligent device for multimodal learning data fusion and analysis. It has the following beneficial effects:
[0046] I. Efficiently collect multimodal feature information through a data management platform to provide real data support for accurately identifying multimodal data application scenarios; intelligently identify multimodal data fusion application scenario types based on multimodal feature data and combined with intelligent search algorithms and standard multimodal data from different application scenarios based on big data, achieving scientific identification of multimodal data fusion application scenario information and improving the accuracy of multimodal data fusion processing; efficiently collect multimodal data attribute information based on multimodal feature data and intelligent search algorithms, and autonomously analyze the data form information of multimodal data by combining standard multimodal data attribute data of different data forms, achieving intelligent analysis of the composition form of multimodal data; realize data composition classification analysis based on text, images, audio, and video, improving the scientific nature of multimodal data fusion processing.
[0047] Second, by combining multimodal data fusion application scenario information, multimodal data fusion data format information, and intelligent search algorithms with scientifically preset multimodal data fusion semantic alignment algorithm feature information, the semantic alignment algorithm feature information of multimodal data fusion is accurately screened. This enables precise matching of multimodal data fusion semantic alignment algorithms based on application scenarios and data formats, and improves the quality of multimodal data fusion processing by combining the visualization output on the display screen.
[0048] Third, by intelligently selecting the optimal data fusion time strategy based on multimodal data fusion application scenario information, multimodal data fusion data format information, and scientifically preset multimodal data fusion time strategy information, the system comprehensively analyzes the application scenarios and data formats of multimodal data fusion to determine the optimal multimodal data fusion time strategy, thereby improving the reliability of multimodal data fusion processing. Furthermore, by combining the optimal multimodal data fusion time strategy information with intelligent search algorithms and scientifically stored multimodal data fusion algorithm feature information, the system accurately matches the optimal data fusion algorithm feature information required for multimodal data fusion, enabling the scientific and dynamic selection of multimodal data fusion processing algorithms and improving the efficiency and accuracy of multimodal data fusion. Attached Figure Description
[0049] Figure 1 A schematic diagram of the modules of the multimodal learning data fusion and analysis system provided by the present invention;
[0050] Figure 2 The flowchart shows the multimodal learning data fusion and analysis method provided by this invention. Detailed Implementation
[0051] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0052] Examples of the multimodal learning data fusion analysis method, system, and intelligent device are as follows:
[0053] Example 1:
[0054] Please see Figures 1-2 A multimodal learning data fusion analysis method, which includes the following steps:
[0055] S1. Collect multimodal feature data and identify the application scenario type of multimodal data fusion to obtain multimodal data fusion application scenario information; acquire multimodal data attribute information and analyze the data form information of multimodal data fusion to obtain multimodal data fusion data form information;
[0056] S2. Based on the application scenario information and data format information of multimodal data, analyze the semantic alignment algorithm feature information of multimodal data fusion, obtain the semantic alignment algorithm information of multimodal data fusion and output it;
[0057] S3. Based on the application scenario information and data format information of multimodal data fusion, the optimal data fusion time strategy information for multimodal data fusion is selected to obtain the optimal multimodal data fusion time strategy information; based on the data fusion time strategy information of multimodal data, the feature information of the optimal data fusion algorithm required for multimodal data fusion is matched to obtain the optimal multimodal data fusion algorithm information.
[0058] For further details, please refer to Figures 1-2 The steps for collecting multimodal feature data and identifying the application scenario type of multimodal data fusion to obtain multimodal data fusion application scenario information, and obtaining multimodal data attribute information and analyzing the data format information of multimodal data fusion to obtain multimodal data fusion data format information are as follows:
[0059] S11. Collect heterogeneous data from different data collection terminals online through the data management platform and generate a multimodal feature dataset. ,in Indicates the number of collections Multimodal feature data; multimodal feature data includes domain-specific text data, image data, audio data, and video data;
[0060] S12. Based on the multimodal feature dataset and the standard multimodal datasets of different application scenarios, multimodal data fusion is performed to identify the application scenario type and obtain multimodal data fusion application scenario information.
[0061] S13. The multimodal feature dataset is processed using the KD-tree nearest neighbor search algorithm. Multimodal feature data We search and collect data attribute information to obtain a multimodal data attribute information set. ,in Representing multimodal feature data The corresponding multimodal data attribute information includes data format information, data size information, data location information, and data opening method information; among which, the data format information includes TXT, DOC, HTML, JPEG, PNG, SVG, TIFF, MP3, FLAC, OGG, MP4, AVI, and MKV;
[0062] S14. Based on the multimodal data attribute information set and the standard multimodal data attribute datasets of different data formats, perform multimodal data fusion data format information analysis and processing to obtain multimodal data fusion data format information.
[0063] The steps for identifying application scenario types by fusing multimodal data based on multimodal feature datasets and standard multimodal datasets for different application scenarios to obtain multimodal data fusion application scenario information are as follows:
[0064] S121. Establish standard multimodal datasets for different application scenarios. ,in Indicates the first Standard multimodal data for different application scenarios corresponding to various application scenario types; application scenario types include medical, transportation, security, education, and entertainment; standard multimodal data for different application scenarios are standard heterogeneous data for specific application scenarios set for different application scenario types;
[0065] S122. Use the XGBoost algorithm to process the multimodal feature dataset. Multimodal feature data Standard multimodal datasets for different application scenarios Standard multimodal data for different application scenarios Perform data matching to search for data matching results with multimodal feature datasets. Matching standard multimodal data for different application scenarios The corresponding application scenario type information is used to construct multimodal data fusion application scenario information through data identification.
[0066] The steps for analyzing and processing the data format information of multimodal data fusion based on the multimodal data attribute information set and standard multimodal data attribute datasets of different data formats are as follows:
[0067] S141. Establish standard multimodal data attribute datasets in different data formats. ,in Indicates the first Standard multimodal data attribute data for different data formats corresponding to various data format types; data format types include text data format type, image data format type, audio data format type, and video data format type; standard multimodal data attribute data for different data format types sets standard multimodal data attribute information for different data format types;
[0068] S142. The Rabin-Karp algorithm is used to process the multimodal data attribute information set. Multimodal data attribute information Standard multimodal data attribute datasets in different data formats Chinese standard multimodal data attribute data in different data formats Perform text information matching to search for sets of multimodal data attribute information. Matching all different data formats of standard multimodal data attribute data The corresponding data format type information is used to construct multimodal data fusion data format information through data identification. Multimodal data fusion data format information represents all data format types of text information of the collected multimodal data.
[0069] The multimodal data acquisition unit efficiently collects multimodal feature information using a data management platform, providing real data support for accurately identifying multimodal data application scenarios. The multimodal data fusion application scenario identification unit intelligently identifies the types of multimodal data fusion application scenarios based on multimodal feature data and combined with intelligent search algorithms and standard multimodal data from different application scenarios based on big data. This enables scientific identification of multimodal data fusion application scenario information and improves the accuracy of multimodal data fusion processing. The multimodal data attribute acquisition unit and the multimodal data fusion data form analysis unit work together to efficiently collect multimodal data attribute information based on multimodal feature data and intelligent search algorithms. Simultaneously, they autonomously analyze the data form information of multimodal data by combining standard multimodal data attribute data from different data forms, intelligently analyzing the composition of multimodal data. Finally, they enable data composition classification analysis of multimodal data based on text, images, audio, and video, improving the scientific rigor of multimodal data fusion processing.
[0070] For further details, please refer to Figures 1-2Based on the application scenario and data format information of multimodal data, the semantic alignment algorithm feature information of multimodal data fusion is analyzed, and the operation steps for obtaining and outputting the semantic alignment algorithm information of multimodal data fusion are as follows:
[0071] S21. Based on the application scenario information of multimodal data fusion, the data form information of multimodal data fusion, and the feature information matrix of multimodal data fusion semantic alignment algorithm, perform semantic alignment algorithm feature information analysis and processing of multimodal data fusion to obtain multimodal data fusion semantic alignment algorithm information;
[0072] S22. The generated multimodal data is fused with semantic alignment algorithm information and displayed on the screen.
[0073] The steps for analyzing and processing the semantic alignment algorithm feature information of multimodal data fusion based on multimodal data fusion application scenario information, multimodal data fusion data format information, and multimodal data fusion semantic alignment algorithm feature information matrix are as follows:
[0074] S211. Establish the feature information matrix of the multimodal data fusion semantic alignment algorithm. ,in Indicates the first The multimodal data feature combination type represents the feature information of the semantic alignment algorithm for multimodal data fusion corresponding to various multimodal data feature combination types. The multimodal data feature combination type represents the index data type, which is mainly composed of application scenario information and data format information of multimodal data, and is used to search for the semantic alignment algorithm required for multimodal data fusion. The semantic alignment algorithms in the multimodal data fusion process include CLIP, ALIGN, WenLan, autoencoder, VITA model, and TIP model. The semantic alignment algorithm feature information of multimodal data fusion represents the name, version, and source text information of the optimal semantic alignment algorithm for multimodal data fusion set for different multimodal data feature combination types.
[0075] S212. A bidirectional search algorithm is used to integrate multimodal data fusion application scenario information, multimodal data fusion data format information, and multimodal data fusion semantic alignment algorithm feature information matrix. Multimodal data fusion semantic alignment algorithm feature information Perform keyword matching based on application scenarios and data formats to search for multimodal data fusion semantic alignment algorithm feature information that matches the application scenario information and data format information of multimodal data fusion. Furthermore, a multimodal data fusion semantic alignment algorithm was constructed.
[0076] By having the multimodal data fusion semantic alignment algorithm analysis unit and the multimodal data fusion semantic alignment algorithm output unit work together, the semantic alignment algorithm feature information of multimodal data fusion is accurately screened based on the application scenario information and data format information of multimodal data fusion, combined with intelligent search algorithms and scientifically preset multimodal data fusion semantic alignment algorithm feature information. This enables accurate matching of multimodal data fusion semantic alignment algorithm based on application scenarios and data formats, and improves the quality of multimodal data fusion processing by combining the visualization output on the display screen.
[0077] For further details, please refer to Figures 1-2 The steps for selecting the optimal data fusion time strategy based on application scenario and data format information of multimodal data fusion are as follows: The optimal data fusion time strategy information is then matched with the feature information of the optimal data fusion algorithm required for multimodal data fusion to obtain the optimal multimodal data fusion algorithm information.
[0078] S31. Based on the multimodal data fusion application scenario information, multimodal data fusion data format information and multimodal data fusion time strategy information matrix, perform the optimal data fusion time strategy screening process for multimodal data fusion to obtain the optimal multimodal data fusion time strategy information;
[0079] S311. Establish a multimodal data fusion time strategy information matrix. ,in Indicates the first Multimodal data fusion time strategy information corresponding to various multimodal data feature combination types. The multimodal data fusion time strategy information represents the optimal selected time stage scheme for multimodal data fusion based on the application scenario and data format settings. The multimodal data fusion time strategy information includes early data fusion, late data fusion, and mixed data fusion.
[0080] S312, Integrate multimodal data fusion application scenario information, multimodal data fusion data format information, and multimodal data fusion time strategy information matrix. Multimodal data fusion time strategy information Perform character matching based on application scenarios and data formats to search for multimodal data fusion time strategy information that matches the application scenario information and data format information of multimodal data fusion. And construct the optimal fusion time strategy information for multimodal data;
[0081] S32. Based on the optimal fusion time strategy information of multimodal data and the feature information matrix of multimodal data fusion algorithm, perform feature information matching processing of the optimal data fusion algorithm required for multimodal data fusion to obtain the optimal multimodal data fusion algorithm information.
[0082] The steps for matching the optimal multimodal data fusion algorithm feature information required for multimodal data fusion based on the optimal multimodal data fusion time strategy information and the feature information matrix of the multimodal data fusion algorithm are as follows:
[0083] S321. Establish the feature information matrix of the multimodal data fusion algorithm. ,in Indicates multimodal data fusion timing strategy information The corresponding multimodal data fusion algorithm feature information; the name, version, and source text information of the optimal multimodal data fusion algorithm set for different multimodal data fusion time strategies; multimodal data fusion algorithms include PCA, ICA, NMF, weight-based fusion methods, maximum value fusion, Progressive Fusion, and ALBEF;
[0084] S322. Integrate the optimal fusion time strategy information of multimodal data with the feature information of the multimodal data fusion algorithm. Perform text information matching of data fusion time strategy to search for multimodal data fusion algorithm feature information that matches the optimal fusion time strategy information of multimodal data. And construct the optimal multimodal data fusion algorithm information.
[0085] The multimodal data fusion timing strategy selection unit intelligently filters the optimal multimodal data fusion timing strategy based on multimodal data fusion application scenario information, multimodal data fusion data format information, and scientifically preset multimodal data fusion timing strategy information. This enables a comprehensive analysis of multimodal data fusion application scenarios and data formats to determine the optimal multimodal data fusion timing strategy, thereby improving the reliability of multimodal data fusion processing. The multimodal data fusion algorithm matching unit accurately matches the optimal multimodal data fusion algorithm feature information required for multimodal data fusion based on the optimal multimodal data fusion timing strategy information, combined with an intelligent search algorithm and scientifically stored multimodal data fusion algorithm feature information. This enables the scientific and dynamic selection of multimodal data fusion processing algorithms, improving the efficiency and accuracy of multimodal data fusion.
[0086] Example 2:
[0087] Please see Figures 1-2A multimodal learning data fusion and analysis system is used to implement multimodal learning data fusion and analysis methods. The system includes a multimodal data fusion pre-analysis module, a multimodal data fusion semantic alignment scheme analysis module, and a multimodal data fusion scheme analysis module.
[0088] The multimodal data fusion pre-analysis module includes a multimodal data acquisition unit, a standard multimodal data storage unit for different application scenarios, a multimodal data fusion application scenario identification unit, a multimodal data attribute acquisition unit, a standard multimodal data attribute storage unit for different data formats, and a multimodal data fusion data format analysis unit.
[0089] The system comprises the following components: a multimodal data acquisition unit, which collects multimodal feature data through a data management platform; a standard multimodal data storage unit for different application scenarios, which stores standard multimodal data attribute data in different data formats; a multimodal data fusion application scenario identification unit, which identifies the application scenario type of multimodal data fusion based on multimodal feature data and standard multimodal data for different application scenarios to obtain multimodal data fusion application scenario information; a multimodal data attribute acquisition unit, which collects multimodal data attribute information based on multimodal feature data; a standard multimodal data attribute storage unit for different data formats, which stores standard multimodal data attribute data in different data formats; and a multimodal data fusion data format analysis unit, which analyzes and processes the multimodal data fusion data format information based on multimodal data attribute information and standard multimodal data attribute data in different data formats to obtain multimodal data fusion data format information.
[0090] The multimodal data fusion semantic alignment scheme analysis module includes a multimodal data fusion semantic alignment algorithm text information storage unit, a multimodal data fusion semantic alignment algorithm analysis unit, and a multimodal data fusion semantic alignment algorithm output unit;
[0091] The system includes: a text information storage unit for multimodal data fusion semantic alignment algorithm, used to store feature information of the multimodal data fusion semantic alignment algorithm; a multimodal data fusion semantic alignment algorithm analysis unit, which analyzes and processes the feature information of the multimodal data fusion semantic alignment algorithm based on the application scenario information, data format information, and feature information of the multimodal data fusion semantic alignment algorithm, to obtain the multimodal data fusion semantic alignment algorithm information; and a multimodal data fusion semantic alignment algorithm output unit, which displays and outputs the multimodal data fusion semantic alignment algorithm information in conjunction with a display screen.
[0092] The multimodal data fusion scheme analysis module includes a multimodal data fusion time strategy information storage unit, a multimodal data fusion time strategy selection unit, a multimodal data fusion algorithm text information storage unit, and a multimodal data fusion algorithm matching unit;
[0093] The system includes: a multimodal data fusion timing strategy information storage unit for storing multimodal data fusion timing strategy information; a multimodal data fusion timing strategy selection unit for filtering the optimal multimodal data fusion timing strategy based on multimodal data fusion application scenario information, multimodal data fusion data format information, and multimodal data fusion timing strategy information, to obtain the optimal multimodal data fusion timing strategy information; a multimodal data fusion algorithm text information storage unit for storing multimodal data fusion algorithm feature information; and a multimodal data fusion algorithm matching unit for matching the optimal multimodal data fusion algorithm feature information required for multimodal data fusion based on the optimal multimodal data fusion timing strategy information and the multimodal data fusion algorithm feature information, to obtain the optimal multimodal data fusion algorithm information.
[0094] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A multimodal learning data fusion and analysis method, characterized in that, The method includes the following steps: S1. Collect multimodal feature data and identify the application scenario type of multimodal data fusion to obtain multimodal data fusion application scenario information; acquire multimodal data attribute information and analyze the data form information of multimodal data fusion to obtain multimodal data fusion data form information; S2. Based on the application scenario information and data format information of multimodal data, analyze the semantic alignment algorithm feature information of multimodal data fusion, obtain the semantic alignment algorithm information of multimodal data fusion and output it; S3. Based on the application scenario information and data format information of multimodal data fusion, the optimal data fusion time strategy information for multimodal data fusion is selected to obtain the optimal multimodal data fusion time strategy information; based on the data fusion time strategy information of multimodal data, the feature information of the optimal data fusion algorithm required for multimodal data fusion is matched to obtain the optimal multimodal data fusion algorithm information; The operation steps of S3 are as follows: S31. Based on the multimodal data fusion application scenario information, the multimodal data fusion data format information, and the multimodal data fusion time strategy information matrix, perform the optimal data fusion time strategy screening process for multimodal data fusion to obtain the optimal multimodal data fusion time strategy information; S311. Establish a multimodal data fusion time strategy information matrix. The include ;in Indicates the first Information on multimodal data fusion time strategies corresponding to various multimodal data feature combination types; S312, combine the multimodal data fusion application scenario information, the multimodal data fusion data format information, and the... The above Perform character matching based on application scenario and data format to search for characters that match the multimodal data fusion application scenario information and the multimodal data fusion data format information. And construct the optimal fusion time strategy information for multimodal data; S32. Based on the optimal fusion time strategy information of the multimodal data and the feature information matrix of the multimodal data fusion algorithm, perform the matching processing of the feature information of the optimal data fusion algorithm required for multimodal data fusion to obtain the optimal multimodal data fusion algorithm information; S321. Establish the feature information matrix of the multimodal data fusion algorithm. The include ;in Indicates the Corresponding multimodal data fusion algorithm feature information; S322, Combine the optimal fusion time strategy information of the multimodal data with the... Perform data fusion time strategy text information matching to search for the optimal fusion time strategy information of the multimodal data. And construct the optimal multimodal data fusion algorithm information.
2. The multimodal learning data fusion and analysis method according to claim 1, characterized in that: The operation steps of S1 are as follows: S11. Collect heterogeneous data from different data collection terminals online through the data management platform and generate a multimodal feature dataset. The include ;in Indicates the number of collections Multimodal feature data; S12. Based on the multimodal feature dataset and standard multimodal datasets for different application scenarios, perform multimodal data fusion to identify application scenario types and obtain multimodal data fusion application scenario information. S13. The KD-tree nearest neighbor search algorithm is used to perform the... The above We search and collect data attribute information to obtain a multimodal data attribute information set. The included ;in Indicates the Corresponding multimodal data attribute information; S14. Based on the multimodal data attribute information set and the standard multimodal data attribute datasets of different data formats, perform multimodal data fusion data format information analysis and processing to obtain multimodal data fusion data format information.
3. The multimodal learning data fusion and analysis method according to claim 2, characterized in that: The operation steps of S12 are as follows: S121. Establish standard multimodal datasets for different application scenarios. The include ;in Indicates the first Standard multimodal data for different application scenarios corresponding to various application scenario types; S122, The XGBoost algorithm is used to process the... The above With the The above Perform data matching and search for results matching the above. The matching The corresponding application scenario type information is used to construct multimodal data fusion application scenario information through data identification.
4. The multimodal learning data fusion and analysis method according to claim 3, characterized in that: The operation steps of S14 are as follows: S141. Establish standard multimodal data attribute datasets in different data formats. The include ;in Indicates the first Different data format standards for multimodal data attribute data corresponding to various data format types; S142, The Rabin-Karp algorithm is used to process the... The above With the The above Perform text information matching to search for results matching the given information. All of the matching statements The corresponding data format type information is used to construct multimodal data fusion data format information through data identification.
5. The multimodal learning data fusion and analysis method according to claim 4, characterized in that: The operation steps of S2 are as follows: S21. Based on the multimodal data fusion application scenario information, the multimodal data fusion data form information, and the multimodal data fusion semantic alignment algorithm feature information matrix, perform multimodal data fusion semantic alignment algorithm feature information analysis and processing to obtain multimodal data fusion semantic alignment algorithm information; S22. The generated multimodal data is displayed and output using the semantic alignment algorithm information fused with the display screen.
6. The multimodal learning data fusion and analysis method according to claim 5, characterized in that: The operation steps of S21 are as follows: S211. Establish the feature information matrix of the multimodal data fusion semantic alignment algorithm. The include ;in Indicates the first Feature information of multimodal data fusion semantic alignment algorithm corresponding to various multimodal data feature combination types; S212. A bidirectional search algorithm is used to combine the multimodal data fusion application scenario information, the multimodal data fusion data format information, and the... The above Perform keyword matching based on application scenarios and data formats to search for keywords that match the multimodal data fusion application scenario information and the multimodal data fusion data format information. Furthermore, a multimodal data fusion semantic alignment algorithm was constructed.
7. A multimodal learning data fusion and analysis system, used to implement the multimodal learning data fusion and analysis method according to any one of claims 1-6, characterized in that: The system includes a multimodal data fusion pre-analysis module, a multimodal data fusion semantic alignment scheme analysis module, and a multimodal data fusion scheme analysis module.
8. A multimodal learning data fusion and analysis intelligent device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: The processor executes the computer program to implement the steps of the multimodal learning data fusion analysis method according to any one of claims 1-6.
Citation Information
Patent Citations
Data intelligent fusion analysis system and method
CN117828539B
Optimization method and system for adaptive multi-stage fine-tuning multi-modal large model
CN120105345A
Multi-modal knowledge graph construction method in cross-media retrieval
CN120611774A