Big data analysis method based on multi-modal data fusion
Through the combined design of high-dimensional and low-dimensional feature extraction units and the dynamic weight allocation mechanism, combined with the distributed computing environment, the problem of excessive computing resource consumption in multimodal data fusion is solved, and efficient and accurate data analysis is achieved.
Patent Information
- Application Number
- CN202510555151.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-08-15
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art consumes too high computing resources in multimodal data fusion, and the analysis speed is slow, making it difficult to take into account both real-time and accuracy.
The combination design of high-dimensional feature extraction unit and low-dimensional feature extraction unit is adopted, combined with the dynamic weight allocation mechanism, features are extracted through deep neural networks and linear dimensionality reduction algorithms, and data processing is performed in a distributed computing environment, using the collaborative working mechanism between the master and slave nodes and the publish-subscribe mode of the message queue.
It significantly reduces the computational complexity, improves the efficiency and accuracy of data analysis, enhances the scalability and fault tolerance of the system, and optimizes the real-time and accuracy of multimodal big data analysis.
Smart Images

Figure CN120492915A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data analysis and data fusion, and in particular to a big data analysis method based on multimodal data fusion. Background Art
[0002] With the continuous development of multimodal data fusion technology, the requirements for data analysis accuracy and efficiency are increasing, leading to a significant increase in computing resource requirements. In order to effectively integrate multi-source heterogeneous data and avoid information loss, complex fusion algorithms must be introduced, resulting in increased processing time and hardware costs.
[0003] Currently, optimizing the data fusion process through distributed architectures reduces resource consumption and addresses some bottlenecks in multimodal data processing. However, this technology suffers from high-dimensional data complexity, resulting in slower analysis speeds and difficulty balancing real-time performance and accuracy. Therefore, to overcome these shortcomings, a more efficient big data analysis method is needed. Summary of the Invention
[0004] The purpose of the invention is to provide a big data analysis method based on multimodal data fusion, which solves the problems mentioned in the background technology.
[0005] The present invention is implemented as follows: a big data analysis method based on multimodal data fusion, comprising: obtaining at least two heterogeneous data sources, and generating at least two feature extraction modules based on each data source, wherein the feature extraction module includes at least one high-dimensional feature extraction unit and at least one low-dimensional feature extraction unit; constructing a data fusion model according to each feature extraction module, wherein the data fusion model includes at least one feature integration layer, and the feature integration layer includes at least one high-dimensional feature extraction unit and at least one low-dimensional feature extraction unit; obtaining a target sample data set, and training the data fusion model based on the target sample data set until the data fusion model training termination condition is met, wherein the target sample data set is a data set used to train the data fusion model.
[0006] In the above method, the high-dimensional feature extraction unit is implemented using a deep neural network, with its input directly connected to the original data source and its output connected to the feature integration layer. The low-dimensional feature extraction unit is implemented using a linear dimensionality reduction algorithm, with its input connected to the preprocessed data source and its output also connected to the feature integration layer. The feature integration layer weights and integrates the outputs of the high-dimensional and low-dimensional feature extraction units using a dynamic weight allocation mechanism that adjusts the weights in real time based on the distribution characteristics of the current input data.
[0007] According to a second aspect of an embodiment of the present invention, a method for processing multi-source heterogeneous data is provided, comprising: obtaining multimodal data to be analyzed; inputting the multimodal data into a data fusion model, and obtaining analysis results output by the data fusion model, wherein the data fusion model is trained by the above-mentioned big data analysis method based on multimodal data fusion.
[0008] In the above method, the multimodal data includes structured data, unstructured data, and semi-structured data. Structured data is acquired through a relational database interface, unstructured data is acquired through a file system interface, and semi-structured data is acquired through a key-value store interface. The input end of the data fusion model is connected to a multimodal data preprocessing module, which includes a data cleaning unit, a data normalization unit, and a data labeling unit, which respectively perform noise filtering, numerical standardization, and label mapping on the input data.
[0009] According to a third aspect of an embodiment of the present invention, a multimodal data fusion method is provided, which is applied to a distributed computing environment, and includes: receiving a multimodal data fusion request sent by an edge node, wherein the multimodal data fusion request includes multimodal data to be analyzed; inputting the multimodal data into a data fusion model, and obtaining an analysis result output by the data fusion model, wherein the data fusion model is trained by the above-mentioned big data analysis method based on multimodal data fusion; and returning the analysis result to the edge node.
[0010] In the above method, the distributed computing environment includes a master node and multiple slave nodes. The master node is responsible for task scheduling and resource allocation, while the slave nodes are responsible for executing specific data processing tasks. The data fusion model is deployed on the master node, which communicates with the slave nodes via a message queue. The message queue is implemented using a publish-subscribe model, with the master node acting as the publisher and the slave nodes as subscribers. Each slave node is equipped with a local cache module for temporarily storing intermediate computation results. The master node monitors the status of the slave nodes through a heartbeat detection mechanism to ensure the continuity of task execution.
[0011] The solution of the embodiment of the present invention can significantly reduce the computational complexity while ensuring the integrity of data features by introducing a combined design of a high-dimensional feature extraction unit and a low-dimensional feature extraction unit. The high-dimensional feature extraction unit captures subtle changes in the data, while the low-dimensional feature extraction unit focuses on the extraction of core features. The two work together through a dynamic weight allocation mechanism to avoid information loss or redundancy problems that may be caused by a single feature extraction method. In addition, the design of the feature integration layer enables the model to adaptively adjust the feature weights according to the distribution characteristics of the input data, thereby achieving optimal feature expression in different scenarios.
[0012] In a distributed computing environment, the collaborative working mechanism of master and slave nodes enables efficient processing and analysis of multimodal data. The master node's task scheduling strategy, combined with the slave node's local caching mechanism, not only reduces data transmission overhead but also improves the system's fault tolerance. The publish-subscribe model of the message queue further enhances the system's scalability, supporting the dynamic addition and removal of slave nodes to adapt to varying computing needs.
[0013] The technical solution of this embodiment of the present invention solves the existing problems of excessive computing resource consumption and slow analysis speed in multimodal data fusion by combining multi-level feature extraction with a dynamic weight allocation mechanism. Furthermore, the design of a distributed computing environment optimizes the data processing process, improving the system's real-time performance and accuracy, and providing a new solution for multimodal big data analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 A schematic diagram of a process for a big data analysis method based on multimodal data fusion provided by an embodiment of the present invention;
[0015] Figure 2 A detailed block diagram of multimodal data and processing modules of a big data analysis method based on multimodal data fusion provided by an embodiment of the present invention;
[0016] Figure 3 A detailed block diagram of a feature extraction module of a big data analysis method based on multimodal data fusion provided by an embodiment of the present invention;
[0017] Figure 4 Detailed block diagram of the distributed computing environment for the big data analysis method based on multimodal data fusion provided in an embodiment of the present invention.
[0018] The accompanying drawings are marked as follows: 1. Heterogeneous data source; 2. Feature extraction module; 3. High-dimensional feature extraction unit; 4. Low-dimensional feature extraction unit; 5. Data fusion model; 6. Feature integration layer; 7. Target sample data set; 8. Dynamic weight allocation mechanism; 9. Multimodal data preprocessing module; 10. Distributed computing environment. DETAILED DESCRIPTION
[0019] The embodiment of the present invention provides a big data analysis method based on multimodal data fusion, combining Figure 1The flowchart shown in the figure is described in detail below. The components and their numbers are marked in the figure, including a heterogeneous data source 1, a feature extraction module 2, a high-dimensional feature extraction unit 3, a low-dimensional feature extraction unit 4, a data fusion model 5, a feature integration layer 6, a target sample dataset 7, a dynamic weight allocation mechanism 8, a multimodal data preprocessing module 9, and a distributed computing environment 10. These components implement the technical solution of the present invention through specific connection relationships and collaborative methods.
[0020] In practical applications, it is first necessary to obtain at least two heterogeneous data sources 1. These data sources can be structured data, unstructured data, and semi-structured data from different fields. For example, in a smart city scenario, structured data may come from the relational database of a traffic flow monitoring system, unstructured data may come from the file storage interface of a video surveillance system, and semi-structured data may come from the key-value storage interface of a sensor network. The heterogeneous data source 1 is connected to the multimodal data preprocessing module 9 through a data transmission channel. The module includes a data cleaning unit, a data normalization unit, and a data labeling unit. The data cleaning unit is responsible for removing noise data, the data normalization unit performs numerical standardization on the data, and the data labeling unit adds labels to the data for subsequent analysis. The preprocessed data is passed to the feature extraction module 2.
[0021] The feature extraction module 2 consists of a high-dimensional feature extraction unit 3 and a low-dimensional feature extraction unit 4. The high-dimensional feature extraction unit 3 is implemented through a deep neural network, with its input directly connected to the original data source 1 and its output connected to the feature integration layer 6. The low-dimensional feature extraction unit 4 is implemented through a linear dimensionality reduction algorithm, with its input connected to the preprocessed data source and its output also connected to the feature integration layer 6. There is no direct physical connection between the high-dimensional feature extraction unit 3 and the low-dimensional feature extraction unit 4, but their output results are weighted and integrated in the feature integration layer 6. The feature integration layer 6 adjusts the output results of the two types of feature extraction units in real time through a dynamic weight allocation mechanism 8. This adjustment depends on the distribution characteristics of the current input data, thereby ensuring that the model can adaptively optimize feature expression according to different data scenarios.
[0022] The data fusion model 5 is composed of multiple feature extraction modules 2 and a feature integration layer 6. When constructing the data fusion model 5, it is first necessary to define the specific parameters of each feature extraction module 2, such as the number of neural network layers and the number of nodes in each layer of the high-dimensional feature extraction unit 3, and the dimensionality reduction dimension of the low-dimensional feature extraction unit 4. Subsequently, these modules are assembled according to a predetermined topological structure to form a complete model architecture. In order to train the data fusion model 5, a target sample data set 7 is required. The target sample data set 7 is a data set specifically for training, which contains a large amount of labeled multimodal data. During the training process, the target sample data set 7 is sent to the data fusion model 5 through the data input interface. The model will gradually adjust the internal parameters according to the input data until the training termination condition is met. The training termination condition is usually set as the prediction error of the model is lower than a certain threshold or reaches a predetermined number of iterations.
[0023] In the distributed computing environment 10, the deployment and operation process of the data fusion model 5 is as follows: the master node is responsible for task scheduling and resource allocation, and the slave node is responsible for executing specific data processing tasks. The master node and the slave node communicate through a message queue, which is implemented using a publish-subscribe model. The master node acts as a publisher, sending task instructions to the message queue, and the slave node acts as a subscriber, receiving task instructions and executing the corresponding computing tasks. Each slave node is configured with a local cache module for temporarily storing intermediate calculation results to reduce data transmission overhead. The master node monitors the status of the slave node through a heartbeat detection mechanism. If a slave node fails, the master node will reallocate the node's tasks to other available nodes to ensure the continuity of task execution. After receiving the user's multimodal data fusion request, the edge node sends the data in the request to the master node. The master node inputs the data into the data fusion model 5 for analysis and returns the analysis results to the edge node.
[0024] Throughout the system, the positional and coordinated relationships between various components are crucial. For example, heterogeneous data source 1 sits at the forefront of the system, providing raw data. Multimodal data preprocessing module 9 follows closely behind, cleaning, normalizing, and labeling the raw data. Feature extraction module 2 and feature integration layer 6 together form the core of data fusion model 5, responsible for extracting and integrating features from preprocessed data. Dynamic weight allocation mechanism 8, embedded within feature integration layer 6, adjusts feature weights based on the distribution characteristics of the input data. Target sample dataset 7, serving as the source of training data, is closely tied to the training process of data fusion model 5. In a distributed computing environment 10, master and slave nodes communicate via message queues. The slave node's local cache module works in conjunction with the master node's heartbeat detection mechanism to ensure efficient system operation.
[0025] In actual operation, suppose a smart factory needs to monitor and analyze the status of equipment on its production line in real time. In this case, heterogeneous data sources 1 include vibration and temperature data from a sensor network, as well as video data from cameras. These data are cleaned and normalized by the multimodal data preprocessing module 9 before being fed into the feature extraction module 2. The high-dimensional feature extraction unit 3 uses a deep neural network to capture subtle changes in the video data, such as tiny cracks on the equipment surface. The low-dimensional feature extraction unit 4 uses a linear dimensionality reduction algorithm to extract core features from the vibration and temperature data, such as frequency distribution and temperature trends. The feature integration layer 6 weights and integrates these two types of features using a dynamic weight allocation mechanism 8 to generate the final feature representation. Subsequently, the data fusion model 5 is trained using the target sample dataset 7 to learn how to predict the health status of the equipment based on the input data. In a distributed computing environment 10, the master node coordinates the work of multiple slave nodes, which use local cache modules to temporarily store intermediate results, thereby accelerating computation and improving the system's fault tolerance.
[0026] As can be seen from the above implementation, this invention solves the existing problems of excessive computing resource consumption and slow analysis speed in multimodal data fusion by combining multi-level feature extraction with a dynamic weight allocation mechanism. Furthermore, the design of a distributed computing environment optimizes the data processing flow, improving the system's real-time performance and accuracy, and providing a new solution for multimodal big data analysis.
[0027] In order to better enable relevant personnel in this technical field to fully understand and implement the present invention, the specific implementation principle of the present invention is further supplemented below with reference to a specific application scenario.
[0028] In the real-time monitoring scenario of equipment status in a smart factory, it is first necessary to obtain heterogeneous data sources 1. These data include vibration data and temperature data from the sensor network and video data from the camera. Vibration data and temperature data are transmitted through a key-value storage interface, while video data is transmitted through a file system interface. Subsequently, this raw data is sent to the multimodal data preprocessing module 9 through a data transmission channel. The data cleaning unit in the multimodal data preprocessing module 9 is responsible for removing noise data, such as removing outliers or filling missing values; the data normalization unit standardizes data of different dimensions to ensure consistency in the numerical range; and the data annotation unit adds labels to the data, such as marking normal or abnormal status based on historical records. The data processed through the above steps is passed to the feature extraction module 2.
[0029] Feature extraction module 2 consists of a high-dimensional feature extraction unit 3 and a low-dimensional feature extraction unit 4. High-dimensional feature extraction unit 3 is implemented using a deep neural network, with its input directly connected to the video data source 1 and its output connected to the feature integration layer 6. In actual operation, the deep neural network uses multi-layer convolution operations to capture subtle changes in video data, such as the morphological characteristics of small cracks on the device surface. Low-dimensional feature extraction unit 4 is implemented using a linear dimensionality reduction algorithm. Its input is connected to preprocessed vibration and temperature data, and its output is also connected to the feature integration layer 6. The linear dimensionality reduction algorithm extracts core features that reflect the health status of the device by performing principal component analysis on high-frequency vibration signals and temperature trends. The outputs of high-dimensional feature extraction unit 3 and low-dimensional feature extraction unit 4 are weighted and integrated in the feature integration layer 6. A dynamic weight allocation mechanism 8 adjusts the weights of the two feature types in real time based on the distribution characteristics of the current input data. For example, if cracks on the device surface are more obvious, the dynamic weight allocation mechanism 8 will increase the weight of the high-dimensional features; if the vibration frequency is significantly abnormal, the dynamic weight allocation mechanism 8 will increase the weight of the low-dimensional features.
[0030] The data fusion model 5 consists of multiple feature extraction modules 2 and a feature integration layer 6. When constructing the data fusion model 5, the specific parameters of each feature extraction module 2 must be defined. For example, the deep neural network of the high-dimensional feature extraction unit 3 is configured to include five convolutional layers, each containing 64, 128, 256, 512, and 1024 nodes, respectively. The dimensionality reduction of the low-dimensional feature extraction unit 4 is configured to retain the first three principal components. These modules are then assembled according to a predetermined topological structure to form a complete model architecture. To train the data fusion model 5, a target sample dataset 7 is fed into the model through a data input interface. The target sample dataset 7 contains a large amount of labeled multimodal data, such as historical equipment data marked as normal or abnormal. During training, the model gradually adjusts its internal parameters until the prediction error falls below a set threshold or reaches a predetermined number of iterations. The prediction error is calculated using a mean squared error loss function, and the model parameters are optimized using a backpropagation algorithm.
[0031] In the distributed computing environment 10, the deployment and operation process of the data fusion model 5 is as follows: the master node receives the multimodal data fusion request sent by the edge node and distributes the data in the request to the slave node for processing. The master node and the slave node communicate through the message queue, which is implemented in a publish-subscribe mode. The master node acts as a publisher and sends task instructions to the message queue; the slave node acts as a subscriber and receives task instructions and executes the corresponding computing tasks. Each slave node is configured with a local cache module for temporarily storing intermediate calculation results, such as partial feature extraction results or model parameter update values, thereby reducing data transmission overhead. The master node monitors the status of the slave node through a heartbeat detection mechanism. If a slave node fails, the master node will reallocate its tasks to other available nodes to ensure the continuity of task execution. Finally, the master node returns the analysis results to the edge node for users to view real-time monitoring results of the device status.
[0032] It can be seen from the above steps that the technical solution of the present invention solves the problems of excessive consumption of computing resources and slow analysis speed in the process of multimodal data fusion in the prior art by combining multi-level feature extraction with a dynamic weight allocation mechanism. The high-dimensional feature extraction unit 3 is able to capture subtle changes in the data, while the low-dimensional feature extraction unit 4 focuses on the extraction of core features. The two work together through a dynamic weight allocation mechanism to avoid information loss or redundancy problems that may be caused by a single feature extraction method. In addition, the design of the distributed computing environment optimizes the data processing flow. The collaborative working mechanism of the master node and the slave node combined with the application of the local cache module not only reduces the data transmission overhead, but also improves the fault tolerance of the system, thereby improving the real-time performance and accuracy of the system.
[0033] Any information not described in detail in this specification belongs to the prior art known to those skilled in the art, and the model parameters of each electrical appliance are not specifically limited; conventional equipment can be used. Electrical control components not mentioned in this technical solution are not shown in the figures because they belong to the prior art and will not be described here.
[0034] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
[0035] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A big data analysis method based on multimodal data fusion, characterized in that: Acquire at least two heterogeneous data sources, and generate at least two feature extraction modules based on each heterogeneous data source, wherein the feature extraction module includes at least one high-dimensional feature extraction unit and at least one low-dimensional feature extraction unit; Constructing a data fusion model according to each feature extraction module, wherein the data fusion model includes at least one feature integration layer, and the feature integration layer includes at least one high-dimensional feature extraction unit and at least one low-dimensional feature extraction unit; Obtain a target sample data set and train the data fusion model based on the target sample data set until the data fusion model training termination condition is met.
2. The big data analysis method based on multimodal data fusion according to claim 1, characterized in that: The high-dimensional feature extraction unit is implemented through a deep neural network, whose input is directly connected to the original data source and the output is connected to the feature integration layer; The low-dimensional feature extraction unit is implemented through a linear dimensionality reduction algorithm, with its input end connected to the preprocessed data source and its output end connected to the feature integration layer.
3. The big data analysis method based on multimodal data fusion according to claim 1, characterized in that: The feature integration layer performs weighted integration on the output results of the high-dimensional feature extraction unit and the low-dimensional feature extraction unit through a dynamic weight allocation mechanism. The dynamic weight allocation mechanism adjusts the weight value in real time based on the distribution characteristics of the current input data.
4. A method for processing multi-source heterogeneous data, comprising: Obtain multimodal data to be analyzed; The multimodal data is input into a data fusion model to obtain an analysis result output by the data fusion model, wherein the data fusion model is trained by the method described in any one of claims 1 to 3.
5. The method according to claim 4, wherein: Multimodal data includes structured data, unstructured data and semi-structured data. Structured data is obtained through the relational database interface, unstructured data is obtained through the file system interface, and semi-structured data is obtained through the key-value storage interface.
6. A multimodal data fusion method applied to a distributed computing environment, comprising: receiving a multimodal data fusion request sent by an edge node, wherein the multimodal data fusion request includes multimodal data to be analyzed; Input the multimodal data into a data fusion model to obtain an analysis result output by the data fusion model, wherein the data fusion model is trained by the method described in any one of claims 1 to 3; and return the analysis result to the edge node.
7. The method according to claim 6, wherein: The distributed computing environment consists of a master node and multiple slave nodes. The master node is responsible for task scheduling and resource allocation, while the slave nodes are responsible for executing specific data processing tasks. The data fusion model is deployed on the master node, which communicates with the slave nodes through a message queue. The message queue is implemented in a publish-subscribe mode, with the master node as the publisher and the slave nodes as subscribers. Each slave node is equipped with a local cache module for temporarily storing intermediate calculation results. The master node monitors the status of the slave nodes through a heartbeat detection mechanism.
Citation Information
Cited By
Multi-source heterogeneous data fusion processing method and system
CN121009504A