Big data acquisition system applying artificial intelligence analysis

By introducing attention mechanisms and dynamic weight adjustments in multimodal data fusion, the problem of insufficient accuracy of data fusion in the existing technology is solved, more accurate data analysis and secure data sharing are achieved, and data quality and acquisition efficiency are improved.

CN120336755AInactive Publication Date: 2025-07-18CHANGCHUN YOUWANG INTELLIGENT TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510416055.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When processing text, image, and audio data, the existing multimodal data fusion method ignores the differences in importance of different modal data, resulting in insufficient accuracy of data fusion.

Method used

The multimodal fusion module of text, images, and audio is used to calculate the attention weight through feature extraction, attention mechanism and softmax function, dynamically adjust the weight of each mode, and build an artificial intelligence analysis model through fully connected neural networks and recurrent neural networks for data analysis.

Benefits of technology

The accuracy and effectiveness of multimodal data fusion are improved, allowing the artificial intelligence analysis module to analyze data more accurately, and guarantee data security and privacy through the data sharing module, improving data quality and collection efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336755A_ABST
    Figure CN120336755A_ABST
Patent Text Reader

Abstract

The invention discloses a big data acquisition system applying artificial intelligence analysis, which comprises a data acquisition module, relates to the technical field of data analysis, and is characterized in that when multi-modal data including texts, images and audios are processed, feature extraction is carried out on the data of different modalities, then an attention mechanism is introduced, and the data of different modalities are acquired; the method comprises the following steps: according to the correlation of different modal features, dynamically allocating weights, calculating attention weights through a softmax function, and finally fusing weighted multi-modal features to obtain fused feature vectors as comprehensive feature representations, and inputting the fused feature vectors into an artificial intelligence analysis module for data analysis; according to a traditional multi-modal fusion method, features are generally simply spliced or averagely fused, the importance difference of data of different modals is ignored, a text, image and audio multi-modal fusion module dynamically adjusts the weight of each modal, the accuracy and effectiveness of data fusion are improved, and therefore an artificial intelligence analysis module analyzes data more accurately.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data analysis, and particularly to a big data acquisition system applying artificial intelligence analysis. Background Art

[0002] Artificial intelligence is a cutting-edge scientific and technological field that crosses multiple domains, aiming to endow machines with capabilities related to human intelligence, enabling them to simulate human thinking, behavior, and decision-making processes. Data acquisition refers to the process of collecting and obtaining data from various data sources.

[0003] The publication number CN117312759B discloses an RFID data acquisition system based on artificial intelligence, including an acquisition module for reading electronic tags to acquire data on products; a statistical module connected to the acquisition module for statistically analyzing the qualification rate of the data acquired by the acquisition module; an analysis module for determining whether the product stability is qualified based on the qualification rate of the acquired data; a region division module for dividing the production line into several production regions in the case where the analysis module determines that the product stability is unqualified; a control module including an evaluation unit and a calculation unit connected to each other; the calculation unit calculates the qualification rate of the data acquired in each production region within a preset time and calculates the variance of the data acquired in the production region; the evaluation unit determines the reason for the unqualified product stability based on the variance. This improves the stability of the production line and the production qualification rate of products.

[0004] However, the above application still has the following problems: The data acquisition in existing enterprise data includes multi-modal data such as text, images, and audio. Traditional multi-modal fusion methods often simply splice or average-fuse features, ignoring the importance differences of different modal data, and the accuracy of data fusion needs to be improved. Summary of the Invention

[0005] To solve the technical problems in the background art, the present invention proposes a big data acquisition system applying artificial intelligence analysis.

[0006] The big data acquisition system applying artificial intelligence analysis proposed by the present invention includes:

[0007] A data acquisition module: used to acquire various required data from inside and outside the enterprise;

[0008] A data preprocessing module: receives the data acquired by the data acquisition module, removes duplicate data, outliers, and noise data, and performs standardization processing on the data;

[0009] Multimodal Fusion Module for Text, Image, and Audio: For the multimodal data containing text data, image data, and audio data collected by the data acquisition module, when processing the multimodal data containing text data, image data, and audio data, the multimodal fusion module for text, image, and audio is used to construct a fused feature vector for text, image, and audio;

[0010] Artificial Intelligence Analysis Module: An artificial intelligence analysis model is constructed through a fully connected neural network and a recurrent neural network. The artificial intelligence analysis model performs data analysis on the data preprocessed by the data preprocessing module and the fused feature vector of the multimodal fusion module for text, image, and audio. The artificial intelligence analysis model is used to perform data classification, data prediction, and associated data analysis on the data;

[0011] Preferably, in the multimodal fusion module for text, image, and audio, the fused feature vector for text, image, and audio is constructed as follows:

[0012] When processing multimodal data containing text data, image data, and audio data:

[0013] For text data, the word embedding technique is used to convert it into a vector representation. The text data is word-embedded to obtain a sequence of word vectors where t = 1, 2,..., T, T is the length of the text, and d w is the dimension of the word vector, representing the set of real numbers;

[0014] Then, a convolutional neural network is used for feature extraction to obtain a text feature vector d c is the number of output channels of the convolutional kernel in the convolutional neural network;

[0015] For image data, a convolutional neural network is used to extract image features. Let the output of the last convolutional layer of the image passing through the network be where H and W are the height and width of the feature map, and C is the number of channels;

[0016] An image feature vector is obtained through global average pooling operation

[0017]

[0018] The audio data is Fourier-transformed to the frequency domain to obtain spectral features where a = 1, 2,..., A, A is the number of frames of the audio, and d s is the dimension of the spectral feature;

[0019] LSTM is used for temporal feature extraction. The output of LSTM is:

[0020] h a = LSTM(s a , h a-1 );

[0021] Among them, is the hidden state at the a-th time step, and h0 is the initial hidden state;

[0022] LSTM is the long short-term memory network;

[0023] Take the hidden state of the last time step as the audio feature vector d h is the dimension of the LSTM hidden state vector;

[0024] Perform a linear transformation on the feature vectors of the three modalities of text data, image data, and audio data:

[0025]

[0026] Among them, is the learnable weight matrix;

[0027] Calculate the correlation score between modalities:

[0028] The correlation score s between text and image TI = q T · q I ;

[0029] The correlation score s between text and audio TA = q T · q A ;

[0030] The correlation score s between image and audio IA = q I · q A ;

[0031] Concatenate the correlation scores: s = [s TI , s TA , s IA ;

[0032] Calculate the attention weights through the softmax function:

[0033] α = softmax(W s s + b s ), where is the learnable weight matrix, is the bias term, and α = [α TI , α TA , α IA is the attention weight vector;

[0034] The Softmax function is a function in deep learning, mainly used for multi-classification problems, which converts a real number vector into a probability distribution;

[0035] Concatenate f T and f I to form [f T ; f I ;

[0036] Concatenate f T and f A to form [f T ; f A ;

[0037] Concatenate f I and f A to form [f I ; f A ;

[0038] Weighted fusion of modal features according to attention weights:

[0039]

[0040] where, is a learnable weight matrix, is the fused feature vector, d fusion is the dimension of the fused feature vector f fusion ;

[0041] A learnable weight matrix refers to a matrix whose element values can be continuously adjusted and updated through an optimization algorithm during model training;

[0042] Input the fused feature vector d fusion into the subsequent artificial intelligence analysis module for data analysis.

[0043] Preferably, it further includes:

[0044] Data sharing module: used to share data with other systems within the enterprise or external partners. During the data sharing process, data encryption and desensitization technologies are adopted to protect the privacy and security of the data; at the same time, audit and record of the data sharing are carried out to trace the flow and usage of the data.

[0045] Preferably, the data acquisition module adopts multi-threaded and distributed acquisition technologies to simultaneously acquire data from multiple data sources in parallel, improving the efficiency of data acquisition.

[0046] Preferably, when cleaning the data, the data preprocessing module uses a clustering algorithm to perform clustering analysis on the data and identify outliers in the data as anomalies for removal.

[0047] Preferably, it further includes:

[0048] Data storage module: By constructing a distributed database, classify and store the data collected by the data collection module and analyzed by the artificial intelligence analysis module;

[0049] Adopt data backup and recovery strategies in the data storage module; ensure the security and reliability of the data and prevent data loss.

[0050] Preferably, it further includes:

[0051] Visualization display module: Display the data classification, data prediction and associated data information of the artificial intelligence analysis module through a visualization interface.

[0052] A big data collection method applying artificial intelligence analysis includes the following steps:

[0053] S1. Collect various required data from inside and outside the enterprise;

[0054] S2. Receive the data collected in S1, remove duplicate data, outliers and noise data, and perform standardization processing on the data;

[0055] S3. For the multi-modal data containing text data, image data and audio data collected in S1, when processing the multi-modal data containing text data, image data and audio data, construct a fused feature vector of text, image and audio;

[0056] S4. Obtain the fused feature vector in S3, construct an artificial intelligence analysis model through a fully connected neural network and a recurrent neural network, and the artificial intelligence analysis model performs data analysis on the data preprocessed in S2 and the fused feature vector in S3. The artificial intelligence analysis model is used for data classification, data prediction and associated data analysis of the data.

[0057] In the present invention, the proposed big data collection system applying artificial intelligence analysis has the following beneficial technical effects:

[0058] 1. When processing multi-modal data containing text, images, and audio, first extract features from data of different modalities separately. Then, introduce an attention mechanism, dynamically allocate weights according to the correlation of features of different modalities, calculate attention weights through the softmax function. Finally, fuse the weighted multi-modal features to obtain a fused feature vector as a comprehensive feature representation, and input it into the subsequent artificial intelligence analysis module for data analysis. Traditional multi-modal fusion methods often simply splice or average the features, ignoring the importance differences of data of different modalities. The multi-modal fusion module for text, images, and audio dynamically adjusts the weights of each modality, improving the accuracy and effectiveness of data fusion, so that the artificial intelligence analysis module can analyze data more accurately.

[0059] 2. In the data preprocessing stage, there are not only cleaning, filling, and standardization processes, but also the use of clustering algorithms to identify outliers, which can better improve data quality. There is a data sharing module, and encryption and desensitization technologies are used to ensure privacy security during data sharing, and an audit record is made for the sharing, providing a safe and reliable way for data sharing inside and outside the enterprise.

[0060] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the present invention. Brief Description of the Drawings

[0061] Figure 1 is a schematic block diagram of the system of the present invention;

[0062] Figure 2 is a flowchart of the method of the present invention. Detailed Description of the Embodiments

[0063] The embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the drawings, where the same or similar reference signs denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the drawings are exemplary only for explaining the present invention and should not be construed as limiting the present invention.

[0064] As Figure 1 shown, a big data acquisition system applying artificial intelligence analysis includes:

[0065] Data acquisition module: used to collect various required data from inside and outside the enterprise;

[0066] Data preprocessing module: receives the data collected by the data acquisition module, removes duplicate data, outliers, and noise data, and performs standardization processing on the data;

[0067] Multimodal Fusion Module for Text, Image, and Audio: For the multimodal data collected by the data acquisition module, which includes text data, image data, and audio data, when processing the multimodal data containing text data, image data, and audio data, the multimodal fusion module for text, image, and audio is used to construct a fused feature vector of text, image, and audio;

[0068] Artificial Intelligence Analysis Module: An artificial intelligence analysis model is constructed through a fully connected neural network and a recurrent neural network. The artificial intelligence analysis model performs data analysis on the data preprocessed by the data preprocessing module and the fused feature vector of the multimodal fusion module for text, image, and audio. The artificial intelligence analysis model is used to perform data classification, data prediction, and associated data analysis on the data;

[0069] Data Storage Module: By constructing a distributed database, the data collected by the data acquisition module and analyzed by the artificial intelligence analysis module are classified and stored;

[0070] The data storage module adopts a data backup and recovery strategy; to ensure the security and reliability of the data and prevent data loss;

[0071] Data Sharing Module: It is used to share data with other systems within the enterprise or external partners. During the data sharing process, data encryption and desensitization technologies are adopted to protect the privacy and security of the data; at the same time, the data sharing is audited and recorded to trace the flow and usage of the data. The data shared by the data sharing module includes the data in the data storage module;

[0072] There is a data sharing module, and encryption and desensitization technologies are used during data sharing to ensure privacy and security, and the sharing is audited and recorded, providing a safe and reliable way for internal and external data sharing within the enterprise.

[0073] Visualization Display Module: Displays the data classification, data prediction, and associated data information of the artificial intelligence analysis module through a visualization interface.

[0074] Furthermore, in the multimodal fusion module for text, image, and audio, the fused feature vector of text, image, and audio is constructed as follows:

[0075] When processing multimodal data containing text data, image data, and audio data:

[0076] For text data, the word embedding technique is used to convert it into a vector representation, and the text data is word-embedded to obtain a sequence of word vectors where t = 1, 2,..., T, T is the length of the text, d w is the dimension of the word vector, represents the set of real numbers;

[0077] Then, a convolutional neural network is used for feature extraction to obtain a text feature vector d c is the number of output channels of the convolutional kernel in the convolutional neural network;

[0078] For image data, a convolutional neural network is used to extract image features. Let the output of the last convolutional layer of the network for the image be where H and W are the height and width of the feature map, and C is the number of channels;

[0079] An image feature vector is obtained through global average pooling operation

[0080]

[0081] The audio data is Fourier-transformed to the frequency domain to obtain spectral features where a = 1, 2,..., A, A is the number of frames of the audio, and d s is the dimension of the spectral features;

[0082] LSTM is used for temporal feature extraction. The output of the LSTM is:

[0083] h a = LSTM(s a , h a-1 );

[0084] where, is the hidden state at the a-th time step, and h0 is the initial hidden state;

[0085] LSTM is the long short-term memory network;

[0086] The hidden state at the last time step is taken as the audio feature vector d h is the dimension of the LSTM hidden state vector;

[0087] The feature vectors of the three modalities of text data, image data, and audio data are linearly transformed:

[0088]

[0089] where, is the learnable weight matrix;

[0090] The correlation scores between modalities are calculated:

[0091] The correlation score s between text and image TI = q T ·q I ;

[0092] The relevance score s between text and audio TA = q T ·q A ;

[0093] The relevance score s between image and audio IA = q I ·q A ;

[0094] Concatenate the relevance scores: s = [s TI , s TA , s IA ;

[0095] Calculate the attention weights through the softmax function:

[0096] α = softmax(W s s + b s ), where is a learnable weight matrix, is a bias term, and α = [α TI , α TA , α IA is the attention weight vector;

[0097] The Softmax function is a function in deep learning, mainly used for multi-classification problems, which converts a real number vector into a probability distribution;

[0098] Concatenate f T and f I to form [f T ; f I ;

[0099] Concatenate f T and f A to form [f T ; f A ;

[0100] Concatenate f I and f A to form [f I ; f A ;

[0101] Weighted fusion of modal features according to attention weights:

[0102]

[0103] where is a learnable weight matrix, is the fused feature vector, and d fusion is the fused feature vector f fusionDimension;

[0104] A learnable weight matrix refers to a matrix whose element values can be continuously adjusted and updated through an optimization algorithm during model training.

[0105] Input the fused feature vector f fusion Into the subsequent artificial intelligence analysis module for data analysis.

[0106] When processing multi-modal data containing text, images, and audio, first extract features from different modalities of data separately. Then, introduce an attention mechanism, dynamically allocate weights according to the correlation of different modality features, calculate attention weights through the softmax function. Finally, fuse the weighted multi-modal features to obtain a fused feature vector as a comprehensive feature representation, and input it into the subsequent artificial intelligence analysis module for data analysis. Traditional multi-modal fusion methods often simply splice or average fuse features, ignoring the importance differences of different modality data. The multi-modal fusion module for text, images, and audio dynamically adjusts the weights of each modality, improving the accuracy and effectiveness of data fusion, thus enabling the artificial intelligence analysis module to analyze data more accurately.

[0107] Furthermore, the data acquisition module adopts multi-threading and distributed acquisition technologies to simultaneously acquire data in parallel from multiple data sources, improving the efficiency of data acquisition.

[0108] Furthermore, when cleaning data, the data preprocessing module uses a clustering algorithm to perform clustering analysis on the data and identify outliers in the data as anomalies to be removed.

[0109] In the data preprocessing link, not only cleaning, filling, and standardization processing are performed, but also a clustering algorithm is used to identify outliers, which can better improve data quality.

[0110] Such as Figure 2 A big data acquisition method applying artificial intelligence analysis as shown, includes the following steps:

[0111] S1. Collect various required data from inside and outside the enterprise;

[0112] S2. Receive the data collected in S1, remove duplicate data, outliers, and noise data, and perform standardization processing on the data;

[0113] S3. For the multi-modal data containing text data, image data, and audio data collected in S1, when processing the multi-modal data containing text data, image data, and audio data, construct a fused feature vector for text, images, and audio.

[0114] S4. Obtain the fused feature vectors in S3, construct an artificial intelligence analysis model through a fully connected neural network and a recurrent neural network. The artificial intelligence analysis model performs data analysis on the preprocessed data in S2 and the fused feature vectors in S3. The artificial intelligence analysis model is used for data classification, data prediction, and correlation data analysis of the data.

[0115] Meanwhile, the content not detailedly described in this specification belongs to the prior art well-known to those skilled in the art.

[0116] In the embodiments provided by the present invention, it should be understood that the disclosed system or method can be implemented in other ways. For example, the above-described invention embodiments are merely illustrative. For example, the division of modules is only a logical function division, and there can be other division methods in actual implementation.

[0117] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules. They can be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0118] In addition, in each embodiment of the present invention, the functional modules can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module. The above integrated modules can be implemented in the form of hardware or in the form of a combination of hardware and software functional modules.

[0119] For those skilled in the field of operation and maintenance, it is obvious that the present invention is not limited to the details of the above-described exemplary embodiments, and can be implemented in other specific forms without departing from the basic characteristics of the present invention.

[0120] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent replacements or changes, and all should be covered within the protection scope of the present invention.

Claims

1. A big data collection system applying artificial intelligence analysis, characterized in that, Including: Data acquisition module: Collect various required data from both inside and outside the enterprise; Data preprocessing module: Receive the data collected by the data acquisition module, remove duplicate data, outliers and noise data, and perform standardization processing on the data; Multi-modal fusion module for text, image and audio: For the multi-modal data containing text data, image data and audio data collected by the data acquisition module, when processing the multi-modal data containing text data, image data and audio data, the multi-modal fusion module for text, image and audio is used to construct the fused feature vectors of text, image and audio; Artificial intelligence analysis module: Construct an artificial intelligence analysis model through a fully connected neural network and a recurrent neural network. The artificial intelligence analysis model performs data analysis on the data preprocessed by the data preprocessing module and the fused feature vectors of the multi-modal fusion module for text, image and audio. The artificial intelligence analysis model is used to perform data classification, data prediction and associated data analysis on the data.

2. The big data acquisition system for artificial intelligence analysis according to claim 1, characterized in that, In the multi-modal fusion module for text, image and audio, the fused feature vectors of text, image and audio are constructed as follows: When processing multi-modal data containing text data, image data and audio data: For text data, the word embedding technique is used to convert it into a vector representation. The text data is word-embedded to obtain a sequence of word vectors where t = 1, 2,..., T, T is the length of the text, and d w is the dimension of the word vector, representing the set of real numbers; Then, a convolutional neural network is used for feature extraction to obtain a text feature vector d c is the number of output channels of the convolutional kernel in the convolutional neural network; For the image data, a convolutional neural network is used to extract image features. Let the output of the last convolutional layer of the image passing through the network be where H and W are the height and width of the feature map, and C is the number of channels; Obtain the image feature vector through the global average pooling operation Perform Fourier transform on the audio data to convert it to the frequency domain and obtain spectral features where a = 1, 2, ..., A, A is the number of frames of the audio, and d s is the dimension of the spectral features; Use LSTM for time series feature extraction, and the output of LSTM is: h a = LSTM(s a , h a-1 ); where, is the hidden state at the a-th time step, and h0 is the initial hidden state; Take the hidden state of the last time step as the audio feature vector d h is the dimension of the LSTM hidden state vector; Perform linear transformation on the feature vectors of the three modalities of text data, image data and audio data: Among them, is a learnable weight matrix; Calculate the correlation score between modalities: The relevance score s between text and image TI = q T ·q I ; The relevance score s between text and audio TA = q T ·q A ; The correlation score s between the image and the audio IA = q I ·q A ; Concatenate the relevance scores: s = [s TI , s TA , s IA ; Calculate the attention weights through the softmax function: α = softmax(W s s + b s ), where is a learnable weight matrix, is a bias term, and α = [α TI , α TA , α IA is an attention weight vector; Concatenate f T and f I to form a vector [f T ; f I ; Concatenate f T and f A to form a vector [f T ; f A ; Concatenate f I and f A to form a vector [f I ; f A ; Perform weighted fusion on the modality features according to the attention weights: Among them, is a learnable weight matrix, is the fused feature vector, d fusion is the fused feature vector f fusion is the dimension of; Input the fused feature vector f fusion into the artificial intelligence analysis module for data analysis.

3. The big data collection system for artificial intelligence analysis according to claim 1, characterized in that Also including: Data sharing module: Used for data sharing with other systems inside the enterprise or external partners. During the data sharing process, data encryption and desensitization technologies are adopted; at the same time, audit and recording of data sharing are carried out.

4. The big data collection system using artificial intelligence analysis according to claim 1, wherein The data acquisition module adopts multi-threaded and distributed acquisition technologies to collect data in parallel from multiple data sources.

5. The big data collection system using artificial intelligence analysis according to claim 1, characterized in that, When the data preprocessing module cleans the data, it uses a clustering algorithm to perform clustering analysis on the data and identifies the outliers in the data as outliers for removal.

6. The big data collection system for artificial intelligence analysis according to claim 1, characterized in that Also including: Data storage module: Classify and store the data collected by the data acquisition module and analyzed by the artificial intelligence analysis module by constructing a distributed database; Data backup and recovery strategies are adopted in the data storage module.

7. The big data collection system using artificial intelligence analysis according to claim 1, characterized in that, Also including: Visualization display module: Display the data classification, data prediction and associated data information of the artificial intelligence analysis module through a visualization interface.

8. The big data collection method using artificial intelligence analysis according to any one of claims 1-7, characterized in that Including the following steps: S1. Collect various required data from both inside and outside the enterprise; S2. Receive the data collected in S1, remove duplicate data, outliers and noise data, and perform standardization processing on the data; S3. For the multi-modal data containing text data, image data and audio data collected in S1, when processing the multi-modal data containing text data, image data and audio data, construct the fused feature vectors of text, image and audio; S4. Obtain the fused feature vectors in S3, construct an artificial intelligence analysis model through a fully connected neural network and a recurrent neural network. The artificial intelligence analysis model performs data analysis on the preprocessed data in S2 and the fused feature vectors in S3. The artificial intelligence analysis model is used for data classification, data prediction, and correlation data analysis of the data.

Citation Information

Patent Citations

  • RFID data collection system based on artificial intelligence

    CN117312759B