Multi-modal data analysis method and system based on tensor decomposition
Through the multimodal data analysis method based on tensor decomposition, a variety of physiological signal data are acquired and processed, and the problem of low credibility of emotion and intention recognition in the prior art is solved, achieving higher accuracy and more comprehensive user physiological state reflection.
Patent Information
- Application Number
- CN202510466182.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-29
AI Technical Summary
The emotional recognition method based on facial expressions and speech in the prior art is low in credibility under specific circumstances and cannot accurately identify the user's emotions and intentions.
A multimodal data analysis method based on tensor decomposition is adopted to obtain multiple types of multimodal physiological signal data (such as image data, speech data, and EEG signal data), pre-processing, modal alignment, tensor analysis and data fusion, and a meaning mining model is constructed for processing.
It realizes a more comprehensive reflection of the user's physiological status and behavioral characteristics, provides richer information and higher recognition accuracy, eliminates data noise and inconsistency, and improves the accuracy and efficiency of recognition.
Smart Images

Figure CN120386451A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and more particularly to a multi-modal data analysis method and system based on tensor decomposition. Background Art
[0002] At present, the recognition of users' emotions and intentions is an important research direction in the fields of artificial intelligence and human-computer interaction, and is of great significance for improving the user experience and the intelligence of applications. The recognition of emotions and intentions based on multi-modal user physiological data can capture the user's state more accurately due to the objectivity and diversity of its data sources, and has become a research hotspot in recent years.
[0003] However, since people will hide their emotions in certain situations, there will be a low credibility in identifying a person's current emotions based on information such as facial expressions and voice-based emotion recognition.
[0004] Therefore, how to provide a multi-modal data analysis method that can solve the above problems is an urgent problem to be solved by those skilled in the art. Summary of the Invention
[0005] In view of this, the present invention provides a multi-modal data analysis method and system based on tensor decomposition. By acquiring and processing various types of multi-modal physiological signal data (image data, voice data, electroencephalogram signal data), it can more comprehensively reflect the user's physiological state and behavioral characteristics, and provide richer information and higher recognition accuracy.
[0006] In order to achieve the above object, the present invention adopts the following technical solutions:
[0007] A multi-modal data analysis method based on tensor decomposition, comprising the following steps:
[0008] Acquire various types of multi-modal physiological signal data of a user, and preprocess the multi-modal physiological signal data;
[0009] Align the modalities of the preprocessed multi-modal physiological signal data;
[0010] Perform tensor analysis and data fusion on the multi-modal physiological signal data after modality alignment to obtain a corresponding data fusion result;
[0011] Construct a meaning mining model, input the data fusion result into the meaning mining model for processing, and obtain a corresponding physiological signal recognition result.
[0012] Preferably, the multi-modal physiological signal data specifically includes: the user's image data, voice data, and electroencephalogram signal data.
[0013] Preferably, the specific process of performing modality alignment on the preprocessed multimodal physiological signal data includes:
[0014] Performing frequency-domain signal conversion on the preprocessed image data, speech data, and electroencephalogram signal data to obtain corresponding conversion results;
[0015] Performing spatial mapping on the conversion results through a machine learning method to complete the alignment.
[0016] Preferably, the specific process of obtaining the corresponding data fusion result includes:
[0017] Respectively performing feature extraction on the image data, speech data, and electroencephalogram signal data after modality alignment to obtain corresponding image features, speech features, and electroencephalogram features;
[0018] Performing tensor representation on the image features, speech features, and electroencephalogram features to obtain corresponding image tensors, speech tensors, and electroencephalogram tensors;
[0019] Fusing the image tensor, speech tensor, and electroencephalogram tensor to obtain the corresponding data fusion result.
[0020] Preferably, the specific process of fusing the image tensor, speech tensor, and electroencephalogram tensor includes:
[0021] Concatenating the image tensor, speech tensor, and electroencephalogram tensor;
[0022] Fusing the concatenation result through a neural network to obtain the corresponding data fusion result.
[0023] The present invention also provides a system for a multimodal data analysis method based on tensor decomposition, which is characterized by including:
[0024] An acquisition module, configured to acquire various types of multimodal physiological signal data of a user and preprocess the multimodal physiological signal data, where the physiological signal data includes the user's image data, speech data, and electroencephalogram signal data;
[0025] An alignment module, configured to perform modality alignment on the preprocessed multimodal physiological signal data;
[0026] A fusion module, configured to perform tensor analysis and data fusion on the multimodal physiological signal data after modality alignment to obtain the corresponding data fusion result;
[0027] A processing module, configured to construct a meaning mining model, input the data fusion result into the meaning mining model for processing, and obtain the corresponding physiological signal recognition result.
[0028] Preferably, the alignment module includes:
[0029] A conversion unit for performing frequency-domain signal conversion on the preprocessed image data, voice data, and electroencephalogram signal data to obtain corresponding conversion results;
[0030] An alignment unit for performing spatial mapping on the conversion results through a machine learning method to complete alignment.
[0031] Preferably, the fusion module includes:
[0032] A feature extraction unit for respectively extracting features from the image data, voice data, and electroencephalogram signal data that have undergone modality alignment to obtain corresponding image features, voice features, and electroencephalogram features;
[0033] A tensor representation unit for tensorizing the image features, voice features, and electroencephalogram features to obtain corresponding image tensors, voice tensors, and electroencephalogram tensors;
[0034] A fusion unit for fusing the image tensors, voice tensors, and electroencephalogram tensors to obtain corresponding data fusion results.
[0035] Through the above technical solutions, compared with the prior art, the present invention discloses a multi-modal data analysis method and system based on tensor decomposition, which has the following beneficial effects:
[0036] 1. By acquiring and processing various types of multi-modal physiological signal data (such as image data, voice data, electroencephalogram signal data), the present invention can more comprehensively reflect the physiological state and behavioral characteristics of users. Compared with single-modal data analysis, multi-modal data fusion can provide richer information and higher recognition accuracy.
[0037] 2. By preprocessing and aligning the multi-modal physiological signal data, the present invention can eliminate noise and inconsistencies in the data, improving the accuracy and reliability of subsequent analysis. Frequency-domain signal conversion and spatial mapping techniques help to convert data of different modalities into the same space or frequency domain, facilitating subsequent data fusion and analysis.
[0038] 3. The present invention extracts features and performs tensor representation on the multi-modal physiological signal data that has undergone modality alignment, and then fuses these tensors. Tensor decomposition can efficiently process high-dimensional data, mine potential relationships and patterns between data, thereby obtaining more accurate data fusion results.
[0039] 4. By constructing a meaning mining model and inputting the fused data into the model for processing, the corresponding physiological signal recognition results can be obtained. This intelligent recognition method can automatically learn and adapt to data changes, improving the accuracy and efficiency of recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other accompanying drawings can be obtained based on the provided drawings without creative efforts.
[0041] Figure 1 FIG. is the overall flowchart of a multi-modal data analysis method based on tensor decomposition provided by the present invention;
[0042] Figure 2 FIG. is the structural schematic diagram of a multi-modal data analysis system based on tensor decomposition provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0044] See Figure 1 As shown, the embodiments of the present invention disclose a multi-modal data analysis method based on tensor decomposition, including the following steps:
[0045] Obtain various types of multi-modal physiological signal data of a user and preprocess the multi-modal physiological signal data, where the preprocessing process may include steps such as filtering and removing abnormal data;
[0046] Align the modalities of the preprocessed multi-modal physiological signal data;
[0047] Perform tensor analysis and data fusion on the multi-modal physiological signal data after modality alignment to obtain corresponding data fusion results;
[0048] Construct a meaning mining model, input the data fusion results into the meaning mining model for processing, and obtain corresponding physiological signal recognition results.
[0049] Specifically, the specific structure of the meaning mining model can be a fusion model of XGBoost and Support Vector Machine (SVM). Considering adopting a multi-model fusion strategy, multiple models of different types are combined to improve the accuracy of recognition.
[0050] In a specific embodiment, the multi-modal physiological signal data specifically includes: the user's image data, voice data, and electroencephalogram signal data.
[0051] In a specific embodiment, the specific process of performing modal alignment on the preprocessed multi-modal physiological signal data includes:
[0052] Perform frequency-domain signal conversion on the preprocessed image data, voice data, and electroencephalogram signal data to obtain the corresponding conversion results;
[0053] Perform spatial mapping on the conversion results through machine learning methods to complete the alignment.
[0054] Specifically, after completing the multi-modal data alignment, the embodiment of the present invention can calculate the correlation or consistency index between different modal data, evaluate the effect of modal alignment, ensure the accuracy of subsequent data fusion, and re-perform data alignment when the threshold requirements are not met.
[0055] In a specific embodiment, the specific process of obtaining the corresponding data fusion result includes:
[0056] Extract features from the image data, voice data, and electroencephalogram signal data that have undergone modal alignment respectively to obtain the corresponding image features, voice features, and electroencephalogram features;
[0057] Perform tensor representation on the image features, voice features, and electroencephalogram features to obtain the corresponding image tensors, voice tensors, and electroencephalogram tensors;
[0058] Fuse the image tensors, voice tensors, and electroencephalogram tensors to obtain the corresponding data fusion result.
[0059] In a specific embodiment, the specific process of fusing the image tensors, voice tensors, and electroencephalogram tensors includes:
[0060] Concatenate the image tensors, voice tensors, and electroencephalogram tensors;
[0061] Fuse the concatenation result through a neural network to obtain the corresponding data fusion result.
[0062] See Figure 2 As shown, the embodiment of the present invention also provides a system for a multi-modal data analysis method based on tensor decomposition, including:
[0063] An acquisition module, configured to acquire multi-modal physiological signal data of various types of a user, and preprocess the multi-modal physiological signal data, where the physiological signal data includes the user's image data, voice data, and electroencephalogram signal data;
[0064] An alignment module, configured to perform modality alignment on the preprocessed multi-modal physiological signal data;
[0065] A fusion module, configured to perform tensor analysis and data fusion on the multi-modal physiological signal data after modality alignment to obtain a corresponding data fusion result;
[0066] A processing module, configured to construct a meaning mining model, input the data fusion result into the meaning mining model for processing, and obtain a corresponding physiological signal recognition result.
[0067] In a specific embodiment, the alignment module includes:
[0068] A conversion unit, configured to perform frequency-domain signal conversion on the preprocessed image data, voice data, and electroencephalogram signal data to obtain corresponding conversion results;
[0069] An alignment unit, configured to perform spatial mapping on the conversion results through a machine learning method to complete the alignment.
[0070] In a specific embodiment, the fusion module includes:
[0071] A feature extraction unit, configured to respectively extract features from the image data, voice data, and electroencephalogram signal data after modality alignment to obtain corresponding image features, voice features, and electroencephalogram features;
[0072] A tensor representation unit, configured to perform tensor representation on the image features, voice features, and electroencephalogram features to obtain corresponding image tensors, voice tensors, and electroencephalogram tensors;
[0073] A fusion unit, configured to fuse the image tensors, voice tensors, and electroencephalogram tensors to obtain a corresponding data fusion result.
[0074] In this specification, each embodiment is described in a progressive manner. The key point of each embodiment is the difference from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0075] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A multi-modal data analysis method based on tensor decomposition, characterized in that It includes the following steps: Obtain multi-modal physiological signal data of various types of users, and preprocess the multi-modal physiological signal data; Perform modal alignment on the preprocessed multi-modal physiological signal data; Perform tensor analysis and data fusion on the multi-modal physiological signal data after modal alignment to obtain corresponding data fusion results; Construct a meaning mining model, input the data fusion results into the meaning mining model for processing, and obtain corresponding physiological signal recognition results.
2. The multimodal data analysis method based on tensor decomposition according to claim 1, wherein The multi-modal physiological signal data specifically includes: image data, voice data, and electroencephalogram signal data of users.
3. A multimodal data analysis method based on tensor decomposition according to claim 2, characterized in that, The specific process of performing modal alignment on the preprocessed multi-modal physiological signal data includes: Perform frequency-domain signal conversion on the preprocessed image data, voice data, and electroencephalogram signal data to obtain corresponding conversion results; Perform spatial mapping on the conversion results through machine learning methods to complete the alignment.
4. A multimodal data analysis method based on tensor decomposition according to claim 2, characterized in that The specific process of obtaining corresponding data fusion results includes: Extract features from the image data, voice data, and electroencephalogram signal data after modal alignment respectively to obtain corresponding image features, voice features, and electroencephalogram features; Perform tensor representation on the image features, voice features, and electroencephalogram features to obtain corresponding image tensors, voice tensors, and electroencephalogram tensors; Fuse the image tensors, voice tensors, and electroencephalogram tensors to obtain corresponding data fusion results.
5. A multimodal data analysis method based on tensor decomposition according to claim 4, characterized in that The specific process of fusing the image tensors, voice tensors, and electroencephalogram tensors includes: Concatenate the image tensors, voice tensors, and electroencephalogram tensors; Fuse the concatenation results through a neural network to obtain corresponding data fusion results.
6. A system for a multi-modal data analysis method based on tensor decomposition, characterized in that, It includes: An acquisition module, configured to obtain multi-modal physiological signal data of various types of users, and preprocess the multi-modal physiological signal data, where the physiological signal data includes image data, voice data, and electroencephalogram signal data of users; An alignment module, configured to perform modal alignment on the preprocessed multi-modal physiological signal data; A fusion module, configured to perform tensor analysis and data fusion on the multi-modal physiological signal data after modal alignment to obtain corresponding data fusion results; A processing module, configured to construct a meaning mining model, input the data fusion results into the meaning mining model for processing, and obtain corresponding physiological signal recognition results.
7. The system of a multi-modal data analysis method based on tensor decomposition according to claim 6, wherein, The alignment module includes: A conversion unit, configured to perform frequency-domain signal conversion on the preprocessed image data, voice data, and electroencephalogram signal data to obtain corresponding conversion results; An alignment unit, configured to perform spatial mapping on the conversion results through machine learning methods to complete the alignment.
8. The system of a multi-modal data analysis method based on tensor decomposition according to claim 6, wherein The fusion module includes: A feature extraction unit, configured to extract features from the image data, voice data, and electroencephalogram signal data after modal alignment respectively to obtain corresponding image features, voice features, and electroencephalogram features; A tensor representation unit, configured to perform tensor representation on the image features, voice features, and electroencephalogram features to obtain corresponding image tensors, voice tensors, and electroencephalogram tensors; A fusion unit, configured to fuse the image tensor, the speech tensor, and the electroencephalogram tensor to obtain a corresponding data fusion result.