Multi-modal data acquisition and feature fusion system, method and server
Through the multimodal data acquisition and feature fusion system, the fusion processing of multimodal data is performed using neural networks and self-supervised learning strategies, the problems of insufficient multimodal data processing, feature fusion and cross-modal correlation in the existing technology are solved, better correlation capture and heterogeneity processing are achieved, and the performance of multimodal tasks is improved.
Patent Information
- Application Number
- CN202510573755.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-06-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When processing multimodal data, the prior art cannot fully utilize the complementary information between different modes, resulting in insufficient features extraction, cross-modal correlation and training optimization, making it difficult to achieve ideal results.
It provides a multimodal data acquisition and feature fusion system, including a data acquisition and feature extraction module, a cross-modal data sampling synchronization module, a data feature fusion and optimization module, and a self-supervised learning and cross-modal training module. The fusion processing of multimodal data is carried out through neural networks, and a self-supervised learning strategy is used to automatically generate labels or loss functions for cross-modal training.
Through innovative cross-modal data synchronization mechanism and feature fusion strategy, it can better capture the potential correlation between different modes, effectively handle the heterogeneity of text, audio and image data, improve the generalization and robustness of multimodal tasks, reduce dependence on manual labeled data, and improve processing efficiency, model training speed and accuracy.
Smart Images

Figure CN120105346A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of multimodal data processing, and in particular to a multimodal data acquisition and feature fusion system, method and server. Background Art
[0002] With the rapid development of artificial intelligence technology, multimodal data can provide more comprehensive information. How to effectively process and integrate these heterogeneous data is a major challenge in multimodal learning.
[0003] At present, traditional unimodal data analysis methods often perform poorly when processing multimodal data due to the inability to fully utilize the complementary information between different modalities. Multimodal data has obvious heterogeneity, and the differences make it a technical challenge to fuse them into a unified representation. In addition, multimodal data usually comes from different sources or time periods. How to align and ensure the relevance of information in the same time dimension is a prerequisite for effective fusion. Although the prior art has proposed a variety of multimodal data fusion methods, most methods still have deficiencies in feature extraction, cross-modal association, and training optimization, which makes it difficult for the model to achieve ideal results in complex applications.
[0004] Therefore, there is an urgent need for a multimodal data fusion solution that can solve the above problems. Summary of the invention
[0005] The present application provides a multimodal data acquisition and feature fusion system, method and server to solve the problems of insufficient multimodal data processing, feature fusion and cross-modal correlation in the prior art.
[0006] The technical effects to be achieved by this application are achieved through the following solutions: According to a first aspect of the present application, a multimodal data acquisition and feature fusion system is provided, comprising: Data collection and feature extraction module, used for collecting multimodal data and extracting features; The cross-modal data sampling synchronization module is used to synchronously sample the collected multi-modal data at different time points, so that the multi-modal data at the same time point can complement each other and be effectively synchronized; The data feature fusion and optimization module uses a neural network to fuse the synchronized multimodal data and generate the same multimodal feature representation; The self-supervised learning and cross-modal training module adopts a self-supervised learning strategy to automatically generate labels or loss functions for cross-modal training.
[0007] Preferably, in the data acquisition and feature extraction module, the multimodal data includes text modality data, audio modality data and image modality data.
[0008] Preferably, feature extraction of multimodal data includes: for multimodal data collected at the same time, text modal data performs feature extraction of text information of subtitles and dialogues through words to generate text feature vectors; audio modal data extracts features of speech content and background music through spectrum analysis to generate audio feature vectors; image modal data performs feature extraction of people and objects in video frames and images through convolutional neural networks to generate image feature vectors.
[0009] Preferably, in the cross-modal data sampling synchronization module, time marking and event synchronization mechanisms are used to compare and fuse multimodal data; The time tag is specifically: when collecting multimodal data, a timestamp in a unified format is assigned to each data frame or event, and the timestamp is a unified time reference generated based on GPS or a system synchronous clock, which is used to achieve time alignment of cross-modal data; The event synchronization mechanism is specifically as follows: constructing cross-modal synchronization anchor points through key events in multimodal data streams, combining sliding window strategy with similarity calculation method, dynamically identifying time matching intervals between modalities, and establishing event alignment graphs to assist the time synchronization process; During the sampling process, a multimodal mechanism is used to ensure the correlation of each modal data.
[0010] Preferably, in the data feature fusion and optimization module, the following formula is used to perform weighted fusion on the text modality data, the audio modality data and the image modality data: ; Among them, F is the multimodal feature, T is the text feature vector, A is the audio feature vector, V is the image feature vector, , are the weights of text feature vector, audio feature vector and image feature vector respectively.
[0011] Preferably, in the data feature fusion and optimization module, the fusion features are further optimized through a multi-layer neural network to improve the distinguishability and robustness of the features, specifically through the following formula: ; in, is the initial fusion feature of the input; , used to extract global semantic associations; Represents a feedforward neural network structure; Representation layer normalization operation; It is the optimized fusion feature of the final output.
[0012] Preferably, in the self-supervised learning and cross-modal training module, a cross-modal loss function is designed to measure the correlation between the features of two different modalities, specifically: ; Among them, A i and B i are two different modal features at the same time i, represents the Euclidean distance, To measure the loss value of the difference or similarity between two different modal features, The smaller the value of , the stronger the correlation between the two modes, and vice versa.
[0013] Preferably, contrastive learning is used to further enhance cross-modal associations, and the following loss function is used for optimization: ; in: represents the similarity function, is the temperature parameter, A i and B i They are two different modal features at the same time i.
[0014] According to a second aspect of the present application, a multimodal data acquisition and feature fusion method using the above-mentioned multimodal data acquisition and feature fusion system is provided, comprising the following steps: Step 1: Collect multimodal data and extract features; Step 2: synchronously sample the collected multimodal data at the time point, so that the multimodal data at the same time point complement each other and are effectively synchronized; Step 3: Use a neural network to fuse the synchronized multimodal data to generate the same multimodal feature representation; Step 4: Use a self-supervised learning strategy to automatically generate labels or loss functions for cross-modal training.
[0015] According to a third aspect of the present application, a server includes: a memory and at least one processor; The memory stores a computer program, and the at least one processor executes the computer program stored in the memory to implement the above-mentioned multimodal data collection and feature fusion method.
[0016] According to an embodiment of the present application, the beneficial effects of adopting the present multimodal data acquisition and feature fusion system are: through the innovative cross-modal data synchronization mechanism and feature fusion strategy, the potential correlation between different modalities can be better captured; Compared with traditional methods, this invention can more effectively handle the heterogeneity of text, audio and image data; by combining deep learning networks with self-supervised learning strategies, it improves the generalization and robustness of multimodal tasks, and has stronger adaptability in various practical applications, especially in cross-modal analysis and reasoning tasks; Through the self-supervised learning mechanism, the reliance on manually labeled data is reduced, which significantly improves the efficiency of multimodal data processing. The model can automatically generate labels and loss functions to optimize the training process, thereby improving the training speed and accuracy of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the embodiments of the present application or the existing technical solutions, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0018] Figure 1 This is a structural block diagram of a multimodal data acquisition and feature fusion system in one embodiment of the present application; Figure 2 This is a flow chart of a multimodal data collection and feature fusion method in one embodiment of the present application; Figure 3 This is a structural block diagram of a server in an embodiment of the present application. DETAILED DESCRIPTION
[0019] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the present application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without creative work are within the scope of protection of the present application.
[0020] like Figure 1 As shown, a multimodal data acquisition and feature fusion system in one embodiment of the present application includes: Data collection and feature extraction module, used for collecting multimodal data and extracting features; In this module, multimodal data includes text modal data, audio modal data and image modal data. Among them, text modal data includes subtitles, dialogues, etc., audio modal data includes voice content, background music, etc., and image modal data includes image video frames, character images, objects, etc.
[0021] In a specific example, text information is extracted from veterinarians’ diagnosis records, medical reports, and breeding logs, such as descriptions of animal symptoms, treatment plans, and medication records. Veterinarians’ diagnostic voice records, animal calls, and environmental background sounds are collected. Veterinarians’ descriptions of the animal’s condition, abnormal calls made by animals due to discomfort, etc. Animal physical signs images are obtained, including appearance photos, behavioral video frames, and close-up images of specific parts (such as skin lesions, eye abnormalities, etc.).
[0022] When in use, text information is extracted from the subtitles or speech-to-text data in the video, the audio signal in the video is analyzed, and the visual features in the video frame are extracted using image recognition technology.
[0023] Then, the text modality data uses words to extract features from the text information of subtitles and dialogues to generate text feature vectors; the audio modality data uses spectrum analysis to extract features of speech content and background music to generate audio feature vectors; the image modality data uses convolutional neural networks to extract features of people and objects in video frames and images to generate image feature vectors.
[0024] Finally, the information of the above modes is combined to ensure that the extracted features can be matched at the same time to form a unified multimodal data representation.
[0025] The cross-modal data sampling synchronization module is used to synchronously sample the collected multi-modal data at different time points, so that the multi-modal data at the same time point can complement each other and be effectively synchronized; In this module, data from different modalities (text, audio, image) are sampled synchronously at the same time point to ensure that the text, audio and image information at the same time point can complement each other and be effectively synchronized. Through specific time marking and event synchronization mechanisms, it is ensured that all modal information can be effectively compared and integrated within the same time period, and that the correlation between the modal data will not be lost during the sampling process.
[0026] Among them, the time tag refers to the allocation of a high-precision timestamp in a unified format to each data frame or event when collecting various modal data. The timestamp is based on a unified time reference generated by GPS or a system synchronization clock, and is used to achieve time alignment of cross-modal data; the event synchronization mechanism refers to the construction of cross-modal synchronization anchor points through key events in multimodal data streams, combining sliding window strategies with similarity calculation methods, dynamically identifying time matching intervals between modalities, and establishing an event alignment map to assist the time synchronization process; During the sampling process, a multimodal mechanism is adopted to ensure the relevance of each modal data. The multimodal mechanism establishes a semantic association channel between modalities and combines a cross-modal attention mechanism or a contrastive learning strategy to achieve modal alignment and weight distribution at the feature level. The retention rate of effective information is improved through modal consistency detection and noise modal rejection strategies, thereby ensuring the integrity and coordination of multimodal information in the subsequent feature fusion stage and avoiding the loss of key semantics during the fusion process.
[0027] This module processes data of different modalities through time acquisition, so that data such as text, audio and image are aligned on the time axis and effectively synchronized. Timestamps are added to each text record, audio clip and image frame to ensure that data at the same time point can correspond. For data with incomplete timestamp matching, interpolation methods are used for alignment to provide a basis for subsequent feature fusion.
[0028] For example, the voice record of a veterinarian describing an animal's symptoms at a certain moment needs to be synchronized with the animal's vital signs image taken at that moment. When the time point of the image frame is not completely consistent with the voice record, the feature representation of the corresponding time point is generated through interpolation. When the text record of a certain time point is missing, supplementary text features are generated through contextual information.
[0029] The data feature fusion and optimization module uses a neural network to fuse the synchronized multimodal data and generate the same multimodal feature representation; In this module, deep learning technology is used to fuse the features of each modality. The synchronized data is weighted and merged through a neural network to generate a unified multimodal feature representation. The fused features are optimized through a multi-layer neural network to improve the distinguishability and robustness of the features. For example, a fully connected layer is used to reduce the dimension and denoise the fused features to generate a more compact feature representation.
[0030] Specifically, the Transformer model is used to capture the potential semantic relationship between text and vision, and to combine image features with text features; the audio-joint model is used to fuse audio features with image feature waveforms, improving the video's ability in emotional expression, semantic consistency, event recognition, and scene restoration, so that the visual content and audio content complement each other semantically and support each other in expression, forming a more complete and coherent multimodal information flow. In order to enable the model to better handle the heterogeneity of cross-modal data, the fused features are optimized through the back propagation algorithm to improve the accuracy and robustness of the fused features.
[0031] The following formula is used to perform weighted fusion on text modality data, audio modality data, and image modality data:
[0032] Among them, F is the multimodal feature, T is the text feature vector, A is the audio feature vector, V is the image feature vector, , are the weights of text feature vector, audio feature vector and image feature vector respectively.
[0033] In this module, the fusion features are further optimized through a multi-layer neural network:
[0034] in is the initial fusion feature of the input; , used to extract global semantic associations; Represents a feedforward neural network structure; Representation layer normalization operation; The optimized fusion features are finally output. The nonlinear enhancement, dimension compression and semantic alignment of features are realized, and the distinguishability and generalization ability of fusion features are improved.
[0035] The self-supervised learning and cross-modal training module adopts a self-supervised learning strategy to automatically generate labels or loss functions for cross-modal training.
[0036] In this module, labels or loss functions are automatically generated by the model without relying on manual supervision data, making the multimodal feature fusion process more efficient. Self-supervised learning enables the model to self-learn the correlation between different mode data by constructing cross-modal pre-training tasks. During the training process, the features between different modes are associated through a common embedding space, so that the system can automatically learn the intrinsic connection between text, audio, and image modes, and optimize the feature fusion effect.
[0037] The model automatically generates pseudo labels for cross-modal association tasks and constructs positive and negative sample pairs. The model is trained to distinguish the associations between different modalities, a cross-modal loss function is designed, and the multimodal fusion capability of the model is optimized.
[0038] In the self-supervised learning and cross-modal training module, a cross-modal loss function is designed to measure the correlation between the features of two different modalities. Specifically: ; Among them, A i and B i are two different modal features at the same time i, Represents the Euclidean distance, which is used to measure the similarity between text and image features. It is a loss value that measures the difference or similarity between two different modal features. The smaller the value of , the stronger the correlation between the two modes, and vice versa.
[0039] In a specific example, let the text modality feature be , the image modality feature is V, then the cross-modal loss function can be defined as: ; In addition, contrastive learning is used to further enhance cross-modal associations, and the following loss function is used for optimization: ; in: represents the similarity function, is the temperature parameter, A i and B i They are two different modal features at the same time i.
[0040] In a specific example, let the text modality feature be , the image modality feature is V, and the loss function is: ; Through the self-modeling learning supervision strategy, the model can automatically learn the potential relationships between different modalities, thereby improving the representation ability of multimodal data.
[0041] like Figure 2 As shown, in one embodiment of the present application, a multimodal data acquisition and feature fusion method using the above-mentioned multimodal data acquisition and feature fusion system is provided, comprising the following steps: Step 1: Collect multimodal data and extract features; Step 2: synchronously sample the collected multimodal data at the time point, so that the multimodal data at the same time point complement each other and are effectively synchronized; Step 3: Use a neural network to fuse the synchronized multimodal data to generate the same multimodal feature representation; Step 4: Use a self-supervised learning strategy to automatically generate labels or loss functions for cross-modal training.
[0042] like Figure 3 As shown, a server in an embodiment of the present application includes: a memory 301 and at least one processor 302; The memory 301 stores computer programs, and the at least one processor 302 executes the computer programs stored in the memory 301 to implement the above-mentioned multimodal data collection and feature fusion method.
[0043] According to an embodiment of the present application, the beneficial effects of adopting the present multimodal data acquisition and feature fusion system are: through the innovative cross-modal data synchronization mechanism and feature fusion strategy, the potential correlation between different modalities can be better captured; Compared with traditional methods, this invention can more effectively handle the heterogeneity of text, audio and image data; by combining deep learning networks with self-supervised learning strategies, it improves the generalization and robustness of multimodal tasks, and has stronger adaptability in various practical applications, especially in cross-modal analysis and reasoning tasks; Through the self-supervised learning mechanism, the reliance on manually labeled data is reduced, which significantly improves the efficiency of multimodal data processing. The model can automatically generate labels and loss functions to optimize the training process, thereby improving the training speed and accuracy of the model.
[0044] It should be noted that the above detailed descriptions are exemplary and are intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the art to which the present application belongs.
[0045] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should also be understood that when the terms "comprise" and / or "include" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.
[0046] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein.
[0047] In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or inherent to these processes, methods, products, or apparatuses.
[0048] For ease of description, spatially relative terms, such as "above", "above", "on the upper surface of", "above", etc., may be used herein to describe the spatial positional relationship between a device or feature and other devices or features as shown in the figure. It should be understood that spatially relative terms are intended to include different orientations of the device in use or operation in addition to the orientation described in the figure. For example, if the device in the accompanying drawings is inverted, the device described as "above other devices or structures" or "above other devices or structures" will be positioned as "below other devices or structures" or "below other devices or structures". Thus, the exemplary term "above" may include both "above" and "below". The device may also be positioned in other different ways, such as rotated 90 degrees or in other orientations, and the spatially relative descriptions used herein are interpreted accordingly.
[0049] In the above detailed description, reference is made to the accompanying drawings, which form a part of this document. In the accompanying drawings, similar symbols typically identify similar components unless the context indicates otherwise. The illustrated embodiments described in the detailed description, drawings, and claims are not meant to be limiting. Other embodiments may be used, and other changes may be made, without departing from the spirit or scope of the subject matter presented herein.
[0050] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A multimodal data acquisition and feature fusion system, characterized in that: include: Data collection and feature extraction module, used for collecting multimodal data and extracting features; The cross-modal data sampling synchronization module is used to synchronously sample the collected multi-modal data at different time points, so that the multi-modal data at the same time point can complement each other and be effectively synchronized; The data feature fusion and optimization module uses a neural network to fuse the synchronized multimodal data and generate the same multimodal feature representation; The self-supervised learning and cross-modal training module adopts a self-supervised learning strategy to automatically generate labels or loss functions for cross-modal training.
2. The multimodal data acquisition and feature fusion system according to claim 1, characterized in that: In the data collection and feature extraction module, the multimodal data includes text modality data, audio modality data and image modality data.
3. The multimodal data acquisition and feature fusion system according to claim 2, characterized in that: Feature extraction of multimodal data includes: for multimodal data collected at the same time, text modal data uses words to extract features from subtitles and text information of dialogues to generate text feature vectors; audio modal data uses spectrum analysis to extract features of speech content and background music to generate audio feature vectors; image modal data uses convolutional neural networks to extract features from people and objects in video frames and images to generate image feature vectors.
4. The multimodal data acquisition and feature fusion system according to claim 1, characterized in that: In the cross-modal data sampling synchronization module, time marking and event synchronization mechanisms are used to compare and fuse multimodal data; The time tag is specifically: when collecting multimodal data, a timestamp in a unified format is assigned to each data frame or event, and the timestamp is a unified time reference generated based on GPS or a system synchronous clock, which is used to achieve time alignment of cross-modal data; The event synchronization mechanism is specifically as follows: constructing cross-modal synchronization anchor points through key events in multimodal data streams, combining sliding window strategy with similarity calculation method, dynamically identifying time matching intervals between modalities, and establishing event alignment graphs to assist the time synchronization process; During the sampling process, a multimodal mechanism is used to ensure the correlation of each modal data.
5. The multimodal data acquisition and feature fusion system according to claim 1, characterized in that: In the data feature fusion and optimization module, the following formula is used to perform weighted fusion on text modality data, audio modality data, and image modality data: ; Among them, F is the multimodal feature, T is the text feature vector, A is the audio feature vector, V is the image feature vector, , are the weights of text feature vector, audio feature vector and image feature vector respectively.
6. The multimodal data acquisition and feature fusion system according to claim 5, characterized in that: In the data feature fusion and optimization module, the fusion features are further optimized through a multi-layer neural network to improve the distinguishability and robustness of the features. Specifically, the following formula is used: ; in, is the initial fusion feature of the input; , used to extract global semantic associations; Represents a feedforward neural network structure; Representation layer normalization operation; It is the optimized fusion feature of the final output.
7. The multimodal data acquisition and feature fusion system according to claim 5, characterized in that: In the self-supervised learning and cross-modal training module, a cross-modal loss function is designed to measure the correlation between the features of two different modalities. Specifically: ; Among them, A i and B i are two different modal features at the same time i, represents the Euclidean distance, To measure the loss value of the difference or similarity between two different modal features, The smaller the value of , the stronger the correlation between the two modes, and vice versa.
8. The multimodal data acquisition and feature fusion system according to claim 6, characterized in that: Contrastive learning is used to further enhance cross-modal associations, and the following loss function is used for optimization: ; in: represents the similarity function, is the temperature parameter, A i and B i They are two different modal features at the same time i, and n is the number of time points sampled for a single modal feature.
9. A multimodal data acquisition and feature fusion method using the multimodal data acquisition and feature fusion system according to any one of claims 1 to 8, characterized in that: The steps include: Step 1: Collect multimodal data and extract features; Step 2: synchronously sample the collected multimodal data at the time point, so that the multimodal data at the same time point complement each other and are effectively synchronized; Step 3: Use a neural network to fuse the synchronized multimodal data to generate the same multimodal feature representation; Step 4: Use a self-supervised learning strategy to automatically generate labels or loss functions for cross-modal training.
10. A server, characterized in that: include: memory and at least one processor; The memory stores a computer program, and the at least one processor executes the computer program stored in the memory to implement the multimodal data acquisition and feature fusion method described in claim 9.
Citation Information
Patent Citations
Multi-modal data driven attention rehabilitation feedback method and system
CN117731287A
Remote sensing cross-modal retrieval method based on deep ternary fusion sensing network
CN119336968A
Cited By
Event aggregated short video information detection method
CN120599525A
Energy storage power station key electrical equipment fault diagnosis method and system based on multi-mode deep learning
CN121051429A