Multi-modal data alignment method and device, electronic equipment and storage medium

Through comparative learning and fusion feature construction, the training parameters of the feature extraction model are optimized, and the problem of poor alignment of multimodal data is solved, and efficient alignment and semantic association of modal features are achieved.

CN120541542APending Publication Date: 2025-08-26BEIJING ZHONGKE JINDEZHU INTELLIGENT TECH CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510643975.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

In the prior art, the alignment effect of multimodal data is poor, resulting in information loss or deviation, and it is impossible to effectively associate data of different modalities.

Method used

By inputting the multimodal training session to feature extraction model, comparative learning and fusion feature construction, the training parameters are adjusted using the first loss and the second loss, and the feature extraction model is optimized to improve the alignment effect of the modal features.

Benefits of technology

The effective alignment of multimodal data is realized, the semantic correlation and feature matching between modal features are improved, and the execution ability of cross-modal tasks is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120541542A_ABST
    Figure CN120541542A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal data alignment method and device, electronic equipment and a storage medium, and the method comprises the steps: inputting a multi-modal training session into a to-be-trained feature extraction model, and obtaining a modal feature of each modal data in the multi-modal training session; performing comparative learning according to the modal features of each modal data in the multi-modal training session to obtain a first loss; according to the modal feature of each modal data in the multi-modal training session, constructing a fusion feature, and inputting the fusion feature into a preset question and answer model to obtain a question answer corresponding to the multi-modal training session; semantic features of the answers to the questions are extracted, the semantic features are compared with the fusion features, and second losses are obtained; and according to the first loss and the second loss, adjusting training parameters of a feature extraction model to obtain a trained feature extraction model, and extracting a plurality of mutually aligned modal features of the multi-modal real-time session through the feature extraction model. According to the invention, the feature alignment effect of the multi-modal data can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to a method and device for aligning multimodal data, an electronic device, and a computer-readable storage medium. Background Art

[0002] Alignment of multimodal data refers to mapping data of different modalities into the same semantic space so that the modal features of different modalities can be correlated with each other.

[0003] In related technologies, due to the obvious heterogeneity of each modal data itself, information loss or deviation often occurs when aligning the features of each modal data, making it impossible to guarantee the alignment effect of multimodal data. Summary of the Invention

[0004] The present disclosure provides a method and device for aligning multimodal data, an electronic device, and a computer-readable storage medium, which can effectively improve the alignment effect of multimodal data.

[0005] In a first aspect, the present disclosure provides a method for aligning multimodal data, which includes: inputting a multimodal training session into a feature extraction model to be trained to obtain modal features of each modal data in the multimodal training session; performing comparative learning based on the modal features of each modal data in the multimodal training session to obtain a first loss; constructing a fusion feature based on the modal features of each modal data in the multimodal training session, inputting the fusion feature into a preset question-answering model to obtain an answer to a question corresponding to the multimodal training session; extracting semantic features of the question answer, comparing the semantic features with the fusion feature, and obtaining a second loss; adjusting the training parameters of the feature extraction model based on the first loss and the second loss to obtain a trained feature extraction model, so as to extract multiple mutually aligned modal features of the multimodal real-time session through the trained feature extraction model.

[0006] In a second aspect, the present disclosure provides a multimodal data alignment device, which includes: an extraction module for inputting a multimodal training session into a feature extraction model to be trained to obtain the modal features of each modal data in the multimodal training session; a comparison module for performing comparative learning based on the modal features of each modal data in the multimodal training session to obtain a first loss; a generation module for constructing a fusion feature based on the modal features of each modal data in the multimodal training session, inputting the fusion feature into a preset question-answering model to obtain an answer to a question corresponding to the multimodal training session; a comparison module for extracting semantic features of the question answer, comparing the semantic features with the fusion feature, and obtaining a second loss; a training module for adjusting the training parameters of the feature extraction model based on the first loss and the second loss to obtain a trained feature extraction model, so as to extract multiple mutually aligned modal features of the multimodal real-time session through the trained feature extraction model.

[0007] In a third aspect, the present disclosure provides an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores one or more computer programs executable by the at least one processor, and one or more of the computer programs are executed by the at least one processor to enable the at least one processor to perform the above-mentioned multimodal data alignment method.

[0008] In a fourth aspect, the present disclosure provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the above-mentioned multimodal data alignment method when executed by a processor.

[0009] The multimodal data alignment method provided by the embodiment of the present disclosure first obtains the modal features of each modal data in the multimodal training session by inputting the multimodal training session into the feature extraction model to be trained. Secondly, comparative learning is performed based on the modal features of each modal data in the multimodal training session to obtain a first loss. Then, based on the modal features of each modal data in the multimodal training session, a fusion feature is constructed, and the fusion feature is input into a preset question-answering model to obtain the answer to the question corresponding to the multimodal training session. The semantic features of the answer to the question are extracted and compared with the fusion feature to obtain a second loss. The first loss is used to measure the semantic alignment effect of the modal features of each modal data, and the loss of the first loss is used to measure the semantic alignment effect of the modal features of each modal data. The smaller the value, the closer the feature distance between the modal features of each modal data and the modal features with semantic correlation is. The second loss is used to measure the feature expression ability of the fusion feature constructed based on the modal features of each modal data. The smaller the loss value of the second loss, the higher the feature matching degree between the fusion feature and the semantic feature, that is, the stronger the feature semantic representativeness of each modal feature used to construct the fusion feature. Correspondingly, the comparative learning based on each modal feature achieves better feature alignment. Finally, according to the first loss and the second loss, the training parameters of the feature extraction model are jointly adjusted to obtain the trained feature extraction model, so as to extract multiple mutually aligned modal features of the multimodal real-time conversation through the trained feature extraction model.

[0010] It can be seen that when training the feature extraction model, the present disclosure can align the modal features of each modal data extracted by the feature extraction model through the comparative learning method of the first loss, and can make the feature matching degree between the fusion features of the multimodal training session and the semantic features of the question answer higher through the second loss, thereby optimizing the feature extraction ability of the feature extraction model for modal features, and then making the alignment effect between the modal features of each modal data better. Therefore, the present disclosure can extract multiple mutually aligned modal features of the multimodal real-time conversation based on the trained feature extraction model, thereby improving the alignment effect of the multimodal data.

[0011] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The accompanying drawings are used to provide a further understanding of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the present disclosure and do not constitute a limitation of the present disclosure. The above and other features and advantages will become more apparent to those skilled in the art by describing detailed example embodiments with reference to the accompanying drawings. In the accompanying drawings:

[0013] Figure 1 A diagram illustrating an application scenario of the multimodal data alignment method and apparatus provided in an embodiment of the present disclosure;

[0014] Figure 2 A flowchart of a multimodal data alignment method provided in an embodiment of the present disclosure;

[0015] Figure 3 A schematic diagram of a flow chart of a multimodal data alignment method provided in an embodiment of the present disclosure;

[0016] Figure 4 Schematic diagram of the training process of the feature extraction model;

[0017] Figure 5 A block diagram of a multimodal data alignment device provided in an embodiment of the present disclosure;

[0018] Figure 6 A block diagram of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0019] To enable those skilled in the art to better understand the technical solutions of the present disclosure, exemplary embodiments of the present disclosure are described below in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0020] In the absence of conflict, the various embodiments of the present disclosure and the various features therein may be combined with each other.

[0021] As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.

[0022] The terms used herein are only used to describe specific embodiments and are not intended to limit the present disclosure. As used herein, the singular forms "a" and "the" are also intended to include the plural forms, unless the context clearly indicates otherwise. It will also be understood that when the terms "comprising" and / or "made of" are used in this specification, the presence of the features, wholes, steps, operations, elements and / or components is specified, but the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or groups thereof is not excluded. Similar words such as "connected" or "connected" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect.

[0023] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and the present disclosure, and will not be interpreted as having an idealized or overly formal meaning unless expressly defined as such herein.

[0024] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals. The use of user data in this technical solution complies with relevant national laws and regulations (for example, the "Information Security Technology Personal Information Security Specification", etc.). For example: corresponding prescribed measures are taken to control access to personal information; the display of personal information is subject to prescribed restrictions; the purpose of using personal information does not exceed the scope of direct or reasonable connection; when using personal information, clear identity reference is eliminated to avoid precise positioning of specific individuals.

[0025] Multimodal data alignment involves mapping data from different modalities (such as text, images, and audio) into the same feature space or semantic space through algorithms or models, thereby enabling effective interaction and integration of information. Aligning multimodal data ensures that information across different modalities can be effectively linked and understood, enabling better cross-modal tasks such as intelligent question answering.

[0026] In related technologies, due to the obvious heterogeneity in feature distribution and information expression of multiple modal data such as text, images, audio, and video, the modal features extracted from multimodal data often result in important information loss or feature deviation, resulting in poor alignment between the modal features of each modal data.

[0027] In view of this, an embodiment of the present disclosure provides a method for aligning multimodal data. First, the modal features of each modal data in the multimodal training session are obtained by inputting the multimodal training session into the feature extraction model to be trained. Secondly, comparative learning is performed based on the modal features of each modal data in the multimodal training session to obtain a first loss. Then, a fusion feature is constructed based on the modal features of each modal data in the multimodal training session. The fusion feature is input into a preset question-answering model to obtain the answer to the question corresponding to the multimodal training session. The semantic features of the answer to the question are extracted and compared with the fusion feature to obtain a second loss. The first loss is used to measure the semantic alignment effect of the modal features of each modal data. The smaller the loss value, the closer the feature distance between the modal features of each modal data and the modal features with semantic correlation is. The second loss is used to measure the feature expression ability of the fusion feature constructed based on the modal features of each modal data. The smaller the loss value of the second loss, the higher the feature matching degree between the fusion feature and the semantic feature, that is, the stronger the feature semantic representativeness of each modal feature used to construct the fusion feature. Correspondingly, the comparative learning based on each modal feature achieves better feature alignment. Finally, according to the first loss and the second loss, the training parameters of the feature extraction model are jointly adjusted to obtain the trained feature extraction model, so as to extract multiple mutually aligned modal features of the multimodal real-time conversation through the trained feature extraction model.

[0028] It can be seen that when training the feature extraction model, the present disclosure can align the modal features of each modal data extracted by the feature extraction model through the comparative learning method of the first loss, and can make the feature matching degree between the fusion features of the multimodal training session and the semantic features of the question answer higher through the second loss, thereby optimizing the feature extraction ability of the feature extraction model for modal features, and then making the alignment effect between the modal features of each modal data better. Therefore, the present disclosure can extract multiple mutually aligned modal features of the multimodal real-time conversation based on the trained feature extraction model, thereby improving the alignment effect of the multimodal data.

[0029] Figure 1 A diagram illustrating an application scenario of the multimodal data alignment method and apparatus provided in an embodiment of the present disclosure.

[0030] like Figure 1 As shown, an application scenario of an embodiment of the present disclosure may include a terminal device 101, a network 103, and a server 102. The network 103 is used as a medium for providing a communication link between the terminal device 101 and the server 102. The network 103 may include various connection types, such as wired or wireless communication links or fiber optic cables.

[0031] The user can use the terminal device 101 to interact with the server 102 via the network 103 to receive or send messages, etc. Various communication client applications can be installed on the terminal device 101, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).

[0032] The terminal device 101 may be any electronic device having a display screen and supporting web browsing, including but not limited to a smart phone, a tablet computer, a laptop computer, a desktop computer, and the like.

[0033] The server 102 may be a server that provides various services, such as a background management server (for example only) that supports websites browsed by users using the terminal device 101. The background management server may analyze and process received user requests and other data, and feed back the processing results (e.g., web pages, information, or data obtained or generated according to user requests) to the terminal device.

[0034] It should be noted that the multimodal data alignment method and apparatus provided in the embodiments of the present disclosure can be executed by the server 102. Accordingly, the multimodal data alignment method and apparatus provided in the embodiments of the present disclosure can be set in the server 102. The multimodal data alignment method and apparatus provided in the embodiments of the present disclosure can also be executed by a server or server cluster that is different from the server 102 and can communicate with the terminal device 101 and / or the server 102. Accordingly, the multimodal data alignment method and apparatus provided in the embodiments of the present disclosure can also be set in a server or server cluster that is different from the server 102 and can communicate with the terminal device 101 and / or the server 102.

[0035] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.

[0036] Figure 2 A flowchart of a multimodal data alignment method provided by an embodiment of the present disclosure. Figure 2 , the method comprising:

[0037] Step S210: input the multimodal training session into the feature extraction model to be trained to obtain the modal features of each modal data in the multimodal training session.

[0038] A multimodal training session refers to multimodal session data used to train a feature extraction model. A multimodal training session may include multiple rounds of training sessions, each of which contains multimodal data for representing multimodal questions raised by users.

[0039] Among them, the multiple modal data in the multimodal training session can be at least two modal data among text data, image data, audio data, video data, etc.

[0040] An example multimodal training session is: “Text data: What color is this bird?

[0041] Image data: contains a picture of a bird;

[0042] Audio data: The user's voice description of the text data.

[0043] Among them, the feature extraction model is used to extract the modal features of each modal data in the multimodal training session.

[0044] In an optional implementation, since there are obvious differences in feature distribution and information expression among the modal data, a corresponding feature extraction module can be set for each modal data in the feature extraction model to improve the extraction effect of the modal features of each modal data.

[0045] Correspondingly, the feature extraction model includes multiple feature extractors corresponding to multiple modality types, and different feature extractors are used to extract modality data corresponding to different modality types.

[0046] Exemplarily, when the modal data is text data, the feature extraction model includes a first feature extractor for extracting text data, such as a Transformer. When the modal data is image data, the feature extraction model includes a second feature extractor for extracting image data, such as a Vision Transformer or a convolutional neural network. When the modal data is audio or video data, the feature extraction model includes a third feature extractor for extracting audio or video data, such as a Transformer or a recurrent neural network.

[0047] Therefore, by inputting the multimodal training session into the feature extraction model to be trained, the modal features of each modal data in the multimodal training session can be obtained.

[0048] Furthermore, after obtaining the modal features of each modal data, the modal features of each modal data can be mapped to the same shared semantic space, so as to facilitate subsequent comparative learning and feature fusion processing of the modal features of each modal data in the shared semantic space.

[0049] It should be noted that semantic alignment of multimodal data refers to extracting modal features of the multimodal data and making the modal features of each modal data semantically equivalent or semantically associated in a shared semantic space.

[0050] Therefore, in the embodiment of the present disclosure, after each feature extractor in the feature extraction model, a cross-modal projection layer is further included to map the modal features of each modal data into the same shared semantic space.

[0051] The cross-modal projection layer can be a linear projection layer, such as using a fully connected layer to linearly project the modal features of each modal data, thereby mapping each modal feature to the same semantic space. The cross-modal projection layer can also be a nonlinear projection layer, that is, constructed based on a fully connected layer and an activation function, which nonlinearly projects the modal features of each modal data, thereby mapping each modal feature to the same shared semantic space.

[0052] Therefore, in the disclosed embodiment, first, the modal features of each modal data are extracted by the feature extractors of the feature extraction model. Then, the modal features of each modal data are projected into the same shared semantic space by the cross-modal projection layer of the feature extraction model. Thus, the modal features of each modal data are comparatively learned in the shared semantic space. That is, by constructing a contrastive learning loss function, i.e., the first loss, to compare the feature similarity between any modal feature and its corresponding cross-modal positive sample modal feature and negative sample modal feature.

[0053] By minimizing the first loss, the feature similarity between any modal feature and its corresponding cross-modal positive sample modal feature is increased, while the feature similarity between any modal feature and its corresponding cross-modal negative sample modal feature is decreased. This reduces the feature distance between modal features with semantic similarity / semantic relevance, while increases the distance between features without semantic similarity / semantic relevance. Ultimately, through comparative learning, semantic associations can exist between the modal features of each modal data, achieving semantic alignment.

[0054] In addition, since each feature extractor of the feature extraction model may have problems such as semantic deviation in the extracted modal features during feature extraction, this affects the comparative learning of each modal feature to achieve the alignment effect of semantic alignment.

[0055] Therefore, in the disclosed embodiment, feature fusion is also performed based on the modal features of multimodal data. That is, the modal features of multiple modal data are constructed into a fused feature, and the fused feature is input into a preset question-answering model to extract the semantic features of the answer to the question generated by the preset question-answering model. The higher the feature similarity between the semantic feature and the fused feature, the closer the correspondence between the two. Therefore, this indicates that the feature representativeness of the fused feature constructed based on the modal features of the multimodal data is stronger, that is, the feature semantic expression ability of each modal feature extracted by the feature extraction model is stronger.

[0056] Thus, the disclosed embodiment compares the semantic features of the answer to the question with the fused features and constructs a second loss based on the comparison results. By minimizing the second loss, the correspondence between the fused features constructed based on each modal feature and the semantic features is made closer, thereby optimizing the modal features extracted by the feature extraction model to improve the representativeness of each modal feature. This further improves the learning effect of the comparative learning of each modal feature based on the first loss, ultimately improving the semantic alignment effect between the modal features after comparative learning.

[0057] Step S220: performing comparative learning based on the modal features of each modal data in the multimodal training session to obtain a first loss.

[0058] Among them, comparative learning based on the modal features of each modal data in the multimodal training session refers to enabling the feature extraction model to learn the feature relationship of the modal features, so as to effectively distinguish between modal features with similar semantics and modal features with dissimilar semantics.

[0059] Correspondingly, the first loss can also be called feature contrast loss, which is determined based on the feature similarity between different modal features. By reducing the loss value of the first loss, the distance between semantically similar modal features can be reduced, and the distance between semantically dissimilar modal features can be increased. For example, if there is an image of "cat" and the text "What kind of cat is in the picture?", by minimizing the first loss, "cat" can be used as the semantic association core, so that the semantic features about cats in the modal features of the text can be semantically associated with the descriptive features about cats in the modal features of the image, reducing the feature distance between the two. For modal features in the image that are not related to "cats", such as the background features of the image, the feature distance between them and the modal features of the text is increased, thereby achieving a better semantic association between the modal features of the text and the modal features of the image.

[0060] Among them, the first loss can be any contrast loss function constructed according to the modal features of each modal data in the multimodal training session, such as information maximization loss (InfoNCE loss) based on noise contrast estimation, triple loss, symmetric cross entropy loss, etc., which is not limited in the embodiments of the present disclosure.

[0061] Among them, when constructing the first loss, it is necessary to determine, for any modal feature, the positive sample modal features that are semantically related to it, and the negative sample modal features that are semantically irrelevant to it in the modal features of the multimodal training session, so as to construct the first loss based on the similarity between the modal feature and the positive sample modal features and the negative sample modal features. By reducing the loss value of the first loss, the similarity between the modal feature and the positive sample modal features is greater, and the similarity between the modal feature and the negative sample modal features is reduced.

[0062] It should be noted that when selecting positive sample modal features and negative sample modal features, adaptive selection can be performed according to actual application needs, and the embodiments of the present disclosure do not limit this.

[0063] For example, a multimodal training session includes multiple rounds of training sessions, each of which includes multiple modal data. For any modal data, if the modal feature of the modal data is the first modal feature, i.e., the global feature, the modal features of the other modal data in the training session to which it belongs are the positive sample modal features. Correspondingly, the negative sample modal features can be selected from the modal features of the modal data in the remaining rounds of training sessions.

[0064] In the case where the modal feature of the modal data is the second modal feature, that is, a local feature, among the modal features of other modal data in the training session to which it belongs, the modal feature that has semantic relevance with the modal feature of the modal data is the positive sample modal feature, and the modal feature that has no semantic relevance with the modal feature of the modal data can be the negative sample modal feature. The semantic relevance of each modal feature can be determined based on the semantic category label pre-marked during training.

[0065] Step S230: Construct a fusion feature based on the modal features of each modal data in the multimodal training session, input the fusion feature into a preset question-answering model, and obtain the answer to the question corresponding to the multimodal training session.

[0066] The fused feature is a feature representation generated by fusing the modal features of each modal data in a multimodal training session. The fused feature integrates the comprehensive semantic information of each modal data, such as text, image, and audio / video.

[0067] In an optional implementation, the fusion feature can be obtained by weighting the modal features of each modal data in the multimodal training session. The weight value of the modal feature of each modal data can be a preset value, or can be determined based on the data information quality of the modal data (such as information relevance, signal-to-noise ratio), or the modal features of each modal data can be input into a preset attention mechanism (such as a channel attention mechanism, a self-attention mechanism, etc.), and the preset attention mechanism dynamically assigns the weight value of the modal feature of each modal data. The embodiments of the present disclosure are not limited to this.

[0068] The pre-trained question-answering model is a large, pre-trained language model used to answer multimodal conversations. It incorporates a pre-trained knowledge base and possesses contextual understanding capabilities, enabling comprehensive analysis and semantic understanding of the modal features of multimodal data.

[0069] Therefore, by inputting the fused features into the preset question-answering model, the preset question-answering model can capture the feature connections of the fused features, generate feature semantics that conform to the fused features, and expand the explanation of questions based on external knowledge, so as to meet the question-answering needs in complex scenarios.

[0070] Step S240: Extract the semantic features of the answer to the question, compare the semantic features with the fusion features, and obtain the second loss.

[0071] Among them, the question answer can be feature extracted by a dedicated answer encoder (such as a convolutional neural network) to obtain the semantic features of the question answer. Accordingly, the semantic features of the question answer can be compared with the fusion features to measure the semantic matching degree between the generated question answer and the input multimodal training session data. Correspondingly, the second loss is a loss function used to measure the feature matching degree between the semantic features and the fusion features. The smaller the loss value of the second loss, the higher the feature matching degree between the semantic features and the fusion features, that is, the stronger the feature representativeness of the fusion features constructed based on each modal feature, and the better the feature expression ability of each modal feature. For example, the second loss can be cosine similarity loss, Euclidean distance loss, Manhattan distance loss, etc., which is not limited in the embodiments of the present disclosure.

[0072] Step S250: Adjust the training parameters of the feature extraction model according to the first loss and the second loss to obtain a trained feature extraction model, so as to extract multiple mutually aligned modal features of the multimodal real-time conversation through the trained feature extraction model.

[0073] Among them, corresponding loss weights can be assigned to the first loss and the second loss, so as to construct a joint loss based on the loss weight, the first loss and the second loss. By minimizing the joint loss, the training parameters of the feature extraction model are adjusted to obtain the trained feature extraction model.

[0074] As can be seen from the above description, the first loss is used to compare and learn the modal features of each modal data in the shared semantic space. The first loss is determined according to the similarity between any modal feature and its positive sample modal feature and negative sample modal feature. By minimizing the first loss, the similarity value between the modal feature and the positive sample modal feature can be made larger, and the similarity value between the modal feature and the negative sample modal feature can be made smaller, thereby making the feature connection between the modal features with semantic relevance closer, and realizing the semantic alignment of the modal features in the shared semantic space.

[0075] The second loss is used to compare semantic features and fused features. By minimizing the second loss, the degree of feature matching between semantic features and fused features is improved, thereby making the feature representativeness of the fused features constructed based on the features of each modality stronger. This optimizes the feature semantics of each modal feature extracted by the feature extraction model. Accordingly, the learning effect of contrastive learning based on the optimized modal features will be better, resulting in better semantic alignment of the modal features after contrastive learning.

[0076] Therefore, the embodiment of the present disclosure performs joint training based on the first loss and the second loss, which can improve the semantic alignment effect between the modal features extracted by the feature extraction model for multimodal data.

[0077] After the feature extraction model training is completed, the modal features of each modal data in the multimodal real-time conversation can be extracted based on the trained feature extraction model, thereby obtaining multiple mutually aligned modal features, that is, making the multiple modal features comparable and semantically related, thereby achieving collaborative understanding and complementary enhancement between multiple modal features.

[0078] Among them, the multimodal real-time conversation is the multimodal question data input by the user in real time, so that the multimodal real-time conversation is feature extracted through the trained feature extraction model to obtain multiple mutually aligned modal features, so that the preset question-and-answer model can generate the conversation answer corresponding to the multimodal real-time conversation based on the multiple mutually aligned modal features.

[0079] In an optional implementation, after adjusting the training parameters of the feature extraction model based on the first loss and the second loss to obtain the trained feature extraction model, the method is further used to: input the multimodal real-time conversation into the trained feature extraction model to obtain the modal features of each modal data in the multimodal real-time conversation; construct a fusion feature corresponding to the multimodal real-time conversation based on the modal features of each modal data in the multimodal real-time conversation; and input the fusion feature corresponding to the multimodal real-time conversation into a preset question-and-answer model to obtain a conversation answer corresponding to the multimodal real-time conversation.

[0080] When constructing the fusion features corresponding to the multimodal real-time session, the fusion features corresponding to the multimodal real-time session can be obtained by weighting the modal features of each modal data in the multimodal real-time session. The weight values ​​of the modal features of each modal data can be determined based on the data information quality of the modal data (such as information relevance and signal-to-noise ratio), or the modal features of each modal data can be input into a preset attention mechanism, and the preset attention mechanism dynamically assigns the weight values ​​of the modal features of each modal data. The embodiments of the present disclosure are not limited to this.

[0081] In the embodiment of the present disclosure, since the feature extraction model is jointly trained based on the first loss and the second loss, the trained feature extraction model can improve the feature alignment effect between the modal features of each modal data extracted. Accordingly, the feature quality of the fusion feature constructed based on multiple mutually aligned modal features is higher, and thus the fusion feature is input into the preset question-answering model, which enables the preset question-answering model to fully understand the semantic connection of the multimodal real-time conversation, so as to improve the accuracy of the generated conversation answers.

[0082] The embodiment of the present disclosure provides a method for aligning multimodal data. First, the modal features of each modal data in the multimodal training session are obtained by inputting the multimodal training session into a feature extraction model to be trained. Secondly, comparative learning is performed based on the modal features of each modal data in the multimodal training session to obtain a first loss. Then, a fusion feature is constructed based on the modal features of each modal data in the multimodal training session. The fusion feature is input into a preset question-answering model to obtain the answer to the question corresponding to the multimodal training session. The semantic features of the answer to the question are extracted and compared with the fusion feature to obtain a second loss. The first loss is used to measure the semantic alignment effect of the modal features of each modal data, and the loss of the first loss is the sum of the modal features of the first loss and the sum of the modal features of the first loss. The smaller the value, the closer the feature distance between the modal features of each modal data and the modal features with semantic correlation is. The second loss is used to measure the feature expression ability of the fusion feature constructed based on the modal features of each modal data. The smaller the loss value of the second loss, the higher the feature matching degree between the fusion feature and the semantic feature, that is, the stronger the feature semantic representativeness of each modal feature used to construct the fusion feature. Correspondingly, the comparative learning based on each modal feature achieves better feature alignment. Finally, according to the first loss and the second loss, the training parameters of the feature extraction model are jointly adjusted to obtain the trained feature extraction model, so as to extract multiple mutually aligned modal features of the multimodal real-time conversation through the trained feature extraction model.

[0083] It can be seen that when training the feature extraction model, the present disclosure can align the modal features of each modal data extracted by the feature extraction model through the comparative learning method of the first loss, and can make the feature matching degree between the fusion features of the multimodal training session and the semantic features of the question answer higher through the second loss, thereby optimizing the feature extraction ability of the feature extraction model for modal features, and then making the alignment effect between the modal features of each modal data better. Therefore, the present disclosure can extract multiple mutually aligned modal features of the multimodal real-time conversation based on the trained feature extraction model, thereby improving the alignment effect of the multimodal data.

[0084] In an optional implementation, in order to improve the feature extraction effect of the feature extraction model for the modal features of each modal data, the feature extraction model can extract the global features and local features of the modal data respectively. Accordingly, the feature extraction model includes a global feature extraction module and a local feature extraction module, and the multimodal training session is input into the feature extraction model to be trained to obtain the modal features of each modal data in the multimodal training session, including: according to the global feature extraction module, global feature extraction is performed on the multimodal training session to obtain the first modal feature of each modal data in the multimodal training session; according to the local feature extraction module, local feature extraction is performed on the multimodal training session to obtain the second modal feature of each modal data in the multimodal training session; wherein the modal features include the first modal feature and the second modal feature.

[0085] The global feature extraction module is used to extract global features of the modal data. The local feature extraction module is used to extract local features of the modal data. Accordingly, the first modal feature can also be called a global feature, and the second modal feature can also be called a local feature.

[0086] It should be noted that the disclosed embodiments do not restrict the module positional relationship between the global feature extraction module and the local feature extraction module in the feature extraction model. For example, the global feature extraction module and the local feature extraction module may be in a first positional relationship, in which case they are two parallel feature extraction branches. Thus, by inputting the multimodal training session into the global feature extraction module and the local feature extraction module, respectively, first modal features and second modal features are obtained.

[0087] For another example, the global feature extraction module and the local feature extraction module may be in a second positional relationship, where the global feature extraction module is located after the local feature extraction module. The second modal feature is obtained by inputting the multimodal training session into the local feature extraction module, and the second modal feature is input into the global feature extraction module for feature fusion to obtain the first modal feature.

[0088] For another example, the global feature extraction module and the local feature extraction module can be in a third position relationship. In this case, the global feature extraction module is located before the local feature extraction module. By inputting the multimodal training session into the global feature extraction module, the first modal feature is obtained, and the first modal feature is input into the local feature extraction module for local feature extraction to obtain the second modal feature.

[0089] In an optional implementation, since different modal data have different data representation methods, the feature extraction model may include a feature extractor for extracting features from each modal data, and each feature extractor includes a global feature extraction module and a local feature extraction module. In addition, the module position relationship between the global feature extraction module and the local feature extraction module in the feature extractor needs to be determined based on the modal type of the modal data corresponding to the feature extractor. The module position relationship between the global feature extraction module and the local feature extraction module corresponding to different modal types is different.

[0090] For example, when the modal data is of the first type, the module position relationship between the global feature extraction module and the local feature extraction module in the first feature extractor corresponding to the modal data is a first position relationship. When the modal data is of the second type, the module position relationship between the global feature extraction module and the local feature extraction module in the second feature extractor corresponding to the modal data is a second position relationship. When the modal data is of the third type, the module position relationship between the global feature extraction module and the local feature extraction module in the third feature extractor corresponding to the modal data is a third position relationship. The first type, the second type, and the third type can be any one of a text type, an image type, and a multimedia type (i.e., an audio and video type), respectively.

[0091] It should also be noted that the module structure of the global feature extraction module and the local feature extraction module can be adaptively set according to the actual feature extraction needs, and the embodiments of the present disclosure do not limit this. For example, the global feature extraction module is composed of a global average pooling layer, which obtains the global features of the modal data based on spatial averaging. The global feature extraction module can also be composed of a convolutional layer and a fully connected layer, and the features extracted by the convolutional layer are integrated and transformed by the fully connected layer to obtain global features. The global feature extraction module can also be composed of an LSTM layer, which extracts the global features of the modal data by integrating the modal data into the hidden state.

[0092] For another example, the local feature extraction module can be composed of convolution layers, such as performing convolution processing on the modal data through convolution kernels of different scales to obtain multiple local features of the modal data, or processing multiple blocks corresponding to the modal data through a multi-layer Transformer to extract local features, or extracting the initial local features of the modal data through a convolution layer, and further extracting the initial local features through an LSTM layer to obtain the final local features. In an optional implementation, in order to improve the feature representation capability of local features, an attention mechanism layer can also be set in the local feature extraction module, such as a multi-head attention mechanism layer, a self-attention mechanism layer, a channel attention mechanism layer, etc. The modal data is convolved through a convolution layer and pooled through a pooling layer to obtain initial local features. The attention mechanism layer assigns weights to different initial local features to obtain the final local features, i.e., the second modal features.

[0093] For ease of understanding, the specific model structure of the feature extraction model in the embodiment of the present disclosure is described below: the feature extraction model includes: an input layer, a feature extractor, and a cross-modal projection layer. The input layer is used to input multiple modal data of a multimodal training session. Thus, the multiple modal data input by the input layer are input into the feature extractor to extract the modal features of the modal data. For different modal data, the feature extraction model may include corresponding feature extractors, for example, feature extractors corresponding to text modal data, feature extractors corresponding to image data, and feature extractors corresponding to audio and video data. The output of each feature extractor is the modal feature of the modal data. By inputting the modal features into the cross-modal projection layer, each modal feature is mapped to a shared semantic space.

[0094] It should be noted that each feature extractor may include the aforementioned global feature extraction module and local feature extraction module. The specific structure of the module can be referred to above and will not be elaborated here. Thus, each feature extractor can extract the first modal features of the corresponding modal data according to the global feature extraction module and extract the second modal features of the corresponding modal data according to the local feature extraction module. Furthermore, the first modal features and the second modal features are input into a cross-modal projection layer, for example, where the features are mapped to a shared semantic space through an activation function.

[0095] It should also be noted that the feature extraction model can also include a local attention mechanism layer. For example, a multi-head attention mechanism layer can be used, which can be located after the feature extractor to further perform local attention processing on the second modal features extracted by the feature extractor, or located after the cross-modal projection layer to perform local attention processing on the second modal features feature mapped in the cross-modal projection layer, thereby improving the local feature expression ability of each second modal feature.

[0096] Correspondingly, the output layer of the feature extraction model can directly output the first modal features and second modal features corresponding to each modal data obtained after the final processing of each module in the model. The embodiments of the present disclosure are not limited to this.

[0097] Furthermore, the training parameters of the feature extraction model are the training parameters corresponding to each module in the feature extraction model, for example, the training parameters corresponding to each feature extractor, cross-modal projection layer, and local attention mechanism layer. Accordingly, by inputting the multimodal training session into the feature extraction model, a first loss can be constructed based on the second modal features extracted by the feature extraction model, and a fusion feature can be constructed based on the first modal features corresponding to the feature extraction model. Then, a second loss is constructed based on the semantic features corresponding to the answer generated by the preset question-answering model for the fusion feature. The feature extraction model is trained by configuring the training rounds, training batch size, optimizer, loss rate, etc., and the training parameters of the feature extraction model are adjusted by reducing the first loss and the second loss to obtain a trained feature extraction model.

[0098] In the embodiment of the present disclosure, two types of fine-grained features, global features and local features of modal data, are extracted according to a feature extraction model, thereby improving the feature richness of the modal features, improving the feature representation effect of the modal features, and further improving the feature alignment capability of the modal features.

[0099] In an optional implementation, the number of second modal features of each modal data is multiple, and comparative learning is performed based on the modal features of each modal data in the multimodal training session to obtain a first loss, including: for any second modal feature of any modal data, determining the positive sample modal features and the negative sample modal features from the second modal features of the remaining modal data in the multimodal training session; calculating the similarity between any second modal feature and the positive sample modal feature to obtain a first similarity value; calculating the similarity between any second modal feature and the negative sample modal feature to obtain a second similarity value; and obtaining a first loss based on the first similarity value and the second similarity value.

[0100] Among them, positive sample modal features refer to cross-modal features that correspond to the same semantics, categories, or labels as the second modal features. Negative sample modal features refer to cross-modal features that correspond to different semantics, categories, or labels than the second modal features.

[0101] For example, a second modal feature of text data represents the text attribute "color," while a second modal feature of image data focuses on the color region within the image attribute. Therefore, these two second modal features are each other's positive modal features. Another second modal feature of image data focuses on the shape and texture of objects in the image. Therefore, it and the second modal feature of the text data are each other's negative modal features.

[0102] It should be noted that corresponding category labels can be pre-set for the second modal features of each modal data. Thus, for any second modal feature of any modal data, second modal features with the same category labels can be selected from the second modal features of the remaining modal data in the multimodal training session as positive sample modal features, and second modal features with different category labels can be selected as negative sample modal features. It should also be noted that the embodiments of the present disclosure do not impose any specific restrictions on the number of positive and negative sample modal features.

[0103] Among them, the similarity between any second modal feature and the positive sample modal feature and the negative sample modal feature can be calculated by any distance function, for example, the Euclidean distance, Manhattan distance, cosine similarity, etc. between the two can be calculated, and the embodiments of the present disclosure are not limited to this.

[0104] Thus, the first similarity value represents the similarity between the second modal feature and the positive sample modal feature, and the second similarity value represents the similarity between the second modal feature and the negative sample modal feature. Based on the first and second similarities, a first loss, namely the contrastive loss, can be constructed.

[0105] For example, the loss formula for the first loss is:

[0106]

[0107] Among them, L global is the first loss, sim() represents the similarity between the two modal features, z i is any second modal feature, z j is the modal feature of the positive sample, z k is the negative sample modal feature, N is the number of negative sample modal features, and τ is the temperature coefficient.

[0108] In the disclosed embodiment, the second modal features, i.e., local features, of each modal data are compared with the similarity between the modal features of the cross-modal positive samples and the modal features of the negative samples to construct a first loss, thereby effectively improving the granularity of feature contrast learning and enabling effective alignment between the local features of each modal data, thereby improving the feature alignment effect.

[0109] In an optional implementation, a multimodal training session includes multiple rounds of training sessions, each round of training session includes multiple modal data, and a fusion feature is constructed based on the modal features of each modal data in the multimodal training session, including: for any round of training session, based on the first modal feature of each modal data in any round of training session, determining the data information quality of each modal data; based on the data information quality of each modal data, determining the data weight of each modal data; and constructing the fusion feature corresponding to any round of training session based on the data weight of each modal data and the first modal feature of each modal data.

[0110] The data information quality of modal data is used to characterize the data reliability and data validity of the modal data. The data information quality includes information relevance and signal-to-noise ratio.

[0111] Information relevance is used to characterize the correlation between each modal data item in the current conversation and the corresponding type of modal data in the contextual conversation. For example, if a user raises multiple rounds of questions regarding a particular issue or series of questions, the information relevance between the modal data items in these rounds is relatively high.

[0112] Among them, the signal-to-noise ratio is used to characterize the proportional relationship between the effective data intensity and the noise data of each modal data, which reflects the data clarity of the modal data and the degree of interference of the noise. A high signal-to-noise ratio indicates that the effective data is dominant, while the influence of noise is small, so that the data quality of the modal data is better. For text data, the signal-to-noise ratio can be obtained by calculating the ratio of useful information to noise information in the first modal feature. For image data, the noise level can be calculated using the mean square error and the signal-to-noise ratio can be calculated based on the signal strength of the first modal feature. For audio data, the signal-to-noise ratio can be obtained by using the power spectral density or directly calculating the ratio of the mean square value of the signal to the noise. The embodiments of the present disclosure are not limited to this.

[0113] Therefore, after determining the data information quality of each modal data, the data weight of each modal data can be determined according to the numerical value of the data information quality, and then the fusion feature corresponding to the training session can be constructed according to the data weight of each modal data and the first modal feature of each modal data.

[0114] It should also be noted that in order to improve the construction effect of fusion features, the fusion features corresponding to the training dialogue can also be jointly constructed based on the data weight of each modal data, the first modal features and the second modal features of each modal data. The embodiments of this disclosure do not limit this.

[0115] In the embodiment of the present disclosure, the data weight of each modal data is dynamically determined based on the signal-to-noise ratio and information relevance of each modal data, so that the weight value of each modal data is dynamically adjusted according to its actual contribution in the conversation, ensuring that the most valuable modal data in the fusion feature dominates the decision, thereby improving the feature construction accuracy of the fusion feature.

[0116] In an optional implementation, when data quality is information relevance, the data information quality of each modal data is determined based on the first modal feature of each modal data in any round of training session, including: determining a context session related to any round of training session from multiple rounds of training sessions; for any modal data in any round of training session, based on the data type of any modal data, determining the target modal data corresponding to the data type from the modal data of the context session, comparing the first modal feature of any modal data with the first modal feature of the target modal data, and determining the information relevance of any modal data.

[0117] The contextual conversation related to any round of training conversation refers to the relevant historical conversation rounds (such as the i-th round of conversation) before the current round of conversation (such as the i-th round of conversation).

[0118] Specifically, the corresponding session time information and / or session round number are pre-marked for each round of training session, so that the context session related to any round of training session can be determined according to the session time information and / or session round number.

[0119] Thus, after determining the contextual conversation, for any modal data in the current training conversation, the contextual conversation can be used to search for target modal data corresponding to its data type. For example, if the modal data is text-type modal data, the text-type modal data can be searched from the contextual conversation as the target modal data. Furthermore, by comparing the first modal feature (i.e., global feature) of the modal data with the first modal feature (i.e., global feature) of the target modal data, for example by calculating feature similarity between the two, the information relevance of the modal data can be determined based on the comparison result, such as the magnitude of the similarity value.

[0120] In an embodiment of the present disclosure, by searching for a context session related to a training session and comparing the context session with the first modal feature of the modal data corresponding to the same data type in the training session, the information relevance of the modal data of the data type in the training session is determined, so that the historical training session information can be effectively combined to effectively reflect the information coherence of the modal data, thereby improving the accuracy of determining the information relevance.

[0121] In an optional implementation, in order to improve the accuracy of constructing the second loss, the semantic features can be decoded to obtain reconstructed fusion features, and the second loss can be obtained based on the comparison result between the reconstructed fusion features and the fusion features.

[0122] Correspondingly, the semantic feature is compared with the fusion feature to obtain a second loss, including: decoding the semantic feature to obtain a reconstructed fusion feature corresponding to the semantic feature; comparing the reconstructed fusion feature with the fusion feature for similarity to obtain a first comparison result; comparing the semantic feature with the fusion feature for similarity to obtain a second comparison result; and obtaining a second loss based on the first comparison result and the second comparison result.

[0123] The semantic features can be decoded by a preset generator to obtain the reconstructed fusion features corresponding to the semantic features. The preset generator can be a neural network structure with feature decoding capabilities, such as a generative adversarial network structure or an autoencoder structure, and the embodiments of the present disclosure are not limited to this.

[0124] Among them, the reconstructed fusion feature is the fusion feature generated by reverse reasoning based on the semantic feature. The generator can decode the semantic feature to generate corresponding modal features, such as text features, image features, and audio features, and then fuse the above features to obtain the reconstructed fusion feature.

[0125] Among them, the feature distance between the fused feature and the reconstructed fused feature and the semantic feature can be calculated, so as to obtain the first comparison result and the second comparison result according to the feature distance, and then obtain the second loss according to the first comparison result and the second comparison result.

[0126] For example, corresponding loss weight values ​​may be assigned to the first comparison result and the second comparison result, so that the second loss is obtained by weighting the first comparison result and the second comparison result.

[0127] For example, the loss formula of the second loss is:

[0128]

[0129] Among them, a is the semantic feature, x f is the fusion feature, z h To reconstruct the fusion features, the semantic representation after multimodal fusion, α1 and α2 are the loss weights.

[0130] In the embodiment of the present disclosure, the semantic features are decoded, and the reconstructed fusion features after decoding are compared with the original fusion features to obtain a first comparison result, and the original fusion features are compared with the semantic features to obtain a second comparison result, so as to measure the feature matching degree between the semantic features and the fusion features from multiple dimensions to improve the construction accuracy of the second loss and enhance the optimization effect of the feature extraction capability of the feature extraction model.

[0131] To facilitate understanding, the following describes the specific implementation details of the above embodiment using a specific example:

[0132] In recent years, with the rapid development of artificial intelligence, big data, and deep learning technologies, large multimodal language models have become an important research direction in natural language processing and computer vision. Leveraging large-scale pre-trained models and massive amounts of data, researchers have successfully built systems capable of simultaneously processing text, images, audio, and even video information. They have made significant progress in areas such as intelligent question answering, cross-modal retrieval, automatic summarization, and intelligent customer service.

[0133] Although current multimodal large language models have achieved certain success, they still face many technical challenges and bottlenecks in practical applications. First, the data of each modality is obviously heterogeneous. The feature distribution and noise level of data such as text, images, and audio vary greatly. Information loss or deviation often occurs when aligning the features of each modality. Second, for multimodal data, traditional feature fusion methods usually use fixed weights or simple weighting strategies. It is difficult to adaptively adjust the contribution of each modality according to different scenarios or the specific conditions of the input data, thus limiting the system's performance in complex scenarios. In addition, when faced with information-rich and dynamically changing input, current question-answering systems are prone to mismatches or misunderstandings, which in turn affects the overall accuracy and user experience.

[0134] In light of this, this example provides a multimodal data alignment method that builds an end-to-end multimodal processing framework, including preprocessing modules for various modal data, such as text, images, audio, and video data, dedicated encoders, and shared semantic embedding construction modules. By mapping each modal data into a unified semantic space and utilizing global and local alignment strategies, basic cross-modal information fusion is achieved. Furthermore, this example further introduces an adaptive weighting mechanism for dynamic semantic alignment. Based on real-time evaluation of contextual information and signal-to-noise ratio, this mechanism automatically adjusts the weight distribution of each modality during the fusion process, thereby achieving more accurate feature fusion. Furthermore, to address the semantic bias that may occur after question and answer generation, this example designs a self-supervised semantic calibration feedback loop. By mapping the generated answers into a shared semantic space, comparing them with the original multimodal input data in real time, and using consistency loss for feedback tuning, this system achieves closed-loop optimization, significantly improving the accuracy of the question and answer results and the overall robustness of the system.

[0135] Overall, this solution adopts a multi-level fusion strategy for the alignment of multimodal data, which not only ensures the consistency and integrity of information from different modalities, but also enhances the system's adaptability to complex and changing scenarios. The main innovations of this solution include:

[0136] (1) Dynamic adaptive weighting mechanism: Based on context and signal-to-noise ratio evaluation, the weight of each modality data is adjusted dynamically in real time;

[0137] (2) Self-supervised semantic calibration feedback loop: Build a closed-loop feedback system to compare the generated answers with the input multimodal data and perform real-time tuning of the question-answering results through consistency loss;

[0138] (3) Multi-level fusion strategy: In terms of feature alignment of multimodal data, it takes into account both overall consistency and fine-grained feature matching to improve the robustness and adaptability of the system.

[0139] This demonstrates that the multimodal data alignment method used in this example achieves basic fusion of data from different modalities by constructing an end-to-end multimodal processing framework based on global and local alignment strategies. Furthermore, a dynamic adaptive weighting mechanism and a self-supervised semantic calibration feedback loop are introduced to achieve real-time adjustment and closed-loop optimization of the weights of each modal data. Through this multi-level fusion strategy, this example not only ensures the overall consistency of information from each modality, but also strengthens the matching capabilities of fine-grained features and significantly improves the accuracy, robustness, and adaptability of the question-answering system.

[0140] Figure 3 A flowchart of a multimodal data alignment method provided for this example, see Figure 3 , the method comprising:

[0141] Step S301: inputting the multimodal training session into the feature extraction model to be trained to obtain the modal features of each modal data in the multimodal training session, wherein the modal features of each modal data include a first modal feature and multiple second modal features.

[0142] Before extracting features from the multimodal training session, a preprocessing module may be used to preprocess the multiple modal data in the multimodal training session to improve the data quality of the multimodal training session.

[0143] Figure 4 This is a diagram of the training process of the feature extraction model, refer to Figure 4 Multimodal training sessions are collected by the multimodal input layer, which supports the parallel input of text, image, audio, and video modal data. Each data type is uploaded through a dedicated interface to ensure the integrity and diversity of data collection. In addition, at the input layer, multimodal data is labeled with timestamps, sources, and data categories, so that the contextual relationships between different modal data can be accurately associated during multimodal fusion.

[0144] Then, the data preprocessing module is used to preprocess the multimodal training session data.

[0145] Reference Figure 4 ,For modal data of different data types, the pre-processing module uses ,specialized technical means to clean and standardize the modal data, ,ensuring that the subsequent feature extraction model can receive high-quality ,input.

[0146] (1) Text data preprocessing: The text modal data is divided into independent words or subword units through word segmentation technology; then, the denoising method is used to filter out irrelevant information and noise words to improve the accuracy of subsequent semantic modeling.

[0147] (2) Image data preprocessing: Normalization technology is used to map the pixel values ​​of the image modality data to a standard range; then, the sample data volume of the image modality data is expanded through data enhancement methods such as random cropping, rotation, and flipping, thereby improving the generalization ability of the model; finally, Gaussian filtering, median filtering and other filtering denoising techniques are used to denoise the image modality data, reduce the noise and interference information in the image, and further optimize the image quality.

[0148] (3) Audio / video data preprocessing: First, the audio and video modal data are filtered using a frequency domain filter to eliminate background noise and retain valid signals; second, representative frames are selected from the audio and video modal data through key frame extraction, which not only reduces the amount of calculation but also retains the main visual information; finally, the audio and video modal data are time-aligned using frame rate synchronization technology to ensure the consistency of the two modalities in the time dimension.

[0149] Furthermore, in the feature extraction model, a dedicated feature extractor can be used for modal data of different data types to encode the preprocessed multiple modal data.

[0150] Reference Figure 4 The encoder for text modality data can be based on the Transformer structure, the encoder for image modality data can adopt the Vision Transformer (ViT) structure or the Convolutional Neural Network (CNN) structure, and the encoder for audio / video modality data can be based on the Transformer or the Recurrent Neural Network (RNN) structure. It should be noted that different encoders all include a global feature extraction module and a local feature extraction module. For the Transformer-based encoder, the local feature extraction module can include Transformer modules of different levels to process data windows of different sizes for local feature extraction. The global feature extraction module can perform average pooling on all local features output by the Transformer module in the spatial dimension by a global average pooling layer. For encoders based on recurrent neural networks or LSTM, the output features extracted from modal data by neurons in each layer are global features. In order to extract local features of modal data, a one-dimensional convolution kernel can be used to perform convolution operations on the modal data to obtain multiple initial local features, which are then further input into neurons in each layer to extract the final local features. The local feature extraction module can also be set as a sliding window or recurrent neural networks at different levels to extract local features of the modal data in a sliding window or layered manner.

[0151] Thus, based on the encoder extraction, the first modal features of each modal data, namely the global features, and multiple second modal features, namely the local features, are obtained.

[0152] Furthermore, the vectors output by each modality-specific encoder can be mapped to the same shared semantic embedding space through a cross-modal projection layer, laying the foundation for feature alignment between different modal data, thereby eliminating the obstacles caused by differences in data format and dimension between different modal data, and ensuring that the representation of the same semantic concept in different modalities is as close as possible, facilitating cross-modal retrieval and matching.

[0153] Step S302: performing comparative learning based on the second modality features of each modality data in the multimodal training session to obtain a first loss.

[0154] Specifically, in this step, for any second modal feature of any modal data, the positive sample modal feature and the negative sample modal feature can be determined respectively from the second modal features of the remaining modal data of the multimodal training session; the similarity between any second modal feature and the positive sample modal feature is calculated to obtain a first similarity value; the similarity between any second modal feature and the negative sample modal feature is calculated to obtain a second similarity value; and the first loss is obtained based on the first similarity value and the second similarity value.

[0155] Furthermore, it should be noted that in order to improve the feature alignment effect for each second modality feature, i.e., local feature, the feature extraction model may further include a multi-head attention mechanism module to further capture the detailed information of the second modality feature in the shared semantic space, such as entity relationships in the text and local features in the image, thereby obtaining the processed second modality feature, and performing comparative learning based on the processed second modality feature, i.e., performing global semantic alignment to obtain the first loss. The formula for the multi-head attention mechanism is:

[0156]

[0157] Where Q is the query, K is the key, V is the value, and dk is the dimension of the key / query vector. Used to scale the dot product result to prevent the gradient from disappearing.

[0158] Reference Figure 4 For each second modal feature in the shared semantic embedding space, it is first locally semantically aligned, that is, the second modal feature is processed through the local multi-head attention mechanism to obtain the processed second modal feature, so that the second modal feature can pay more attention to the detailed information of the modal data.

[0159] Then, the processed second modal features are subjected to cross-modal contrast information to construct the first loss, that is, cross-modal contrast learning based on each second modal feature, i.e., local features, to achieve global semantic alignment of each modal feature. In this step, the feature extraction model uses global contrast learning based on local features to ensure overall semantic consistency, while using a local attention mechanism to further align key points of image and text features. For example, the second modal features of the text data focus on the text portion of "the color of the bird", while the second modal features of the image data focus on the color area of ​​the bird in the image, so that the preset question-answering model can generate accurate question-answering answers based on the aligned fusion features, such as generating the answer to the question "This is a red bird."

[0160] Step S303: For any round of training session in the multimodal training session, determine the data information quality of each modal data based on the first modal feature of each modal data in any round of training session; determine the data weight of each modal data based on the data information quality of each modal data; and construct the fusion feature corresponding to any round of training session based on the data weight of each modal data and the first modal feature of each modal data.

[0161] Reference Figure 4 In this step, a dynamic weighting mechanism is adopted to assign corresponding weights to the modal data, so as to construct corresponding fusion features according to the data weight of the modal data and the first modal features of the modal data.

[0162] Specifically, the modal features of each modal data are weighted and superimposed by adopting a dynamic weighting formula, the data weight of the modal data is determined based on the information relevance and signal-to-noise ratio of each modal data, and based on the data weight, the fusion feature is constructed so that the final fusion feature can fuse the comprehensive semantic information of text, image and audio / video.

[0163] It should be noted that since the feature extraction model in this application can extract two fine-grained feature representations, namely the first modal features and the second modal features, a fusion feature can be constructed based on the first modal features, the second modal features and the data weights of the modal data. The fusion feature not only maintains the consistency of the global semantics, but also retains the local fine-grained features, providing high-quality input for the subsequent preset question-answering model.

[0164] Given three modal data: text data (t), image data (v), and audio data (a), the dynamic weighting formula is as follows:

[0165]

[0166] Among them, SNR m is the signal-to-noise ratio of the modal data, context is the information correlation of the modal data, and f() is a data processing function constructed based on the signal-to-noise ratio and information correlation, such as accumulation, multiplication, or exponential transformation of the two.

[0167] Assuming that the total value of the SNR and information relevance of text is 10, the total value of the SNR and information relevance of image is 5, and the total value of the SNR and information relevance of audio is 1, the weights are calculated using Softmax normalization:

[0168] w v ≈0.018, w a ≈0.02

[0169] Therefore, in the case of high-quality text input, the system will automatically tend to favor text modal features when performing feature fusion, while weakening audio modal features with greater noise.

[0170] Step S304: Input the fused features into a preset question-answering model to obtain answers to questions corresponding to the multimodal training session.

[0171] Reference Figure 4 The pre-trained question-answering model is a large pre-trained language model. By inputting the fused features into the pre-trained question-answering model, the answer output layer can output the answers to the questions corresponding to the multimodal training session. The answer output layer can post-process the answers generated by the pre-trained question-answering model to ensure that their format and content meet user expectations. Specifically, the answers are formatted and output, and the layout and punctuation are adjusted to improve readability. Secondly, a confidence assessment is performed based on the probability scores within the pre-trained question-answering model to provide users with a reference to the credibility of the answers.

[0172] Step S305: Extract the semantic features of the answer to the question, compare the semantic features with the fusion features, and obtain the second loss.

[0173] Among them, the semantic features of the answer to the question can be extracted through a dedicated answer encoder, and the generated answer can be mapped back to the shared semantic space. The second loss can be a consistency loss, and the specific formula of the second loss is:

[0174]

[0175] Among them, a is the semantic feature of the answer to the question, z f For fusion features.

[0176] Step S306: According to the first loss and the second loss, the training parameters of the feature extraction model are adjusted to obtain a trained feature extraction model.

[0177] In this step, a joint loss is constructed through the first loss and the second loss, so that during the training process, the loss function is minimized to adjust and optimize the training parameters of the feature extraction model, so that the feature extraction model can improve the feature alignment effect between the modal features of the multimodal training session.

[0178] Reference Figure 4 Based on the first and second losses, we can optimize the training parameters of the feature extraction model by minimizing the loss function during training, performing backpropagation to achieve online dynamic tuning and continuously improve the accuracy and robustness of the question-answering system. For example, we can optimize the training parameters of each feature encoder and the training parameters of the local multi-head attention mechanism to obtain the trained feature extraction model.

[0179] Step S307: Input the multimodal real-time conversation into the trained feature extraction model to obtain the modal features of each modal data in the multimodal real-time conversation.

[0180] Step S308: constructing a fusion feature corresponding to the multimodal real-time conversation based on the modal features of each modal data in the multimodal real-time conversation.

[0181] Among them, the data weight of the modal data can be determined by calculating the information relevance and signal-to-noise ratio of each modal data, so as to construct the fusion features corresponding to the multimodal real-time conversation based on the modal characteristics and modal weights of the modal data.

[0182] Step S309: Input the fusion features corresponding to the multimodal real-time conversation into a preset question-answering model to obtain a conversation answer corresponding to the multimodal real-time conversation.

[0183] In related technologies, the alignment of multimodal data is often limited to static mapping and simple data splicing, lacking adaptive adjustment to data quality and dynamic changes in actual scenarios, making it difficult to capture fine-grained semantic differences. This example achieves intelligent optimization of multimodal data alignment through a dynamic adaptive weighting mechanism, a self-supervised semantic calibration feedback loop, and a multi-level fusion strategy. This method can dynamically adjust the data weights of each modal data according to the data quality of real-time modal data, ensuring the effective use of high-quality information, and through a self-supervised learning mechanism, provide real-time feedback on the semantic deviation between the generated answer and the fused features, thereby improving the consistency and accuracy of feature extraction and generated answers. At the same time, this example combines global contrastive learning with a local attention mechanism, so that the feature extraction model can not only improve the feature representativeness of the extracted second modal features, but also achieve global contrastive learning based on local features, taking into account both overall semantic consistency and fine-grained feature matching, thereby enhancing the robustness and feature alignment accuracy in complex scenarios.

[0184] It is understood that the above-mentioned various method embodiments mentioned in this disclosure can be combined with each other to form combined embodiments without violating the principle logic. Due to space limitations, this disclosure will not go into details. It is understood by those skilled in the art that in the above-mentioned methods of specific implementation, the specific execution order of each step should be determined by its function and possible internal logic.

[0185] In addition, the present disclosure also provides a multimodal data alignment device, an electronic device, and a computer-readable storage medium, all of which can be used to implement any multimodal data alignment method provided by the present disclosure. The corresponding technical solutions and descriptions are referred to the corresponding records in the method section and will not be repeated here.

[0186] Figure 5 A block diagram of a multimodal data alignment device provided in an embodiment of the present disclosure.

[0187] Reference Figure 5 , an embodiment of the present disclosure provides a multimodal data alignment device, the multimodal data alignment device comprising:

[0188] An extraction module 51 is configured to input the multimodal training session into a feature extraction model to be trained to obtain modal features of each modal data in the multimodal training session;

[0189] a comparison module 52 for performing comparative learning based on the modal features of each modal data in the multimodal training session to obtain a first loss;

[0190] A generating module 53 is configured to construct a fusion feature based on the modal features of each modal data in the multimodal training session, input the fusion feature into a preset question-answering model, and obtain an answer to a question corresponding to the multimodal training session;

[0191] a comparison module 54, configured to extract semantic features of the answer to the question, and compare the semantic features with the fusion features to obtain a second loss;

[0192] The training module 55 is used to adjust the training parameters of the feature extraction model according to the first loss and the second loss to obtain a trained feature extraction model, so as to extract multiple mutually aligned modal features of the multimodal real-time conversation through the trained feature extraction model.

[0193] In an optional implementation, the feature extraction model includes a global feature extraction module and a local feature extraction module. The extraction module 51 inputs the multimodal training session into the feature extraction model to be trained to obtain the modal features of each modal data in the multimodal training session, including:

[0194] Performing global feature extraction on the multimodal training session according to the global feature extraction module to obtain a first modal feature of each modal data in the multimodal training session;

[0195] performing local feature extraction on the multimodal training session according to the local feature extraction module to obtain a second modal feature of each modal data in the multimodal training session;

[0196] The modal features include the first modal features and the second modal features.

[0197] In an optional implementation, the number of the second modal features of each modal data is multiple, and the comparison module 52 performs comparative learning based on the modal features of each modal data in the multimodal training session to obtain the first loss, including:

[0198] For any second modal feature of any modal data, determine the positive sample modal feature and the negative sample modal feature respectively from the second modal features of the remaining modal data of the multimodal training session; calculate the similarity between the any second modal feature and the positive sample modal feature to obtain a first similarity value; calculate the similarity between the any second modal feature and the negative sample modal feature to obtain a second similarity value; and obtain the first loss based on the first similarity value and the second similarity value.

[0199] In an optional implementation, the multimodal training session includes multiple rounds of training sessions, each round of training session includes multiple modal data, and the generation module 53 constructs a fusion feature based on the modal features of each modal data in the multimodal training session, including:

[0200] For any round of training session, determining the data information quality of each modal data according to the first modal feature of each modal data in the any round of training session;

[0201] Determining a data weight for each modal data according to the data information quality of each modal data;

[0202] Constructing a fusion feature corresponding to any round of training session according to the data weight of each modal data and the first modal feature of each modal data;

[0203] The data information quality includes information relevance and signal-to-noise ratio.

[0204] In an optional implementation, when the data quality is information relevance, determining the data information quality of each modal data according to the first modal feature of each modal data in any round of training session includes:

[0205] Determining, from the multiple rounds of training sessions, a context session related to any round of training session;

[0206] For any modal data in any round of training session, according to the data type of any modal data, determine the target modal data corresponding to the data type from the modal data of the context session, compare the first modal feature of any modal data with the first modal feature of the target modal data, and determine the information relevance of any modal data.

[0207] In an optional implementation, comparing the semantic feature with the fusion feature to obtain a second loss includes:

[0208] Decoding the semantic features to obtain reconstructed fusion features corresponding to the semantic features;

[0209] Comparing the reconstructed fusion feature with the fusion feature for similarity to obtain a first comparison result;

[0210] Comparing the semantic feature with the fusion feature for similarity to obtain a second comparison result;

[0211] The second loss is obtained according to the first comparison result and the second comparison result.

[0212] In an optional implementation, after adjusting the training parameters of the feature extraction model according to the first loss and the second loss to obtain the trained feature extraction model, the method further includes:

[0213] Inputting the multimodal real-time conversation into the trained feature extraction model to obtain modal features of each modal data in the multimodal real-time conversation;

[0214] Constructing a fusion feature corresponding to the multimodal real-time session according to the modal features of each modal data in the multimodal real-time session;

[0215] The fusion features corresponding to the multimodal real-time conversation are input into the preset question-answering model to obtain a conversation answer corresponding to the multimodal real-time conversation.

[0216] Each module in the multimodal data alignment device described above can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in hardware form, or can be stored in a memory in the computer device in software form, so that the processor can call and execute the corresponding operations of each module.

[0217] Figure 6 A block diagram of an electronic device provided in an embodiment of the present disclosure.

[0218] Reference Figure 6 An embodiment of the present disclosure provides an electronic device, which includes: at least one processor 601; at least one memory 602, and one or more I / O interfaces 603 connected between the processor 601 and the memory 602; wherein the memory 602 stores one or more computer programs that can be executed by the at least one processor 601, and the one or more computer programs are executed by the at least one processor 601 to enable the at least one processor 601 to perform the above-mentioned multimodal data alignment method.

[0219] Each module in the above-mentioned electronic device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each module.

[0220] The present disclosure also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the above-mentioned multimodal data alignment method. The computer-readable storage medium may be a volatile or non-volatile computer-readable storage medium.

[0221] An embodiment of the present disclosure also provides a computer program product, including a computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above-mentioned multimodal data alignment method.

[0222] It will be understood by those skilled in the art that all or some of the steps, systems, and functional modules / units in the methods disclosed above may be implemented as software, firmware, hardware, and appropriate combinations thereof. In a hardware implementation, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed by several physical components in cooperation. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or may be implemented as hardware, or may be implemented as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable storage medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or temporary medium).

[0223] As is well known to those skilled in the art, the term computer storage media includes volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information (such as computer-readable program instructions, data structures, program modules or other data). Computer storage media includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), static random access memory (SRAM), flash memory or other memory technology, portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical disc storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those skilled in the art, communication media typically contains computer-readable program instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.

[0224] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.

[0225] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, and conventional procedural programming languages ​​such as "C" language or similar programming languages. Computer-readable program instructions may be executed entirely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., utilizing an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be personalized by utilizing the state information of the computer-readable program instructions. The electronic circuit may execute the computer-readable program instructions, thereby realizing various aspects of the present disclosure.

[0226] The computer program product described herein may be implemented in hardware, software, or a combination thereof. In one embodiment, the computer program product is implemented as a computer storage medium. In another embodiment, the computer program product is implemented as a software product, such as a software development kit (SDK).

[0227] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0228] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0229] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0230] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and the part of the module, program segment or instruction contains one or more executable instructions for realizing the prescribed logical function. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the prescribed function or action, or can be implemented by a combination of dedicated hardware and computer instructions.

[0231] Example embodiments have been disclosed herein, and although specific terms are employed, they are used and should be interpreted only in a general illustrative sense and not for purposes of limitation. In some instances, it will be apparent to those skilled in the art that, unless otherwise expressly indicated, features, characteristics, and / or elements described in conjunction with a particular embodiment may be used alone or in combination with features, characteristics, and / or elements described in conjunction with other embodiments. Therefore, it will be understood by those skilled in the art that various changes in form and detail may be made without departing from the scope of the present disclosure as set forth in the appended claims.

Claims

1. A method for aligning multimodal data, characterized in that: include: Inputting the multimodal training session into a feature extraction model to be trained to obtain modal features of each modal data in the multimodal training session; performing comparative learning based on the modal features of each modal data in the multimodal training session to obtain a first loss; Constructing a fusion feature based on the modal features of each modal data in the multimodal training session, inputting the fusion feature into a preset question-answering model to obtain an answer to a question corresponding to the multimodal training session; Extracting semantic features of the answer to the question, and comparing the semantic features with the fusion features to obtain a second loss; According to the first loss and the second loss, the training parameters of the feature extraction model are adjusted to obtain a trained feature extraction model, so as to extract multiple mutually aligned modal features of the multimodal real-time conversation through the trained feature extraction model.

2. The method according to claim 1, characterized in that The feature extraction model includes a global feature extraction module and a local feature extraction module. Inputting the multimodal training session into the feature extraction model to be trained to obtain the modal features of each modal data in the multimodal training session includes: Performing global feature extraction on the multimodal training session according to the global feature extraction module to obtain a first modal feature of each modal data in the multimodal training session; performing local feature extraction on the multimodal training session according to the local feature extraction module to obtain a second modal feature of each modal data in the multimodal training session; The modal features include the first modal features and the second modal features.

3. The method according to claim 2, characterized in that The number of the second modal features of each modal data is multiple, and the comparative learning is performed according to the modal features of each modal data in the multimodal training session to obtain the first loss, including: For any second modal feature of any modal data, determine the positive sample modal feature and the negative sample modal feature respectively from the second modal features of the remaining modal data of the multimodal training session; calculate the similarity between the any second modal feature and the positive sample modal feature to obtain a first similarity value; calculate the similarity between the any second modal feature and the negative sample modal feature to obtain a second similarity value; and obtain the first loss based on the first similarity value and the second similarity value.

4. The method according to claim 2, characterized in that The multimodal training session includes multiple rounds of training sessions, each round of training session includes multiple modal data, and constructing fusion features based on the modal features of each modal data in the multimodal training session includes: For any round of training session, determining the data information quality of each modal data according to the first modal feature of each modal data in the any round of training session; Determining a data weight for each modal data according to the data information quality of each modal data; Constructing a fusion feature corresponding to any round of training session according to the data weight of each modal data and the first modal feature of each modal data; The data information quality includes information relevance and signal-to-noise ratio.

5. The method according to claim 4, characterized in that In a case where the data quality is information relevance, determining the data information quality of each modal data according to the first modal feature of each modal data in any round of training session includes: Determining, from the multiple rounds of training sessions, a context session related to any round of training session; For any modal data in any round of training session, according to the data type of any modal data, determine the target modal data corresponding to the data type from the modal data of the context session, compare the first modal feature of any modal data with the first modal feature of the target modal data, and determine the information relevance of any modal data.

6. The method according to any one of claims 1 to 5, characterized in that The comparing the semantic feature with the fusion feature to obtain a second loss includes: Decoding the semantic features to obtain reconstructed fusion features corresponding to the semantic features; Comparing the reconstructed fusion feature with the fusion feature for similarity to obtain a first comparison result; Comparing the semantic feature with the fusion feature for similarity to obtain a second comparison result; The second loss is obtained according to the first comparison result and the second comparison result.

7. The method according to any one of claims 1 to 5, characterized in that After adjusting the training parameters of the feature extraction model according to the first loss and the second loss to obtain a trained feature extraction model, the method further includes: Inputting the multimodal real-time conversation into the trained feature extraction model to obtain modal features of each modal data in the multimodal real-time conversation; Constructing a fusion feature corresponding to the multimodal real-time session according to the modal features of each modal data in the multimodal real-time session; The fusion features corresponding to the multimodal real-time conversation are input into the preset question-answering model to obtain a conversation answer corresponding to the multimodal real-time conversation.

8. A multimodal data alignment device, characterized in that: include: an extraction module, configured to input the multimodal training session into a feature extraction model to be trained, and obtain modal features of each modal data in the multimodal training session; a comparison module, configured to perform comparative learning based on the modal features of each modal data in the multimodal training session to obtain a first loss; a generation module, configured to construct a fusion feature based on the modal features of each modal data in the multimodal training session, input the fusion feature into a preset question-answering model, and obtain an answer to a question corresponding to the multimodal training session; a comparison module, configured to extract semantic features of the answer to the question, and compare the semantic features with the fusion features to obtain a second loss; A training module is used to adjust the training parameters of the feature extraction model according to the first loss and the second loss to obtain a trained feature extraction model, so as to extract multiple mutually aligned modal features of the multimodal real-time conversation through the trained feature extraction model.

9. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores one or more computer programs that can be executed by the at least one processor, and the one or more computer programs are executed by the at least one processor to enable the at least one processor to perform the multimodal data alignment method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the computer program implements the multimodal data alignment method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Multi-modal generative dialogue task processing method, device and equipment

    CN120932648A

  • Multi-modal data multi-model combined training method and system

    CN121144858A

  • Data processing method and device, storage medium and electronic equipment

    CN121327198A

  • Data processing method and device, storage medium and electronic equipment

    CN121327198B