Question and answer data processing method and system based on multi-modal large model

By deeply learning processing of the text problems and image modal contexts entered by users, feature alignment and fusion are achieved, and problems are solved in traditional multimodal question-and-answer systems are difficult to capture deep correlations, and answers that are logically complete and related to the question are generated, which improves the multimodal information processing capability of the question-and-answer system.

CN120030132AInactive Publication Date: 2025-05-23DATATANG(BEIJING)TECH CO LTD

Patent Information

Application Number
CN202510510268.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-05-23
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional multimodal question and answer systems are difficult to effectively process and fuse multimodal data, resulting in the inability of the model to capture deep-seated relationships across modalities, and the generated answers are often out of touch with the image content or are logically incomplete.

Method used

The multimodal data processing technology based on deep learning is adopted to perform semantic analysis of the text problems and image modal contexts input by users, extract semantic features, and realize feature alignment and fusion through linear projection and cross-modal feature global correlation interaction mechanism to generate text answers related to text problems.

Benefits of technology

Significantly improve the understanding and processing ability of the Q&A system for multimodal information, generate text answers that are closely related to text questions and are logically complete, and meet users' information needs in multimodal Q&A scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030132A_ABST
    Figure CN120030132A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent questions and answers, and particularly discloses a question and answer data processing method and system based on a multi-modal large model, and the method comprises the steps: carrying out the semantic analysis of a text question and an image modal context inputted by a user through employing a multi-modal data processing technology based on deep learning; the method comprises the following steps: respectively extracting semantic features of a text problem and an image modal context, then carrying out linear projection on the text problem and the image modal context to realize feature alignment, introducing a cross-modal feature global association interaction mechanism, mining deep semantic association between the text problem and the image modal context, and carrying out feature extraction on the text problem and the image modal context; effective fusion of the text question and the image modal context information is realized, and then a text answer related to the text question is generated by utilizing the reasoning ability of a large language model. In this way, the ability of the question-answering system to understand and process the multi-modal information can be significantly improved, the text answers closely related to the text questions and logically complete are generated, and the information requirements of the user in the multi-modal question-answering scene are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent question answering technology, and more specifically, to a question answering data processing method and system based on a multimodal large model. Background Art

[0002] In today's digital age, with the rapid development of information technology, people's demand for information acquisition has become increasingly diversified and efficient. As a tool that can directly provide users with accurate information, question-answering systems have been widely used in many fields, such as intelligent customer service, medical diagnosis assistance, and educational intelligent tutoring. However, most traditional question-answering systems can only process single-modal data, such as questions and answers in plain text, which greatly limits their ability to understand and respond to complex real-world scenarios. Taking the medical field as an example, when diagnosing a disease, doctors will not only refer to the patient's medical record text (symptom description, medical history, etc.), but also check medical images (X-rays, CT images, etc.), and make accurate diagnoses through comprehensive analysis of multimodal information. In this context, how to effectively process and fuse multimodal data has become a key challenge to improve the performance and intelligence level of question-answering systems. In recent years, multimodal big models have emerged to integrate data information from multiple modalities. Through powerful model architecture and training algorithms, they learn the intrinsic associations and semantic connections between different modal data, thereby achieving in-depth understanding and processing of multimodal information. In question-answering scenarios, multimodal big models can accept multimodal question-related information input by users, such as text questions combined with image context, and then generate accurate and comprehensive answers based on their learned knowledge and semantic understanding capabilities.

[0003] However, traditional multimodal question-answering systems usually use simple feature concatenation or weighted fusion to achieve information interaction between modalities. Since there is usually a semantic gap between text and image features in high-dimensional space, simple concatenation or linear operations make it difficult to model the global semantic relationship of multimodal information, resulting in the model being unable to capture deep cross-modal associations. In addition, large language models (such as GPT and PaLM) perform well in pure text reasoning tasks, but when they are directly transferred to multimodal scenarios, due to the lack of deep semantic alignment representation of cross-modal information, it is difficult to effectively associate visual clues in the image context with the implicit intent of the text question, and often generates answers that are disconnected from the image content or logically incomplete.

[0004] Therefore, we look forward to an optimized question-answering data processing method and system based on a multimodal large model. Summary of the invention

[0005] In order to solve the above technical problems, the present application is proposed. The embodiment of the present application provides a question-answering data processing method and system based on a multimodal large model, which uses a multimodal data processing technology based on deep learning to perform semantic analysis on the text questions and image modal contexts input by the user, so as to extract the semantic features of the text questions and image modal contexts respectively, and then further linearly project the two to achieve feature alignment, and introduce a cross-modal feature global correlation interaction mechanism to explore the deep semantic association between text questions and image modal contexts, and achieve effective fusion of text questions and image modal context information, and then on this basis, use the reasoning ability of the large language model to generate text answers related to text questions. In this way, the question-answering system's ability to understand and process multimodal information can be significantly improved, and text answers that are closely related to text questions and logically complete can be generated to meet users' information needs in multimodal question-answering scenarios.

[0006] According to one aspect of the present application, a method for processing question-answering data based on a multimodal large model is provided, which includes: Get the text question and image modal context entered by the user; Performing semantic alignment encoding on the text question and the image modal context to obtain an aligned text question semantic embedding encoding vector and an aligned image modal context semantic embedding encoding vector; Performing fine-grained cross-modal fusion based on a global association graph on the aligned text question semantic embedding coding vector and the aligned image modality context semantic embedding coding vector to obtain a question multimodal fusion coding vector; Generate a text answer based on the multimodal fusion encoding vector of the question.

[0007] According to another aspect of the present application, a question-answering data processing system based on a multimodal large model is provided, comprising: A user input module, used to obtain text questions and image modal contexts input by the user; A semantic alignment coding module, used for performing semantic alignment coding on the text question and the image modal context to obtain an aligned text question semantic embedding coding vector and an aligned image modal context semantic embedding coding vector; A fine-grained cross-modal fusion module is used to perform fine-grained cross-modal fusion on the aligned text question semantic embedding coding vector and the aligned image modality context semantic embedding coding vector based on a global association graph to obtain a question multimodal fusion coding vector; The text answer generation module is used to generate a text answer based on the multimodal fusion encoding vector of the question.

[0008] Compared with the prior art, the question-answering data processing method and system based on a multimodal large model provided by the present application uses a multimodal data processing technology based on deep learning to perform semantic analysis on the text questions and image modal contexts input by the user, so as to extract the semantic features of the text questions and image modal contexts respectively, and then further linearly project the two to achieve feature alignment, and introduce a cross-modal feature global correlation interaction mechanism to mine the deep semantic association between text questions and image modal contexts, and achieve effective fusion of text questions and image modal context information, and then on this basis, use the reasoning ability of the large language model to generate text answers related to text questions. In this way, the question-answering system's ability to understand and process multimodal information can be significantly improved, and text answers that are closely related to text questions and logically complete can be generated to meet users' information needs in multimodal question-answering scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] By describing the embodiments of the present application in more detail in conjunction with the accompanying drawings, the above and other purposes, features and advantages of the present application will become more apparent. The accompanying drawings are used to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the accompanying drawings, the same reference numerals generally represent the same components or steps.

[0010] Figure 1 It is a flowchart of a question and answer data processing method based on a multimodal large model according to an embodiment of the present application.

[0011] Figure 2 Schematic diagram of data flow of a question-and-answer data processing method based on a multimodal large model according to an embodiment of the present application.

[0012] Figure 3 This is a flowchart of sub-step S2 of the question and answer data processing method based on a multimodal large model according to an embodiment of the present application.

[0013] Figure 4 This is a flowchart of sub-step S3 of the question and answer data processing method based on a multimodal large model according to an embodiment of the present application.

[0014] Figure 5 It is a block diagram of a question and answer data processing system based on a multimodal large model according to an embodiment of the present application. DETAILED DESCRIPTION

[0015] As shown in this application and claims, unless the context clearly indicates an exception, the words "a", "an", "an" and / or "the" do not refer to the singular and may also include the plural. Generally speaking, the terms "include" and "comprise" only indicate the inclusion of the steps and elements that have been clearly identified, and these steps and elements do not constitute an exclusive list. The method or device may also include other steps or elements.

[0016] Although the present application makes various references to certain modules in the system according to the embodiments of the present application, any number of different modules can be used and run on the user terminal and / or server. The modules are only illustrative, and different aspects of the system and method can use different modules.

[0017] Flowcharts are used in the present application to illustrate the operations performed by the system according to the embodiments of the present application. It should be understood that the preceding or following operations are not necessarily performed accurately in order. On the contrary, various steps may be processed in reverse order or simultaneously as required. Meanwhile, other operations may also be added to these processes, or a certain step or several steps of operations may be removed from these processes.

[0018] Below, the exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described here.

[0019] It is worth noting that in this application, all actions to obtain data are carried out in compliance with the relevant data protection laws and policies of the country where the data is located, and with the authorization given by the owner of the corresponding device.

[0020] In response to the technical problems described in the above background technology, this application proposes a question-answering data processing method based on a multimodal large model, which uses a multimodal data processing technology based on deep learning to perform semantic analysis on the text questions and image modal contexts input by the user, so as to extract the semantic features of the text questions and image modal contexts respectively, and then further linearly project the two to achieve feature alignment, and introduce a cross-modal feature global correlation interaction mechanism to explore the deep semantic association between text questions and image modal contexts, and achieve effective fusion of text questions and image modal context information. On this basis, the reasoning ability of the large language model is used to generate text answers related to the text questions. In this way, the question-answering system's ability to understand and process multimodal information can be significantly improved, and text answers that are closely related to text questions and logically complete can be generated to meet users' information needs in multimodal question-answering scenarios.

[0021] Figure 1It is a flowchart of a question and answer data processing method based on a multimodal large model according to an embodiment of the present application. Figure 2 Schematic diagram of data flow of a question-answering data processing method based on a multimodal large model according to an embodiment of the present application. Figure 1 and Figure 2 As shown, the question-answering data processing method based on the multimodal large model includes the steps of: S1, obtaining a text question and an image modal context input by a user; S2, performing semantic alignment encoding on the text question and the image modal context to obtain an aligned text question semantic embedding encoding vector and an aligned image modal context semantic embedding encoding vector; S3, performing fine-grained cross-modal fusion based on a global association graph on the aligned text question semantic embedding encoding vector and the aligned image modal context semantic embedding encoding vector to obtain a question multimodal fusion encoding vector; S4, generating a text answer based on the question multimodal fusion encoding vector.

[0022] In the above-mentioned question-answering data processing method based on a multimodal large model, the step S1 obtains the text question and image modal context input by the user. It should be understood that in practical applications, when users seek answers to questions, they often use a variety of information carriers to accurately express their needs. Taking the smart shopping assistant scenario as an example, the user wants to buy a favorite piece of clothing, but it may not be possible to fully convey what he wants through text description alone. For example, the user may describe "need a loose-fitting dress with a retro floral pattern", but there are certain limitations in the language description of the specific style of "retro floral pattern". At this time, if the user can provide a picture of a similar style as an image modal context, it can greatly supplement and refine the demand information. Therefore, supporting users to input questions in a combination of text and images can significantly improve the information acquisition efficiency and answer accuracy of the question-answering system.

[0023] In the specific implementation process, when the user starts the question-answering system and prepares to enter questions, the system interface should provide clear and explicit prompts and instructions to let the user know that text and image information can be entered at the same time to express the question. The user enters a text description of the question in the text input box, such as "I need a loose-fitting dress with a retro floral pattern." At this time, the system will receive and store this text information in real time. At the same time, the system interface should also provide an image upload entry. The user can click the upload button to select a picture related to the question from the local device as the image modal context. After the user selects the picture, the system will start the image upload process and transmit the image data to the server for subsequent processing.

[0024] When users upload images, the system needs to perform preliminary processing and verification on the image data. First, it is necessary to check whether the image format meets the requirements, such as whether it is a common image format such as JPEG, PNG, etc. If it does not meet the requirements, the user will be prompted to re-upload. For images that meet the format requirements, the size and resolution of the image will be further detected to ensure that the image data will not affect performance and effects due to being too large or too small during subsequent processing. If the image is too large, it needs to be compressed automatically to meet the needs of system processing; if the image is too small, it may cause loss of details and affect subsequent semantic parsing and fusion. At this time, the system will prompt the user to upload a clearer picture, or try to enlarge the image to a certain extent, but it is necessary to ensure that the image quality will not be significantly reduced.

[0025] In the process of obtaining the text questions and image modal context entered by the user, the interactive design of the system is also crucial. For example, the text input box should be large enough to facilitate users to enter longer text descriptions; the image upload button should be obvious and easy to operate, and a clear progress prompt should be provided during the upload process to let users know the upload status. In addition, the system should also have a certain fault tolerance. When there are problems with the text entered by the user or the uploaded image, the system should be able to give clear prompt information in time and guide the user to perform the correct operation instead of simply rejecting or reporting an error. For example, when the image format uploaded by the user is not supported, the system can prompt the user to support the image format and suggest that the user convert the format and re-upload; when the text entered by the user is too short or unclear, the system can prompt the user to add more detailed information so that the system can understand the problem more accurately.

[0026] In the above-mentioned question-answering data processing method based on a multimodal large model, the step S2 performs semantic alignment encoding on the text question and the image modal context to obtain an aligned text question semantic embedding encoding vector and an aligned image modal context semantic embedding encoding vector. Figure 3 FIG. 1 is a flowchart of sub-step S2 of the question-answer data processing method based on a multimodal large model according to an embodiment of the present application. Figure 3 As shown, the step S2 includes the steps of: S21, using a text encoder to perform semantic embedding encoding on the text question to obtain a text question semantic embedding encoding vector; S22, using an image encoder to perform visual semantic embedding encoding on the image modal context to obtain an image modal context semantic embedding encoding vector; S23, based on a linear projection layer, performing cross-modal feature alignment and mapping on the text question semantic embedding encoding vector and the image modal context semantic embedding encoding vector to obtain the aligned text question semantic embedding encoding vector and the aligned image modal context semantic embedding encoding vector.

[0027] Specifically, in step S21, the text question is semantically embedded and encoded using a text encoder to obtain a text question semantic embedding encoding vector. Specifically, considering that the semantic analysis of the text question is the key to understanding the user's intention. In this regard, the present application uses a text encoder based on deep learning technology to semantically embed the text question to convert it into a dense vector representation in a high-dimensional semantic space, that is, a text question semantic embedding encoding vector. In a specific example of the present application, the text encoder is a Bert model. As a pre-trained deep bidirectional model, the Bert model can learn rich language knowledge and context dependencies through training on a large-scale corpus, thereby achieving accurate semantic analysis of text questions. For example, for the question "What is the action of the character in the picture?", the model will extract high-dimensional features of core words such as "character" and "action", and encode their dependencies (such as the question word "what" points to the target to be identified). In this way, through the processing of the Bert model, the text question input by the user is converted into a text question semantic embedding encoding vector containing rich semantic information, which provides an effective information basis for subsequent multimodal fusion.

[0028] Specifically, in step S22, the image modal context is visually semantically embedded and encoded using an image encoder to obtain an image modal context semantic embedding encoding vector. It should be understood that since the original image data exists in the form of a pixel matrix, although it contains rich visual information, it is difficult to be directly used for cross-modal information fusion. Therefore, the present application further uses an image encoder to extract high-level semantic features of the image modal context. In a specific example of the present application, the image encoder is a ViT (Vision Transformer) model. The ViT model divides the image into a series of small patches, linearly embeds these small patches into a high-dimensional space, and then uses the Transformer architecture to learn the global dependencies between the embedded features of each image patch, thereby achieving high-level semantic understanding of the image. In this way, the image modal context is converted into an image modal context semantic embedding encoding vector containing rich visual semantic information, which provides an important visual information foundation for subsequent multimodal fusion.

[0029] Specifically, the step S23 performs cross-modal feature alignment and mapping on the text question semantic embedding coding vector and the image modal context semantic embedding coding vector based on the linear projection layer to obtain the aligned text question semantic embedding coding vector and the aligned image modal context semantic embedding coding vector. Specifically, since text and image features are usually located in different high-dimensional spaces (the text question semantic embedding coding vector is based on the word embedding space, and the image modal context semantic embedding coding vector is based on the visual feature space), there is a certain semantic gap between the two, which makes it difficult to effectively capture the deep cross-modal associations by directly performing feature fusion. In this regard, the present application introduces a linear projection layer, which maps the text question semantic embedding coding vector and the image modal context semantic embedding coding vector to the same semantic space through a set of learnable linear mapping matrices to force the alignment of the feature distributions of the two, so that the text question semantic embedding coding vector and the image modal context semantic embedding coding vector have higher comparability and fusion in the new feature space.

[0030] In the above-mentioned question-answering data processing method based on a multimodal large model, the step S3 performs fine-grained cross-modal fusion based on the global association map on the aligned text question semantic embedding coding vector and the aligned image modal context semantic embedding coding vector to obtain a question multimodal fusion coding vector. It should be understood that in complex question-answering scenarios, single modality information is often difficult to fully reflect the real needs of users. Although the traditional feature splicing or simple addition method realizes the fusion of cross-modal information to a certain extent, it is difficult to effectively model the global semantic association between the visual clues in the image context and the text question intent, and the information fusion effect is not good, which may cause the subsequently generated text answers to be out of touch with the image content. In this regard, the present application proposes a fine-grained cross-modal fusion mechanism based on a global association graph, which performs feature principal component analysis on the aligned text question semantic embedding coding vector and the aligned image modal context semantic embedding coding vector to achieve feature principal component kernel semantic association coding between the two, and further constructs a global association graph based on the principal component feature association fusion information of the two, and performs correlation topological modeling on each principal component interaction feature between the text question and the image modal context to achieve global context perception of the associated interaction information between the two, thereby focusing on the essential semantic commonalities of cross-modal information, generating a multimodal fusion coding vector for the question, and providing strong support for the subsequent generation of accurate text answers. Among them, Figure 4 FIG. 1 is a flowchart of sub-step S3 of the question-answer data processing method based on a multimodal large model according to an embodiment of the present application. Figure 4As shown, the step S3 includes the steps of: S31, performing feature principal component-based kernel association coding on the aligned text question semantic embedding coding vector and the aligned image modal context semantic embedding coding vector to obtain a set of text question-image context principal component feature kernel association coding vectors; S32, constructing a text question-image context feature performance operator association topology matrix based on the correlation topological structure within the set of text question-image context principal component feature kernel association coding vectors; S33, performing global significant association interactive aggregation coding on the set of text question-image context principal component feature kernel association coding vectors based on the text question-image context feature performance operator association topology matrix to obtain the question multimodal fusion coding vector.

[0031] Specifically, in a specific example of the present application, the step S31 includes: first, performing feature principal component analysis on the aligned text question semantic embedding coding vector and the aligned image modality context semantic embedding coding vector to obtain a set of text question feature principal component coding vectors and a set of image modality context feature principal component coding vectors, which are expressed as follows:

[0032]

[0033] in, represents the semantic embedding encoding vector of the aligned text question, represents the principal component analysis network, is the covariance matrix of the aligned text question semantic embedding encoding vector, Express The matrix composed of the set of principal component encoding vectors of text question features obtained by eigenvalue decomposition, represents transpose, Express The diagonal matrix composed of the set of principal component eigenvalues ​​of the text question obtained by eigenvalue decomposition, and Respectively represent the first and the second in the set of principal component eigenvalues ​​of the text question eigenvalues, , and Respectively The first, second and text question feature principal component encoding vector, is the number of principal component encoding vectors of the text question features, represents the aligned image modality context semantic embedding encoding vector, is the covariance matrix of the aligned image modality context semantic embedding encoding vector, Express The matrix composed of the set of principal component encoding vectors of the image modality context features obtained by eigenvalue decomposition, represents transpose, Express The diagonal matrix consisting of the set of eigenvalues ​​of the image context principal component obtained by eigenvalue decomposition, and Respectively represent the first and the second in the set of image context principal component eigenvalues eigenvalues, , and Respectively The first, second and The principal component encoding vector of image modality context features.

[0034] That is, through feature principal component analysis, the aligned text question semantic embedding encoding vector and the aligned image modal context semantic embedding encoding vector can be orthogonally transformed, and the high-dimensional features can be mapped into a set of linearly independent principal component spaces, thereby eliminating the redundant correlation between features and achieving feature decoupling. Specifically, through dimensionality reduction and feature structure reorganization, the non-redundant, low-dimensional semantic association patterns implicit in the text and image modal features are revealed, so that the subsequent cross-modal fusion stage can more accurately mine the global correlation between text questions and image contexts based on the deep semantic structure represented by the principal component encoding vector. Based on the generated set of principal component encoding vectors of text question features and the set of principal component encoding vectors of image modal context features, the key principal components are retained, the feature dimension is reduced, and the computational burden brought by the direct fusion of high-dimensional features is alleviated. At the same time, the collinear interference between features is suppressed through orthogonal transformation, thereby enhancing the model's ability to cross the cross-modal semantic gap and improving the logical integrity of answer generation and the robustness of multimodal information fusion.

[0035] Then, each corresponding set of the text question feature principal component coding vector and the image modality context feature principal component coding vector in the set of the text question feature principal component coding vector and the set of the image modality context feature principal component coding vector is respectively input into the principal component kernel association coding network to obtain the set of kernel association coding vectors between the text question and image context principal component features, which is expressed as follows:

[0036] in, express The text question feature principal component encoding vector, express The The principal component encoding vector of image modality context features , represents the norm of a vector, and For different weighting parameters, express and Kernel correlation encoding vector between principal component features of text question-image context.

[0037] That is, through the principal component kernel association coding network, the nonlinear expression ability of its network structure is utilized to map the low-dimensional text question feature principal component coding vector and the image modal context feature principal component coding vector to a high-dimensional feature space, where the potential association patterns between principal components can be explicitly decoupled and enhanced for modeling. Specifically, through deep nonlinear transformation, the representation limitations of linear space are broken through, and a semantic association feature representation between principal components across text and image modalities is constructed, so that the cross-modal association relationship that was originally difficult to directly measure in the low-dimensional principal component space is converted into a quantifiable and computable vector encoding form, generating a set of kernel association coding vectors between text question-image context principal component features, providing a feature foundation with high-order semantic association for the subsequent construction of a cross-modal global association map.

[0038] Specifically, in a specific example of the present application, the step S32 includes: calculating the performance operator between any two kernel association coding vectors between the text question and the image context principal component features in the set of kernel association coding vectors between the text question and the image context principal component features to obtain the text question-image context feature performance operator association topology matrix composed of multiple performance operators. More specifically, the position difference vector between any two kernel association coding vectors between the text question and the image context principal component features in the set of kernel association coding vectors between the text question and the image context principal component features is calculated, and the square root of the sum of the squares of the eigenvalues ​​in the position difference vector is used as the performance operator between the two kernel association coding vectors between the text question and the image context principal component features. The above step S32 is expressed by the formula:

[0039]

[0040] in, , , and Respectively represent the first and the second kernel association encoding vectors between the text question and the image context principal component features. , and A kernel correlation encoding vector between the principal component features of the text question-image context, express Middle The eigenvalues ​​at the positions, express Middle The eigenvalues ​​at the positions, is the feature scale value of the kernel correlation encoding vector between the text question-image context principal component features, represents the performance operator calculation function, Represents the text question-image context feature performance operator association topology matrix.

[0041] That is, by introducing the square root of the sum of squares of position-based differential vectors as a performance operator (i.e., the Euclidean distance metric), the similarity differences between different feature combinations can be quantified from the perspective of geometric space. Specifically, by converting the discrete text-image feature association relationship in the high-dimensional feature space into a computable text question-image context feature performance operator association topological matrix structure, the subsequent graph neural network can perform information propagation and aggregation based on the adjacency relationship of the text question-image context feature performance operator association topological matrix, thereby exploring the potential synergy between multimodal features.

[0042] In particular, in a preferred example of the present application, the step S33 includes: first, based on the text question-image context feature performance operator association topology matrix, each text question-image context principal component feature kernel association coding vector in the set of text question-image context principal component feature kernel association coding vectors is subjected to feature space distribution adaptation calibration to obtain a set of optimized text question-image context principal component feature kernel association coding vectors, which is expressed as follows:

[0043]

[0044]

[0045]

[0046] in, represents the matrix multiplication operation, express The initialization space transition calibration matrix, express The corresponding compactified text problem - kernel correlation encoding vector between image context principal component features, Represents the kernel correlation encoding matrix between the compactified text problem-image context principal component features, represents the difference by position, represents the F-norm of the matrix, express and The variance of the set of all matrix values ​​of , for The corresponding iterative spatial transition calibration matrix, express Corresponding optimization text problem - kernel correlation encoding vector between image context principal component features.

[0047] That is, by constructing an initialized spatial transition calibration matrix as the spatial standard transition specification field, the kernel correlation coding vector between the text question-image context principal component features is multiplied by the initialized spatial transition calibration matrix to generate a compactified text question-image context principal component feature kernel correlation coding vector, and then a compactified text question-image context principal component feature kernel correlation coding matrix is ​​formed through two-dimensional splicing, and then its Gaussian correlation coefficient with the text question-image context feature performance operator correlation topological matrix is ​​calculated and iteratively optimized to eliminate the spatial distribution disorder caused by the random potential field.

[0048] Through this iterative optimization method, the spatial transition calibration matrix is ​​used to remap the discrete and disordered kernel correlation coding vectors between the principal component features of the text question-image context to a canonical space compatible with the topological matrix, so that the spatial distribution of the kernel correlation coding vectors between the principal component features of the text question-image context is geometrically aligned with the text question-image context feature performance operator correlation topological matrix, so as to facilitate information propagation and aggregation in a unified topological form space based on the calibrated optimized kernel correlation coding vectors between the principal component features of the text question-image context. This not only solves the problem of information transmission distortion caused by dimensionality breakage, but also strengthens the continuity of cross-modal features in the topological space through Gaussian correlation coefficient constraints, so that the subsequent data processing process can capture the deep semantic topological relationship implicit in the multimodal data, thereby improving the model's decoding accuracy of complex cross-modal interaction logic.

[0049] Then, the set of kernel association encoding vectors between the optimized text question-image context principal component features and the text question-image context feature performance operator association topology matrix are input into the graph convolutional neural network model to obtain a set of global significant interaction encoding vectors of text question-image context features, which is expressed as follows:

[0050] in , is a graph convolutional neural network model, A collection of vectors encoding global significant interactions between text question and image context features.

[0051] That is, by optimizing the set of kernel correlation encoding vectors between the principal component features of the text question-image context as graph node features, and inputting the text question-image context feature performance operator correlation topological matrix as the adjacency matrix into the GCN model, it is possible to utilize the unique message passing mechanism of the graph structure to perform multi-hop interactions between each feature node and its topological neighbors. In this way, the generated set of global significant interaction encoding vectors of the text question-image context features, through the hierarchical information aggregation of graph convolution, integrates the global correlation pattern of cross-modal features while retaining the semantics of local features, and captures the implicit collaborative relationship between cross-modal feature nodes.

[0052] Finally, the set of global significant interaction coding vectors of the text question-image context features is feature cascaded to obtain the multimodal fusion coding vector of the question. It should be understood that through the feature cascade operation, the global significant interaction coding vectors of text question-image context features of different dimensions can be spliced ​​along the feature axis, so as to achieve lossless aggregation of multi-granular cross-modal correlation information. Specifically, by fusing feature expressions at different levels, a joint representation with both local details and global consistency can be formed, providing sufficient multimodal semantic support for the subsequent large language model to generate answers.

[0053] In the above-mentioned question-answering data processing method based on the multimodal large model, the step S4 generates a text answer based on the multimodal fusion coding vector of the question. In a specific example of the present application, the step S4 includes: inputting the multimodal fusion coding vector of the question into the reasoning module based on the large language model to obtain the text answer. That is, using the reasoning ability of the large language model, accurate and coherent text answers are generated based on multimodal fusion information to answer user questions. Specifically, the large language model (such as the GPT series model) is based on the Transformer architecture, and learns the grammar, semantics and logical structure of the language through pre-training of a large amount of text data. After receiving the multimodal fusion coding vector of the question, the model uses its own powerful language generation ability and context understanding ability to deeply analyze and reason the multimodal fusion coding vector of the question through the internal multi-layer attention mechanism and feedforward neural network, and gradually constructs the semantic structure and logical chain related to the question, thereby generating a text sequence that conforms to the language rules and question logic to meet the user's information acquisition needs.

[0054] In summary, a question-and-answer data processing method based on a multimodal large model according to an embodiment of the present application is explained, which uses a multimodal data processing technology based on deep learning to perform semantic analysis on the text questions and image modal contexts input by the user, so as to extract the semantic features of the text questions and image modal contexts respectively, and then further linearly project the two to achieve feature alignment, and introduce a cross-modal feature global correlation interaction mechanism to explore the deep semantic association between text questions and image modal contexts, and achieve effective fusion of text questions and image modal context information, and then on this basis, use the reasoning ability of the large language model to generate text answers related to text questions. In this way, the question-and-answer system's ability to understand and process multimodal information can be significantly improved, and text answers that are closely related to text questions and logically complete can be generated to meet users' information needs in multimodal question-and-answer scenarios.

[0055] Furthermore, a question-answering data processing system based on a multimodal large model is also provided.

[0056] Figure 5 FIG. 1 is a block diagram of a question-answering data processing system based on a multimodal large model according to an embodiment of the present application. Figure 5 As shown, according to an embodiment of the present application, a question-answering data processing system 100 based on a multimodal large model includes: a user input module 110, used to obtain a text question and an image modal context input by a user; a semantic alignment encoding module 120, used to perform semantic alignment encoding on the text question and the image modal context to obtain an aligned text question semantic embedding encoding vector and an aligned image modal context semantic embedding encoding vector; a fine-grained cross-modal fusion module 130, used to perform fine-grained cross-modal fusion based on a global association graph on the aligned text question semantic embedding encoding vector and the aligned image modal context semantic embedding encoding vector to obtain a question multimodal fusion encoding vector; and a text answer generation module 140, used to generate a text answer based on the question multimodal fusion encoding vector.

[0057] Here, those skilled in the art can understand that the specific operations of each module in the above-mentioned question-answering data processing system based on the multimodal large model have been described in the above reference. Figures 1 to 4 It has been introduced in detail in the description of the question and answer data processing method based on the multimodal large model, and therefore, its repeated description will be omitted. The basic principle of the present invention is described above in conjunction with specific embodiments. However, it should be pointed out that the advantages, strengths, effects, etc. mentioned in the present invention are only examples and not limitations, and it cannot be considered that these advantages, strengths, effects, etc. must be possessed by each embodiment of the present invention. In addition, the specific details of the above embodiments are only for the purpose of illustration and facilitation of understanding, rather than limitation, and the above details do not limit the present invention to being implemented by adopting the above specific details.

[0058] In the above embodiments, the description of each embodiment has its own emphasis. For the parts that are not described or recorded in detail in a certain embodiment, please refer to the relevant description of other embodiments. In the several embodiments provided by the present invention, it should be understood that the disclosed system and method can be implemented in other ways. For example, the system embodiment described above is only schematic. For example, the unit division is only a logical function division, and there may be other division methods in actual implementation. The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the scheme of this embodiment.

[0059] It will be apparent to those skilled in the art that the invention is not limited to the details of the exemplary embodiments described above and that the invention can be implemented in other specific forms without departing from the spirit or essential features of the invention. Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description, and it is intended that all variations falling within the meaning and scope of the equivalent elements of the claims be included in the invention. Any reference to a figure in a claim should not be considered as limiting the claim to which it relates.

[0060] In addition, it is obvious that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units stated in the system claims can also be implemented by one unit through software or hardware.

[0061] Finally, it should be noted that the above description has been given for the purpose of illustration and description. In addition, the above embodiments are only used to illustrate the technical solution of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, it should be understood by those skilled in the art that the technical solution of the present invention can be modified or replaced by equivalents without departing from the spirit and scope of the technical solution of the present invention.

Claims

1. A method for processing question-answering data based on a multimodal large model, characterized in that: include: Get the text question and image modal context entered by the user; Performing semantic alignment encoding on the text question and the image modal context to obtain an aligned text question semantic embedding encoding vector and an aligned image modal context semantic embedding encoding vector; Performing fine-grained cross-modal fusion based on a global association graph on the aligned text question semantic embedding coding vector and the aligned image modality context semantic embedding coding vector to obtain a question multimodal fusion coding vector; Generate a text answer based on the multimodal fusion encoding vector of the question.

2. The method for processing question and answer data based on a multimodal large model according to claim 1, characterized in that: The text question and the image modal context are semantically aligned and encoded to obtain an aligned text question semantic embedding encoding vector and an aligned image modal context semantic embedding encoding vector, including: Using a text encoder to perform semantic embedding encoding on the text question to obtain a text question semantic embedding encoding vector; Using an image encoder to perform visual semantic embedding encoding on the image modality context to obtain an image modality context semantic embedding encoding vector; Based on a linear projection layer, cross-modal feature alignment and mapping are performed on the text question semantic embedding coding vector and the image modal context semantic embedding coding vector to obtain the aligned text question semantic embedding coding vector and the aligned image modal context semantic embedding coding vector.

3. The method for processing question and answer data based on a multimodal large model according to claim 2, characterized in that: The text encoder is a Bert model, and the image encoder is a ViT model.

4. The method for processing question and answer data based on a multimodal large model according to claim 3, characterized in that: The aligned text question semantic embedding coding vector and the aligned image modality context semantic embedding coding vector are subjected to fine-grained cross-modal fusion based on a global association graph to obtain a question multimodal fusion coding vector, including: Performing kernel association coding based on feature principal components on the aligned text question semantic embedding coding vector and the aligned image modality context semantic embedding coding vector to obtain a set of kernel association coding vectors between text question-image context principal component features; Based on the correlation topological structure within the set of kernel correlation coding vectors between the text question-image context principal component features, construct a text question-image context feature performance operator correlation topological matrix; Based on the text question-image context feature performance operator association topology matrix, the set of kernel association coding vectors between the text question-image context principal component features is subjected to global significant association interaction aggregation coding to obtain the question multimodal fusion coding vector.

5. The method for processing question and answer data based on a multimodal large model according to claim 4, characterized in that: The aligned text question semantic embedding coding vector and the aligned image modality context semantic embedding coding vector are subjected to kernel association coding based on feature principal components to obtain a set of kernel association coding vectors between text question and image context principal component features, including: Performing feature principal component analysis on the aligned text question semantic embedding coding vector and the aligned image modality context semantic embedding coding vector to obtain a set of text question feature principal component coding vectors and a set of image modality context feature principal component coding vectors; Each corresponding group of text question feature principal component coding vectors and image modal context feature principal component coding vectors in the set of text question feature principal component coding vectors and the set of image modal context feature principal component coding vectors are respectively input into the principal component kernel association coding network to obtain the set of kernel association coding vectors between the text question-image context principal component features.

6. The method for processing question and answer data based on a multimodal large model according to claim 5, characterized in that: Based on the internal correlation topological structure of the set of kernel correlation coding vectors between the text question-image context principal component features, a text question-image context feature performance operator correlation topological matrix is ​​constructed, including: Calculate the performance operator between any two text question-image context principal component feature kernel association coding vectors in the set of the text question-image context principal component feature kernel association coding vectors to obtain the text question-image context feature performance operator association topology matrix composed of multiple performance operators.

7. The method for processing question and answer data based on a multimodal large model according to claim 6, characterized in that: Calculating the performance operator between any two text question-image context principal component feature kernel association encoding vectors in the set of the text question-image context principal component feature kernel association encoding vectors to obtain the text question-image context feature performance operator association topology matrix composed of multiple performance operators, including: Calculate the positional difference vector between any two kernel association coding vectors between text question and image context principal component features in the set of kernel association coding vectors between text question and image context principal component features, and use the square root of the sum of squares of each eigenvalue in the positional difference vector as the performance operator between the two kernel association coding vectors between text question and image context principal component features.

8. The method for processing question and answer data based on a multimodal large model according to claim 7, characterized in that: Based on the text question-image context feature performance operator association topology matrix, a set of kernel association coding vectors between the text question-image context principal component features is subjected to global significant association interaction aggregation coding to obtain the question multimodal fusion coding vector, including: Based on the text question-image context feature performance operator association topology matrix, each text question-image context principal component feature kernel association coding vector in the set of text question-image context principal component feature kernel association coding vectors is subjected to feature space distribution adaptation calibration to obtain a set of optimized text question-image context principal component feature kernel association coding vectors; Inputting the set of kernel association coding vectors between the optimized text question-image context principal component features and the text question-image context feature performance operator association topology matrix into a graph convolutional neural network model to obtain a set of global significant interaction coding vectors of text question-image context features; The set of global significant interaction encoding vectors of the text question-image context features is feature concatenated to obtain the multimodal fusion encoding vector of the question.

9. The method for processing question and answer data based on a multimodal large model according to claim 8, characterized in that: Based on the multimodal fusion encoding vector of the question, a text answer is generated, including: The multimodal fusion encoding vector of the question is input into a reasoning module based on a large language model to obtain the text answer.

10. A question-answering data processing system based on a multimodal large model, characterized in that: include: A user input module, used to obtain text questions and image modal contexts input by the user; A semantic alignment coding module, used for performing semantic alignment coding on the text question and the image modal context to obtain an aligned text question semantic embedding coding vector and an aligned image modal context semantic embedding coding vector; A fine-grained cross-modal fusion module is used to perform fine-grained cross-modal fusion on the aligned text question semantic embedding coding vector and the aligned image modality context semantic embedding coding vector based on a global association graph to obtain a question multimodal fusion coding vector; The text answer generation module is used to generate a text answer based on the multimodal fusion encoding vector of the question.

Citation Information

Patent Citations

  • Question and answer method and device based on multi-modal information and application of question and answer method and device

    CN117828142A

  • Medical visual question and answer method and system based on multi-task modeling

    CN119202334A

  • Multi-modal large model training data acquisition method and system

    CN119380144A

  • Combined visual question and answer method based on core-to-global semantic fusion reasoning

    CN119397384A

  • AI dialogue system based on multi-modal input

    CN119831058A

Cited By

  • Visual question and answer method and device based on multiple document images

    CN120508686A

  • Insurance risk dynamic analysis system for multi-modal biological data

    CN120636805A

  • Large model sensitive large content filtering method in learning scene

    CN120821836A

  • Multi-source data-based spinning melt rheological control method and system

    CN120848389A