Multi-modal model analysis enhancement method, system and equipment and storage medium
By building a type knowledge base and combining image and text information processing methods, the problem that multimodal models are difficult to utilize domain knowledge in handling specific domain problems is solved, and more accurate and in-depth answer generation is achieved, improving user experience and system intelligence.
Patent Information
- Application Number
- CN202510063122.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-13
AI Technical Summary
Existing multimodal model analysis techniques are difficult to effectively utilize domain knowledge, resulting in low accuracy and reliability of answers when dealing with specific domain questions.
By building a type knowledge base in the target field, combining image and text information, using neural network models to identify the target type to which the image belongs, and matching relevant knowledge in the type knowledge base, filling it into the original text to generate more accurate answers.
It improves the accuracy and depth of the answers to questions, enhances the automation and intelligence level of user experience and system, and can make more effective use of domain knowledge and improves the application effect of specific fields.
Smart Images

Figure CN119988868A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of data analysis, and specifically to a method, system, device and storage medium for enhancing multimodal model analysis. Background Art
[0002] In recent years, multimodal model analysis technology has been widely used in various fields, especially in intelligent question answering, image recognition and natural language processing. With the rapid development of deep learning technology, multimodal models can better understand and process complex multi-source data, thereby improving the accuracy and robustness of the system. This technology not only improves user experience, but also brings significant economic benefits to enterprises, especially in medical diagnosis, financial analysis and educational counseling.
[0003] In existing multimodal model analysis technologies, image and text data are generally processed directly through a single neural network model. This method is simple and intuitive, but has limited effect when processing complex and diverse data; or image and text data are processed separately first, and then the results of the two are fused. This method can make full use of the advantages of each, but there are also problems such as information loss and insufficient fusion. These existing multimodal model analysis methods generally have the problem of being unable to effectively utilize domain knowledge when dealing with problems in specific fields. Due to the lack of in-depth understanding of the target field, the model often performs poorly when dealing with specific problems, resulting in low accuracy and reliability of the answers.
[0004] Therefore, how to effectively integrate domain knowledge into multimodal models and improve the application effect of models in specific fields is a technical problem that needs to be solved urgently. Summary of the invention
[0005] The present application provides a method, system, device and storage medium for enhancing multimodal model analysis, which improves the accuracy and depth of problem answering by combining image and text information and using a type knowledge base to provide professional knowledge support, while enhancing the user experience and the automation and intelligence level of the system.
[0006] In a first aspect of the present application, a method for enhancing multimodal model analysis is provided, which is applied to a model analysis platform, and the method comprises: Constructing a type knowledge base of a target domain, wherein the type knowledge base includes classification types of the target domain and knowledge corresponding to the classification types; When receiving a question input by a user, split the question into an image and original text, input the image into a preset neural network model to obtain the target type to which the image belongs, and match the corresponding target knowledge in the type knowledge base according to the target type; The target knowledge is filled into a preset position in the original text to obtain a filled text, and the filled text and the image are input into a preset multimodal large model to obtain an answer to the question.
[0007] Optionally, inputting the image into a preset neural network model to obtain the target type to which the image belongs includes: Extracting preliminary features of the image through a plurality of convolutional layers in the preset neural network model, wherein the preliminary features include edge, texture or color information of the image; Capturing intermediate layer features of different scales of the image according to the preliminary features through multiple parallel Inception modules; Using an auxiliary classifier to classify the plurality of intermediate layer features to obtain a plurality of categories; Through the fully connected layer and the softmax layer, the probability of each category is output, and the category with the highest probability is selected as the target type to which the image belongs.
[0008] Optionally, outputting the probability of each category through the fully connected layer and the softmax layer, and selecting the category with the highest probability as the target type to which the image belongs includes: Generate a feature vector according to the intermediate layer features, input the feature vector into a fully connected layer, and generate an original score vector equal to the number of categories; Normalizing the original score vector using a softmax function to generate a normalized probability vector; The category corresponding to the maximum value in the probability vector is selected as the classification result of the image.
[0009] Optionally, filling the target knowledge into a preset position in the original text to obtain a filled text, and inputting the filled text and the image into a preset multimodal macro model to obtain an answer to the question includes: Traversing the original text to determine preset identifiers, and filling the target knowledge into positions between the preset identifiers to obtain a filling text; Integrating the filler text and the image into a unified data structure, and inputting the data structure into the preset multimodal macro model; Acquire text features through a text encoder, acquire image features through an image encoder, and fuse the text features and the image features to generate a cross-modal joint representation; An answer to the question is formed based on the joint representation.
[0010] Optionally, the method further includes training the preset multimodal large model, specifically including: The target element set of historical filler text and historical image is extracted through the preset multimodal large model, the similarity between the target element set and the preset element set is calculated, and the parameters of the multimodal large model are adjusted according to the similarity. The preset element set is a set of standard elements corresponding to the historical filler text and the historical image.
[0011] Optionally, the calculating the similarity between the target feature set and the preset feature set includes: All first elements in the target element set are concatenated to obtain a first result, all second elements in the preset element set are concatenated to obtain a second result, and a first similarity between the first result and the second result is calculated according to cosine similarity.
[0012] Optionally, the calculating the similarity between the target feature set and the preset feature set includes: For each first element in the target element set, a second element with the highest similarity value is matched from the preset element set to form an element group, and the similarity values of all element groups are summed and averaged to obtain the similarity.
[0013] In a second aspect of the present application, a system for enhancing multimodal model analysis is provided, comprising a construction module, a matching module and an execution module, wherein: A construction module configured to construct a type knowledge base of a target domain, wherein the type knowledge base includes classification types of the target domain and knowledge corresponding to the classification types; A matching module configured to, when receiving a question input by a user, split the question into an image and an original text, input the image into a preset neural network model to obtain a target type to which the image belongs, and match corresponding target knowledge in the type knowledge base according to the target type; An execution module is configured to fill the target knowledge into a preset position in the original text to obtain a filled text, and input the filled text and the image into a preset multimodal large model to obtain an answer to the question.
[0014] In the third aspect of the present application, an electronic device is provided, including a processor, a memory, a user interface and a network interface, the memory is used to store instructions, the user interface and the network interface are both used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device executes any one of the methods described above.
[0015] In a fourth aspect of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores instructions, and when the instructions are executed, any of the methods described above is executed.
[0016] In summary, one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: 1. By building a type knowledge base in the target field, more accurate and in-depth knowledge support can be provided for questions in specific fields. This helps ensure that the answer is not only based on the direct content of the question, but also combines professional knowledge and classification information in the field; 2. Combining information from both image and text modalities, the target type to which the image belongs is identified through a neural network model, and relevant knowledge is matched in the type knowledge base. This multimodal processing method can more comprehensively understand the problem and improve the completeness and accuracy of the answer to the question; 3. The questions input by the user are split into images and original texts, and by filling the target knowledge into the preset positions in the original texts, richer and more specific filling texts are generated. This not only helps the multimodal large model to better understand the questions, but also provides users with more detailed and useful answers, thereby improving the user experience; 4. The type knowledge base can be updated and expanded as needed to meet the needs of different fields or specific problems. This flexibility enables the method to be widely used in various scenarios to meet the needs of different users; 5. The entire process from question splitting, image recognition, knowledge matching to answer generation is automated, reducing the need for human intervention. This improves processing efficiency and enables the method to respond to user questions more quickly. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a flowchart of the method for enhancing multimodal model analysis disclosed in the embodiment of the present application; Figure 2 It is an overall schematic diagram of the method for enhancing multimodal model analysis disclosed in the embodiment of the present application; Figure 3 It is a module schematic diagram of a system for enhancing multimodal model analysis disclosed in an embodiment of the present application; Figure 4 It is a structural schematic diagram of an electronic device disclosed in an embodiment of the present application.
[0018] Explanation of the reference numerals: 301, construction module; 302, matching module; 303, execution module; 401, processor; 402, communication bus; 403, user interface; 404, network interface; 405, memory. DETAILED DESCRIPTION
[0019] In order to enable technicians in this field to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments.
[0020] In the description of the embodiments of the present application, words such as "for example" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "for example" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "for example" or "for example" is intended to present related concepts in a specific way.
[0021] In the description of the embodiments of the present application, the meaning of the term "multiple" refers to two or more. For example, multiple systems refer to two or more systems, and multiple screen terminals refer to two or more screen terminals. In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance or implicitly indicating the indicated technical features. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. The terms "include", "comprise", "have" and their variations all mean "including but not limited to", unless otherwise specifically emphasized.
[0022] This embodiment discloses a method for enhancing multimodal model analysis, which is applied to a model analysis platform. Figure 1 is a flow chart of the method for enhancing multimodal model analysis disclosed in the embodiment of the present application, such as Figure 1 As shown, the method comprises the following steps: S101, constructing a type knowledge base of a target domain, wherein the type knowledge base includes classification types of the target domain and knowledge corresponding to the classification types; S102, when receiving a question input by a user, splitting the question into an image and original text, inputting the image into a preset neural network model to obtain a target type to which the image belongs, and matching corresponding target knowledge in the type knowledge base according to the target type; S103, filling the target knowledge into a preset position in the original text to obtain a filled text, and inputting the filled text and the image into a preset multimodal large model to obtain an answer to the question.
[0023] For each target area, such as healthcare, auto repair, legal knowledge, etc., collect and define classification types (such as disease types, auto parts types, legal provisions types, etc.) and detailed knowledge corresponding to these types (such as disease symptoms, treatment methods; auto parts functions, repair methods; legal provisions interpretation, cases, etc.). Organize the collected classification types and corresponding knowledge into a structured knowledge base. This can be a knowledge graph based on a graph database or a database in the form of key-value pairs. The key is to be able to retrieve and match efficiently.
[0024] The following table is a comparison table of the classification types of target projects and corresponding knowledge: When a user enters a question, the question is split into an image part and an original text part. For example, a user may upload a picture of a car breakdown with a text description. The split image is input into a preset neural network model (such as a convolutional neural network CNN, a deep learning model, etc.), which is trained to identify the target type in the image (such as car parts, disease symptoms, etc.). According to the target type identified in the image, the corresponding target knowledge is found and matched in the type knowledge base. This step may involve complex retrieval and matching algorithms to ensure that the most relevant knowledge is accurately found. The matched target knowledge is filled into the preset position in the original text. This may require further text processing, such as natural language generation (NLG) technology, to ensure that the filled text is fluent and meaningful. The filled text and the original image are input into the preset multimodal large model together. The preset multimodal large model can process and understand the association between text and image, so as to generate more accurate and comprehensive answers. Finally, the answer generated by the multimodal large model is output to the user. This may be a simple text answer, or it may contain pictures, charts or other multimedia elements to show the answer more intuitively.
[0025] By building a type knowledge base for the target domain, it is possible to classify specific domains and store knowledge related to these classifications. This structured knowledge base helps to quickly and accurately find knowledge related to user questions. When the user enters a question, the method can identify the target type in the image and match the corresponding knowledge in the knowledge base, thereby improving the accuracy and efficiency of information retrieval. The user-entered question is split into images and original texts, which reflects the ability to process multimodal information. By processing images and texts separately, these two information sources can be fully utilized to answer questions. By filling the target knowledge into the original text and combining it with image information, richer and more accurate answers can be generated. This answer generation method that combines image and text information helps to provide more comprehensive and specific answers. The preset multimodal large model can process complex input information and generate high-quality answers. This model usually has strong language generation and understanding capabilities and can generate coherent and accurate answers based on input information.
[0026] Figure 2 is an overall schematic diagram of the method for enhancing multimodal model analysis disclosed in the embodiment of the present application, such as Figure 2As shown, the user inputs the original question, the original question is split into the original image and the original text, the original image is classified, the corresponding domain knowledge is matched according to the classification result, and the domain knowledge is filled into the original text to obtain the filled text. Then, the question filled with domain knowledge is obtained based on the original image and the filled text, and the question is input into the multimodal large model to generate an answer, and the answer is sent to the user.
[0027] Optionally, inputting the image into a preset neural network model to obtain the target type to which the image belongs includes: Extracting preliminary features of the image through a plurality of convolutional layers in the preset neural network model, wherein the preliminary features include edge, texture or color information of the image; Capturing intermediate layer features of different scales of the image according to the preliminary features through multiple parallel Inception modules; Using an auxiliary classifier to classify the plurality of intermediate layer features to obtain a plurality of categories; Through the fully connected layer and the softmax layer, the probability of each category is output, and the category with the highest probability is selected as the target type to which the image belongs.
[0028] The embodiment of the present application uses a neural network model of the Inception v3 structure as a classifier. When an image is input into the neural network model, it passes through one or more convolutional layers. The main function of these convolutional layers is to extract preliminary features of the image, which generally include edge, texture and color information of the image. The convolutional layer captures the local features of the image through convolution operations (i.e., using convolution kernels to slide on the image and perform dot product operations). As the convolutional layer goes deeper, the neural network model can learn more complex and abstract feature representations. After the preliminary feature extraction, the image data enters the core part of Inception v3, i.e., multiple parallel Inception modules. Each Inception module consists of multiple parallel convolution paths, which can contain convolution kernels of different sizes (such as 1x1, 3x3, 5x5, etc.) and a maximum pooling layer. Such a design allows the neural network model to capture image features at different scales, thereby improving the neural network model's ability to understand and generalize image content. By stacking multiple Inception modules, the neural network model can gradually learn more advanced and abstract feature representations, which are essential for subsequent image classification tasks. In the middle of Inception v3, auxiliary classifiers were introduced. These classifiers can classify the features of the middle layer and output the probability of each category. The output of the auxiliary classifiers is used to calculate additional losses, which helps regularize the model and prevent overfitting. By introducing auxiliary classifiers, the neural network model can learn useful feature representations more stably during training. After the Inception module, the neural network model usually contains one or more fully connected layers. These fully connected layers map the extracted features to the final classification results. Finally, the neural network model outputs the probability of each category through a softmax layer. The role of the softmax layer is to convert the output of the fully connected layer into a probability distribution so that the sum of the probabilities of all categories is 1. After obtaining the probability of each category, the model will select the category with the highest probability as the target type to which the image belongs. This category is the final classification result of the model for the input image.
[0029] The preliminary features of the image are extracted through multiple convolutional layers. These features cover basic elements such as the edge, texture and color information of the image. This multi-level and multi-scale feature extraction method helps the model understand the image content more comprehensively. The design of the convolutional layer enables the model to automatically learn the key features in the image without manually designing the feature extractor, thereby improving the accuracy and efficiency of feature extraction. The Inception module achieves the ability to capture image features at different scales through multiple parallel convolution paths, each of which contains convolution kernels and pooling layers of different sizes. This design helps the model better adapt to image elements of different sizes and shapes, and improves the generalization ability and robustness of the model. Introducing auxiliary classifiers to classify the intermediate layer features not only helps the model learn richer feature representations during training, but also regularizes the model by calculating additional losses to prevent overfitting. This regularization strategy helps improve the generalization performance of the model, so that it can still maintain good classification results when facing unseen images. Through the fully connected layer and the softmax layer, the model can map the extracted features to the final classification results and output the probability of each category. This design enables the model to quickly complete the classification task while ensuring the accuracy and reliability of the classification results. Selecting the category with the highest probability as the target type to which the image belongs is both intuitive and meets the needs of practical applications.
[0030] The following table shows the categories of fields involved in the target projects and the number of pictures of each category: Optionally, outputting the probability of each category through the fully connected layer and the softmax layer, and selecting the category with the highest probability as the target type to which the image belongs includes: Generate a feature vector according to the intermediate layer features, input the feature vector into a fully connected layer, and generate an original score vector equal to the number of categories; Normalizing the original score vector using a softmax function to generate a normalized probability vector; The category corresponding to the maximum value in the probability vector is selected as the classification result of the image.
[0031] After being processed by multiple convolutional layers and Inception modules, the neural network model generates a series of intermediate layer features. These features are usually in the form of feature maps, which contain information about the image at different scales and levels. In order to convert these feature maps into feature vectors that can be used for classification, the model uses a series of operations (such as global average pooling or global maximum pooling) to compress the spatial dimensions of the feature map to obtain a fixed-length feature vector. The generated feature vector is input into the fully connected layer. The fully connected layer is a linear transformation layer that maps the feature vector to a new space where each dimension corresponds to a possible category. Specifically, the fully connected layer calculates the product of the feature vector and the weight matrix and adds a bias term to obtain a raw score vector equal to the number of categories. Each element in this vector represents the raw score of the corresponding category. Since the elements in the raw score vector are unnormalized, they cannot be directly interpreted as probabilities. In order to obtain the probability of each category, the model normalizes the raw score vector using the softmax function. The role of the Softmax function is to convert the elements in the original score vector into a probability distribution so that the sum of the probabilities of all categories is 1. Specifically, the softmax function calculates the ratio of each element to the sum of all elements to obtain a normalized probability vector. After obtaining the normalized probability vector, the model selects the category with the highest probability as the classification result of the image. This is because the category with the highest probability is most likely to be the true category to which the image belongs. To achieve this, the model traverses the elements in the probability vector and finds the maximum value. Then, it looks for the category index corresponding to this maximum value and outputs this category as the final classification result.
[0032] The intermediate layer features are effectively converted into feature vectors through the fully connected layer. This process maps the high-dimensional intermediate layer features to a low-dimensional space equal to the number of categories, that is, the original score vector. This conversion not only retains key information, but also simplifies subsequent calculations. The original score vector is normalized using the softmax function to generate a normalized probability vector. The softmax function can convert any real-valued vector into a probability distribution so that the sum of the probabilities of all categories is 1. This normalization process makes the probabilities of each category comparable, providing a basis for subsequent category selection. The category corresponding to the maximum value in the probability vector is selected as the classification result of the image. Since the softmax function has ensured the rationality of the probability distribution, it is intuitive and reasonable to select the category with the largest probability as the classification result. This method not only improves the accuracy of classification, but also reduces the risk of misclassification. The calculation process of the fully connected layer and the softmax layer is relatively simple and efficient, and can quickly output the probability of each category. This efficient calculation method enables the model to have better real-time performance and response speed in practical applications. At the same time, the normalization of the softmax function also enhances the stability of the model. When faced with input data of different scales and distributions, the model can maintain stable output performance and is not easily disturbed by extreme values or noise.
[0033] Optionally, filling the target knowledge into a preset position in the original text to obtain a filled text, and inputting the filled text and the image into a preset multimodal macro model to obtain an answer to the question includes: Traversing the original text to determine preset identifiers, and filling the target knowledge into positions between the preset identifiers to obtain a filling text; Integrating the filler text and the image into a unified data structure, and inputting the data structure into the preset multimodal macro model; Acquire text features through a text encoder, acquire image features through an image encoder, and fuse the text features and the image features to generate a cross-modal joint representation; An answer to the question is formed based on the joint representation.
[0034] In the original text, the preset identifier is "{}". Traverse the original text to find the positions of the preset identifier in the text so that the target knowledge can be accurately filled in these positions. Fill the target knowledge into the positions between the "{}" identifiers in the original text to obtain the filler text. In order to input the filler text and the image together into the multimodal large model, they need to be integrated into a unified data structure. This may be a JSON object containing two fields, text and image, or a specially designed data format to support the input of multimodal data. When the filler text and image are integrated into a unified data structure, the data structure can be input into the preset multimodal large model. The preset multimodal large model is designed to process text and image data at the same time and output answers related to the question. Inside the multimodal large model, there is a text encoder responsible for processing the filler text. It converts the text data into a series of vector representations that capture the semantic information in the text. At the same time, there is an image encoder responsible for processing the image data. It converts the image into a series of vector representations that capture the visual features in the image. In order to generate a cross-modal joint representation, the multimodal large model fuses text features and image features. This can be achieved in many ways, such as splicing, weighted summation, attention mechanism, etc. The fused features contain both textual information and image information, providing rich context for subsequent answer generation. The multimodal large model generates an answer to the question based on the fused joint representation. This may be a short text answer or a complex answer with multiple information points. In any case, the answer will be based on the information in the filled text and image and answer the original question as accurately as possible. For example, the original text is as follows: "You are an experienced police officer. The above pictures are some clues you have collected so far, which may help you get more information.\ {}Please follow the steps below:\ 1. Describe the content in the picture in detail. In this step, please do not make unnecessary associations. \ 2. According to the description of step 1, further analyze and extract the highly sensitive elements (including pornography, violence, etc.). For the extracted elements, please perform semantic matching on them. For those with too high similarity, just keep one, and then arrange these elements from high to low in terms of sensitivity. \ 3. Return the results of step 1 and step 2 respectively. The final result is returned in JSON format: it only contains two variables "description" and "sensitive_information". \ Among them, "description" is a plain text description of the image without other sub-variables and control characters such as '\n'; and sensitive_information is the sensitive information in the image, which is a detailed description of the sensitive information. " After filling the target knowledge corresponding to the result returned by the neural network in “{}”, the final text is input into the multimodal large model.
[0035] By traversing the original text to determine the preset identifiers and accurately filling the target knowledge into the positions between these identifiers, the accuracy and completeness of the knowledge are ensured. This filling method is not only applicable to fixed knowledge templates, but also can flexibly respond to the needs of different scenarios and problems. The filling text and image are integrated into a unified data structure, so that multimodal data can be processed and analyzed under a unified framework. This integration method fully utilizes the complementarity between text and image, and improves the accuracy of information extraction and understanding. Text features and image features are obtained through text encoders and image encoders respectively, and these features are fused to generate a cross-modal joint representation. This fusion method not only retains the key information in the original data, but also captures the correlation and consistency between different modalities, providing rich context for subsequent answer generation. The answer to the question is formed based on the cross-modal joint representation. This process fully utilizes the powerful computing power and intelligent reasoning ability of the multimodal large model. The model can automatically learn and generate answers related to the question based on the input data without manual intervention or additional rule design. At the same time, because the model can process text and image data at the same time, it can generate more accurate answers in a shorter time.
[0036] Optionally, the method further includes training the preset multimodal large model, specifically including: The target element set of historical filler text and historical image is extracted through the preset multimodal large model, the similarity between the target element set and the preset element set is calculated, and the parameters of the multimodal large model are adjusted according to the similarity. The preset element set is a set of standard elements corresponding to the historical filler text and the historical image.
[0037] Historical filler texts and historical images are data collected from practical applications. They contain a variety of possible input situations, such as different text descriptions, image content, etc. These data will be used to train the multimodal large model to enable it to recognize and process various inputs. The preset feature set is a set of standard features that defines the key information that the multimodal large model should focus on when processing inputs. These features may include keywords in the text, specific objects or scenes in the image, etc. The preset feature set is an important reference in the training process, which helps the model learn to recognize and extract these key information. During the training process, the multimodal large model will first extract the target feature set from the historical filler text and historical images. These target features are the key information that the model automatically learns and recognizes based on the input data. The process of extracting the target feature set is an important step for the model to understand the input data. The multimodal large model will calculate the similarity between the extracted target feature set and the preset feature set. The calculation of similarity can be achieved through a variety of methods, such as cosine similarity, Jaccard similarity, etc. The purpose of this step is to evaluate the degree of match between the target features extracted by the model and the standard features, so as to judge the performance of the model. Based on the calculated similarity, the multimodal large model adjusts its parameters to improve performance. If the target feature set has a high similarity with the preset feature set, it means that the model is already able to identify and extract key information well. At this time, the learning rate can be reduced or fine-tuned to avoid overfitting. If the similarity is low, it means that the model needs further learning. At this time, the learning rate can be increased or more training data can be introduced to improve the performance of the model. The above process will be iterated multiple times until the performance of the model on the validation set is stable or no longer significantly improved. During the iteration process, the model will continue to learn and adjust its parameters to better adapt to the input data and extract key information.
[0038] By extracting the target feature set of historical filler text and historical images and calculating the similarity with the preset feature set (i.e., a set of standard features), the performance of the model in identifying and understanding multimodal data (text and images) can be accurately evaluated. Adjusting the model parameters based on the feedback of similarity helps the model to more accurately capture and interpret the key information in the data, thereby improving the accuracy of its prediction or generation. Using historical data for training allows the model to be exposed to a variety of input situations, thereby enhancing its ability to handle multimodal data of different types and styles. This training method helps the model to apply the learned knowledge more flexibly when facing new data, thereby improving the generalization performance of the model. By calculating the similarity between the target feature set and the preset feature set, clear guidance is provided for the adjustment of model parameters. This similarity-based training method is more efficient than the traditional trial and error method, and can find the optimal model parameter configuration more quickly, thereby shortening training time and reducing training costs.
[0039] Optionally, the calculating the similarity between the target feature set and the preset feature set includes: All first elements in the target element set are concatenated to obtain a first result, all second elements in the preset element set are concatenated to obtain a second result, and a first similarity between the first result and the second result is calculated according to cosine similarity.
[0040] For example, the target feature set P = [p_1, p_2, ..., p_n], the preset feature set R = [r_1, r_2, ..., r_m], all the first elements in the target feature set (i.e., p_1, p_2, ..., p_n) are spliced to obtain a continuous text sequence, which is called the first result (P_c). Similarly, all the second elements in the preset feature set (i.e., r_1, r_2, ..., r_m) are spliced to obtain a continuous text sequence, which is called the second result (R_c). After obtaining the first result and the second result, the cosine similarity is used to measure the similarity between the two. Cosine similarity is a commonly used text similarity calculation method, which evaluates the similarity between two vectors by calculating the cosine value of the angle between them. Here, the first result and the second result are regarded as two vectors, and the cosine similarity between them is calculated. Through the calculation of cosine similarity, a value between -1 and 1 can be obtained, which represents the similarity between the target feature set and the preset feature set. When the similarity value is close to 1, it means that the two feature sets are very similar; when the similarity value is close to -1, it means that the two feature sets are very different; when the similarity value is close to 0, it means that the two feature sets are independent to some extent.
[0041] By concatenating all the features in the target feature set and the preset feature set and calculating the cosine similarity between them, a specific value can be obtained to measure the similarity between the two. This quantitative evaluation method is more accurate and objective than a simple qualitative description, and helps to understand the performance of the model in feature extraction more accurately. The process of concatenating feature sets and calculating similarity is relatively simple and efficient, and a large amount of evaluation work can be completed in a short time. This helps to improve the efficiency of evaluation, especially when dealing with large-scale data sets, which can significantly reduce the time and labor costs required for evaluation. By concatenating all the features in the feature set, an overall representation containing all feature information can be obtained. This overall representation can more comprehensively reflect the characteristics of the feature set and support simultaneous comparison and analysis of multiple features. By calculating the similarity, the performance evaluation results of the model in feature extraction can be obtained. These results can be used as the basis for model optimization, such as adjusting model parameters, improving algorithms, or increasing training data, to improve the accuracy and generalization ability of the model.
[0042] Optionally, the calculating the similarity between the target feature set and the preset feature set includes: For each first element in the target element set, a second element with the highest similarity value is matched from the preset element set to form an element group, and the similarity values of all element groups are summed and averaged to obtain the similarity.
[0043] For each first element in the target element set P, it is necessary to find a second element with the highest similarity to it in the preset element set R to form an element pair. This similarity can be calculated by a text similarity algorithm (such as cosine similarity, Jaccard similarity, etc.). It should be noted that the matching here is not necessarily a strict one-to-one matching, because the number of elements in the target element set and the preset element set may be different, and a target element may have a certain similarity with multiple preset elements. But in this step, only the preset element with the highest similarity is selected for matching. Preset element set R: This is a specific set of elements that should be extracted for each image manually, such as R = ["dog", "grass", "sunny day"]. Target element set P: This is a set of elements extracted by the model from user input that has been filled with specific domain knowledge, such as P = ["pet", "green plants", "sunny weather"]. For the first element "pet" in the target element set P, it is necessary to find the element with the highest similarity to it in the preset element set R. In this example, it can be considered that "pet" has a certain similarity with "dog" (although "dog" is a subset of "pet", in order to simplify the explanation, it is assumed that there is a certain text similarity between them). Therefore, "pet" and "dog" are grouped into one element group, and the similarity between them is calculated (assuming it is 0.7, which is obtained through a certain text similarity algorithm). For the second element "green plants" in the target element set P, the element with the highest similarity in the preset element set R is "grass" (assuming the similarity is 0.8). Finally, for the third element "clear weather" in the target element set P, the element with the highest similarity in the preset element set R is "sunny day" (assuming the similarity is 0.9). After obtaining the similarities of all element groups, these similarities are summed up: 0.7+0.8+0.9=2.4. Then, the average of these similarities is calculated as the similarity: 2.4 / 3=0.8.
[0044] By finding the most similar elements in the preset element set for each element in the target element set and forming an element group, the correlation between the elements of the two element sets can be accurately measured. This fine-grained matching method helps to capture the subtle differences between the elements, thereby providing a more accurate similarity assessment. By summing and averaging the similarity values of all element groups, the overall similarity of the two element sets can be fully reflected. This global evaluation method helps to grasp the overall correlation and consistency between the element sets. The embodiment of the present application does not rely on a specific element representation or similarity calculation algorithm and has strong adaptability. It can select suitable element representation methods and similarity calculation algorithms according to different application scenarios and requirements to achieve more accurate similarity assessment.
[0045] The following table shows the correlation results between the elements extracted by the embodiment of the present application and the elements extracted manually: Calculation method 1 refers to concatenating all the first elements in the target element set to obtain a first result, concatenating all the second elements in the preset element set to obtain a second result, and calculating the first similarity between the first result and the second result according to the cosine similarity; Calculation method 2 refers to for each first element in the target element set, matching the second element with the highest similarity value from the preset element set to form an element group, and summing and averaging the similarity values of all element groups to obtain the similarity. It can be seen from the above table that the multimodal model specific field analysis enhancement method proposed in the embodiment of the present application has obtained good correlation in the three categories.
[0046] This embodiment also discloses a system for enhancing multimodal model analysis. Figure 3 is a module diagram of a system for enhancing multimodal model analysis disclosed in an embodiment of the present application, such as Figure 3 As shown, the system includes a construction module 301, a matching module 302 and an execution module 303, wherein: A construction module 301 is configured to construct a type knowledge base of a target domain, wherein the type knowledge base includes classification types of the target domain and knowledge corresponding to the classification types; The matching module 302 is configured to, when receiving a question input by a user, split the question into an image and an original text, input the image into a preset neural network model to obtain a target type to which the image belongs, and match corresponding target knowledge in the type knowledge base according to the target type; The execution module 303 is configured to fill the target knowledge into a preset position in the original text to obtain a filled text, and input the filled text and the image into a preset multimodal large model to obtain an answer to the question.
[0047] Optionally, the matching module 302 is configured to: Extracting preliminary features of the image through a plurality of convolutional layers in the preset neural network model, wherein the preliminary features include edge, texture or color information of the image; Capturing intermediate layer features of different scales of the image according to the preliminary features through multiple parallel Inception modules; Using an auxiliary classifier to classify the plurality of intermediate layer features to obtain a plurality of categories; Through the fully connected layer and the softmax layer, the probability of each category is output, and the category with the highest probability is selected as the target type to which the image belongs.
[0048] Optionally, the matching module 302 is configured to: Generate a feature vector according to the intermediate layer features, input the feature vector into a fully connected layer, and generate an original score vector equal to the number of categories; Normalizing the original score vector using a softmax function to generate a normalized probability vector; The category corresponding to the maximum value in the probability vector is selected as the classification result of the image.
[0049] Optionally, the execution module 303 is configured to: Traversing the original text to determine preset identifiers, and filling the target knowledge into positions between the preset identifiers to obtain a filling text; Integrating the filler text and the image into a unified data structure, and inputting the data structure into the preset multimodal macro model; Acquire text features through a text encoder, acquire image features through an image encoder, and fuse the text features and the image features to generate a cross-modal joint representation; An answer to the question is formed based on the joint representation.
[0050] Optionally, the system further comprises a training module, wherein the training module is configured to: The target element set of historical filler text and historical image is extracted through the preset multimodal large model, the similarity between the target element set and the preset element set is calculated, and the parameters of the multimodal large model are adjusted according to the similarity. The preset element set is a set of standard elements corresponding to the historical filler text and the historical image.
[0051] Optionally, the training module is configured to: All first elements in the target element set are concatenated to obtain a first result, all second elements in the preset element set are concatenated to obtain a second result, and a first similarity between the first result and the second result is calculated according to cosine similarity.
[0052] Optionally, the training module is configured to: For each first element in the target element set, a second element with the highest similarity value is matched from the preset element set to form an element group, and the similarity values of all element groups are summed and averaged to obtain the similarity.
[0053] It should be noted that: when the device provided in the above embodiment realizes its function, only the division of the above functional modules is used as an example. In actual application, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device and method embodiments provided in the above embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.
[0054] This embodiment also discloses an electronic device, referring to Figure 4 The electronic device may include: at least one processor 401 , at least one communication bus 402 , a user interface 403 , a network interface 404 , and at least one memory 405 .
[0055] The communication bus 402 is used to realize the connection and communication between these components.
[0056] The user interface 403 may include a display screen (Display) and a camera (Camera), and the optional user interface 403 may also include a standard wired interface and a wireless interface.
[0057] The network interface 404 may optionally include a standard wired interface or a wireless interface (such as a WI-FI interface).
[0058] Among them, the processor 401 may include one or more processing cores. The processor 401 uses various interfaces and lines to connect various parts in the entire server, and executes various functions of the server and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory 405, and calling data stored in the memory 405. Optionally, the processor 401 can be implemented in at least one hardware form of digital signal processing (Digital Signal Processing, DSP), field programmable gate array (Field-Programmable Gate Array, FPGA), and programmable logic array (Programmable Logic Array, PLA). The processor 401 can integrate one or more combinations of a central processing unit (Central Processing Unit, CPU), a graphics processing unit (Graphics Processing Unit, GPU) and a modem. Among them, the CPU mainly processes the operating system, user interface and application programs; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; the modem is used to process wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor 401, and it can be implemented by a single chip.
[0059] Among them, the memory 405 may include a random access memory (Random Access Memory, RAM) and may also include a read-only memory (Read-Only Memory). Optionally, the memory 405 includes a non-transitory computer-readable storage medium. The memory 405 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 405 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store data involved in the above-mentioned various method embodiments, etc. The memory 405 may optionally be at least one storage device located away from the aforementioned processor 401. As Figure 4 As shown, the memory 405 as a computer storage medium may include an operating system, a network communication module, a user interface module, and an application program of the method for enhancing multimodal model analysis.
[0060] exist Figure 4In the electronic device shown, the user interface 403 is mainly used to provide an input interface for the user and obtain data input by the user; and the processor 401 can be used to call the application program for the multimodal model analysis enhancement method stored in the memory 405. When executed by one or more processors 401, the electronic device executes one or more methods in the above-mentioned embodiments.
[0061] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the order of the actions described, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required for the present application.
[0062] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0063] In the several embodiments provided in this application, it should be understood that the disclosed devices can be implemented in other ways. For example, the device embodiments described above are only schematic, such as the division of units, which is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some service interfaces, and the indirect coupling or communication connection of devices or units can be electrical or other forms.
[0064] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0065] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0066] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a memory 405, including a number of instructions for a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of the various embodiments of the present application. The aforementioned memory 405 includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk.
[0067] The above is only an exemplary embodiment of the present disclosure and cannot be used to limit the scope of the present disclosure. That is, any equivalent changes and modifications made according to the teachings of the present disclosure are still within the scope of the present disclosure. After considering the disclosure of the specification, those skilled in the art will easily think of other embodiments of the present disclosure. This application is intended to cover any modification, use or adaptation of the present disclosure, which follows the general principles of the present disclosure and includes common knowledge or customary technical means in the technical field that are not recorded in the present disclosure. The description and examples are only regarded as exemplary, and the scope and spirit of the present disclosure are defined by the claims.
Claims
1. A method for enhancing multimodal model analysis, characterized in that: Applied to the model analysis platform, the method comprises: Constructing a type knowledge base of a target domain, wherein the type knowledge base includes classification types of the target domain and knowledge corresponding to the classification types; When receiving a question input by a user, split the question into an image and original text, input the image into a preset neural network model to obtain the target type to which the image belongs, and match the corresponding target knowledge in the type knowledge base according to the target type; The target knowledge is filled into a preset position in the original text to obtain a filled text, and the filled text and the image are input into a preset multimodal large model to obtain an answer to the question.
2. The method for enhancing multimodal model analysis according to claim 1, characterized in that: The step of inputting the image into a preset neural network model to obtain the target type to which the image belongs includes: Extracting preliminary features of the image through a plurality of convolutional layers in the preset neural network model, wherein the preliminary features include edge, texture or color information of the image; Capturing intermediate layer features of different scales of the image according to the preliminary features through multiple parallel Inception modules; Using an auxiliary classifier to classify the plurality of intermediate layer features to obtain a plurality of categories; Through the fully connected layer and the softmax layer, the probability of each category is output, and the category with the highest probability is selected as the target type to which the image belongs.
3. The method for enhancing multimodal model analysis according to claim 2, characterized in that: The method of outputting the probability of each category through the fully connected layer and the softmax layer, and selecting the category with the highest probability as the target type to which the image belongs includes: Generate a feature vector according to the intermediate layer features, input the feature vector into a fully connected layer, and generate an original score vector equal to the number of categories; Normalizing the original score vector using a softmax function to generate a normalized probability vector; The category corresponding to the maximum value in the probability vector is selected as the classification result of the image.
4. The method for enhancing multimodal model analysis according to claim 1, characterized in that: The step of filling the target knowledge into a preset position in the original text to obtain a filled text, and inputting the filled text and the image into a preset multimodal macro model to obtain an answer to the question includes: Traversing the original text to determine preset identifiers, and filling the target knowledge into positions between the preset identifiers to obtain a filling text; Integrating the filler text and the image into a unified data structure, and inputting the data structure into the preset multimodal macro model; Acquire text features through a text encoder, acquire image features through an image encoder, and fuse the text features and the image features to generate a cross-modal joint representation; An answer to the question is formed based on the joint representation.
5. The method for enhancing multimodal model analysis according to claim 1, characterized in that: The method further includes training the preset multimodal large model, specifically including: The target element set of historical filler text and historical image is extracted through the preset multimodal large model, the similarity between the target element set and the preset element set is calculated, and the parameters of the multimodal large model are adjusted according to the similarity. The preset element set is a set of standard elements corresponding to the historical filler text and the historical image.
6. The method for enhancing multimodal model analysis according to claim 5, characterized in that: The calculating the similarity between the target element set and the preset element set includes: All first elements in the target element set are concatenated to obtain a first result, all second elements in the preset element set are concatenated to obtain a second result, and a first similarity between the first result and the second result is calculated according to cosine similarity.
7. The method for enhancing multimodal model analysis according to claim 5, characterized in that: The calculating the similarity between the target element set and the preset element set includes: For each first element in the target element set, a second element with the highest similarity value is matched from the preset element set to form an element group, and the similarity values of all element groups are summed and averaged to obtain the similarity.
8. A system for enhancing multimodal model analysis, characterized in that: It includes building module, matching module and execution module, among which: A construction module configured to construct a type knowledge base of a target domain, wherein the type knowledge base includes classification types of the target domain and knowledge corresponding to the classification types; A matching module configured to, when receiving a question input by a user, split the question into an image and an original text, input the image into a preset neural network model to obtain a target type to which the image belongs, and match corresponding target knowledge in the type knowledge base according to the target type; An execution module is configured to fill the target knowledge into a preset position in the original text to obtain a filled text, and input the filled text and the image into a preset multimodal large model to obtain an answer to the question.
9. An electronic device, characterized in that: It includes a processor, a memory, a user interface and a network interface, the memory is used to store instructions, the user interface and the network interface are both used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device executes the method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, and when the instructions are executed, the method according to any one of claims 1 to 7 is executed.