Method for improving model picture recognition capability based on multi-modal large model

By using a multimodal large model infrastructure and fine-tuning training, the shortcomings of traditional image recognition technology in terms of accuracy and efficiency are solved, achieving a high-efficiency and low-cost improvement in image recognition capabilities.

CN120833540APending Publication Date: 2025-10-24BEIJING INFORMATION TECH BOTE INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202510995517.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2025-10-24

AI Technical Summary

Technical Problem

Traditional image recognition technology is difficult to meet the needs of diverse scenarios in terms of accuracy and efficiency, and the training process is time-consuming and costly, requiring the retraining of all model parameters.

Method used

A multimodal large model is adopted. By building the basic architecture of the multimodal large model, the basic model is trained with massive images. Image and text data are preprocessed, and instruction fine-tuning datasets are designed for fine-tuning training. The model parameters are optimized to enhance the model's generalization ability and recognition accuracy.

Benefits of technology

It significantly improves the accuracy and efficiency of image recognition, reduces the number of training images, enhances the model's stability and generalization ability in unseen scenes, and reduces training costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120833540A_ABST
    Figure CN120833540A_ABST
Patent Text Reader

Abstract

The invention discloses a method for improving the model picture recognition capability based on a multi-modal large model, and relates to the technical field of image recognition, and the method comprises the steps: building an infrastructure of the multi-modal large model, and training the multi-modal large model as a basic model through a large number of pictures; newly added images are collected, and corresponding standard questions and answers describing picture contents are prepared in combination with the images; and according to a picture recognition task requirement, designing a corresponding instruction, namely a cue word, and guiding the large model to carry out picture recognition. The multi-modal large model is trained by using massive pictures, the accuracy of a picture recognition technology can be optimized by fully utilizing a picture feature library of the multi-modal large model when the multi-modal large model is applied to the field of picture recognition, and the magnitude order of training pictures can be greatly reduced by using fine tuning training of the multi-modal large model, so that the accuracy of the picture recognition technology is improved. After composite fine tuning training is carried out, the multi-modal large model can have the recognition capability of generalizing trained pictures, and the number of pictures needing to be used for training can be greatly reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image recognition, in particular to a method for improving model picture recognition capability based on a multi-modal large model. BACKGROUND

[0002] In today's era of rapid development of digitization and intelligence, image data is growing explosively and is widely used in social media, medical imaging, security monitoring, autonomous driving, industrial detection and many other fields. Accurate and efficient picture recognition capability has become a key technical requirement for promoting the development of various fields. With the widespread application of artificial intelligence technology, picture recognition technology has shown great potential in various fields. However, different scenarios have higher requirements for the accuracy and efficiency of picture recognition, and traditional methods are difficult to meet these needs.

[0003] The optimization process of traditional picture recognition technology for accuracy is to continuously add new pictures for labeling and then perform model training. Each time, all pictures need to be retrained and all model parameters need to be updated. Since all pictures need to be retrained, the training time is long and the cost is high, and it is repetitive work. Therefore, the present application proposes a method for improving model picture recognition capability based on a multi-modal large model to solve the above-mentioned problems. SUMMARY

[0004] The present application aims to provide a method for improving model picture recognition capability based on a multi-modal large model to solve the problems mentioned in the background.

[0005] To solve the above technical problems, the technical solution adopted by the present application is: The method for improving model picture recognition capability based on a multi-modal large model comprises the following steps: S1, building a basic framework of a multi-modal large model, and training the multi-modal large model as a basic model using a large amount of pictures; S2, collecting new images, and preparing corresponding standard questions and answers describing the content of the pictures in combination with the images; S3, designing corresponding instructions, i.e. prompt words, according to the requirements of the picture recognition task to guide the large model to perform picture recognition; S4, combining the instructions, new images, standard questions and answers describing the content of the pictures into a one-to-one instruction fine-tuning data set, and adding prompt word content in combination with the prompt word engineering to mine more capabilities of the large model and reduce the number of fine-tuning data sets; S5, fine-tuning the multi-modal large model using the edited instruction fine-tuning data set, optimizing the model parameters, avoiding recognition capability deviation, and improving the explainability; S6, evaluate the fine-tuned multi-modal large model, calculate the estimated recognition coefficient, analyze the picture recognition accuracy and generalization ability of the model, and then determine whether to apply it to the actual scene.

[0006] Further improvement of the technical scheme of the application is that S1 specifically comprises: The model structure of a multi-modal large model supporting multi-modal data input and fusion is selected, and the encoder and decoder parts of the model are constructed, the encoder is responsible for extracting the features of images and texts, and the decoder generates output according to the features, that is, a cross-modal encoder combining a convolutional neural network (CNN) for processing images and a Transformer architecture for processing texts; Picture data is extracted from a multi-source data set containing a large amount of labeled images, covering various scenes, styles and object categories, and text data related to the pictures is collected, including picture descriptions and labels, etc., and at the same time, the collected data is preprocessed, including image cropping, scaling and normalization operations, and text segmentation and encoding processing, so that the data can adapt to the input requirements of the model; The multi-modal large model is trained using the preprocessed large amount of picture data and text data, during the training process, the Adam optimizer is used to adjust the parameters of the model, to learn the association between images and texts and the feature representation of images, at the same time, training parameters including learning rate, batch size, etc. are set to ensure that the model can converge stably, and during the training process, the performance of the model is evaluated regularly, and the model is adjusted and optimized according to the evaluation results, and finally a basic model with a rich picture feature library is obtained.

[0007] Further improvement of the technical scheme of the application is that S2 specifically comprises: According to the specific requirements of the picture recognition task, newly added images are collected from public data sets, industry databases and real scene shooting, to ensure that the images cover various scenes, styles and object categories, cover 100+ subdivided scenes under different lighting, angle, shielding conditions, to improve the generalization ability of the model, and through an automated script, low-quality images, i.e. blurred and repeated samples, are filtered, and then the images are preliminarily classified to ensure data diversity, at the same time, the meta information of the retained image annotation source and shooting parameters is retained; For each newly added image, a domain expert designs standard questions and corresponding answers to guide the model to perform picture recognition, the questions need to cover object detection, attribute judgment, spatial relationship, etc., a "question-answer-reason" triple annotation method is used, and the answer needs to clearly quote the image features; Through a cross-validation mechanism, two annotators independently generate question and answer pairs, and when there is a conflict, a third person arbitrates, so that the annotation consistency reaches more than 95%.

[0008] Further improvement of the technical scheme of the application is that S3 specifically comprises: According to the specific goal of the picture recognition task, the core recognition dimension is disassembled, including object detection, attribute judgment, spatial relationship and domain specification, a basic instruction template is designed for each dimension to ensure that the instruction covers the whole scene of the task; Embedding domain knowledge prompt words in the basic instruction, guiding through prompt word combination in multiple modes, and adopting a dynamic prompt word generation strategy to automatically adjust the instruction according to the input image features, reduces the data dependence while improving the instruction accuracy; Construct a verification set containing 5000+ samples to simulate real task scenarios to test the instruction effect, and calculate the recognition accuracy, recall rate and false alarm rate, compare different instruction versions through A / B test, select the optimal instruction combination, and establish an instruction iteration mechanism to optimize the prompt words according to the model output deviation, and finally form a standardized instruction library covering more than 95% of the task scenarios.

[0009] The further improvement of the technical scheme of the application is that the S4 specifically comprises: The newly added image is paired with the corresponding standard question and the answer describing the picture content one by one, so that each image has a clear question and a detailed answer, forming a complete data unit, which helps to accurately define the features and avoid model deviation in the recognition process; Embed the designed unique corresponding instruction in each data unit to form a "instruction-image-question-answer" four-tuple, and check the data consistency to ensure that each four-tuple is logically self-consistent, and embed the domain knowledge prompt words in the instruction, automatically add auxiliary instructions according to the image features, and guide through prompt word combination in multiple modes; Based on the optimization results of the prompt words, high-value samples covering the core task scenarios are selected, redundant data such as repeated images of similar angles are deleted, the data set size is compressed by 30%-50%, and then a visual quality inspection tool is constructed, samples are randomly inspected and "instruction-image-answer" triplets are displayed, and when the annotations are inconsistent, the rollback mechanism is triggered, and finally a refined and high-quality instruction fine-tuning data set is formed.

[0010] The further improvement of the technical scheme of the application is that the process of automatically adding auxiliary instructions according to image features is: For each newly added image, a four-tuple is generated according to the "instruction-image-question-answer" structure, wherein the instruction is preset by a domain expert, the question needs to point directly to the core content of the image, and the answer needs to describe the features in detail, and the four-tuple consistency is automatically checked through a script, including: domain relevance, image path and format, and uniqueness; Embed the domain knowledge prompt words in the preset instruction to guide the model to understand the domain semantics of the image, and automatically add auxiliary instructions through image feature analysis, and then adopt a prompt word combination strategy to guide the model to focus on the whole image and details at the same time; The script secondary check quadruple logic self-consistency includes the matching of the answer and the question and the compatibility of the instruction and the prompt word, wherein, by analyzing the matching of the answer and the question, it is checked whether the answer directly answers the question, by analyzing the compatibility of the instruction and the prompt word, it is verified whether the prompt word is consistent with the instruction field, for the quadruple that does not meet the requirements, it is modified or regenerated, and it is ensured that the logic of each quadruple is clear and the semantics is accurate.

[0011] The further improvement of the technical scheme of the application is that the S5 specifically comprises: The final check is performed on the edited instruction fine-tuning dataset to ensure the integrity and accuracy of the data, the logical self-consistency of each "instruction-image-question-answer" quadruple in the instruction fine-tuning dataset is verified, it is checked whether the image path and file format are correct, the compatibility of the instruction and the prompt word in the dataset is ensured, the dataset is formatted by an automatic script to adapt to the input requirements of the multi-modal large model, and then the instruction fine-tuning dataset is divided into a training set (80%), a verification set (15%) and a test set (5%), and at the same time, the images are standardized, the text instructions and answers are segmented and encoded, and converted into a tensor format that can be input into the model; A training framework suitable for the multi-modal large model is selected, distributed training parameters are configured, a pre-trained model weight is loaded, and an optimizer and a loss function are initialized to ensure that the training environment is stable and efficient, and then the multi-modal large model is fine-tuned and the parameters are optimized in two stages, wherein in the first stage of fine-tuning, the model backbone parameters are fixed, only the instruction encoding layer and the output head are fine-tuned, the model is quickly adapted to the instruction-image alignment task, in the second stage of fine-tuning, the backbone parameters are gradually unfrozen, the small learning rate full parameter fine-tuning is adopted, the model's ability to analyze complex scenes is strengthened, the loss weight is dynamically adjusted according to the task difficulty to avoid model recognition deviation caused by data imbalance, and after each round of training, the accuracy, recall rate and F1 score are calculated on the verification set, if the indicators do not improve for 3 consecutive rounds, the early stopping mechanism is triggered to prevent overfitting; The Grad-CAM tool is used to generate the attention heat map of the model on the key area of the image to verify whether the model focuses on the instruction-related features, and the SHAP value is used to quantify the contribution of the instruction and the image to the output result to identify potential bias, i.e., over-reliance on the background rather than the subject, and then the accuracy and robustness of the model in the core scene are evaluated on the test set, if the performance of the model in a specific scene is found to be decreased, the data is supplemented and the model is fine-tuned again to form a "training-verification-optimization" closed loop, and finally a fine-tuned multi-modal large model with strong explainability and balanced recognition ability is output.

[0012] The further improvement of the technical scheme of the application is that the S6 specifically comprises: Prepare a test dataset for evaluating the fine-tuned multimodal large model, containing diverse image samples, and define evaluation metrics including precision, recall, F1 score, and decay rate; Run the fine-tuned multimodal large model on the test dataset and calculate the accuracy, recall, F1 score, and decay rate. Then, analyze the basic recognition ability using the weighted F1 score, analyze the generalization ability using the normalized decay rate, and clarify the domain knowledge adaptability by analyzing the prediction probability. Combined with the basic recognition ability, generalization ability, and domain knowledge adaptability, calculate the estimated recognition coefficient. Compare it to the baseline model before fine-tuning to verify the improved efficiency of instruction fine-tuning. Also, analyze the error types using the confusion matrix to locate the model's weak links. Based on the calculated estimated recognition coefficient, a performance threshold T is set. If the model meets the performance threshold, the deployment process begins. If it does not meet the standard, failure cases are analyzed, and targeted data is supplemented or the model structure is adjusted. The setting of the performance threshold is adjusted according to the task type and requirements, and continuous analysis is then carried out to ensure that the decision logic of the multimodal large model conforms to domain knowledge and is applicable to actual scenarios.

[0013] A further improvement of the technical solution of the present invention is that the calculation formula of the estimated recognition coefficient is as follows: ; ; ; ; Where, To estimate the recognition coefficient, As a basic recognition ability, is the basic recognition ability weight, , is the weighted F1 score, is the number of categories, For the Class weights, For the The F1 score of the class, For generalization ability, is the generalization ability weight, , is the normalized decay rate, is the original decay rate, and are the minimum and maximum decay rates in the historical data, For domain knowledge adaptability, is the domain knowledge weight, , is the domain knowledge adaptability, For all decision outcomes, and are the prediction probabilities of the expert and model pair results respectively.

[0014] Due to the adoption of the above technical solutions, the technical progress achieved by the present application relative to the prior art is: 1. The present application provides a method for improving the picture recognition capability of a model based on a multi-modal large model. The multi-modal large model is trained using a large number of pictures, and when applied to the field of picture recognition, the picture feature library of the multi-modal large model can be fully utilized to optimize the accuracy of the picture recognition technology. Furthermore, by using fine-tuning training of the multi-modal large model, the order of magnitude of the training pictures can be greatly reduced. After composite fine-tuning training, the multi-modal large model can have the recognition capability of generalizing the trained pictures, and the number of pictures needed for training can be greatly reduced.

[0015] 2. The present application provides a method for improving the picture recognition capability of a model based on a multi-modal large model. By integrating multi-source information such as images and texts, a rich feature representation space is constructed. In the instruction fine-tuning stage, the model can adapt to the recognition needs of different scenes through a dynamic prompt word generation strategy. In addition, the cross-modal pre-training task ensures the high alignment of semantics between modalities, so that the model can still maintain stable performance in unseen scenes, significantly enhancing the generalization capability. BRIEF DESCRIPTION OF DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments described in the present application, and other drawings can also be obtained by those skilled in the art based on these drawings.

[0017] Figure 1 is a workflow diagram of the present application; Figure 2 is a method flow diagram of the present application. DETAILED DESCRIPTION

[0018] In order to make the purpose, technical solutions and advantages of the embodiments of the present application more clear, the technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0019] Embodiment 1, as Figure 1 , Figure 2As shown, the present application provides a method for improving the picture recognition capability of a large multi-modal model, comprising the following steps: S1, build the basic framework of the multi-modal large model, and train the multi-modal large model as the basic model by using the massive pictures, provide the basis for subsequent optimization of recognition accuracy by means of the rich picture feature library accumulated by the multi-modal large model, select a model structure of the multi-modal large model supporting multi-modal data input and fusion, construct the encoder and decoder parts of the model, the encoder is responsible for extracting the features of the image and the text, and the decoder generates the output according to the features, that is, the cross-modal encoder combining the convolutional neural network (CNN) processing image and the Transformer architecture processing text, wherein the pre-training task of image text matching and mask language modeling ensures the semantic alignment between modalities, constructs an extensible underlying framework, extracts picture data from a multi-source data set containing massive annotated images, covers various scenes, styles and object categories, and collects text data related to the pictures, including picture description and labels, etc., at the same time, the collected data is preprocessed, including image cropping, scaling and normalization operations, and text segmentation and encoding processing, so that the data can adapt to the input requirements of the model, the multi-modal large model is trained by using the preprocessed massive picture data and text data, in the training process, the Adam optimizer is used to adjust the parameters of the model, learn the association between the image and the text and the feature representation of the image, at the same time, set the training parameters including learning rate, batch size, etc., to ensure that the model can converge stably, and during the training process, the performance of the model is evaluated regularly, the model is adjusted and optimized according to the evaluation results, and finally the basic model with rich picture feature library is obtained; S2, collect new images, prepare corresponding standard questions and answers describing picture content in combination with images, according to the specific requirements of the picture recognition task, collect new images from public data sets, industry databases and real scene shooting, ensure that the images cover various scenes, styles and object categories, cover 100+ sub-scenes under different lighting, angle, shielding conditions, in order to improve the generalization ability of the model, and filter low-quality images, that is, fuzzy and repeated samples through automatic script, and then classify the images preliminarily to ensure data diversity, at the same time, the meta information of the reserved image annotation source and shooting parameters, for each new image, the domain expert designs standard questions and corresponding answers to guide the model to recognize pictures, the questions need to cover object detection, attribute judgment, spatial relationship and other dimensions, adopt the "question-answer-reason" three-link labeling method, require the answer to clearly quote the image features, through the cross-validation mechanism, two annotators independently generate question and answer pairs, in case of conflict, a third person arbitrates, so that the consistency of labeling reaches more than 95%, effectively reduces the labeling error, improves the data quality, provides accurate labeled data for model training, and thus improves the recognition accuracy and generalization ability of the model; S3. According to the requirements of the picture recognition task, design the corresponding instructions, i.e. prompt words, to guide the large model to perform picture recognition. According to the specific goal of the picture recognition task, disassemble the core recognition dimensions, including object detection, attribute judgment, spatial relationship and domain specification. Design basic instruction templates for each dimension to ensure that the instructions cover all scenarios of the task. Embed domain knowledge prompt words in the basic instructions, combine the prompt words for multi-modal guidance, and use a dynamic prompt word generation strategy to automatically adjust the instructions according to the input image features, reducing data dependence while improving instruction accuracy. Construct a validation set containing 5000+ samples to simulate real task scenarios and test the instruction effect. Calculate the recognition accuracy, recall rate and false positive rate. Compare different instruction versions through A / B testing to select the optimal instruction combination. Establish an instruction iteration mechanism to optimize the prompt words based on the model output bias. Finally, form a standardized instruction library covering more than 95% of the task scenarios. S4. Combine the instructions, new images, standard questions and answers describing the content of the pictures into a one-to-one instruction fine-tuning dataset, and add prompt word content to the prompt word engineering to tap the capabilities of the large model and reduce the size of the fine-tuning dataset. Pair the new images with the corresponding standard questions and answers describing the content of the pictures in a one-to-one manner, so that each image has a clear question and a detailed answer, forming a complete data unit. Embed the designed unique corresponding instructions in each data unit to form a "instruction-image-question-answer" four-tuple, and verify the data consistency to ensure that each four-tuple is logically self-consistent. At the same time, embed domain knowledge prompt words in the instructions, automatically add auxiliary instructions based on image features, and use prompt word combinations for multi-modal guidance. Based on the prompt word optimization results, select high-value samples that cover core task scenarios and delete redundant data such as similar angle repeat images to compress the dataset size by 30%-50%. Then, build a visual quality inspection tool, randomly sample and display "instruction-image-answer" triplets, and trigger the rollback mechanism when the annotations are inconsistent. Finally, form a simplified and high-quality instruction fine-tuning dataset. In addition, the process of automatically adding auxiliary instructions based on image features is as follows: For each new image, a quadruple is generated according to the "instruction-image-question-answer" structure, where the instruction is preset by a domain expert, the question must clearly point to the core content of the image, and the answer must describe the features in detail. The consistency of the quadruple is automatically verified by a script, including: domain relevance, image path and format, and uniqueness. By analyzing the domain relevance, it is ensured that the instruction and the question / answer belong to the same domain. By analyzing the image path and format, it is verified whether the image path exists and whether the file format is supported. By analyzing the uniqueness, it is avoided that multiple images use the same instruction, ensuring that each image corresponds to a unique instruction to prevent semantic confusion. Domain knowledge prompts are embedded in the preset instructions to guide the model to understand the domain semantics of the image, and auxiliary instructions are automatically added through image feature analysis. For example, when a low-contrast image is detected, the model will track the image. Add "enhanced edge recognition, ignoring reflective areas" and "prioritize identification of anomalies 5mm in size" when discovering small target objects. In the presence of occlusion, add "infer the features of the occluded part based on the context". Then adopt a prompt word combination strategy to guide the model to focus on both the overall image and details, reduce dependence on large-scale annotated data, and improve the ability to parse complex scenes (such as occlusion and lighting changes). Use scripts to verify the logical consistency of the quadruple, including the matching of answers and questions and the compatibility of instructions and prompt words. Specifically, by analyzing the matching of answers and questions, check whether the answers directly respond to the questions. By analyzing the compatibility of instructions and prompt words, verify whether the prompt words are consistent with the instruction domain. For quadruple groups that do not meet the requirements, correct or regenerate them to ensure that each quadruple has clear logic and accurate semantics. S5. Use the edited instruction fine-tuning dataset to fine-tune the multimodal large model, optimize model parameters, avoid recognition capability deviation, and improve interpretability; S6. Evaluate the fine-tuned multimodal large model, calculate the estimated recognition coefficient, analyze the model's image recognition accuracy and generalization ability, and then determine whether it can be applied to actual scenarios.

[0020] Example 2, as Figure 1 、 Figure 2 As shown, based on Example 1, the present invention provides a technical solution: preferably, S5 specifically includes: The final check is performed on the polished instruction fine-tuning dataset to ensure data integrity and accuracy, verify the logical consistency of each "instruction-image-question-answer" quadruple in the instruction fine-tuning dataset, check the correctness of the image path and file format, ensure the compatibility of the instructions in the dataset with the prompt words, format the dataset through an automated script to adapt to the input requirements of the multi-modal large model, and then divide the instruction fine-tuning dataset into training set (80%), validation set (15%) and test set (5%). At the same time, the images are standardized, the text instructions and answers are segmented, encoded, and converted into tensor format that can be input into the model, the training framework that adapts to the multi-modal large model is selected, the distributed training parameters are configured, the pre-trained model weights are loaded, and the optimizer and loss function are initialized to ensure stable and efficient training environment. Then, the multi-modal large model is fine-tuned and parameter optimized in two stages. In the first stage of fine-tuning, the model backbone parameters are fixed, only the instruction encoding layer and output head are fine-tuned, so that the model quickly adapts to the instruction-image alignment task. In the second stage of fine-tuning, the backbone parameters are gradually unfrozen, and the model is fine-tuned with a small learning rate to strengthen the model's ability to analyze complex scenes. Then, the loss weight is dynamically adjusted according to the task difficulty to avoid model recognition bias caused by data imbalance. After each round of training, the accuracy, recall rate and F1 score are calculated on the validation set. If the indicators do not improve for 3 consecutive rounds, the early stopping mechanism is triggered to prevent overfitting. The Grad-CAM tool is used to generate the attention heat map of the model on the key areas of the image to verify whether the model focuses on the instruction-related features. The SHAP value is used to quantify the contribution of each part of the instruction and image to the output result to identify potential biases, i.e. over-reliance on background rather than subject. Then, the model's accuracy and robustness in core scenarios are evaluated on the test set. If the performance of specific scenarios is found to be declining, additional data is added and the model is fine-tuned again to form a "training-validation-optimization" closed loop. Finally, an interpretable multi-modal large model with balanced recognition ability is output. S6 specifically comprises: Prepare a test dataset for evaluating the fine-tuned multi-modal large model, containing diverse image samples covering various scenarios and conditions encountered by the model, ensuring that the test dataset has some difference in distribution from the training dataset to better evaluate the model's generalization ability. Define evaluation metrics, including accuracy, recall, F1 score, and decay rate. Run the fine-tuned multi-modal large model on the test dataset, calculate accuracy, recall, F1 score, and decay rate, then analyze the basic recognition ability through weighted F1 score, analyze the generalization ability through normalized decay rate, and analyze the domain knowledge adaptation degree through the analysis of prediction probability. Combine the basic recognition ability, generalization ability, and domain knowledge adaptability to calculate the estimated recognition coefficient. Compare the baseline model before fine-tuning to verify the improvement efficiency of the instruction fine-tuning. At the same time, analyze the error types through the confusion matrix to locate the weak links of the model. According to the calculated estimated recognition coefficient, set a performance threshold T. If the model meets the performance threshold, it enters the deployment process. If it does not meet the standard, analyze the failure cases and supplement data or adjust the model structure accordingly. The performance threshold is set according to the task type and requirements, and then continuous analysis is carried out to make the decision logic of the multi-modal large model conform to the domain knowledge and applicable to actual scenarios. The calculation formula of the estimated recognition coefficient is as follows: ; ; ; ; In the formula, is the estimated recognition coefficient, is the basic recognition ability, is the basic recognition ability weight, , is the weighted F1 score, is the number of categories, is the weight of the th category, is the F1 score of the th category, is the generalization ability, is the generalization ability weight, , is the normalized decay rate, is the original decay rate, and are the minimum and maximum decay rates in the historical data, is the domain knowledge adaptability, is the domain knowledge weight, , is the domain knowledge adaptation degree, is all the decision results, and the prediction probability of the expert and model pair results The smaller the value is, the higher the adaptability is. The higher the value is, the stronger the basic recognition ability is. The higher the value is, the stronger the generalization ability is. The higher the value is, the better the domain knowledge adaptability is. The closer the value is to 1, the better the comprehensive performance of the model is.

[0021] The above merely provides a specific implementation of the present application, but the protection scope of the present application is not limited thereto, any person skilled in the art can easily think of changes or replacements within the technical range disclosed by the present application, which shall be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.​

Claims

1. A method for improving the picture recognition capability of a model based on a multi-modal large model, characterized in that, The method comprises the following steps: S1, building a basic framework of a multi-modal large model, and training the multi-modal large model as a basic model by using a large amount of pictures; S2, collecting newly added images, combining the images to prepare corresponding standard questions and answers describing the content of the images; S3, designing corresponding instructions, i.e., prompt words, according to the requirements of the image recognition task, to guide the large model to perform image recognition; S4, combining the instructions, newly added images, standard questions and answers describing the content of the images into an instruction fine-tuning data set in a one-to-one manner, and adding prompt word content in combination with the prompt word engineering; S5, fine-tuning the multi-modal large model by using the instruction fine-tuning data set arranged and edited, and optimizing the model parameters; S6, evaluating the fine-tuned multi-modal large model, calculating the estimated recognition coefficient, analyzing the image recognition accuracy and generalization ability of the model, and then determining whether to apply it to the actual scene.

2. The method for improving the model picture recognition ability based on the multi-modal large model according to claim 1, characterized in that: The S1 specifically comprises: selecting a model structure of a multi-modal large model supporting multi-modal data input and fusion, constructing an encoder and a decoder part of the model, the encoder being responsible for extracting features of images and texts, and the decoder generating output according to the features, i.e., a cross-modal encoder combining a convolutional neural network processing images and a Transformer architecture processing texts; extracting picture data from a multi-source data set containing a large amount of labeled images, and collecting text data related to the pictures, at the same time, pre-processing the collected data, including image cropping, scaling and normalization operations, and text segmentation and encoding processing; training the multi-modal large model by using the pre-processed large amount of picture data and text data, adjusting the parameters of the model by using an Adam optimizer, learning the association between images and texts and the feature representation of images, at the same time, setting training parameters including learning rate and batch size, and regularly evaluating the performance of the model during the training process, adjusting and optimizing the model according to the evaluation results, and finally obtaining a basic model with a rich picture feature library.

3. The method of claim 1, wherein the method is based on a multi-modal large model to improve the picture recognition capability of the model. The S2 specifically comprises: According to the specific requirements of the image recognition task, directionally collect newly added images from public data sets, industry databases and real scene shooting, cover 100+ sub-scenarios under different light, angle and shielding conditions, and filter low-quality images, i.e., blurred and repeated samples, through an automatic script, and then preliminarily classify the images, at the same time, reserve the meta information of image annotation source and shooting parameters; For each newly added image, a domain expert designs a standard question and a corresponding answer for guiding the model to perform image recognition, adopts a "question-answer-reason" triple annotation method, and requires the answer to clearly quote image features; Through a cross-validation mechanism, two annotators independently generate question and answer pairs, and a third person arbitrates in case of conflict, so that the consistency of annotation reaches more than 95%.

4. The method for improving the model picture recognition ability based on the multi-modal large model according to claim 1, characterized in that: The S3 specifically comprises: According to the specific target of the image recognition task, disassemble the core recognition dimensions, including object detection, attribute judgment, spatial relationship and domain specification, and design a basic instruction template for each dimension; Embedding domain knowledge prompts in basic instructions, guiding through prompt combination, and adopting a dynamic prompt generation strategy to automatically adjust instructions according to image features; A verification set containing 5000+ samples is constructed to simulate real task scenarios and test instruction effectiveness. The recognition accuracy, recall rate, and false positive rate are calculated. Different instruction versions are compared through A / B testing to select the optimal instruction combination. An instruction iteration mechanism is established to optimize prompts based on model output bias. Finally, a standardized instruction library covering more than 95% of task scenarios is formed.

5. The method for improving model image recognition capability based on a multimodal large model according to claim 1, characterized in that: The S4 specifically includes: The newly added images are paired one-to-one with corresponding standard questions and answers describing the content of the pictures, so that each image has a clear question and a detailed answer, forming a complete data unit; Embedding the designed unique corresponding instructions in each data unit forms a "instruction-image-question-answer" four-tuple, and checks the data consistency. At the same time, embed domain knowledge prompts in the instructions, automatically add auxiliary instructions according to image features, and guide through prompt combination. Based on the prompt optimization results, high-value samples covering core task scenarios are selected, and redundant data is deleted to construct a visual quality inspection tool. Randomly sample and display "instruction-image-answer" triads, and trigger rollback mechanism when inconsistent. Finally, a fine-tuning data set of instructions is formed.

6. The method of claim 5, wherein the method is based on a multi-modal large model to improve the picture recognition capability of the model. The process of automatically adding auxiliary instructions according to image features is as follows: For each newly added image, generate a four-tuple according to the "instruction-image-question-answer" structure, and automatically check the consistency of the four-tuple through scripts, including domain relevance, image path and format, and uniqueness; Embed domain knowledge prompts in the preset instructions to guide the model to understand the domain semantics of the image, and automatically add auxiliary instructions through image feature analysis. Then, use the prompt combination strategy to guide the model to focus on both the overall and the details of the image. Check the logical self-consistency of the four-tuple through scripts, including the matching of answers and questions and the compatibility of instructions and prompts. For four-tuples that do not meet the requirements, modify or regenerate them.

7. The method of claim 1, wherein the method is based on a multi-modal large model to improve the picture recognition capability of the model. The S5 specifically includes: Finally check the fine-tuning data set of instructions, verify the logical self-consistency of each "instruction-image-question-answer" four-tuple in the fine-tuning data set of instructions, check whether the image path and file format are correct, perform data set formatting through automated scripts, and then divide the fine-tuning data set of instructions into training set, verification set and test set. At the same time, standardize the images, tokenize and encode the text instructions and answers, and convert them into tensor format that can be input into the model. The training framework of the adaptive multi-modal large model is selected, the distributed training parameters are configured, the pre-trained model weight is loaded, and the optimizer and loss function are initialized, and then the two-stage fine-tuning training and parameter optimization of the multi-modal large model are performed, wherein in the first stage fine-tuning training, the model backbone parameters are fixed, and only the instruction encoding layer and the output head are fine-tuned, and in the second stage fine-tuning training, the backbone parameters are gradually unfrozen, and the small learning rate full parameter fine-tuning is adopted, and then the loss weight is dynamically adjusted according to the task difficulty, and after each training, the accuracy, recall rate and F1 score are calculated on the validation set, if the index does not improve for 3 consecutive rounds, the early stopping mechanism is triggered to prevent overfitting; The attention heat map of the model to the key area of the image is generated by the Grad-CAM tool, the model whether focuses on the instruction related features is verified, the SHAP value is used to quantify the contribution of the instruction and the image to the output result, the potential bias is identified, and then the accuracy and robustness of the model in the core scene are evaluated on the test set, if the performance of a specific scene is found to be decreased, the data is supplemented and re-fine-tuned, forming a "training-verification-optimization" closed loop, and finally the fine-tuned multi-modal large model is output.

8. The method for improving the model picture recognition ability based on the multi-modal large model according to claim 1, characterized in that: The S6 specifically comprises: Prepare a test data set for evaluating the fine-tuned multi-modal large model, containing diversified image samples, and define evaluation indicators including accuracy, recall rate, F1 score and decay rate; Run the fine-tuned multi-modal large model on the test data set, calculate the accuracy, recall rate, F1 score and decay rate, then analyze the basic recognition ability through the weighted F1 score, analyze the generalization ability through the normalized decay rate, and determine the domain knowledge adaptation degree through the analysis of the prediction probability, and combine the basic recognition ability, generalization ability and domain knowledge adaptability to calculate the estimated recognition coefficient, compare the baseline model before fine-tuning to verify the improvement efficiency of instruction fine-tuning, and analyze the error type through the confusion matrix to locate the weak link of the model; According to the calculated estimated recognition coefficient, set a performance threshold T, if the model meets the performance threshold, enter the deployment process, if it does not meet the standard, analyze the failure case, supplement the data or adjust the model structure, wherein the performance threshold is set according to the task type and demand, and then continuous analysis is performed to make the decision logic of the multi-modal large model conform to the domain knowledge and applicable to the actual scene.

9. The method for improving the model picture recognition ability based on the multi-modal large model according to claim 8, characterized in that: The calculation formula of the estimated recognition coefficient is as follows: ; ; ; ; wherein, is the estimated recognition coefficient, is the basic recognition ability, is the basic recognition ability weight, , is the weighted F1 score, is the number of classes, is the weight of the th class, is the F1 score of the th class, is the generalization ability, is the generalization ability weight, , is the normalized decay rate, is the original decay rate, and are the minimum and maximum decay rates in the historical data, is the domain knowledge adaptability, is the domain knowledge weight, , is the domain knowledge adaptability degree, is the final decision result, and are the prediction probabilities of the expert and the model on the result respectively.

Citation Information

Patent Citations

  • Image fine-grained description method and system of instruction fine-tuning multi-mode large model

    CN117423108A

  • Multi-modal data exception identification method and device, electronic equipment and storage medium

    CN118379755A

  • Sample pair generation method and device, large model training method and device, image retrieval method and device, equipment and medium

    CN118643342A

  • Instruction fine tuning data set construction method and device

    CN118966379A

  • Image-based cue word generation method and device, equipment and storage medium

    CN119229443A

Cited By

  • Large language model comprehensive capability assessment method and related device

    CN121581225A

  • Method and device for detecting state of clamping device of automatic alignment device

    CN121685913A

  • Multi-modal model-based thrown object detection method and system

    CN122024186A

  • Large model construction method for vertical field of audiology

    CN122174988A