Model training method and device, image processing method and device, electronic equipment and storage medium

By training a multimodal large language model, combining the relationship between image description information, scenes and elements, the problem of difficulty in extracting multi-dimensional information in the prior art is solved, and high-accuracy multi-label marking and image description are achieved.

CN120236109APending Publication Date: 2025-07-01HANGZHOU ALICLOUD FEITIAN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311860076.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-29
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The prior art is difficult to effectively extract information from multiple dimensions in image data processing, especially in multi-label marking and image description, and the training process is cumbersome.

Method used

By obtaining sample images and their corresponding image description information, scene sets and feature sets, a training set is constructed and a multi-modal large language model is trained. The model includes a scene recognition module, a feature recognition module and an image description module. The model is trained through multiple rounds of dialogue and adjusting model parameters to improve prediction accuracy.

Benefits of technology

The accuracy of multi-label marking and image description of image data is improved, which reduces the need for training data, simplifies the training process, and improves the efficiency and accuracy of image processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236109A_ABST
    Figure CN120236109A_ABST
Patent Text Reader

Abstract

The invention provides a model training method and device, an image processing method and device, electronic equipment and a storage medium, and the model training method comprises the steps: obtaining a plurality of sample data, the sample data comprises a sample image, the image description information of the sample image, and a plurality of scene sets and a plurality of element sets corresponding to the sample image; the plurality of scene sets are obtained by permutation and combination of scene labels labeled in the sample image, and the plurality of element sets are obtained by permutation and combination of picture element labels labeled in the sample image; the image description information is used for describing scenes and elements contained in the sample image and the relation between the elements and the scenes; constructing a training set based on the multiple pieces of sample data; and training the multi-modal large language model based on the training set. According to the method, the multi-modal large language model is trained based on the sample image, the scene set and the element set of the sample image and the image description information, so that the trained model has higher accuracy on image description and multi-label marking.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of image processing, and specifically relates to a model training, image processing method, device, electronic device, and storage medium. Background Art

[0002] With the continuous development of Internet of Things technology, the Internet of Things platform carries image data reported by tens of thousands of front-end image acquisition devices. For this image data, from the perspective of the platform, it is necessary to further explore the value of the image data and output it to users to improve the competitiveness of products. From the perspective of users, in addition to viewing images, they also need to recognize and manage this image data from a richer dimension, so they hope to obtain more intuitive information from the image data.

[0003] It should be noted that the above statements are only used to provide background technical information related to this application, and do not necessarily constitute prior art. Summary of the Invention

[0004] This application proposes a model training, image processing method, device, electronic device, and storage medium, which can train a multi-modal large language model based on sample images, their scene sets and element sets, and image description information, so that the trained model has higher accuracy in image description and multi-label tagging.

[0005] The first aspect of the embodiments of this application proposes a model training method, including:

[0006] Obtain a plurality of sample data, where the sample data includes a sample image, the image description information of the sample image, and a plurality of scene sets and a plurality of element sets corresponding to the sample image; the plurality of scene sets are obtained by permuting and combining the respective scene labels annotated in the sample image, and the plurality of element sets are obtained by permuting and combining the respective picture element labels annotated in the sample image; the image description information is used to describe the scenes and elements included in the sample image, as well as the relationship between the elements and the scenes;

[0007] Construct a training set based on the plurality of sample data;

[0008] Train a multi-modal large language model based on the training set.

[0009] In some embodiments of this application, the constructing a training set based on the plurality of sample data includes:

[0010] For each sample data, combine the sample image and the image description information in the same sample data with one scene set in each of the scene sets and one element set in each of the element sets in this sample data to form a training sample;

[0011] The obtained multiple training samples are combined to form a training set.

[0012] In some embodiments of the present application, the multimodal large language model includes a scene recognition module, a feature recognition module, and an image description module; training the multimodal large language model based on the training set includes:

[0013] Input the training samples in the training set into the scene recognition module, and output the scene prediction results corresponding to the sample images included in the training samples;

[0014] Input the scene prediction results into the feature recognition module, and output the feature prediction results corresponding to the sample images included in the training samples;

[0015] Input the scene prediction results and the feature prediction results into the image description module, and output the image description prediction results of the sample images included in the training samples.

[0016] In some embodiments of the present application, the method further includes:

[0017] Obtain the device description information corresponding to the sample image, where the device description information is the information of the acquisition device that acquired the sample image;

[0018] Based on the device description information, determine the device features of the acquisition device, where the device features include at least one of the device type or the installation scene.

[0019] In some embodiments of the present application, the multimodal large language model includes a scene recognition module, a feature recognition module, and an image description module; training the multimodal large language model based on the training set includes:

[0020] Input the training samples in the training set and the device features corresponding to the sample images in the training samples into the scene recognition module, and output the scene prediction results in the samples included in the training samples;

[0021] Input the scene prediction results into the feature recognition module, and output the feature prediction results in the sample images included in the training samples;

[0022] Input the device features, the scene prediction results, and the feature prediction results into the image description module, and output the image description prediction results of the sample images included in the training samples.

[0023] In some embodiments of the present application, training the multimodal large language model based on the training set further includes:

[0024] Calculate a first loss value based on the predicted results of each scenario and the set of scenarios included in the training samples; adjust the first model parameters of the scenario recognition module based on the first loss value;

[0025] Calculate a second loss value based on the predicted results of each element and the set of elements included in the training samples; adjust the second model parameters of the element recognition module based on the second loss value;

[0026] Calculate a third loss value based on the predicted result of the image description and the image description information included in the training samples; adjust the third model parameters of the image description module based on the third loss value;

[0027] Based on the adjusted first model parameters, second model parameters, and third model parameters, return and loop the step of obtaining the predicted results of each scenario until a preset convergence condition is reached, and then obtain a trained multi-modal large language model.

[0028] In some embodiments of the present application, the method further includes:

[0029] Based on a preset label library, obtain the predicted results of each scenario and the predicted results of each element through the multi-modal large language model;

[0030] If there are target scenario labels and / or target picture element labels in the multiple training samples that do not exist in the preset label library, add the target scenario labels and / or the target picture element labels to the preset label library.

[0031] An embodiment of the second aspect of the present application provides an image processing method, including:

[0032] Obtain a target image to be processed;

[0033] Input the target image into a pre-trained multi-modal large language model, and output the image description information of the target image;

[0034] Wherein, the multi-modal large language model is obtained by the model training method described in the first aspect.

[0035] In some embodiments of the present application, the obtaining of the target image to be processed includes:

[0036] Obtain image data to be processed, and perform abnormal picture detection on the image data;

[0037] Determine the image data without detected abnormal pictures as the target image.

[0038] An embodiment of the third aspect of the present application provides a model training device, including:

[0039] A sample data acquisition module for acquiring a plurality of sample data, where the sample data includes sample images, image description information of the sample images, and a plurality of scene sets and a plurality of element sets corresponding to the sample images; the plurality of scene sets are obtained by arranging and combining the respective scene labels marked in the sample images, and the plurality of element sets are obtained by arranging and combining the respective picture element labels marked in the sample images; the image description information is used to describe the content of the sample images;

[0040] A training set construction module for constructing a training set based on the plurality of sample data;

[0041] A model training module for training a multi-modal large language model based on the training set.

[0042] An embodiment of the fourth aspect of the present application provides an image processing device, including:

[0043] An image acquisition module for acquiring a target image to be processed;

[0044] An image processing module for inputting the target image into a pre-trained multi-modal large language model and outputting the image description information of the target image;

[0045] Wherein, the multi-modal large language model is obtained by using the model training method described in the first aspect.

[0046] An embodiment of the fifth aspect of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor runs the computer program to implement the method described in the first aspect or the second aspect above.

[0047] An embodiment of the sixth aspect of the present application provides a computer-readable storage medium, on which a computer program is stored, and the program is executed by a processor to implement the method described in the first aspect or the second aspect above.

[0048] The technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:

[0049] In the embodiments of the present application,

[0050] The additional aspects and advantages of the present application will be partially given in the following description, partially will become obvious from the following description, or will be understood through the practice of the present application. Description of the Drawings

[0051] Various other advantages and benefits will become apparent to those of ordinary skill in the art upon reading the following detailed description of the preferred embodiments. The drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present application. Also, throughout the drawings, the same reference numerals are used to denote the same components.

[0052] In the drawings:

[0053] Figure 1 A schematic flowchart of a model training method provided by some embodiments of the present application is shown;

[0054] Figure 2 A schematic flowchart of the specific process of step S2 in some embodiments of the present application is shown;

[0055] Figure 3 A schematic flowchart of the specific process of step S3 in some embodiments of the present application is shown;

[0056] Figure 4 A schematic flowchart of an image processing method provided by some embodiments of the present application is shown;

[0057] Figure 5 A schematic flowchart of the specific process of the image processing method provided by some embodiments of the present application is shown;

[0058] Figure 6 A schematic structural diagram of a model training apparatus provided by some embodiments of the present application is shown;

[0059] Figure 7 A schematic structural diagram of an image processing apparatus provided by some embodiments of the present application is shown;

[0060] Figure 8 A schematic diagram of an electronic device provided by an embodiment of the present application is shown;

[0061] Figure 9 A schematic diagram of a storage medium provided by an embodiment of the present application is shown. Detailed Embodiments

[0062] The exemplary embodiments of the present application will be described in more detail below with reference to the drawings. Although the exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present application can be more thoroughly understood and the scope of the present application can be fully conveyed to those skilled in the art.

[0063] It should be noted that unless otherwise specified, the technical terms or scientific terms used in the present application should have the ordinary meaning understood by those skilled in the art to which the present application belongs.

[0064] In the related art, in order to obtain more intuitive information from image data, the input image data can be labeled by a model for generating labels. However, most of the existing models for generating labels have better single-label labeling effects, but the multi-label labeling effects are not ideal, the accuracy is relatively low, and a large amount of supervised data is required for training, and the training process is relatively cumbersome. Moreover, the existing models for generating labels generally only support the labeling function and do not support generative image descriptions, which cannot meet the requirements of users to recognize and manage data from multiple dimensions.

[0065] Based on the above problems, the embodiments of the present application provide a model training method and an image processing method. The image processing method processes images using the model obtained by training with the model training method. The model training method first obtains a plurality of sample data, where the sample data includes sample images, image description information of the sample images, and a plurality of scene sets and a plurality of element sets corresponding to the sample images. The plurality of scene sets are obtained by permuting and combining the respective scene labels annotated in the sample images, and the plurality of element sets are obtained by permuting and combining the respective picture element labels annotated in the sample images; the image description information is used to describe the scenes and elements included in the sample images, as well as the relationships between the elements and the scenes. Then, a multimodal large language model is trained based on the training set constructed from these sample data. One sample data in the model training method includes a sample image and a plurality of scene sets and a plurality of element sets of the image. In this way, even with a small amount of sample data, the trained model can not only learn the ownership relationship and dependency relationship between the scenes and elements, but also avoid the model learning unnecessary sequential / hierarchical relationships between the scenes and elements, so that an image processing model with more accurate labeling effects and more appropriate image description information can be obtained.

[0066] The model training method provided by the embodiments of the present application will be elaborated in detail below with reference to the accompanying drawings.

[0067] Please refer to Figure 1 , which is a schematic flowchart of the model training method provided by the embodiments of the present application. As Figure 1 shown, the model training method may include the following steps:

[0068] Step S1, obtain a plurality of sample data, where the sample data includes sample images, image description information of the sample images, and a plurality of scene sets and a plurality of element sets corresponding to the sample images.

[0069] Among them, the scene tags include but are not limited to: warehouse, inside the drone nest, elevator entrance, elevator car, security control room, hotel, equipment room (the machine room above the elevator), office, parking lot, park, factory, lake, corridor, playground, street store, school, residential area, park, river, road, countryside, stairway entrance, nine small places, alley, cafeteria, square, shopping mall, etc. Multiple scene sets are obtained by permuting and combining the various scene tags annotated in the sample images. For example, for sample image A, based on the results of manual annotation, it can be known that sample image A contains 3 scene tags: road, street store, and square. When forming the sample data, the annotated data can be expanded by permuting and combining the multiple scene tags in sample image A to form multiple scene sets. For example, set 1 is {road, street store, square}, set 2 is {street store, square, road}, set 3 is {square, street store, road}, and so on.

[0070] The element tags include but are not limited to: parking space, no-parking area, entrance (specifically referring to the scene with a door, a lifting pole, or an access control), green belt, sports facilities (including basketball hoops, elliptical machines, parallel bars, swings, slides...), intersection, zebra crossing, sidewalk, etc. Multiple element sets are obtained by permuting and combining the various element tags annotated in the sample images. For example, for sample image A, based on the results of manual annotation, it can be known that sample image A contains 4 element tags: no-parking area, sidewalk,

[0071] and sports facilities. When forming the sample data, the annotated data can be expanded by permuting and combining the multiple element tags in sample image A to form multiple element sets. For example, set 1 is {no-parking area, sidewalk, zebra crossing, sports facilities}, set 2 is {sidewalk, no-parking area, zebra crossing, sports facilities}, set 3 is {zebra crossing, sidewalk, sports facilities, no-parking area}, and so on.

[0072] The image description information is used to describe the scenes and elements included in the sample image, as well as the relationship between the elements and the scenes. When generating the description information, it is necessary to focus on all the generated tag information and make a reasonable description in combination with the dependency relationship between the elements and the scenes learned by the model. For example, for the above sample image A that includes 3 scene tags and 4 element tags, the corresponding image description information can be that image A shows a road with street stores and squares on both sides of the road, there is a zebra crossing on the road, there are no-parking areas and sidewalks beside the road, and there are sports facilities in the square.

[0073] Step S2, construct a training set based on multiple sample data.

[0074] After obtaining the above sample data in this embodiment, a training set can be constructed based on these sample data in a reasonable manner for training and testing the above multi-modal large language model.

[0075] In some embodiments, as shown in Figure 2, step S2 may specifically include the following processes: Step S21, for each sample data, combine the sample image and the image description information in the same sample data with one scene set and one element set in each scene set and each element set in the sample data to form a training sample; Step S22, form a training set with the obtained multiple training samples.

[0076] Among them, the training sample can be understood as the data input to the model when training the model for the current time.

[0077] In this embodiment, the sample image and the image description information in the same sample data are respectively combined with one scene set and one element set in each scene set and each element set in the sample data to form a training sample. In this way, based on a set of sample data of the same sample image, multiple training samples can be constructed. Using multiple training samples of the same sample image to train the model, the model can learn the belonging relationship and dependency relationship between the scene and the element. And the scene labels and element labels in each training sample are arranged in the order of natural combination, which can prevent the model from learning unnecessary order / hierarchy relationships between the scene and the element.

[0078] For example, for the above sample image A, according to the traditional training method, it is expected that the model will generate the following three scene labels during training: road, street-side store, and square. In this way, the model may learn the following order / hierarchy relationship: the street-side store must appear when the road appears. However, in actual applications, only the street-side store (such as the storefront image of a small nine-place venue) is captured in the sample picture, and the road is not captured. In this case, when there is no road in the scene label, the model may not output the label of the street-side store. Based on the above implementation method of constructing the training set in this embodiment, this unnecessary order / hierarchy relationship between the scene and the element can be avoided.

[0079] Step S3, based on the training set, train the multi-modal large language model.

[0080] Among them, the multi-modal large language model can be any model that has been trained by the above training set and can perform multi-label tagging on the input image and generate image description information. For example, but not limited to, the open-source model Qwen-VL, which is a large-scale vision language model (Large Vision Language Model, LVLM), or other models based on multi-modal Transformer (based on the attention mechanism). This embodiment does not make specific limitations on this, as long as it can achieve a high multi-label tagging and image description accuracy.

[0081] In some embodiments, such as Figure 3As shown in the figure, the multimodal large language model may include a scene recognition module, an element recognition module, and an image description module. Based on this, step S3 above may specifically include the following steps: Step S31, input the training samples in the training set into the scene recognition module, and output the scene prediction results corresponding to the sample images included in the training samples; Step S32, input the scene prediction results into the element recognition module, and output the element prediction results corresponding to the sample images included in the training samples; Step S33, input the scene prediction results and the element prediction results into the image description module, and output the image description prediction results of the sample images included in the training samples.

[0082] Among them, the scene prediction result is to generate a scene label, and the sample image corresponding to this scene label is the sample image included in the training sample. Similarly, the element prediction result is to generate an element label, and the sample image corresponding to this element label is also the sample image included in the training sample. It can be understood that the input data of the element recognition module includes the scene prediction result and also the sample image. The input information of the image description module includes the scene prediction result, each element prediction result, and also the sample image.

[0083] The multimodal large language model provided in this embodiment includes a scene recognition module, an element recognition module, and an image description module. These three modules are connected in sequence, and the output of the previous module is used as the input of the next module, forming a coherent multimodal large language model. The trained multimodal large language model can directly output image description information based on the input image, with a relatively fast image processing speed and a relatively high label output accuracy.

[0084] It should be noted that the above scene recognition module, element recognition module, and image description module can all operate independently to achieve corresponding functions. That is, this embodiment can output the scene label or the element label alone.

[0085] In some other embodiments, the above multimodal large language model may also be trained based on the device description information. Then, before training the model, the following processing may also be performed: Obtain the device description information corresponding to the sample image; Based on the device description information, determine the device characteristics of the acquisition device.

[0086] Among them, the device description information is the information of the acquisition device that acquires the sample image, and may include information such as the "deployment location" and "scene" description of the acquisition device. For example, the device description information may be "South of the intersection of Jianshe Road, Xiejiacun (panoramic)". The device characteristics may include at least one of the device type or the installation scene. Here, the device type refers to the type determined according to the picture size and position captured by the device, such as panoramic, barrier gate access control, etc.

[0087] After obtaining the device description information corresponding to the sample image in this embodiment, based on the device description information, the device characteristics of the acquisition device can be determined. The device characteristics can include scene-related information, and using it in the model training process can improve the accuracy of scene prediction.

[0088] In the case of obtaining the above device characteristics, the model can be trained based on the above training set and the device characteristics together. Correspondingly, step S3 above can specifically include the following processing: input the training samples in the training set and the device characteristics corresponding to the sample images in the training samples into the scene recognition module, and output the scene prediction results of each scene in the samples included in the training samples; input the scene prediction results into the element recognition module, and output the element prediction results of each element in the sample images included in the training samples; input the device characteristics, the scene prediction results, and the element prediction results into the image description module, and output the image description prediction results of the sample images included in the training samples.

[0089] It can be understood that the model training process here is the same as the model training process without adding device characteristics above (except for the different input data), and the information contained in the scene prediction results, element prediction results, and image description prediction results is also the same as the scene prediction results, element prediction results, and image description prediction results in the model training process without adding device characteristics above, which will not be elaborated here.

[0090] In the process of training the model in this embodiment, the device characteristics are added to the input data. The device type can be judged based on the device characteristics, and more accurate scene information can be obtained based on the device type. Therefore, the accuracy of the model predicting the scene label can be improved.

[0091] In addition, a step of detecting abnormal images can also be added in the model training process, that is, the quality of the image is detected before scene prediction to determine whether there are abnormal phenomena such as "green screen" and "ghosting" in the sample image. If so, the sample image is ignored and the subsequent training process is stopped. The subsequent training process is only carried out when the sample image has no abnormalities.

[0092] Specifically, as shown in Table 1 below, the model can be trained in the form of multiple rounds of conversations. Assuming that the two ends of the conversation are R and U, the training process of multiple rounds of conversations is as follows. In the conversation, the content in () is the content that needs to be generated by the large model, and the rest is the input content.

[0093] R: Now there is an image of x, and the details of the image are as follows: Image A;

[0094] U: Please judge whether there is an abnormal image problem;

[0095] R: The current image (has no abnormal image problem);

[0096] U: Please determine which device captured the image. The available device types are "...";

[0097] R: The current image was captured by the (xx) camera;

[0098] U: Please determine which scenes exist in the image. The available scene labels for reference are "...";

[0099] R: Based on the content of the picture and the text information, the scenes (xxx, xxx) exist in the picture;

[0100] U: Please continue to detect which picture elements exist in the image. The available element labels for reference are "...";

[0101] R: Based on the content of the picture and the text information, the elements (xxxx, xxxx) exist in the picture;

[0102] U: Based on the above content, describe the content of the picture in detail;

[0103] R: (……).

[0104] Table 1

[0105]

[0106]

[0107] As described above, the model training is carried out in a multi-round dialogue manner. Only one set of training samples is required, and all modules of the multi-modal large language model can be trained through one data input, thereby improving the training efficiency, significantly reducing the time consumed in the training process, and ensuring the coherence of the inference in the training process.

[0108] In some other embodiments, the model training method can also adjust the model parameters based on the loss value. Specifically, the first loss value can be calculated based on the prediction results of each scene and the scene set included in the training samples; the first model parameters of the scene recognition module can be adjusted based on the first loss value; the second loss value can be calculated based on the prediction results of each element and the element set included in the training samples; the second model parameters of the element recognition module can be adjusted based on the second loss value; the third loss value can be calculated based on the prediction results of the image description and the image description information included in the training samples; the third model parameters of the image description module can be adjusted based on the third loss value; based on the adjusted first model parameters, second model parameters, and third model parameters, the step of obtaining the prediction results of each scene is repeatedly executed until the training is completed when the preset convergence condition is reached to obtain the multi-modal large language model.

[0109] Among them, the first model parameter, the second model parameter, and the third model parameter do not refer to a single data alone, and can each be a single data or a set of data. The first model parameter can be regarded as the model parameter of the scene recognition module, the second model parameter can be regarded as the model parameter of the feature recognition module, and the third model parameter can be regarded as the model parameter of the image description module.

[0110] In this embodiment, the loss values are calculated based on the scene recognition module, the feature recognition module, and the image description module respectively, and the model parameters of each module can be adjusted independently. When the loss values of each module of the multi-modal large prediction model reach the preset convergence conditions, it can be considered that the multi-modal large prediction model as a whole reaches the convergence conditions. In this way, while obtaining the trained multi-modal large prediction model, an independent labeling model and an image description model can also be obtained.

[0111] It should be noted that the specific calculation processes of the above first loss value, second loss value, and third loss value, as well as the loss functions used, are not specifically limited in this embodiment, as long as the difference between the prediction result and the corresponding feature set can be calculated. For example, the loss function can be a cross-entropy loss function, a regression loss function, a classification loss function, etc.

[0112] In some other embodiments, the model training method may further include the following processing: based on a preset label library, obtaining each scene prediction result and each feature prediction result through the multi-modal large language model; if there are target scene labels and / or target picture element labels in the multiple training samples that do not exist in the preset label library, then add the target scene labels and / or target picture element labels to the preset label library.

[0113] Among them, the preset label library is a vocabulary library required during the output process of the multi-modal large language model.

[0114] In this embodiment, before training the multi-modal large language model, it is equipped with a corresponding vocabulary library. When outputting labels, the corresponding vocabulary can be searched based on an identified word to achieve the purpose of fast output. However, there may be target scene labels and / or target picture element labels in the training samples that do not exist in the preset label library. At this time, the target scene labels and / or target picture element labels can be directly added to the preset label library to improve the output efficiency and output accuracy of the multi-modal large language model.

[0115] This embodiment also compares the labeling accuracy of the existing multi-label labeling model with the labeling accuracy of the multi-modal large language model trained using the above model training method. When tested on 2000 random samples, the average accuracy rate of labeling using the multi-modal large language model obtained in this embodiment can reach 94.89%, and the highest average accuracy rate of labeling using other multi-label labeling models only reaches 82.4%.

[0116] In summary, for the model training method provided in this embodiment, multiple sample data are first obtained. The sample data include sample images, image description information of the sample images, and multiple scene sets and multiple element sets corresponding to the sample images. Then, a multi-modal large language model is trained based on the training set constructed from these sample data. One sample data in this model training method includes one sample image and multiple scene sets and multiple element sets of this image. In this way, even with a small amount of sample data, the trained model can not only learn the belonging relationship and dependency relationship between scenes and elements, but also avoid the model learning unnecessary sequential / hierarchical relationships between scenes and elements, so that an image processing model with more accurate labeling effect and more appropriate image description information can be obtained.

[0117] Some embodiments of this application also provide an image processing method, as Figure 4 shown. This image processing method includes the following steps:

[0118] Step S10: Obtain a target image to be processed;

[0119] Step S20: Input the target image into a pre-trained multi-modal large language model, and output the image description information of the target image;

[0120] Among them, the multi-modal large language model is obtained by the above model training method.

[0121] In some embodiments, when obtaining the target image to be processed, a large amount of image data to be processed can be first obtained, and abnormal picture detection is performed on these image data; the image data without detected abnormal pictures is determined as the target image.

[0122] Specifically, as Figure 5 shown, it is the specific process of using the above-trained multi-modal large language model to perform image description. The implementation process is as follows:

[0123] Step1: Input the image into the trained multi-modal large language model, perform picture abnormality diagnosis on the input image data, determine whether the input image picture contains abnormal phenomena such as "green screen" and "ghosting", and can output the image data without abnormal phenomena and delete the images with abnormal phenomena.

[0124] Step2: Input the device description information and the image data without abnormal phenomena into the trained multi-modal large language model together to perform device type judgment, and output the device type of the acquisition device of this image data, such as panoramic or barrier gate access control, etc.

[0125] Step 3: Directly based on the output result of Step 2, continue with the scene labeling process of the image frame and output the scene prediction result of the image data. For example, the scenes in the frame can include: parking lot, street store, intersection, etc.

[0126] Step 4: Directly based on the output result of Step 3, continue with the element labeling process of the image frame and output the element prediction result of the image data. For example, the elements in the frame can include: no-parking area, parking space, etc.

[0127] Step 5: Generate image description information based on the output results of Step 2 - Step 4.

[0128] The image processing method provided in this embodiment, using the model trained by the model training method provided in the embodiments of this application, can generate scene labels and element labels with higher accuracy and describe the image more accurately.

[0129] Some embodiments of this application also provide a model training device, which is used to execute the model training method provided in any of the above embodiments. Figure 6 A schematic diagram of the model training device is shown, as Figure 6 shown, the model training device includes:

[0130] A sample data acquisition module, which is used to acquire multiple sample data. The sample data includes sample images, image description information of the sample images, and multiple scene sets and multiple element sets corresponding to the sample images; the multiple scene sets are obtained by permuting and combining the scene labels marked in the sample images, and the multiple element sets are obtained by permuting and combining the frame element labels marked in the sample images; the image description information is used to describe the content of the sample images.

[0131] A training set construction module, which is used to construct a training set based on multiple sample data.

[0132] A model training module, which is used to train a multi-modal large language model based on the training set.

[0133] It can be understood that the model training device provided in this embodiment and the model training method provided in the embodiments of this application are based on the same inventive concept, and can at least achieve the same beneficial effects as the model training method. Moreover, all the implementation manners of the model training method embodiments are also equally applicable to the embodiments of this model training device, and will not be elaborated here.

[0134] Some embodiments of this application also provide an image processing device, which is used to execute the image processing method provided in any of the above embodiments. Figure 7 A schematic diagram of the image processing device is shown, as Figure 7As shown, the image processing device includes:

[0135] An image acquisition module for acquiring a target image to be processed;

[0136] An image processing module for inputting the target image into a pre-trained multi-modal large language model and outputting image description information of the target image;

[0137] Among them, the multi-modal large language model is obtained by using the above-mentioned model training method.

[0138] It can be understood that the image processing device provided in this embodiment and the image processing method provided in the embodiments of the present application are based on the same inventive concept, and can at least achieve the same beneficial effects as the image processing method. Moreover, various implementation manners of the embodiments of the image processing method are also equally applicable to the embodiments of this image processing device, and will not be elaborated here.

[0139] It should be noted that the data involved in the present application (including but not limited to data for model training, stored data, displayed data, etc.) are all information and data authorized by users or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.

[0140] The embodiments of the present application also provide an electronic device to execute the above-mentioned model training method or model training method. Please refer to Figure 8 , which shows a schematic diagram of an electronic device provided by some embodiments of the present application. As Figure 8 shown, the electronic device 4 includes: a processor 400, a memory 401, a bus 402, and a communication interface 403. The processor 400, the communication interface 403, and the memory 401 are connected through the bus 402; a computer program that can run on the processor 400 is stored in the memory 401, and when the processor 400 runs the computer program, it executes the model training method or model training method provided in any of the foregoing embodiments of the present application.

[0141] Among them, the memory 401 may include a high-speed random access memory (RAM: Random Access Memory), and may also include a non-volatile memory, such as at least one disk memory. Through at least one communication interface 403 (which can be wired or wireless), a communication connection is realized between this device network element and at least one other network element, and the Internet, wide area network, local area network, metropolitan area network, etc. can be used.

[0142] The bus 402 can be an ISA bus, a PCI bus, an EISA bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. Among them, the memory 401 is used to store a program. After receiving an execution instruction, the processor 400 executes the program. Any implementation manner of the model training method disclosed in any implementation manner of the embodiments of the present application or the model training method can be applied to the processor 400 or implemented by the processor 400.

[0143] The processor 400 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 400 or by an instruction in the form of software. The above-mentioned processor 400 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 401, and the processor 400 reads the information in the memory 401 and combines its hardware to complete the steps of the above method.

[0144] The electronic device provided by the embodiments of the present application and the model training method or model training method provided by the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run, or implemented by them.

[0145] The embodiments of the present application also provide a computer-readable storage medium corresponding to the model training method or model training method provided in the foregoing embodiments. Please refer to Figure 9 , which shows that the computer-readable storage medium is an optical disc 50, on which a computer program (i.e., a program product) is stored. When the computer program is run by a processor, it will execute the model training method or model training method provided in any of the foregoing embodiments.

[0146] It should be noted that examples of computer-readable storage media may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other optical and magnetic storage media, which will not be elaborated here one by one.

[0147] The embodiments of the present application also provide a computer program product, including a computer program, which is executed by a processor to implement the model training method or the model training method of any of the above embodiments.

[0148] The computer-readable storage media and computer program products provided in the above embodiments of the present application are all based on the same inventive concept as the model training method or the model training method provided in the embodiments of the present application, and have the same beneficial effects as the methods adopted, run, or implemented by the application programs stored therein.

[0149] It should be noted that:

[0150] In the specification provided here, a large number of specific details are described. However, it can be understood that the embodiments of the present application can be practiced without these specific details. In some instances, well-known structures and technologies are not shown in detail so as not to obscure the understanding of this specification.

[0151] Similarly, it should be understood that, in order to streamline the present application and help understand one or more of the various inventive aspects, in the above description of the exemplary embodiments of the present application, the various features of the present application are sometimes grouped together into a single embodiment, figure, or description thereof. However, the disclosed method should not be construed as reflecting the following schematic: that the claimed present application requires more features than those expressly recited in each of the claims. Rather, as reflected in the following claims, the inventive aspect lies in less than all the features of the single embodiment disclosed above. Therefore, the claims following the detailed description are hereby expressly incorporated into the detailed description, where each claim itself serves as a separate embodiment of the present application.

[0152] In addition, those skilled in the art can understand that, although some of the embodiments described herein include certain features included in other embodiments but not other features, the combination of features of different embodiments means that it is within the scope of the present application and forms different embodiments. For example, in the following claims, any of the claimed embodiments can be used in any combination.

[0153] As described above, it is only the preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present application should be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims described above.

Claims

1. A model training method, characterized in that, The method includes: Obtaining a plurality of sample data, where the sample data includes a sample image, image description information of the sample image, and a plurality of scene sets and a plurality of element sets corresponding to the sample image; the plurality of scene sets are obtained by permuting and combining the respective scene labels annotated in the sample image, and the plurality of element sets are obtained by permuting and combining the respective picture element labels annotated in the sample image; the image description information is used to describe the scenes and elements included in the sample image, as well as the relationship between the elements and the scenes; Constructing a training set based on the plurality of sample data; Training a multi-modal large language model based on the training set.

2. The method according to claim 1, characterized in that, The constructing a training set based on the plurality of sample data includes: For each sample data, combining the sample image and the image description information in the same sample data with one scene set from each of the scene sets and one element set from each of the element sets in the sample data to form a training sample; Combining the obtained plurality of training samples to form a training set.

3. The method according to claim 2, characterized in that, The multi-modal large language model includes a scene recognition module, an element recognition module, and an image description module; the training the multi-modal large language model based on the training set includes: Inputting the training samples in the training set into the scene recognition module to output respective scene prediction results corresponding to the sample images included in the training samples; Inputting the respective scene prediction results into the element recognition module to output respective element prediction results corresponding to the sample images included in the training samples; Inputting the respective scene prediction results and the respective element prediction results into the image description module to output an image description prediction result of the sample image included in the training sample.

4. The method according to claim 2, characterized in that, The method further includes: Obtaining device description information corresponding to the sample image, where the device description information is information about the acquisition device that acquired the sample image; Determining the device characteristics of the acquisition device based on the device description information, where the device characteristics include at least one of the device type or the installation scene.

5. The method according to claim 4, wherein The multi-modal large language model includes a scene recognition module, an element recognition module, and an image description module; the training the multi-modal large language model based on the training set includes: Inputting the training samples in the training set and the device characteristics corresponding to the sample images in the training samples into the scene recognition module to output respective scene prediction results in the samples included in the training samples; Inputting the respective scene prediction results into the element recognition module to output respective element prediction results in the sample images included in the training samples; Inputting the device characteristics, the respective scene prediction results, and the respective element prediction results into the image description module to output an image description prediction result of the sample image included in the training sample.

6. The method according to claim 3 or 5, characterized in that, The training the multi-modal large language model based on the training set further includes: Calculating a first loss value based on the respective scene prediction results and the scene sets included in the training samples; adjusting the first model parameters of the scene recognition module based on the first loss value; Calculate a second loss value based on the prediction results of the respective elements and the set of elements included in the training samples; adjust the second model parameters of the element recognition module based on the second loss value; Calculate a third loss value based on the prediction result of the image description and the image description information included in the training samples; adjust the third model parameters of the image description module based on the third loss value; Based on the adjusted first, second, and third model parameters, return and repeatedly execute the steps of obtaining the prediction results for each scenario until a preset convergence condition is reached, and then obtain a trained multi-modal large language model.

7. The method according to claim 3 or 5, characterized in that, The method further includes: Based on a preset label library, obtain the prediction results for each scenario and the prediction results for each element through the multi-modal large language model; If there are target scenario labels and / or target picture element labels in the multiple training samples that do not exist in the preset label library, add the target scenario labels and / or the target picture element labels to the preset label library.

8. An image processing method, characterized in that, It includes: Obtain a target image to be processed; Input the target image into a pre-trained multi-modal large language model to output the image description information of the target image; Wherein, the multi-modal large language model is obtained by using the model training method according to any one of claims 1-7.

9. The method according to claim 8, wherein The obtaining of the target image to be processed includes: Obtain image data to be processed and perform abnormal picture detection on the image data; Determine the image data without detected abnormal pictures as the target image.

10. A model training device, characterized in that, It includes: A sample data acquisition module, configured to acquire a plurality of sample data, where the sample data includes a sample image, the image description information of the sample image, and a plurality of scenario sets and a plurality of element sets corresponding to the sample image; the plurality of scenario sets are obtained by permuting and combining the respective scenario labels annotated in the sample image, and the plurality of element sets are obtained by permuting and combining the respective picture element labels annotated in the sample image; the image description information is used to describe the content of the sample image; A training set construction module, configured to construct a training set based on the plurality of sample data; A model training module, configured to train a multi-modal large language model based on the training set.

11. An image processing apparatus, characterized in that, It includes: An image acquisition module, configured to acquire a target image to be processed; An image processing module, configured to input the target image into a pre-trained multi-modal large language model to output the image description information of the target image; Wherein, the multi-modal large language model is obtained by using the model training method according to any one of claims 1-7.

12. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor executes the program to implement the method according to any one of claims 1-9.

13. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to implement the method according to any one of claims 1-9.