Image processing method, device and equipment based on large model and readable storage medium
Through the large-model-based image processing method, image annotation is automatically generated, which solves the problem of low image annotation efficiency and accuracy in the prior art, and realizes a more efficient and accurate image annotation process.
Patent Information
- Application Number
- CN202311805455.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-25
- Publication Date
- 2025-06-27
AI Technical Summary
In the prior art, image annotation relies on manual operation, which is time-consuming, labor-intensive and error-prone, resulting in low efficiency and accuracy.
The image processing method based on the big model is adopted to generate the annotation image by obtaining the image to be processed, generating feature map information, and automatically labeling using the annotation model.
Improve the efficiency and accuracy of image annotation, and realize automatic annotation through pre-trained models, reducing manual intervention, and improving processing speed and accuracy of results.
Smart Images

Figure CN120219781A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to an image processing method, apparatus, device, and readable storage medium based on a large model. Background Art
[0002] Image annotation is a key step in machine learning. However, existing image annotations are all manually performed by professionals on images. Manual annotation is usually time-consuming, labor-intensive, and prone to errors, resulting in low efficiency and accuracy of data annotation. Therefore, how to improve the efficiency and accuracy of image annotation is an urgent problem to be solved. Summary of the Invention
[0003] Embodiments of this application provide an image processing method, apparatus, device, and readable storage medium based on a large model, which can improve the efficiency and accuracy of image annotation.
[0004] In a first aspect, embodiments of this application provide an image processing method based on a large model, the method including:
[0005] Obtain an image to be processed;
[0006] Generate first feature map information based on the image to be processed and a feature map generation model;
[0007] Perform annotation processing on the image to be processed based on an annotation model and the first feature map information to obtain an annotated image.
[0008] In a second aspect, embodiments of this application provide an image processing apparatus based on a large model,
[0009] An obtaining unit, configured to obtain an image to be processed;
[0010] A first generation unit, configured to generate first feature map information based on the image to be processed and a feature map generation model;
[0011] A second generation unit, configured to perform annotation processing on the image to be processed based on an annotation model and the first feature map information to obtain an annotated image.
[0012] In a third aspect, embodiments of this application further provide an image processing device based on a large model, including a memory storing multiple instructions; a processor loads instructions from the memory to execute the steps of any one of the image processing methods based on a large model provided by embodiments of this application.
[0013] In a fourth aspect, embodiments of this application further provide a readable storage medium, the readable storage medium storing multiple instructions, the instructions being suitable for being loaded by a processor to execute the steps of any one of the image processing methods based on a large model provided by embodiments of this application.
[0014] In a fifth aspect, an embodiment of the present application further provides a computer program product, including a computer program or instructions, which when executed by a processor, implement the steps in any one of the image processing methods based on a large model provided by the embodiments of the present application.
[0015] Adopting the solution of the application embodiment, obtain the image to be processed; generate first feature map information based on the image to be processed and the feature map generation model; perform annotation processing on the image to be processed based on the annotation model and the first feature map information to obtain an annotated image. Automatically annotating the image to be processed based on a pre-trained model improves the efficiency of image processing based on a large model. Generate the first feature map information corresponding to the image to be processed through a pre-trained feature map generation model, perform annotation on the image to be processed based on the first feature map information through a pre-trained annotation model to generate an annotated image, and perform automatic annotation on the image to be processed based on the first feature map information, improving the accuracy of image processing based on a large model. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0017] Figure 1 It is a schematic flowchart of the first embodiment of the image processing method based on a large model provided by the present application;
[0018] Figure 2 It is a schematic flowchart of the training process of the feature map generation model of the image processing method based on a large model provided by the present application;
[0019] Figure 3 It is a schematic flowchart of the third embodiment of the image processing method based on a large model provided by the present application;
[0020] Figure 4 It is a schematic structural diagram of the image processing device based on a large model provided in the embodiments of the present application;
[0021] Figure 5 It is a schematic structural diagram of the image processing device based on a large model provided in the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application. At the same time, in the description of the embodiments of the present application, terms such as "first" and "second" are only used for differential description and cannot be understood as indicating or implying relative importance. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more features. In the description of the embodiments of the present application, the meaning of "a plurality" is two or more, unless otherwise clearly and specifically defined.
[0023] The embodiments of the present application provide an image processing method, device, equipment and readable storage medium based on a large model.
[0024] Specifically, this embodiment will be described from the perspective of an image processing device based on a large model. The image processing device based on a large model can be specifically integrated in an image processing device based on a large model, that is, the image processing method based on a large model in the embodiments of the present application can be executed by an image processing device based on a large model.
[0025] The image processing method based on a large model provided by the embodiments of the present application can be applied to, for example, an image processing device based on a large model.
[0026] The following will be described in detail with reference to the accompanying drawings. In this embodiment, the execution subject is an image processing device based on a large model as an example. It should be noted that the description order of the following embodiments does not limit the preferred order of the embodiments. Although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order from that shown in the drawings.
[0027] Please refer to Figure 1 , and a first embodiment of the present application is proposed. The specific process of the image processing method based on a large model includes the following steps:
[0028] Step 101, obtain an image to be processed;
[0029] Step 102, generate first feature map information based on the image to be processed and a feature map generation model;
[0030] Step 103, perform annotation processing on the image to be processed based on an annotation model and the first feature map information to obtain an annotated image.
[0031] In this embodiment, when it is necessary to annotate a certain type of image, the user uploads these images to the image processing device based on the large model. The image processing device based on the large model then obtains these images as the images to be processed. The image processing device based on the large model generates the first feature map information corresponding to the image to be processed based on the image to be processed and the pre-trained feature map generation model, and then inputs the first feature map information and the image to be processed into the pre-trained annotation model. The annotation model performs annotation processing on the image to be processed based on the first feature map information to generate an annotated image. It should be noted that the first feature map information is a guiding feature map, which indicates the annotation target and the distribution of annotation target features, provides guidance for the annotation model, helps the annotation model understand the requirements for annotating the image to be processed, and enables the annotation model to accurately annotate the image to be processed to generate an annotated image.
[0032] The image processing device based on the large model in this embodiment obtains the image to be processed; generates the first feature map information based on the image to be processed and the feature map generation model; and performs annotation on the image to be processed based on the first feature map information through the annotation model to generate an annotated image. Automatic annotation of the image to be processed is realized based on the pre-trained model, which improves the efficiency of image processing based on the large model. The first feature map information corresponding to the image to be processed is generated through the pre-trained feature map generation model, and the image to be processed is annotated based on the first feature map information through the pre-trained annotation model to generate an annotated image. Automatic annotation of the image to be processed is performed based on the first feature map information, which improves the accuracy of image processing based on the large model.
[0033] Specifically, each step is described in detail as follows:
[0034] Step 101, obtain the image to be processed;
[0035] In this step, an interactive interface or an interactive interface is set in the image processing device based on the large model. When it is necessary to annotate a certain type of image, the user uploads these images to be annotated to the image processing device based on the large model through the interactive interface or the interactive interface. Then, the image processing device based on the large model obtains these images as the images to be processed. It should be noted that the user can upload a large number of images to be processed, and the image processing device based on the large model can annotate all the images to be processed to improve the annotation efficiency. Moreover, all the uploaded images to be processed are of the same type, such as all images contain animals, all images contain circuit board defects, or all images contain human faces.
[0036] Step 102, generate the first feature map information based on the image to be processed and the feature map generation model;
[0037] In this step, after the large model-based image processing device obtains the image to be processed, according to the image to be processed, it obtains the corresponding target image, inputs the target image into the pre-trained feature map generation model, and generates the first feature map information corresponding to the image to be processed through the feature map generation model. It should be noted that the role of the first feature map information is to provide a reference or guidance for the annotation model to help the annotation model generate accurate and relevant annotation results. The first feature map information is usually an image similar to or related to the image to be processed, which can be images of different angles or scales of the same scene, or images containing similar objects or backgrounds. The first feature map information in this solution also includes the annotation target and the annotation target feature distribution. The first feature map information can have the following functions:
[0038] 1. Provide context and information: Through the first feature map information, the annotation model can obtain more information about the scene, objects, background, etc., so as to better understand the image to be processed. This helps the annotation model generate more accurate and precise annotation results.
[0039] 2. Guide the model's attention: The first feature map information can guide the attention of the annotation model, making it focus on specific regions or objects. The annotation model can learn to obtain key features from the first feature map information and apply these features to the annotation process of the image to be processed.
[0040] 3. Provide sample diversity: By using different first feature map information, the annotation model can be exposed to more sample diversity. This helps the annotation model better generalize to different image scenes and generate more robust annotation results.
[0041] Furthermore, referring to Figure 2 , before step 102, it includes:
[0042] Step a, obtain the pre-trained feature map generation model and the training annotation images;
[0043] In this step, before the large model-based image processing device performs large model-based image processing, it needs to train the feature map generation model; the large model-based image processing device obtains the pre-trained feature map generation model and the training annotation images, where the training annotation images contain real annotation information.
[0044] Step b, process the training annotation images through the pre-trained feature map generation model to obtain random mask images;
[0045] In this step, the large model-based image processing device processes the training annotation image through the pre-trained feature map generation model, randomly determines a certain number of masked pixel points in the training annotation image through the pre-trained feature map generation model, and covers or masks the masked pixel points, thereby introducing a certain degree of noise and variation to obtain a random masked image. Optionally, the large model-based image processing device randomly selects a certain number of masked pixel points in the training annotation image according to the task and data type through the pre-trained feature map generation model to obtain a random masked image; optionally, the large model-based image processing device randomly selects some pixel points to be masked completely through the pre-trained feature map generation model, or can also mask some pixel points according to certain rules or probabilities. In the training annotation image, some regions can be randomly selected and their pixel values can be set to zero or other fixed values to obtain a random masked image.
[0046] Step c, obtaining the first annotation information corresponding to the training annotation image through the pre-trained feature map generation model, and predicting the second annotation information corresponding to the random masked image;
[0047] In this step, the large model-based image processing device obtains the first annotation information corresponding to the training annotation image through the pre-trained feature map generation model according to the true annotation information contained in the training annotation image, and based on ViT, the large model-based image processing device divides the random masked image into a series of random masked sub-images through the pre-trained feature map generation model, extracts vectors representing the semantic and visual features of each random masked sub-image, inputs the vectors of each random masked sub-image into the Transformer, and predicts the second annotation information corresponding to the random masked image.
[0048] Step d, calculating a first loss value based on the first annotation information and the second annotation information until the first loss value meets the first convergence condition to obtain the feature map generation model.
[0049] In this step, the large model-based image processing device calculates a first loss value based on the first annotation information and the second annotation information, determines whether the first loss value meets the first convergence condition. If the first loss value does not meet the first convergence condition, the parameters in the pre-trained feature map generation model are adjusted based on the first loss value, and a new training annotation image is re-input to train the pre-trained feature map generation model until the calculated first loss value meets the first convergence condition, completing the training of the pre-trained feature map generation model to obtain the feature map generation model. Among them, the first convergence condition can be that the first loss value is less than the target loss value after continuous multiple predictions, or the first loss value remains unchanged after continuous multiple predictions, etc., which are not limited here.
[0050] Before performing image processing based on a large model, an image processing device based on a large model trains a feature map generation model, which facilitates automatically generating first feature map information corresponding to an image to be processed during subsequent image processing based on a large model. Compared with the current manual generation of first feature map information, the automatic generation of first feature map information based on the feature map generation model not only improves the generation efficiency of the first feature map information but also improves the generation accuracy of the first feature map information.
[0051] Specifically, step 102 includes:
[0052] Step 1021: Determine a target image based on the type of the image to be processed.
[0053] In this step, the image processing device based on a large model obtains the type of the image to be processed and then selects a target image corresponding to this type from the image library according to the type. It can be understood that the image library is preset in the image processing device based on a large model in advance. Various types of images are stored in the image library in advance, such as images containing animals, images containing circuit board defects, images containing human faces, etc. When the image processing device based on a large model obtains the type of the image to be processed, it can obtain an image of the corresponding type in the image library as the target image.
[0054] Furthermore, multiple images of the same type are stored in the image library. The image processing device based on a large model can first select a preset number of target images to be processed from all the images to be processed, and compare the multiple images of the same type as the images to be processed stored in the image library with the selected preset number of target images to be processed to determine the similarity between the images in the image library and each target image to be processed. Then, according to the similarity, select the image with the highest similarity from the multiple images of the same type as the images to be processed as the target image. Selecting the target image from multiple images of the same type as the images to be processed through similarity can improve the accuracy of determining the target image, facilitate obtaining the target image with the highest similarity to the image to be processed, and improve the accuracy of subsequent annotation of the image to be processed.
[0055] Step 1022: Input the target image into the feature map generation model to generate annotation feature information of the target image.
[0056] In this step, the image processing device based on the large model inputs the target image into the feature map generation model to generate the labeled feature information of the target image. Specifically, the feature map generation model includes ViT (Vision Transformer), which is a new type of vision model based on Transformer. The feature map generation model divides the target image into a series of target sub-images based on ViT, and then converts these target sub-images into vector representations through a linear transformation. These vectors are used to represent the semantic and visual features of each target sub-image. These vectors are used as the input sequence and input into the Transformer. The Transformer consists of multiple encoder layers, and each encoder layer contains a multi-head self-attention mechanism and a feed-forward neural network. The self-attention mechanism is used to calculate the correlation degree of each position in the input sequence, so as to obtain the context information of the target image. Through the stacking of multiple encoder layers, the Transformer gradually extracts and integrates the features in the target image based on the pre-trained model parameters to obtain the labeled feature information of the target image. The labeled feature information refers to the image features of the labeled target and the position features of these image features in the target image.
[0057] Exemplarily, the target image is an image containing circuit board defects. The feature map generation model divides the target image into a series of target sub-images through ViT, extracts the vectors representing the semantic and visual features of each target sub-image, and inputs the vectors of each target sub-image into the Transformer with pre-trained model parameters. The pre-trained model parameters are the parameters that can match the semantic and visual features of the circuit board defects. Furthermore, the Transformer outputs the image features corresponding to the circuit board defects in the target image and the position features of these image features in the target image.
[0058] Step 1023: Generate the first feature map information based on the labeled feature information and the target image.
[0059] In this step, the image processing device based on the large model performs annotation processing and binarization processing on the target image based on the labeled feature information through the feature map generation model to obtain a binarized annotation image, and then generates the first feature map information according to the binarized annotation image and the target image.
[0060] Specifically, step 1023 includes:
[0061] Step 10231: Determine the pixels to be annotated in the target image based on the labeled feature information, and perform annotation on the pixels to be annotated to obtain an annotated target image;
[0062] Step 10232: Binarize the labeled target image to obtain a binarized labeled image, and generate first feature map information based on the binarized labeled image and the target image.
[0063] In steps 10231 to 10232, the image processing device based on the large model determines the pixels to be labeled among all the pixels in the target image according to the labeled feature information through the feature map generation model, and then labels each pixel to be labeled in the target image to obtain a labeled target image. The image processing device based on the large model then binarizes the labeled target image to obtain a binarized labeled image, and generates first feature map information based on the binarized labeled image and the target image. That is, the first feature map information includes the binarized labeled image and the target image.
[0064] Step 103: Based on the labeling model and the first feature map information, perform labeling processing on the image to be processed to obtain a labeled image.
[0065] In this step, the image processing device based on the large model inputs the first feature map information and the image to be processed into the labeling model. The labeling model obtains feature information based on the first feature map information and adjusts the model attention. The labeling model comprehensively considers the features and context information of the image to be processed, as well as the correlation between the image to be processed and the first feature map information, and performs labeling on the image to be processed to generate a labeled image. Preferably, the labeling model uses the SegmentEverything large model as the basic model, and then fine-tunes the SegmentEverything large model with a large amount of image data of the same type as the image to be processed to improve the understanding ability of the SegmentEverything large model for image data of the same type as the image to be processed, and obtains a labeling model that can automatically label the image to be processed. The labeling model has a strong ability to understand the first feature map information and can perform labeling on the image to be processed according to the input first feature map information (one or more) to generate a labeled image.
[0066] Specifically, step 103 includes:
[0067] Step 1031: Obtain the labeled target information and the labeled target feature distribution information corresponding to the first feature map information through the labeling model;
[0068] In this step, after the large model-based image processing device inputs the first feature map information into the annotation model, it obtains the annotation target information and the annotation target feature distribution information corresponding to the first feature map information through the annotation model. Among them, the annotation target information is the information of the object to be annotated in the first feature map information, and the annotation target feature distribution information is the distribution of semantic and visual features among the pixel points corresponding to the annotation target in the first feature map information. Exemplarily, the first feature map information is the first feature map information containing circuit board defects, the annotation target information is the circuit board defect information, including the type, size, location, etc. of the circuit board defects, and the annotation target feature distribution information is the distribution of semantic and visual features among the pixel points corresponding to the circuit board defect location in the image containing the circuit board defects.
[0069] Step 1032: Extract the feature information of the image to be processed based on the annotation model and the annotation target information.
[0070] In this step, the large model-based image processing device predicts the annotation target in the image to be processed through the annotation model based on the annotation target information of the first feature map information, intercepts the corresponding position of the annotation target in the image to be processed to obtain the annotation target image, extracts the feature map of the annotation target image, and the feature information of the feature map. Specifically, the annotation model includes ViT. After the large model-based image processing device obtains the annotation target image corresponding to the image to be processed through the annotation model, it divides the annotation target image into a series of annotation target sub-images through the annotation model based on ViT. These annotation target sub-images are the feature maps of the image to be processed, and then these annotation target sub-images are converted into vector representations through a linear transformation. These vector representations are used to represent the semantic and visual features of each annotation target sub-image.
[0071] Step 1033: Annotate the image to be processed based on the annotation model, the feature information, and the annotation target feature distribution information to obtain an annotated image.
[0072] In this step, after the large model-based image processing device determines the feature information of the feature map, it annotates the image to be processed through the annotation model based on the feature information and the annotation target feature distribution information to generate an annotated image. Specifically, the large model-based image processing device calculates the first correlation degree between the semantic and visual features corresponding to the annotation target in the first feature map information through the annotation model according to the annotation target feature distribution information and a pre-designed calculation method, and calculates the second correlation degree between the semantic and visual features corresponding to the annotation target in the image to be processed through the annotation model according to the feature information and the pre-designed calculation method. Then, based on the first correlation degree and the second correlation degree, it determines the pixel points to be annotated in the image to be processed, and then annotates the pixel points to be annotated to generate an annotated image.
[0073] It should be noted that to calculate the correlation degree between multiple different features, a covariance matrix or a correlation coefficient matrix can be used. These matrices provide information on the correlation degree between features. Covariance matrix: For a dataset with n features, the covariance matrix is an n×n matrix, where the (i,j)-th element represents the covariance between the i-th and j-th features. The calculation formula for the covariance matrix C is as follows: C = cov(X), where X is an n-dimensional dataset and cov() is the covariance function. Correlation coefficient matrix: The correlation coefficient matrix is also an n×n matrix, where the (i,j)-th element represents the correlation coefficient between the i-th and j-th features. The most commonly used is the Pearson correlation coefficient matrix. The calculation formula for the correlation coefficient matrix R is as follows: R = corr(X), where X is an n-dimensional dataset and corr() is the correlation coefficient function.
[0074] Furthermore, step 1033 includes:
[0075] Step 10331, based on the annotation model, the feature information, and the annotation target feature distribution information, perform annotation scoring on each pixel point in the feature map of the image to be processed, and obtain the scoring value of each pixel point;
[0076] In this step, after the image processing device based on the large model obtains the feature information corresponding to the feature map obtained by segmenting the image to be processed and the annotation target feature distribution information corresponding to the first feature map information through the annotation model, the annotation model performs annotation scoring on each pixel point in the feature map based on the feature information and the annotation target feature distribution information. Specifically, the annotation target feature distribution information is the distribution of features such as the shape, texture, and color of the annotation target in the first feature map information in the pixel points corresponding to the annotation target. The image processing device based on the large model calculates the first correlation degree between the features such as the shape, texture, and color corresponding to the annotation target in the first feature map information through the annotation model according to the annotation target feature distribution information and the pre-designed calculation method. The image processing device based on the large model calculates the second correlation degree between the features such as the shape, texture, and color included in each pixel point of each feature map obtained by segmenting the image to be processed through the annotation model according to the feature information and the pre-designed calculation method, and then compares the first correlation degree and the second correlation degree to perform annotation scoring on each pixel point in each feature map.
[0077] Optionally, the large model-based image processing device calculates the difference between the first correlation degree and the second correlation degree corresponding to each pixel point in each feature map through the annotation model. If the difference meets the preset condition, the corresponding pixel point is scored 1; if the difference does not meet the preset condition, the corresponding pixel point is scored 0. Optionally, the large model-based image processing device calculates the difference between the first correlation degree and the second correlation degree corresponding to each pixel point in each feature map through the annotation model, determines which preset difference range threshold the difference falls into according to the difference and the preset difference range threshold, and then uses the score corresponding to the preset difference range threshold as the annotation score of the corresponding pixel point.
[0078] Step 10332: Determine the pixel points to be annotated in the feature map based on the annotation model and the score values of each pixel point;
[0079] Step 10333: Annotate the image to be processed based on the annotation model and the pixel points to be annotated to obtain an annotated image.
[0080] In steps 10332 to 10333, the large model-based image processing device determines corresponding pixels to be labeled in each feature map obtained by segmenting the image to be processed based on the labeling scores corresponding to each pixel point in each feature map and a preset labeling score threshold, and then determines the pixels to be labeled on the image to be processed. Then, the image to be processed is labeled according to the pixels to be labeled to generate a labeled image. Optionally, the large model-based image processing device calculates the difference between the first correlation degree and the second correlation degree corresponding to each pixel point in each feature map through the labeling model. If the difference meets the preset condition, the score for the corresponding pixel point is set to 1; if the difference does not meet the preset condition, the score for the corresponding pixel point is set to 0. At this time, the preset labeling score threshold is set to 1. When the pixel point score is 1, the labeling model determines the pixel point as a pixel to be labeled; when the pixel point score is 0, the labeling model determines that the pixel point is not a pixel to be labeled. After the labeling model determines all the pixels to be labeled, the image to be processed is labeled according to the pixels to be labeled to generate a labeled image. Optionally, the large model-based image processing device calculates the difference between the first correlation degree and the second correlation degree corresponding to each pixel point in each feature map through the labeling model, determines which preset difference range threshold the difference falls into according to the difference and the preset difference range threshold, and then uses the score corresponding to the preset difference range threshold as the labeling score for the corresponding pixel point. At this time, the preset labeling score threshold is set to a specific score value, such as 95. When the pixel point score is greater than 95, the labeling model determines the pixel point as a pixel to be labeled; when the pixel point score is not greater than 95, the labeling model determines that the pixel point is not a pixel to be labeled. After the labeling model determines all the pixels to be labeled, the image to be processed is labeled according to the pixels to be labeled to generate a labeled image.
[0081] The large model-based image processing device in this embodiment obtains the image to be processed, generates first feature map information based on the image to be processed and the feature map generation model, and labels the image to be processed based on the first feature map information through the labeling model to generate a labeled image. The automatic labeling of the image to be processed is realized based on the pre-trained model, which improves the efficiency of large model-based image processing. The first feature map information corresponding to the image to be processed is generated through the pre-trained feature map generation model, and the image to be processed is labeled based on the first feature map information through the pre-trained labeling model to generate a labeled image. The automatic labeling of the image to be processed based on the first feature map information improves the accuracy of large model-based image processing.
[0082] Further, referring to Figure 3, the second embodiment of the present application is proposed. The difference between the second embodiment and the first embodiment is that before inputting the first feature map information and the image to be processed into the annotation model and generating an annotated image by annotating the image to be processed based on the first feature map information through the annotation model, it includes:
[0083] Step e, obtain a pre-trained annotation model, and obtain training images and training first feature map information according to the type of the image to be processed;
[0084] In this step, before the image processing device based on the large model performs image processing based on the large model, it is necessary to train the annotation model. The image processing device based on the large model obtains a pre-trained annotation model, and obtains training images and training first feature map information according to the type of the image to be processed input by the relevant personnel. Among them, the training images do not contain real annotation information, and the training first feature map information is obtained by a trained feature map generation model.
[0085] It should be noted that the pre-trained annotation model is the Segment Everything large model for image segmentation. The image segmentation large model has a strong ability to understand the first feature map information and can annotate the training images according to the input first feature map information (one or more) to generate predicted annotation images. Then, use a large number of training images of the same type as the image to be processed to segment the large model to improve the understanding ability of the image segmentation large model for the type of the image to be processed.
[0086] Step f, generate a predicted annotation image corresponding to the training image based on the training first feature map information through the pre-trained annotation model;
[0087] In this step, the image processing device based on the large model uses the pre-trained annotation model to segment the training image into a series of training sub-images according to the training first feature map information through ViT, extracts vectors representing the semantics and visual features of each training sub-image, and inputs the vectors of each training sub-image into the Transformer to generate a predicted annotation image corresponding to the training image.
[0088] Step g, obtain the real annotation image of the training image, and calculate a second loss value based on the predicted annotation image and the real annotation image until the second loss value meets the second convergence condition to obtain an annotation model.
[0089] In this step, when the large model-based image processing device obtains the training images, it simultaneously obtains the corresponding ground truth annotation images. After obtaining the predicted annotation images corresponding to the training images, the large model-based image processing device calculates the second loss value based on the predicted annotation images and the ground truth annotation images through the pre-trained annotation model, and determines whether the second loss value meets the second convergence condition. If the second loss value does not meet the second convergence condition, the parameters in the pre-trained annotation model are adjusted based on the second loss value, and new training images are re-input to train the pre-trained annotation model until the calculated second loss value meets the second convergence condition, completing the training of the pre-trained annotation model and obtaining the annotation model. Among them, the second convergence condition can be that the second loss value is less than the target loss value after consecutive multiple predictions, or the second loss value remains unchanged after consecutive multiple predictions, etc., which are not limited here.
[0090] Further, after obtaining the annotation model, it includes:
[0091] Step h, obtain the validation images and the preset standard first feature map information, and obtain the target verification condition from the verification condition information according to the type of the validation images;
[0092] In this step, after the large model-based image processing device trains to obtain the annotation model, it needs to further verify the annotation model; the large model-based image processing device obtains the validation images and the preset standard first feature map information, and obtains the target verification condition from the verification condition information according to the type of the validation images. It can be understood that the large model-based image processing device stores the verification condition information corresponding to various types of images in the database in advance, and the target verification condition can be obtained from the verification condition information according to the type of the validation images.
[0093] Step i, generate the validation annotation images corresponding to the validation images through the annotation model based on the preset standard first feature map information;
[0094] In this step, the large model-based image processing device inputs the preset standard first feature map information and the validation images into the annotation model. Through the annotation model, according to the preset standard first feature map information, the validation images are segmented into a series of validation sub-images by ViT, vectors representing the semantic and visual features of each validation sub-image are extracted, and the vectors of each validation sub-image are input into the Transformer to generate the validation annotation images corresponding to the validation images.
[0095] Step j, if the validation annotation images do not meet the target verification condition, re-train the annotation model.
[0096] In this step, after the image processing device based on the large model obtains the verification annotation image corresponding to the verification image, it determines whether the verification annotation image meets the target verification condition. If the verification annotation image does not meet the target verification condition, the annotation model is retrained. If the verification annotation image meets the target verification condition, the image to be processed is annotated based on the annotation model.
[0097] Before performing image processing based on the large model, the image processing device based on the large model in this embodiment trains an annotation model, which facilitates automatically generating the annotation image corresponding to the image to be processed when performing image processing based on the large model later. Compared with the current manual generation of annotation images, automatically generating annotation images based on the annotation model not only improves the generation efficiency of annotation images but also improves the generation accuracy of annotation images.
[0098] This embodiment also provides an image processing device based on a large model. This image processing device based on a large model can be specifically integrated in devices such as an image processing device based on a large model, as Figure 4 shown. This image processing device based on a large model may include:
[0099] An acquisition unit 1001, configured to acquire an image to be processed;
[0100] A first generation unit 1002, configured to generate first feature map information based on the image to be processed and a feature map generation model;
[0101] A second generation unit 1003, configured to perform annotation processing on the image to be processed based on the annotation model and the first feature map information to obtain an annotation image.
[0102] In an optional example, the first generation unit is configured to:
[0103] Determine a target image based on the type of the image to be processed;
[0104] Input the target image into the feature map generation model to generate the annotation feature information of the target image;
[0105] Generate first feature map information based on the annotation feature information and the target image.
[0106] In an optional example, the first generation unit is further configured to:
[0107] Determine the pixels to be annotated in the target image based on the annotation feature information, and annotate the pixels to be annotated to obtain an annotated target image;
[0108] Perform binarization processing on the annotated target image to obtain a binarized annotation image, and generate first feature map information based on the binarized annotation image and the target image.
[0109] In an optional example, the second generating unit is further configured to:
[0110] Obtain the labeled target information and the labeled target feature distribution information corresponding to the first feature map information through the labeling model;
[0111] Extract the feature information of the image to be processed based on the labeling model and the labeled target information;
[0112] Label the image to be processed based on the labeling model, the feature information, and the labeled target feature distribution information to obtain a labeled image.
[0113] In an optional example, the second generating unit is further configured to:
[0114] Perform a labeling score on each pixel point in the feature map of the image to be processed based on the labeling model, the feature information, and the labeled target feature distribution information to obtain the score value of each pixel point;
[0115] Determine the pixel points to be labeled in the feature map based on the labeling model and the score values of each pixel point;
[0116] Label the image to be processed based on the labeling model and the pixel points to be labeled to obtain a labeled image.
[0117] In an optional example, the image processing device based on the large model further includes a training unit, and the training unit is configured to:
[0118] Obtain a pre-trained labeling model, and obtain training images and training first feature map information according to the type of the image to be processed;
[0119] Generate a predicted labeled image corresponding to the training image based on the training first feature map information through the pre-trained labeling model;
[0120] Obtain the true labeled image of the training image, and calculate a second loss value based on the predicted labeled image and the true labeled image until the second loss value satisfies the second convergence condition to obtain a labeling model.
[0121] In an optional example, the image processing device based on the large model further includes a verification unit, and the verification unit is configured to:
[0122] Obtain verification images and preset standard first feature map information, and obtain target verification conditions in the verification condition information according to the type of the verification images;
[0123] Generate a verification labeled image corresponding to the verification image based on the preset standard first feature map information through the labeling model;
[0124] If the verified labeled image does not meet the target verification condition, retrain the labeling model.
[0125] Adopting the solution of this embodiment, an image to be processed is obtained; based on the image to be processed and a feature map generation model, first feature map information is generated; based on a labeling model and the first feature map information, the image to be processed is labeled to obtain a labeled image. Automatically labeling the image to be processed based on a pre-trained model improves the efficiency of image processing based on a large model. The first feature map information corresponding to the image to be processed is generated through the pre-trained feature map generation model, and the image to be processed is labeled based on the first feature map information through the pre-trained labeling model to generate a labeled image. Automatically labeling the image to be processed based on the first feature map information improves the efficiency of image processing based on a large model.
[0126] Correspondingly, an embodiment of the present application further provides an image processing device based on a large model, as Figure 5 shown Figure 5 is a schematic structural diagram of the image processing device based on a large model provided by an embodiment of the present application. The image processing device 1100 based on a large model includes a processor 1101 having one or more processing cores, a memory 1102 having one or more readable storage media, and a computer program stored in the memory 1102 and executable on the processor. Among them, the processor 1101 is electrically connected to the memory 1102. Those skilled in the art can understand that the structural diagram of the image processing device based on a large model shown in the figure does not constitute a limitation on the image processing device based on a large model, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0127] The processor 1101 is the control center of the image processing device 1100 based on a large model, connecting various parts of the entire image processing device 1100 through various interfaces and lines, and by running or loading software programs and / or units stored in the memory 1102, and calling data stored in the memory 1102, performing various functions of the image processing device 1100 and processing data, so as to monitor the entire image processing device 1100 based on a large model. The processor 1101 may be a central processing unit (CPU), a graphics processing unit (GPU), a network processor (NP), etc., and may implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application.
[0128] In an embodiment of the present application, the processor 1101 in the large model-based image processing device 1100 will load the instructions corresponding to the processes of one or more application programs into the memory 1102 according to the following steps, and the processor 1101 will run the application programs stored in the memory 1102 to implement various functions.
[0129] For the specific implementation of each of the above operations, reference may be made to the previous embodiments and will not be elaborated herein.
[0130] Optionally, as Figure 5 shown, the large model-based image processing device 1100 further includes: a touch display screen 1103, a radio frequency circuit 1104, an audio circuit 1105, an input unit 1106, and a power supply 1107. Among them, the processor 1101 is electrically connected to the touch display screen 1103, the radio frequency circuit 1104, the audio circuit 1105, the input unit 1106, and the power supply 1107 respectively. Those skilled in the art can understand that Figure 5 the structure of the large model-based image processing device shown in
[0131] The touch display screen 1103 can be used to display a graphical user interface and receive operation instructions generated by a user's interaction with the graphical user interface. The touch display screen 1103 may include a display panel and a touch panel. Among them, the display panel can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of the image processing device. These graphical user interfaces can be composed of graphics, text, icons, videos, and any combination thereof. Optionally, the display panel can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc. The touch panel can be used to collect touch operations of the user on or near it (such as operations of the user using any suitable object or accessory such as a finger or a stylus on or near the touch panel), and generate corresponding operation instructions, and the operation instructions execute the corresponding program. Optionally, the touch panel can include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch position of the user and detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor 1101, and can receive and execute the commands sent by the processor 1101. The touch panel can cover the display panel. When the touch panel detects a touch operation on or near it, it transmits it to the processor 1101 to determine the type of touch event. Subsequently, the processor 1101 provides a corresponding visual output on the display panel according to the type of touch event. In the embodiments of the present application, the touch panel and the display panel can be integrated into the touch display screen 1103 to implement input and output functions. However, in some embodiments, the touch panel and the touch panel can be implemented as two independent components to implement input and output functions. That is, the touch display screen 1103 can also be used as part of the input unit 1106 to implement the input function.
[0132] The radio frequency circuit 1104 can be used to transmit and receive radio frequency signals to establish wireless communication with a network device or other image processing devices through wireless communication, and transmit and receive signals with the network device or other image processing devices.
[0133] The audio circuit 1105 can be used to provide an audio interface between the user and the image processing device through a speaker and a microphone. The audio circuit 1105 can transmit the electrical signal converted from the received audio data to the speaker, which converts it into a sound signal for output; on the other hand, the microphone converts the collected sound signal into an electrical signal, which is received by the audio circuit 1105 and then converted into audio data. After the audio data is output and processed by the processor 1101, it is sent through the radio frequency circuit 1104 to, for example, another image processing device, or the audio data is output to the memory 1102 for further processing. The audio circuit 1105 may also include an earphone jack to provide communication between the peripheral earphone and the image processing device.
[0134] The input unit 1106 can be used to receive input digital, character information or user feature information (such as fingerprint, iris, face information, etc.), and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.
[0135] The power supply 1107 is used to supply power to each component of the image processing device 1100. Optionally, the power supply 1107 can be logically connected to the processor 1101 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 1107 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.
[0136] Although Figure 5 not shown in the figure, the image processing device 1100 may also include a camera, a sensor, a Wi-Fi module, a Bluetooth module, etc., which will not be elaborated here.
[0137] In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0138] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions, or by controlling relevant hardware through instructions. These instructions can be stored in a readable storage medium and loaded and executed by a processor.
[0139] Therefore, an embodiment of the present application provides a readable storage medium, in which multiple computer programs are stored. These computer programs can be loaded by a processor to execute any one of the image processing methods based on a large model provided by the embodiments of the present application.
[0140] For the specific implementation of the above operations, reference can be made to the previous embodiments, which will not be elaborated here.
[0141] Among them, the readable storage medium may include: read-only memory (ROM), random access memory (RAM), magnetic disk, optical disk, etc.
[0142] Since the computer program stored in the readable storage medium can execute any one of the large model-based image processing methods provided in the embodiments of the present application, the beneficial effects achievable by any one of the large model-based image processing methods provided in the embodiments of the present application can be realized. For details, please refer to the previous embodiments and will not be elaborated here.
[0143] According to an aspect of the present application, there is also provided a computer program product or computer program. The computer program product or computer program includes computer instructions, and the computer instructions are stored in a readable storage medium. The processor of the image processing device reads the computer instructions from the readable storage medium, and the processor executes the computer instructions, so that the image processing device executes the methods provided in various alternative implementations in the above embodiments.
[0144] In the above embodiments of the large model-based image processing device, readable storage medium, large model-based image processing device, and computer program product, the descriptions of each embodiment have their own focuses. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments. Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes and beneficial effects of the above-described large model-based image processing device, readable storage medium, computer program product, large model-based image processing device, and their corresponding units can refer to the description of the large model-based image processing method in the above embodiments and will not be elaborated here specifically.
[0145] The above has introduced in detail a large model-based image processing method, device, device, readable storage medium, and computer program product provided by the embodiments of the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. An image processing method based on a large model, characterized in that The large model-based image processing method includes: Obtain the image to be processed; Generate first feature map information based on the image to be processed and the feature map generation model; Perform annotation processing on the image to be processed based on the annotation model and the first feature map information to obtain an annotated image.
2. The image processing method based on a large model according to claim 1, wherein The generating of the first feature map information based on the image to be processed and the feature map generation model includes: Determine the target image based on the type of the image to be processed; Input the target image into the feature map generation model to generate the annotation feature information of the target image; Generate the first feature map information based on the annotation feature information and the target image.
3. The image processing method based on a large model according to claim 2, wherein The generating of the first feature map information based on the annotation feature information and the target image includes: Determine the pixels to be annotated in the target image based on the annotation feature information, and annotate the pixels to be annotated to obtain an annotated target image; Perform binarization processing on the annotated target image to obtain a binarized annotated image, and generate the first feature map information based on the binarized annotated image and the target image.
4. The image processing method based on a large model according to claim 1, wherein The performing of the annotation processing on the image to be processed based on the annotation model and the first feature map information to obtain an annotated image includes: Obtain the annotation target information and the annotation target feature distribution information corresponding to the first feature map information through the annotation model; Extract the feature information of the image to be processed based on the annotation model and the annotation target information; Perform annotation on the image to be processed based on the annotation model, the feature information, and the annotation target feature distribution information to obtain an annotated image.
5. The image processing method based on a large model according to claim 4, wherein The performing of the annotation on the image to be processed based on the annotation model, the feature information, and the annotation target feature distribution information to obtain an annotated image includes: Perform annotation scoring on each pixel in the feature map of the image to be processed based on the annotation model, the feature information, and the annotation target feature distribution information to obtain the scoring value of each pixel; Determine the pixels to be annotated in the feature map based on the annotation model and the scoring value of each pixel; Perform annotation on the image to be processed based on the annotation model and the pixels to be annotated to obtain an annotated image.
6. The image processing method based on a large model according to claim 1, wherein Before the performing of the annotation processing on the image to be processed based on the annotation model and the first feature map information to obtain an annotated image, it includes: Obtain a pre-trained annotation model, and obtain the training image and the training first feature map information according to the type of the image to be processed; Generate a predicted annotated image corresponding to the training image based on the training first feature map information through the pre-trained annotation model; Obtain the true annotated image of the training image, and calculate the second loss value based on the predicted annotated image and the true annotated image until the second loss value meets the second convergence condition to obtain the annotation model.
7. The image processing method based on a large model according to claim 6, wherein, After obtaining the annotation model, it includes: Obtain the verification image and the standard first feature map information, and obtain the target verification condition in the verification condition information according to the type of the verification image; Generate a verification annotated image corresponding to the verification image based on the preset standard first feature map information through the annotation model; If the verified annotated image does not meet the target verification condition, retrain the annotation model.
8. An image processing device based on a large model, characterized in that, The device includes: An acquisition unit, configured to acquire an image to be processed; A first generation unit, configured to generate first feature map information based on the image to be processed and a feature map generation model; A second generation unit, configured to perform annotation processing on the image to be processed based on an annotation model and the first feature map information to obtain an annotated image; Preferably, the first generation unit generates first feature map information based on the image to be processed and a feature map generation model, including: Determining a target image based on the type of the image to be processed; Inputting the target image into the feature map generation model to generate annotation feature information of the target image; Generating first feature map information based on the annotation feature information and the target image; Preferably, the first generation unit generates first feature map information based on the annotation feature information and the target image, including: Determining pixels to be annotated in the target image based on the annotation feature information, and annotating the pixels to be annotated to obtain an annotated target image; Performing binarization processing on the annotated target image to obtain a binarized annotated image, and generating first feature map information based on the binarized annotated image and the target image; Preferably, the second generation unit performs annotation processing on the image to be processed based on an annotation model and the first feature map information to obtain an annotated image, including: Obtaining annotation target information and annotation target feature distribution information corresponding to the first feature map information through the annotation model; Extracting feature information of the image to be processed based on the annotation model and the annotation target information; Performing annotation on the image to be processed based on the annotation model, the feature information, and the annotation target feature distribution information to obtain an annotated image; Preferably, the second generation unit performs annotation on the image to be processed based on the annotation model, the feature information, and the annotation target feature distribution information to obtain an annotated image, including: Performing annotation scoring on each pixel in the feature map of the image to be processed based on the annotation model, the feature information, and the annotation target feature distribution information to obtain a scoring value for each pixel; Determining pixels to be annotated in the feature map based on the annotation model and the scoring value of each pixel; Performing annotation on the image to be processed based on the annotation model and the pixels to be annotated to obtain an annotated image; Preferably, the annotation model is obtained by the second generation unit by performing the following steps: Obtaining a pre-trained annotation model, and acquiring training images and training first feature map information according to the type of the image to be processed; Generating a predicted annotated image corresponding to the training image based on the training first feature map information through the pre-trained annotation model; Obtaining the true annotated image of the training image, and calculating a second loss value based on the predicted annotated image and the true annotated image until the second loss value meets a second convergence condition to obtain an annotation model; Preferably, after the second generation unit obtains the annotation model, it includes: Obtain the verification image and the standard first feature map information, and obtain the target verification condition from the verification condition information according to the type of the verification image; Generate the verification annotation image corresponding to the verification image by the annotation model based on the preset standard first feature map information; If the verification annotation image does not meet the target verification condition, retrain the annotation model.
9. An image processing device based on a large model, characterized in that, It includes a processor and a memory, and the memory stores multiple instructions; the processor loads the instructions from the memory to execute the steps of the large model-based image processing method according to any one of claims 1-7.
10. A readable storage medium, characterized in that, The readable storage medium stores multiple instructions, and the instructions are suitable for being loaded by the processor to execute the steps of the large model-based image processing method according to any one of claims 1-7.