Model Training and Image Detection Methods, Devices, and Equipment

Through the semi-supervised learning method, the image detection model is trained using labeled and labelless training samples to generate pseudo-labels, which solves the problem of high labor cost in image detection model training and realizes efficient model training and detection.

CN115049892BActive Publication Date: 2025-07-25SHENZHEN HIVT TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210529423.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-16
Publication Date
2025-07-25
Estimated Expiration
2042-05-16

AI Technical Summary

Technical Problem

Existing image detection models require a large number of labeling personnel in training, which leads to high labor costs, especially in some scenarios where expert labeling is required.

Method used

The semi-supervised learning method is used to train the image detection model using labeled and labelless training samples, and generate pseudo-labels through the labeling model to reduce the number of labels on the sample images.

Benefits of technology

With the unchanged training sample size, the labor cost of model training is reduced, and the training efficiency and model accuracy are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115049892B_ABST
    Figure CN115049892B_ABST
Patent Text Reader

Abstract

The present application provides a model training and image detection method, apparatus and device. The method includes: obtaining training samples, where the training samples include multiple first sample images, the sample detection images of the first sample images, and multiple second sample images, and the sample detection images are images obtained by annotating sample objects on the first sample images; performing at least one semi-supervised model training on a first model according to the multiple first sample images, the sample detection images of the first sample images, and the multiple second sample images to obtain a trained first model. Without changing the scale of the training samples, since unlabeled training samples are added for training, the number of images to be annotated for the sample images is reduced, and the labor cost of model training is lowered.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of image processing technology, and in particular, to a model training and image detection method, apparatus, and device. Background Art

[0002] Image detection refers to the process of identifying whether there is a target object in an image and the position of the target object in the image. Image detection has a wide range of application scenarios, such as goods recognition, intelligent healthcare, and so on.

[0003] Currently, image detection is mainly implemented based on an image detection model. During the training process of the image detection model, it is necessary to obtain sample images and the annotation information of the sample images, and then train the image detection model with the sample images and the annotation information. After training is completed, the image detection model has the ability to detect images. Inputting the image to be detected into the image detection model can obtain the corresponding detection result.

[0004] Since the training samples required in the training of the image detection model are very large, a large number of annotators are often needed to annotate the sample images, so the labor cost of training the image detection model is relatively high. Summary of the Invention

[0005] The present application provides a model training and image detection method, apparatus, and device to solve the above problems.

[0006] In a first aspect, the present application provides a model training method, including:

[0007] Obtaining training samples, where the training samples include multiple first sample images, the sample detection images of the first sample images, and multiple second sample images, and the sample detection images are images obtained by annotating the sample objects on the first sample images;

[0008] Performing at least one semi-supervised model training on a first model according to the multiple first sample images, the sample detection images of the first sample images, and the multiple second sample images to obtain a trained first model.

[0009] In a possible implementation manner, performing at least one semi-supervised model training on a first model according to the multiple first sample images, the sample detection images of the first sample images, and the multiple second sample images to obtain a trained first model includes:

[0010] Obtaining a trained annotation model according to the multiple first sample images and the sample detection images of the first sample images;

[0011] Inputting the second sample images into the trained annotation model to obtain the sample detection images of the second sample images;

[0012] Train the first model based on the first sample image, the sample detection image of the first sample image, the second sample image, and the sample detection image of the second sample image to obtain the trained first model.

[0013] In a possible implementation, obtaining a trained annotation model according to the multiple first sample images and the sample detection images of the first sample images includes:

[0014] For any one of the multiple first sample images, input the first sample image into the annotation model to obtain an annotation detection image output by the annotation model;

[0015] Adjust the parameters of the annotation model according to the annotation detection image and the sample detection image of the first sample image;

[0016] When the difference value between the annotation detection image and the sample detection image of the first sample image is less than or equal to a first preset value, obtain the trained annotation model.

[0017] In a possible implementation, the first model includes a segmentation sub-model and an identification sub-model; training the first model according to the first sample image, the sample detection image of the first sample image, the second sample image, and the sample detection image of the second sample image to obtain the trained first model includes:

[0018] For any sample image among the multiple first sample images and the multiple second sample images, input the sample image into the segmentation sub-model to obtain a segmentation image output by the segmentation sub-model. The segmentation image includes a foreground region and a background region. The pixel values of the pixel points in the foreground region are first pixel values, and the pixel values of the pixel points in the background region are second pixel values;

[0019] Input the sample image into the identification sub-model to obtain a candidate result output by the identification sub-model. The candidate result indicates that the sample image includes a target object, or does not include the target object;

[0020] Obtain a candidate detection image of the sample image according to the segmentation image and the candidate result;

[0021] Adjust the parameters of the first model according to the difference value between the candidate detection image and the sample detection image of the sample image until the difference value between the detection image and the sample detection image is less than or equal to a second preset value, and then obtain the trained first model.

[0022] In a possible implementation, inputting the sample image into the segmentation sub-model to obtain a segmentation image output by the segmentation sub-model includes:

[0023] Performing feature extraction processing on the sample image to obtain a feature vector of the sample image;

[0024] Performing semantic segmentation processing on the feature vector according to a first convolutional layer to obtain a semantic segmentation vector of the sample image;

[0025] Performing superpixel segmentation processing on the feature vector according to a second convolutional layer to obtain a superpixel segmentation vector of the sample image;

[0026] Obtaining the segmentation image according to the semantic segmentation vector and the superpixel segmentation vector.

[0027] In a possible implementation, obtaining a candidate detection image of the sample image according to the segmentation image and the candidate result includes:

[0028] If the candidate result indicates that the target object is included in the sample image, determining the segmentation image as the candidate detection image;

[0029] If the candidate result indicates that the target object is not included in the sample image, updating pixel values of pixel points in the foreground region of the segmentation image to the second pixel value to obtain the candidate detection image.

[0030] In a second aspect, the present application provides an image detection method, including:

[0031] Obtaining a first image to be detected;

[0032] Inputting the first image into a first model to obtain a detection image of the first image, where the detection image indicates that the target object is included in the first image and the position of the target object on the first image, or the detection image indicates that the target object is not included in the first image, where the first model is a model trained according to the model training method according to any item in the first aspect.

[0033] In a third aspect, the present application provides a model training device, including:

[0034] An acquisition module, configured to acquire training samples, where the training samples include multiple first sample images, sample detection images of the first sample images, and multiple second sample images, and the sample detection images are images obtained by annotating sample objects on the first sample images;

[0035] A processing module, configured to perform at least one semi-supervised model training on a first model according to the multiple first sample images, the sample detection image of the first sample image, and the multiple second sample images, so as to obtain a trained first model.

[0036] In a possible implementation manner, the processing module is specifically configured to:

[0037] Obtain a trained annotation model according to the multiple first sample images and the sample detection image of the first sample image;

[0038] Input the second sample image into the trained annotation model to obtain the sample detection image of the second sample image;

[0039] Train the first model according to the first sample image, the sample detection image of the first sample image, the second sample image, and the sample detection image of the second sample image, so as to obtain a trained first model.

[0040] In a possible implementation manner, the processing module is specifically configured to:

[0041] For any one of the multiple first sample images, input the first sample image into the annotation model to obtain an annotation detection image output by the annotation model;

[0042] Adjust the parameters of the annotation model according to the annotation detection image and the sample detection image of the first sample image;

[0043] When the difference value between the annotation detection image and the sample detection image of the first sample image is less than or equal to a first preset value, obtain the trained annotation model.

[0044] In a possible implementation manner, the first model includes a segmentation sub-model and an identification sub-model; the processing module is specifically configured to:

[0045] For any sample image among the multiple first sample images and the multiple second sample images, input the sample image into the segmentation sub-model to obtain a segmentation image output by the segmentation sub-model, where the segmentation image includes a foreground region and a background region, the pixel values of the pixel points in the foreground region are first pixel values, and the pixel values of the pixel points in the background region are second pixel values;

[0046] Input the sample image into the identification sub-model to obtain a candidate result output by the identification sub-model, where the candidate result indicates that the sample image includes the target object, or does not include the target object;

[0047] Obtain a candidate detection image of the sample image according to the segmented image and the candidate result;

[0048] Adjust the parameters of the first model according to the difference value between the candidate detection image and the sample detection image of the sample image until the difference value between the detection image and the sample detection image is less than or equal to a second preset value, and then obtain the trained first model.

[0049] In a possible implementation manner, the processing module is specifically configured to:

[0050] Perform feature extraction processing on the sample image to obtain a feature vector of the sample image;

[0051] Perform semantic segmentation processing on the feature vector according to the first convolutional layer to obtain a semantic segmentation vector of the sample image;

[0052] Perform superpixel segmentation processing on the feature vector according to the second convolutional layer to obtain a superpixel segmentation vector of the sample image;

[0053] Obtain the segmented image according to the semantic segmentation vector and the superpixel segmentation vector.

[0054] In a possible implementation manner, the processing module is specifically configured to:

[0055] If the candidate result indicates that the target object is included in the sample image, determine the segmented image as the candidate detection image;

[0056] If the candidate result indicates that the target object is not included in the sample image, update the pixel values of the pixel points in the foreground region of the segmented image to the second pixel value to obtain the candidate detection image.

[0057] In a fourth aspect, the present application provides an image detection device, including:

[0058] An acquisition module, configured to acquire a first image to be detected;

[0059] A detection module, configured to input the first image into a first model to obtain a detection image of the first image, where the detection image indicates that the first image includes a target object and the position of the target object on the first image, or the detection image indicates that the first image does not include the target object, where the first model is a model trained according to the model training method according to any item in the first aspect.

[0060] In a fifth aspect, the present application provides an electronic device, including: at least one processor and a memory;

[0061] The memory stores computer-executable instructions;

[0062] The at least one processor executes the computer-executable instructions stored in the memory, such that the at least one processor executes the model training method according to any one of the first aspect, or such that the at least one processor executes the image detection method according to the second aspect.

[0063] In a sixth aspect, the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the model training method according to any one of the first aspect is implemented, or the image detection method according to the second aspect is implemented.

[0064] The model training and image detection methods, apparatuses, and devices provided by the embodiments of the present application first obtain training samples. The training samples include multiple first sample images, sample detection images of the first sample images, and multiple second sample images. The sample detection images are images obtained by annotating sample objects on the first sample images. Then, at least one semi-supervised model training is performed on the first model according to the multiple first sample images, the sample detection images of the first sample images, and the multiple second sample images to obtain a trained first model. Since the training samples used for training the first model in the embodiments of the present application include labeled training samples composed of the first sample images and the sample detection images of the first sample images, and unlabeled training samples composed of the second sample images, and semi-supervised model training is performed through the labeled training samples and the unlabeled training samples. Without changing the scale of the training samples, since unlabeled training samples are added for training, the number of images to be annotated for the sample images is reduced, and the labor cost of model training is lowered. Description of the Drawings

[0065] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0066] Figure 1 It is a schematic diagram of an application scenario provided by an embodiment of the present application;

[0067] Figure 2 It is a schematic flowchart of the model training method provided by an embodiment of the present application;

[0068] Figure 3 It is a schematic diagram of the model training process provided by an embodiment of the present application;

[0069] Figure 4 Schematic diagram of semi-supervised model training provided by an embodiment of the present application;

[0070] Figure 5 Schematic diagram of the processing flow of the first model provided by an embodiment of the present application;

[0071] Figure 6 Schematic diagram of the network module of the segmentation sub-model provided by an embodiment of the present application;

[0072] Figure 7 Schematic diagram of the processing architecture of the segmentation sub-model provided by an embodiment of the present application;

[0073] Figure 8 Schematic diagram of the fusion of the segmented image and the candidate result provided by an embodiment of the present application;

[0074] Figure 9 Schematic diagram of the process of the image detection method provided by an embodiment of the present application;

[0075] Figure 10 Schematic diagram of the structure of the model training device provided by an embodiment of the present application;

[0076] Figure 11 Schematic diagram of the structure of the image detection device provided by an embodiment of the present application;

[0077] Figure 12 Schematic diagram of the hardware structure of the electronic device provided by an embodiment of the present application. Detailed implementation manners

[0078] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0079] For ease of understanding, first, the concepts involved in the present application will be described.

[0080] Multi-task learning: Multi-task learning is a machine learning method corresponding to single-task learning. In the field of machine learning, the standard algorithm theory is to learn one task at a time, that is, in the case where the system output is a real number. First, a complex learning problem is decomposed into theoretically independent sub-problems, then each sub-problem is learned separately, and finally, a mathematical model of the complex problem is established by combining the learning results of the sub-problems. Multi-task learning is a joint learning method where multiple tasks are learned in parallel and the learning results influence each other. Multi-task learning mainly utilizes the domain-specific information contained in the training data of related tasks to improve the generalization effect of the main task by deriving biases. Multi-task learning includes the simultaneous parallel learning of multiple related tasks and the simultaneous backpropagation of gradients. Multiple tasks help each other learn through the underlying shared representation to improve the generalization effect.

[0081] Semi-supervised learning: Also known as self-supervised learning, it mainly uses auxiliary tasks to mine its own supervision information from large-scale unsupervised data, and trains the network through this constructed supervision information, so as to learn valuable representations for downstream tasks. The biggest advantage of this learning method is that only a small amount of labeled data is required to guide the training of a large amount of unlabeled data, which is very suitable for the characteristics of difficult image data annotation and small data volume.

[0082] Figure 1 The following is a schematic diagram of the application scenario provided by the embodiments of the present application. As Figure 1 shown, it includes a client 11 and a server 12. Among them, the client 11 and the server 12 are connected through a wired or wireless network.

[0083] The client 11 can send an image to be detected 13 to the server 12. The server 12 performs image detection processing on the image to be detected 13 to obtain a detected image 14 of the image to be detected 13, and then obtains the target object on the image to be detected 13 according to the detected image 14, thus realizing the image detection process.

[0084] In some embodiments, the client 11 and the server 12 can be two independent devices. In other embodiments, the functions of the client 11 and the server 12 can also be integrated in the same device. The embodiments of the present application do not make any limitations in this regard.

[0085] The model training and image detection methods provided by the embodiments of this application can be applied to different technical fields and application scenarios, such as cargo detection, target detection, auxiliary pathological judgment, etc. Taking cargo detection as an example, after taking a cargo image of a certain area, the method provided by the embodiments of this application can be used to detect whether a specific cargo is included in the cargo image; taking target vehicle detection as an example, after taking an image on the road, the method provided by the embodiments of this application can be used to detect whether a target vehicle is included in the image; taking auxiliary pathological judgment as an example, after obtaining a pathological image, the method provided by the embodiments of this application can be used to detect whether corresponding lesions are included in the pathological image to help determine the corresponding disease, etc.

[0086] Image detection is usually completed based on a model, and the model needs to be trained with a large number of training samples before application. Generally, the larger the scale of the training samples, the higher the accuracy of the trained model, and as the model structure becomes more and more complex, the scale of the required training samples also increases accordingly. And the annotation of sample images is usually completed manually. The larger the scale of the required training samples, the higher the labor cost of annotation. And in some cases, the annotation needs to be completed by specialized experts. For example, pathological images can only be annotated by experienced doctors, which further increases the labor cost of annotation.

[0087] Based on this, the embodiments of this application provide a model training method to complete the training of the model with only a small number of labeled samples. Next, according to Figure 1 the application scenarios of the examples, combined with Figure 2 the solutions of this application will be introduced.

[0088] Figure 2 is a schematic flowchart of the model training method provided by the embodiments of this application. As Figure 2 shown, the method may include:

[0089] S21, obtaining training samples, where the training samples include multiple first sample images, sample detection images of the first sample images, and multiple second sample images, and the sample detection images are images obtained by annotating sample objects on the first sample images.

[0090] The execution subject of each embodiment in this application can be, for example, a server, a processor, a microprocessor, a chip, or other devices with data processing functions. The specific execution subject of this embodiment is not limited, and it can be selected and set according to actual needs. As long as it is a device with data processing functions, it can be used as the execution subject of each embodiment in this application.

[0091] The first model can be trained with multiple sets of training samples, where the training samples include multiple first sample images, the sample detection images of the first sample images, and multiple second sample images. Among them, a first sample image and the sample detection image of the first sample image can form a set of labeled training samples, and a second sample image can form a set of unlabeled training samples. The multiple sets of training samples include these labeled training samples and unlabeled training samples.

[0092] In the embodiments of the present application, after annotating the sample objects on the first sample image, the sample detection image of the first sample image can be obtained. For example, a possible implementation is to set the pixel values of the pixel points corresponding to the sample objects on the first sample image to the first pixel value, and set the pixel values of the pixel points in other areas except the sample objects on the first sample image to the second pixel value. Optionally, after annotating the first sample image, a series of image enhancement methods such as image flipping, contrast transformation, and grayscale transformation can also be performed on the first sample image to obtain multiple sample images, so as to achieve the effect of expanding the training samples and enhance the robustness of the model.

[0093] The sample detection image also includes the category information of the sample object, and this category information indicates whether the sample object is the target object. For example, it can be set that when the category information is 0, it means that the sample object is not the target object, and when the category information is 1, it means that the sample object is the target object; for example, it can be set that when the category information is 1, it indicates that the sample object is not the target object, and when the category information is 0, it means that the sample object is the target object, and so on.

[0094] S22. Perform at least one semi-supervised model training on the first model according to the multiple first sample images, the sample detection images of the first sample images, and the multiple second sample images to obtain the trained first model.

[0095] After obtaining multiple sets of training samples, since the multiple sets of training samples include multiple first sample images, the sample detection images of the first sample images, and multiple second sample images, that is, the multiple sets of training samples include both labeled training samples and unlabeled training samples, so the first model can be subjected to at least one semi-supervised model training according to these labeled training samples and unlabeled training samples. In the embodiments of the present application, semi-supervised model training refers to a method of performing deep learning model training by simultaneously using labeled training samples and unlabeled training samples.

[0096] After performing at least one semi-supervised model training on the first model based on multiple first sample images, the sample detection images of the first sample images, and multiple second sample images, a trained first model is obtained. The first model has the ability to detect images and can detect whether a target object is included in the image. If a target object is included, the position of the target object in the image can also be obtained according to the first model.

[0097] The model training method provided in the embodiments of the present application first obtains training samples. The training samples include multiple first sample images, the sample detection images of the first sample images, and multiple second sample images. The sample detection images are images obtained by annotating the sample objects on the first sample images. Then, at least one semi-supervised model training is performed on the first model according to the multiple first sample images, the sample detection images of the first sample images, and the multiple second sample images to obtain a trained first model. Since the training samples used in the embodiments of the present application for training the first model include labeled training samples composed of the first sample images and the sample detection images of the first sample images, and unlabeled training samples composed of the second sample images, and semi-supervised model training is performed through the labeled training samples and the unlabeled training samples. Without changing the scale of the training samples, since unlabeled training samples are added for training, the number of images to be annotated for the sample images is reduced, and the labor cost of model training is reduced.

[0098] Based on any of the above embodiments, the solution of the present application will be further introduced in detail below with reference to the accompanying drawings.

[0099] Figure 3 For the model training process schematic diagram provided in the embodiments of the present application, as Figure 3 shown, it includes:

[0100] S31, obtaining a trained annotation model according to multiple first sample images and the sample detection images of the first sample images.

[0101] In the embodiments of the present application, at least one semi-supervised model training can be performed on the first model according to multiple first sample images, the sample detection images of the first sample images, and multiple second sample images. Specifically, first, an annotation model can be trained according to multiple first sample images and the sample detection images of the first sample images. Among them, the structure of the annotation model and the structure of the first model can be the same or different. The structure of the annotation model can adopt common convolutional neural network structures, deep learning network structures, etc. This embodiment does not limit this.

[0102] For any one of the multiple first sample images, the first sample image can be input into the annotation model, and the annotation model processes the first sample image to obtain an annotated detection image output by the annotation model. Then, the parameters of the annotation model are adjusted according to the difference value between the annotated detection image and the sample detection image of the first sample image. After the parameter adjustment, the next first sample image can be input into the annotation model, and the annotation model processes the next first sample image to obtain an annotated detection image output by the annotation model. Similarly, the parameters of the annotation model are still adjusted according to the difference value between the annotated detection image and the sample detection image of the first sample image.

[0103] When the difference value between the annotated detection image and the sample detection image of the first sample image is greater than the first preset value, the above operation steps are repeatedly executed for training the annotation model until the difference value between the annotated detection image and the sample detection image of the first sample image is less than or equal to the first preset value, and then the training process is stopped to obtain a trained annotation model.

[0104] S32. Input the second sample image into the trained annotation model to obtain a sample detection image of the second sample image.

[0105] The annotation model can be used to annotate a sample image and generate a pseudo-label. Specifically, since the second sample image is not annotated, after obtaining the trained annotation model, the second sample image can be input into the trained annotation model, and the trained annotation model processes the second sample image to obtain a sample detection image of the second sample image, and the sample detection image of the second sample image is the pseudo-label of the second sample image. Through the above method, for the unannotated second sample image, the annotation of the second sample image is realized based on the trained annotation model, and there is no need to manually annotate the second sample image, further saving the labor cost of image annotation.

[0106] S33. Train the first model according to the first sample image, the sample detection image of the first sample image, the second sample image, and the sample detection image of the second sample image to obtain a trained first model.

[0107] After obtaining the sample detection image of the second sample, the first model can be trained according to the first sample image, the sample detection image of the first sample image, the second sample image, and the sample detection image of the second sample image to obtain a trained first model.

[0108] Specifically, for any first sample image, the first sample image and the sample detection image of the first sample image can form a set of labeled training samples; for any second sample image, the second sample image and the sample detection image of the second sample image can form a set of labeled training samples. Therefore, multiple sets of labeled training samples can be obtained based on the first sample image, the sample detection image of the first sample image, the second sample image, and the sample detection image of the second sample image, and these multiple sets of labeled training samples can be used for the training of the first model.

[0109] Reference can be made to Figure 4 understand the above semi-supervised model training. Figure 4 FIG. is a schematic diagram of semi-supervised model training provided by an embodiment of the present application. As Figure 4 shown, the training samples include unlabeled training samples (i.e., multiple second sample images) and labeled training samples (i.e., multiple first sample images and the sample detection images of the first sample images).

[0110] Through the labeled training samples, the teacher network (i.e., the annotation model in the present application) can be trained. After the training is completed, the sample detection image of the second sample image can be output through the annotation model. The second sample image and the sample detection image of the second sample image together form a set of labeled training samples, and together with the labeled training samples formed by the first sample image and the sample detection image of the first sample image, the student network (i.e., the first model in the present application) is jointly trained.

[0111] In the above embodiment, a solution for training the first model by means of semi-supervised model training is introduced. Next, the specific processing process of the first model for sample images will be introduced.

[0112] In an embodiment of the present application, the first model includes a segmentation sub-model and an identification sub-model. Next, the specific processing processes of the segmentation sub-model and the identification sub-model will be introduced in conjunction with Figure 5 FIG..

[0113] Figure 5 FIG. is a schematic diagram of the processing flow of the first model provided by an embodiment of the present application. As Figure 5 shown, it includes:

[0114] S51. For any sample image among multiple first sample images and multiple second sample images, input the sample image into the segmentation sub-model to obtain a segmentation image output by the segmentation sub-model. The segmentation image includes a foreground region and a background region. The pixel values of the pixel points in the foreground region are the first pixel values, and the pixel values of the pixel points in the background region are the second pixel values.

[0115] Reference can be made to Figure 6 and Figure 7Understand the processing process of the segmentation sub-model. Figure 6 It is a schematic diagram of the network module of the segmentation sub-model provided by the embodiment of the present application. Figure 7 It is a schematic diagram of the processing architecture of the segmentation sub-model provided by the embodiment of the present application.

[0116] As Figure 6 shown, the segmentation sub-model includes a feature vectorization module, a segmentation module, and a Softmax layer. For any sample image among multiple first sample images and multiple second sample images, first, the feature extraction process is performed on the sample image through the feature vectorization module to obtain the feature vector of the sample image. Then, the feature vector is processed by the segmentation module to obtain the segmentation image. The segmentation module includes a residual convolutional neural network module and an attention mechanism module, which are alternately arranged, and the number can be set as needed and is not limited here. The Softmax layer is used to calculate the difference value between the segmentation image and the sample detection image, calculate the loss function, and is used for subsequent adjustment of model parameters.

[0117] In the embodiment of the present application, a multi-task learning framework is constructed to realize the segmentation and recognition of target objects. As Figure 7 shown, the segmentation sub-model constructs a segmentation framework combined with supervoxels. This segmentation framework is a dual-branch network model structure mainly composed of two-dimensional convolutional layers, which can perform semantic segmentation and supervoxel segmentation simultaneously. These two network model structures are the first convolutional layer and the second convolutional layer. According to the first convolutional layer, semantic segmentation processing is performed on the feature vector to obtain the semantic segmentation vector of the sample image. According to the second convolutional layer, superpixel segmentation processing is performed on the feature vector to obtain the superpixel segmentation vector of the sample image.

[0118] Specifically, the two-branch network model structures share multi-scale convolutional layers to extract corresponding feature information. The fine-grained details captured by multi-directional attention, such as horizontal, vertical, and depth information, are combined with prior indications. Then, a superpixel segmentation iterative algorithm is established to convert the extracted feature information into superpixel data, and the obtained superpixel segmentation data is input into the superpixel layer for pooling operation to reduce the feature vector of the image to one feature vector for each superpixel. By adding a fully connected layer, the superpixel activation function will be mapped to the output space. Finally, the results of the two branches are combined in a pixel-level form to obtain the final segmentation result. In addition, during the process of using the gradient descent algorithm for model optimization iteration, two different loss functions will be constructed according to the pixel-level segmentation task and the superpixel segmentation task to obtain better training effects. By performing spatial modeling from different directions and levels, fine local details are provided to achieve the category positioning of low-level features. Further, an attention mechanism and a residual network module are added to the feature extraction module to further improve the segmentation ability of the segmentation sub-module.

[0119] After obtaining the semantic segmentation vector and the superpixel segmentation vector, a segmented image can be obtained based on the semantic segmentation vector and the superpixel segmentation vector. For example, the semantic segmentation vector and the superpixel segmentation vector can be fused to obtain the segmented image. The fusion can be, for example, vector addition or concatenation, etc.

[0120] S52. Input the sample image into the recognition sub-model to obtain the candidate result output by the recognition sub-model. The candidate result indicates that the target object is included in the sample image, or the target object is not included.

[0121] The recognition sub-model is used to recognize whether the target object is included in the sample image. The target object is the object to be detected. There may be various different objects in the sample image, and only the target object is the object to be detected. By performing recognition processing on the sample image through the recognition sub-model, the candidate result output by the recognition sub-model is obtained, and this candidate result is used to indicate whether the target object is included in the sample image.

[0122] S53. Obtain the candidate detection image of the sample image according to the segmented image and the candidate result.

[0123] After obtaining the segmented image output by the segmentation sub-model and the candidate result output by the recognition sub-model, the candidate detection image of the sample image can be obtained according to the fusion of the segmented image and the candidate result.

[0124] Specifically, when the candidate result indicates that the target object is included in the sample image, directly determine the segmented image as the candidate detection image. The candidate detection image indicates that the target object is included in the sample image, and the position of the foreground region corresponding to the first pixel value on the subsequent detection image is the position of the target object in the sample image. When the candidate result indicates that the target object is not included in the sample image, update the pixel values of the pixel points in the foreground region of the segmented image to the second pixel value to obtain the candidate detection image, that is, the pixel values of the pixel points on the candidate detection image are all the second pixel value, and the candidate detection image indicates that the target object is not included in the sample image.

[0125] Figure 8 It is a schematic diagram of the fusion of the segmented image and the candidate result provided by the embodiment of the present application. As Figure 8 shown, the sample image is respectively input into the segmentation sub-model and the recognition sub-model. The segmentation sub-model processes the sample image to obtain the segmented image, and the recognition sub-model processes the sample image to obtain the candidate result.

[0126] Then, perform fusion processing on the segmented image and the candidate result to obtain the candidate detection image.

[0127] When the difference value between the candidate detection image and the sample detection image of the sample image is less than or equal to the second preset value, a trained first model is obtained.

[0128] After obtaining the candidate detection image, the parameters of the first model are adjusted according to the difference value between the candidate detection image and the sample detection image of the sample image.

[0129] When the difference value between the candidate detection image and the sample detection image of the sample image is greater than the second preset value, the above operation steps are repeatedly executed for training the first model until the difference value between the candidate detection image and the sample detection image of the sample image is less than or equal to the second preset value, and then the training process is stopped to obtain a trained first model.

[0130] In the above embodiment, the training process of the first model is introduced. After the first model is trained, image detection can be performed through the first model. The following is combined with Figure 9 for introduction.

[0131] Figure 9 The flowchart of the image detection method provided by the embodiment of the present application is shown in Figure 9 as shown, and the method may include:

[0132] S91, obtaining a first image to be detected.

[0133] The first image is the image to be detected, and the first image may or may not include a target object.

[0134] S92, inputting the first image into the first model to obtain a detection image of the first image, where the detection image indicates that the first image includes a target object and the position of the target object on the first image, or the detection image indicates that the first image does not include a target object.

[0135] The first model is Figures 2 - 8 the first model trained by the method exemplified in the embodiment. After inputting the first model into the first model, the first model processes the first image to obtain a detection image of the first image, and the detection image indicates that the first image includes a target object and the position of the target object on the first image, or the detection image indicates that the first image does not include a target object.

[0136] Specifically, the first model includes a segmentation sub-model and an identification sub-model. The first image is processed by the segmentation sub-model to obtain a segmented image corresponding to the first image. The pixel points on the segmented image correspond one-to-one with the pixel points on the first image. The segmented image includes a foreground region and a background region. The pixel values of the pixel points in the foreground region are the first pixel values, and the pixel values of the pixel points in the background region are the second pixel values. The foreground region and the background region are distinguished by the pixel values of the pixel points. The foreground region is the object on the first image segmented by the segmentation sub-model.

[0137] The first image is processed by the identification sub-model to obtain a candidate result of the first image. The candidate result indicates that the first image includes a target object, or the first image does not include a target object.

[0138] After obtaining the segmented image and the candidate result, the segmented image and the candidate result can be fused to obtain a segmented image corresponding to the first image. Specifically, if the candidate result indicates that the first image includes a target object, the segmented image is determined as the detection image of the first image, and the object corresponding to the foreground region on the detection image is the target object. If the candidate result indicates that the first image does not include a target object, the pixel values of the pixel points in the foreground region of the segmented image are updated to the second pixel values to obtain a detection image. At this time, the detection image does not include a target object.

[0139] The model training and image detection method provided by the embodiments of the present application first obtains training samples. The training samples include multiple first sample images, the sample detection images of the first sample images, and multiple second sample images. The sample detection image is an image obtained by annotating the sample objects on the first sample image. Then, at least one semi-supervised model training is performed on the first model according to the multiple first sample images, the sample detection images of the first sample images, and the multiple second sample images to obtain a trained first model. Since the training samples used in the embodiments of the present application for training the first model include labeled training samples composed of the first sample images and the sample detection images of the first sample images, and unlabeled training samples composed of the second sample images, and semi-supervised model training is performed through the labeled training samples and the unlabeled training samples. Without changing the scale of the training samples, since unlabeled training samples are added for training, the number of images that need to be annotated for the sample images is reduced, and the labor cost of model training is reduced.

[0140] Figure 10 This is a schematic structural diagram of the model training device provided by the embodiments of the present application, as Figure 10 shown, including:

[0141] An acquisition module 101 for acquiring training samples, where the training samples include multiple first sample images, a sample detection image of the first sample image, and multiple second sample images, and the sample detection image is an image obtained by annotating a sample object on the first sample image;

[0142] A processing module 102 for performing at least one semi-supervised model training on a first model according to the multiple first sample images, the sample detection image of the first sample image, and the multiple second sample images to obtain a trained first model.

[0143] In a possible implementation manner, the processing module 102 is specifically configured to:

[0144] Obtain a trained annotation model according to the multiple first sample images and the sample detection image of the first sample image;

[0145] Input the second sample image into the trained annotation model to obtain a sample detection image of the second sample image;

[0146] Train the first model according to the first sample image, the sample detection image of the first sample image, the second sample image, and the sample detection image of the second sample image to obtain a trained first model.

[0147] In a possible implementation manner, the processing module 102 is specifically configured to:

[0148] For any one of the multiple first sample images, input the first sample image into the annotation model to obtain an annotation detection image output by the annotation model;

[0149] Adjust the parameters of the annotation model according to the annotation detection image and the sample detection image of the first sample image;

[0150] When the difference value between the annotation detection image and the sample detection image of the first sample image is less than or equal to a first preset value, obtain the trained annotation model.

[0151] In a possible implementation manner, the first model includes a segmentation sub-model and an identification sub-model; the processing module 102 is specifically configured to:

[0152] For any sample image among the multiple first sample images and the multiple second sample images, input the sample image into the segmentation sub-model to obtain a segmentation image output by the segmentation sub-model, where the segmentation image includes a foreground region and a background region, the pixel value of the pixel points in the foreground region is a first pixel value, and the pixel value of the pixel points in the background region is a second pixel value;

[0153] Input the sample image into the recognition sub-model to obtain a candidate result output by the recognition sub-model, where the candidate result indicates that the target object is included in the sample image, or the target object is not included in the sample image;

[0154] Obtain a candidate detection image of the sample image according to the segmentation image and the candidate result;

[0155] Adjust the parameters of the first model according to the difference value between the candidate detection image and the sample detection image of the sample image until the difference value between the detection image and the sample detection image is less than or equal to a second preset value, and then obtain the trained first model.

[0156] In a possible implementation manner, the processing module 102 is specifically configured to:

[0157] Perform feature extraction processing on the sample image to obtain a feature vector of the sample image;

[0158] Perform semantic segmentation processing on the feature vector according to a first convolutional layer to obtain a semantic segmentation vector of the sample image;

[0159] Perform superpixel segmentation processing on the feature vector according to a second convolutional layer to obtain a superpixel segmentation vector of the sample image;

[0160] Obtain the segmentation image according to the semantic segmentation vector and the superpixel segmentation vector;

[0161] In a possible implementation manner, the processing module 102 is specifically configured to:

[0162] If the candidate result indicates that the target object is included in the sample image, determine the segmentation image as the candidate detection image;

[0163] If the candidate result indicates that the target object is not included in the sample image, update the pixel values of the pixel points in the foreground region of the segmentation image to the second pixel value to obtain the candidate detection image.

[0164] The device provided in this embodiment can be used to execute the technical solutions of the above method embodiments, and its implementation principles and technical effects are similar, which will not be elaborated here in this embodiment.

[0165] Figure 11 For the structural schematic diagram of the image detection device provided in the embodiments of the present application, as Figure 11 shown, it includes:

[0166] An acquisition module 111, configured to acquire a first image to be detected;

[0167] The detection module 112 is configured to input the first image into a first model to obtain a detection image of the first image, where the detection image indicates that the first image includes a target object and the position of the target object on the first image, or the detection image indicates that the first image does not include the target object. The first model is a model trained according to the model training method described in the foregoing embodiments.

[0168] The device provided in this embodiment can be used to execute the technical solutions of the foregoing method embodiments. The implementation principles and technical effects are similar, and will not be elaborated here.

[0169] Figure 12 It is a schematic hardware structure diagram of an electronic device provided in an embodiment of the present application. As Figure 12 shown, the electronic device in this embodiment includes: a processor 121 and a memory 122; where

[0170] The memory 122 is configured to store computer execution instructions;

[0171] The processor 121 is configured to execute the computer execution instructions stored in the memory to implement each step executed by the model training method or the image detection method in the foregoing embodiments. Specifically, reference can be made to the relevant descriptions in the foregoing method embodiments.

[0172] Optionally, the memory 122 can be either independent or integrated with the processor 121.

[0173] When the memory 122 is independently provided, the electronic device further includes a bus 123 for connecting the memory 122 and the processor 121.

[0174] An embodiment of the present application further provides a computer-readable storage medium, in which computer execution instructions are stored. When the processor executes the computer execution instructions, the model training method or the image detection method executed by the above electronic device is implemented.

[0175] An embodiment of the present application can further provide a computer program product, which can be executed by a processor. When the computer program product is executed, the model training method or the image detection method shown above can be implemented.

[0176] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be indirect couplings or communication connections through some interfaces, devices or modules, and can be in electrical, mechanical or other forms.

[0177] The integrated modules implemented in the form of software function modules as described above can be stored in a computer-readable storage medium. The above software function modules are stored in a storage medium and include several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor (English: processor) to execute some steps of the methods described in various embodiments of the present application.

[0178] It should be understood that the above processor may be a central processing unit (English: Central Processing Unit, abbreviated as: CPU), and may also be other general-purpose processors, digital signal processors (English: Digital Signal Processor, abbreviated as: DSP), application specific integrated circuits (English: Application Specific Integrated Circuit, abbreviated as: ASIC), etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the invention can be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.

[0179] The memory may include high-speed RAM memory, and may also include non-volatile storage NVM, such as at least one disk memory, and may also be a USB flash drive, a mobile hard disk, a read-only memory, a disk or an optical disc, etc.

[0180] The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, the bus in the drawings of the present application is not limited to only one bus or one type of bus.

[0181] The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk or an optical disc. The storage medium can be any available medium accessible by a general-purpose or special-purpose computer.

[0182] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes various media that can store program codes, such as ROM, RAM, magnetic disks, or optical discs.

[0183] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A model training method, characterized in that, Including: Obtain training samples, where the training samples include multiple first sample images, the sample detection images of the first sample images, and multiple second sample images, and the sample detection images are images obtained by annotating sample objects on the first sample images; Perform at least one semi-supervised model training on the first model according to the multiple first sample images, the sample detection images of the first sample images, and the multiple second sample images to obtain a trained first model; Performing at least one semi-supervised model training on the first model according to the multiple first sample images, the sample detection images of the first sample images, and the multiple second sample images to obtain a trained first model includes: Obtain a trained annotation model according to the multiple first sample images and the sample detection images of the first sample images; Input the second sample images into the trained annotation model to obtain the sample detection images of the second sample images; Train the first model according to the first sample images, the sample detection images of the first sample images, the second sample images, and the sample detection images of the second sample images to obtain the trained first model; The first model includes a segmentation sub-model and an identification sub-model; training the first model according to the first sample images, the sample detection images of the first sample images, the second sample images, and the sample detection images of the second sample images to obtain the trained first model includes: For any sample image among the multiple first sample images and the multiple second sample images, input the sample image into the segmentation sub-model to obtain a segmentation image output by the segmentation sub-model, where the segmentation image includes a foreground region and a background region, the pixel values of the pixel points in the foreground region are the first pixel values, and the pixel values of the pixel points in the background region are the second pixel values; Input the sample image into the identification sub-model to obtain a candidate result output by the identification sub-model, where the candidate result indicates that the sample image includes a target object, or does not include the target object; Obtain a candidate detection image of the sample image according to the segmentation image and the candidate result; Adjust the parameters of the first model according to the difference value between the candidate detection image and the sample detection image of the sample image until the difference value between the detection image and the sample detection image is less than or equal to a second preset value, then obtain the trained first model.

2. The method according to claim 1, characterized in that Obtaining a trained annotation model according to the multiple first sample images and the sample detection images of the first sample images includes: For any one of the multiple first sample images, input the first sample image into the annotation model to obtain an annotation detection image output by the annotation model; Adjust the parameters of the annotation model according to the annotation detection image and the sample detection image of the first sample image; When the difference value between the labeled detection image and the sample detection image of the first sample image is less than or equal to a first preset value, the trained labeling model is obtained.

3. The method according to claim 1, characterized in that, Inputting the sample image into the segmentation sub-model to obtain a segmentation image output by the segmentation sub-model, including: Performing feature extraction processing on the sample image to obtain a feature vector of the sample image; Performing semantic segmentation processing on the feature vector according to a first convolutional layer to obtain a semantic segmentation vector of the sample image; Performing superpixel segmentation processing on the feature vector according to a second convolutional layer to obtain a superpixel segmentation vector of the sample image; Obtaining the segmentation image according to the semantic segmentation vector and the superpixel segmentation vector.

4. The method according to claim 1, wherein Obtaining a candidate detection image of the sample image according to the segmentation image and the candidate result, including: If the candidate result indicates that the target object is included in the sample image, determining the segmentation image as the candidate detection image; If the candidate result indicates that the target object is not included in the sample image, updating pixel values of pixel points in the foreground region of the segmentation image to a second pixel value to obtain the candidate detection image.

5. An image detection method, characterized in that, Including: Obtaining a first image to be detected; Inputting the first image into a first model to obtain a detection image of the first image, where the detection image indicates that the target object is included in the first image and a position of the target object on the first image, or the detection image indicates that the target object is not included in the first image, where the first model is a model trained according to the model training method according to any one of claims 1-4.

6. A model training device, characterized in that, Including: An obtaining module, configured to obtain training samples, where the training samples include multiple first sample images, sample detection images of the first sample images, and multiple second sample images, and the sample detection image is an image obtained by labeling a sample object on the first sample image; A processing module, configured to perform at least one semi-supervised model training on a first model according to the multiple first sample images, the sample detection images of the first sample images, and the multiple second sample images to obtain a trained first model; The processing module is configured to: Obtain a trained labeling model according to the multiple first sample images and the sample detection images of the first sample images; Input the second sample image into the trained labeling model to obtain a sample detection image of the second sample image; Train the first model according to the first sample image, the sample detection image of the first sample image, the second sample image, and the sample detection image of the second sample image to obtain the trained first model; The first model includes a segmentation sub-model and an identification sub-model; The processing module is specifically configured to: For any of the multiple first sample images and the multiple second sample images, input the sample image into the segmentation sub-model to obtain a segmented image output by the segmentation sub-model. The segmented image includes a foreground region and a background region. Pixel values of pixel points in the foreground region are first pixel values, and pixel values of pixel points in the background region are second pixel values; Input the sample image into the recognition sub-model to obtain a candidate result output by the recognition sub-model. The candidate result indicates that the sample image includes a target object, or does not include the target object; Obtain a candidate detection image of the sample image according to the segmented image and the candidate result; Adjust parameters of the first model according to a difference value between the candidate detection image and a sample detection image of the sample image until the difference value between the detection image and the sample detection image is less than or equal to a second preset value, and then obtain the trained first model.

7. An image detection device, characterized in that, Comprising: An acquisition module, configured to acquire a first image to be detected; A detection module, configured to input the first image into a first model to obtain a detection image of the first image. The detection image indicates that the first image includes a target object and a position of the target object on the first image, or the detection image indicates that the first image does not include the target object, where the first model is a model trained by the model training method according to any one of claims 1-4.

8. An electronic device, characterized in that, Comprising: At least one processor and a memory; The memory stores computer-executable instructions; The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the model training method according to any one of claims 1-4, or so that the at least one processor executes the image detection method according to claim 5.

Citation Information

Patent Citations

  • Image annotation method and device, image semantic segmentation method and device and model training method and device

    CN112734775A

  • Image annotation method and device, model training method and device, electronic equipment and medium

    CN114266896A