Image processing model training method, image segmentation method and electronic equipment
By training the image processing model, erasing and filling the edge areas of the portrait, and combining it with quality evaluation, the complexity problem of portrait segmentation is solved, and high-precision and natural and beautiful segmentation effects are achieved.
Patent Information
- Application Number
- CN202510560500.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-09-23
AI Technical Summary
Existing image segmentation technology cannot achieve accurate portrait segmentation, which is limited by factors such as the complexity of portrait edges, lighting and background color.
By training the image processing model, the preset areas in the image that are difficult to segment are first erased, and then filled with reference images, and the quality evaluation function is added to achieve joint learning of image segmentation and quality evaluation.
High-precision portrait segmentation is achieved, generating natural, beautiful and easily distinguishable segmented images from the background, improving the segmentation accuracy and quality of the image processing model.
Smart Images

Figure CN120689693A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to an image processing model training method, an image segmentation method, and an electronic device. Background Art
[0002] Portrait segmentation has always been a hot topic. Distinguishing people from backgrounds at the pixel level is a classic image segmentation task with a wide range of applications.
[0003] However, due to the influence of factors such as the complexity of portrait edges, lighting and background color, current image segmentation technology cannot achieve accurate portrait segmentation. Summary of the Invention
[0004] The purpose of the embodiments of the present application is to provide an image processing model training method, an image segmentation method and an electronic device for achieving high-precision image segmentation and ensuring that the edges of the foreground in the segmented image are natural and accurate.
[0005] In order to achieve the above objectives, the embodiments of the present application adopt the following technical solutions: In a first aspect, an embodiment of the present application provides a method for training an image processing model, comprising: Acquire a first image of a first object, acquire an area in the first image corresponding to a preset part of the first object, and erase the area to obtain a second image; Acquiring a reference image corresponding to the preset part; Filling a preset portion of the second image based on the reference image using an image processing model to obtain a third image; Segmenting the third image using the image processing model to obtain a first segmented image, and performing a quality evaluation on the first segmented image to obtain a score of the first segmented image; Parameters of the image processing model are adjusted based on the first segmented image and the score of the first segmented image.
[0006] The training method of the image processing model provided in this embodiment erases the area corresponding to the preset part of the first object in the first image to obtain a second image, and uses the reference image of the preset part to fill the preset part of the second image to obtain a third image, and segments the third image to obtain a first segmented image; at the same time, a quality evaluation function is added to the image processing model, and the quality of the first segmented image is evaluated by the image processing model to obtain the score of the first segmented image; further, based on the first segmented image and the score of the first segmented image, the parameters of the image processing model are adjusted. In this way, the joint learning of image segmentation and quality evaluation by the image processing model is realized. On the one hand, since the edges of the area corresponding to the preset part in the first image are generally more complex, compared with the first segmented image, the image processing model can be used to evaluate the quality of the first segmented image. The areas corresponding to other parts in the first image are not easy to distinguish from the background, while the areas corresponding to the preset parts in the third image are obtained by filling in the reference image of the preset parts, and the quality of the reference image is high. Therefore, the areas corresponding to the preset parts in the third image are more refined, natural and beautiful, and easy to distinguish from the background than the areas corresponding to the preset parts in the first image, so that the areas corresponding to the entire first object in the first image are easy to distinguish from the background, so that the image processing model can better learn how to distinguish the areas corresponding to the preset parts and the background in the image, so as to accurately distinguish and mark the two, and obtain a natural, beautiful, and high-precision segmented image; on the other hand, it can enable the image processing model to learn the quality of the segmented image, and thus tend to generate higher quality segmented images.
[0007] In a second aspect, an embodiment of the present application provides an image segmentation method, comprising: Acquire a first image of a second object, and acquire a region corresponding to a preset portion of the second object in the first image of the second object, and erase the region to obtain a second image of the second object; Through the image processing model, the second image of the second object is filled with a preset part to obtain a third image of the second object, and the third image of the second object is segmented to obtain a first segmented image of the second object.
[0008] The image segmentation method provided in the embodiment of the present application is that, for the first image of the second object, the edge of the area corresponding to the preset part of the second object is generally more complex and difficult to distinguish from the background compared with the areas corresponding to other parts, and the trained image processing model has learned the characteristics of the preset part. The second image of the second object is obtained by first erasing the area corresponding to the preset part in the first image, and then the image processing model uses the learned characteristics of the preset part to fill the preset part of the second image to obtain the third image of the second object, so that the area corresponding to the preset part in the third image is more refined, natural and beautiful and easy to distinguish from the background compared with the area corresponding to the preset part in the first image, so that the area corresponding to the entire first object in the first image is easy to distinguish from the background, providing data support for subsequent image segmentation; further, by segmenting the third image through the image processing model, the areas corresponding to the background and the second object in the third image can be accurately distinguished and marked, and a natural, beautiful and high-precision segmented image can be obtained.
[0009] In a third aspect, an embodiment of the present application provides a training device for an image processing model, comprising: an erasing module, configured to acquire a first image of a first object, acquire an area in the first image corresponding to a preset portion of the first object, and erase the area to obtain a second image; An acquisition module, configured to acquire a reference image corresponding to the preset part; a filling module, configured to fill a preset portion of the second image based on the reference image using an image processing model to obtain a third image; a processing module, configured to segment the third image using the image processing model to obtain a first segmented image, and perform a quality evaluation on the first segmented image to obtain a score of the first segmented image; An adjustment module is used to adjust parameters of the image processing model based on the first segmented image and the score of the first segmented image.
[0010] In a fourth aspect, an embodiment of the present application provides an image segmentation device, comprising: an erasing module, configured to acquire a first image of a second object, acquire a region corresponding to a preset portion of the second object in the first image of the second object, and erase the region to obtain a second image of the second object; a filling module, configured to fill a preset portion of the second image of the second object using an image processing model to obtain a third image of the second object; A processing module is configured to segment the third image of the second object to obtain a first segmented image of the second object.
[0011] In a fifth aspect, an embodiment of the present application provides an electronic device, including: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the image processing model training method provided in the first aspect or the image segmentation method provided in the second aspect.
[0012] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium. When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device can execute the training method of the image processing model provided in the first aspect or the image segmentation method provided in the second aspect.
[0013] In the seventh aspect, an embodiment of the present application provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to enable a computer to execute part or all of the steps in the training method of the image processing model provided in the first aspect or the image segmentation method provided in the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings: Figure 1 A schematic diagram of an implementation environment provided for one embodiment of the present application; Figure 2 A flowchart of a method for training an image processing model provided in accordance with an embodiment of the present application; Figure 3 A flowchart of a method for training an image processing model provided in another embodiment of the present application; Figure 4 A flowchart of an image segmentation method provided in one embodiment of the present application; Figure 5 A flowchart of an image segmentation method provided in another embodiment of the present application; Figure 6 A schematic structural diagram of a training device for an image processing model provided in one embodiment of the present application; Figure 7 A schematic structural diagram of an image segmentation device provided in one embodiment of the present application; Figure 8 A schematic structural diagram of an electronic device provided in accordance with an embodiment of the present application. DETAILED DESCRIPTION
[0015] To make the objectives, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the specific embodiments of this application and the corresponding drawings. It should be understood that the embodiments described are only some of the embodiments of this application, and are not intended to be exhaustive. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of this application without inventive effort are intended to fall within the scope of protection of this application. The terms "first," "second," and so on, used in the specification and claims of this application and in the accompanying drawings, are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that such terms are interchangeable where appropriate and are merely used to distinguish objects with the same properties when describing the embodiments of this application. Furthermore, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus comprising a list of elements is not necessarily limited to those elements but may include other elements not expressly listed or inherent to those elements.
[0016] In order to solve the influence of factors such as the complexity of portrait edges, lighting and background color, and achieve high-precision portrait segmentation, the inventive concept of the image segmentation scheme proposed in the embodiment of the present application is: first erase the area in the image of the object that is difficult to segment, such as the area in the image corresponding to the preset part of the object, and then use the features of the part learned by the image processing model to fill the erased area, so that the image data of the area after filling is more refined, natural and beautiful, and easy to distinguish from the background compared to the image data before filling, so that the areas corresponding to the object and background in the image can be accurately identified, providing data support for subsequent image segmentation; on this basis, the filled image is channel compressed, so that the areas corresponding to the background and object in the image can be accurately distinguished and marked, and a natural, beautiful and high-precision segmented image can be obtained.
[0017] Based on the above-mentioned inventive concept, an embodiment of the present application proposes a training method for an image processing model, which erases the area corresponding to the preset part of the first object in the first image to obtain a second image, and uses the reference image of the preset part to fill the preset part of the second image to obtain a third image, and segment the third image to obtain a first segmented image; at the same time, a quality evaluation function is added to the image processing model, and the quality of the first segmented image is evaluated by the image processing model to obtain the score of the first segmented image; further, based on the first segmented image and the score of the first segmented image, the parameters of the image processing model are adjusted. In this way, the joint learning of image segmentation and quality evaluation by the image processing model is realized. On the one hand, since the edge of the area corresponding to the preset part in the first image is generally relatively large, the image processing model can be used to evaluate the quality of the first segmented image. The first image is complex and difficult to distinguish from the background compared to the areas corresponding to other parts in the first image. The area corresponding to the preset part in the third image is obtained by filling in the reference image of the preset part, and the quality of the reference image is high. Therefore, the area corresponding to the preset part in the third image is more refined, natural and beautiful, and easy to distinguish from the background compared to the area corresponding to the preset part in the first image, so that the area corresponding to the entire first object in the first image is easy to distinguish from the background, so that the image processing model can better learn how to distinguish the area corresponding to the preset part and the background in the image, so as to accurately distinguish and mark the two, and obtain a natural, beautiful and high-precision segmented image; on the other hand, it can enable the image processing model to learn the quality of the segmented image, and thus tend to generate a higher quality segmented image.
[0018] Based on the image processing model trained by the above training method, the embodiment of the present application also proposes an image segmentation method. For the first image of the second object, the edge of the area corresponding to the preset part of the second object is generally more complex and difficult to distinguish from the background compared with the areas corresponding to other parts. The trained image processing model has learned the characteristics of the preset part. By first erasing the area corresponding to the preset part in the first image, a second image of the second object is obtained. Then, the image processing model uses the learned characteristics of the preset part to fill the preset part of the second image to obtain a third image of the second object. The area corresponding to the preset part in the third image is more refined, natural and beautiful than the area corresponding to the preset part in the first image. It is easy to distinguish from the background, so that the area corresponding to the entire first object in the first image is easy to distinguish from the background, providing data support for subsequent image segmentation; further, by segmenting the third image through the image processing model, the areas corresponding to the background and the second object in the third image can be accurately distinguished and marked, and a natural, beautiful and high-precision segmented image can be obtained.
[0019] It should be understood that the image processing model training method and image segmentation method proposed in the embodiments of the present application can be executed by an electronic device. As an example, it can be executed by software in an electronic device. The so-called electronic devices here can include terminal devices, such as smart phones, tablet computers, laptops, desktop computers, intelligent voice interaction devices, smart home appliances, smart watches, vehicle terminals, aircraft, etc.; or, the electronic device can also include a server, such as an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud computing services.
[0020] Before introducing the training method of the image processing model and the image segmentation method provided by the embodiment of the present application in detail, a brief introduction to the implementation environment involved in the embodiment of the present application is first given. Figure 1 , is a schematic diagram of an implementation environment provided by an embodiment of the present application, the implementation environment includes a terminal 10, or the implementation environment includes a terminal 10 and an image processing platform 20. The terminal 10 is connected to the image processing platform 20 via a wireless network or a wired network.
[0021] The terminal 10 may be a terminal device, such as at least one of a smartphone, a tablet computer, a laptop computer, a desktop computer, an intelligent voice interaction device, a smart home appliance, a smart watch, a vehicle-mounted terminal, and an aircraft. The terminal 10 may have installed and run an application that supports image processing. For example, the application may be a system application, an online video application, a conference application, a social application, and the like.
[0022] For example, the terminal 10 can obtain an image of a sample object and train an image processing model based on the image of the sample object. After the training is completed, an image processing model with good segmentation accuracy and robustness is obtained. The trained image processing model can then be used to perform image segmentation on images of various objects to obtain corresponding segmented images. The terminal 10 can complete this task independently or through the image processing platform 20 to provide data services for it. This embodiment of the present application is not limited to this.
[0023] The image processing platform 20 comprises at least one of a single server, multiple servers, a cloud computing platform, and a virtualization center. The image processing platform 20 provides backend services for applications that support image processing. Optionally, the image processing platform 20 performs primary processing, while the terminal 10 performs secondary processing. Alternatively, the image processing platform 20 performs secondary processing, while the terminal 10 performs primary processing. Alternatively, either the image processing platform 20 or the terminal 10 can operate independently.
[0024] Exemplarily, the terminal 10 obtains an image of a sample object and sends the obtained image to the image processing platform 20, which trains the image processing model. After the training, an image processing model with good segmentation accuracy and robustness is obtained. The trained image processing model can be used to perform image segmentation on images of various objects sent by the terminal 10, obtain corresponding segmented images and return them to the terminal 10; alternatively, the image processing platform 20 can also send the trained image processing model to the terminal 10, which uses the trained image processing model to perform image segmentation on images of various objects and obtain corresponding segmented images.
[0025] Based on the implementation environment introduced above, the training method of the image processing model and the image segmentation method provided in the embodiments of the present application are introduced in detail with reference to the accompanying drawings.
[0026] Please refer to Figure 2 , is a flow chart of a method for training an image processing model provided in one embodiment of the present application, the method comprising the following steps: S202 : Acquire a first image of a first object, acquire a region in the first image corresponding to a preset portion of the first object, and erase the region to obtain a second image.
[0027] The first object can be any object, such as a person, a person's head, etc., which is not limited in the embodiments of the present application. The preset part can be a part of the first object that is difficult to distinguish from the background, such as hair, etc., which is not limited in the embodiments of the present application. Depending on the application scenario, the first object can be different, and the preset part can also be different. For example, in scenes such as video calls and meetings, the first object is the upper body of a person. Since the clothes of the person in this scene are difficult to distinguish from the background of the same color, the preset part can include the clothes of the person. For another example, in scenes such as character cutout, the first object is a person. Since the edges of the area of the person's hair in the image are highly complex, it is often difficult to accurately cut out the image. Therefore, the preset part includes the person's hair.
[0028] In the above S202 , the first image of the first object may be acquired through various appropriate methods, which are not limited in this embodiment of the present application.
[0029] In one implementation, an image capturing device is used to capture an image of a first object to obtain a first image of the first object.
[0030] In another implementation, any image of the first object is selected from an image library as the first image of the first object, wherein the image library contains images of different objects.
[0031] In S202 above, the region corresponding to the preset portion of the first object in the first image can be obtained by various appropriate methods, which are not limited in this embodiment of the present application. In one embodiment, the first image is input into a trained image recognition model, and the image recognition model performs image recognition based on image features of the first image to obtain the region corresponding to the preset portion of the first object.
[0032] In another embodiment, key point detection is performed on the first image to obtain first key point data; then, based on the first key point, an area in the first image corresponding to the preset part is determined.
[0033] The first key point data includes key points of the preset part in the first image. The first key point data is used to describe the positional distribution of the preset part of the first object in the first image, specifically including but not limited to the position coordinates of the key points of the preset part in the first image. The first key point data can be obtained by performing key point detection on the first image using various key point detection technologies, which are not limited in the embodiments of the present application. Because the first key point data objectively and accurately reflects the positional distribution of the preset part in the first image, the area corresponding to the preset part in the first image can be accurately located based on the first key point data.
[0034] In the above 202, the area corresponding to the preset part in the first image can be masked by using a masking technology commonly used in the art, so that the pixel values of the pixels in the area are 0, thereby achieving erasure of the area.
[0035] S204: Acquire a reference image corresponding to a preset part.
[0036] The reference image can be any image that contains the preset part and whose image quality (e.g., clarity, angle, etc.) of the area where the preset part is located meets the requirements. The reference image can be manually selected from a library of images of the preset part, or a pre-trained quality assessment model can be used to evaluate the quality of each image in the library and then determine the image in the library that meets the quality requirements as the reference image.
[0037] In practical applications, the reference image and the first image correspond to the same image ID; or, the reference image and the first image correspond to different image IDs.
[0038] Of course, the number of reference images can also be multiple, including reference images with the same ID as the first image, as well as reference images with different IDs. In this way, compared to the method of training only with reference images with the same ID, it can prevent the image processing model from becoming lazy and allow the image processing model to fully learn the characteristics of the preset parts.
[0039] S206 , using an image processing model, filling a preset portion of the second image based on the reference image to obtain a third image.
[0040] The image processing model can obtain the features of the preset part from the reference image, such as texture features, and use the features to fill the area corresponding to the preset part in the first image, which is equivalent to regenerating the image data of the area, so that the image data of the area after filling is more refined, natural and beautiful, and easy to distinguish from the background compared to the image data before filling, thereby enabling the object and background to be accurately identified in the third area of the image respectively, providing data support for subsequent image segmentation.
[0041] In one embodiment, the above S206 includes the following steps: S261: Downsample the features of the second image to obtain a first latent variable.
[0042] The features of the second image can be obtained by extracting features from the second image using an image processing model. Then, the second image is downsampled using the image processing model to extract high-level abstract semantic features and obtain the first latent variable.
[0043] Latent codes are high-dimensional features that contain rich information and are typically used to represent the underlying characteristics or structure of data. They are a compact representation of data that helps the model find the underlying relationships hidden beneath surface features, allowing for more accurate subsequent classification and other tasks. In this embodiment of the present application, the first latent variable contains high-level abstract semantic features of the second image, such as the object's identity (ID), background, brightness, skin color, hair color, etc. This information helps the image processing model more accurately perform subsequent image segmentation tasks.
[0044] For example, Figure 3 As shown, after the second image is input into the image processing model, the feature extraction layer of the image processing model performs feature extraction, and the features of the second image can be obtained. The feature extraction layer can adopt various network structures with feature extraction functions, and the embodiments of the present application are not limited to this. In an optional manner, the feature extraction layer can be MobileNetV3. MobileNetV3 is a lightweight convolutional neural network (CNN) designed for mobile devices and embedded devices, aiming to achieve efficient image recognition and feature extraction under resource-constrained conditions. In this way, the image processing model can be widely deployed on web pages (Web), mobile phones, edge devices, etc., and is suitable for portrait segmentation in scenarios such as video calls and conferences, especially when it is necessary to take into account both segmentation accuracy and speed and portrait segmentation.
[0045] Afterwards, the downsampling of the features of the second image is achieved by stacking multiple convolutional layers and multiple pooling layers, with each convolutional layer followed by a pooling layer ( Figure 3 Not shown in ). By stacking multiple convolutional layers ( Figure 3 (illustrated in the figure using four convolutional layers), the convolution and pooling operations are continuously performed on the features of the second image, gradually reducing the size of the features of the second image while increasing the number of feature channels to capture feature information of different sizes, thereby extracting high-level abstract semantic features and obtaining the first latent variable. Each convolutional layer used in this implementation can be a 3*3 convolutional layer with a stride of 2 and a kernel size of 64. The pooling layer can use either the PreLu or ReLu function for pooling.
[0046] S262: Fusing the features of the region corresponding to the preset part in the reference image with the first latent variable to obtain a second latent variable.
[0047] The features of the region corresponding to the preset part in the reference image can be obtained by extracting features from the region using an image processing model. The features of the region include texture features of the preset part. The texture features reflect the slowly changing or periodically changing surface structural organization and arrangement properties of the surface of the preset part, and can intuitively and accurately reflect the characteristics of the preset part. By fusing the features of the region corresponding to the preset part in the reference image with the first latent variable, the texture features of the preset part are added to the first latent variable. The resulting second latent variable helps generate fine, natural, beautiful, and easily distinguishable image data for the erased area from the background.
[0048] For example, Figure 3 As shown, after the reference image is input into the image processing model, the feature extraction layer of the image processing model is used to extract the features of the reference image, and the features of the reference image can be obtained. Afterwards, the features of the reference image are converted into a feature vector with the same dimension as the first latent variable by a multilayer perceptron (MLP), and the feature vector is added to the first latent variable to achieve effective fusion of the two. Among them, the feature extraction layer can adopt various network structures with feature extraction functions, and the multilayer perceptron can adopt various network structures with dimensionality transformation functions, which are not limited in this embodiment of the present application. In an optional manner, the feature extraction layer can be MobileNetV3, and the multilayer perceptron can include 4 fully connected layers.
[0049] S263: Upsample the second latent variable to obtain a third image.
[0050] By upsampling the second latent variable, the second latent variable can be restored to the size of the second image, and the high-level semantic features implicit in the second latent variable and the supplementary texture features of the preset location are fully utilized to generate high-quality image data for the erased area, resulting in a third image. In this way, the area of the preset location in the third image is more detailed, natural, and beautiful than that in the second image, and is easily distinguishable from the background. This allows the image processing model to better learn how to distinguish between the object and background areas in the image, thereby accurately distinguishing and marking the two, resulting in a natural, beautiful, and high-precision segmented image.
[0051] For example, the upsampling of the second latent variable is achieved by stacking multiple convolutional layers and multiple pooling layers, with each convolutional layer followed by a pooling layer ( Figure 3 Not shown in ). By stacking multiple convolutional layers ( Figure 3 (illustrated by four convolutional layers in the figure), the convolution and pooling operations are continuously performed on the second latent variable, gradually increasing its size while reducing the number of feature channels. High-level semantic features implicit in the second latent variable and supplemented by texture features of the preset location are used to generate high-quality image data for the erased area, resulting in a third image. Furthermore, skip connections are used between the convolutional layers used for upsampling and the convolutional layers used for downsampling. This connects the upsampling convolutional layer to a convolutional layer of the same size used in downsampling to fuse information from different levels. This helps the image processing model restore the size while combining detailed information from the early downsampling phase with abstract information from the later phases, improving the quality of the third image.
[0052] The above describes some implementation methods of the above S206. Of course, it should be understood that the above S206 can also be implemented in other ways, and the present embodiment of the application does not limit this.
[0053] S208 , segmenting the third image using an image processing model to obtain a first segmented image.
[0054] Since the number of channels in the third image may not match the actual segmentation task, by performing channel compression on the third image, the channel data of the third image is compressed to be consistent with the actual segmentation task. This allows us to determine the category to which each pixel in the third image belongs (for example, object or background), and distinguish and label the background and object areas in the third image. This type of image is the first segmented image.
[0055] For example, Figure 3As shown in the figure, after the third image is obtained by stacking 4 convolutional layers, a convolution operation is performed on the third image through a 3*3 convolutional layer to achieve smoothing of the third image to reduce noise and highlight important structures. Finally, a convolution operation is performed on the third image through a 1*1*3 convolutional layer to achieve channel compression of the third image and obtain the first segmented image.
[0056] S210 , performing quality evaluation on the first segmented image using an image processing model to obtain a score of the first segmented image.
[0057] The score of the first segmented image reflects the quality of the first segmented image. By introducing the quality evaluation function of the first segmented image into the image processing model, the image processing model can learn what is a good segmented image and what is a bad segmented image.
[0058] In one embodiment, the above S210 includes the following steps: performing channel compression on the third image to obtain a fourth image; performing a pooling operation based on the fourth image to obtain a first vector; and performing nonlinear mapping on the first vector to obtain a score of the first segmented image.
[0059] Since the quality of the first segmented image depends on the quality of the third image, the quality evaluation of the first segmented image can be converted into the quality evaluation of the third image. Moreover, the quality evaluation task of the third image can be regarded as a regression task. The number of channels required for the regression task is low, for example, 1 channel is required, while the number of channels of the third image is large and regression prediction cannot be achieved. Therefore, after channel compression of the third channel data, it can be used for subsequent regression tasks.
[0060] As an optional method, the third image obtained by the last upsampling convolution layer can be channel compressed through a 1*1*1 convolution layer to obtain a fourth image with a channel number of 1; then, the fourth image is mapped, for example, the fourth image is first converted into a 1*1 vector (i.e., the first vector) through a global maximum pooling layer (Global Max Pooling layer, GLP), and then the sigmoid function is used to map the vector to a score, which is the score of the first segmented image.
[0061] As another optional method, the number of third images is n, and the n third images are obtained by performing n-level upsampling on the second latent variable. The second latent variable is obtained based on the features of the reference image and the features of the second image, and the score of the first segmented image includes the first score and the second score. Accordingly, the above-mentioned channel compression of the third image to obtain the fourth image includes: performing channel compression on the n-1th third image to obtain the fourth image. Furthermore, the above-mentioned nonlinear mapping of the first vector to obtain the score of the first segmented image includes: performing nonlinear mapping on the first vector to obtain the first score; performing channel compression on the product of the n-1th third image and the first score and fusing the nth third image to obtain the fifth image; performing pooling operation on the fifth image to obtain the second vector; and performing nonlinear mapping on the second vector to obtain the second score.
[0062] For example, Figure 3 As shown in the figure, the second latent variable is upsampled to four levels through four stacked convolutional layers. Each level of upsampling is a convolution operation. Each level of upsampling corresponds to a third image, and there are four third images in total. The first third image is obtained by upsampling the second latent variable to the first level, and each subsequent third image is obtained by upsampling the previous third image accordingly.
[0063] Afterwards, for the third third image, a 1*1*1 convolution layer is used to perform channel compression on the third image to obtain a fourth image with a channel number of 1; then, the fourth image is first converted into a 1*1 vector (i.e., the first vector) through GLP, and the sigmoid function is used to map the vector to a score, which is the first score; then, the product of the first score and the third third image and the third third image are added to realize the fusion of the two, and a 1*1*1 convolution layer is used to perform channel compression on the fusion result to obtain a fifth image with a channel number of 1; finally, the fourth image is first converted into a 1*1 vector (i.e., the second vector) through GLP, and the sigmoid function is used to map the vector to a score, which is the second score.
[0064] In practical applications, the first score may be 0, which will result in the information contained in the n-1th third image being unable to be fused into the nth third image, thereby affecting the final image segmentation effect. To this end, a preset decimal a, such as a=0.05, can be added to the first score to obtain a new first score, and then the product of the new first score and the third third image and the third third image are added to achieve effective fusion of the two.
[0065] This second approach allows for a more refined quality assessment of the first segmented image, specifically reflecting the degree of influence of the third image, obtained by the final two levels of upsampling, on the image segmentation result. The first score reflects the influence of the (n-1)th third image on the image segmentation result, while the second score reflects the influence of the nth third image on the image segmentation result. These two scores help better guide the image processing model in generating detailed, natural-looking image data for the erased area that is easily distinguishable from the background, ultimately outputting a beautiful, natural-looking, and highly accurate segmented image.
[0066] The above describes some implementation methods of the above S210. Of course, it should be understood that the above S210 can also be implemented in other ways, and the present embodiment of the application does not limit this.
[0067] S212: Adjust parameters of the image processing model based on the first segmented image and the score of the first segmented image.
[0068] In one embodiment, the above S212 includes the following steps: S2121 : Determine a first loss based on the first segmented image and a reference segmented image corresponding to the second image.
[0069] The reference segmented image refers to a segmented image used as a label, which can be obtained by performing image segmentation on the first image through manual segmentation or other automated methods. For example, the reference segmented image can include one or more of the following segmented images: Category 1: The second segmented image of the reference image.
[0070] The second category: a third segmented image obtained by adding interference information around the area corresponding to the preset part in the second segmented image.
[0071] For example, in the second segmented image, image data of small objects or circles and dots are added around the area corresponding to the preset part to interfere with the area. This helps the image processing model better learn and distinguish the interference information, thereby obtaining a high-precision segmented image.
[0072] The third category: a fourth segmented image obtained by dividing the second segmented image into multiple regions and randomly erasing some regions.
[0073] For example, the second segmented image is divided into 9*9 areas, and 20% to 30% of the areas are randomly erased to obtain a fourth segmented image, so as to prevent the erased area from not including the area where the preset part is located.
[0074] Each type of reference segmented image has a corresponding score, with the second segmented image having the highest score, while the third and fourth segmented images have lower scores. This provides rich information for the image processing model, allowing it to fully learn the quality of different reference segmented images, and thus tend to generate high-quality segmented images.
[0075] The first loss reflects the image segmentation effect of the image processing model. As an example, the first loss can be calculated using the Binary Cross-Entropy (BCE) function, that is, ,in, represents the first loss, represents the image processing model, represents the second image, represents the first segmented image, represents the reference segmentation image.
[0076] S2122: Determine a second loss based on the score of the first segmented image and the score of the reference segmented image.
[0077] The score of the reference segmented image represents the quality of the reference segmented image, which may be obtained through manual annotation or other automated annotation methods, and is not limited in this embodiment of the present application.
[0078] The second loss reflects the quality evaluation accuracy of the image processing model. As an example, the second loss can also be calculated using the BCE function. Specifically, if the score of the first segmented image includes a first score and a second score, the second loss includes: based on the first score (denoted as ) and the second loss determined by the score of the reference segmented image (denoted as ), based on the second score (denoted as ) and the second loss determined by the score of the reference segmented image (denoted as ).
[0079] S2123: Adjust parameters of the image processing model based on the first loss and the second loss.
[0080] For example, based on the weighted sum of the first loss and the second loss, the total loss of the image processing model is obtained, which is recorded as ,in, represents the total loss, Represents the weight of the first loss, which can be set according to actual needs and is not limited in the embodiments of the present application; then, reducing the total loss of the image processing model is taken as the optimization goal, and the back propagation algorithm is used to adjust the parameters of the image processing model.
[0081] Through the above method, the image processing model can better realize the joint learning of image segmentation and quality evaluation. On the one hand, since the area corresponding to the preset part in the third image is more delicate, natural and beautiful and easy to distinguish from the background than the area corresponding to the preset part in the second image, the image processing model can better learn how to distinguish the areas corresponding to the object and the background in the image, so as to accurately distinguish and mark the two, and obtain a natural, beautiful and high-precision segmented image; on the other hand, the image processing model can learn the quality of the segmented image, and thus tend to generate higher quality segmented images.
[0082] In another embodiment, if the score of the first segmented image is less than a preset score, it indicates that the effect of the first segmented image is not good. Then, increasing the score of the first segmented image can be used as an optimization goal, and the parameters of the image processing model can be adjusted through the back propagation algorithm.
[0083] It is worth noting that the above S202 to S212 is only a training process for the image processing model. In actual applications, the image processing model can be trained multiple times, that is, the above S202 to S212 are repeatedly executed multiple times until the preset training stop condition is met. Among them, the training stop condition can be set according to actual needs, such as the number of training times reaches a preset number threshold, or the weighted sum of the first loss and the second loss is less than the loss threshold, or the score of the first segmented image is greater than or equal to the preset score, etc., and the embodiment of the present application does not limit this. In addition, the first image and the reference image in each training process can be randomly extracted from the image set used as training samples.
[0084] The above describes some implementation methods of the above S212. Of course, it should be understood that the above S212 can also be implemented in other ways, and the present embodiment of the application does not limit this.
[0085] The training method of the image processing model provided in the embodiment of the present application erases the area corresponding to the preset part of the first object in the first image to obtain a second image, and uses the reference image of the preset part to fill the preset part of the second image to obtain a third image, and segments the third image to obtain a first segmented image; at the same time, a quality evaluation function is added to the image processing model, and the quality of the first segmented image is evaluated by the image processing model to obtain the score of the first segmented image; further, based on the first segmented image and the score of the first segmented image, the parameters of the image processing model are adjusted. In this way, the joint learning of image segmentation and quality evaluation by the image processing model is realized. On the one hand, since the edges of the area corresponding to the preset part in the first image are generally more complex, the quality of the first segmented image is relatively good. Compared with the areas corresponding to other parts in the first image, it is not easy to distinguish from the background, and the area corresponding to the preset part in the third image is obtained by filling the reference image of the preset part, and the quality of the reference image is high. Therefore, the area corresponding to the preset part in the third image is more refined, natural and beautiful, and easy to distinguish from the background compared with the area corresponding to the preset part in the first image, so that the area corresponding to the entire first object in the first image is easy to distinguish from the background, so that the image processing model can better learn how to distinguish the area corresponding to the preset part and the background in the image, so as to accurately distinguish and mark the two, and obtain a natural, beautiful and high-precision segmented image; on the other hand, it can enable the image processing model to learn the quality of the segmented image, and thus tend to generate a higher quality segmented image.
[0086] Based on the image processing model obtained by the above training method, the embodiment of the present application also provides an image segmentation method. Figure 4 , is a flow chart of an image segmentation method provided in one embodiment of the present application, the method comprising the following steps: S402: Acquire a first image of the second object, and acquire a region in the first image of the second object corresponding to a preset portion of the second object.
[0087] The second object can be any object, such as a person, a person's head, etc., which is not limited in this embodiment of the present application. In the above S402, the first image of the second object can be obtained by various appropriate methods, which is not limited in this embodiment of the present application.
[0088] In one implementation, an image capturing device is used to capture an image of the second object to obtain a first image of the second object.
[0089] In another implementation, any image of the second object is selected from an image library as the first image of the second object, wherein the image library contains images of different objects.
[0090] In S402 above, the region corresponding to the preset portion of the second object in the first image can be obtained by various appropriate methods, which are not limited in this embodiment of the present application. In one embodiment, the first image is input into a trained image recognition model, and the image recognition model performs image recognition based on image features of the first image to obtain the region corresponding to the preset portion of the second object.
[0091] In another embodiment, key point detection is performed on the first image to obtain first key point data; then, based on the first key point, an area in the first image corresponding to the preset part is determined.
[0092] The first key point data includes key points of the preset part in the first image. The first key point data is used to describe the positional distribution of the preset part of the second object in the first image, specifically including but not limited to the position coordinates of the key points of the preset part in the first image. The first key point data can be obtained by performing key point detection on the first image using various key point detection technologies, which are not limited in this embodiment of the present application. Because the first key point data objectively and accurately reflects the positional distribution of the preset part in the first image, the area corresponding to the preset part in the first image can be accurately located based on the first key point data.
[0093] S404: Erasing the area to obtain a second image of the second object.
[0094] Erasing the area corresponding to the preset part in the first image of the second object can be implemented in various appropriate ways, which is not limited in this embodiment of the present application.
[0095] In one embodiment, a masking technique commonly used in the art may be used to mask an area corresponding to a preset portion in the first image of the second object so that the pixel values of the pixels in the area are 0, thereby erasing the area.
[0096] S406: Filling a preset portion of the second image of the second object using an image processing model to obtain a third image of the second object.
[0097] The specific implementation of the above S406 is the same as the above Figure 2 The specific implementation of S206 in the illustrated embodiment is similar and will not be repeated here.
[0098] For example, Figure 5 As shown, after the second image of the second object is input into the image processing model, the feature extraction layer of the image processing model performs feature extraction to obtain the features of the second image. After that, through the stacking of 4 convolutional layers and the pooling layer after each convolutional layer ( Figure 5(not shown in the figure), continuously performing convolution and pooling operations on the features of the second image, gradually reducing the size of the features of the second image, and increasing the number of feature channels to capture feature information of different sizes, thereby extracting high-level abstract semantic features and obtaining the third latent variable.
[0099] Since the image processing model has learned the texture features of the preset parts during the training phase, after obtaining the third latent variable in the application phase, the texture features can be added to the third latent variable to obtain the fourth latent variable; then, through the stacking of 4 convolutional layers and the pooling layer after each convolutional layer ( Figure 5 ), continuously performing convolution and pooling operations on the fourth latent variable, gradually increasing the size of the fourth latent variable while reducing the number of feature channels, and utilizing the high-level semantic features implicit in the fourth latent variable and the texture features of the supplementary preset parts to generate high-quality image data for the erased area, thereby obtaining a third image of the second object.
[0100] S408 : Segment the third image of the second object using the image processing model to obtain a first segmented image of the second object.
[0101] The specific implementation of the above S408 is the same as the above Figure 2 The specific implementation of S208 in the illustrated embodiment is similar and will not be repeated here.
[0102] For example, Figure 5 As shown in the figure, a convolution operation is performed on the third image obtained by the fourth convolution layer through a 3*3 convolution layer to achieve smoothing of the third image to reduce noise and highlight important structures. Finally, a convolution operation is performed on the third image through a 1*1*3 convolution layer to achieve channel compression of the third image and obtain the first segmented image of the second object.
[0103] The image segmentation method provided in the embodiment of the present application is that, for the first image of the second object, the edge of the area corresponding to the preset part of the second object is generally more complex and difficult to distinguish from the background compared with the areas corresponding to other parts, and the trained image processing model has learned the characteristics of the preset part. The second image of the second object is obtained by first erasing the area corresponding to the preset part in the first image, and then the image processing model uses the learned characteristics of the preset part to fill the preset part of the second image to obtain the third image of the second object, so that the area corresponding to the preset part in the third image is more refined, natural and beautiful and easy to distinguish from the background compared with the area corresponding to the preset part in the first image, so that the area corresponding to the entire first object in the first image is easy to distinguish from the background, providing data support for subsequent image segmentation; further, by segmenting the third image through the image processing model, the areas corresponding to the background and the second object in the third image can be accurately distinguished and marked, and a natural, beautiful and high-precision segmented image can be obtained.
[0104] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0105] Based on the same inventive concept, the present application also provides a training device for an image processing model. Figure 6 , which is a structural diagram of a training device 600 for an image processing model provided in an embodiment of the present application, the device 600 includes: an erasing module 610, an acquisition module 620, a filling module 630, a processing module 640 and an adjustment module 650.
[0106] The erasing module 610 is configured to acquire a first image of a first object, acquire an area in the first image corresponding to a preset part of the first object, and erase the area to obtain a second image.
[0107] An acquisition module 620 is configured to acquire a reference image corresponding to the preset part; The filling module 630 is used to fill the preset parts of the second image based on the reference image through an image processing model to obtain a third image;
[0108] The processing module 640 is configured to segment the third image to obtain a first segmented image, and perform quality evaluation on the first segmented image to obtain a score of the first segmented image.
[0109] The adjustment module 650 is configured to adjust parameters of the image processing model based on the first segmented image and the score of the first segmented image.
[0110] In another embodiment, the filling module is used to: Downsampling the features of the second image to obtain a first latent variable; fusing the features of the region corresponding to the preset part in the reference image with the first latent variable to obtain a second latent variable; The second latent variable is upsampled to obtain a third image.
[0111] In another embodiment, the processing module is configured to: performing channel compression on the third image to obtain a fourth image; Performing a pooling operation on the fourth image to obtain a first vector; Nonlinear mapping is performed on the first vector to obtain a score of the first segmented image.
[0112] In another embodiment, the number of third images is n, the n third images are obtained by performing n-level upsampling on the second latent variable, the second latent variable is obtained based on features of the reference image and features of the second image; and the score of the first segmented image includes a first score and a second score. When the processing module performs channel compression on the third image to obtain the fourth image, the processing module performs the following steps: performing channel compression on the (n-1)th third image to obtain the fourth image; In another embodiment, when the processing module performs nonlinear mapping on the first vector to obtain the score of the first segmented image, the processing module performs the following steps: Performing nonlinear mapping on the first vector to obtain the first score; performing channel compression on the product of the (n-1)th third image and the first score and the (n)th third image, and obtaining a fifth image; performing a pooling operation on the fifth image to obtain a second vector; Perform nonlinear mapping on the second vector to obtain the second score.
[0113] In another embodiment, the adjustment module is configured to: determining a first loss based on the first segmented image and a reference segmented image corresponding to the first image; determining a second loss based on the score of the first segmented image and the score of the reference segmented image; Based on the first loss and the second loss, parameters of the image processing model are adjusted.
[0114] In another embodiment, the reference segmented image includes one or more of the following segmented images: a second segmented image of the reference image; a third segmented image obtained by adding interference information to a region around the preset part in the second segmented image; After dividing the second segmented image into a plurality of regions, a fourth segmented image is obtained by randomly erasing some regions.
[0115] Obviously, the training device of the image processing model provided in the embodiment of the present application can be used as Figure 6 The execution body of the training method of the image processing model shown, for example Figure 6 In the training method of the image processing model shown in FIG, step S202 can be performed by Figure 6 The erasing module 610 in the training device of the image processing model shown in FIG. 1 is executed, and step S204 can be performed by Figure 6 The acquisition module 620 in the training device of the image processing model shown in FIG. 1 is executed, and step S206 can be performed by Figure 6 The filling module 630 in the training device of the image processing model shown in FIG. 1 is executed, and steps S208 and S210 can be performed by Figure 6 The processing module 640 in the training device of the image processing model shown in FIG. 1 is executed, and step S212 can be performed by Figure 6 The training apparatus of the image processing model is shown as being performed by the adjustment module 650 .
[0116] According to another embodiment of the present application, Figure 6 The various modules in the training device of the image processing model shown can be individually or completely combined into one or several other modules to form a whole, or one (or more) of the modules can be further divided into multiple functionally smaller modules to form a whole, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application. The above-mentioned units are divided based on logical functions. In actual applications, the functions of a module can also be implemented by multiple modules, or the functions of multiple modules can be implemented by one module. In the embodiments of the present application, the training device of the image processing model can also include other modules. In actual applications, these modules can also be implemented with the assistance of other modules, and can be implemented by the collaboration of multiple modules.
[0117] According to another embodiment of the present application, a general computing device such as a computer including a central processing unit (CPU), a random access memory (RAM), a read-only memory (ROM) and other processing elements and storage elements can be run to execute the following operations: Figure 2 A computer program (including program code) for each step involved in the corresponding method shown in FIG. Figure 6 The image processing model training device shown, and the image processing model training method for implementing the embodiment of the present application. The computer program can be recorded on, for example, a computer-readable storage medium, and transferred to an electronic device through the computer-readable storage medium and run therein.
[0118] Based on the same inventive concept, the present application also provides an image segmentation device. Figure 7 , is a structural diagram of an image segmentation device 700 provided in an embodiment of the present application. The device 700 includes: an erasing module 710, a filling module 720 and a processing module 730.
[0119] The erasing module 710 is configured to acquire a first image of a second object, acquire an area corresponding to a preset portion of the second object in the first image of the second object, and erase the area to obtain a second image of the second object.
[0120] The filling module 720 is configured to fill a preset portion of the second image of the second object using an image processing model to obtain a third image of the second object.
[0121] The processing module 730 is configured to segment the third image of the second object to obtain a first segmented image of the second object.
[0122] Obviously, the image segmentation device provided in the embodiment of the present application can be used as Figure 7 The execution body of the image segmentation method shown is, for example Figure 7 In the image segmentation method shown in FIG, steps S402 and S404 can be performed by Figure 7 The erasing module 710 in the image segmentation apparatus shown in FIG. 1 is executed, and step S406 can be performed by Figure 7 The filling module 720 in the image segmentation apparatus shown in FIG. 1 is executed, and step S4068 can be performed by Figure 7 The processing module 730 in the image segmentation device shown is executed.
[0123] According to another embodiment of the present application, Figure 7The various modules in the image segmentation device shown can be individually or entirely combined into one or several other modules to form a whole, or one (or more) of the modules can be further divided into multiple functionally smaller modules to form a whole, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application. The above-mentioned units are divided based on logical functions. In actual applications, the functions of a module can also be implemented by multiple modules, or the functions of multiple modules can be implemented by one module. In the embodiments of the present application, the image segmentation device may also include other modules. In actual applications, these modules can also be implemented with the assistance of other modules, and can be implemented by the collaboration of multiple modules.
[0124] According to another embodiment of the present application, the system can be executed on a general computing device such as a computer including processing elements such as a CPU, RAM, ROM and storage elements. Figure 2 A computer program (including program code) for each step involved in the corresponding method shown in FIG. Figure 7 The image segmentation device shown in the figure is used to implement the image segmentation method of the embodiment of the present application. The computer program can be recorded on a computer-readable storage medium, for example, and transferred to an electronic device through the computer-readable storage medium and run therein.
[0125] Figure 8 This is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Figure 8 At the hardware level, the electronic device includes a processor and, optionally, an internal bus, a network interface, and memory. The memory may include internal memory, such as high-speed random-access memory (RAM), or non-volatile memory, such as at least one disk drive. Of course, the electronic device may also include other hardware required for its services.
[0126] The processor, network interface, and memory can be interconnected via an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus. The bus can be divided into an address bus, a data bus, a control bus, and the like. For ease of representation, Figure 8 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0127] The memory is used to store programs. Specifically, the program may include program code, which includes computer operating instructions. The memory may include internal memory and non-volatile memory, and provides instructions and data to the processor.
[0128] The processor reads the corresponding computer program from the non-volatile memory into the internal memory and then runs it, forming a training device for the image processing model at the logical level. The processor executes the program stored in the memory and is specifically used to perform the following operations: Acquire a first image of a first object, acquire an area in the first image corresponding to a preset part of the first object, and erase the area to obtain a second image; Acquiring a reference image corresponding to the preset part; Filling a preset portion of the second image based on the reference image using an image processing model to obtain a third image; Segmenting the third image using the image processing model to obtain a first segmented image, and performing a quality evaluation on the first segmented image to obtain a score of the first segmented image; Parameters of the image processing model are adjusted based on the first segmented image and the score of the first segmented image.
[0129] Alternatively, the processor reads the corresponding computer program from the non-volatile memory into the internal memory and then runs it, forming an image segmentation device at the logical level. The processor executes the program stored in the memory and is specifically used to perform the following operations: Acquire a first image of a second object, and acquire a region corresponding to a preset portion of the second object in the first image of the second object, and erase the region to obtain a second image of the second object; Through the image processing model, the second image of the second object is filled with a preset part to obtain a third image of the second object, and the third image of the second object is segmented to obtain a first segmented image of the second object.
[0130] The above application Figure 2 The training device of the image processing model disclosed in the embodiment shown or the above-mentioned application Figure 4The methods performed by the image segmentation device disclosed in the illustrated embodiments can be applied to or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be performed by hardware integrated logic circuits in the processor or by software instructions. The above processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The methods, steps, and logic block diagrams disclosed in the embodiments of this application can be implemented or executed. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above method.
[0131] The electronic device may also perform Figure 2 Method, and implement the training device of the image processing model in Figure 2 、 Figure 3 Alternatively, the electronic device may also perform the functions of the embodiment shown. Figure 4 Method and image segmentation device are implemented in Figure 4 、 Figure 5 The functions of the illustrated embodiment will not be described in detail in the embodiments of the present application.
[0132] Of course, in addition to software implementation, the electronic device of this application does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0133] The embodiment of the present application also provides a computer-readable storage medium, which stores one or more programs, wherein the one or more programs include instructions, which, when executed by an electronic device including multiple application programs, can enable the electronic device to execute Figure 2 The method of the embodiment shown is specifically used to perform the following operations: Acquire a first image of a first object, acquire an area in the first image corresponding to a preset part of the first object, and erase the area to obtain a second image; Acquiring a reference image corresponding to the preset part; Filling a preset portion of the second image based on the reference image using an image processing model to obtain a third image; Segmenting the third image using the image processing model to obtain a first segmented image, and performing a quality evaluation on the first segmented image to obtain a score of the first segmented image; Adjusting parameters of the image processing model based on the first segmented image and the score of the first segmented image Alternatively, when the instruction is executed by an electronic device including multiple applications, the electronic device can execute Figure 4 The method of the embodiment shown is specifically used to perform the following operations: Acquire a first image of a second object, and acquire a region corresponding to a preset portion of the second object in the first image of the second object, and erase the region to obtain a second image of the second object; Through the image processing model, the second image of the second object is filled with a preset part to obtain a third image of the second object, and the third image of the second object is segmented to obtain a first segmented image of the second object.
[0134] An embodiment of the present application also provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to enable a computer to execute part or all of the steps in the training method of the image processing model or the image segmentation method provided in the embodiment of the present application.
[0135] In short, the above description is only a preferred embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
[0136] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0137] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0138] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0139] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the description of the method embodiments.
Claims
1. A training method for an image processing model, characterized in that: include: Acquire a first image of a first object, acquire an area in the first image corresponding to a preset part of the first object, and erase the area to obtain a second image; Acquiring a reference image corresponding to the preset part; Filling a preset portion of the second image based on the reference image using an image processing model to obtain a third image; Segmenting the third image using the image processing model to obtain a first segmented image, and performing a quality evaluation on the first segmented image to obtain a score of the first segmented image; Parameters of the image processing model are adjusted based on the first segmented image and the score of the first segmented image.
2. The method according to claim 1, characterized in that Filling a preset portion of the second image based on the reference image to obtain a third image includes: Downsampling the features of the second image to obtain a first latent variable; fusing the features of the region corresponding to the preset part in the reference image with the first latent variable to obtain a second latent variable; The second latent variable is upsampled to obtain a third image.
3. The method according to claim 1, characterized in that The performing quality evaluation on the first segmented image to obtain a score of the first segmented image includes: performing channel compression on the third image to obtain a fourth image; Performing a pooling operation on the fourth image to obtain a first vector; Nonlinear mapping is performed on the first vector to obtain a score of the first segmented image.
4. The method according to claim 3, characterized in that The number of the third images is n, and the n third images are obtained by performing n-level upsampling on the second latent variable, and the second latent variable is obtained based on the features of the reference image and the features of the second image; The score of the first segmented image includes a first score and a second score; The performing channel compression on the third image to obtain a fourth image includes: Perform channel compression on the (n-1)th third image to obtain a fourth image; The performing nonlinear mapping on the first vector to obtain a score of the first segmented image includes: Performing nonlinear mapping on the first vector to obtain the first score; performing channel compression on the product of the (n-1)th third image and the first score and the (n)th third image, and obtaining a fifth image; performing a pooling operation on the fifth image to obtain a second vector; Perform nonlinear mapping on the second vector to obtain the second score.
5. The method according to claim 1, wherein The adjusting the parameters of the image processing model based on the first segmented image and the score of the first segmented image includes: determining a first loss based on the first segmented image and a reference segmented image corresponding to the first image; determining a second loss based on the score of the first segmented image and the score of the reference segmented image; Based on the first loss and the second loss, parameters of the image processing model are adjusted.
6. The method according to claim 5, characterized in that The reference segmented image includes one or more of the following segmented images: a second segmented image of the reference image; a third segmented image obtained by adding interference information to a region around the preset part in the second segmented image; After dividing the second segmented image into a plurality of regions, a fourth segmented image is obtained by randomly erasing some regions.
7. An image segmentation method, characterized in that: include: Acquire a first image of a second object, and acquire a region corresponding to a preset portion of the second object in the first image of the second object, and erase the region to obtain a second image of the second object; Through the image processing model, the second image of the second object is filled with preset parts to obtain the third image of the second object, and the third image of the second object is segmented to obtain the first segmented image of the second object; the image processing model is trained by the image processing model training method described in any one of claims 1-6.
8. A training device for an image processing model, characterized in that: include: an erasing module, configured to acquire a first image of a first object, acquire an area in the first image corresponding to a preset portion of the first object, and erase the area to obtain a second image; An acquisition module, configured to acquire a reference image corresponding to the preset part; a filling module, configured to fill a preset portion of the second image based on the reference image using an image processing model to obtain a third image; a processing module, configured to segment the third image using the image processing model to obtain a first segmented image, and perform a quality evaluation on the first segmented image to obtain a score of the first segmented image; An adjustment module is used to adjust parameters of the image processing model based on the first segmented image and the score of the first segmented image.
9. An image segmentation device, characterized in that: include: an erasing module, configured to acquire a first image of a second object, acquire a region corresponding to a preset portion of the second object in the first image of the second object, and erase the region to obtain a second image of the second object; a filling module, configured to fill a preset portion of the second image of the second object using an image processing model to obtain a third image of the second object; A processing module is configured to segment the third image of the second object to obtain a first segmented image of the second object.
10. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the image processing model training method according to any one of claims 1 to 6 or the image segmentation method according to claim 7.