Image processing method and device, readable medium and electronic equipment
By determining a second image of a preset size in image processing, feature maps are obtained to reduce model complexity, solving the problem of low efficiency and high computational cost in existing technologies for image matting, and achieving efficient image processing results.
Patent Information
- Application Number
- CN202311477321.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-07
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2043-11-07
AI Technical Summary
Existing image matting techniques based on deep learning models are computationally intensive and inefficient, making it difficult to meet the needs of efficient image processing.
By determining the preset size requirements of the second image, a first feature map and a second feature map are obtained. These feature maps are then used to determine the opacity of edge and non-edge regions in the image, thereby reducing model complexity and improving processing efficiency.
While reducing model complexity, the efficiency of image processing was improved, achieving an optimized effect from coarse segmentation to fine matting, thus improving the quality and efficiency of image processing.
Smart Images

Figure CN119963589B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and more specifically, to an image processing method, apparatus, readable medium, and electronic device. Background Technology
[0002] With the advancement of computer technology, image matting has become increasingly widely used. Image matting refers to extracting the target of interest from an image. An image can be composed of a foreground and a background, while image matting separates the foreground and background of a given image to extract the foreground target.
[0003] In related technologies, image matting can be achieved based on deep learning models, but the computational load of these models is large and the efficiency is low. Summary of the Invention
[0004] This summary section is provided to briefly introduce the concepts, which will be described in detail in the detailed description section below. This summary section is not intended to identify key or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0005] According to a first aspect of the present disclosure, an image processing method is provided, the method comprising:
[0006] A second image is determined based on the first image to be processed, and the image size of the second image meets the preset size requirements;
[0007] A first feature map and a second feature map are obtained based on the second image; wherein, the first feature map is used to determine the opacity of the edge region in the second image, the edge region being the region where the foreground image and the background image meet in the second image, and the second feature map is used to determine the opacity of the non-edge region in the second image;
[0008] Based on the first image, the first feature map, and the second feature map, obtain the target opacity feature map corresponding to the first image;
[0009] The first image is processed based on the target opacity feature map.
[0010] According to a second aspect of the present disclosure, an image processing apparatus is provided, the apparatus comprising:
[0011] The determining module is used to determine a second image based on the first image to be processed, wherein the image size of the second image meets the preset size requirements;
[0012] A first processing module is used to obtain a first feature map and a second feature map based on the second image input; wherein, the first feature map is used to determine the opacity of edge regions in the second image, the edge regions being the regions where the foreground image and the background image meet in the second image, and the second feature map is used to determine the opacity of non-edge regions in the second image;
[0013] The acquisition module is used to acquire a target opacity feature map corresponding to the first image based on the first image, the first feature map, and the second feature map;
[0014] The second processing module is used to process the first image based on the target opacity feature map.
[0015] According to a third aspect of the present disclosure, a computer-readable medium is provided having a computer program stored thereon, which, when executed by a processing device, implements the steps of the method described in the first aspect of the present disclosure.
[0016] According to a fourth aspect of the present disclosure, an electronic device is provided, comprising:
[0017] A storage device on which computer programs are stored;
[0018] A processing device for executing the computer program in the storage device to implement the steps of the method described in the first aspect of this disclosure.
[0019] Using the above technical solution, a second image is determined based on the first image to be processed, and the image size of the second image meets the preset size requirements. A first feature map and a second feature map are obtained based on the second image. The first feature map is used to determine the opacity of edge regions in the second image, where the edge regions are the areas where the foreground and background images meet. The second feature map is used to determine the opacity of non-edge regions in the second image. Based on the first image, the first feature map, and the second feature map, a target opacity feature map corresponding to the first image is obtained. The first image is then processed based on the target opacity feature map. In this way, the core target model inference calculation is performed on the second image that meets the preset size requirements, which can reduce model complexity, improve model efficiency, and thus improve image processing efficiency.
[0020] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0021] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale. In the drawings:
[0022] Figure 1 This is a flowchart illustrating an image processing method according to an embodiment of the present disclosure.
[0023] Figure 2 It is based on Figure 1 The illustrated embodiment shows a flowchart of step S103.
[0024] Figure 3 This is a schematic diagram illustrating an image processing method according to an embodiment of the present disclosure.
[0025] Figure 4 This is a flowchart illustrating a training method for a target model according to an embodiment of the present disclosure.
[0026] Figure 5 This is a schematic diagram illustrating an image processing method according to an embodiment of the present disclosure.
[0027] Figure 6 This is a block diagram of an image processing apparatus according to an embodiment of the present disclosure.
[0028] Figure 7 This is a block diagram of an electronic device according to an embodiment of the present disclosure. Detailed Implementation
[0029] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0030] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0031] The term "comprising" and its variations as used in this disclosure are open-ended, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Relevant definitions for other terms will be given in the description below.
[0032] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0033] It should be noted that the terms "one" and "multiple" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that they should be understood as "one or more" unless explicitly stated in the context. In the description of this disclosure, unless otherwise stated, "multiple" means two or more, and other quantifiers are similar; "at least one," "one or more," or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one 'a' can represent any number of 'a's; as another example, one or more of a, b, and c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple; "and / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone, where A and B can be singular or plural. The character " / " indicates that the objects before and after it are in an "or" relationship. The singular forms "a," "a kind," "an item," "the," and "the" are also intended to include the plural forms unless the context clearly indicates otherwise.
[0034] Although operations or steps are described in a specific order in the accompanying drawings in the embodiments of this disclosure, it should not be construed as requiring these operations or steps to be performed in the specific order or serial order shown, or requiring all of the shown operations or steps to be performed to obtain the desired result. In the embodiments of this disclosure, these operations or steps may be performed serially; they may be performed in parallel; or a portion of these operations or steps may be performed.
[0035] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0036] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0037] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0038] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0039] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0040] Meanwhile, it is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0041] The embodiments disclosed herein can be applied to image processing scenarios, such as image matting, image segmentation, or target extraction from an image. A target image (I) can be formed by fusing the foreground (F) and background (B) based on the foreground opacity (γ). Image matting is the process of separating the foreground and background of a given image, and can be achieved by calculating the foreground opacity.
[0042] It should be noted that the foreground opacity γ can also be represented by α (alpha).
[0043] In some embodiments, end-to-end matting refers to an algorithm system that only requires the original image as input to calculate and output a foreground mask. The definition of the foreground can be categorized into different vertical tasks. For example, in portrait matting, the foreground is defined as the person in the image; in pet matting, the foreground is defined as small animals such as cats and dogs; and in salient matting (subject matting), the foreground is defined as the most central subject in the image.
[0044] Refined image matting requires a certain level of detail in the estimated opacity map, such as the ability to depict very fine hairs and fine lines. Common refined image matting methods place high demands on the model, involve complex calculations, and are relatively inefficient.
[0045] In some embodiments of this disclosure, the target image can be divided into a foreground image and a background image. For example, the foreground image and background image can be fused together to obtain the target image based on the foreground opacity. Conversely, to obtain the foreground image in the target image, the target image can be segmented based on the foreground opacity.
[0046] In some embodiments, the foreground opacity of the target image can be obtained directly.
[0047] In other embodiments, the foreground opacity can be decomposed into a first guiding parameter a and a second guiding parameter b, wherein the first guiding parameter a plays a dominant role in estimating the foreground opacity of edge regions, while the second guiding parameter b plays a dominant role in estimating the foreground opacity of non-edge regions (such as defined foreground, background, and locations with indistinct edges).
[0048] In this way, the first guiding parameter a and the second guiding parameter b can be obtained first, and then the foreground opacity of the target image can be determined based on the linear relationship, according to the first guiding parameter a, the second guiding parameter b and the target image.
[0049] Furthermore, based on local linear modeling, the foreground and background colors of the target image are considered to be smooth, making it less sensitive to operations such as scaling. Therefore, this method still works after image scaling, and the small-sized segmentation estimation map (feature map) will not introduce obvious blurring, jagged edges, or burrs that interfere with the edge quality of the foreground opacity even after magnification.
[0050] Figure 1 This is a flowchart illustrating an image processing method according to an embodiment of the present disclosure. The method can be applied to electronic devices, which may include terminal devices such as smartphones, smart wearable devices, smart speakers, smart tablets, PDAs (Personal Digital Assistants), CPEs (Customer Premise Equipment), personal computers, in-vehicle terminals, etc.; the electronic device may also include a server, such as a local server or a cloud server. Figure 1 As shown, the method may include:
[0051] S101. Determine the second image based on the first image to be processed.
[0052] The second image has a preset size requirement. The first image may or may not have a preset size requirement.
[0053] In some embodiments, if the image size of the first image meets the preset size requirements, the first image can be directly used as the second image.
[0054] In other embodiments, if the image size of the first image does not meet the preset size requirement, the first image can be resized to obtain a second image. The resizing operation can be a size reduction operation or a size enlargement operation.
[0055] In one implementation, the size adjustment can be a size reduction operation, for example, the image size of the first image can be reduced by downsampling to obtain the second image.
[0056] Optionally, the size of the first image can be reduced to a range that meets preset size requirements to obtain the second image. For example, the range that meets the preset size requirements can be a size range that is less than or equal to a first preset size and greater than or equal to a second preset size; the first preset size is greater than the second preset size, and the first and second preset sizes can be pre-set sizes. Alternatively, the range that meets the preset size requirements can also be a size range that is less than or equal to the aforementioned first preset size. Yet another example is that the range that meets the preset size requirements can also be equal to the aforementioned first preset size.
[0057] This allows the image size of the input model to be within a preset size range, which reduces the complexity of model processing.
[0058] In one implementation, the size adjustment can be a size enlargement operation, for example, the image size of the first image can be enlarged by upsampling to obtain the second image.
[0059] Optionally, the size of the first image can be enlarged to meet the above-mentioned preset size requirements to obtain the second image.
[0060] In some embodiments, the image size can be characterized by length and width, and the units of length and width can be meters, decimeters, centimeters, millimeters, or pixels.
[0061] It should be noted that the aforementioned first image can be a picture or a video. The first image can be an image captured in real time by a camera device, or a pre-stored image, or something else entirely. This disclosure does not limit the type of the first image. The first image can be an image captured in real time, a pre-stored image, or an image received from another device. This disclosure does not limit the method of acquiring the first image.
[0062] S102. Obtain the first feature map and the second feature map based on the second image.
[0063] In some embodiments, the second image can be input into a pre-generated target model to obtain a first feature map and a second feature map output by the target model.
[0064] In some embodiments, the target model can be a neural network model, such as a convolutional neural network (CNN) or a recursive neural network (RNN).
[0065] In some embodiments, the target model may include an encoder and a decoder.
[0066] In some embodiments, the image size of the first feature map, the image size of the second feature map, and the image size of the second image are all the same.
[0067] In other embodiments, the image size of the first feature map is the same as that of the second feature map, but the image size of the first feature map may be smaller than that of the second image, and similarly, the image size of the second feature map may also be smaller than that of the second image.
[0068] In some embodiments, the first feature map can be used to determine the opacity of edge regions in the second image, where the edge regions can be the areas where the foreground and background images meet in the second image. The second feature map can be used to determine the opacity of non-edge regions in the second image.
[0069] For example, the first feature map may include a first guiding parameter a corresponding to each pixel of the second image, and the second feature map may include a second guiding parameter b corresponding to each pixel of the second image. For a detailed description of the first guiding parameter a and the second guiding parameter b, please refer to the description of the foregoing embodiments of this disclosure, which will not be repeated here.
[0070] In some embodiments, the first feature map may be a three-channel feature map. For example, the first feature map may include feature maps with three channels: red (R), green (G), and blue (B).
[0071] In some embodiments, the second feature map may be a single-channel feature map.
[0072] S103. Based on the first image, the first feature map, and the second feature map, obtain the target opacity feature map corresponding to the first image.
[0073] For example, the first image, the first feature map, and the second feature map can be fused to obtain the target opacity feature map.
[0074] In this way, the impact of size changes can be reduced by using the first feature map and the second feature map, thereby improving the accuracy of the target opacity feature map.
[0075] S104. Process the first image based on the target opacity feature map.
[0076] For example, the first image can be matted or segmented based on the target opacity feature map to obtain the foreground target in the first image.
[0077] Using the above method, a second image is determined based on the first image to be processed, and the image size of the second image meets the preset size requirements. A first feature map and a second feature map are obtained based on the second image as input. The first feature map is used to determine the opacity of edge regions in the second image, where the edge regions are the areas where the foreground and background images meet. The second feature map is used to determine the opacity of non-edge regions in the second image. Based on the first image, the first feature map, and the second feature map, a target opacity feature map corresponding to the first image is obtained. The first image is then processed based on the target opacity feature map. In this way, the core target model inference calculation is performed on the second image that meets the preset size requirements, which can reduce model complexity, improve model efficiency, and thus improve image processing efficiency.
[0078] Figure 2 It is based on Figure 1 The illustrated embodiment shows a flowchart of step S103. For example... Figure 2 As shown, step S103 may include the following sub-steps:
[0079] S1031. Determine the third feature map based on the first feature map.
[0080] The image size of the third feature map is the same as that of the first image.
[0081] S1032. Determine the fourth feature map based on the second feature map.
[0082] The fourth feature map has the same image size as the first image.
[0083] In some embodiments, if the image size of the first image meets the preset size requirement, and the first image is directly used as the second image in the aforementioned step S101, then in step S1031, the first feature image can be directly used as the third feature image, and in step S1032, the second feature image can be directly used as the fourth feature image.
[0084] In other embodiments, the first feature map can be resized to obtain a third feature map. The second feature map can also be resized to obtain a fourth feature map.
[0085] Optionally, the size adjustment operation can be a size reduction operation (e.g., downsampling) or a size enlargement operation (e.g., upsampling).
[0086] S1033. Perform feature operations based on the first image, the third feature map, and the fourth feature map to obtain the initial opacity feature map corresponding to the first image.
[0087] In some embodiments, the initial opacity feature map can be obtained through multiplication and addition operations.
[0088] For example, the third feature map and the first image can be multiplied first to obtain the fifth feature map; then the fifth feature map and the fourth feature map can be added together to obtain the initial opacity feature map.
[0089] S1034. Activate the initial opacity feature map to obtain the target opacity feature map.
[0090] This target opacity feature map can also be called the target foreground opacity feature map.
[0091] In some embodiments, an activation function can be used to activate the initial opacity feature map to obtain the target opacity feature map. This activation function may include the sigmoid function and / or the sigmoid midline linear transformation function.
[0092] For example, the activation function may include the following formula (1):
[0093]
[0094] Where Y represents the target opacity feature map, X represents the initial opacity feature map, and Sigmoid is the activation function.
[0095] It should be noted that the image matting task can be considered a segmentation task in terms of region, following the pixel classification task, i.e., foreground and background classification. It is often paired with a cross-entropy type loss function for supervision, such as using a class-balanced cross-entropy loss function in combination with the sigmoid activation function in single-channel mode. This can effectively avoid the gradient vanishing problem when the opacity estimate is 0 or 1, and improve the integrity and accuracy of the segmented region. However, for edge positions such as hair, it is mainly a regression problem, which requires a more accurate estimated opacity value. It is expected that the original image multiply-accumulate guided calculation will have a linear relationship. Therefore, the activation function after the final guided multiply-accumulate operation is designed with a linear modification of the middle segment of the sigmoid. The linear modification of the middle segment of the sigmoid is achieved through the above formula (1), thereby improving the accuracy of the output target opacity feature map.
[0096] In some embodiments, different activation functions can be used without distinguishing between different scenarios. For example, in scenarios with complex backgrounds and higher priority for regional accuracy, a general sigmoid function can be used; for another example, in simple scenarios with higher priority for hair quality, the sigmoid mid-segment linear modification of the above formula (1) can be used.
[0097] In this way, different activation functions can be adapted to different scenarios, further improving the accuracy of the output target opacity feature map.
[0098] Figure 3 This is a schematic diagram illustrating an image processing method according to an embodiment of the present disclosure. Figure 3 As shown, the input to the target model 311 is a second image 302, which is obtained based on the first image 301 to be processed. For example, if the first image 301 is a large-size or high-resolution image, the first image 301 can be reduced in size to obtain a second image 302 that meets the preset size requirements. In this embodiment, the image processing method may include:
[0099] First, the second image 302 is input into the target model 311 to obtain the output feature 321 of the target model 311. The output feature 321 may include a first feature map and a second feature map. The image size of the first feature map, the image size of the second feature map, and the image size of the second image are all the same. The first feature map may include a first guiding parameter a corresponding to each pixel of the second image, and the second feature map may include a second guiding parameter b corresponding to each pixel of the second image.
[0100] Secondly, after sampling the output feature 321 into the sampling module, a sampling output 322 can be obtained. This sampling output 322 can include a third feature map and a fourth feature map, both with the same image size as the first image. The third feature map can be obtained by sampling (upsampling or downsampling) the first feature map, and it can include a third guiding parameter A corresponding to each pixel of the first image. The fourth feature map can be obtained by sampling (upsampling or downsampling) the second feature map, and it can include a fourth guiding parameter B corresponding to each pixel of the first image.
[0101] Next, after performing a multiplication and addition operation on the sampled feature 322 and the first image 302, the initial opacity feature map corresponding to the first image is obtained. For example, the third feature map and the first image can be multiplied first to obtain the fifth feature map; then the fifth feature map and the fourth feature map are added together to obtain the initial opacity feature map.
[0102] Finally, the initial opacity feature map is activated using activation function 313 to obtain the target opacity feature map 323. This activation function 313 may include the sigmoid function and / or a linear transformation function of the sigmoid mid-segment.
[0103] In this way, the core target model inference calculation can be performed on a second image that meets the preset size requirements. This allows for the optimization of the image from coarse segmentation to fine matting without increasing computational costs, thus improving the efficiency of image processing. Furthermore, by calculating linear guiding parameters across the entire network based on the first and second feature maps, the "jagged" problem caused by edges and background colors being close together is reduced when optimizing the matting quality of hair, further improving the overall image processing quality.
[0104] Figure 4 This is a flowchart illustrating a training method for a target model according to an embodiment of this disclosure. Figure 4 As shown, the training method may include:
[0105] S401. Obtain the first sample set and the second sample set.
[0106] The first sample set includes multiple coarse-grained types of first sample data, and the second sample set includes multiple fine-grained types of second sample data. The annotation accuracy of the fine-grained types is higher than that of the coarse-grained types.
[0107] Optionally, fine-grained annotation can be a detailed annotation of the edge regions of the foreground and background images. For example, it can be a hair-level annotation, which can finely annotate the edge hair of the foreground image. Coarse-grained annotation can be a contour-level annotation. For example, it can only annotate the edge contour without implementing hair-level annotation.
[0108] In some embodiments, the first sample data of the coarse-grained type may also be referred to as coarse standard data, and the second sample data of the fine-grained type may also be referred to as finely labeled data.
[0109] S402. Pre-train the second model based on the first sample set to obtain the first model.
[0110] In some embodiments, the second model may be a neural network model, such as a CNN or an RNN.
[0111] In some embodiments, this pre-training can also be called segmentation pre-training. In image matting tasks, it is necessary to ensure the accuracy of the segmented regions to guarantee the overall quality of the matting. Therefore, pre-training can be performed on small-sized sample images using a first sample set (coarsely labeled data).
[0112] In one implementation, the pre-training process may use only a single-channel output, such as outputting only the second feature map, or outputting any one of the three RGB channels of the first feature map.
[0113] In another implementation, multi-channel output can also be used during the pre-training process, such as outputting a first feature map and a second feature map.
[0114] Pre-training helps improve the accuracy of the entire model in estimating the foreground region in vertical matting tasks.
[0115] In some embodiments, step S402 can be omitted. For example, an untrained second model can be used as the first model. Another example is that a pre-defined neural network model can be used as the first model.
[0116] S403. Train the first model based on the first sample set and the second sample set to obtain the target model.
[0117] In some embodiments, the first model can be trained by repeatedly executing the model training steps until it is determined that the trained first model meets the preset stopping iteration condition, and then the trained first model is used as the target model.
[0118] Optionally, the preset stopping iteration condition may include the target loss value being less than or equal to a preset loss threshold, or the change in the target loss value within a certain number of iterations being less than a preset change threshold, or the number of iterations being greater than or equal to a preset number, or it may be a stopping iteration condition commonly used in related technologies, which is not limited in this embodiment. The aforementioned preset loss threshold or preset change threshold can both be pre-set values.
[0119] In addition, if the first trained model is determined to meet the preset stopping iteration condition based on the target loss value, the training step of the model can be stopped.
[0120] In some embodiments, the model training step may include:
[0121] S11. Perform N rounds of iterative training on the first model based on the first loss function set and the first sample set to obtain the iterative first model, where N is a positive integer.
[0122] S12. Perform M rounds of iterative training on the first model after iteration based on the second loss function set and the second sample set to obtain the first model after training. M is a positive integer, and M is less than or equal to N.
[0123] In some embodiments, M can be equal to 1, and N can be a positive integer greater than 1, such as 5 or 10. In this way, alternating between N training epochs with a coarse-grained first sample set and 1 training epoch with a fine-grained second sample set during the training process can improve training efficiency and the accuracy of the trained model.
[0124] In some embodiments of this disclosure, the target sample data may include a first target sample image, a second target sample image, and target annotation information. The second target sample image may be an image that meets preset size requirements, determined based on the first target sample image. The image sizes of the first and second target sample images may be different or the same. The second target sample image serves as input to the first model. The target sample data may be either the first sample data or the second sample data.
[0125] The input to the first model can be a target second sample image that meets the preset size requirements. The output of the first model can include a first feature map and a second feature map of the sample. During training, the sample opacity feature map corresponding to the target first sample image can be obtained based on the target first sample image, the first feature map of the sample, and the second feature map of the sample.
[0126] In some embodiments, the first set of loss functions described above can be used to constrain the loss of at least one of the sample first feature map, sample second feature map, and sample opacity feature map.
[0127] For example, the first set of loss functions may include at least one of a first loss function, a second loss function, and a third loss function, wherein:
[0128] The first loss function can be used to determine the first loss value of the entire graph of the target sample data. The first loss value is used to characterize the difference between the labeled opacity and the predicted opacity of the target sample data in the entire graph. The predicted opacity is the opacity predicted by the first model for the target sample data. The target sample data is either the first sample data or the second sample data.
[0129] Optionally, the first loss function can be a class-balanced cross-entropy loss function, a weighted balanced cross-entropy loss function, a loss function based on overlap, a cross-union ratio loss function, or a region mutual information loss function, etc. The first loss function can be used for full-graph equal-weight supervision.
[0130] The second loss function is used to determine the second loss value of the target sample data in the edge region. The second loss value is used to characterize the difference between the labeled first feature map and the predicted first feature map corresponding to the target sample data in the edge region. The predicted first feature map is the first feature map predicted and output by the first model for the target sample data.
[0131] Optionally, the second loss function can supervise the edge regions, for example, it can supervise only the edge regularization term.
[0132] In some embodiments, the weight of the loss value of the first guiding parameter corresponding to the pixel in the non-edge region of the first feature map can be reduced, for example, reduced to N times the original loss value, where N is a decimal less than 1, such as 0.1, 0.01 or 0.001. In this way, the weight of the non-edge region can be reduced, thereby achieving supervision of the edge region.
[0133] Optionally, the aforementioned third loss function is used to determine the third loss value of the full image of the target sample data. The third loss value is used to characterize the difference between the labeled second feature map and the predicted second feature map corresponding to the target sample data in the full image. The predicted second feature map is the second feature map predicted and output by the first model for the target sample data.
[0134] The second feature map serves as the channel output for the dominant segmentation tendency (pixel classification). The third loss function used to supervise the second feature map can be similar to the first loss function. For example, loss functions such as class-balanced cross-entropy loss function and region mutual information loss function can be used for supervision.
[0135] In one implementation, for sample data with binary labels, only the class-balanced cross-entropy loss function can be used for supervision, avoiding the jagged effect introduced by the region mutual information loss function in learning the marginal distribution.
[0136] In some embodiments, the second set of loss functions may include the first loss function, the second loss function, the third loss function, and the fourth loss function, wherein:
[0137] The fourth loss function can be used to determine the fourth loss value of the edge region of the target sample data. The fourth loss value is used to characterize the difference between the labeled opacity and the predicted opacity of the target sample data in the edge region. The predicted opacity is the opacity predicted by the first model for the target sample data.
[0138] Optionally, the edge region can be subject to weighted loss supervision through this fourth loss function. For example, the edge region can be obtained by erosion and dilation based on the label, achieving a weight of K times, where K can be a positive integer greater than 1, such as 10.
[0139] Optionally, the fourth loss function can be: L1 loss or Laplacian loss. The Laplacian loss can be used to enhance hair texture.
[0140] Thus, when performing M rounds of iterative training on the first model after iteration based on the second sample set, a fourth loss function can be added on the basis of the first loss function set to perform edge-weighted loss supervision on opacity, thereby further improving the accuracy of target opacity in edge regions.
[0141] Figure 5 This is a schematic diagram illustrating an image processing effect according to an embodiment of the present disclosure. For example... Figure 5 As shown, based on steps S101 to S103 of the foregoing embodiments of this disclosure, after processing the first image 501, a target opacity feature map 511 can be obtained; similarly, based on steps S101 to S103 of the foregoing embodiments of this disclosure, after processing the first image 502, a target opacity feature map 512 can be obtained. It can be seen that this target opacity feature map can achieve hair-level segmentation.
[0142] Figure 6 This is a block diagram of an image processing apparatus 1100 according to an embodiment of the present disclosure, such as... Figure 6 As shown, the device 1100 may include:
[0143] The determining module 1101 is used to determine a second image based on the first image to be processed, wherein the image size of the second image meets the preset size requirements;
[0144] The first processing module 1102 is used to obtain a first feature map and a second feature map based on the second image; wherein, the first feature map is used to determine the opacity of the edge region in the second image, the edge region being the region where the foreground image and the background image meet in the second image, and the second feature map is used to determine the opacity of the non-edge region in the second image;
[0145] The acquisition module 1103 is used to acquire a target opacity feature map corresponding to the first image based on the first image, the first feature map, and the second feature map;
[0146] The second processing module 1104 is used to process the first image according to the target opacity feature map.
[0147] In some embodiments, the acquisition module 1103 is configured to: determine a third feature map based on the first feature map, wherein the image size of the third feature map is the same as that of the first image; determine a fourth feature map based on the second feature map, wherein the image size of the fourth feature map is the same as that of the first image; perform feature operations based on the first image, the third feature map, and the fourth feature map to obtain an initial opacity feature map corresponding to the first image; and perform activation processing on the initial opacity feature map to obtain the target opacity feature map.
[0148] In some embodiments, the acquisition module 1103 is used to perform a dot product operation on the third feature map and the first image to obtain a fifth feature map; and to perform an addition operation on the fifth feature map and the fourth feature map to obtain the initial opacity feature map.
[0149] In some embodiments, the activation function includes the sigmoid function and / or the sigmoid midline linear transformation function.
[0150] In some embodiments, the first processing module 1102 is used to input the second image into a pre-generated target model to obtain a first feature map and a second feature map output by the target model; wherein the image size of the first feature map, the image size of the second feature map, and the image size of the second image are the same.
[0151] In some embodiments, the apparatus further includes:
[0152] The model generation module 1105 is used to acquire a first sample set and a second sample set; wherein the first sample set includes multiple coarse-grained first sample data, and the second sample set includes multiple fine-grained second sample data, wherein the annotation accuracy of the fine-grained type is higher than that of the coarse-grained type; the first model is trained based on the first sample set and the second sample set to obtain the target model.
[0153] In some embodiments, the model generation module 1105 is used to repeatedly execute the model training steps until it is determined that the first trained model meets the preset stopping iteration condition, and then the first trained model is used as the target model.
[0154] The model training steps include:
[0155] The first model is trained through N rounds of iterative training based on the first loss function set and the first sample set to obtain the iterative first model, where N is a positive integer.
[0156] The first model after iteration is trained by performing M rounds of iterative training based on the second loss function set and the second sample set, where M is a positive integer and M is less than or equal to N.
[0157] In some embodiments, the first set of loss functions includes a first loss function, a second loss function, and a third loss function, wherein:
[0158] The first loss function is used to determine the first loss value of the entire image of the target sample data. The first loss value is used to characterize the difference between the label opacity and the predicted opacity of the target sample data in the entire image. The predicted opacity is the opacity predicted by the first model for the target sample data. The target sample data is either the first sample data or the second sample data.
[0159] The second loss function is used to determine the second loss value of the target sample data in the edge region. The second loss value is used to characterize the difference between the labeled first feature map and the predicted first feature map corresponding to the target sample data in the edge region. The predicted first feature map is the first feature map predicted and output by the first model for the target sample data.
[0160] The third loss function is used to determine the third loss value of the full image of the target sample data. The third loss value is used to characterize the difference between the labeled second feature map and the predicted second feature map corresponding to the target sample data in the full image. The predicted second feature map is the second feature map predicted and output by the first model for the target sample data.
[0161] In some embodiments, the second set of loss functions includes the first loss function, the second loss function, the third loss function, and the fourth loss function, wherein:
[0162] The fourth loss function is used to determine the fourth loss value of the edge region of the target sample data. The fourth loss value is used to characterize the difference between the labeled opacity and the predicted opacity of the target sample data in the edge region. The predicted opacity is the opacity predicted by the first model for the target sample data.
[0163] In some embodiments, the model generation module 1105 is further configured to pre-train the second model based on the first sample set to obtain the first model.
[0164] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0165] The following is for reference. Figure 7This document illustrates a structural diagram of an electronic device 2000 (e.g., a terminal device or a server) suitable for implementing embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. The server in the embodiments of the present disclosure may include, but is not limited to, local servers, cloud servers, single servers, and distributed servers. Figure 7 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0166] like Figure 7 As shown, electronic device 2000 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 2001, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 2002 or a program loaded from storage device 2008 into random access memory (RAM) 2003. RAM 2003 also stores various programs and data required for the operation of electronic device 2000. Processing device 2001, ROM 2002, and RAM 2003 are interconnected via bus 2004. Input / output (I / O) interface 2005 is also connected to bus 2004.
[0167] Typically, the following devices can be connected to the input / output interface 2005: input devices 2006 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 2007 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 2008 including, for example, magnetic tapes, hard disks, etc.; and communication devices 2009. Communication device 2009 allows electronic device 2000 to communicate wirelessly or wiredly with other devices to exchange data. Although... Figure 7 An electronic device 2000 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0168] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 2009, or installed from storage device 2008, or installed from ROM 2002. When the computer program is executed by processing device 2001, it performs the functions defined in the methods of embodiments of this disclosure.
[0169] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0170] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0171] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0172] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: determine a second image based on a first image to be processed, wherein the image size of the second image meets a preset size requirement; obtain a first feature map and a second feature map based on the second image; wherein the first feature map is used to determine the opacity of an edge region in the second image, the edge region being the region where the foreground image and the background image meet in the second image, and the second feature map is used to determine the opacity of a non-edge region in the second image; obtain a target opacity feature map corresponding to the first image based on the first image, the first feature map, and the second feature map; and process the first image based on the target opacity feature map.
[0173] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0174] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0175] The modules described in the embodiments of this disclosure can be implemented in software or in hardware. The names of the modules are not necessarily limiting in certain circumstances; for example, a determining module can also be described as "determining a second image based on a first image to be processed".
[0176] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0177] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0178] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0179] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0180] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative forms of implementing the claims. Regarding the apparatus in the above embodiments, the specific manner in which the various modules perform their operations has been described in detail in the embodiments relating to the method, and will not be elaborated upon here.
Claims
1. An image processing method, characterized by, The method comprises: determining a second image according to a first image to be processed, the image size of the second image meeting preset size requirements; obtaining a first feature map and a second feature map based on the second image; wherein the first feature map is used to determine the opacity of an edge region in the second image, the edge region being a region where a foreground image and a background image in the second image meet, and the second feature map is used to determine the opacity of a non-edge region in the second image; obtaining a target opacity feature map corresponding to the first image according to the first image, the first feature map and the second feature map; processing the first image according to the target opacity feature map.
2. The method of claim 1, wherein, The obtaining of the target opacity feature map corresponding to the first image according to the first image, the first feature map and the second feature map comprises: determining a third feature map according to the first feature map, the image size of the third feature map being the same as that of the first image; determining a fourth feature map according to the second feature map, the image size of the fourth feature map being the same as that of the first image; performing feature operation on the first image, the third feature map and the fourth feature map to obtain an initial opacity feature map corresponding to the first image; performing activation processing on the initial opacity feature map to obtain the target opacity feature map.
3. The method of claim 2, wherein, The performing of the feature operation on the first image, the third feature map and the fourth feature map to obtain the initial opacity feature map corresponding to the first image comprises: performing point multiplication operation on the third feature map and the first image to obtain a fifth feature map; performing addition operation on the fifth feature map and the fourth feature map to obtain the initial opacity feature map.
4. The method according to any one of claims 1 to 3, characterized in that, The obtaining of the first feature map and the second feature map based on the second image comprises: inputting the second image into a target model generated in advance to obtain the first feature map and the second feature map output by the target model; wherein the image size of the first feature map, the image size of the second feature map and the image size of the second image are the same.
5. The method of claim 4, wherein, The target model is a model generated based on the following manner: obtaining a first sample set and a second sample set; wherein the first sample set comprises a plurality of first sample data of a coarse-grained type, and the second sample set comprises a plurality of second sample data of a fine-grained type, the annotation accuracy of the fine-grained type being higher than that of the coarse-grained type; training a first model according to the first sample set and the second sample set to obtain the target model.
6. The method of claim 5, wherein, The training of the first model according to the first sample set and the second sample set to obtain the target model comprises: recursively performing a model training step until it is determined that the trained first model meets preset stopping iteration conditions, and then taking the trained first model as the target model; the model training step comprises: performing N rounds of iterative training on the first model according to a first loss function set and the first sample set to obtain an iterated first model, N being a positive integer; Perform M rounds of iterative training on the first model after iteration according to the second loss function set and the second sample set, to obtain a trained first model, M is a positive integer, and M is less than or equal to N.
7. The method of claim 6, wherein, The first loss function set includes a first loss function, a second loss function, and a third loss function, wherein: The first loss function is used to determine a first loss value of a full graph of target sample data, the first loss value is used to represent the difference between the annotated opacity and the predicted opacity corresponding to the target sample data in the full graph, the predicted opacity is the opacity predicted and output by the first model for the target sample data, and the target sample data is the first sample data or the second sample data; The second loss function is used to determine a second loss value of the target sample data in the edge region, and the second loss value is used to represent the difference between the annotated first feature map and the predicted first feature map corresponding to the target sample data in the edge region, and the predicted first feature map is the first feature map predicted and output by the first model for the target sample data; The third loss function is used to determine a third loss value of a full graph of the target sample data, and the third loss value is used to represent the difference between the annotated second feature map and the predicted second feature map corresponding to the target sample data in the full graph, and the predicted second feature map is the second feature map predicted and output by the first model for the target sample data.
8. The method of claim 7, wherein, The second loss function set includes the first loss function, the second loss function, the third loss function, and a fourth loss function, wherein: The fourth loss function is used to determine a fourth loss value of the edge region of the target sample data, and the fourth loss value is used to represent the difference between the annotated opacity and the predicted opacity corresponding to the target sample data in the edge region, and the predicted opacity is the opacity predicted and output by the first model for the target sample data.
9. The method of claim 5, wherein, The method further comprises: Pre-training the second model according to the first sample set to obtain the first model.
10. An image processing apparatus characterized by comprising: The device comprises: A determination module configured to determine a second image according to a first image to be processed, the image size of the second image satisfying a preset size requirement; A first processing module configured to obtain a first feature map and a second feature map based on the second image, wherein the first feature map is used to determine the opacity of an edge region in the second image, the edge region is a region where a foreground image and a background image in the second image meet, and the second feature map is used to determine the opacity of a non-edge region in the second image; An acquisition module configured to acquire a target opacity feature map corresponding to the first image according to the first image, the first feature map, and the second feature map; A second processing module configured to process the first image according to the target opacity feature map.
11. A computer readable medium having stored thereon a computer program, characterized in that The computer program is executed by a processing device to implement the steps of the method of any one of claims 1 to 9.
12. An electronic device, comprising: Comprise: A storage device having a computer program stored thereon; processing means for executing the computer program in said storage means to implement the steps of the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Image target area acquisition method and device, equipment, medium and program product
CN114299101A
Part identification method and device based on image identification model
CN115170471A