Image processing method and device, readable medium and electronic equipment
By adjusting the size and obtaining feature maps for the image to be processed, the problem of low cutting processing efficiency of deep learning models in the prior art is solved, and a more efficient image processing effect is achieved.
Patent Information
- Application Number
- CN202311477321.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-07
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2043-11-07
AI Technical Summary
In the prior art, the cutout processing based on the deep learning model is low in efficiency and has a large amount of calculation, making it difficult to meet the needs of efficient image processing.
By determining a second image that meets the preset size requirement, a first feature map and a second feature map are acquired based on the image for determining the opacity of edge and non-edge areas in the image, and processing the original image according to these feature maps.
This reduces the complexity of the model and improves the efficiency of model, thereby improving the efficiency of image processing and achieving more efficient cutout and image segmentation effects.
Smart Images

Figure CN119963589A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular, to an image processing method, device, readable medium and electronic device. Background Art
[0002] With the advancement of computer technology, image matting technology has been used more and more widely. Image matting refers to extracting the target of interest from an image. An image can be composed of a foreground and background, and image matting can distinguish the foreground and background of a given image, thereby extracting the foreground target.
[0003] In related technologies, image clipping can be achieved based on a deep learning model, but the model has a large amount of computation and low efficiency. Summary of the invention
[0004] This summary is provided to introduce concepts in a brief form that will be described in detail in the detailed description below. This summary is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0005] According to a first aspect of an embodiment of the present disclosure, there is provided an image processing method, the method comprising:
[0006] Determine a second image according to the first image to be processed, wherein the image size of the second image meets a preset size requirement;
[0007] Obtaining a first feature map and a second feature map based on the second image; wherein the first feature map is used to determine the opacity of an edge area in the second image, the edge area being an area at the junction of a foreground image and a background image in the second image, and the second feature map is used to determine the opacity of a non-edge area in the second image;
[0008] Acquire a target opacity feature map corresponding to the first image according to the first image, the first feature map, and the second feature map;
[0009] The first image is processed according to the target opacity feature map.
[0010] According to a second aspect of an embodiment of the present disclosure, there is provided an image processing apparatus, the apparatus comprising:
[0011] A determination module, configured to determine a second image according to a first image to be processed, wherein an image size of the second image meets a preset size requirement;
[0012] A first processing module, configured to obtain a first feature map and a second feature map based on the second image input; wherein the first feature map is used to determine the opacity of an edge area in the second image, the edge area being an area at the junction of a foreground image and a background image in the second image, and the second feature map is used to determine the opacity of a non-edge area in the second image;
[0013] an acquisition module, configured to acquire a target opacity feature map corresponding to the first image according to the first image, the first feature map and the second feature map;
[0014] The second processing module is used to process the first image according to the target opacity feature map.
[0015] According to a third aspect of an embodiment of the present disclosure, a computer-readable medium is provided, on which a computer program is stored, and when the computer program is executed by a processing device, the steps of the method described in the first aspect of the present disclosure are implemented.
[0016] According to a fourth aspect of an embodiment of the present disclosure, there is provided an electronic device, including:
[0017] a storage device having a computer program stored thereon;
[0018] A processing device is used to execute the computer program in the storage device to implement the steps of the method described in the first aspect of the present disclosure.
[0019] By adopting the above technical solution, the second image is determined according to the first image to be processed, and the image size of the second image meets the preset size requirement; the first feature map and the second feature map are obtained based on the second image; wherein the first feature map is used to determine the opacity of the edge area in the second image, and the edge area is the area where the foreground image and the background image in the second image intersect, and the second feature map is used to determine the opacity of the non-edge area in the second image; according to the first image, the first feature map and the second feature map, the target opacity feature map corresponding to the first image is obtained; and the first image is processed according to the target opacity feature map. In this way, the core target model reasoning calculation part is performed on the second image that meets the preset size requirement, which can reduce the model complexity, improve the model efficiency, and thus improve the image processing efficiency.
[0020] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the originals and elements are not necessarily drawn to scale. In the drawings:
[0022] Figure 1 It is a flowchart of an image processing method according to an embodiment of the present disclosure.
[0023] Figure 2 is based on Figure 1 The illustrated embodiment shows a flow chart of step S103.
[0024] Figure 3 It is a schematic diagram of an image processing method according to an embodiment of the present disclosure.
[0025] Figure 4 It is a flowchart of a method for training a target model according to an embodiment of the present disclosure.
[0026] Figure 5 It is a schematic diagram showing an image processing according to an embodiment of the present disclosure.
[0027] Figure 6 It is a block diagram of an image processing device according to an embodiment of the present disclosure.
[0028] Figure 7 It is a block diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0029] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein, which are instead provided for a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.
[0030] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0031] The term "including" and its variations used in the present disclosure are open inclusions, i.e., "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.
[0032] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0033] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise clearly indicated in the context, they should be understood as "one or more". In the description of the present disclosure, unless otherwise specified, "multiple" refers to two or more than two, and other quantifiers are similar; "at least one item", "one or more items" or similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one item a can represent any number of a; for another example, one or more items of a, b and c can represent: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or multiple; "and / or" is a kind of association relationship that describes the associated objects, indicating that there can be three relationships, for example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. The character “ / ” indicates that the preceding and following objects are in an “or” relationship. The singular forms “a”, “an”, “an”, “said” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise.
[0034] Although operations or steps are described in a specific order in the drawings in the embodiments of the present disclosure, it should not be understood that it is required to perform these operations or steps in the specific order shown or in a serial order, or to perform all the operations or steps shown to obtain the desired results. In the embodiments of the present disclosure, these operations or steps can be performed in series; these operations or steps can also be performed in parallel; or some of these operations or steps can be performed.
[0035] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0036] It is understandable that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, scope of use, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0037] For example, in response to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested to be performed will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, application, server, or storage medium that performs the operation of the technical solution of the present disclosure according to the prompt message.
[0038] As an optional but non-limiting implementation, in response to receiving an active request from the user, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0039] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that meet the relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0040] At the same time, it is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and relevant provisions.
[0041] The disclosed embodiments can be applied to image processing scenarios, such as image cutout, image segmentation, or target extraction in an image. A target image (I) can be formed by merging a foreground (F) and a background (B) according to the foreground opacity (γ), and image cutout is to distinguish the foreground from the background of a given image, and image cutout or image segmentation can be achieved by calculating the foreground opacity.
[0042] It should be noted that the foreground opacity γ can also be represented by α (alpha).
[0043] In some embodiments, end-to-end cutout means that the algorithm system only needs to input the original image, and the output foreground mask can be calculated. The definition of foreground can be classified into different vertical tasks. For example, for portrait cutout, the foreground is defined as the person in the picture; for pet cutout, the foreground is defined as small animals such as cats and dogs in the picture; for saliency cutout (subject cutout), the foreground is defined as the most core subject in the picture.
[0044] Fine-grained cutouts require that the estimated opacity map has a certain degree of fineness, such as being able to reflect some very fine hairs and filaments. Common fine-grained cutouts have high requirements on the model, and the model calculation is complex and the efficiency is low.
[0045] In some embodiments of the present disclosure, the target image can be divided into a foreground image and a background image. For example, the foreground image and the background image can be merged to obtain the target image according to the foreground opacity. Conversely, in order to obtain the foreground image in the target image, the target image can be segmented according to the foreground opacity.
[0046] In some embodiments, the foreground opacity of the target image may be directly acquired.
[0047] In other embodiments, the foreground opacity can be decomposed into a first guiding parameter a and a second guiding parameter b, wherein the first guiding parameter a can play a dominant role in estimating the foreground opacity of the edge area, while the second guiding parameter b can play a dominant role in estimating the foreground opacity of the non-edge area (such as the determined foreground, background, and the location where the edge is unclear).
[0048] In this way, the first guiding parameter a and the second guiding parameter b may be acquired first, and then based on a linear relationship, the foreground opacity of the target image may be determined according to the first guiding parameter a, the second guiding parameter b and the target image.
[0049] Furthermore, based on local linear model modeling, it can be considered that the local foreground and background colors of the target image are smooth and less sensitive to operations such as resizing. Therefore, after the image is scaled, this method is still effective, and the small-sized segmentation estimation map (feature map) will not introduce obvious blur, jagged edges or burrs that interfere with the edge quality of the foreground opacity after enlarging.
[0050] Figure 1 is a flowchart of an image processing method according to an embodiment of the present disclosure. The method can be applied to electronic devices, which may include terminal devices, such as smart phones, smart wearable devices, smart speakers, smart tablets, PDAs (Personal Digital Assistants), CPEs (Customer Premise Equipment), personal computers, vehicle-mounted terminals, etc.; the electronic devices may also include servers, such as local servers or cloud servers. Figure 1 As shown, the method may include:
[0051] S101. Determine a second image according to a first image to be processed.
[0052] The image size of the second image meets a preset size requirement. The image size of the first image may meet or not meet the preset size requirement.
[0053] In some embodiments, if the image size of the first image meets a preset size requirement, the first image can be directly used as the second image.
[0054] In some other embodiments, if the image size of the first image does not meet the preset size requirement, a size adjustment operation may be performed on the first image to obtain the second image. The size adjustment operation may be a size reduction operation or a size enlargement operation.
[0055] In one implementation, the size adjustment may be a size reduction operation, for example, the image size of the first image may be reduced by downsampling to obtain the second image.
[0056] Optionally, the size of the first image may be reduced to a range that meets a preset size requirement to obtain the second image. For example, the range that meets the preset size requirement may be a size range that is less than or equal to the first preset size and greater than or equal to the second preset size; the first preset size is greater than the second preset size, and the first preset size and the second preset size may be pre-set sizes. For another example, the range that meets the preset size requirement may also be a size range that is less than or equal to the first preset size. For another example, the range that meets the preset size requirement may also be equal to the first preset size.
[0057] In this way, the image size of the input model can be within a preset size range, which can reduce the complexity of model processing.
[0058] In one implementation, the size adjustment may be a size enlargement operation, for example, the image size of the first image may be enlarged by upsampling to obtain the second image.
[0059] Optionally, the size of the first image may be enlarged to a range that meets the above-mentioned preset size requirement to obtain the second image.
[0060] In some embodiments, the image size may be represented by length and width, and the units of the length and width may be meters, decimeters, centimeters, millimeters or pixels.
[0061] It should be noted that the first image may be a picture or a video, and the first image may be an image captured in real time by a camera device, or a pre-stored image, or an image received from other devices, and the present disclosure does not limit the type of the first image. The first image may be an image captured in real time, or a pre-stored image, or an image received from other devices, and the present disclosure does not limit the method for obtaining the first image.
[0062] S102: Obtain a first feature map and a second feature map based on the second image.
[0063] In some embodiments, the second image may be input into a pre-generated target model to obtain a first feature map and a second feature map output by the target model.
[0064] In some embodiments, the target model may be a neural network model, such as a convolutional neural network (CNN) or a recursive neural network (RNN).
[0065] In some embodiments, the target model may include an encoder Encoder and a decoder Decoder.
[0066] In some embodiments, the image size of the first feature map, the image size of the second feature map, and the image size of the second image are all the same.
[0067] In other embodiments, the image size of the first feature map is the same as the image size of the second feature map. The image size of the first feature map may be smaller than the image size of the second image. Similarly, the image size of the second feature map may also be smaller than the image size of the second image.
[0068] In some embodiments, the first feature map can be used to determine the opacity of the edge area in the second image, and the edge area can be the area where the foreground image and the background image in the second image intersect. The second feature map can be used to determine the opacity of the non-edge area in the second image.
[0069] For example, the first feature map may include a first guidance parameter a corresponding to each pixel of the second image, and the second feature map may include a second guidance parameter b corresponding to each pixel of the second image. For a specific description of the first guidance parameter a and the second guidance parameter b, reference may be made to the description of the aforementioned embodiment of the present disclosure, which will not be repeated here.
[0070] In some embodiments, the first feature map may be a three-channel feature map. For example, the first feature map may include feature maps of three channels: red (R), green (G), and blue (B).
[0071] In some embodiments, the second feature map may be a single-channel feature map.
[0072] S103. Acquire a target opacity feature map corresponding to the first image according to the first image, the first feature map, and the second feature map.
[0073] For example, the first image, the first feature map, and the second feature map may be fused to obtain a target opacity feature map.
[0074] In this way, the influence of size change can be reduced by the first feature map and the second feature map, and the accuracy of the target opacity feature map can be improved.
[0075] S104: Process the first image according to the target opacity feature map.
[0076] For example, the first image may be matted or segmented according to the target opacity feature map to obtain a foreground target in the first image.
[0077] By adopting the above method, the second image is determined according to the first image to be processed, and the image size of the second image meets the preset size requirement; the first feature map and the second feature map are obtained based on the second image input; wherein the first feature map is used to determine the opacity of the edge area in the second image, and the edge area is the area where the foreground image and the background image in the second image intersect, and the second feature map is used to determine the opacity of the non-edge area in the second image; according to the first image, the first feature map and the second feature map, the target opacity feature map corresponding to the first image is obtained; and the first image is processed according to the target opacity feature map. In this way, the core target model reasoning calculation part is performed on the second image that meets the preset size requirement, which can reduce the model complexity, improve the model efficiency, and thus improve the image processing efficiency.
[0078] Figure 2 is based on Figure 1 The embodiment shown is a flow chart of step S103. Figure 2 As shown, the step S103 may include the following sub-steps:
[0079] S1031. Determine a third feature map according to the first feature map.
[0080] The image size of the third feature map is the same as that of the first image.
[0081] S1032. Determine a fourth feature map according to the second feature map.
[0082] The image size of the fourth feature map is the same as that of the first image.
[0083] In some embodiments, if the image size of the first image meets the preset size requirement, the first image is directly used as the second image in the aforementioned step S101, then in step S1031, the first feature map can be directly used as the third feature map, and in step S1032, the second feature map can be directly used as the fourth feature map.
[0084] In some other embodiments, the first characteristic map may be resized to obtain a third characteristic map, and the second characteristic map may be resized to obtain a fourth characteristic map.
[0085] Optionally, the resizing operation may be a downsizing operation (eg, downsampling) or an upsizing operation (eg, upsampling).
[0086] S1033. Perform feature operation according to the first image, the third feature map, and the fourth feature map to obtain an initial opacity feature map corresponding to the first image.
[0087] In some embodiments, the initial opacity feature map may be obtained through a multiplication-addition operation.
[0088] For example, the third feature map and the first image may be firstly subjected to a dot product operation to obtain a fifth feature map; and then the fifth feature map and the fourth feature map may be added to obtain an initial opacity feature map.
[0089] S1034: Activate the initial opacity feature map to obtain a target opacity feature map.
[0090] The target opacity feature map may also be referred to as a target foreground opacity feature map.
[0091] In some embodiments, the initial opacity feature map may be activated by an activation function to obtain a target opacity feature map. The activation function may include a sigmoid function and / or a sigmoid mid-segment linear transformation function.
[0092] For example, the activation function may include the following formula (1):
[0093]
[0094] Among them, Y represents the target opacity feature map, X represents the initial opacity feature map, and Sigmoid is the activation function.
[0095] It should be noted that the image cutout task can be considered as a segmentation task from a regional perspective, following the pixel classification task, i.e. foreground and background classification, and is often supervised by a cross entropy type loss function, such as using a class-balanced cross entropy loss function in combination with a sigmoid activation function under a single channel for supervision. This can effectively avoid the gradient vanishing problem when the opacity estimate is 0 or 1 using linear truncation activation, and improve the integrity and accuracy of the segmented region. However, for edge locations such as hair, which are mainly regression problems, the estimated opacity value needs to be more accurate, and it is expected that the original image will be linearly related when multiplied and added to guide the calculation. Therefore, a sigmoid mid-segment linear transformation is designed for the activation function after the final guided multiplication and addition operation, and the sigmoid mid-segment linear transformation is achieved through the above formula (1), thereby improving the accuracy of the output target opacity feature map.
[0096] In some embodiments, different activation functions may be used without distinguishing different scenes. For example, in scenes with complex backgrounds and where regional accuracy is given higher priority, a general sigmoid function may be used. For another example, in scenes with simple scenes and where hair quality is given higher priority, the sigmoid mid-segment linear transformation of the above formula (1) may be used.
[0097] In this way, different activation functions can be adapted according to different scenarios to further improve the accuracy of the output target opacity feature map.
[0098] Figure 3 FIG. 1 is a schematic diagram of an image processing method according to an embodiment of the present disclosure. Figure 3 As shown, the input of the target model 311 is the second image 302, and the second image 302 is obtained according to the first image 301 to be processed. For example, if the first image 301 is a large-size image or a high-resolution image, the first image 301 can be reduced in size to obtain the second image 302 that meets the preset size requirement. In this embodiment, the image processing method may include:
[0099] First, the second image 302 is input into the target model 311 to obtain the output feature 321 of the target model 311. The output feature 321 may include a first feature map and a second feature map, the image size of the first feature map, the image size of the second feature map, and the image size of the second image are all the same, the first feature map may include a first guidance parameter a corresponding to each pixel of the second image, and the second feature map may include a second guidance parameter b corresponding to each pixel of the second image.
[0100] Secondly, the output feature 321 is input into the sampling module for sampling to obtain a sampled output 322. The sampled output 322 may include a third feature map and a fourth feature map, and the image size of the third feature map, the image size of the fourth feature map, and the image size of the first image are all the same. Among them, the third feature map may be a feature map obtained by sampling (up-sampling or down-sampling) the first feature map, and the third feature map may include a third guide parameter A corresponding to each pixel of the first image. The fourth feature map may be a feature map obtained by sampling (up-sampling or down-sampling) the second feature map, and the fourth feature map may include a fourth guide parameter B corresponding to each pixel of the first image.
[0101] Again, the sampling feature 322 and the first image 302 are multiplied and added to obtain an initial opacity feature map corresponding to the first image. For example, the third feature map and the first image may be first multiplied to obtain a fifth feature map; and then the fifth feature map and the fourth feature map may be added to obtain an initial opacity feature map.
[0102] Finally, the initial opacity feature map is activated by the activation function 313 to obtain the target opacity feature map 323. The activation function 313 may include a sigmoid function and / or a sigmoid mid-segment linear transformation function.
[0103] In this way, the core target model reasoning calculation part can be performed on the second image that meets the preset size requirements, and the effect optimization from coarse segmentation to refined cutout can be achieved without increasing the calculation cost, thereby improving the efficiency of image processing. In addition, the linear guidance parameters are calculated based on the first feature map and the second feature map, which reduces the "burr" problem introduced at the position close to the edge and background color when optimizing the cutout quality of hair, and also improves the quality of image processing.
[0104] Figure 4 FIG. 1 is a flow chart of a method for training a target model according to an embodiment of the present disclosure. Figure 4 As shown, the training method may include:
[0105] S401: Acquire a first sample set and a second sample set.
[0106] The first sample set includes a plurality of first sample data of coarse-grained types, and the second sample set includes a plurality of second sample data of fine-grained types, and the annotation accuracy of the fine-grained types is higher than that of the coarse-grained types.
[0107] Optionally, the fine-grained type of annotation can be a detailed annotation of the edge areas of the foreground image and the background image, for example, it can be a hair-level annotation, which finely marks the edge hair of the foreground image; the coarse-grained type of annotation can be a contour-level annotation, for example, only the edge contour can be marked without implementing hair-level annotation.
[0108] In some embodiments, the first sample data of the coarse-grained type may also be referred to as coarse standard data, and the second sample data of the fine-grained type may also be referred to as finely labeled data.
[0109] S402: Pre-train the second model according to the first sample set to obtain the first model.
[0110] In some embodiments, the second model may be a neural network model, such as a CNN or a RNN.
[0111] In some embodiments, the pre-training may also be referred to as segmentation pre-training. In the image cutout task, the accuracy of the segmented area needs to be guaranteed to ensure the overall quality of the image cutout. In this way, the first sample set (coarsely annotated data) may be used to perform pre-training on a small-sized sample image.
[0112] In one implementation, only a single-channel output may be used during the pre-training process, for example, only the second feature map is output, or any one of the three RGB channels of the first feature map is output.
[0113] In another implementation, multi-channel output may also be used during the pre-training process, such as outputting a first feature map and a second feature map.
[0114] Pre-training helps improve the accuracy of the entire model in estimating the foreground area in the vertical cutout task.
[0115] In some embodiments, the step S402 may be omitted, for example, a second model that has not been pre-trained may be used as the first model. For another example, a preset neural network model may be used as the first model.
[0116] S403: Train the first model according to the first sample set and the second sample set to obtain a target model.
[0117] In some embodiments, the first model may be trained by looping through the model training steps until it is determined that the trained first model satisfies a preset condition for stopping iteration, and then using the trained first model as the target model.
[0118] Optionally, the preset stop iteration condition may include that the target loss value is less than or equal to a preset loss threshold, or that the change value of the target loss value within a certain number of iterations is less than a preset change threshold, or that the number of iterations is greater than or equal to a preset number, or that it may be a commonly used condition for stopping iteration in the related art, which is not limited in the embodiments of the present disclosure. The above-mentioned preset loss threshold or preset change threshold may be a preset value.
[0119] In addition, if it is determined according to the target loss value that the trained first model satisfies the preset stop iteration condition, the model training step can be stopped.
[0120] In some embodiments, the model training step may include:
[0121] S11. Perform N rounds of iterative training on the first model according to the first loss function set and the first sample set to obtain a first model after iteration, where N is a positive integer.
[0122] S12. Perform M rounds of iterative training on the iterated first model according to the second loss function set and the second sample set to obtain the trained first model, where M is a positive integer and is less than or equal to N.
[0123] In some embodiments, M may be equal to 1, and N may be a positive integer greater than 1, such as 5 or 10. In this way, N coarse-grained first sample set training rounds and 1 fine-grained second sample set training round are interleaved during the training process, which can improve the training efficiency and the accuracy of the trained model.
[0124] In some embodiments of the present disclosure, the target sample data may include a target first sample image, a target second sample image, and target annotation information, wherein the target second sample image may be an image that meets a preset size requirement determined based on the target first sample image, and the image sizes of the target first sample image and the target second sample image may be different or the same. The target second sample image is used as the input of the first model. The target sample data may be the first sample data or the second sample data.
[0125] The input of the above-mentioned first model can be a target second sample image that meets the preset size requirements. The output of the first model can include a sample first feature map and a sample second feature map. During training, the sample opacity feature map corresponding to the target first sample image can be obtained based on the target first sample image, the sample first feature map and the sample second feature map.
[0126] In some embodiments, the above-mentioned first loss function set can be used to constrain the loss of at least one of the sample first feature map, the sample second feature map and the sample opacity feature map.
[0127] For example, the first loss function set may include at least one of a first loss function, a second loss function, and a third loss function, wherein:
[0128] The first loss function can be used to determine a first loss value of the entire image of the target sample data. The first loss value is used to characterize the difference between the labeled opacity and the predicted opacity corresponding to the target sample data in the entire image. The predicted opacity is the opacity predicted and output by the first model for the target sample data. The target sample data is the first sample data or the second sample data.
[0129] Optionally, the first loss function may be a category-balanced cross entropy loss function, a weighted balanced cross entropy loss function, an overlap-based loss function, an intersection-over-union loss function, or a regional mutual information loss function, etc., and the first loss function may be used to perform equal-weighted supervision on the entire image.
[0130] The second loss function is used to determine the second loss value of the target sample data in the edge area. The second loss value is used to characterize the difference between the labeled first feature map and the predicted first feature map corresponding to the target sample data in the edge area. The predicted first feature map is the first feature map predicted and output by the first model for the target sample data.
[0131] Optionally, the second loss function may further supervise the edge region, for example, may only supervise the edge regularization term.
[0132] In some embodiments, the weight of the loss value of the first guidance parameter corresponding to the pixels in the non-edge area of the first feature map can be reduced, for example, reduced to N times the original loss value, where N is a decimal less than 1, such as 0.1, 0.01 or 0.001. In this way, the weight of the non-edge area can be reduced to achieve supervision of the edge area.
[0133] Optionally, the third loss function is used to determine a third loss value of the entire image of the target sample data, and the third loss value is used to characterize the difference between the labeled second feature map and the predicted second feature map corresponding to the target sample data in the entire image, and the predicted second feature map is the second feature map predicted and output by the first model for the target sample data.
[0134] The second feature map serves as the channel output of the dominant segmentation tendency (pixel classification), and the third loss function for supervising the second feature map can be similar to the first loss function. For example, loss functions such as the category balanced cross entropy loss function and the regional mutual information loss function can be used for supervision.
[0135] In one implementation, for the sample data whose segmentation is annotated with binary labels, only the category-balanced cross entropy loss function can be used for supervision to avoid the sawtooth phenomenon introduced by the regional mutual information loss function to the edge distribution learning.
[0136] In some embodiments, the second loss function set may include the first loss function, the second loss function, the third loss function and the fourth loss function, wherein:
[0137] The fourth loss function can be used to determine a fourth loss value of an edge area of the target sample data, and the fourth loss value is used to characterize the difference between the labeled opacity and the predicted opacity corresponding to the target sample data in the edge area, and the predicted opacity is the opacity predicted and output by the first model for the target sample data.
[0138] Optionally, the fourth loss function can be used to perform weighted loss supervision on the edge area. For example, the edge area is obtained by eroding and dilating the label label to achieve K times the weight, where K can be a positive integer greater than 1, such as 10.
[0139] Optionally, the fourth loss function may be: an L1 loss function (L1 Loss) or a Laplacian Loss (Laplacian Loss). The Laplacian Loss may be used to enhance the hair texture.
[0140] In this way, when performing M rounds of iterative training on the iterated first model according to the second sample set, a fourth loss function can be added on the basis of the first loss function set to perform edge-weighted loss supervision on opacity, thereby further improving the accuracy of the target opacity in the edge area.
[0141] Figure 5 FIG. 1 is a schematic diagram showing an image processing effect according to an embodiment of the present disclosure. Figure 5 As shown, based on the steps S101 to S103 of the aforementioned embodiment of the present disclosure, after processing the first image 501, a target opacity feature map 511 can be obtained; similarly, based on the steps S101 to S103 of the aforementioned embodiment of the present disclosure, after processing the first image 502, a target opacity feature map 512 can be obtained. It can be seen that the target opacity feature map can realize hair-level segmentation.
[0142] Figure 6 is a block diagram of an image processing device 1100 according to an embodiment of the present disclosure, such as Figure 6 As shown, the device 1100 may include:
[0143] A determination module 1101 is used to determine a second image according to a first image to be processed, wherein the image size of the second image meets a preset size requirement;
[0144] A first processing module 1102 is configured to obtain a first feature map and a second feature map based on the second image; wherein the first feature map is used to determine the opacity of an edge region in the second image, the edge region being the region where the foreground image and the background image meet in the second image, and the second feature map is used to determine the opacity of a non-edge region in the second image;
[0145] An acquisition module 1103 is used to acquire a target opacity feature map corresponding to the first image according to the first image, the first feature map and the second feature map;
[0146] The second processing module 1104 is used to process the first image according to the target opacity feature map.
[0147] In some embodiments, the acquisition module 1103 is used to determine a third feature map based on the first feature map, and the image size of the third feature map is the same as that of the first image; determine a fourth feature map based on the second feature map, and the image size of the fourth feature map is the same as that of the first image; perform feature operations based on the first image, the third feature map, and the fourth feature map to obtain an initial opacity feature map corresponding to the first image; and activate the initial opacity feature map to obtain the target opacity feature map.
[0148] In some embodiments, the acquisition module 1103 is used to perform a dot multiplication operation on the third feature map and the first image to obtain a fifth feature map; and perform an addition operation on the fifth feature map and the fourth feature map to obtain the initial opacity feature map.
[0149] In some embodiments, the activation function includes a sigmoid function and / or a sigmoid mid-piece linear transformation function.
[0150] In some embodiments, the first processing module 1102 is used to input the second image into a pre-generated target model to obtain a first feature map and a second feature map output by the target model; wherein the image size of the first feature map, the image size of the second feature map and the image size of the second image are the same.
[0151] In some embodiments, the apparatus further comprises:
[0152] The model generation module 1105 is used to obtain a first sample set and a second sample set; wherein the first sample set includes a plurality of first sample data of coarse-grained types, and the second sample set includes a plurality of second sample data of fine-grained types, and the annotation accuracy of the fine-grained types is higher than that of the coarse-grained types; the first model is trained according to the first sample set and the second sample set to obtain the target model.
[0153] In some embodiments, the model generation module 1105 is used to cyclically execute the model training steps until it is determined that the trained first model meets the preset stop iteration condition, and then the trained first model is used as the target model;
[0154] The model training step includes:
[0155] Perform N rounds of iterative training on the first model according to the first loss function set and the first sample set to obtain an iterated first model, where N is a positive integer;
[0156] Perform M rounds of iterative training on the iterated first model according to the second loss function set and the second sample set to obtain the trained first model, where M is a positive integer and is less than or equal to N.
[0157] In some embodiments, the first loss function set includes a first loss function, a second loss function and a third loss function, wherein:
[0158] The first loss function is used to determine a first loss value of the entire image of the target sample data, the first loss value is used to characterize the difference between the labeled opacity and the predicted opacity corresponding to the target sample data in the entire image, the predicted opacity is the opacity predicted and output by the first model for the target sample data, and the target sample data is the first sample data or the second sample data;
[0159] The second loss function is used to determine a second loss value of the target sample data in the edge area, and the second loss value is used to characterize the difference between the labeled first feature map and the predicted first feature map corresponding to the target sample data in the edge area, and the predicted first feature map is the first feature map predicted and output by the first model for the target sample data;
[0160] The third loss function is used to determine the third loss value of the entire image of the target sample data, and the third loss value is used to characterize the difference between the labeled second feature map and the predicted second feature map corresponding to the target sample data in the entire image, and the predicted second feature map is the second feature map predicted and output by the first model for the target sample data.
[0161] In some embodiments, the second loss function set includes the first loss function, the second loss function, the third loss function and a fourth loss function, wherein:
[0162] The fourth loss function is used to determine a fourth loss value of the edge area of the target sample data, and the fourth loss value is used to characterize the difference between the labeled opacity and the predicted opacity corresponding to the target sample data in the edge area, and the predicted opacity is the opacity predicted and output by the first model for the target sample data.
[0163] In some embodiments, the model generation module 1105 is further used to pre-train the second model according to the first sample set to obtain the first model.
[0164] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0165] Reference below Figure 7, which shows a schematic diagram of the structure of an electronic device 2000 (e.g., a terminal device or a server) suitable for implementing the embodiment of the present disclosure. The terminal device in the embodiment of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (e.g., vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. The server in the embodiment of the present disclosure may include, but is not limited to, local servers, cloud servers, single servers, distributed servers, etc. Figure 7 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0166] like Figure 7 As shown, the electronic device 2000 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 2001, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 2002 or a program loaded from a storage device 2008 into a random access memory (RAM) 2003. In the RAM 2003, various programs and data required for the operation of the electronic device 2000 are also stored. The processing device 2001, the ROM 2002, and the RAM 2003 are connected to each other via a bus 2004. An input / output (I / O) interface 2005 is also connected to the bus 2004.
[0167] Typically, the following devices may be connected to the input / output interface 2005: an input device 2006 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 2007 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 2008 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 2009. The communication device 2009 may allow the electronic device 2000 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 7 The electronic device 2000 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.
[0168] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 2009, or installed from the storage device 2008, or installed from the ROM 2002. When the computer program is executed by the processing device 2001, the above-mentioned functions defined in the method of the embodiment of the present disclosure are executed.
[0169] It should be noted that the computer-readable medium disclosed above may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. Computer-readable storage media may include, but are not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0170] In some embodiments, the client and the server may communicate using any currently known or future developed network protocol such as HTTP (HyperText Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.
[0171] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0172] The above-mentioned computer-readable medium carries one or more programs. When the above-mentioned one or more programs are executed by the electronic device, the electronic device: determines the second image according to the first image to be processed, and the image size of the second image meets the preset size requirement; obtains the first feature map and the second feature map based on the second image; wherein the first feature map is used to determine the opacity of the edge area in the second image, and the edge area is the area where the foreground image and the background image in the second image intersect, and the second feature map is used to determine the opacity of the non-edge area in the second image; obtains the target opacity feature map corresponding to the first image according to the first image, the first feature map and the second feature map; and processes the first image according to the target opacity feature map.
[0173] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or a combination thereof, including, but not limited to, object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0174] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0175] The modules involved in the embodiments described in the present disclosure may be implemented by software or hardware. The name of the module does not limit the module itself in some cases. For example, the determination module may also be described as "determining the second image according to the first image to be processed".
[0176] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0177] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0178] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the above features are replaced with the technical features with similar functions disclosed in the present disclosure (but not limited to) by each other to form a technical solution.
[0179] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0180] Although the subject matter has been described in language specific to structural features and / or method logic actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. On the contrary, the specific features and actions described above are merely example forms of implementing the claims. Regarding the device in the above embodiment, the specific manner in which each module performs the operation has been described in detail in the embodiment related to the method, and will not be elaborated here.
Claims
1. An image processing method, characterized in that: The method comprises: Determine a second image according to the first image to be processed, wherein the image size of the second image meets a preset size requirement; Obtaining a first feature map and a second feature map based on the second image; wherein the first feature map is used to determine the opacity of an edge area in the second image, the edge area being an area at the junction of a foreground image and a background image in the second image, and the second feature map is used to determine the opacity of a non-edge area in the second image; Acquire a target opacity feature map corresponding to the first image according to the first image, the first feature map, and the second feature map; The first image is processed according to the target opacity feature map.
2. The method according to claim 1, characterized in that The acquiring, according to the first image, the first feature map and the second feature map, a target opacity feature map corresponding to the first image comprises: Determine a third feature map according to the first feature map, wherein the image size of the third feature map is the same as that of the first image; Determine a fourth feature map according to the second feature map, wherein the image size of the fourth feature map is the same as that of the first image; Performing feature operation according to the first image, the third feature map, and the fourth feature map to obtain an initial opacity feature map corresponding to the first image; The initial opacity feature map is activated to obtain the target opacity feature map.
3. The method according to claim 2, characterized in that The performing feature operation according to the first image, the third feature map and the fourth feature map to obtain an initial opacity feature map corresponding to the first image comprises: Performing a dot product operation on the third feature map and the first image to obtain a fifth feature map; The fifth feature map and the fourth feature map are added to obtain the initial opacity feature map.
4. The method according to any one of claims 1 to 3, characterized in that The obtaining of the first feature map and the second feature map based on the second image comprises: The second image is input into a pre-generated target model to obtain a first feature map and a second feature map output by the target model; wherein the image size of the first feature map, the image size of the second feature map and the image size of the second image are the same.
5. The method according to claim 4, characterized in that The target model is a model generated based on the following method: Acquire a first sample set and a second sample set; wherein the first sample set includes a plurality of first sample data of coarse-grained types, and the second sample set includes a plurality of second sample data of fine-grained types, and the annotation accuracy of the fine-grained types is higher than that of the coarse-grained types; The first model is trained according to the first sample set and the second sample set to obtain the target model.
6. The method according to claim 5, characterized in that The training of the first model according to the first sample set and the second sample set to obtain the target model comprises: The model training step is executed cyclically until it is determined that the first model after training meets the preset stop iteration condition, and the first model after training is used as the target model; The model training step includes: Perform N rounds of iterative training on the first model according to the first loss function set and the first sample set to obtain an iterated first model, where N is a positive integer; Perform M rounds of iterative training on the iterated first model according to the second loss function set and the second sample set to obtain the trained first model, where M is a positive integer and is less than or equal to N.
7. The method according to claim 6, characterized in that The first loss function set includes a first loss function, a second loss function and a third loss function, wherein: The first loss function is used to determine a first loss value of the entire image of the target sample data, the first loss value is used to characterize the difference between the labeled opacity and the predicted opacity corresponding to the target sample data in the entire image, the predicted opacity is the opacity predicted and output by the first model for the target sample data, and the target sample data is the first sample data or the second sample data; The second loss function is used to determine a second loss value of the target sample data in the edge area, and the second loss value is used to characterize the difference between the labeled first feature map and the predicted first feature map corresponding to the target sample data in the edge area, and the predicted first feature map is the first feature map predicted and output by the first model for the target sample data; The third loss function is used to determine the third loss value of the entire image of the target sample data, and the third loss value is used to characterize the difference between the labeled second feature map and the predicted second feature map corresponding to the target sample data in the entire image, and the predicted second feature map is the second feature map predicted and output by the first model for the target sample data.
8. The method according to claim 7, characterized in that The second loss function set includes the first loss function, the second loss function, the third loss function and the fourth loss function, wherein: The fourth loss function is used to determine a fourth loss value of the edge area of the target sample data, and the fourth loss value is used to characterize the difference between the labeled opacity and the predicted opacity corresponding to the target sample data in the edge area, and the predicted opacity is the opacity predicted and output by the first model for the target sample data.
9. The method according to claim 5, characterized in that The method further comprises: The second model is pre-trained according to the first sample set to obtain the first model.
10. An image processing device, characterized in that: The device comprises: A determination module, configured to determine a second image according to a first image to be processed, wherein an image size of the second image meets a preset size requirement; A first processing module, configured to obtain a first feature map and a second feature map based on the second image; wherein the first feature map is used to determine the opacity of an edge area in the second image, the edge area being an area at the junction of a foreground image and a background image in the second image, and the second feature map is used to determine the opacity of a non-edge area in the second image; an acquisition module, configured to acquire a target opacity feature map corresponding to the first image according to the first image, the first feature map and the second feature map; The second processing module is used to process the first image according to the target opacity feature map.
11. A computer readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processing device, the steps of the method according to any one of claims 1 to 9 are implemented.
12. An electronic device, characterized in that: include: a storage device having a computer program stored thereon; A processing device, configured to execute the computer program in the storage device to implement the steps of the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Image processing method, device and equipment and storage medium
CN111369581A
Image processing method and device, electronic equipment and storage medium
CN112308866A
Image target area acquisition method and device, equipment, medium and program product
CN114299101A
Part identification method and device based on image identification model
CN115170471A
Image processing apparatus, image processing method and storage medium
US20170039691A1
Cited By
Deep learning-based waste acid resource utilization method and system
CN120833516A