Image processing method and device, network model training method and device, network model application method and device and storage medium

By designing a prediction method of integrated decoding network and combining head modules, the problems of insufficient generalization ability of portrait cutouts and unsatisfactory segmentation results in the prior art are solved, and high-precision and high-quality cutouts and segmentation results are achieved.

CN120070486APending Publication Date: 2025-05-30CANON KK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311610645.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-28
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Due to the lack of semantic guidance in portrait cutouts, the generalization capability on complex real images is insufficient, and the segmentation decoder fails to fully utilize the encoder features, resulting in unsatisfactory global segmentation results.

Method used

A training method for neural networks is designed. By building an integrated decoding network, multiple decoding modules are used to decode decoding features with the same resolution as the input image, and combined with the head module to predict the output image, the rich features of the cutout and segmentation tasks are achieved simultaneously.

Benefits of technology

The accuracy of portrait cutouts and the quality of segmentation results are improved, the model's generalization ability of complex images is enhanced, and under the limited search overhead conditions, a high-precision neural network model that meets the computational overhead constraints can be searched.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070486A_ABST
    Figure CN120070486A_ABST
Patent Text Reader

Abstract

The invention provides an image processing method and device, a network model training method and device, a network model application method and device and a storage medium. The image processing method includes: an encoding step of generating a plurality of encoding features having different resolutions based on an input image and an encoder; a decoding step: decoding a decoding feature having the same resolution as the input image based on the plurality of coding features and a decoder cascaded with a plurality of decoding modules; and a prediction step of predicting an output image having the same resolution as the input image based on the decoding feature and the head module.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image processing method, a neural network model training method, an image processing apparatus, a neural network model training apparatus and an applied method, and a storage medium. Background Art

[0002] Image matting is a fundamental problem in computer vision, and its purpose is to predict an opacity for each pixel to accurately cut out the target image region. Estimating the foreground, background, and opacity from a single image is a typical ill-posed problem, but it has a wide range of applications in the fields of image and video editing, virtual reality, augmented reality, entertainment, etc. In particular, portrait image matting refers to a specific image matting task, where the input image is a portrait without any external guidance.

[0003] Common image matting methods all require a trimap as an additional input, and on this basis, predict or estimate the opacity of each pixel in the uncertain region. In recent years, due to the success of convolutional neural networks, researchers have begun to study matting methods that do not require any external guidance. These matting models capture semantics and details through end-to-end training on large-scale datasets. However, due to the lack of semantic guidance, these methods face challenges in generalization when tested on complex real images.

[0004] In 2021, Jizhizi Li, Sihan Ma, Jing Zhang, Dacheng Tao, in Privacy-Preserving Portrait Matting, proposed a portrait matting network P3M-Net based on the U-Net structure that does not require additional guidance. It uses a unified end-to-end multi-task framework to achieve semantic perception and detail matting, and particularly emphasizes the interaction between semantic perception and detail matting and the encoder to promote the matting process. Among them, the network includes an encoder, a segmentation decoder, and a matting decoder, which can explicitly model semantic segmentation and detail matting and jointly optimize them. At the same time, the network also includes TFI, dBFI, and sBFI modules to strengthen the interaction effect between encoding and decoding, so as to be able to extract better features.

[0005] The prior art methods generate the final matting result by fusing the global segmentation result and the local matting result. If the global segmentation result is poor, it directly affects the final matting result. And the segmentation decoder of the prior art methods only uses the high-dimensional semantic features of the encoder to obtain the final decoded features, without making full use of the features of the encoder, so the global segmentation result will not be ideal. Summary of the Invention

[0006] The present invention provides a method for training a neural network, which can search for a high-precision neural network model that meets the computational overhead constraint under the condition of limited search overhead.

[0007] According to one aspect of the present invention, there is provided an image processing method, characterized in that the method includes: an encoding step of generating a plurality of encoded features with different resolutions based on an input image and an encoder; a decoding step of decoding a decoded feature with the same resolution as the input image based on the plurality of encoded features and a decoder cascaded with a plurality of decoding modules; and a prediction step of predicting an output image with the same resolution as the input image based on the decoded feature and a head module.

[0008] According to another aspect of the present invention, there is provided a method for training a neural network model, including: a construction step of constructing a neural network model used in the method according to one aspect of the present invention; a prediction step of calculating a predicted output result based on the constructed neural network model and data obtained from a training dataset; and an update step of calculating a loss based on a loss function and the predicted output result to update the parameters of the current neural network.

[0009] According to another aspect of the present invention, there is provided an image processing apparatus, the apparatus including: an encoding unit configured to generate a plurality of encoded features with different resolutions based on an input image and an encoder; a decoding unit configured to decode a decoded feature with the same resolution as the input image based on the plurality of encoded features and a decoder cascaded with a plurality of decoding modules; and a prediction unit configured to predict an output image with the same resolution as the input image based on the decoded feature and a head module.

[0010] According to another aspect of the present invention, there is provided a training apparatus for a neural network model, including: a construction unit configured to construct a neural network model used in the method according to one aspect of the present invention; a prediction unit configured to calculate a predicted output result based on the constructed neural network model and data obtained from a training dataset; and an update unit configured to calculate a loss based on a loss function and the predicted output result to update the parameters of the current neural network.

[0011] Other features of the present invention will become clear from the following description of exemplary embodiments with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The accompanying drawings incorporated in and constituting a part of this specification illustrate exemplary embodiments of the present invention and, together with the description of the exemplary embodiments, are used to explain the principles of the present invention.

[0013] Figure 1 A block diagram illustrating a hardware configuration according to an exemplary embodiment of the present invention.

[0014] Figure 2 Illustrates an integrated decoding network according to an exemplary embodiment of the present invention.

[0015] Figure 3 Illustrates an integrated decoding module for different inputs according to an exemplary embodiment of the present invention.

[0016] Figure 4 Illustrates skip connection sub-modules with different fusion methods according to an exemplary embodiment of the present invention.

[0017] Figure 5 Illustrates splitting sub-modules with different guiding methods according to an exemplary embodiment of the present invention.

[0018] Figure 6 Illustrates matte extraction sub-modules with different guiding methods according to an exemplary embodiment of the present invention.

[0019] Figure 7 Illustrates a flowchart for training a convolutional neural network model using an integrated decoder 1 according to an exemplary embodiment of the present invention.

[0020] Figure 8 Illustrates a flowchart for training a convolutional neural network model using an integrated decoder 2 according to an exemplary embodiment of the present invention.

[0021] Figure 9 Illustrates a flowchart for training a convolutional neural network model using an integrated decoder 3 according to an exemplary embodiment of the present invention.

[0022] Figure 10 Illustrates a flowchart for training a convolutional neural network model using an integrated decoder 4 according to an exemplary embodiment of the present invention.

[0023] Figure 11 Illustrates a flowchart for training a convolutional neural network model using an integrated decoder 5 according to an exemplary embodiment of the present invention.

[0024] Figure 12 Illustrates a schematic diagram of a training system according to an exemplary embodiment of the present invention. Detailed implementation

[0025] The exemplary embodiments of the present invention will be described below with reference to the accompanying drawings. For clarity and conciseness, not all features of the embodiments are described in the specification. However, it should be understood that many implementation-specific settings must be made during the implementation of the embodiments in order to achieve the specific goals of the developer, for example, to comply with those limitations related to the device and the business, and these limitations may vary with different implementations. In addition, it should also be understood that although the development work may be very complex and time-consuming, for those skilled in the art who benefit from the content of the present invention, such development work is only a routine task.

[0026] Here, it should also be noted that in order to avoid obscuring the present invention with unnecessary details, only the processing steps and / or system structures that are closely related to at least the solution of the present invention are shown in the drawings, while other details that are not closely related to the present invention are omitted.

[0027] In the context of the present invention, the data set can represent data including any image, such as a color image, a grayscale image, etc. The type and format of the image are not particularly limited.

[0028] (Hardware Configuration)

[0029] First, reference will be made to Figure 1 describe the hardware configuration that can implement the technology described below.

[0030] The hardware configuration 100 includes, for example, a central processing unit (CPU) 110, a random access memory (RAM) 120, a read-only memory (ROM) 130, a hard disk 140, an input device 150, an output device 160, a network interface 170, and a system bus 180. In one implementation, the hardware configuration 100 can be implemented by a computer, such as a tablet computer, a laptop computer, a desktop computer, or other suitable electronic devices.

[0031] In one implementation, the device for training the neural network model according to the present invention is constructed by hardware or firmware and serves as a module or component of the hardware configuration 100. In another implementation, the method for training the neural network model according to the present invention is constructed by software stored in the ROM 130 or the hard disk 140 and executed by the CPU 110.

[0032] The CPU 110 is any suitable programmable control device (such as a processor), and can perform various functions to be described below by executing various application programs stored in the ROM 130 or the hard disk 140 (such as a memory). The RAM 120 is used to temporarily store programs or data loaded from the ROM 130 or the hard disk 140, and is also used as a space where the CPU 110 executes various processes and other available functions. The hard disk 140 stores various information such as an operating system (OS), various applications, control programs, sample images, trained neural network models, predefined data (e.g., thresholds (THs)), etc.

[0033] In one implementation, the input device 150 is used to allow a user to interact with the hardware configuration 100. In one example, the user can input a sample image and a label of the sample image (e.g., region information of an object, category information of an object, etc.) through the input device 150. In another example, the user can trigger the corresponding processing of the present invention through the input device 150. In addition, the input device 150 can take various forms, such as buttons, keyboards, or touchscreens.

[0034] In one implementation, the output device 160 is used to store the finally trained neural network model in, for example, the hard disk 140 or to output the finally generated neural network model to subsequent image processing such as object detection, object classification, image segmentation, etc.

[0035] The network interface 170 provides an interface for connecting the hardware configuration 100 to a network. For example, the hardware configuration 100 can perform data communication with other electronic devices connected via the network through the network interface 170. Optionally, a wireless interface can be provided for the hardware configuration 100 to perform wireless data communication. The system bus 180 can provide a data transmission path for mutually transmitting data among the CPU 110, the RAM 120, the ROM 130, the hard disk 140, the input device 150, the output device 160, and the network interface 170, etc. Although called a bus, the system bus 180 is not limited to any specific data transmission technology.

[0036] The above-mentioned hardware configuration 100 is merely illustrative and is in no way intended to limit the present invention, its applications, or uses. Moreover, for the sake of simplicity, Figure 1 only one hardware configuration is shown. However, multiple hardware configurations can also be used as needed, and multiple hardware configurations can be connected via a network. In this case, the multiple hardware configurations can be implemented, for example, by a computer (such as a cloud server), or can also be implemented by an embedded device, such as a camera, a video camera, a personal digital assistant (PDA), or other suitable electronic devices.

[0037] Next, various aspects of the present invention will be described.

[0038] <The first exemplary embodiment>

[0039] The purpose of portrait image matting is to predict the transparency of the portrait area (foreground), that is, to determine that the transparency value of the foreground is 1, the transparency value of the background is 0, and the transparency value of the unknown area between 1 and 0 represents the possibility of a portrait. In the matting task of the exemplary embodiment of the present invention, it can be simplified into a segmentation problem. Specifically, it is divided into three typical segmentation tasks. The prior art uses two decoders for segmentation and matting respectively. The purpose of the present invention is to design a decoder that can simultaneously recover the rich features of the matting task and the segmentation task, so as to obtain better matting results.

[0040] The following will refer to Figure 7 Describe the processing in the training process of the neural network model according to the exemplary embodiment of the present invention. In this exemplary embodiment, the integrated decoder consists of five integrated decoding modules with different guiding feature inputs and two head modules for segmentation and matting respectively. The specific description is as follows.

[0041] Step S1010: Use the encoding network to extract the encoded features of each layer from the training data.

[0042] In this step, the input is a part of the training data randomly selected from the matting database. The encoding network can select existing multi-layer neural networks such as ResNet, Transformer, and MLP. When the training data is input, the encoding network generates five encoded features with gradually decreasing resolutions, named E0, E1, E2, E3, and E4 respectively.

[0043] In the context of the present invention, the training data may refer to data including any image, such as a color image, a grayscale image, etc. The type and format of the image samples are not particularly limited. In addition, the image can be an original image or its processed version, such as a version of the image that has undergone preliminary filtering or preprocessing before performing the operations of the present application on the image.

[0044] Step S1020: Decode the features through the integrated decoding module 4 without skip connections and guiding features.

[0045] In this step, as Figure 3 shown, the integrated decoding module 4 is the simplest first decoding module without a skip sub-module. Among them, the segmentation sub-module is designed as a series connection of at least two convolutional operations and one upsampling operation, as Figure 5 shown in the first design of Figure 6As shown in the first design. After obtaining the final encoded feature E4 from the encoding network, the segmentation sub-module and the matting sub-module respectively decode the features required for their respective tasks from it, and finally fuse the two decoded features by concatenation as the final decoded feature D4 of this decoding module.

[0046] Step S1030: Decode the feature through the integrated decoding module 3 that contains the skip connection feature E3 input and the guiding features E0 and E4 inputs.

[0047] In this step, the integrated decoding module 3 is designed as the most complex third decoding module, as Figure 3 shown. Among them, the skip connection sub-module is specifically designed as in Figure 4 the first way, the segmentation sub-module is specifically designed as in Figure 5 the second way, and the matting sub-module is specifically designed as in Figure 6 the second way. First, the skip connection sub-module performs feature transformation and fusion on the input decoded feature D4 and the skip connection feature E3 to obtain an enhanced feature; then in the segmentation sub-module, the enhanced feature is first partially decoded and then the guiding feature from the encoded feature E4 with high-level semantic information is added, and then further decoded and upsampled to obtain the decoded feature of the segmentation task; at the same time, in the matting sub-module, the enhanced feature is first partially decoded and then the guiding feature from the encoded feature E0 with low-level texture information is added, and then further decoded and upsampled to obtain the decoded feature of the matting task; finally, the two decoded features are fused by concatenation as the final decoded feature D3 of this decoding module.

[0048] Step S1040: Decode the feature through the integrated decoding module 2 that contains the skip connection feature E2 input and the guiding features E0 and E4 inputs.

[0049] In this step, the integrated decoding module 2 is designed as the most complex third decoding module, as Figure 3 shown. Among them, the skip connection sub-module is specifically designed as in Figure 4 the first way, the segmentation sub-module is specifically designed as in Figure 5 the second way, and the matting sub-module is specifically designed as in Figure 6The second method. First, the skip connection sub-module performs feature transformation and fusion on the input decoded feature D3 and the skip connection feature E2 to obtain an enhanced feature. Then, in the segmentation sub-module, the enhanced feature is first partially decoded and then the guiding feature from the encoded feature E4 with high-level semantic information is added, and then further decoded and upsampled to obtain the decoded feature for the segmentation task. At the same time, in the matting sub-module, the enhanced feature is first partially decoded and then the guiding feature from the encoded feature E0 with low-level texture information is added, and then further decoded and upsampled to obtain the decoded feature for the matting task. Finally, the two decoded features are fused by concatenation as the final decoded feature D2 of this decoding module.

[0050] Step S1050: Decode the feature through the integrated decoding module 1 with the input of the skip connection feature E1 and the guiding features E0 and E4

[0051] In this step, the integrated decoding module 1 is designed as the most complex third decoding module, as Figure 3 shown. Among them, the skip connection sub-module is specifically designed as the first method as Figure 4 , the segmentation sub-module is specifically designed as the second method as Figure 5 , and the matting sub-module is specifically designed as the second method as Figure 6 . First, the skip connection sub-module performs feature transformation and fusion on the input decoded feature D2 and the skip connection feature E1 to obtain an enhanced feature. Then, in the segmentation sub-module, the enhanced feature is first partially decoded and then the guiding feature from the encoded feature E4 with high-level semantic information is added, and then further decoded and upsampled to obtain the decoded feature for the segmentation task. At the same time, in the matting sub-module, the enhanced feature is first partially decoded and then the guiding feature from the encoded feature E0 with low-level texture information is added, and then further decoded and upsampled to obtain the decoded feature for the matting task. Finally, the two decoded features are fused by concatenation as the final decoded feature D1 of this decoding module.

[0052] Step S1060: Decode the feature through the integrated decoding module 0 with the input of the guiding features E0 and E4.

[0053] In this step, the integrated decoding module 0 is designed as the second decoding module without a skip connection sub-module, as Figure 3 shown. Among them, the segmentation sub-module is specifically designed as the second method as Figure 5 , and the matting sub-module is specifically designed as the second method as Figure 6The second method. First, the input feature is the decoded feature D1; then in the segmentation sub-module, the input feature is partially decoded first, and then the guiding feature from the encoded feature E4 with high-level semantic information is added, and then further decoded and upsampled to obtain the decoded feature for the segmentation task; at the same time, in the matting sub-module, the input feature is partially decoded first, and then the guiding feature from the encoded feature E0 with low-level texture information is added, and then further decoded and upsampled to obtain the decoded feature for the matting task; finally, the two decoded features are fused by concatenation as the final decoded feature D0 of this decoding module.

[0054] Step S1070: Use the segmentation head module to predict the segmentation result.

[0055] In this step, the segmentation task is defined as a three-class classification task, the input feature is the decoded feature D0, and the segmentation head module composed of convolution operations is directly used to estimate the probability of each pixel belonging to each classification as the prediction result of the segmentation, that is

[0056] Step S1080: Use the matting head module to predict the matting result.

[0057] In this step, the matting task is defined as a transparency regression task, the input feature is the decoded feature D0, and the matting head module is a combination of a convolution operation and an activation function to achieve the dense prediction of the transparency of each pixel, and finally the prediction result of the transparency, that is, α M . The activation function can be the sigmoid function.

[0058] Step S1090: Calculate the loss of the current prediction result according to the defined loss function.

[0059] In this step, the input is the segmentation prediction result and the matting prediction result, and the defined loss function is the weighted sum of all defined segmentation loss functions and matting loss functions. The loss of the current prediction result is calculated according to this defined loss function.

[0060] For the segmentation task, given the segmentation prediction result and the ground truth segmentation label image G ∈ R 3 ×H×W, the cross-entropy loss function L CE is used to calculate the classification loss between the two, that is:

[0061]

[0062] where c represents the number of segmentation classification categories, H is the height of the input image, and W is the width of the input image.

[0063] For the matting task, given the matting prediction result α M∈R 1 ×H×W and the true transparency image α ∈ R 1 ×H×W. On the one hand, an alpha loss function is defined in the whole image region, and the Laplacian loss function and the composition loss function are combined to calculate the loss of the matting task on the whole image; at the same time, the alpha loss function and the Laplacian loss function are used in the uncertain image region to calculate the loss of the local region to further optimize the details. The specific definitions are as follows:

[0064] (1) The alpha loss function L of the whole image region α is defined as the root mean square error between the predicted transparency and the true transparency corresponding to all pixels in the whole image region:

[0065]

[0066] where i refers to the pixel index of the whole image, h is the height of the input image, w is the width of the input image, and ε = 10 -6 is a very small value to ensure the stability of loss calculation.

[0067] (2) The Laplacian loss function L of the whole image region 1ap is defined as the L1 distance between the predicted transparency and the true transparency corresponding to all pixels in the whole image region on multiple Laplacian pyramid images:

[0068]

[0069] where i refers to the pixel index of the whole image, and Lap k represents the k-th layer Laplacian pyramid image.

[0070] (3) The composition loss function L of the whole image region comp is defined as the root mean square error between the RGB image generated from the predicted transparency, the foreground image and the background image and the true RGB image

[0071]

[0072] corresponding to all pixels in the whole image region: -6 where i refers to the pixel index of the whole image, h is the height of the input image, w is the width of the input image, and ε = 10

[0073] (4) The alpha loss function in the uncertain region Defined as the root mean square error between the predicted transparency and the true transparency corresponding to all pixels within the uncertain region:

[0074]

[0075] where i refers to the pixel index of the entire image, indicates whether the current pixel belongs to the uncertain region, h is the height of the input image, w is the width of the input image, and ε = 10 -6 is a very small value to ensure the stability of loss calculation.

[0076] (5) Laplacian loss function for the uncertain region Defined as the L1 distance between the predicted transparency and the true transparency corresponding to all pixels within the uncertain region on multiple Laplacian pyramid images:

[0077]

[0078] where i refers to the pixel index of the entire image, Lap k represents the k-th layer Laplacian pyramid image, indicates whether the current pixel belongs to the uncertain region.

[0079] Finally, define the final loss function L as the weighted sum of all defined segmentation loss functions and matting loss functions:

[0080]

[0081] where λ s = 1, λ m = 1, λ mu = 2 are weight parameters.

[0082] Step S1100: Update the parameters of the entire neural network according to the calculated loss.

[0083] In this step, update the parameters of the entire neural network using the backpropagation algorithm according to the loss calculated in step S1090.

[0084] Step S1110: Determine whether to end the training process.

[0085] In this step, it can be judged whether to end the training through some preset thresholds, such as whether the current loss is less than the given threshold, or whether the current training iteration number reaches the given maximum training cycle number. If the conditions are met, end the training of the network model and enter step S1120; otherwise, return to step S1010 and continue the next training iteration process.

[0086] The training process of the neural network model is a cyclic and repetitive process. Each training includes three processes: forward propagation, backward propagation, and parameter update. Among them, the forward propagation process described in this disclosure can be a known forward propagation process, and the quantization process of weights and feature maps of any bit can be included in the forward propagation process, and this disclosure does not limit this. If the difference between the actual output result and the expected output result of the neural network model does not exceed a predetermined threshold, it means that the weights in the neural network model are optimal solutions, and the performance of the trained neural network model has reached the expected performance, and the training of the neural network model is completed. On the contrary, if the difference between the actual output result and the expected output result of the neural network model exceeds the predetermined threshold, the backward propagation process needs to be continued, that is, based on the difference between the actual output result and the expected output result, operations are performed layer by layer from bottom to top in the neural network model to update the parameters in the model, so that the performance of the network model after weight update is closer to the expected performance.

[0087] The neural network model applicable to the present invention can be any known model, such as a convolutional neural network model, a recurrent neural network model, and a graph neural network model, etc. The present invention does not limit the type of the network model.

[0088] The computing precision applicable to the neural network model of the present disclosure can be any precision, both high precision and low precision are acceptable. The terms "high precision" and "low precision" are relative levels of precision and do not limit specific values. For example, high precision can be 32-bit floating point type, and low precision can be 1-bit fixed point type. Of course, other precisions such as 16-bit, 8-bit, 4-bit, and 2-bit are also included in the computing precision range applicable to the solution of the present disclosure. The term "computing precision" can refer to the precision of the weights in the neural network model or the precision of the input to be trained. This disclosure does not limit this. The neural network model described in this disclosure can be a binary neural network model (BNNs), and of course, it is not limited to neural network models with other computing precisions.

[0089] Step S1012: Output the trained convolutional neural network model.

[0090] In this step, the current parameters of all layers in the convolutional neural network structure are the trained network model, and the network model and the corresponding parameter information can be output.

[0091] According to the technical solution of the exemplary embodiment of the present invention, taking portrait segmentation as a simplification of portrait matting, an integrated decoding network is designed to simultaneously decode the corresponding information of segmentation and matting, and more refined portrait segmentation results and portrait matting results can be obtained simultaneously.

[0092] <Variant Example 1>

[0093] The following will refer to Figure 8Describe this exemplary embodiment. The integrated decoder 2 involved in this exemplary embodiment consists of 5 integrated decoding modules with different guiding feature inputs and a matte head module. Compared with the previous embodiment, in this exemplary embodiment, the designed integrated decoding network has only one head module for regressing the human portrait transparency, which simplifies the model training and prediction processes. The difference lies in including step S2080.

[0094] Step S2080: Calculate the loss of the current prediction result according to the defined loss function.

[0095] In this step, the input is the matte prediction result, and the defined loss function is the weighted sum of all defined matte loss functions. Calculate the loss of the current prediction result according to this defined loss function.

[0096] For the matte task, given the matte prediction result α M ∈R 1 ×H×W and the ground truth transparency image α ∈ R 1 ×H×W, on the one hand, define the alpha loss function, Laplacian loss function, and composition loss function in the whole image area to calculate the loss of the matte task on the whole image; at the same time, use the alpha loss function and Laplacian loss function in the uncertain image area to calculate the loss of the local area to further optimize the details. Its specific definition is the same as S1090. Please refer to the introduction of S1090. The defined final loss function L is the weighted sum of all defined matte loss functions:

[0097]

[0098] where, λ m = 1, λ mu = 2 are weight parameters.

[0099] According to the method of this exemplary embodiment, the model training and prediction processes can be simplified. The remaining steps are similar to those of S2010 - 2070 and S2090 - 2110 in the first exemplary embodiment, and will not be elaborated here.

[0100] <Variant Example 2>

[0101] The integrated decoder involved in this exemplary embodiment consists of 5 integrated decoding modules with different guiding feature inputs and three head modules respectively for segmentation, matte, and fusion results. As Figure 9 shown, the specific description is as follows.

[0102] Step S3010: Extract the encoding features of each layer from the training data using the encoding network. The training data can be from the matte training database. This step can adopt a method similar to that in Step S1010.

[0103] Step S3020: Decode the features through the integrated decoding module 4 without skip features and guidance features. This step can adopt a method similar to that in Step S1020.

[0104] Step S5030: Decode the features through the integrated decoding module 3 with the input of skip feature E3 and the inputs of guidance features E0 and E4.

[0105] In this step, the integrated decoding module 3 is designed as the most complex third decoding module, as Figure 3 shown. Among them, the skip sub-module is specifically designed as a skip sub-module based on concatenation, as Figure 4 shown in the second way; the segmentation sub-module is specifically designed as a segmentation sub-module guided by concatenated encoding features, as Figure 5 shown in the third way; the matte sub-module is specifically designed as a matte sub-module guided by concatenated encoding features, as Figure 6 shown in the third way. First, the skip sub-module performs feature transformation and concatenation on the input decoded feature D4 and the skip feature E3 to obtain enhanced features; then in the segmentation sub-module, the enhanced features are first partially decoded and then concatenated with the guidance features from the encoding feature E4 with high-level semantic information, and then further decoded and upsampled to obtain the decoded features for the segmentation task; at the same time, in the matte sub-module, the enhanced features are first partially decoded and then concatenated with the guidance features from the encoding feature E0 with low-level texture information, and then further decoded and upsampled to obtain the decoded features for the matte task; finally, the two decoded features are fused by concatenation as the final decoded feature D3 of this decoding module.

[0106] Step S3040: Decode the features through the integrated decoding module 2 with the input of skip feature E2 and the inputs of guidance features E0 and E4. The specific processing of this step is similar to that in Step S3030.

[0107] Step S3050: Decode the features through the integrated decoding module 1 with the input of skip feature E1 and the inputs of guidance features E0 and E4. The specific processing of this step is similar to that in Step S3030.

[0108] Step S3060: Decode the features through the integrated decoding module 0 without skip features and guidance features. The specific processing of this step is similar to that in Step S1020.

[0109] Step S3070: Predict the segmentation result using the segmentation head module. This step can adopt a method similar to that in Step S1070.

[0110] Step S3080: Predict the matte result using the matte head module. This step can adopt a method similar to that in Step S1080.

[0111] Step S3090: Generate a fused matte result based on the segmentation prediction result and the matte prediction result.

[0112] In this step, the inputs are the segmentation prediction result and the matte prediction result. Based on the segmentation prediction result, the transparency of the regions segmented as foreground is set to 1, the transparency of the regions segmented as background is set to 0, and the transparency of other regions is the corresponding matte prediction result, thereby generating the final fused matte result α. F 。

[0113] Step S3100: Calculate the loss of the current prediction result according to the defined loss function.

[0114] In this step, the inputs are the segmentation prediction result, the matte prediction result, and the fused matte result of Steps S3070 - 3090. The defined loss function is the weighted sum of all defined segmentation loss functions and matte loss functions, and the loss of the current prediction result is calculated according to this defined loss function. The specific calculation process is described below.

[0115] For the segmentation task, the defined loss function is the same as that in S1090.

[0116] For the matte task, given the fused matte result α F ∈R 1 ×H×W and the ground - truth transparency image α ∈ R 1 =H×W, a combination of the alpha loss function, the Laplacian loss function, and the composition loss function is defined in the whole - image region to calculate the loss of the matte task on the entire image. The specific definitions are as follows:

[0117] (1) The alpha loss function L of the whole - image region α is defined as the root - mean - square error between the predicted transparency and the ground - truth transparency corresponding to all pixels in the entire image region:

[0118]

[0119] where i refers to the pixel index of the whole image, h is the height of the input image, w is the width of the input image, and ε = 10 -6 is a very small value to ensure the stability of loss calculation.

[0120] The Laplacian loss function L of the whole - image region lapDefine the L1 distance between the fusion transparency and the true transparency corresponding to all pixels in the entire image area on multiple Laplacian pyramid images:

[0121]

[0122] where i refers to the pixel index of the entire image, and Lap k represents the k-th layer Laplacian pyramid image.

[0123] (3) The composition loss function L for the entire image area comp Define the RGB image generated from the fusion transparency, the foreground image, and the background image and the true RGB image as the root mean square error between all pixels in the entire image area:

[0124]

[0125] where i refers to the pixel index of the entire image, h is the height of the input image, w is the width of the input image, and ε = 10 -6 is a very small value to ensure the stability of loss calculation.

[0126] For the matting task, given the matting prediction result α M ∈R 1 ×H×W and the true transparency image α ∈ R 1 ×H×W, only use the alpha loss function and the Laplacian loss function in the uncertain image area to calculate the loss of the local area to further optimize the details. The specific definitions are as follows:

[0127] (1) The alpha loss function in the uncertain area Define it as the root mean square error between the predicted transparency and the true transparency corresponding to all pixels in the uncertain area:

[0128]

[0129] where i refers to the pixel index of the entire image, indicates whether the current pixel belongs to the uncertain area, h is the height of the input image, w is the width of the input image, and ε = 10 -6 is a very small value to ensure the stability of loss calculation.

[0130] (2) The Laplacian loss function in the uncertain area Define it as the L1 distance between the predicted transparency and the true transparency corresponding to all pixels in the uncertain area on multiple Laplacian pyramid images:

[0131]

[0132] Among them, i refers to the pixel index of the entire image, and Lap k represents the Laplacian pyramid image of the k-th layer, indicating whether the current pixel belongs to the uncertain region.

[0133] Finally, define the final loss function L as the weighted sum of all defined segmentation loss functions and matting loss functions:

[0134]

[0135] Among them, λ s = 1, λ f = 1, λ m = 2 are weight parameters.

[0136] Step S5110: Update the parameters of the entire neural network according to the calculated loss. The specific processing of this step is similar to that in step S1100.

[0137] Step S3120: Determine whether to end the training process. The specific processing of this step is similar to that in step S1110.

[0138] Step S3130: Output the trained convolutional neural network model. The specific processing of this step is similar to that in step S1120.

[0139] According to this exemplary embodiment, the human portrait segmentation is regarded as an object estimation equally important as the human portrait matting, and the parsing features of the integrated decoding module 0 are simplified. Since the human portrait segmentation is a relatively simple task, even with the simplified integrated decoding module 0, accurate segmentation results can be obtained through training. Therefore, an integrated decoding network is designed to simultaneously decode the corresponding information of segmentation and matting, and fuse the two to generate the final result with both segmentation and matting.

[0140] <Variant 3>

[0141] The following will refer to Figure 10 to describe this exemplary embodiment. The integrated decoder of this exemplary embodiment consists of 5 integrated decoding modules with different guiding feature inputs and three head modules respectively for segmentation, matting, and fusion results, as Figure 10 shown, and the specific description is as follows.

[0142] Step S4010: Use the encoding network to extract the encoding features of each layer from the training data. This step can adopt a method similar to that in step S1010.

[0143] Step S4020: Decode the features through the integrated decoding module 4 without skip connection and guidance features. This step can adopt a method similar to that in step S1020.

[0144] Step S4030: Decode the features through the integrated decoding module 3 with skip feature E3 input and guidance features E0 and E4 input.

[0145] In this step, the integrated decoding module 3 is designed as the fourth decoding module as shown in Figure 3 wherein, the segmentation sub-module is specifically designed as a segmentation sub-module guided by concatenated encoded features, as shown in the third way in Figure 5 ; the matte extraction sub-module is specifically designed as a matte extraction sub-module guided by concatenated encoded features, as shown in the third way in Figure 6 ; the skip connection sub-module is specifically designed as a skip connection sub-module based on concatenation, as shown in the second way in Figure 4 First, in the segmentation sub-module, the input feature D4 is partially decoded and then the guidance feature from the encoded feature E4 with high-level semantic information is added, and then further decoded and upsampled to obtain the decoded feature for the segmentation task; at the same time, in the matte extraction sub-module, the input feature D4 is partially decoded and then the guidance feature from the encoded feature E0 with low-level texture information is added, and then further decoded and upsampled to obtain the decoded feature for the matte extraction task; then the two decoded features are fused into the intermediate decoded feature of the decoding module 3 by concatenation; finally, the skip connection sub-module performs feature transformation and fusion on the input intermediate decoded feature and the skip feature E3 to obtain the enhanced decoded feature D3.

[0146] Step S4040: Decode the features through the integrated decoding module 2 with skip feature E2 input and guidance features E0 and E4 input. The specific processing of this step is similar to that in step S4030.

[0147] Step S4050: Decode the features through the integrated decoding module 1 with skip feature E1 input and guidance features E0 and E4 input. The specific processing of this step is similar to that in step S4030.

[0148] Step S4060: Decode the features through the integrated decoding module 0 with skip feature E0 input and guidance features E0 and E4 input. The specific processing of this step is similar to that in step S4030.

[0149] Step S4070: Predict the segmentation result using the segmentation head module. The specific processing of this step is similar to that in step S1070.

[0150] Step S4080: Predict the matte extraction result using the matte extraction head module. The specific processing of this step is similar to that in step S1080.

[0151] Step S4090: Generate a fused matte result based on the segmentation prediction result and the matte prediction result. The specific processing of this step is similar to that in step S3090.

[0152] Step S4100: Calculate the loss of the current prediction result according to the defined loss function. The specific processing of this step is similar to that in step S3100.

[0153] Step S4110: Update the parameters of the entire neural network according to the calculated loss. The specific processing of this step is similar to that in step S1010.

[0154] Step S4120: Determine whether to end the training process. The specific processing of this step is similar to that in step S1110.

[0155] Step S4130: Output the trained convolutional neural network model. The specific processing of this step is similar to that in step S1120.

[0156] According to this exemplary embodiment, the human segmentation is regarded as an object estimation as important as the human matte extraction, and the parsing features of the complex integrated decoding module 0 are adopted, so that the matte extraction can obtain more low-level texture information to decode more refined features. Therefore, an integrated decoding network is designed to decode the corresponding information of segmentation and matte extraction simultaneously, and fuse the two to generate a refined result with both segmentation and matte extraction.

[0157] <Variant Example 4>

[0158] The following will refer to Figure 11 describe this exemplary embodiment. The integrated decoder of this exemplary embodiment is composed of 5 integrated decoding modules with different guiding feature inputs and three head modules for segmentation, matte extraction, and depth estimation respectively.

[0159] Step S5010: Use the encoding network to extract encoded features at each level from the training data.

[0160] In this step, the input is a part of the training data randomly selected from the matte extraction database. The encoding network can select existing multi-layer neural networks such as ResNet, Transformer, and MLP. When the training data is input, the encoding network will generate five levels of encoded features with gradually decreasing resolutions, named E0, E1, E2, E3, and E4 respectively.

[0161] Step S5020: Decode the features through the integrated decoding module 4 without skip connections and guiding features.

[0162] In this step, the integrated decoding module 4 is designed to be similar to the simplest first decoding module, such as Figure 3As shown, there is no skip connection sub-module, but the feature decoding is completed in parallel by the segmentation sub-module, the matting sub-module, and the depth sub-module for their respective tasks, and then the decoded features are concatenated into the final decoded feature. Among them, the segmentation sub-module is designed as a series connection of at least two convolutional operations and one upsampling operation, as shown in Figure 5 the first design of Figure 6 ; the matting sub-module also adopts the same structure, as shown in Figure 6 the first design of

[0163] ; the depth sub-module also adopts the same structure, as shown in

[0164] the first design of Figure 3 . When the final encoded feature E4 obtained from the encoding network is input, the segmentation sub-module, the matting sub-module, and the depth sub-module respectively decode the features required for their respective tasks, and finally fuse the three decoded features by concatenation as the final decoded feature D4 of this decoding module. Figure 4 Figure 5 Figure 6 Figure 6

[0165]

[0166] Step S5030: Decode the feature through the integrated decoding module 3 with the input of the skip feature E3 and the guiding features E0 and E4.

[0164] In this step, the integrated decoding module 3 is designed similar to the most complex third decoding module, as shown in Figure 3 . Among them, the skip connection sub-module is specifically designed in the first way as shown in Figure 4 ; the segmentation sub-module is specifically designed in the second way as shown in Figure 5 ; the matting sub-module is specifically designed in the second way as shown in Figure 6 ; the depth sub-module is the same as the matting sub-module and is specifically designed in the second way as shown in Figure 6 . First, the skip connection sub-module performs feature transformation and fusion on the input decoded feature D4 and the skip feature E3 to obtain enhanced features; then in the segmentation sub-module, the enhanced features are first partially decoded and then the guiding features from the encoded feature E4 with high-level semantic information are added, and then further decoded and upsampled to obtain the decoded feature for the segmentation task; at the same time, in the matting sub-module, the enhanced features are first partially decoded and then the guiding features from the encoded feature E0 with low-level texture information are added, and then further decoded and upsampled to obtain the decoded feature for the matting task; at the same time, in the depth sub-module, the enhanced features are first partially decoded and then the guiding features from the encoded feature E0 with low-level texture information are added, and then further decoded and upsampled to obtain the decoded feature for the depth task; finally, the three decoded features are fused by concatenation as the final decoded feature D3 of this decoding module.

[0165] Step S5040: Decode the feature through the integrated decoding module 2 with the input of the skip feature E2 and the guiding features E0 and E4. The specific processing of this step is similar to that in step S5030.

[0166] Step S5050: Decode the features through the integrated decoding module 1 with the skip connection feature E1 input and the guidance features E0 and E4 inputs. The specific processing of this step is similar to that in step S5030.

[0167] Step S5060: Decode the features through the integrated decoding module 0 with the guidance features E0 and E4 inputs.

[0168] In this step, the integrated decoding module 0 is designed as a similar second decoding module without a skip sub-module, as Figure 3 shown. Among them, the segmentation sub-module is specifically designed in the second way as Figure 5 shown, the matte extraction sub-module is specifically designed in the second way as Figure 6 shown, and the depth sub-module is the same as the matte extraction sub-module, specifically designed in the second way as Figure 6 shown. First, the input feature is the decoded feature D1; then in the segmentation sub-module, the input feature is partially decoded first and then the guidance feature from the encoded feature E4 with high-level semantic information is added, and then further decoded and upsampled to obtain the decoded feature for the segmentation task; at the same time, in the matte extraction sub-module, the input feature is partially decoded first and then the guidance feature from the encoded feature E0 with low-level texture information is added, and then further decoded and upsampled to obtain the decoded feature for the matte extraction task; at the same time, in the depth sub-module, the input feature is partially decoded first and then the guidance feature from the encoded feature E0 with low-level texture information is added, and then further decoded and upsampled to obtain the decoded feature for the depth task; finally, the three decoded features are fused by concatenation as the final decoded feature D0 of this decoding module.

[0169] Step S5070: Predict the segmentation result using the segmentation head module. The specific processing of this step is similar to that in step S1070.

[0170] Step S5080: Predict the matte extraction result using the matte extraction head module. The specific processing of this step is similar to that in step S1080.

[0171] Step S5090: Predict the depth result using the depth head module. The specific processing of this step is similar to that in step S1080.

[0172] Step S5100: Calculate the loss of the current prediction result according to the defined loss function.

[0173] In this step, the inputs are the segmentation prediction result and the matte extraction prediction result, and the defined loss function is the weighted sum of all defined segmentation loss functions, matte extraction loss functions, and depth loss functions. Calculate the loss of the current prediction result according to this defined loss function. Among them, the definitions of the segmentation loss function and the matte extraction loss function are the same as those in step S1090.

[0174] For the depth estimation task, given the depth prediction result d p ∈R 1 ×H×W and the ground truth depth image d ∈ R 1 ×H×W, the alpha loss function and the Laplacian loss function are used to calculate the loss of depth estimation within the entire image region.

[0175] Finally, the final loss function L is defined as the weighted sum of all defined segmentation loss functions, matting loss functions, and depth loss functions:

[0176]

[0177] where λ s = 1, λ m = 1, λ mu = 2 are weight parameters.

[0178] Step S5110: Update the parameters of the entire neural network according to the calculated loss. The specific processing of this step is similar to that in step S1100.

[0179] Step S5120: Determine whether to end the training process. The specific processing of this step is similar to that in step S1110.

[0180] Step S5130: Output the trained convolutional neural network model. The specific processing of this step is similar to that in step S1120.

[0181] According to this exemplary embodiment, depth estimation is also used as a task objective, and an integrated decoding network is designed to simultaneously decode the corresponding information of segmentation, matting, and depth, and finally obtain the portrait segmentation result, portrait matting result, and depth estimation result.

[0182] According to the method of the exemplary embodiment of the present invention, based on the encoding-integrated decoding network of the U-Net structure, rich features that can simultaneously express segmentation and matting are parsed from the encoded features, thereby improving the accuracy of portrait matting.

[0183] Specifically, as Figure 2 shown, the integrated decoding network according to the exemplary embodiment of the present invention is composed of a cascade of multiple integrated decoder modules. Each decoder module uses different features from the encoder to parse out the integrated features, and the final integrated features can be respectively used for the segmentation task and the matting task.

[0184] In the integrated decoding network according to the exemplary embodiment of the present invention, as Figure 3As shown, different integrated decoding modules are designed for different input encoding features. Specifically, the input features are first optionally enhanced by a skip sub-module and specific features from the encoder, and then the enhanced feature maps are respectively subjected to feature transformation by a segmentation sub-module and a matting sub-module to restore their respective different features, and finally integrated into the final output features. Among them, the segmentation sub-module and the matting sub-module also optionally use different encoding features to guide the feature transformation. In the integrated decoding network according to an exemplary embodiment of the present invention, the integrated decoding module 4 uses the Figure 3 first design shown in, the integrated decoding module 0 uses the second design, and the integrated decoding modules 1, 2, and 3 use the third design.

[0185] In the integrated decoding module according to an exemplary embodiment of the present invention, for different task objectives, in order to be able to effectively decode the features required for different task objectives, a skip sub-module, a segmentation sub-module, and a matting sub-module are further designed. As Figure 4 shown, the optional skip sub-module uses specific encoding features skipped from the encoder and the current input features for feature transformation and fusion to obtain enhanced input features. In the segmentation sub-module, the optional encoding feature E4 comes from the output of the bottom layer of the encoder, representing the high-level semantic information that the encoder can obtain, and can provide important information for the segmentation task. Therefore, the segmentation sub-module adds the guiding feature from the encoding feature E4 after partially decoding the input features, and then further decodes and up-samples to obtain the final output features, as Figure 5 shown. In the matting sub-module, the optional encoding feature E0 comes from the output of the top layer of the encoder, representing the low-level texture information that the encoder can obtain, and can provide important detail information for the matting task. Therefore, the matting sub-module adds the guiding feature from the encoding feature E0 after partially decoding the input features, and then further decodes and up-samples to obtain the final output features, as Figure 6 shown.

[0186] In addition, for the two different tasks of segmentation and matting, it is necessary to obtain the prediction results of their respective tasks from the decoded features. Therefore, a segmentation head module and a matting head module are respectively designed. Since the segmentation task is actually a three-class classification task, a convolution operation is directly used to form the segmentation head module to realize the probability estimation of each pixel belonging to each classification. The matting task is actually a regression problem, and the matting head module is designed as a combination of a convolution operation and an activation function to complete the dense prediction of the transparency of each pixel.

[0187] Furthermore, for the segmentation task, a common cross-entropy loss function is used to calculate the classification loss. For the matting task, in the full-image region, a combination of a common alpha loss function, a Laplacian loss function, and a composition loss function is used to calculate the full-image loss; meanwhile, in the uncertain region, the alpha loss function and the Laplacian loss function are used to further optimize the details.

[0188] According to this exemplary embodiment, the integrated decoding module is divided into two branch modules for segmentation and matting to decode different feature representations, and then the two are merged into integrated features to play a role of influencing each other. The integrated decoding module according to this exemplary embodiment can choose to utilize skip features from the encoding module to enhance the intermediate integrated decoding features, which can be beneficial for information recovery in both segmentation and matting. The segmentation and matting decoding modules according to this exemplary embodiment can use smaller channel numbers to recover information without affecting the accuracy.

[0189] The performance of the present invention and the prior art will be compared through experiments below.

[0190] Experiment: Validation is performed on the P3M-10k training set

[0191] Training set: P3M-10k (9421 images) and 906 available images selected from other datasets.

[0192] Test sets: RealWorldPortrait_636, P3M_500_NP, P3M_500_P

[0193] Evaluation criteria: MSE (10-3, full image), SAD (full image), MSE (10-3, uncertain region), SAD (uncertain region), model size (MB), MACs (GB), average MACs (KB).

[0194] Convolutional neural network architecture: RestNet34-mp

[0195] Comparison with the prior art: Prior art (P3M-Net)

[0196] Experimental results:

[0197] Table 1 shows the comparison of the accuracy results between the present invention and the prior art on a general dataset.

[0198] Table 1

[0199]

[0200] Table 2 shows the comparison of the model scales obtained by training the present invention and the prior art.

[0201] Table 2

[0202]

[0203] Experimental results show that, compared with the prior art, according to the exemplary embodiments of the present invention, taking portrait segmentation as a simplification of portrait matting, an integrated decoding module is designed to decode the corresponding information of segmentation and matting simultaneously, and finally a more refined portrait matting result can be obtained.

[0204] <Second Exemplary Embodiment>

[0205] Based on the foregoing first exemplary embodiment, the second exemplary embodiment of the present invention describes a network model training system. The training system includes a terminal, a communication network, and a server. The terminal and the server communicate with each other through the communication network. The server uses the network model stored locally to train the network model stored in the terminal online, so that the terminal can use the trained network model for real-time services. The following describes each part in the training system of the second exemplary embodiment of the present invention.

[0206] The terminal in the training system can be an embedded image acquisition device such as a security camera, or a device such as a smart phone or a PAD. Of course, the terminal can also not be a terminal with relatively weak computing power such as an embedded device, but other terminals with relatively strong computing power. The number of terminals in the training system can be determined according to actual needs. For example, if the training system is to train the security cameras in a shopping mall, all the security cameras in the shopping mall can be regarded as terminals. At this time, the number of terminals in the training system is fixed. For another example, if the training system is to train the smart phones of users in a shopping mall, the smart phones accessing the wireless local area network of the shopping mall can be regarded as terminals. At this time, the number of terminals in the training system is not fixed. In the second exemplary embodiment of the present invention, the type and number of terminals in the training system are not limited, as long as the terminal can store and train a network model.

[0207] The server in the training system can be a high-performance server with relatively strong computing power, such as a cloud server. The number of servers in the training system can be determined according to the number of terminals it serves. For example, if the number of terminals to be trained in the training system is small or the geographical range of the terminal distribution is small, the number of servers in the training system is small, such as only one server. If the number of terminals to be trained in the training system is large or the geographical range of the terminal distribution is large, the number of servers in the training system is large, such as establishing a server cluster. In the second exemplary embodiment of the present invention, the type and number of servers in the training system are not limited, as long as the server can store at least one network model and provide information for training the network model stored in the terminal.

[0208] The communication network in the second exemplary embodiment of the present invention is a wireless network or a wired network for realizing information transmission between a terminal and a server. Any network currently available for up / downlink transmission between a network server and a terminal can be used as the communication network in this embodiment. The second exemplary embodiment of the present invention does not limit the type and communication method of the communication network. Of course, the second exemplary embodiment of the present invention is not limited to other communication methods either. For example, a third-party storage area is allocated for this training system. When the terminal and the server need to transmit information to each other, the information to be transmitted is stored in the third-party storage area, and the terminal and the server regularly read the information in the third-party storage area to realize information transmission between the two.

[0209] The following will Figure 12 , in conjunction with, describe in detail the online training process of the training system according to the second exemplary embodiment of the present invention. Figure 12 Fig. shows an example of the training system. It is assumed that the training system includes a terminal and a server. The terminal can perform real-time shooting. It is assumed that the terminal stores a network model that can be trained and can process pictures, and the same network model is stored in the server. The training process of the training system is described as follows.

[0210] Step S201: The terminal sends a training request to the server via the communication network.

[0211] The terminal sends a training request to the server via the communication network. This request includes information such as the terminal identifier. The terminal identifier is information that uniquely represents the identity of the terminal (for example, the ID or IP address of the terminal, etc.).

[0212] This step S201 is described by taking a single terminal sending a training request as an example. Of course, multiple terminals can also send training requests in parallel. The processing process for multiple terminals is similar to that for a single terminal and will not be elaborated here.

[0213] Step S202: The server receives the training request.

[0214] In Figure 12 the shown training system, only one server is included. Therefore, the communication network can transmit the training request initiated by the terminal to this server. If the training system includes multiple servers, the training request can be transmitted to a relatively idle server according to the idle status of the servers.

[0215] Step S203: The server responds to the received training request.

[0216] The server determines the terminal that initiated the request based on the terminal identifier included in the received training request, and then determines the network model to be trained stored in the terminal. One optional method is that the server determines the network model to be trained stored in the terminal that initiated the request based on a comparison table of terminals and network models to be trained; another optional method is that the training request contains information about the network model to be trained, and the server can determine the network model to be trained based on the information. Here, determining the network model to be trained includes but is not limited to determining the network architecture, hyperparameters, and other information that characterizes the network model.

[0217] After the server determines the network model to be trained, the method of the first exemplary embodiment of the present invention can be used to train the network model stored in the terminal that initiates the request using the same network model stored locally in the server. Specifically, the server updates the weights in the network model locally according to the method in the first exemplary embodiment, and transmits the updated weights to the terminal, so that the terminal synchronizes the network model to be trained stored in the terminal according to the received updated weights. Here, the network model in the server and the network model to be trained in the terminal can be the same network model, or the network model in the server is more complex than the network model in the terminal, but the outputs of the two are close. The present disclosure does not limit the types of network models used for training in the server and the network models to be trained in the terminal, as long as the updated weights output from the server can synchronize the network model in the terminal, so that the output of the synchronized network model in the terminal is closer to the expected output.

[0218] exist Figure 12 In the training system shown, the terminal actively initiates the training request. Optionally, the second exemplary embodiment of the present invention is not limited to the server broadcasting an inquiry message and then the terminal responding to the inquiry message to perform the above training process.

[0219] Through the training system described in the second exemplary embodiment of the present invention, the server can perform online training on the network model in the terminal, which improves the flexibility of training; at the same time, it also greatly enhances the business processing capabilities of the terminal and expands the business processing scenarios of the terminal. The above second exemplary embodiment describes the training system by taking online training as an example, but the present invention is not limited to the offline training process, which will not be repeated here.

[0220] <Third Exemplary Embodiment>

[0221] The third exemplary embodiment of the present invention describes a training device for a neural network model, which can execute the training method described in the first exemplary embodiment, and when the device is applied in an online training system, it can be the device in the server described in the second exemplary embodiment.

[0222] The training device of this embodiment also has modules that implement the functions of the server in the training system, such as the function of identifying the received data, the data encapsulation function, the network communication function, etc., which will not be elaborated here.

[0223] Other embodiments

[0224] The present invention can be used in many applications. For example, the present invention can be used to monitor, identify, and track objects in static images or moving videos captured by a camera, and is particularly advantageous for portable devices equipped with a camera, (camera-based) mobile phones, and the like.

[0225] It should be noted that the methods and devices described herein can be implemented as software, firmware, hardware, or any combination thereof. Some components can be implemented, for example, as software running on a digital signal processor or a microprocessor. Other components can be implemented, for example, as hardware and / or an application-specific integrated circuit.

[0226] In addition, the methods and systems of the present invention can be implemented in various ways. For example, the methods and systems of the present invention can be implemented by software, hardware, firmware, or any combination thereof. The order of the steps of the method described above is merely illustrative, and unless otherwise specifically stated, the steps of the method of the present invention are not limited to the order specifically described above. Furthermore, in some embodiments, the present invention can also be embodied as a program recorded in a recording medium, including machine-readable instructions for implementing the method according to the present invention. Therefore, the present invention also encompasses a recording medium storing a program for implementing the method according to the present invention.

[0227] Those skilled in the art should be aware that the boundaries between the above operations are merely illustrative. Multiple operations can be combined into a single operation, a single operation can be distributed among additional operations, and operations can be performed at least partially overlapping in time. Moreover, alternative embodiments can include multiple instances of a particular operation, and the order of operations can be changed in various other embodiments. However, other modifications, variations, and substitutions are also possible. Therefore, this specification and the drawings should be regarded as illustrative rather than restrictive.

[0228] Although some specific embodiments of the present invention have been described in detail by way of examples, those skilled in the art should understand that the above examples are only for illustration and not for limiting the scope of the present invention. The embodiments of this invention can be combined arbitrarily without departing from the spirit and scope of the present invention. Those skilled in the art should also understand that various modifications can be made to the embodiments without departing from the scope and spirit of the present invention.

Claims

1. An image processing method, the method comprises: an encoding step of generating a plurality of encoded features with different resolutions based on an input image and an encoder; a decoding step of decoding, by a decoder cascaded with the plurality of encoded features and a plurality of decoding modules, a decoded feature having the same resolution as the input image; a prediction step of predicting, based on the decoded feature and a head module, an output image having the same resolution as the input image.

2. The method according to claim 1, wherein the output image can be a transparency image, a segmentation image, a depth image, or a combination thereof.

3. The method according to claim 1, wherein the encoder can be a multi-layer neural network capable of generating encoded features with gradually decreasing resolutions, wherein, the multi-layer neural network can be ResNet, Transformer, or MLP.

4. The method according to claim 1, wherein the plurality of decoding modules of the decoder respectively generate decoded features having resolutions consistent with the encoded features generated in the corresponding encoding steps.

5. The method according to claim 1, each of the decoding modules has at least one input feature, wherein, the at least one input feature comes from the output features of other decoding modules or the output features of the encoding step, and each of the decoding modules includes at least one upsampling operation and at least one convolution operation, and generates output features.

6. The method according to claim 5, the input feature includes at least one of the output feature of the previous decoding module and the encoded feature output by the encoding step, wherein, the input feature can also include encoded features with high-level semantics, encoded features containing low-level details, and corresponding encoded features.

7. The method according to claim 5, the decoding module includes at least one sub-module composed of an upsampling operation and a convolution operation to decode features for different targets.

8. The method according to claim 7, wherein, one of the sub-modules is a segmentation sub-module corresponding to a segmentation task, and the segmentation sub-module can include operations for guiding segmentation.

9. The method according to claim 8, wherein, the operations for guiding segmentation include a conversion operation and an upsampling operation, and are used to calculate intermediate features for guiding segmentation from high-level semantic encoded features, and add the intermediate features for guiding segmentation to the features of the segmentation sub-module through a first operation.

10. The method according to claim 7, wherein one of the sub-modules is a matting sub-module corresponding to a matting task, wherein, the matting sub-module includes at least one convolution operation and at least one upsampling operation, and the matting sub-module can include operations for guiding matting.

11. The method according to claim 10, wherein, the operations for guiding matting include a conversion operation and a downsampling operation, and are used to calculate intermediate features for guiding matting from low-level detail encoded features, and add the intermediate features for guiding matting to the features of the matting sub-module through a first operation.

12. The method according to claim 10, wherein one of the sub-modules is a depth sub-module corresponding to a depth estimation task, and the depth sub-module adopts the same structure as the matting sub-module.

13. The method according to claim 5, generating the output feature by integrating operations based on addition or concatenation of different decoded features.

14. The method according to claim 7, wherein, the decoding module includes a skip connection sub-module located before the shown sub-module.

15. The method according to claim 7, the decoding module includes a skip connection sub-module located after the sub-module.

16. The method according to claim 14 or 15, wherein, the skip connection sub-module includes: a plurality of feature transformation operations for transforming the input feature and the corresponding skip connection encoded feature to obtain the transformed feature, a first operation for enhancing the input feature according to the transformed feature, a convolution operation for further fusing the enhanced feature into an output feature.

17. The method according to claim 9 or 11, wherein the first operation is to merge specified features using an addition operation or a concatenation method.

18. The method according to claim 1, wherein, the head module includes a matting head module, and the head module includes at least one convolution operation and one activation operation.

19. The method according to claim 1, the head module includes a segmentation head module, and the shown segmentation head module includes at least one convolution operation and is used to generate a segmentation image from the final encoded feature.

20. The method according to claim 19, wherein, the head module includes a fusion operation, and the fusion operation is used to fuse the segmentation image into the output image to generate a final output image.

21. The method according to claim 16, wherein the first operation is to merge specified features using an addition operation or a concatenation method.

22. A training method of a neural network model, including: a construction step of constructing a neural network model used according to the method of claim 1; a prediction step of calculating a predicted output result according to the constructed neural network model and the data obtained from the training dataset; and an update step of calculating a loss according to the loss function and the predicted output result to update the parameters of the current neural network.

23. The method according to claim 22, in the case that the neural network model updated in the update step does not meet a specific condition, re-entering the prediction step.

24. The method according to claim 22, wherein, the loss function is the sum of multiple losses defined according to the prediction result and different tasks.

25. The method according to claim 23, wherein, the specific condition can be defined as the training cycle reaching a predetermined maximum number of training cycle times.

26. The method according to claim 23, wherein, the specific condition is that the calculated loss is lower than a predetermined threshold.

27. An image processing device, the device including: an encoding unit configured to generate a plurality of encoded features with different resolutions based on an input image and an encoder; A decoding unit configured to decode, based on the multiple encoded features and a decoder formed by cascading multiple decoding modules, a decoded feature having the same resolution as the input image; A prediction unit configured to predict, based on the decoded feature and a header module, an output image having the same resolution as the input image.

28. A training apparatus for a neural network model, comprising: A construction unit configured to construct a neural network model used in the method according to claim 1; A prediction unit configured to calculate a predicted output result according to the constructed neural network model and data obtained from a training dataset; An update unit configured to calculate a loss according to a loss function and the predicted output result to update the parameters of the current neural network.

29. An application method for a neural network model, characterized in that the application method comprises: Storing a neural network model trained by the training method according to any one of claims 22 to 26; Receiving a dataset corresponding to a task requirement executable by the stored neural network model; Performing operations on the dataset layer by layer from top to bottom in the stored neural network model, and outputting a result.

30. An application apparatus for a neural network model, characterized in that the application apparatus comprises: A storage module configured to store a neural network model trained by the training method according to any one of claims 22 to 26; A receiving module configured to receive a dataset corresponding to a task requirement executable by the stored neural network model; A processing module configured to perform operations on the dataset layer by layer from top to bottom in the stored neural network model, and output a result.

31. A non-transitory computer-readable storage medium storing instructions that, when executed by a computer, cause the computer to perform the image processing method according to any one of claims 1 to 21.

32. A non-transitory computer-readable storage medium storing instructions that, when executed by a computer, cause the computer to perform the training method for a neural network model according to any one of claims 22 to 26.