Optimization Method of Network Model, Image Processing Method and Electronic Device
Through the methods of knowledge distillation and image blocking error calculation, the knowledge forgetting problem of deep learning models when adding labeled data is solved, and the optimization efficiency and cutout accuracy of the network model are improved.
Patent Information
- Application Number
- CN202110008820.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-01-05
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-01-05
AI Technical Summary
In the prior art, deep learning models can easily lead to poor cutout effects of the original sample when learning newly labeled data, and there may be errors in the incremental images manually labeled, affecting the optimization effect of the network model. The existing optimization algorithms have not effectively solved this problem.
The prediction results of the original model are extracted every time a new sample is learned, and the fusion results are used to guide the model weight iteration. The errors of the original cutout model mask and incremental image mask are calculated in combination with image chunking, the labeling error is corrected, and the model weight update stability is maintained.
It improves the optimization efficiency of the network model, corrects the impact of labeling errors in incremental images on model optimization, maintains the stability of model weight updates, and improves the accuracy of the cutout model on new data.
Smart Images

Figure CN114782461B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of network model optimization. Specifically, it relates to an optimization method for a network model, an image processing method, and an electronic device. Background Art
[0002] The task of an image segmentation algorithm is to perform semantic analysis on the image content and extract the target object, which is a basic operation in processes such as image beautification, poster production, and film and television special effects. After the first version of the image segmentation algorithm model is developed, continuous data annotation and model training will surely be carried out to continuously optimize the processing effect of the model on new samples.
[0003] However, deep learning models have the characteristic of catastrophic forgetting. When learning newly added annotated data, it is easy to cause the matte extraction effect of previous samples to deteriorate. In addition, during the annotation process of image data, since it is a manual operation, the edges of objects cannot be completely and accurately matte-extracted, and there is no completely accurate ground truth standard.
[0004] Therefore, it is acceptable that there are slight differences between the results predicted by the existing model and the manually annotated results. When the error between the two is less than a certain range, there is no need to force the model to learn the manually annotated results. Moreover, there may be errors in the incrementally annotated images, which affect the optimization effect of the network model. However, the existing technology usually does not consider this point when using the matte iteration optimization algorithm, which will bring instability to the edge effect during the training process.
[0005] In view of the above problems, no effective solution has been proposed yet. Summary of the Invention
[0006] Embodiments of this application provide an optimization method for a network model, an image processing method, and an electronic device to at least solve the technical problem of low optimization efficiency in the existing solutions for optimizing network models.
[0007] According to one aspect of the embodiments of this application, an optimization method for a network model is provided, including: determining a first training model and a second training model based on a network model to be optimized, where the network model to be optimized is used for matte extraction of a target image, the range of weight parameters in the first training model satisfies a first threshold, and the range of weight parameters in the second training model satisfies a second threshold; inputting an incremental image into the first training model to output a first mask; performing mask fusion processing on the first mask and a second mask to obtain a target mask, where the second mask corresponds to the incremental image; and updating the weight parameters in the second training model using the target mask to obtain an optimized network model.
[0008] According to another aspect of the embodiments of the present application, another method for optimizing a network model is provided, including: determining a first training model and a second training model based on the network model to be optimized, wherein the network model to be optimized is used for matte extraction processing of a target image to obtain a foreground target object included in the target image, the weight parameters in the first training model remain unchanged, and the weight parameters in the second training model can be continuously updated; inputting an incremental image into the first training model to output a first mask; performing mask fusion processing on the first mask and a second mask to obtain a target mask, wherein the second mask corresponds to the incremental image; and updating the weight parameters in the second training model by using the target mask to obtain an optimized network model.
[0009] According to another aspect of the embodiments of the present application, an image processing method is further provided, including: receiving a currently input target image; sending the target image to a server; receiving a matte extraction result from the server, wherein the matte extraction result is used to describe a foreground target object included in the target image, and the matte extraction result is obtained by the server through matte extraction processing of the target image by using an optimized network model; and displaying the matte extraction result locally on the client.
[0010] According to another aspect of the embodiments of the present application, an image processing method is further provided, including: receiving a target image from a client; performing matte extraction processing on the target image by using an optimized network model to obtain a matte extraction result, wherein the matte extraction result is used to describe a foreground target object included in the target image; and returning the matte extraction result to the client and displaying the matte extraction result locally on the client.
[0011] According to another aspect of the embodiments of the present application, an image processing method is further provided, including: receiving a currently input target commodity image, wherein the target commodity image includes a foreground commodity image and a background layout image; sending the target commodity image to a server; receiving a matte extraction result from the server, wherein the matte extraction result is the foreground commodity image extracted from the target commodity image by the server through matte extraction processing of the target commodity image by using an optimized network model; and displaying the matte extraction result locally on the client.
[0012] According to another aspect of the embodiments of the present application, a non-volatile storage medium is further provided. The non-volatile storage medium includes a stored program, wherein when the program runs, it controls the device where the non-volatile storage medium is located to execute any one of the above-mentioned network model optimization methods and the above-mentioned image processing methods.
[0013] According to another aspect of the embodiments of the present application, an electronic device is further provided, including: a processor; and a memory connected to the above-mentioned processor for providing instructions for the above-mentioned processor to process the following processing steps: determining a first training model and a second training model based on a network model to be optimized, wherein the network model to be optimized is used for performing matte extraction processing on a target image to obtain a foreground target object included in the target image, the weight parameters in the first training model remain unchanged, and the weight parameters in the second training model can be continuously updated; inputting an incremental image into the first training model to output a first mask; performing mask fusion processing on the first mask and a second mask to obtain a target mask, wherein the second mask corresponds to the incremental image; and updating the weight parameters in the second training model by using the target mask to obtain an optimized network model.
[0014] In the embodiments of the present application, a first training model and a second training model are determined based on a network model to be optimized, wherein the network model to be optimized is used for performing matte extraction processing on a target image to obtain a foreground target object included in the target image, the weight parameters in the first training model remain unchanged, and the weight parameters in the second training model can be continuously updated; inputting an incremental image into the first training model to output a first mask; performing mask fusion processing on the first mask and a second mask to obtain a target mask, wherein the second mask corresponds to the incremental image; and updating the weight parameters in the second training model by using the target mask to obtain an optimized network model.
[0015] It is easy to notice that the embodiments of the present application adopt the method of knowledge distillation. When learning new samples each time, the prediction results of the original model are extracted, and the fusion results of the original model are used to guide the iteration of the model weights, reducing the forgetting of the original knowledge by the model, maintaining the stability of the model weight update, and not requiring stock images to participate in the training process, improving the efficiency of optimizing the network model. Moreover, the embodiments of the present application adopt the method of calculating the error between the mask of the original matte extraction model and the mask of the incremental image in image blocks. When the error between the two is less than the error threshold, it is considered that the prediction result of the original matte extraction model is available, which can correct the influence of the annotation error in the incremental image on the optimization of the network model.
[0016] Thus, the embodiments of the present application achieve the purpose of improving the optimization efficiency of the optimized network model, thereby realizing the technical effect of correcting the influence of the annotation error in the incremental image on the optimization of the network model, and further solving the technical problem of low optimization efficiency in the existing solutions for optimizing network models. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The accompanying drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0018] Figure 1 is a hardware structural block diagram of a computer terminal (or mobile device) for implementing an optimization method of a network model according to an embodiment of the present application;
[0019] Figure 2 is a flowchart of an optimization method of a network model according to an embodiment of the present application;
[0020] Figure 3 is a schematic diagram of a scenario of an optimization method of a network model according to an embodiment of the present application;
[0021] Figure 4 is a flowchart of an image processing method according to an embodiment of the present application;
[0022] Figure 5 is a flowchart of another image processing method according to an embodiment of the present application;
[0023] Figure 6 is a flowchart of another optimization method of a network model according to an embodiment of the present application;
[0024] Figure 7 is a flowchart of another image processing method according to an embodiment of the present application;
[0025] Figure 8 is a schematic structural diagram of an optimization device of a network model according to an embodiment of the present application;
[0026] Figure 9 is a schematic structural diagram of an image processing device according to an embodiment of the present application;
[0027] Figure 10 is a schematic structural diagram of another image processing device according to an embodiment of the present application;
[0028] Figure 11 is a structural block diagram of another computer terminal according to an embodiment of the present application. Detailed Embodiments
[0029] To enable those skilled in the art to better understand the solution of this application, the following will clearly and completely describe the technical solution in the embodiments of this application in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.
[0030] It should be noted that the terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0031] First, some nouns or terms that appear in the process of describing the embodiments of this application are applicable to the following explanations:
[0032] Mask: It refers to an image with the same size as the original image, used to mark whether each pixel belongs to the foreground or the background.
[0033] Incremental learning: It refers to a learning system that can continuously learn new knowledge from incremental new samples and can retain most of the knowledge learned previously.
[0034] Knowledge distillation: It refers to a method of transferring the knowledge in one model to a new model.
[0035] Embodiment 1
[0036] According to the embodiments of this application, an embodiment of an optimization method for a network model is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from here.
[0037] The method embodiment provided by Embodiment 1 of this application can be executed on a mobile terminal, a computer terminal or a similar computing device. Figure 1 A hardware structure block diagram of a computer terminal (or mobile device) for implementing the optimization method of the network model is shown. As Figure 1As shown, the computer terminal 10 (or mobile device 10) may include one or more processors 102 (shown as 102a, 102b, ……, 102n in the figure) (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission module 106 for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 may further include more or fewer components than Figure 1 shown in, or have a different configuration from Figure 1 that shown.
[0038] It should be noted that the above one or more processors 102 and / or other data processing circuits are generally referred to as "data processing circuits" herein. The data processing circuit may be embodied in software, hardware, firmware, or any combination thereof, in whole or in part. In addition, the data processing circuit may be a single independent processing module, or be incorporated in whole or in part into any one of the other elements in the computer terminal 10 (or mobile device). As involved in the embodiments of the present application, the data processing circuit is a kind of processor control (such as the selection of a variable resistance terminal path connected to an interface).
[0039] The memory 104 may be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the optimization method of the network model in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implements the above-mentioned optimization method of the network model. The memory 104 may include a high-speed random access memory, and may further include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely set relative to the processor 102, and these remote memories may be connected to the computer terminal 10 through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an enterprise internal network, a local area network, a mobile communication network, and combinations thereof.
[0040] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by the communication provider of the computer terminal 10. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0041] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables the user to interact with the user interface of the computer terminal 10 (or mobile device).
[0042] Under the above operating environment, the present application provides an optimization method for a network model as Figure 2 shown. Figure 2 is a flowchart of an optimization method for a network model according to an embodiment of the present application. As Figure 2 shown, the above optimization method for the network model includes:
[0043] Step S202, determining a first training model and a second training model based on the network model to be optimized;
[0044] Step S204, inputting the incremental image into the above first training model and outputting a first mask;
[0045] Step S206, performing mask fusion processing on the above first mask and a second mask to obtain a target mask, where the second mask corresponds to the incremental image;
[0046] Step S208, using the above target mask to update the weight parameters in the above second training model to obtain an optimized network model.
[0047] In the embodiment of the present application, by determining a first training model and a second training model based on the network model to be optimized, where the network model to be optimized is used to perform matte extraction on a target image to obtain the foreground target object included in the target image, the weight parameters in the first training model remain unchanged, and the weight parameters in the second training model can be continuously updated; inputting the incremental image into the first training model and outputting a first mask; performing mask fusion processing on the first mask and the second mask to obtain a target mask, where the second mask corresponds to the incremental image; using the target mask to update the weight parameters in the second training model to obtain an optimized network model.
[0048] It is easy to notice that the embodiments of the present application adopt the method of knowledge distillation. When learning new samples each time, the prediction results of the original model are extracted, and the fusion results of the original model are used to guide the iteration of the model weights, reducing the model's forgetting of the original knowledge, maintaining the stability of the model weight update, and not requiring the stock images to participate in the training process, thus improving the efficiency of optimizing the network model. Moreover, the embodiments of the present application adopt the method of calculating the error between the mask of the original matte model and the mask of the incremental image in blocks. When the error between the two is less than the error threshold, it is considered that the prediction result of the original matte model is available, which can correct the influence of the annotation error in the incremental image on the optimization of the network model.
[0049] Thus, the embodiments of the present application achieve the purpose of improving the optimization efficiency of the network model, thereby realizing the technical effect of correcting the influence of the annotation error in the incremental image on the optimization of the network model, and further solving the technical problem of low optimization efficiency in the existing solutions for optimizing the network model.
[0050] Optionally, in the embodiments of the present application, the network model to be optimized is used to perform matte processing on the target image to obtain the foreground target object included in the target image. The weight parameters in the first training model remain unchanged, and the weight parameters in the second training model can be continuously updated.
[0051] Optionally, the network model to be optimized can be an image segmentation algorithm model, for example, a matte model.
[0052] Optionally, the method for optimizing the network model provided by the embodiments of the present application can be, but is not limited to, applied to the application scenario of optimizing the image segmentation algorithm model.
[0053] It should be noted that the task of the image segmentation algorithm model is to perform semantic analysis on the image content and extract the target object, which is the basic operation in processes such as image beautification, poster production, and film and television special effects. It can be widely applied to film and television post-production, automatic generation of online static and dynamic advertisements in e-commerce, and image segmentation services on intelligent vision platforms, and has a very strong role in AI empowerment in industries such as interactive entertainment (such as live broadcast, beauty camera app), film and television post-production, photo retouching, and e-commerce.
[0054] In the embodiments of the present application, taking the network model to be optimized as a matte model as an example, on the pre-trained matte model, a smaller learning rate is adopted, and all or part of the layers in the network are trained on the new annotation dataset, and the number of iterations is controlled.
[0055] Since the manual annotation of the incremental image may not be perfectly accurate, the matting ability of the original matting model is used to correct the annotation error in the incremental image. In the embodiments of the present application, when adopting the matting model optimized by the incremental image, unnecessary modification of the model weights can be avoided, so that the matting model can keep the original parameters stable while learning new data samples.
[0056] In an optional embodiment, the initial assignments of the weight parameters in the to-be-optimized network model, the first training model, and the second training model are the same.
[0057] In the embodiments of the present application, the weight parameters of the matting model to be optimized are assigned to two new network training models with the same structure as the original matting model, and the first training model is set as a non-trainable model (i.e., the weights are not updated through backpropagation), and the second training model is set as a trainable model (i.e., the model weights are updated through backpropagation).
[0058] Taking this incremental image as an example of a lady's handbag, as Figure 3 shown, when the matting model to be optimized is trained using the incremental image, the incremental image is input into the first training model and the second training model at the same time. Then, the first training model will output a first mask, which is the prediction result of the original matting model. By fusing the first mask (i.e., the prediction result of the original matting model) with the second mask of the incremental image, a new target mask is generated. Here, the mask refers to an image with the same size as the original image, which is used to mark whether each pixel belongs to the foreground or the background. Finally, the target mask is used to guide the weight update of the trainable model. After the training is completed, the weights of the trainable model are the optimized model weights.
[0059] It should be noted that the to-be-optimized model in the embodiments of the present application can be a matting model, which predicts the masks of foreground objects and background regions, and is used to expand the segmentation ability of an existing semantic segmentation model for new categories, and can improve the accuracy of the matting model for more images.
[0060] In the embodiments of the present application, during the training process of the matting model, an image pair dataset (two paired images, namely the original image and the mask) needs to be provided to the model. The original image is input into the matting model, and the matting model predicts and generates a mask, and calculates the error loss between the predicted mask and the true mask. Then, the backpropagation algorithm is applied to the weights of the matting model, and the weight parameters of the matting model are updated layer by layer, so as to optimize the trained matting model.
[0061] It should be noted that the network model optimization method provided by the embodiments of the present application requires that the old data set and the new data set have the same distribution. Otherwise, it is easy to overfit to a small number of incremental images, and it is difficult to stably improve the model effect because how to set the learning rate, how to select the trainable layers and the number of iterations all depend on experience.
[0062] In an optional embodiment, performing mask fusion processing on the first mask and the second mask to obtain the target mask includes:
[0063] Step S302: Perform image block processing of the same size on the first mask and the second mask, divide the first mask into a plurality of first block regions, and divide the second mask into a plurality of second block regions;
[0064] Step S304: Perform the same numbering process on the same block regions in the first mask and the second mask to obtain a numbering result;
[0065] Step S306: Calculate the error between the same-numbered block regions in the first mask and the second mask based on an error calculation function to obtain an error result;
[0066] Step S308: Calculate a plurality of target block regions by using a truncation function, the error result, the plurality of first block regions, and the plurality of second block regions;
[0067] Step S310: Stitch the plurality of target block regions to obtain the target mask.
[0068] In an optional embodiment, the error calculation function includes one of the following: mean squared error loss function (L2 loss function), mean absolute error loss function (L1 loss function), cross-entropy loss function (cross-entropy loss function).
[0069] In the above optional embodiment of the present application, by performing image block processing of the same size on the first mask M and the second mask M', the image is cut into n*m block regions R_i (i = 1,.., n*m); specifically, the first mask is divided into a plurality of first block regions, and the second mask is divided into a plurality of second block regions.
[0070] In the above optional embodiment of the present application, based on the error calculation function Loss(), the error between the block regions M(R_i) and M'(R_i) with the same number in the first mask and the second mask is calculated to obtain an error result loss = Loss(M(R_i), M'(R_i)); multiple target block regions are calculated by using a truncation function, the error result, the multiple first block regions, and the multiple second block regions; and the multiple target block regions are spliced to obtain the target mask.
[0071] In an optional embodiment, calculating the error between the block regions with the same number in the first mask and the second mask based on the above error calculation function, and obtaining the above error result includes:
[0072] Step S402, when the calculated error is less than the error threshold, the value of the error result is set to 0;
[0073] Step S404, when the calculated error is greater than or equal to the error threshold, the value of the error result remains the calculated error.
[0074] In the above optional embodiment, by presetting an error threshold T, when the error loss between the block regions with the same number in the first mask and the second mask is less than the error threshold T, the error result loss = 0 is set; when the error loss between the block regions with the same number in the first mask and the second mask is greater than or equal to the error threshold T, the value of the error result remains the calculated error.
[0075] In an optional embodiment, calculating the multiple target block regions by using the truncation function, the error result, the multiple first block regions, and the multiple second block regions includes:
[0076] Step S502, setting the error result, the error threshold, an error control coefficient, a first value, and a second value as input parameters of the truncation function, and outputting a calculation result, where the first value is less than the second value;
[0077] Step S504, calculating the multiple target block regions by using the calculation result, the multiple first block regions, and the multiple second block regions.
[0078] In an optional embodiment, in the above step S502, setting the error result, the error threshold, the error control coefficient, the first value, and the second value as the input parameters and outputting the calculation result includes:
[0079] Step S602: Calculate a third value using the above error result, the above error threshold, and the above error control coefficient.
[0080] Step S604: When the above third value is less than the above first value, set the value of the above calculation result to the above first value; when the above third value is greater than the above second value, set the value of the above calculation result to the above second value; when the above third value is greater than or equal to the above first value and less than or equal to the above second value, set the value of the above calculation result to the above third value.
[0081] Optionally, clip is the above truncation function, clip(a, 0, 1) means that when a < 0, the value is 0; when a > 1, the value is 1; in other cases, the value is a; beta is the error control coefficient, and this beta can be set to any number greater than 1. The larger beta is, the more trust is placed in the effect of the original matte extraction model.
[0082] In the embodiment of the present application, the above calculation result M_new(R_i) = M(R_i) + clip(loss / (T * beta), 0, 1) * (M'(R_i) - M(R_i)). By calculating all M_new(R_i) image blocks when i ranges from 0 to n * m, and then splicing them, the target mask M_new is obtained.
[0083] In the embodiment of the present application, due to the use of knowledge distillation, when learning new samples each time, the prediction results of the original model are extracted, and the fusion results of the original model are used to guide the iteration of the model weights, reducing the model's forgetting of the original knowledge and maintaining the stability of the model weight update. And the method of calculating the error between the mask of the original matte extraction model and the mask of the incremental image in image blocks. When the error between the two is less than the error threshold, it is considered that the prediction result of the original matte extraction model is available, and no stock data needs to participate in the training, so as to correct the influence of the annotation error in the incremental image on the model optimization and improve the optimization efficiency of the network model.
[0084] According to the embodiment of the present application, provided is another Figure 4 image processing method as shown. Figure 4 is a flowchart of an image processing method according to the embodiment of the present application. As Figure 4 shown, the above image processing method includes:
[0085] Step S702: Receive the currently input target image.
[0086] Step S704: Send the above target image to the server.
[0087] Step S706: Receive the matting result from the above-mentioned server. The matting result is used to describe the foreground target object included in the above-mentioned target image, and the matting result is obtained by the above-mentioned server through matting the above-mentioned target image with the optimized network model.
[0088] Step S708: Display the above-mentioned matting result locally on the client.
[0089] It should be noted that the execution subject of the embodiments of this application is the SaaS client. In the embodiments of this application, the client receives the currently input target image, sends the above-mentioned target image to the server, receives the matting result from the above-mentioned server, and displays the above-mentioned matting result locally on the client.
[0090] It is easy to notice that the embodiments of this application adopt the method of knowledge distillation. When learning new samples each time, the prediction results of the original model are extracted, and the fusion results of the original model are used to guide the iteration of the model weights, reducing the model's forgetting of the original knowledge, maintaining the stability of the model weight update, and not requiring stock images to participate in the training process, improving the efficiency of network model optimization. Moreover, the embodiments of this application adopt the method of calculating the error between the mask of the original matting model and the mask of the incremental image in blocks. When the error between the two is less than the error threshold, it is considered that the prediction result of the original matting model is available, which can correct the influence of the annotation error in the incremental image on the network model optimization.
[0091] Therefore, the embodiments of this application achieve the purpose of improving the optimization efficiency of the optimized network model, thus realizing the technical effect of correcting the influence of the annotation error in the incremental image on the network model optimization, and further solving the technical problem of low optimization efficiency in the existing solutions for optimizing network models.
[0092] In an optional embodiment, the above-mentioned optimized network model is a network model obtained by optimizing the weight parameters in the network model to be optimized. The network model to be optimized is used to determine a first training model and a second training model. The range of change of the weight parameters in the first training model satisfies a first threshold, and the range of change of the weight parameters in the second training model satisfies a second threshold. The first training model is used to determine a first mask based on the incremental image. The first mask is fused with a second mask to obtain a target mask. The second mask corresponds to the incremental image. The target mask is used to update the weight parameters in the second training model to obtain the above-mentioned optimized network model.
[0093] In an alternative embodiment, the optimized network model is a network model obtained by optimizing the weight parameters in the network model to be optimized. The network model to be optimized is used to determine a first training model and a second training model. The weight parameters in the first training model remain unchanged, and the weight parameters in the second training model can be continuously updated. The first training model is used to determine a first mask based on an incremental image. The first mask is fused with a second mask to obtain a target mask. The second mask corresponds to the incremental image. The target mask is used to update the weight parameters in the second training model to obtain the optimized network model.
[0094] According to an embodiment of the present application, there is provided another image processing method as Figure 5 shown. Figure 5 FIG. is a flowchart of an image processing method according to an embodiment of the present application. As Figure 5 shown, the image processing method includes:
[0095] Step S802, receiving a target image from a client;
[0096] Step S804, performing matte extraction processing on the target image through an optimized network model to obtain a matte extraction result;
[0097] Step S806, returning the matte extraction result to the client and locally displaying the matte extraction result on the client.
[0098] In the above step S804, the matte extraction result is used to describe the foreground target object included in the target image.
[0099] It should be noted that the execution subject of the embodiment of the present application is the SaaS server. In the embodiment of the present application, the target image is received from the client through the server; the target image is subjected to matte extraction processing through the optimized network model to obtain a matte extraction result; the matte extraction result is returned to the client and locally displayed on the client.
[0100] It is easy to notice that the embodiment of the present application adopts the method of knowledge distillation. When learning new samples each time, the prediction results of the original model are extracted, and the fusion results of the original model are used to guide the iteration of the model weights, reducing the model's forgetting of the original knowledge, maintaining the stability of the model weight update, and not requiring the participation of stock images in the training process, improving the efficiency of network model optimization. Moreover, the embodiment of the present application adopts the method of calculating the error between the matte of the original matte extraction model and the matte of the incremental image in image blocks. When the error between the two is less than the error threshold, it is considered that the prediction result of the original matte extraction model is available, which can correct the influence of the annotation error in the incremental image on the optimization of the network model.
[0101] Thus, the embodiments of the present application achieve the purpose of improving the optimization efficiency of the network model, thereby realizing the technical effect of correcting the influence of the annotation error in the incremental image on the optimization of the network model, and further solving the technical problem of low optimization efficiency in the existing solutions for optimizing the network model.
[0102] In an alternative embodiment, the optimized network model is a network model obtained by optimizing the weight parameters in the network model to be optimized. The network model to be optimized is used to determine a first training model and a second training model. The range of change of the weight parameters in the first training model satisfies a first threshold, and the range of change of the weight parameters in the second training model satisfies a second threshold. The first training model is used to determine a first mask based on the incremental image. The first mask is fused with a second mask to obtain a target mask. The second mask corresponds to the incremental image. The target mask is used to update the weight parameters in the second training model to obtain the optimized network model.
[0103] In an alternative embodiment, the optimized network model is a network model obtained by optimizing the weight parameters in the network model to be optimized. The network model to be optimized is used to determine a first training model and a second training model. The weight parameters in the first training model remain unchanged, and the weight parameters in the second training model can be continuously updated. The first training model is used to determine a first mask based on the incremental image. The first mask is fused with a second mask to obtain a target mask. The second mask corresponds to the incremental image. The target mask is used to update the weight parameters in the second training model to obtain the optimized network model.
[0104] According to the embodiments of the present application, there is provided another method for optimizing a network model as Figure 6 shown. Figure 6 It is a flowchart of another method for optimizing a network model according to the embodiments of the present application. As Figure 6 shown, the method for optimizing the network model includes:
[0105] Step S902: Determine a first training model and a second training model based on the network model to be optimized, where the network model to be optimized is used to perform matte extraction on a target image, the range of change of the weight parameters in the first training model satisfies a first threshold, and the range of change of the weight parameters in the second training model satisfies a second threshold;
[0106] Step S904: Input the incremental image into the first training model and output a first mask;
[0107] Step S906: Perform mask fusion processing on the above-mentioned first mask and second mask to obtain a target mask, where the above-mentioned second mask corresponds to the incremental image;
[0108] Step S908: Update the weight parameters in the above-mentioned second training model using the above-mentioned target mask to obtain an optimized network model.
[0109] In the embodiment of the present application, by determining a first training model and a second training model based on the network model to be optimized, where the network model to be optimized is used for matte extraction processing of a target image, the change range of the weight parameters in the first training model satisfies a first threshold, and the change range of the weight parameters in the second training model satisfies a second threshold; input the incremental image into the first training model and output a first mask; perform mask fusion processing on the first mask and the second mask to obtain a target mask, where the second mask corresponds to the incremental image; update the weight parameters in the second training model using the target mask to obtain an optimized network model.
[0110] It is easy to notice that the embodiment of the present application adopts the method of knowledge distillation. When learning new samples each time, the prediction results of the original model are extracted, and the fusion results of the original model are used to guide the iteration of the model weights, reducing the model's forgetting of the original knowledge, maintaining the stability of the model weight update, and not requiring stock images to participate in the training process, improving the efficiency of network model optimization. Moreover, since the change range of the weight parameters in the first training model satisfies the first threshold and the change range of the weight parameters in the second training model satisfies the second threshold, the embodiment of the present application adopts the method of calculating the error between the matte of the original matte extraction model and the matte of the incremental image in blocks. When the error between the two is less than the error threshold, it is considered that the prediction result of the original matte extraction model is available, which can correct the influence of the annotation error in the incremental image on the optimization of the network model.
[0111] Thus, the embodiment of the present application achieves the purpose of improving the optimization efficiency of the optimized network model, thereby realizing the technical effect of correcting the influence of the annotation error in the incremental image on the optimization of the network model, and further solving the technical problem of low optimization efficiency in the existing solutions for optimizing network models.
[0112] Optionally, in the embodiment of the present application, the network model to be optimized is used for matte extraction processing of a target image, and the network model to be optimized can be an image segmentation algorithm model, for example, a matte extraction model.
[0113] Optionally, the network model optimization method provided by the embodiment of the present application can be but is not limited to being applied in the application scenario of optimizing an image segmentation algorithm model.
[0114] It should be noted that the task of the image segmentation algorithm model is to perform semantic analysis on the image content and extract the target object, which is a basic operation in processes such as image beautification, poster production, and film and television special effects. It can be widely applied to post-production of film and television, automatic generation of online static and dynamic advertisements in e-commerce, and image segmentation services on intelligent vision platforms. It plays a very important role in AI empowerment in industries such as interactive entertainment (such as live streaming, beauty camera apps), post-production of film and television, photo retouching, and e-commerce.
[0115] In the embodiment of the present application, taking the above-mentioned network model to be optimized as the matte extraction model as an example, on the pre-trained matte extraction model, a smaller learning rate is adopted, and all or part of the layers in the network are trained on the new labeled data set, and the number of iterations is controlled.
[0116] Since the manual annotation of the incremental image is not necessarily perfectly accurate, the matte extraction ability of the original matte extraction model is used to correct the annotation error in the incremental image. In the embodiment of the present application, the change range of the weight parameters in the above-mentioned first training model satisfies the first threshold, and the change range of the weight parameters in the second training model satisfies the second threshold. By avoiding unnecessary modification of the model weights when using the incremental image to optimize and train the matte extraction model, the matte extraction model can keep the original parameters stable as much as possible while learning new data samples.
[0117] In an optional embodiment, the initial assignments of the weight parameters in the above-mentioned network model to be optimized, the above-mentioned first training model, and the above-mentioned second training model may but are not limited to be the same.
[0118] In the embodiment of the present application, the weight parameters of the matte extraction model to be optimized are assigned to two new network training models with the same structure as the original matte extraction model, and the first training model is set as an untrainable model (i.e., the weights are not updated through backpropagation), and the second training model is set as a trainable model (i.e., the model weights are updated through backpropagation).
[0119] Taking this incremental image as an example of a lady's handbag, still as Figure 3 shown, when the network model to be optimized uses the incremental image for training, the incremental image is input into the first training model and the second training model at the same time. Then, the first training model will output a first mask, which is the prediction result of the original matte extraction model. By fusing the first mask (i.e., the prediction result of the original matte extraction model) and the second mask of the incremental image, a new target mask is generated. Here, the mask refers to an image with the same size as the original image, which is used to mark whether each pixel belongs to the foreground or the background. Finally, the target mask is used to guide the weight update of the trainable model. After the training is completed, the weights of the trainable model are the optimized model weights.
[0120] It should be noted that the model to be optimized in the embodiments of the present application can be a matting model, which predicts the foreground object and the background area mask and is used to expand the segmentation ability of an existing semantic segmentation model for new categories, and can improve the accuracy of the model in matting more images.
[0121] In the embodiments of the present application, during the training process of the matting model, an image pair dataset (two paired images, namely the original image and the mask) needs to be provided to the model. The original image is input into the matting model, which predicts and generates a mask, and calculates the error loss between the predicted mask and the true mask. Then, the backpropagation algorithm is applied to the weights of the matting model and the weight parameters of the matting model are updated layer by layer, thereby realizing the optimization of the trained matting model.
[0122] It should be noted that the network model optimization method provided in the embodiments of the present application requires that the old data set and the new data set have the same distribution. Otherwise, it is easy to overfit to a small number of incremental images, and the situations of how to set the learning rate, how to select the trainable layers and the number of iterations all depend on experience, and it is difficult to stably improve the model effect.
[0123] According to the embodiments of the present application, another image processing method is provided as Figure 7 shown. Figure 7 It is a flowchart of another image processing method according to the embodiments of the present application. As Figure 7 shown, the above image processing method includes:
[0124] Step S1002, receiving the currently input target commodity image, where the above target commodity image includes: a foreground commodity image and a background layout image;
[0125] Step S1004, sending the above target commodity image to the server;
[0126] Step S1006, receiving the matting result from the above server, where the above matting result is the above foreground commodity image extracted from the above target commodity image after the above server performs matting processing on the above target commodity image through the optimized network model;
[0127] Step S1008, locally displaying the above matting result on the client.
[0128] It should be noted that the execution subject of the embodiments of this application is the SaaS client. In the embodiments of this application, the client receives the currently input target commodity image, where the above target commodity image includes: a foreground commodity image and a background layout image; the above target commodity image is sent to the server; the client receives the matting result from the above server, where the above matting result is the foreground commodity image extracted from the above target commodity image after the server performs matting processing on the above target commodity image through an optimized network model; the above matting result is displayed locally on the client.
[0129] Optionally, the image processing method provided by the embodiments of this application can be but is not limited to being applied in application scenarios for optimizing image segmentation processing.
[0130] It should be noted that the task of image segmentation is to perform semantic analysis on the image content and extract the target object, which is a basic operation in processes such as image beautification, poster production, and film and television special effects. It can be widely applied in film and television post-production, automatic generation of online static and dynamic advertisements in e-commerce, and image segmentation services on intelligent vision platforms. It plays a very important role in AI empowerment in industries such as interactive entertainment (such as live streaming, beauty apps), film and television post-production, photo retouching, and e-commerce.
[0131] In the embodiments of this application, taking the above optimized network model as the matting model as an example, on the pre-trained matting model, a smaller learning rate is adopted, and all or part of the layers in the network are trained on a new labeled data set, and the number of iterations is controlled.
[0132] Since the manual annotation of the incremental image may not be perfectly accurate, the matting ability of the original matting model is used to correct the annotation error in the incremental image. In the embodiments of this application, the change range of the weight parameters in the above first training model satisfies the first threshold, and the change range of the weight parameters in the above second training model satisfies the second threshold. By avoiding unnecessary modification of the model weights when using the incremental image to optimize the trained matting model, the matting model can maintain the stability of the original parameters as much as possible while learning new data samples.
[0133] In the embodiments of this application, the weight parameters of the matting model to be optimized are assigned to two new network training models with the same structure as the original matting model, and the first training model is set as an untrainable model (i.e., the weights are not updated through backpropagation), and the second training model is set as a trainable model (i.e., the model weights are updated through backpropagation).
[0134] It should be noted that the optimized network model in the embodiments of the present application can be a matting model, which predicts the masks of foreground objects and background regions, and is used to expand the segmentation ability of an existing semantic segmentation model on new categories, and can improve the accuracy of the model in matting more images.
[0135] It is easy to notice that the embodiments of the present application adopt the method of knowledge distillation. When learning new samples each time, the prediction results of the original model are extracted, and the fusion results of the original model are used to guide the iteration of the model weights, reducing the model's forgetting of the original knowledge, maintaining the stability of the model weight update, and not requiring stock images to participate in the training process, improving the efficiency of network model optimization. Moreover, the embodiments of the present application adopt the method of calculating the error between the mask of the original matting model and the mask of the incremental image in image blocks. When the error between the two is less than the error threshold, it is considered that the prediction result of the original matting model is available, which can correct the influence of the annotation error in the incremental image on the network model optimization.
[0136] Thus, the embodiments of the present application achieve the purpose of improving the image processing efficiency by optimizing the network model, thereby realizing the technical effect of correcting the influence of the annotation error in the incremental image on the network model optimization, and further solving the technical problem of low optimization efficiency in the existing solutions for optimizing the network model.
[0137] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.
[0138] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present application.
[0139] Embodiment 2
[0140] According to the embodiments of the present application, there is also provided an apparatus embodiment for implementing the above network model optimization method.Figure 8 It is a schematic structural diagram of an optimization device for a network model according to an embodiment of the present application. As Figure 8 shown, the device includes: a determination module 600, a training module 602, a processing module 604, and an optimization module 606, where:
[0141] The determination module 600 is configured to determine a first training model and a second training model based on the network model to be optimized. Wherein, the network model to be optimized is used to perform matte extraction processing on a target image to obtain a foreground target object included in the target image. The weight parameters in the first training model remain unchanged, and the weight parameters in the second training model can be continuously updated; the training module 602 is configured to input an incremental image into the first training model and output a first mask; the processing module 604 is configured to perform mask fusion processing on the first mask and a second mask to obtain a target mask, where the second mask corresponds to the incremental image; the optimization module 606 is configured to update the weight parameters in the second training model by using the target mask to obtain an optimized network model.
[0142] It should be noted here that the above determination module 600, training module 602, processing module 604, and optimization module 606 correspond to steps S202 to S210 in Embodiment 1. The examples and application scenarios implemented by the four modules and the corresponding steps are the same, but are not limited to the content disclosed in the above Embodiment 1. It should be noted that the above modules, as part of the device, can run in the computer terminal 10 provided in Embodiment 1.
[0143] According to an embodiment of the present application, an apparatus embodiment for implementing the above image processing method is further provided. Figure 9 It is a schematic structural diagram of an image processing device according to an embodiment of the present application. As Figure 9 shown, the device includes: a first receiving module 700, a sending module 702, a second receiving module 704, and a display module 706, where:
[0144] A first receiving module 700, configured to receive a current input target image; a sending module 702, configured to send the target image to a server; a second receiving module 704, configured to receive a matting result from the server, where the matting result is used to describe a foreground target object included in the target image, and the matting result is obtained by the server performing matting processing on the target image through an optimized network model, and the optimized network model is a network model obtained by optimizing weight parameters in a network model to be optimized, and the network model to be optimized is used to determine a first training model and a second training model, weight parameters in the first training model remain unchanged, weight parameters in the second training model can be continuously updated, the first training model is used to determine a first mask based on an incremental image, the first mask is fused with a second mask to obtain a target mask, the second mask corresponds to the incremental image, and the target mask is used to update the weight parameters in the second training model to obtain the optimized network model; a display module 706, configured to locally display the matting result on a client.
[0145] It should be noted here that the first receiving module 700, the sending module 702, the second receiving module 704, and the display module 706 correspond to steps S702 to S708 in Embodiment 1. The functions of the four modules and the corresponding steps are the same in terms of implementation examples and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above modules, as part of the device, can run in the computer terminal 10 provided in Embodiment 1.
[0146] According to an embodiment of the present application, there is also provided an apparatus embodiment for implementing the above image processing method. Figure 10 It is a schematic structural diagram of an image processing apparatus according to an embodiment of the present application, as Figure 10 shown. The apparatus includes: a third receiving module 800, a matting module 802, and a returning module 804, where:
[0147] A third receiving module 800, configured to receive a target image from a client; a matting module 802, configured to perform matting processing on the target image through an optimized network model to obtain a matting result, where the matting result is used to describe a foreground target object included in the target image, and the optimized network model is a network model obtained by optimizing weight parameters in a network model to be optimized. The network model to be optimized is used to determine a first training model and a second training model. Weight parameters in the first training model remain unchanged, and weight parameters in the second training model can be continuously updated. The first training model is used to determine a first mask based on an incremental image. The first mask is fused with a second mask to obtain a target mask. The second mask corresponds to the incremental image. The target mask is used to update weight parameters in the second training model to obtain the optimized network model; a return module 804, configured to return the matting result to the client and display the matting result locally on the client.
[0148] It should be noted here that the third receiving module 800, the matting module 802, and the return module 804 correspond to steps S802 to S806 in Embodiment 1. The instances and application scenarios implemented by the three modules and the corresponding steps are the same, but are not limited to the content disclosed in Embodiment 1 above. It should be noted that the above modules, as part of the device, can run in the computer terminal 10 provided in Embodiment 1.
[0149] It should be noted that the preferred implementation manners of this embodiment can refer to the relevant descriptions in Embodiment 1, and will not be elaborated here.
[0150] Embodiment 3
[0151] According to an embodiment of the present application, an embodiment of an electronic device is further provided. The electronic device can be any one of the computing devices in a computing device cluster. The electronic device includes: a processor and a memory, where:
[0152] A processor; and a memory, connected to the processor, configured to provide instructions for the processor to perform the following processing steps: determining a first training model and a second training model based on a network model to be optimized, where the network model to be optimized is used to perform matting processing on a target image to obtain a foreground target object included in the target image, weight parameters in the first training model remain unchanged, and weight parameters in the second training model can be continuously updated; inputting an incremental image into the first training model to output a first mask; performing mask fusion processing on the first mask and a second mask to obtain a target mask, where the second mask corresponds to the incremental image; and using the target mask to update weight parameters in the second training model to obtain an optimized network model.
[0153] In an embodiment of the present application, a first training model and a second training model are determined based on a network model to be optimized. The network model to be optimized is used to perform matte extraction on a target image to obtain a foreground target object included in the target image. The weight parameters in the first training model remain unchanged, and the weight parameters in the second training model can be continuously updated. An incremental image is input into the first training model to output a first mask. The first mask and a second mask are subjected to mask fusion processing to obtain a target mask, where the second mask corresponds to the incremental image. The weight parameters in the second training model are updated using the target mask to obtain an optimized network model.
[0154] It is easy to notice that the embodiment of the present application adopts the method of knowledge distillation. Each time new samples are learned, the prediction results of the original model are extracted, and the fusion results of the original model are used to guide the iteration of the model weights, reducing the model's forgetting of the original knowledge, maintaining the stability of the model weight update, and not requiring stock images to participate in the training process, improving the efficiency of network model optimization. Moreover, the embodiment of the present application adopts the method of calculating the error between the matte of the original matte extraction model and the matte of the incremental image in blocks. When the error between the two is less than the error threshold, it is considered that the prediction result of the original matte extraction model is available, which can correct the influence of the annotation error in the incremental image on the network model optimization.
[0155] Thus, the embodiment of the present application achieves the purpose of improving the optimization efficiency of the optimized network model, thereby realizing the technical effect of correcting the influence of the annotation error in the incremental image on the network model optimization, and further solving the technical problem of low optimization efficiency in the existing solutions for optimizing network models.
[0156] It should be noted that the preferred implementation manner of this embodiment can refer to the relevant description in Embodiment 1 and will not be elaborated here.
[0157] Embodiment 4
[0158] According to an embodiment of the present application, an embodiment of a computer terminal is also provided. The computer terminal can be any computer terminal device in a computer terminal group. Optionally, in this embodiment, the computer terminal can also be replaced with a terminal device such as a mobile terminal.
[0159] Optionally, in this embodiment, the computer terminal can be located in at least one of multiple network devices in a computer network.
[0160] In this embodiment, the above computer terminal may execute the program code of the following steps in the optimization method of the network model of the application program: determining a first training model and a second training model based on the network model to be optimized, wherein the network model to be optimized is used for matte processing of a target image to obtain the foreground target object included in the target image, the weight parameters in the first training model remain unchanged, and the weight parameters in the second training model can be continuously updated; inputting the incremental image into the first training model to output a first mask; performing mask fusion processing on the first mask and a second mask to obtain a target mask, wherein the second mask corresponds to the incremental image; and updating the weight parameters in the second training model by using the target mask to obtain an optimized network model.
[0161] Optionally, Figure 11 is a structural block diagram of another computer terminal according to an embodiment of the present application, as Figure 11 shown, the computer terminal may include: one or more (only one is shown in the figure) processors 902, a memory 904, and a peripheral interface 906.
[0162] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the optimization method and device of the network model in the embodiment of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implements the above optimization method of the network model. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory may further include a memory remotely provided with respect to the processor, and these remote memories may be connected to the computer terminal through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0163] The processor may call the information and application program stored in the memory through a transmission device to execute the following steps: determining a first training model and a second training model based on the network model to be optimized, wherein the network model to be optimized is used for matte processing of a target image to obtain the foreground target object included in the target image, the weight parameters in the first training model remain unchanged, and the weight parameters in the second training model can be continuously updated; inputting the incremental image into the first training model to output a first mask; performing mask fusion processing on the first mask and a second mask to obtain a target mask, wherein the second mask corresponds to the incremental image; and updating the weight parameters in the second training model by using the target mask to obtain an optimized network model.
[0164] Optionally, the above-mentioned processor may also execute the program code of the following steps: perform image block processing of the same size on the above-mentioned first mask and the above-mentioned second mask, divide the above-mentioned first mask into a plurality of first block regions and divide the above-mentioned second mask into a plurality of second block regions; perform the same numbering process on the same block regions in the above-mentioned first mask and the above-mentioned second mask to obtain a numbering result; calculate the error between the same-numbered block regions in the above-mentioned first mask and the above-mentioned second mask based on an error calculation function to obtain an error result; calculate a plurality of target block regions by using a truncation function, the above-mentioned error result, the above-mentioned plurality of first block regions, and the above-mentioned plurality of second block regions; splice the above-mentioned plurality of target block regions to obtain the above-mentioned target mask.
[0165] Optionally, the above-mentioned processor may also execute the program code of the following steps: when the calculated error is less than the error threshold, set the value of the above-mentioned error result to 0; when the calculated error is greater than or equal to the above-mentioned error threshold, keep the value of the above-mentioned error result as the calculated error.
[0166] Optionally, the above-mentioned processor may also execute the program code of the following steps: set the above-mentioned error result, the above-mentioned error threshold, an error control coefficient, a first value, and a second value as input parameters of the above-mentioned truncation function, and output a calculation result, where the above-mentioned first value is less than the above-mentioned second value; calculate the above-mentioned plurality of target block regions by using the above-mentioned calculation result, the above-mentioned plurality of first block regions, and the above-mentioned plurality of second block regions.
[0167] Optionally, the above-mentioned processor may also execute the program code of the following steps: calculate a third value by using the above-mentioned error result, the above-mentioned error threshold, and the above-mentioned error control coefficient; when the above-mentioned third value is less than the above-mentioned first value, set the value of the above-mentioned calculation result to the above-mentioned first value; when the above-mentioned third value is greater than the above-mentioned second value, set the value of the above-mentioned calculation result to the above-mentioned second value; when the above-mentioned third value is greater than or equal to the above-mentioned first value and less than or equal to the above-mentioned second value, set the value of the above-mentioned calculation result to the above-mentioned third value.
[0168] The processor can call the information and application programs stored in the memory through the transmission device to execute the following steps: receive the current input target image; send the above target image to the server; receive the matte result from the above server, where the matte result is used to describe the foreground target object included in the above target image, and the matte result is obtained by the server through the optimized network model for matte processing of the above target image. The optimized network model is a network model obtained by optimizing the weight parameters in the network model to be optimized. The network model to be optimized is used to determine the first training model and the second training model. The weight parameters in the first training model remain unchanged, and the weight parameters in the second training model can be continuously updated. The first training model is used to determine the first mask based on the incremental image. The first mask is fused with the second mask to obtain the target mask. The second mask corresponds to the incremental image. The target mask is used to update the weight parameters in the second training model to obtain the optimized network model; display the matte result locally on the client side.
[0169] The processor can call the information and application programs stored in the memory through the transmission device to execute the following steps: receive the target image from the client; perform matte processing on the above target image through the optimized network model to obtain the matte result, where the matte result is used to describe the foreground target object included in the above target image, and the optimized network model is a network model obtained by optimizing the weight parameters in the network model to be optimized. The network model to be optimized is used to determine the first training model and the second training model. The weight parameters in the first training model remain unchanged, and the weight parameters in the second training model can be continuously updated. The first training model is used to determine the first mask based on the incremental image. The first mask is fused with the second mask to obtain the target mask. The second mask corresponds to the incremental image. The target mask is used to update the weight parameters in the second training model to obtain the optimized network model; return the matte result to the above client and display the matte result locally on the above client.
[0170] An optimization solution for a network model is provided by using the embodiments of the present application. By determining a first training model and a second training model based on the network model to be optimized, wherein the network model to be optimized is used for matte extraction processing of a target image to obtain a foreground target object included in the target image, the weight parameters in the first training model remain unchanged, and the weight parameters in the second training model can be continuously updated; inputting an incremental image into the first training model to output a first mask; performing mask fusion processing on the first mask and a second mask to obtain a target mask, wherein the second mask corresponds to the incremental image; and updating the weight parameters in the second training model by using the target mask to obtain an optimized network model.
[0171] It is easy to notice that the embodiments of the present application adopt the method of knowledge distillation. When learning new samples each time, the prediction results of the original model are extracted, and the fusion results of the original model are used to guide the iteration of the model weights, reducing the forgetting of the original knowledge by the model, maintaining the stability of the model weight update, and not requiring stock images to participate in the training process, improving the efficiency of network model optimization. Moreover, the embodiments of the present application adopt the method of calculating the error between the mask of the original matte extraction model and the mask of the incremental image in a block-by-block manner. When the error between the two is less than the error threshold, it is considered that the prediction result of the original matte extraction model is available, which can correct the influence of the annotation error in the incremental image on the network model optimization.
[0172] Therefore, the embodiments of the present application achieve the purpose of improving the optimization efficiency of the optimized network model, thereby realizing the technical effect of correcting the influence of the annotation error in the incremental image on the network model optimization, and further solving the technical problem of low optimization efficiency existing in the solutions for optimizing network models in the prior art.
[0173] Those of ordinary skill in the art can understand that Figure 11 the structure shown is only for illustration, and the computer terminal can also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a palm computer, and terminal devices such as Mobile Internet Devices (MID), PAD, etc. Figure 11 It does not limit the structure of the above electronic device. For example, the computer terminal may further include more or fewer components (such as a network interface, a display device, etc.) than those shown in Figure 11 or have a different configuration from that shown in Figure 11 the figure.
[0174] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, and the storage medium can include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, an optical disk, etc.
[0175] Embodiment 5
[0176] According to an embodiment of the present application, an embodiment of a non-volatile storage medium is further provided. Optionally, in this embodiment, the above non-volatile storage medium can be used to store the program code executed by the optimization method of the network model provided in the above Embodiment 1 and the above image processing method.
[0177] Optionally, in this embodiment, the above storage medium can be located in any one of the computer terminals in the computer terminal group in the computer network, or in any one of the mobile terminals in the mobile terminal group.
[0178] Optionally, in this embodiment, the storage medium is set to store the program code for performing the following steps: determining a first training model and a second training model based on the network model to be optimized, where the network model to be optimized is used to perform matte processing on a target image to obtain the foreground target object included in the target image, the weight parameters in the first training model remain unchanged, and the weight parameters in the second training model can be continuously updated; inputting the incremental image into the first training model to output a first mask; performing mask fusion processing on the first mask and a second mask to obtain a target mask, where the second mask corresponds to the incremental image; and updating the weight parameters in the second training model using the target mask to obtain an optimized network model.
[0179] Optionally, in this embodiment, the storage medium is set to store the program code for performing the following steps: performing image block processing of the same size on the first mask and the second mask, dividing the first mask into a plurality of first block regions and dividing the second mask into a plurality of second block regions; performing the same numbering process on the same block regions in the first mask and the second mask to obtain a numbering result; calculating the error between the same numbered block regions in the first mask and the second mask based on an error calculation function to obtain an error result; calculating a plurality of target block regions using a truncation function, the error result, the plurality of first block regions, and the plurality of second block regions; and splicing the plurality of target block regions to obtain the target mask.
[0180] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: when the calculated error is less than the error threshold, the value of the error result is set to 0; when the calculated error is greater than or equal to the error threshold, the value of the error result remains the calculated error.
[0181] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: setting the error result, the error threshold, the error control coefficient, the first value, and the second value as input parameters of the truncation function, and outputting a calculation result, where the first value is less than the second value; calculating the plurality of target block regions by using the calculation result, the plurality of first block regions, and the plurality of second block regions.
[0182] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: calculating a third value by using the error result, the error threshold, and the error control coefficient; when the third value is less than the first value, setting the value of the calculation result to the first value; when the third value is greater than the second value, setting the value of the calculation result to the second value; when the third value is greater than or equal to the first value and less than or equal to the second value, setting the value of the calculation result to the third value.
[0183] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: receiving a currently input target image; sending the target image to a server; receiving a matting result from the server, where the matting result is used to describe a foreground target object included in the target image, and the matting result is obtained by the server performing matting processing on the target image through an optimized network model; and locally displaying the matting result on the client.
[0184] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: receiving a target image from a client; performing matting processing on the target image through an optimized network model to obtain a matting result, where the matting result is used to describe a foreground target object included in the target image; and returning the matting result to the client and locally displaying the matting result on the client.
[0185] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: determining a first training model and a second training model based on a network model to be optimized, where the network model to be optimized is used for matte extraction of a target image, the range of weight parameters in the first training model satisfies a first threshold, and the range of weight parameters in the second training model satisfies a second threshold; inputting an incremental image into the first training model to output a first mask; performing mask fusion processing on the first mask and a second mask to obtain a target mask, where the second mask corresponds to the incremental image; and updating the weight parameters in the second training model using the target mask to obtain an optimized network model.
[0186] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: receiving a currently input target commodity image, where the target commodity image includes a foreground commodity image and a background layout image; sending the target commodity image to a server; receiving a matte extraction result from the server, where the matte extraction result is the foreground commodity image extracted from the target commodity image by the server through matte extraction processing of the target commodity image using the optimized network model; and locally displaying the matte extraction result on the client.
[0187] The serial numbers of the embodiments of the present application are only for description and do not represent the advantages or disadvantages of the embodiments.
[0188] In the above embodiments of the present application, the descriptions of the respective embodiments have their own focuses. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0189] In the several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the above unit division is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces, and the indirect coupling or communication connection of units or modules can be in an electrical or other form.
[0190] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0191] In addition, in each embodiment of the present application, the functional units may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0192] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs that can store program codes.
[0193] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
Claims
1. An optimization method for a network model, characterized in that, Including: Determine a first training model and a second training model based on the network model to be optimized. Among them, the network model to be optimized is used for matte extraction of a target image. The variation range of the weight parameters in the first training model meets a first threshold, and the variation range of the weight parameters in the second training model meets a second threshold. The model structure of the first training model and the model structure of the second training model are the same as the model structure of the network model to be optimized; Input the incremental image into the first training model and output a first mask; Perform mask fusion processing on the first mask and a second mask to obtain a target mask. Among them, the second mask corresponds to the incremental image, the second mask is a preset true mask, the target mask is obtained by splicing a plurality of target block regions, the plurality of target block regions are obtained by using calculation results, the first mask and the second mask, and the calculation results are obtained by inputting an error result, an error threshold, an error control coefficient, a first value and a second value into a truncation function. The error result is obtained by performing error calculation on the first mask and the second mask based on an error calculation function; Update the weight parameters in the second training model by using the target mask to obtain an optimized network model.
2. An optimization method for a network model, characterized in that, Including: Determine a first training model and a second training model based on the network model to be optimized. Among them, the network model to be optimized is used for matte extraction of a target image to obtain a foreground target object included in the target image. The weight parameters in the first training model remain unchanged, and the weight parameters in the second training model can be continuously updated. The model structure of the first training model and the model structure of the second training model are the same as the model structure of the network model to be optimized; Input the incremental image into the first training model and output a first mask; Perform mask fusion processing on the first mask and a second mask to obtain a target mask. Among them, the second mask corresponds to the incremental image, the second mask is a preset true mask, the target mask is obtained by splicing a plurality of target block regions, the plurality of target block regions are obtained by using calculation results, the first mask and the second mask, and the calculation results are obtained by inputting an error result, an error threshold, an error control coefficient, a first value and a second value into a truncation function. The error result is obtained by performing error calculation on the first mask and the second mask based on an error calculation function; Update the weight parameters in the second training model by using the target mask to obtain an optimized network model.
3. The optimization method of the network model according to claim 2, wherein The initial assignments of the weight parameters in the network model to be optimized, the first training model and the second training model are the same.
4. The optimization method of the network model according to claim 2, characterized in that Performing mask fusion processing on the first mask and the second mask to obtain the target mask includes: Perform image block processing of the same size on the first mask and the second mask, divide the first mask into a plurality of first block regions and divide the second mask into a plurality of second block regions; Perform the same numbering process on the same block regions in the first mask and the second mask to obtain a numbering result; Calculate the error between the block regions with the same number in the first mask and the second mask based on the error calculation function to obtain the error result; Calculate the multiple target block regions by using the truncation function, the error result, the multiple first block regions, and the multiple second block regions; Stitch the multiple target block regions to obtain the target mask.
5. The optimization method of the network model according to claim 4, wherein Calculating the error between the block regions with the same number in the first mask and the second mask based on the error calculation function to obtain the error result includes: When the calculated error is less than the error threshold, the value of the error result is set to 0; When the calculated error is greater than or equal to the error threshold, the value of the error result remains the calculated error.
6. The optimization method of the network model according to claim 5, characterized in that Calculating the multiple target block regions by using the truncation function, the error result, the multiple first block regions, and the multiple second block regions includes: Set the error result, the error threshold, the error control coefficient, the first value, and the second value as the input parameters of the truncation function, and output the calculation result, where the first value is less than the second value; Calculate the multiple target block regions by using the calculation result, the multiple first block regions, and the multiple second block regions.
7. The optimization method of the network model according to claim 6, characterized in that, Setting the error result, the error threshold, the error control coefficient, the first value, and the second value as the input parameters and outputting the calculation result includes: Calculate a third value by using the error result, the error threshold, and the error control coefficient; When the third value is less than the first value, set the value of the calculation result to the first value; when the third value is greater than the second value, set the value of the calculation result to the second value; when the third value is greater than or equal to the first value and less than or equal to the second value, set the value of the calculation result to the third value.
8. The optimization method of the network model according to claim 4, wherein The error calculation function includes one of the following: Mean squared error loss function, mean absolute error loss function, cross-entropy loss function.
9. An image processing method, characterized in that, Includes: Receive the currently input target image; Send the target image to the server; Receive the matting result from the server, where the matting result is used to describe the foreground target object included in the target image. The matting result is obtained by the server through matting the target image with an optimized network model. The optimized network model is a network model obtained by optimizing the weight parameters in the network model to be optimized. The network model to be optimized is used to determine a first training model and a second training model. The model structures of the first training model and the second training model are the same as the model structure of the network model to be optimized. The first training model is used to determine a first mask based on an incremental image. The first mask is fused with a second mask to obtain a target mask. The second mask corresponds to the incremental image, and the second mask is a preset true mask. The target mask is used to update the weight parameters in the second training model to obtain the optimized network model. The target mask is obtained by splicing a plurality of target block regions. The plurality of target block regions are obtained using a calculation result, the first mask, and the second mask. The calculation result is obtained by inputting an error result, an error threshold, an error control coefficient, a first value, and a second value into a truncation function. The error result is obtained by calculating the error between the first mask and the second mask based on an error calculation function; Display the matting result locally on the client.
10. The method according to claim 9, wherein The change range of the weight parameters in the first training model satisfies a first threshold, and the change range of the weight parameters in the second training model satisfies a second threshold.
11. The method according to claim 9, wherein The weight parameters in the first training model remain unchanged, and the weight parameters in the second training model can be continuously updated.
12. An image processing method, characterized in that, Include: Receive a target image from the client; Perform matting processing on the target image through an optimized network model to obtain a matting result, where the matting result is used to describe the foreground target object included in the target image. The optimized network model is a network model obtained by optimizing the weight parameters in the network model to be optimized. The network model to be optimized is used to determine a first training model and a second training model. The model structures of the first training model and the second training model are the same as the model structure of the network model to be optimized. The first training model is used to determine a first mask based on an incremental image. The first mask is fused with a second mask to obtain a target mask. The second mask corresponds to the incremental image, and the second mask is a preset true mask. The target mask is used to update the weight parameters in the second training model to obtain the optimized network model. The target mask is obtained by splicing a plurality of target block regions. The plurality of target block regions are obtained using a calculation result, the first mask, and the second mask. The calculation result is obtained by inputting an error result, an error threshold, an error control coefficient, a first value, and a second value into a truncation function. The error result is obtained by calculating the error between the first mask and the second mask based on an error calculation function; Return the matte result to the client and display the matte result locally on the client.
13. The method according to claim 12, wherein The range of change of the weight parameters in the first training model satisfies a first threshold, and the range of change of the weight parameters in the second training model satisfies a second threshold.
14. The method according to claim 12, wherein The weight parameters in the first training model remain unchanged, and the weight parameters in the second training model can be continuously updated.
15. An image processing method, characterized in that, Comprising: Receive the current input target commodity image, where the target commodity image includes: a foreground commodity image and a background layout image; Send the target commodity image to the server; Receive the matte result from the server, where the matte result is the foreground commodity image extracted from the target commodity image after the server performs matte processing on the target commodity image through an optimized network model. The optimized network model is a network model obtained by optimizing the weight parameters in the network model to be optimized. The network model to be optimized is used to determine a first training model and a second training model. The model structures of the first training model and the second training model are the same as the model structure of the network model to be optimized. The first training model is used to determine a first mask based on an incremental image. The first mask is fused with a second mask to obtain a target mask. The second mask corresponds to the incremental image, and the second mask is a preset true mask. The target mask is used to update the weight parameters in the second training model to obtain the optimized network model. The target mask is obtained by splicing a plurality of target block regions. The plurality of target block regions are obtained using a calculation result, the first mask, and the second mask. The calculation result is obtained by inputting an error result, an error threshold, an error control coefficient, a first value, and a second value into a truncation function. The error result is obtained by calculating the error between the first mask and the second mask based on an error calculation function; Display the matte result locally on the client.
16. A non-volatile storage medium, characterized in that, The non-volatile storage medium includes a stored program, where when the program runs, it controls the device where the non-volatile storage medium is located to execute the optimization method of the network model as described in any one of claims 1 to 7, and the image processing method as described in any one of claims 9 to 15.
17. An electronic device, characterized in that, Comprising: A processor; And A memory, connected to the processor, for providing instructions for the processor to perform the following processing steps: Determine a first training model and a second training model based on a network model to be optimized, where the network model to be optimized is used to perform matte processing on a target image to obtain the foreground target object included in the target image. The weight parameters in the first training model remain unchanged, and the weight parameters in the second training model can be continuously updated. The model structures of the first training model and the second training model are the same as the model structure of the network model to be optimized; Input the incremental image into the first training model and output a first mask; Perform mask fusion processing on the first mask and the second mask to obtain a target mask. Among them, the second mask corresponds to the incremental image, the second mask is a preset true mask, the target mask is obtained by splicing a plurality of target block regions, the plurality of target block regions are obtained by using the calculation result, the first mask and the second mask, the calculation result is obtained by inputting the error result, the error threshold, the error control coefficient, the first value and the second value into a truncation function, and the error result is obtained by performing error calculation on the first mask and the second mask based on an error calculation function; Use the target mask to update the weight parameters in the second training model to obtain an optimized network model.
Citation Information
Patent Citations
Certificate photo generation method, client and server
CN111383176A
Image processing method and related product
CN111724407A