A cutout network training method and a cutout method
The lightweight cutout student network is trained through the multi-knowledge distillation method, which solves the problems of large-scale neural network models on edge devices and achieves real-time cutout effect on low-computing devices.
Patent Information
- Application Number
- CN202210282649.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-22
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2042-03-22
AI Technical Summary
Large neural network models are computationally expensive and time-consuming in edge mobile devices, making it difficult to achieve real-time cutouts; while lightweight neural network models have fast inference speed, they have a large loss of accuracy.
The multi-knowledge distillation method is used to train the teacher network and student network based on the annotated and unlabeled sample images to obtain a lightweight student network for target cutouts. This method includes offline distillation, semi-supervised distillation and self-supervised distillation. Through multiple knowledge distillation, the knowledge of large teacher networks is transferred to small student networks, and the accuracy and reasoning speed of student networks are improved.
While ensuring the inference accuracy of the cutout network, it significantly improves the inference speed of the network, so that the cutout network can achieve real-time operation on low-computing devices.
Smart Images

Figure CN114565769B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a cutout network training method and a cutout method. Background Art
[0002] In the field of image processing technology, image cutting is a commonly used processing method.
[0003] With the development of deep learning, the development of image cutout technology has been greatly promoted. In the image cutout method based on deep learning, the neural network model is used as an image feature extraction module, which can better represent the image and thus achieve good image cutout effects.
[0004] However, large neural network models require a lot of computational resources and are time-consuming in edge mobile devices with low computational workloads, making it difficult to achieve real-time image cutouts. Lightweight (or small) neural network models have fewer model parameters and improve reasoning speed, but their accuracy is greatly reduced due to the reduction in parameters. Summary of the invention
[0005] In view of this, the embodiments of the present application provide a cutout network training method and a cutout method, which can solve at least one technical problem in the related art.
[0006] In a first aspect, an embodiment of the present application provides a method for training a cutout network, comprising: obtaining labeled and unlabeled sample images, wherein the sample images include a target object, and the label is a result of cutout of the target object of the sample image; and using multiple knowledge distillation methods to train a teacher network and a student network based on the sample images to obtain a lightweight student network for cutout of the target object.
[0007] The embodiment of the present application utilizes multiple knowledge distillations to obtain a lightweight student network for target object cutout based on a teacher network, thereby improving the speed of network reasoning while also ensuring the accuracy of network reasoning.
[0008] In some embodiments, the method of utilizing multiple knowledge distillations to train the teacher network and the student network based on the sample image to obtain a lightweight student network for target object cutout includes: first using offline distillation and then using self-supervised distillation to train the teacher network and the student network based on the sample image to obtain a lightweight student network for target object cutout.
[0009] In some embodiments, the method of first using offline distillation and then using self-supervised distillation to train the teacher network and the student network based on the sample images to obtain a lightweight student network for target object cutout, includes: first using offline distillation to train the teacher network and the student network using labeled sample images to obtain a trained student network; then using self-supervised distillation to replace the teacher network with the student network trained by offline distillation, randomly initializing the parameters of the student network, and training the teacher network and the student network based on labeled and unlabeled sample images to obtain a student network for target object cutout.
[0010] In some embodiments, the method of utilizing multiple knowledge distillations to train a teacher network and a student network based on the sample image to obtain a lightweight student network for target object cutout includes: first using semi-supervised distillation and then using self-supervised distillation to train a teacher network and a student network based on the sample image to obtain a lightweight student network for target object cutout.
[0011] In some embodiments, the method of first using semi-supervised distillation and then self-supervised distillation to train the teacher network and the student network based on the sample images to obtain a lightweight student network for target object cutout, including: first using semi-supervised distillation to train the teacher network and the student network using unlabeled sample images to obtain a trained student network; then using self-supervised distillation to replace the teacher network with the student network trained by the semi-supervised distillation method, randomly initializing the parameters of the student network, and training the teacher network and the student network based on labeled and unlabeled sample images to obtain a student network for target object cutout.
[0012] In some embodiments, the method of utilizing multiple knowledge distillations to train the teacher network and the student network based on the sample image to obtain a lightweight student network for target object cutout includes: first using offline distillation, then using semi-supervised distillation, and finally using self-supervised distillation to train the teacher network and the student network based on the sample image to obtain a lightweight student network for target object cutout.
[0013] In some embodiments, the method of first using offline distillation, then using semi-supervised distillation, and finally using self-supervised distillation to train the teacher network and the student network based on the sample images to obtain a lightweight student network for target object cutout, includes: first using offline distillation to train the teacher network and the student network using labeled sample images to obtain a trained student network; then using semi-supervised distillation to keep the teacher network unchanged, updating the parameters of the student network to the parameters of the student network learned by offline distillation, and using unlabeled sample images to train the teacher network and the student network to obtain a trained student network; finally using self-supervised distillation to replace the teacher network with the student network trained by semi-supervised distillation, randomly initializing the parameters of the student network, and training the teacher network and the student network based on labeled and unlabeled sample images to obtain a student network for target object cutout.
[0014] In a second aspect, an embodiment of the present application provides a cutout method, comprising: acquiring a current image including a target object; inputting the current image into a student network for cutout of the target object obtained by the cutout network training method described in any embodiment of the first aspect, and outputting a cutout result corresponding to the current image.
[0015] In a third aspect, an embodiment of the present application provides a cutout network training device, comprising: an acquisition unit, for acquiring labeled and unlabeled sample images, wherein the sample images include a target object, and the annotation is a result of cutout of the target object of the sample image; a training execution unit, for training a teacher network and a student network based on the sample images by using a multiple knowledge distillation method, to obtain a lightweight student network for cutout of the target object.
[0016] In a fourth aspect, an embodiment of the present application provides a cutout device, comprising: an acquisition unit, for acquiring a current image including a target object; a cutout execution unit, for inputting the current image into a student network for cutout of the target object obtained by the cutout network training device described in the embodiment of the third aspect, and outputting a cutout result corresponding to the current image.
[0017] In a fifth aspect, an embodiment of the present application provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, it implements the cutout network training method as described in any embodiment of the first aspect; or implements the cutout method as described in any embodiment of the second aspect.
[0018] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for training a cutout network as described in any embodiment of the first aspect is implemented; or the method for cutting out as described in any embodiment of the second aspect is implemented.
[0019] In the seventh aspect, an embodiment of the present application provides a computer program product. When the computer program product runs on an electronic device, the electronic device executes the cutout network training method as described in any embodiment of the first aspect; or executes the cutout method as described in any embodiment of the second aspect.
[0020] It should be understood that the beneficial effects of the second to seventh aspects can be found in the relevant description of the embodiment of the first aspect and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0022] Figure 1 It is a schematic diagram of the implementation process of a cutout network training method provided in one embodiment of the present application;
[0023] Figure 2 This is a schematic diagram of an implementation flow of step S120 provided in an embodiment of the present application;
[0024] Figure 3 is a schematic diagram of an implementation flow of step S120 provided in another embodiment of the present application;
[0025] Figure 4 is a schematic diagram of an implementation flow of step S120 provided in another embodiment of the present application;
[0026] Figure 5 This is a schematic diagram of an implementation process of a cutout method provided in an embodiment of the present application;
[0027] Figure 6 It is a structural schematic diagram of a cutout network training device provided in one embodiment of the present application;
[0028] Figure 7 It is a structural schematic diagram of a cutout device provided in one embodiment of the present application;
[0029] Figure 8 It is a structural schematic diagram of an electronic device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0030] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.
[0031] The term "and / or" as used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0032] "One embodiment" or "some embodiments" described in the specification of this application means that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the sentences "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0033] In addition, in the description of the present application, "plurality" means two or more. The terms "first" and "second" are only used to distinguish the description and cannot be understood as indicating or implying relative importance.
[0034] In order to illustrate the technical solution described in this application, a specific embodiment is provided below for illustration.
[0035] At present, large neural network models have a large amount of computation and require a lot of computing resources. In edge mobile devices with low computational capacity, it is time-consuming and difficult to achieve real-time image cutout. Although lightweight (or small) neural network models have fewer model parameters and improve reasoning speed, their accuracy is greatly reduced due to the reduction in parameters.
[0036] In order to solve the technical problems of related technologies, the embodiments of the present application provide a method and device for training a cutout network, which realizes a lightweight cutout network based on multiple knowledge distillation. Specifically, by using multiple knowledge distillation, a small number of labeled images and a large number of unlabeled images are used for training to obtain a lightweight cutout network, which improves the inference speed of the cutout network while ensuring the inference accuracy.
[0037] Figure 1This is a schematic diagram of the implementation process of a cutout network training method provided in an embodiment of the present application. The cutout network training method in this embodiment can be executed by an electronic device. The electronic device includes but is not limited to a computer, a tablet computer, a server, a mobile phone, a camera or a wearable device. Among them, the server includes but is not limited to an independent server or a cloud server. Figure 1 As shown, the cutout network training method may include steps S110 to S130.
[0038] S110, obtaining labeled and unlabeled sample images.
[0039] The sample image includes a target object, and the annotation is a result of clipping the target object of the sample image.
[0040] In the embodiment of the present application, the sample image includes two parts, the first part is the image data with annotations, and the other part is the image data without annotations. The image data with annotations and the image data without annotations together constitute the sample image set.
[0041] In some embodiments, the labeled sample images account for a small part of the sample image set, while the unlabeled image data accounts for the majority of the sample image set, that is, the labeled image data is far less than the unlabeled image data. This can reduce the labeling cost, reduce the difficulty of obtaining the sample image set, and improve efficiency.
[0042] In order to better describe the technical solution of the embodiment of the present application, the target object is a portrait, and the cutout network is a portrait cutout network as an exemplary description in the embodiment of the present application. It should be understood that the portrait cutout is only an exemplary description. More generally, the cutout network can be a network that realizes the cutout of any target object, and the present application does not specifically limit this.
[0043] In some possible implementations, annotated portrait images and unannotated portrait images can be collected by a collection device. Specifically, images including portraits are collected by a collection device such as a camera or a webcam, and are manually or automatically annotated to obtain annotated sample images; if no annotation is performed, an unannotated sample image is obtained.
[0044] In some other possible implementations, the annotated portrait images and the unannotated portrait images may be readily available images on the Internet. Specifically, portrait images may be crawled from the Internet, such as the World Wide Web, using a crawler tool. The portrait images may include annotated portrait images, thereby obtaining annotated sample images; the portrait images may include unannotated portrait images, thereby obtaining unannotated sample images.
[0045] In some other possible implementations, the labeled sample images and the unlabeled sample images may be collected by a collection device or may be ready-made images obtained online.
[0046] S120, using multiple knowledge distillation methods, train the teacher network and the student network based on the sample images to obtain a lightweight student network for target object cutout.
[0047] Step S120 is the training stage.
[0048] The teacher network is a pre-trained teacher network for target object cutout. The teacher network can use a model based on portrait semantic segmentation, such as FCN net, deeplab v3+, etc. The teacher network has two outputs, namely the target object cutout result and the feature map of the last convolution layer, i.e. the feature result.
[0049] The student network is a lightweight model for object cutout. During the training phase, the network has two outputs, namely the portrait cutout result and the feature map of the last convolutional layer, i.e., the feature result; during the inference phase, the network output is only the object cutout result.
[0050] In the embodiment of the present application, at least two knowledge distillations are used to train the teacher network and the student network based on sample images to obtain a lightweight student network for target object cutout. Information is transferred from the large teacher network to the small student network for training, so that the student network imitates the teacher network and obtains similar or even superior performance. While reducing the amount of network model calculation, the accuracy of network model reasoning is improved, which facilitates the deployment of the cutout network in low-computing electronic devices.
[0051] In some embodiments, knowledge distillation includes but is not limited to offline distillation, semi-supervised distillation, and self-supervised distillation, etc.
[0052] In one embodiment, Figure 2 As shown, step S120 includes: first adopting offline distillation and then adopting self-supervised distillation to train the teacher network and the student network based on the sample image to obtain a lightweight student network for target object cutout.
[0053] As a possible implementation method, an offline distillation method is first used to train the teacher network and the student network using labeled sample images to obtain a trained student network; then a self-supervised distillation method is used to replace the teacher network with the trained student network obtained by the offline distillation method, the parameters of the student network are randomly initialized, and the teacher network and the student network are trained based on labeled and unlabeled sample images to obtain a student network for target object cutout.
[0054] As a non-limiting example, an offline distillation training process is first performed, specifically including: initializing the parameters of the teacher network to pre-trained parameters and fixing the model parameters; randomly initializing the parameters of the student network. The sample image is input into the teacher network, and after the teacher network is inferred, the feature results and the target object cutout results are obtained. The same sample image is input into the student network, and after the student network is inferred, the feature results and the target object cutout results are also obtained. The first loss is calculated based on the feature results inferred by the teacher network and the student network, and the second loss is calculated based on the target object cutout results inferred by the teacher network and the student network. The third loss is also calculated based on the target object cutout results and annotations inferred by the student network. The three losses are added or weighted summed as the total loss of training. Using the total loss, back propagation is performed to learn the model parameters of the student network.
[0055] Then, the training process of the self-supervised distillation method is carried out, which specifically includes: first replacing the teacher network (including the model and its parameters) with the student network (including the model and its parameters) learned by the offline distillation method, keeping the model of the student network unchanged, and randomly initializing its parameters. Then, the sample image is randomly input into the teacher network, and after the inference of the teacher network, the feature result and the target object cutout result are obtained. The same sample image is input into the student network, and after the inference of the student network, the feature result and the target object cutout result are also obtained. If the input sample image belongs to a labeled sample image, the first loss is calculated according to the feature result inferred by the teacher network and the student network, the second loss is calculated according to the target object cutout result inferred by the teacher network and the student network, and the third loss is calculated according to the target object cutout result and the label inferred by the student model, and the three losses are added or weighted summed as the total loss of training; if the input sample image belongs to an unlabeled sample image, the first loss is calculated according to the feature result inferred by the teacher network and the student network, and the second loss is calculated according to the portrait cutout result inferred by the teacher network and the student network, and the two losses are added or weighted summed as the total loss of training. Using the total loss, back propagation is performed to learn the model parameters of the student network.
[0056] In another embodiment, Figure 3 As shown, step S120 includes: first using semi-supervised distillation and then using self-supervised distillation to train the teacher network and the student network based on the sample image to obtain a lightweight student network for target object cutout.
[0057] As a possible implementation method, a semi-supervised distillation method is first used to train the teacher network and the student network using unlabeled sample images to obtain a trained student network; then a self-supervised distillation method is used to replace the teacher network with the student network trained by the semi-supervised distillation method, the parameters of the student network are randomly initialized, and the teacher network and the student network are trained based on labeled and unlabeled sample images to obtain a student network for target object cutout.
[0058] As a non-limiting example, a semi-supervised distillation training process is first performed, specifically including: initializing the parameters of the teacher network to pre-trained parameters and fixing the model parameters; randomly initializing the parameters of the student network. The sample image is input into the teacher network, and after the teacher network is inferred, the feature results and the target object cutout results are obtained. The same sample image is input into the student network, and after the student network is inferred, the feature results and the target object cutout results are also obtained. The first loss is calculated based on the feature results inferred by the teacher network and the student network, and the second loss is calculated based on the target object cutout results inferred by the teacher network and the student network. The two losses are added or weighted summed as the total loss of training. Using the total loss, back propagation is performed to learn the model parameters of the student network.
[0059] Then, the training process of the self-supervised distillation method is carried out, which specifically includes: first replacing the teacher network (including the model and its parameters) with the student network (including the model and parameters) learned by the semi-supervised distillation method, keeping the model of the student network unchanged, and randomly initializing its parameters. Then, the sample image is randomly input into the teacher network, and after the inference of the teacher network, the feature result and the target object cutout result are obtained. The same sample image is input into the student network, and after the inference of the student network, the feature result and the target object cutout result are also obtained. If the input sample image belongs to a labeled sample image, the first loss is calculated according to the feature result inferred by the teacher network and the student network, the second loss is calculated according to the target object cutout result inferred by the teacher network and the student network, and the third loss is calculated according to the target object cutout result and the label inferred by the student model, and the three losses are added or weighted summed as the total loss of training; if the input sample image belongs to an unlabeled sample image, the first loss is calculated according to the feature result inferred by the teacher network and the student network, and the second loss is calculated according to the portrait cutout result inferred by the teacher network and the student network, and the two losses are added or weighted summed as the total loss of training. Using the total loss, back propagation is performed to learn the model parameters of the student network.
[0060] In another embodiment, Figure 4As shown, step S120 includes: first using offline distillation, then using semi-supervised distillation, and finally using self-supervised distillation to train the teacher network and the student network based on the sample image to obtain a lightweight student network for target object cutout.
[0061] As a possible implementation method, an offline distillation method is first used to train the teacher network and the student network using labeled sample images to obtain a trained student network; then a semi-supervised distillation method is used to keep the teacher network (including the model and its parameters) unchanged, and the parameters of the student network are updated to the parameters of the student network learned by the offline distillation method, and then the teacher network and the student network are trained using unlabeled sample images to obtain a trained student network; finally, a self-supervised distillation method is used to replace the teacher network with the student network trained by the semi-supervised distillation method, the parameters of the student network are randomly initialized, and the teacher network and the student network are trained based on labeled and unlabeled sample images to obtain a student network for target object cutout.
[0062] As a non-limiting example, an offline distillation training process is first performed, specifically including: initializing the parameters of the teacher network to pre-trained parameters and fixing the model parameters; randomly initializing the parameters of the student network. The sample image is input into the teacher network, and after the teacher network is inferred, the feature results and the target object cutout results are obtained. The same sample image is input into the student network, and after the student network is inferred, the feature results and the target object cutout results are also obtained. The first loss is calculated based on the feature results inferred by the teacher network and the student network, and the second loss is calculated based on the target object cutout results inferred by the teacher network and the student network. The third loss is also calculated based on the target object cutout results and annotations inferred by the student network. The three losses are added or weighted summed as the total loss of training. Using the total loss, back propagation is performed to learn the model parameters of the student network.
[0063] Then, a semi-supervised distillation training process is carried out, which specifically includes: first keeping the teacher network (including the model and its parameters) unchanged, and updating the parameters of the student network to the parameters of the student network learned by the offline distillation method. Then the sample image is input into the teacher network, and after the teacher network is inferred, the feature results and the target object cutout results are obtained. The same sample image is input into the student network, and after the student network is inferred, the feature results and the target object cutout results are also obtained. The first loss is calculated based on the feature results inferred by the teacher network and the student network, and the second loss is calculated based on the target object cutout results inferred by the teacher network and the student network. The two losses are added or weighted summed as the total loss of training. Using the total loss, back propagation is performed to learn the model parameters of the student network.
[0064] Finally, the training process of the self-supervised distillation method is carried out, which specifically includes: first replacing the teacher network (including the model and its parameters) with the student network (including the model and its parameters) learned by the semi-supervised distillation method, keeping the model of the student network unchanged, and randomly initializing its parameters. Then the sample image is randomly input into the teacher network, and after the inference of the teacher network, the feature result and the target object cutout result are obtained. The same sample image is input into the student network, and after the inference of the student network, the feature result and the target object cutout result are also obtained. If the input sample image belongs to a labeled sample image, the first loss is calculated according to the feature result inferred by the teacher network and the student network, the second loss is calculated according to the target object cutout result inferred by the teacher network and the student network, and the third loss is calculated according to the target object cutout result and the label inferred by the student model, and the three losses are added or weighted summed as the total loss of training; if the input sample image belongs to an unlabeled sample image, the first loss is calculated according to the feature result inferred by the teacher network and the student network, and the second loss is calculated according to the portrait cutout result inferred by the teacher network and the student network, and the two losses are added or weighted summed as the total loss of training. Using the total loss, back propagation is performed to learn the model parameters of the student network.
[0065] Another embodiment of the present application provides a cutout method. The cutout method can be applied to electronic devices. Figure 5 As shown, the cutout method may include steps S510 to S520. It should be noted that the cutout method embodiment is the same or similar to the above embodiment, which will not be described here in detail, please refer to the above.
[0066] S510, capturing a current image including a target object.
[0067] The electronic device collects the current image including the target object through the collection device. In some embodiments, the collection device may include a camera or a camera.
[0068] In some embodiments, the collection device may be independent of the electronic device. In other embodiments, the collection device may also be integrated into the electronic device. This application is not limited to this.
[0069] S520, input the current image into the cutout network, and output the cutout result of the target object corresponding to the current image.
[0070] The electronic device is pre-deployed with a cutout network, which can be a student network for target object cutout obtained by the cutout network training method of the aforementioned embodiment.
[0071] The embodiment of the present application implements an end-to-end target object cutout solution based on a cutout network.
[0072] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0073] An embodiment of the present application also provides a cutout network training device. For details not described in the cutout network training device, please refer to the description of the cutout network training method embodiment above.
[0074] See also Figure 6 , Figure 6 6 is a schematic block diagram of a cutout network training device provided in an embodiment of the present application. The cutout network training device comprises: an acquisition unit 61 and a training execution unit 62.
[0075] The acquisition unit 61 is used to acquire sample images with or without annotations, the sample images include target objects, and the annotations are the results of clipping the target objects in the sample images.
[0076] The training execution unit 62 is used to train the teacher network and the student network based on the sample images by using multiple knowledge distillation methods to obtain a lightweight student network for target object cutout.
[0077] In some embodiments, the training execution unit 62 is specifically used to: first use offline distillation and then use self-supervised distillation to train the teacher network and the student network based on the sample image to obtain a lightweight student network for target object cutout.
[0078] As a possible implementation method, the training execution unit 62 is specifically used to: first adopt an offline distillation method to train the teacher network and the student network using labeled sample images to obtain a trained student network; then adopt a self-supervised distillation method to replace the teacher network with the student network trained by the offline distillation method, randomly initialize the parameters of the student network, and train the teacher network and the student network based on labeled and unlabeled sample images to obtain a student network for target object clipping.
[0079] In some embodiments, the training execution unit 62 is specifically used to: first use semi-supervised distillation and then use self-supervised distillation to train the teacher network and the student network based on the sample image to obtain a lightweight student network for target object cutout.
[0080] As a possible implementation method, the training execution unit 62 is specifically used to: first adopt a semi-supervised distillation method to train the teacher network and the student network using unlabeled sample images to obtain a trained student network; then adopt a self-supervised distillation method to replace the teacher network with the student network trained by the semi-supervised distillation method, randomly initialize the parameters of the student network, and train the teacher network and the student network based on labeled and unlabeled sample images to obtain a student network for target object clipping.
[0081] In some embodiments, the training execution unit 62 is specifically used to: first use offline distillation, then use semi-supervised distillation, and finally use self-supervised distillation to train the teacher network and the student network based on the sample image to obtain a lightweight student network for target object cutout.
[0082] As a possible implementation method, the training execution unit 62 is specifically used to: first adopt an offline distillation method to train the teacher network and the student network using labeled sample images to obtain a trained student network; then adopt a semi-supervised distillation method to keep the teacher network unchanged, update the parameters of the student network to the parameters of the student network learned by the offline distillation method, and use unlabeled sample images to train the teacher network and the student network to obtain a trained student network; finally, adopt a self-supervised distillation method to replace the teacher network with the student network trained by the semi-supervised distillation method, randomly initialize the parameters of the student network, and train the teacher network and the student network based on labeled and unlabeled sample images to obtain a student network for target object clipping.
[0083] An embodiment of the present application also provides a cutout device. For details not described in the cutout device, please refer to the description of the cutout method embodiment above.
[0084] See also Figure 7 , Figure 7 1 is a schematic block diagram of a cutout device provided in an embodiment of the present application. The cutout device comprises: a collection unit 71 and a cutout execution unit 72.
[0085] The acquisition unit 71 is used to acquire the current image including the target object.
[0086] The cutout execution unit 72 inputs the current image into the student network for target object cutout obtained by the aforementioned cutout network training device, and outputs the cutout result corresponding to the current image.
[0087] An embodiment of the present application also provides an electronic device, such as Figure 8 As shown, the electronic device may include one or more processors 100 ( Figure 8Only one is shown), a memory 101 and a computer program 102 stored in the memory 101 and executable on one or more processors 100, for example, a program for training a cutout network and / or a program for cutting out an image. When one or more processors 100 execute the computer program 102, the various steps in the cutout network training method and / or the cutout method embodiment can be implemented. Alternatively, when one or more processors 100 execute the computer program 102, the functions of the various modules / units in the cutout network training device and / or the cutout device embodiment can be implemented, which is not limited here.
[0088] Those skilled in the art will understand that Figure 8 The electronic device is merely an example of an electronic device and does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the electronic device may also include input and output devices, network access devices, buses, etc.
[0089] In one embodiment, the processor 100 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0090] In one embodiment, the memory 101 may be an internal storage unit of an electronic device, such as a hard disk or memory of the electronic device. The memory 101 may also be an external storage device of the electronic device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device. Further, the memory 101 may also include both an internal storage unit of the electronic device and an external storage device. The memory 101 is used to store computer programs and other programs and data required by the electronic device. The memory 101 may also be used to temporarily store data that has been output or is to be output.
[0091] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.
[0092] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the cutout network training method and / or the cutout method embodiment can be implemented.
[0093] An embodiment of the present application provides a computer program product. When the computer program product is run on an electronic device, the electronic device can implement the steps in the cutout network training method and / or the cutout method embodiment.
[0094] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0095] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0096] In the embodiments provided in the present application, it should be understood that the disclosed devices / electronic devices and methods can be implemented in other ways. For example, the device / electronic device embodiments described above are merely schematic. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0097] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0098] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0099] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, and the computer program that can be completed by instructing the relevant hardware through a computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device that can carry computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electric carrier signals and telecommunication signals.
[0100] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A method for training a cutout network. It is characterized in that include: Acquire sample images with annotations and without annotations, wherein the sample images include a target object, and the annotations are the target object cutout results of the sample images; By using multiple knowledge distillation methods, the teacher network and the student network are trained based on the sample image to obtain a lightweight student network for target object cutout, including: Firstly, offline distillation is adopted and then self-supervised distillation is adopted to train the teacher network and the student network based on the sample image, so as to obtain a lightweight student network for target object cutout; Alternatively, semi-supervised distillation is first used and then self-supervised distillation is used to train the teacher network and the student network based on the sample image to obtain a lightweight student network for target object cutout; Or, firstly adopt offline distillation, then adopt semi-supervised distillation, and finally adopt self-supervised distillation to train the teacher network and the student network based on the sample image to obtain a lightweight student network for target object cutout; The method of first adopting offline distillation and then adopting self-supervised distillation to train the teacher network and the student network based on the sample image to obtain a lightweight student network for target object cutout includes: First, the teacher network and the student network are trained using labeled sample images by using the offline distillation method to obtain a trained student network. Then, the teacher network is replaced with the student network trained by the offline distillation method by using the self-supervised distillation method, the parameters of the student network are randomly initialized, and the teacher network and the student network are trained based on labeled and unlabeled sample images to obtain a student network for target object cutout. The method of first using semi-supervised distillation and then using self-supervised distillation to train the teacher network and the student network based on the sample image to obtain a lightweight student network for target object cutout includes: First, the semi-supervised distillation method is used to train the teacher network and the student network using unlabeled sample images to obtain a trained student network; then the self-supervised distillation method is used to replace the teacher network with the student network trained by the semi-supervised distillation method, the parameters of the student network are randomly initialized, and the teacher network and the student network are trained based on labeled and unlabeled sample images to obtain a student network for target object cutout; The method of first adopting offline distillation, then adopting semi-supervised distillation, and finally adopting self-supervised distillation to train the teacher network and the student network based on the sample image to obtain a lightweight student network for target object cutout includes: First, the offline distillation method is used to train the teacher network and the student network using labeled sample images to obtain a trained student network; then the semi-supervised distillation method is used to keep the teacher network unchanged, and the parameters of the student network are updated to the parameters of the student network learned by the offline distillation method, and the teacher network and the student network are trained using unlabeled sample images to obtain a trained student network; finally, the self-supervised distillation method is used to replace the teacher network with the student network trained by the semi-supervised distillation method, and the parameters of the student network are randomly initialized. The teacher network and the student network are trained based on labeled and unlabeled sample images to obtain a student network for target object cutout.
2. A method of cutting out a picture, It is characterized in that include: Acquiring a current image including the target object; The current image is input into the student network for target object cutout obtained by the cutout network training method as claimed in claim 1, and the cutout result corresponding to the current image is output.
3. A cutout network training device, It is characterized in that include: An acquisition unit, used for acquiring sample images with annotations and without annotations, wherein the sample images include a target object, and the annotation is a result of clipping the target object in the sample image; The training execution unit is used to train the teacher network and the student network based on the sample image by using multiple knowledge distillation methods to obtain a lightweight student network for target object cutout, including: Firstly, offline distillation is adopted and then self-supervised distillation is adopted to train the teacher network and the student network based on the sample image, so as to obtain a lightweight student network for target object cutout; Alternatively, semi-supervised distillation is first used and then self-supervised distillation is used to train the teacher network and the student network based on the sample image to obtain a lightweight student network for target object cutout; Or, firstly adopt offline distillation, then adopt semi-supervised distillation, and finally adopt self-supervised distillation to train the teacher network and the student network based on the sample image to obtain a lightweight student network for target object cutout; The method of first adopting offline distillation and then adopting self-supervised distillation to train the teacher network and the student network based on the sample image to obtain a lightweight student network for target object cutout includes: First, the teacher network and the student network are trained using labeled sample images by using the offline distillation method to obtain a trained student network. Then, the teacher network is replaced with the student network trained by the offline distillation method by using the self-supervised distillation method, the parameters of the student network are randomly initialized, and the teacher network and the student network are trained based on labeled and unlabeled sample images to obtain a student network for target object cutout. The method of first using semi-supervised distillation and then using self-supervised distillation to train the teacher network and the student network based on the sample image to obtain a lightweight student network for target object cutout includes: First, the semi-supervised distillation method is used to train the teacher network and the student network using unlabeled sample images to obtain a trained student network; then the self-supervised distillation method is used to replace the teacher network with the student network trained by the semi-supervised distillation method, the parameters of the student network are randomly initialized, and the teacher network and the student network are trained based on labeled and unlabeled sample images to obtain a student network for target object cutout; The method of first adopting offline distillation, then adopting semi-supervised distillation, and finally adopting self-supervised distillation to train the teacher network and the student network based on the sample image to obtain a lightweight student network for target object cutout includes: First, the offline distillation method is used to train the teacher network and the student network using labeled sample images to obtain a trained student network; then the semi-supervised distillation method is used to keep the teacher network unchanged, and the parameters of the student network are updated to the parameters of the student network learned by the offline distillation method, and the teacher network and the student network are trained using unlabeled sample images to obtain a trained student network; finally, the self-supervised distillation method is used to replace the teacher network with the student network trained by the semi-supervised distillation method, and the parameters of the student network are randomly initialized. The teacher network and the student network are trained based on labeled and unlabeled sample images to obtain a student network for target object cutout.
4. A cutout device, It is characterized in that include: A collection unit, used for collecting a current image including a target object; A cutout execution unit is used to input the current image into the student network for target object cutout obtained by the cutout network training device as described in claim 3, and output the cutout result corresponding to the current image.
5. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, It is characterized in that When the processor executes the computer program, it implements the cutout network training method as claimed in claim 1, or implements the cutout method as claimed in claim 2.
6. A computer storage medium storing a computer program, It is characterized in that When the computer program is executed by a processor, it implements the cutout network training method as claimed in claim 1, or implements the cutout method as claimed in claim 2.
Citation Information
Patent Citations
Face feature detection method and device, medium and electronic equipment
CN113011356A