Deep learning model training method, image processing method, device and equipment
By combining the deep learning architecture of the first model and the second model, the training method of the deep learning model is optimized, and the cost-effective manual annotation problem is solved, efficient image segmentation in different scenarios is achieved, and the generalization ability and labeling accuracy of the model are improved.
Patent Information
- Application Number
- CN202210249199.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-14
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-03-14
AI Technical Summary
现有深度学习模型在图像分割任务中存在高成本、低效率的人工标注需求,且难以在不同场景下泛化,导致模型泛化能力差和灾难性遗忘问题。
Using a deep learning model architecture including the first model and the second model, by inputting the initial sample image into the first model for preliminary processing, training the second model using the initial sample image and the first process image, combining the second model with the continuous learning function, the deep learning model is optimized to adapt to different labeling scenarios.
It improves the generalization ability of the model, reduces the cost of manual labeling, improves the efficiency and accuracy of labeling, and can adapt to new task requirements in different scenarios.
Smart Images

Figure CN114627343B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to the field of deep learning, and specifically to a training method, image processing method, device, electronic device and storage medium for a deep learning model. Background Art
[0002] Deep learning, also known as deep structured learning or hierarchical learning, is part of a broader family of machine learning methods based on artificial neural networks. Deep learning architectures, such as deep neural networks, deep belief networks, recurrent neural networks, and convolutional neural networks, have been applied to fields including computer vision, image segmentation, speech recognition, natural language processing, audio recognition, social network filtering, machine translation, bioinformatics, drug design, medical image analysis, material inspection, and board game programs. To ensure the accuracy of the output results in each field, corresponding model training is essential. Summary of the invention
[0003] The present disclosure provides a deep learning model training method, image processing method, device, electronic device and storage medium.
[0004] According to one aspect of the present disclosure, a deep learning model training method is provided, including: a deep learning model training method, the deep learning model includes a first model and a second model, the method includes: inputting an initial sample image into the first model to obtain a first processed image; using the initial sample image and the first processed image to train the second model; and determining the deep learning model based on the first model and the trained second model.
[0005] According to another aspect of the present disclosure, there is provided an image processing method, comprising: inputting an image to be processed into a deep learning model to obtain a third processed image; the deep learning model is trained according to the deep learning model training method of the present disclosure.
[0006] According to another aspect of the present disclosure, a training device for a deep learning model is provided, wherein the deep learning model includes a first model and a second model, and the device includes: a first acquisition module, used to input an initial sample image into the first model to obtain a first processed image; a training module, used to train the second model using the initial sample image and the first processed image; and a determination module, used to determine the deep learning model based on the first model and the trained second model.
[0007] According to another aspect of the present disclosure, an image processing device is provided, comprising: a third acquisition module, used to input an image to be processed into a deep learning model to obtain a third processed image; the deep learning model is trained according to the deep learning model training device of the present disclosure.
[0008] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the deep learning model training method and image processing method of the present disclosure.
[0009] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the deep learning model training method and image processing method of the present disclosure.
[0010] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, which, when executed by a processor, implements the deep learning model training method and image processing method of the present disclosure.
[0011] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.
[0013] Figure 1 An exemplary system architecture of a training method, an image processing method, and an apparatus to which a deep learning model can be applied according to an embodiment of the present disclosure is schematically shown;
[0014] Figure 2 A flowchart of a deep learning model training method according to an embodiment of the present disclosure is schematically shown;
[0015] Figure 3 A flowchart of a training method for an interactive deep learning model according to an embodiment of the present disclosure is schematically shown;
[0016] Figure 4 The flowchart of the image processing method according to the embodiment of the present disclosure is schematically shown;
[0017] Figure 5A schematic diagram schematically shows a schematic diagram of optimizing training of an interactive segmentation model with continuous learning capability according to an embodiment of the present disclosure;
[0018] Figure 6 A block diagram of a deep learning model training device according to an embodiment of the present disclosure is schematically shown;
[0019] Figure 7 A block diagram schematically shows a training device for a deep learning model according to an embodiment of the present disclosure; and
[0020] Figure 8 A schematic block diagram of an example electronic device that can be used to implement embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0021] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0022] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, disclosure and application of user personal information involved comply with the provisions of relevant laws and regulations, take necessary confidentiality measures, and do not violate public order and good morals.
[0023] In the technical solution of the present disclosure, the user's authorization or consent is obtained before obtaining or collecting the user's personal information.
[0024] Image segmentation technology is widely used in daily life, such as in autonomous driving lane segmentation, video conferencing portrait cutouts, medical image segmentation, and many other scenarios. With the continuous development of deep learning technology, segmentation accuracy is also constantly improving, further accelerating the application of image segmentation technology to various fields of life. To improve the accuracy of the image segmentation model, it is necessary to use a large amount of high-quality annotated data to train the neural network used to implement the image segmentation model.
[0025] In the process of realizing the concept of the present disclosure, the inventor found that manual labeling is mainly carried out by manually collecting points and pulling lines, which is costly and inefficient. It is difficult for individuals or small businesses to obtain large-scale high-quality data, which affects the implementation of segmentation technology and raises the entry threshold of the industry. In the case of complex labeling scenarios, labelers also need to spend a lot of time. The interactive segmentation model implemented based on the offline learning method will fix the entire network parameters when the user labels, and only forward prediction is performed to obtain the results, so that the model cannot learn and has poor generalization ability. In this case, the user can only label scenes similar to the training data set, and cannot migrate the model to other scenes with large differences. If the model is generalized to other fields, the problem of being unable to label will arise. The interactive segmentation model implemented based on the online learning method directly optimizes the model parameters or activates information through user clicks, thereby changing the model output. This learning method will update the network parameters of the model regardless of the labeling scenario, and the update is frequent, resulting in catastrophic forgetting problems in the model, and it cannot adapt to labeling tasks with large differences. After the model adapts to the current labeling task, when labeling the categories that were previously able to be labeled and performed well, the performance of the model will decrease.
[0026] The present disclosure provides a training method, an image processing method, an apparatus, an electronic device, and a storage medium for a deep learning model. The deep learning model includes a first model and a second model, and the training method includes: inputting an initial sample image into the first model to obtain a first processed image; using the initial sample image and the first processed image to train the second model; and determining the deep learning model based on the first model and the trained second model.
[0027] Figure 1 An exemplary system architecture of a training method, an image processing method, and an apparatus to which a deep learning model can be applied according to an embodiment of the present disclosure is schematically illustrated.
[0028] It should be noted that Figure 1 The examples shown are only examples of system architectures to which the embodiments of the present disclosure can be applied, in order to help those skilled in the art understand the technical content of the present disclosure, but do not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios. For example, in another embodiment, an exemplary system architecture to which the training method, image processing method and apparatus of a deep learning model can be applied may include a terminal device, but the terminal device may implement the training method, image processing method and apparatus of a deep learning model provided by the embodiments of the present disclosure without interacting with a server.
[0029] like Figure 1As shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104 and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0030] Users can use terminal devices 101, 102, 103 to interact with server 105 through network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients and / or social platform software, etc. (only as examples).
[0031] The terminal devices 101 , 102 , and 103 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, and desktop computers.
[0032] The server 105 may be a server that provides various services, such as a background management server that provides support for the content browsed by users using terminal devices 101, 102, and 103 (for example only). The background management server may analyze and process the received data such as user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services ("Virtual Private Server", or "VPS" for short). The server may also be a server of a distributed system, or a server combined with a blockchain.
[0033] It should be noted that the training method of the deep learning model provided in the embodiment of the present disclosure can generally be executed by the server 105. Accordingly, the training device of the deep learning model provided in the embodiment of the present disclosure can generally be set in the server 105. The training method of the deep learning model provided in the embodiment of the present disclosure can also be executed by a server or server cluster that is different from the server 105 and can communicate with the terminal devices 101, 102, 103 and / or the server 105. Accordingly, the training device of the deep learning model provided in the embodiment of the present disclosure can also be set in a server or server cluster that is different from the server 105 and can communicate with the terminal devices 101, 102, 103 and / or the server 105. The image processing method provided in the embodiment of the present disclosure can generally be executed by the terminal devices 101, 102, or 103. Accordingly, the image processing device provided in the embodiment of the present disclosure can also be set in the terminal devices 101, 102, or 103.
[0034] Alternatively, the training method of the deep learning model provided in the embodiment of the present disclosure may also be generally performed by the terminal device 101, 102, or 103. Accordingly, the training device of the deep learning model provided in the embodiment of the present disclosure may also be arranged in the terminal device 101, 102, or 103. The image processing method provided in the embodiment of the present disclosure may generally be performed by the server 105. Accordingly, the image processing device provided in the embodiment of the present disclosure may generally be arranged in the server 105. The image processing method provided in the embodiment of the present disclosure may also be performed by a server or server cluster that is different from the server 105 and can communicate with the terminal device 101, 102, 103 and / or the server 105. Accordingly, the image processing device provided in the embodiment of the present disclosure may also be arranged in a server or server cluster that is different from the server 105 and can communicate with the terminal device 101, 102, 103 and / or the server 105.
[0035] For example, when it is necessary to train the deep learning model, the terminal devices 101, 102, and 103 can obtain an initial sample image, and then send the obtained initial sample image to the server 105, and the server 105 inputs the initial sample image into the first model to obtain a first processed image, and uses the initial sample image and the first processed image to train the second model, and determine the deep learning model based on the first model and the trained second model. Alternatively, the initial sample image is processed by a server or server cluster that can communicate with the terminal devices 101, 102, 103 and / or the server 105, and the deep learning model is determined.
[0036] For example, when it is necessary to process an image, the terminal devices 101, 102, and 103 can obtain the image to be processed, and then send the obtained image to be processed to the server 105, and the server 105 inputs the image to be processed into the deep learning model to obtain a third processed image. The deep learning model is trained according to the training method of the present disclosure. Alternatively, a server or server cluster that can communicate with the terminal devices 101, 102, 103 and / or the server 105 analyzes the image to be processed and obtains the third processed image.
[0037] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to implementation requirements.
[0038] Figure 2 A flowchart of a deep learning model training method according to an embodiment of the present disclosure is schematically shown.
[0039] According to an embodiment of the present disclosure, a deep learning model includes a first model and a second model. The first model may include a network model that receives click-like interaction information as input and outputs a binary segmentation result, such as hrnet18+ocr64, deeplabv3+, etc. HRnet, i.e., High Resolution network, represents a high-resolution network. HRnet18+OCR64 represents a high-resolution network with a self-attention mechanism. deeplabv3+ represents a semantic segmentation network. The second model may include a model determined based on three hole convolution modules and one conv (convolution) module.
[0040] like Figure 2 As shown, the method includes operations S210 to S230.
[0041] In operation S210, an initial sample image is input into a first model to obtain a first processed image.
[0042] In operation S220, the second model is trained using the initial sample image and the first processed image.
[0043] In operation S230, a deep learning model is determined based on the first model and the trained second model.
[0044] According to an embodiment of the present disclosure, the first model may include a model that has been trained and converged using a sample image set consisting of images of the same modality. General scene images, medical images, and remote sensing building images may belong to images of different modalities, respectively. The sample image set may include any one of a general scene image set, a medical image image set, and a remote sensing building image set. The first model may include an interactive image processing model, such as an image segmentation model and other image processing models, and other image processing models may include, for example, at least one of an image classification model, an image recognition model, and the like.
[0045] According to an embodiment of the present disclosure, the initial sample image may include any one of a general scene image, a medical image, a remote sensing architectural image, and an image of other modalities. For example, the general scene image may include an image with people, animals, and other objects. The first processed image may include an image obtained after preliminary processing of the initial sample image. In the case where the first model is an image segmentation model, the initial sample image may include positive annotation points and negative annotation points. The positive annotation points may be used to annotate the pixel points of the area determined as the foreground in the initial sample image, and the negative annotation points may be used to annotate the pixel points of the area determined as the background in the initial sample image. In the case where the first model is other image processing models, the initial sample image may not include annotation point information.
[0046] According to an embodiment of the present disclosure, the second model may include a model with a continuous learning function. The second model may learn image features of various types of images during training. The trained second model may achieve optimized processing of the first processed image. The second model may be a simpler and lighter model than the first model. The second model may not depend on a specific network. For example, the second model may be trained independently of the deep learning model, and then, when the deep learning model needs to use the second model, the second model is introduced to achieve plug-and-play.
[0047] According to an embodiment of the present disclosure, the second model may be trained using only the initial sample image and the first processed image. Then, a deep learning model is determined based on the first model and the trained second model. The first model may also be trained using the initial sample image, and the second model may be trained using the initial sample image and the first processed image. Then, a deep learning model is determined based on the trained first model and the trained second model.
[0048] Through the above-mentioned embodiments of the present disclosure, for the deep learning model constructed based on the first model and the second model, the second model can be trained, so that the obtained deep learning model can not only learn new features and adapt to new tasks, but also reduce the forgetting of initial features, which can effectively improve the generalization ability of the model.
[0049] In conjunction with specific embodiments, Figure 2 The method shown is further explained.
[0050] According to an embodiment of the present disclosure, the initial sample image may include initial positive annotation points and initial negative annotation points. The initial positive annotation points may be generated in response to receiving an annotation operation for a first pixel in a foreground area. The initial negative annotation points may be generated in response to receiving an annotation operation for a second pixel in a background area.
[0051] According to an embodiment of the present disclosure, the first model may be an interactive segmentation model, the second model may be a model with a continuous learning function, and the deep learning model may be an interactive segmentation model with a continuous learning function. The interactive segmentation model with a continuous learning function can be applied to pixel-level image annotation tasks to achieve image segmentation such as autonomous driving lane line segmentation, medical image segmentation, object segmentation in general scenes, and remote sensing building segmentation.
[0052] According to an embodiment of the present disclosure, the area occupied by the target object to be segmented can be determined as the foreground area, and other areas in the sample image except the foreground area can be determined as the background area. The pixels included in the foreground area can be the first pixel points, and the pixels included in the background area can be the second pixel points. The initial positive annotation points and the initial negative annotation points can be obtained by predefined settings or by user annotation. The operations generated by the user annotation can be used as the above-mentioned annotation operations.
[0053] According to an embodiment of the present disclosure, in the process of determining a deep learning model based on a first model and a trained second model, when it is necessary to train the first model, the first model can be trained using an initial sample image and the positive and negative annotation point information annotated on the initial sample image. When it is necessary to train the second model, the initial sample image including the positive and negative annotation points can be input into the first model to obtain an initial segmentation result obtained by performing preliminary segmentation on the initial sample image. Then, the second model can be trained using the initial segmentation result and the initial sample image including the positive and negative annotation points.
[0054] Through the above-mentioned embodiments of the present disclosure, the deep learning model can be trained based on the positive and negative annotation point information of the received annotation operation, and an interactive deep learning model can be realized, which is conducive to training a model that meets the annotation task required in the current scene for different annotation scenes. In addition, the deep learning model can automatically identify the foreground area and the background area based on the information of the initial positive annotation points and the initial negative annotation points in the initial sample image. When performing task annotation based on the trained model, it can effectively reduce the manual annotation cost, simplify the annotation process, improve the annotation efficiency, and save the annotation cost.
[0055] According to an embodiment of the present disclosure, in the process of performing a training operation for a deep learning model using an initial sample image, the training method of the deep learning model may further include: inputting the initial sample image and the first processed image into a second model to obtain a second processed image. In response to determining that the second processed image satisfies a predetermined condition, updating the initial sample image to a target sample image.
[0056] According to an embodiment of the present disclosure, the second processed image may represent an image that is further optimized after the first processed image obtained by processing the initial sample image by the first model. The predetermined condition may include that the second processed image does not achieve the expected processing effect. In the case where it is determined that the second processed image obtained by processing the initial sample image using the deep learning model does not achieve the expected processing effect, the initial sample image may be updated.
[0057] According to an embodiment of the present disclosure, when it is determined that the second processed image obtained by processing the initial sample image using the deep learning model achieves the expected processing effect, the second processed image can be saved to save the processing result after processing the initial sample image. It should be noted that in this case, the initial sample image can also be updated to further optimize the deep learning model.
[0058] Figure 3 A flowchart of a method for training an interactive deep learning model according to an embodiment of the present disclosure is schematically shown.
[0059] like Figure 3 As shown, the method includes operations S310 to S360.
[0060] In operation S310, an initial sample image is input into a first model to obtain a first processed image.
[0061] In operation S320, the second model is trained using the initial sample image and the first processed image.
[0062] In operation S330, it is determined that the initial sample image and the first processed image are input into the second model to obtain a second processed image.
[0063] In operation S340, does the second processed image meet a predetermined condition? If yes, operations S350 to S360 are performed; if no, operation S360 is performed.
[0064] In operation S350, the second processed image is stored.
[0065] In operation S360, the initial sample image is updated to the target sample image, and operations S310 to S330 are performed using the target sample image as the initial sample image.
[0066] According to an embodiment of the present disclosure, when the first model needs to be trained, the first model can be first trained using the initial sample image. Then, the second model is trained using the first processed image obtained by processing the initial sample image with the first model and the initial sample image.
[0067] According to an embodiment of the present disclosure, a method for updating an initial sample image to a target sample image may include: replacing the initial sample image with another sample image of the same modality to obtain the target sample image. Adding at least one of a positive annotation point and a negative annotation point to the initial sample image to obtain the target sample image.
[0068] Through the above-mentioned embodiments of the present disclosure, an interactive deep learning model with continuous learning capability is realized. The annotation point information and annotation results annotated by the user under the current task can be used to effectively optimize the interactive deep learning model, so that the interactive deep learning model can better adapt to the current annotation task and improve the annotation efficiency and annotation accuracy.
[0069] According to an embodiment of the present disclosure, updating the initial sample image to the target sample image may include: in response to receiving a labeling operation for a third pixel in the foreground area, generating a target positive labeling point. The third pixel includes other pixels in the foreground area except the first pixel. In response to receiving a labeling operation for a fourth pixel in the background area, generating a target negative labeling point. The fourth pixel includes other pixels in the background area except the second pixel. The target sample image including the target positive labeling point and the target negative labeling point is determined as the target sample image.
[0070] According to an embodiment of the present disclosure, when it is determined that the second processed image achieves the expected effect, the user can add annotation information to the initial sample image, which may include adding at least one of positive annotation points and negative annotation points, to obtain an updated target sample image.
[0071] Through the above-mentioned embodiments of the present disclosure, an interactive intelligent labeling method is provided based on the interactive features of the interactive deep learning model. The interactive intelligent labeling can support the user to update the annotations in the initial sample image multiple times during the process of performing training operations on the deep learning model using the initial sample image, thereby updating the initial sample image and optimizing the deep learning model based on the interactive update process.
[0072] According to an embodiment of the present disclosure, the first model may be trained using a sample image set whose image modality is a target modality. Determining the deep learning model based on the first model and the trained second model may include: in response to detecting that the image modality represented by the initial sample image is consistent with the target modality, determining the first model and the trained second model as deep learning models.
[0073] According to an embodiment of the present disclosure, the target modality may include any one of a modality represented by a general scene image, a modality represented by a medical image, and a modality represented by a remote sensing architectural image. When it is determined that the image modality represented by the initial sample image is consistent with the image modality represented by the sample image set for training the first model, the method for training and optimizing the deep learning model may be a plug-in optimization method, which means that only the second model is adjusted, and the first model is not adjusted.
[0074] For example, the first model may be a segmentation model, the second model may be a continuous learning model, and the target modality represented by the sample image set used to train the segmentation model may be a modality represented by a general scene image. When the scene corresponding to the current annotation task is similar to the general scene corresponding to the sample image set based on which the first model is trained, the initial parameters of the segmentation model may be fixed, and only the continuous learning model may be optimized, so that only the continuous learning model is allowed to learn the current annotation behavior and the relevant features of the annotated image. Then, the segmentation model and the optimized continuous learning model are determined as the optimized deep learning model.
[0075] Through the above-mentioned embodiments of the present disclosure, when the annotation scene represented by the initial sample image is similar to the training scene represented by the sample image set for training the first model, based on the above-mentioned optimization method, it is possible to effectively alleviate the problem of large parameter modifications to the deep learning model that already has good generalization ability, causing the deep learning model to fall into a local minimum and causing the entire output result to degrade. The deep learning model can also be enabled to continuously adapt to different image distribution differences and continuously improve the annotation effect.
[0076] According to an embodiment of the present disclosure, in the case where the first model is a model trained using a sample image set with the image modality as the target modality, determining the deep learning model based on the first model and the trained second model may also include: in response to detecting that the image modality represented by the initial sample image is inconsistent with the target modality, fine-tuning the first model using the initial sample image. Determining the fine-tuned first model and the trained second model as deep learning models.
[0077] For example, the first model may be a segmentation model, the second model may be a continuous learning model, and the target modality represented by the sample image set used to train the segmentation model may be a modality represented by a general scene image. When the scene corresponding to the current annotation task is a scene such as a remote sensing image or a medical image, it can be determined that the scene corresponding to the current annotation task is quite different from the general scene corresponding to the sample image set based on which the first model was trained. In this case, the segmentation model and the continuous learning model can be optimized together. The optimization process may, for example, include: modifying the lightweight continuous learning model so that it mainly adapts to the differences brought about by the migration of the scene. Fine-tuning the parameters of the segmentation model to adapt it to the current task changes.
[0078] Through the above-mentioned embodiments of the present disclosure, when there is a significant difference between the annotated scene represented by the initial sample image and the training scene represented by the sample image set for training the first model, the model can continue to learn. By fine-tuning the first model and reducing the optimization of the parameters of the first model, the deep learning model can quickly adapt to the annotation accuracy of the current task and effectively reduce the deep learning model's forgetting of the class features that have been learned based on the first model. By mainly applying the optimization of the deep learning model to the second model and adjusting the parameters of the second model, the model's continuous learning ability can be effectively improved, so that the deep learning model can continuously adapt to the current task in scenarios with large domain migration, thereby improving the model's generalization ability.
[0079] According to an embodiment of the present disclosure, a first learning rate used for fine-tuning the first model is smaller than a second learning rate used for training the second model.
[0080] It should be noted that the learning rate is a hyperparameter that updates the network weights before the loss function gradient. The lower the learning rate, the slower the network weights are updated, the slower the loss function changes, and the longer it takes for the loss function to converge.
[0081] According to an embodiment of the present disclosure, the first learning rate may be less than one percent of the second learning rate.
[0082] Through the above-mentioned embodiments of the present disclosure, the model optimization process can be mainly applied to the second model, thereby reducing the problem of catastrophic forgetting of the class features learned by the first model and effectively improving the generalization ability of the model.
[0083] According to an embodiment of the present disclosure, a deep learning model that is continuously optimized and trained based on the above-mentioned training method can be used for image processing, for example.
[0084] Figure 4 The flowchart of the image processing method according to the embodiment of the present disclosure is schematically shown.
[0085] like Figure 4 The method includes operations S410 to S420.
[0086] In operation S410, an image to be processed is acquired.
[0087] In operation S420, the image to be processed is input into a deep learning model to obtain a third processed image.
[0088] According to an embodiment of the present disclosure, the image to be processed may include any one of an image to be segmented, an image to be classified, an image to be recognized, etc. Accordingly, the obtained third processed image may include any one of a segmented image, a classification result, and a recognition result.
[0089] Through the above-mentioned embodiments of the present disclosure, image processing is performed based on a deep learning model including a first model and a second model, which can adapt to images to be processed in various scenarios, improve the efficiency of image processing, and improve the accuracy of each image processing result.
[0090] According to an embodiment of the present disclosure, the image to be processed may include predetermined positive annotation points and predetermined negative annotation points, wherein the predetermined positive annotation points are generated in response to receiving an annotation operation for pixel points in a foreground area, and the predetermined negative annotation points are generated in response to receiving an annotation operation for pixel points in a background area.
[0091] Through the above-mentioned embodiments of the present disclosure, the deep learning model can be trained based on the positive and negative annotation point information of the received annotation operation, and an interactive deep learning model can be realized, which is conducive to training a model that meets the annotation task required in the current scene for different annotation scenes. In addition, the deep learning model can automatically identify the foreground area and the background area based on the information of the initial positive annotation points and the initial negative annotation points in the initial sample image. When performing task annotation based on the trained model, it can effectively reduce the work of the annotator, simplify the annotation process, improve the annotation efficiency, and save the annotation cost.
[0092] According to an embodiment of the present disclosure, the third processed image includes a segmented image.
[0093] Through the above-mentioned embodiments of the present disclosure, combined with the aforementioned deep learning model, an interactive segmentation model with continuous learning capability can be realized, which can be applied to images to be segmented in various modalities to improve segmentation efficiency and accuracy.
[0094] According to the embodiments of the present disclosure, the above-mentioned deep learning model can be obtained by offline training using a collected sample image set of various modalities, or by online training of the deep learning model using the image to be processed and related annotation information during the process of processing the image to be processed using the deep learning model.
[0095] Figure 5 A schematic diagram schematically illustrates the optimization training of an interactive segmentation model with continuous learning capability according to an embodiment of the present disclosure.
[0096] like Figure 5 As shown, the image to be segmented 511 may include predetermined positive annotation points and predetermined negative annotation points. The first model 520 may be a segmentation model, and the second model 540 may be an adaptation model. The segmentation model 520 may take the image to be segmented 511, and information such as positive annotation points and negative annotation points annotated by the user on the image to be segmented 511 as input, and may output a segmentation result 530. The continuous learning model 540 may take the image to be segmented 511, the positive annotation point information and negative annotation point information annotated by the user on the image to be segmented 511, and the segmentation result 530 as input, and further optimize the segmentation result 530 output by the segmentation model 520, for example, a more optimized segmentation result 550 may be obtained.
[0097] It should be noted that the positive annotation points marked on the image to be segmented 511 can be marked on any pixel point included in the area occupied by the animal body, and the negative annotation points marked on the image to be segmented 511 can be marked on any pixel point included in other areas in the image 511 except the animal body.
[0098] According to an embodiment of the present disclosure, when the image 511 is segmented using the interactive segmentation model constructed by the segmentation model 520 and the continuous learning model 540 to obtain the segmentation result 550, operations S560 to S580 may be performed.
[0099] In operation S560 , is it determined whether the user is satisfied? If yes, operations S570 to S580 are performed; if no, operation S580 is performed. In this operation, the user can judge the satisfaction with the segmentation result 550 output by the continuous learning model 540 .
[0100] In operation S570, the segmentation result is stored. In this operation, if the user is satisfied with the segmentation result 550, the segmentation result 550 may be stored.
[0101] In operation S580, the interactive segmentation model is trained. In this operation, whether the user is satisfied with the segmentation result 550 or not, the interactive segmentation model can be further optimized based on this operation.
[0102] According to an embodiment of the present disclosure, operation S580 may include: the user adds new positive and negative annotation points to the image 511 by clicking interactively. For example, the white points in the image 512 may represent the positive annotation points added to the image 511, and the white points in the image 513 may represent the negative annotation points added to the image 511. Based on this, an updated image different from the annotation result of the original image 511 may be obtained. Then, the updated image, the positive annotation points, the negative annotation points and other information marked in the updated image, and the segmentation result corresponding to the updated image output by the segmentation model 520 may be used to optimize the continuous learning model 540 for training. In this process, a more optimized segmentation result may also be obtained.
[0103] According to an embodiment of the present disclosure, during the process of performing operation S580, when it is determined that the image input to the segmentation model 520 is inconsistent with the image modality of the previously processed image, the image with the new modality and the information such as the positive annotation points and the negative annotation points annotated by the user on the image with the new modality can be used to fine-tune the segmentation model 520. Then, the image with the new modality, the information such as the positive annotation points and the negative annotation points annotated by the user on the image with the new modality, and the segmentation result corresponding to the image with the new modality output by the segmentation model 520 can be used to optimize the continuous learning model 540.
[0104] According to the embodiments of the present disclosure, for each segmentation result output by the continuous learning model 540, the annotation points on the image to be segmented corresponding to the segmentation result can be updated or modified, and then the interactive segmentation model can be further optimized and trained using the image after the annotation points are updated or modified. The training process can continue until the user is satisfied with the segmentation result output by the continuous learning model 540.
[0105] Through the above-mentioned embodiments of the present disclosure, a continuous learning module is added on the basis of the interactive segmentation model, and the continuous learning ability can be introduced into the interactive segmentation model to realize an interactive segmentation model with continuous learning ability, and annotate different modality images for users, providing an efficient continuous learning optimization method. By mainly optimizing the lightweight segmentation model to adapt to the difference between the current task and the training data, and reducing the adjustment of the network parameters in the basic segmentation model, the barrier that the interactive segmentation model can only be applied to scenarios similar to the training data set can be broken, so that the interactive segmentation model can continuously adapt to different segmentation tasks, and can effectively reduce the risk of catastrophic forgetting of the class features that have been learned by the basic segmentation model.
[0106] Figure 6 A block diagram of a deep learning model training device according to an embodiment of the present disclosure is schematically shown.
[0107] According to an embodiment of the present disclosure, the deep learning model includes a first model and a second model.
[0108] like Figure 6 As shown, the training device 600 for the deep learning model includes a first acquisition module 610, a training module 620 and a determination module 630.
[0109] The first obtaining module 610 is used to input the initial sample image into the first model to obtain a first processed image.
[0110] The training module 620 is used to train the second model using the initial sample image and the first processed image.
[0111] The determination module 630 is used to determine the deep learning model according to the first model and the trained second model.
[0112] According to an embodiment of the present disclosure, the initial sample image includes initial positive annotation points and initial negative annotation points. The initial positive annotation points are generated in response to receiving an annotation operation for a first pixel point in a foreground area. The initial negative annotation points are generated in response to receiving an annotation operation for a second pixel point in a background area.
[0113] According to an embodiment of the present disclosure, the training device of the deep learning model also includes a second acquisition module and an update module.
[0114] The second acquisition module is used to input the initial sample image and the first processed image into the second model to obtain the second processed image.
[0115] The updating module is used for updating the initial sample image to a target sample image in response to determining that the second processed image satisfies a predetermined condition.
[0116] According to an embodiment of the present disclosure, the updating module includes a first generating unit, a second generating unit and a first determining unit.
[0117] The first generating unit is configured to generate a target positively labeled point in response to receiving a labeling operation for a third pixel point in the foreground area. The third pixel point includes other pixel points in the foreground area except the first pixel point.
[0118] The second generating unit is configured to generate a target negative annotation point in response to receiving an annotation operation for a fourth pixel point in the background area. The fourth pixel point includes other pixel points in the background area except the second pixel point.
[0119] The first determining unit is used to determine a target sample image including target positive annotation points and target negative annotation points as a target sample image.
[0120] According to an embodiment of the present disclosure, the first model is trained using a sample image set whose image modality is the target modality. The determination module includes a second determination unit.
[0121] The second determination unit is used to determine the first model and the trained second model as deep learning models in response to detecting that the image modality represented by the initial sample image is consistent with the target modality.
[0122] According to an embodiment of the present disclosure, the first model is obtained by training a sample image set using an image modality as a target modality. The determination module includes a fine-tuning module and a third determination unit.
[0123] The fine-tuning module is used to fine-tune the first model using the initial sample image in response to detecting that the image modality represented by the initial sample image is inconsistent with the target modality.
[0124] The third determination unit is used to determine the fine-tuned first model and the trained second model as deep learning models.
[0125] According to an embodiment of the present disclosure, a first learning rate used for fine-tuning the first model is smaller than a second learning rate used for training the second model.
[0126] Figure 7 A block diagram of a deep learning model training device according to an embodiment of the present disclosure is schematically shown.
[0127] like Figure 7 As shown, the image processing apparatus 700 includes a third obtaining module 710 .
[0128] The third obtaining module 710 is used to input the image to be processed into the deep learning model to obtain a third processed image. The deep learning model is trained according to the training device of the deep learning model.
[0129] According to an embodiment of the present disclosure, the image to be processed includes predetermined positive annotation points and predetermined negative annotation points, the predetermined positive annotation points are generated in response to receiving an annotation operation for pixel points in the foreground area, and the predetermined negative annotation points are generated in response to receiving an annotation operation for pixel points in the background area.
[0130] According to an embodiment of the present disclosure, the third processed image includes a segmented image.
[0131] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.
[0132] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the deep learning model training method and image processing method of the present disclosure.
[0133] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute the deep learning model training method and image processing method of the present disclosure.
[0134] According to an embodiment of the present disclosure, a computer program product includes a computer program, and when the computer program is executed by a processor, it implements the training method of the deep learning model and the image processing method of the present disclosure.
[0135] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0136] like Figure 8As shown, the device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 to a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0137] A number of components in the device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the device 800 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0138] The computing unit 801 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 801 performs the various methods and processes described above, such as the training method of the deep learning model and the image processing method. For example, in some embodiments, the training method of the deep learning model and the image processing method may be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as a storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the training method of the deep learning model and the image processing method described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to execute a training method or an image processing method of a deep learning model in any other appropriate manner (for example, by means of firmware).
[0139] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0140] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0141] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0142] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0143] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0144] A computer system may include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises through computer programs running on respective computers and having a client-server relationship to each other. The server may be a cloud server, a server in a distributed system, or a server combined with a blockchain.
[0145] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.
[0146] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A training method for a deep learning model, the deep learning model comprising a first model and a second model, the method comprising: Inputting the initial sample image into the first model to obtain a first processed image; Using the initial sample image and the first processed image, training a second model; as well as Determine the deep learning model according to the first model and the trained second model; The first model is obtained by training a sample image set whose image modality is a target modality; and determining the deep learning model according to the first model and the trained second model includes: In response to detecting that the image modality represented by the initial sample image is inconsistent with the target modality, the first model is fine-tuned using the initial sample image; and the fine-tuned first model and the trained second model are determined as the deep learning models, wherein a first learning rate used for fine-tuning the first model is less than a second learning rate used for training the second model.
2. The method according to claim 1, wherein: The initial sample image includes initial positive annotation points and initial negative annotation points, wherein the initial positive annotation points are generated in response to receiving an annotation operation for a first pixel point in a foreground area, and the initial negative annotation points are generated in response to receiving an annotation operation for a second pixel point in a background area.
3. The method according to claim 2, further comprising: Inputting the initial sample image and the first processed image into the second model to obtain a second processed image; as well as In response to determining that the second processed image satisfies a predetermined condition, the initial sample image is updated to a target sample image.
4. The method according to claim 3, wherein: The updating of the initial sample image to the target sample image comprises: In response to receiving a labeling operation for a third pixel point in the foreground area, generating a target positive labeling point, the third pixel point including other pixel points in the foreground area except the first pixel point; In response to receiving a labeling operation for a fourth pixel point in the background area, generating a target negative labeling point, the fourth pixel point including other pixel points in the background area except the second pixel point; and A target sample image including the target positive annotation points and the target negative annotation points is determined as the target sample image.
5. The method according to any one of claims 1 to 4, wherein: Determining the deep learning model according to the first model and the trained second model further includes: In response to detecting that the image modality represented by the initial sample image is consistent with the target modality, the first model and the trained second model are determined as the deep learning models.
6. An image processing method, comprising: Inputting the image to be processed into the deep learning model to obtain a third processed image; Wherein, the deep learning model is trained according to any one of the training methods according to claims 1-5.
7. The method according to claim 6, wherein: The image to be processed includes predetermined positive annotation points and predetermined negative annotation points, wherein the predetermined positive annotation points are generated in response to receiving an annotation operation for pixel points in a foreground area, and the predetermined negative annotation points are generated in response to receiving an annotation operation for pixel points in a background area.
8. The method according to claim 6, wherein: The third processed image includes a segmented image.
9. A training device for a deep learning model, the deep learning model comprising a first model and a second model, the device comprising: A first acquisition module, used for inputting an initial sample image into a first model to obtain a first processed image; A training module, used for training a second model using the initial sample image and the first processed image; as well as A determination module, configured to determine the deep learning model according to the first model and the trained second model; Wherein, the first model is obtained by training using a sample image set whose image modality is the target modality; and the determination module comprises: a fine-tuning module, configured to, in response to detecting that the image modality represented by the initial sample image is inconsistent with the target modality, fine-tune the first model using the initial sample image; and A third determining unit, configured to determine the fine-tuned first model and the trained second model as the deep learning model; Wherein, a first learning rate used for fine-tuning the first model is smaller than a second learning rate used for training the second model.
10. The device according to claim 9, wherein: The initial sample image includes initial positive annotation points and initial negative annotation points, wherein the initial positive annotation points are generated in response to receiving an annotation operation for a first pixel point in a foreground area, and the initial negative annotation points are generated in response to receiving an annotation operation for a second pixel point in a background area.
11. The apparatus according to claim 10, further comprising: A second acquisition module, used for inputting the initial sample image and the first processed image into the second model to obtain a second processed image; as well as An updating module is configured to update the initial sample image to a target sample image in response to determining that the second processed image satisfies a predetermined condition.
12. The device according to claim 11, wherein The update module includes: A first generating unit, configured to generate a target positive annotation point in response to receiving an annotation operation for a third pixel point in the foreground area, wherein the third pixel point includes other pixel points in the foreground area except the first pixel point; a second generating unit, configured to generate a target negative annotation point in response to receiving an annotation operation for a fourth pixel point in the background area, wherein the fourth pixel point includes other pixel points in the background area except the second pixel point; and The first determining unit is configured to determine a target sample image including the target positive annotation points and the target negative annotation points as the target sample image.
13. The device according to any one of claims 9 to 12, wherein: The determination module also includes: The second determination unit is used to determine the first model and the trained second model as the deep learning models in response to detecting that the image modality represented by the initial sample image is consistent with the target modality.
14. An image processing device, comprising: A third acquisition module is used to input the image to be processed into the deep learning model to obtain a third processed image; Wherein, the deep learning model is trained according to the training device described in any one of claims 9-13.
15. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.
16. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-8.
17. A computer program product, comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 8.