Semantic segmentation method based on image noise reduction, image processing equipment and storage medium

By extracting image features using SwinIR model and combining SAM model for denoising and segmentation processing, the problem of low resource utilization in image processing is solved and efficient image processing is achieved.

CN119963846AActive Publication Date: 2025-05-09SHENZHEN SMARTCITY TECH DEV GRP CO LTD

Patent Information

Application Number
CN202510444969.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-05-09
Estimated Expiration
2045-04-10

AI Technical Summary

Technical Problem

The prior art requires model training during image processing, resulting in low resource utilization.

Method used

Image feature information is extracted through image recovery model SwinIR, and denoising and segmentation processing is performed based on SAM of the Segmentation Model, avoiding the need for model training.

Benefits of technology

It improves resource utilization and processing efficiency during image processing, and reduces resource waste caused by model training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963846A_ABST
    Figure CN119963846A_ABST
Patent Text Reader

Abstract

The invention discloses a semantic segmentation method based on image noise reduction, image processing equipment and a storage medium, relates to the technical field of image processing, and discloses the semantic segmentation method based on image noise reduction, which comprises the following steps: after a to-be-processed image is received, a shallow feature extraction module and a deep feature extraction module based on an image restoration model, extracting feature information of the to-be-processed image; based on the feature information, performing denoising processing on the to-be-processed image according to an image reconstruction module to obtain a denoised image; and according to the everything segmentation model and indication information of the noise-reduced image, performing segmentation processing on the noise-reduced image to obtain an image segmentation result, the indication information being used for identifying an entity instance object of the to-be-processed image. The to-be-processed image is subjected to denoising and segmentation processing based on the image restoration model and the everything segmentation model which do not need to be pre-trained, and the resource utilization rate during image processing is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a semantic segmentation method based on image denoising, an image processing device and a storage medium. Background Art

[0002] The basic principle of deep learning image denoising is to use deep learning models such as convolutional neural networks (CNN) to effectively remove various types of noise by learning local features of the image. Image segmentation is to divide the image into several sub-regions based on the similarity and mutual exclusivity of certain local features of the image (such as grayscale, texture, color or statistical features, etc.).

[0003] In order to improve the performance and generalization ability of the model, it is usually necessary to collect and annotate a large amount of image data for training. However, the long training process will take up a lot of computing resources and time, resulting in low resource utilization.

[0004] The above contents are only used to assist in understanding the technical solution of the present application and do not constitute an admission that the above contents are prior art. Summary of the invention

[0005] The main purpose of this application is to provide a semantic segmentation method, image processing device and storage medium based on image denoising, aiming to solve the technical problem of low resource utilization caused by the need for model training when processing images.

[0006] To achieve the above objectives, the present application proposes a semantic segmentation method based on image denoising, the method comprising: After receiving the image to be processed, extracting feature information of the image to be processed based on the shallow feature extraction module and the deep feature extraction module of the image restoration model SwinIR; After receiving the image to be processed, extracting feature information of the image to be processed based on the shallow feature extraction module and the deep feature extraction module of the image restoration model; Based on the feature information, performing denoising processing on the image to be processed according to the image reconstruction module of the image restoration model to obtain a denoised image; The denoised image is segmented according to the object segmentation model and the indication information of the denoised image to obtain an image segmentation result, wherein the indication information is used to identify the entity instance object of the image to be processed.

[0007] In one embodiment, the denoised image is segmented according to the object segmentation model and the indication information of the denoised image to obtain an image segmentation result, and before the step of using the indication information to identify the entity instance object of the image to be processed, the step further includes: Acquire prompt information received by the prompt editing interface of the object segmentation model, and determine a prompt type of the prompt information, wherein the prompt type includes at least voice, text, point selection, box selection, mask, gesture and template input; The indication information corresponding to the prompt information is generated according to the prompt type, and / or the indication information is updated according to the prompt type.

[0008] In one embodiment, there are multiple prompt types, and the step of generating the indication information corresponding to the prompt information according to the prompt type and / or updating the indication information according to the prompt type includes: Determining a set of indication information associated with a plurality of the prompt types; If the number of subsets of the indication information set does not match the number of the prompt types, the indication information with empty state information is generated according to the prompt type, and the existing indication information is updated according to the prompt type.

[0009] In one embodiment, the denoised image is segmented according to the object segmentation model and the indication information of the denoised image to obtain an image segmentation result, and before the step of using the indication information to identify the entity instance object of the image to be processed, the step further includes: Identify the entity instance object of the image to be processed based on a target detection algorithm, and obtain position information and / or shape information of the entity instance object; The indication information is generated according to the position information and / or the shape information and sent to the object segmentation model.

[0010] In one embodiment, the step of performing segmentation processing on the denoised image according to the object segmentation model and the indication information of the denoised image to obtain an image segmentation result, wherein the indication information is used to identify an entity instance object of the image to be processed comprises: Scaling the denoised image according to a preset image processing size of the object segmentation model, and discretizing the scaled image to obtain a processing vector; Inputting the processing vector into an image encoder to obtain image embedding information, and inputting the indication information into an indication information encoder to obtain indication information embedding information; Inputting the image embedding information and the indication information embedding information into a mask encoder to obtain a segmentation mask; The segmentation mask is used as the image segmentation result.

[0011] In one embodiment, after receiving the image to be processed, the step of extracting feature information of the image to be processed based on the shallow feature extraction module and the deep feature extraction module of the image restoration model includes: After receiving the image to be processed, extracting shallow features of the image to be processed based on the shallow feature extraction module, and inputting the shallow feature map into the deep feature extraction module, wherein the shallow features at least include the edge and contour of the image to be processed; Extracting features of the shallow feature map according to the deep feature extraction module to obtain deep features, wherein the deep features at least include texture and detail information of the image to be processed; Feature fusion processing is performed based on the shallow features and the deep features to obtain the fused feature information.

[0012] In one embodiment, the step of performing denoising on the image to be processed based on the feature information and using the image reconstruction module of the image restoration model to obtain the denoised image comprises: Inputting the fused feature information into the image reconstruction module of the image restoration model, so as to convert the fused feature map into an image space through the reconstruction module; The reconstructed image is denoised according to preset network parameters to obtain the denoised image.

[0013] In one embodiment, before the step of performing denoising on the image to be processed based on the feature information and using the image reconstruction module of the image restoration model to obtain the denoised image, the step further includes: A set of images to be processed is obtained, and the image restoration model is pre-trained based on the set of images to be processed, wherein the set of images to be processed is a denoised image.

[0014] In addition, to achieve the above-mentioned purpose, the present application also proposes an image processing device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the semantic segmentation method based on image denoising as described above.

[0015] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the semantic segmentation method based on image denoising as described above are implemented.

[0016] One or more technical solutions proposed in this application have at least the following technical effects: After receiving the image to be processed, the feature information of the image to be processed is directly extracted through the image restoration model, and based on the extracted feature information such as the edge and contour of the image to be processed, the image is denoised to obtain a denoised image, and then the denoised image is segmented through the object segmentation model and the indication information to obtain the segmentation result. The image to be processed is denoised and segmented based on the image restoration model and the object segmentation model that do not require pre-training, thereby improving the resource utilization during image processing and improving the image processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0018] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0019] Figure 1 A flowchart diagram of the first embodiment of the semantic segmentation method based on image denoising provided by the present application; Figure 2 A flowchart diagram of the second embodiment of the semantic segmentation method based on image denoising provided by the present application; Figure 3 A flowchart diagram of the third embodiment of the semantic segmentation method based on image denoising provided by the present application; Figure 4 A flowchart diagram of a fourth embodiment of a semantic segmentation method based on image denoising provided by the present application; Figure 5 Schematic diagram of the device structure of the hardware operating environment involved in the semantic segmentation method based on image denoising in the embodiment of the present application.

[0020] The purpose, features and advantages of this application will be further described in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0021] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.

[0022] The main solution of the embodiment of the present application is: after receiving the image to be processed, extracting feature information of the image to be processed based on the shallow feature extraction module and the deep feature extraction module of the image restoration model; Based on the feature information, performing denoising processing on the image to be processed according to the image reconstruction module of the image restoration model to obtain a denoised image; The denoised image is segmented according to the object segmentation model and the indication information of the denoised image to obtain an image segmentation result, wherein the indication information is used to identify the entity instance object of the image to be processed.

[0023] In the existing technology, in order to improve the performance and generalization ability of the model, it is usually necessary to collect and annotate a large amount of image data for training. The long training process will take up a lot of computing resources and time, resulting in low resource utilization.

[0024] The present application provides a solution, which performs denoising and segmentation processing on the image to be processed respectively through an image restoration model and an object segmentation model, and there is no need to train the two models in this process, thereby improving the resource utilization rate and the overall efficiency of image processing during the image processing process.

[0025] It should be noted that the execution subject of this embodiment may be a computing service device with data processing, network communication and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device capable of realizing the above functions, an image processing device, etc. The following takes an image processing device as an example to illustrate this embodiment and the following embodiments.

[0026] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.

[0027] The present application embodiment provides a semantic segmentation method based on image denoising, wherein semantic segmentation refers to pixel-level recognition of images, that is, marking the object category to which each pixel in the image belongs, including marking the grass, trees, buildings, and object shapes that appear in the image. Figure 1 , Figure 1 This is a flowchart of the first embodiment of the semantic segmentation method based on image denoising of the present application.

[0028] In this embodiment, the semantic segmentation method based on image denoising includes steps S10 to S30: Step S10, after receiving the image to be processed, extracting feature information of the image to be processed based on the shallow feature extraction module and the deep feature extraction module of the image restoration model.

[0029] It should be noted that the image restoration model is SwinIR, which is an image restoration model based on Swin Transformer, namely Swin Image Restoration model, which is specially used for tasks such as image denoising, super-resolution and deblurring. It extracts features at different scales by introducing hierarchical structure, self-attention mechanism and shift window operation, thereby effectively improving the quality of image restoration. Compared with traditional convolutional neural networks, SwinIR performs well in processing details and complex textures, and is more suitable for high-quality image denoising tasks. Among them, SwinIR consists of three modules: shallow feature extraction, deep feature extraction and high-quality image reconstruction module. The shallow feature extraction module and the deep feature extraction module process the image to be processed respectively, thereby extracting the feature information of the image. The features extracted by the two feature extraction modules are processed by the image reconstruction model, so as to effectively denoise the image.

[0030] Feature information refers to the key attributes or elements in an image that can describe or distinguish the image content. Feature information includes at least the edges and contours, texture, color and brightness, frequency information, and spatial relationships of the image. By obtaining the feature information of the image, SwinIR can more effectively distinguish which parts are noise and which parts are the real content of the image. This helps to retain important details and edge information of the image during the denoising process while removing unnecessary interference.

[0031] Therefore, in this embodiment, when performing noise reduction processing on the image to be processed, the shallow feature extraction module and the deep feature extraction module of SwinIR are used to extract the current feature information of the image, so as to effectively denoise the image to be processed through the feature information.

[0032] Specifically, the shallow features of the image to be processed, including simple information such as the edge and contour of the image to be processed, can be extracted by the shallow feature extraction module at the same time, and the deep feature extraction module can extract the deep features of the image to be processed, and the deep features include the texture and detail information of the image to be processed. In addition, feature extraction can be performed based on the shallow feature extraction module first, and then the feature image output by the shallow feature extraction module can be further extracted by the deep feature extraction module to obtain the deep feature extraction module.

[0033] When processing images, SwinIR can be used directly without training, which can effectively improve resource utilization during image processing. The two feature extraction modules of SwinIR are used to extract feature information of the image to be processed, thus improving image processing efficiency.

[0034] Step S20: Based on the feature information, denoising is performed on the image to be processed according to the image reconstruction module of the image restoration model to obtain a denoised image.

[0035] It is understandable that existing image segmentation solutions based on deep learning often do not pay attention to the interference of image noise, resulting in inaccurate segmentation results. Therefore, in this embodiment, after obtaining the feature information, the image reconstruction module based on SwinIR processes the two features to obtain a denoised image. Among them, the features extracted from the shallow features and the deep features can be fused, and then denoising can be performed based on the fused features.

[0036] As an optional implementation method of performing denoising on the image to be processed according to the image reconstruction module, the fused feature information can be input into the image reconstruction module of SwinIR, so that the fused feature map is converted into the image space through the reconstruction module, and then the reconstructed image is denoised according to the preset network parameters to obtain the denoised image. It can be understood that the image reconstruction module uses a convolutional layer or other types of layers to convert the feature map back to the image space, and then adjusts the network parameters through the optimization algorithm and loss function (such as mean square error MSE, peak signal-to-noise ratio PSNR, etc.) to make the reconstructed image as close as possible to a clear and noise-free image, that is, to obtain a denoised image.

[0037] Based on this, before the image is segmented, a clear and noise-free image is provided to meet the actual image segmentation needs.

[0038] Optionally, in addition to denoising the image to be processed through the SwinIR model, image denoising can also be achieved using algorithms such as DPIR (DeepPlug-and-Play Image Restoration) and Unet (U-type network architecture).

[0039] Step S30, performing segmentation processing on the denoised image according to the object segmentation model and the indication information of the denoised image to obtain an image segmentation result, wherein the indication information is used to identify the entity instance object of the image to be processed.

[0040] It should be noted that the Everything Segmentation Model is SAM, which is a general object segmentation model developed by Meta AI. It is designed to handle the segmentation task of any object in the image without special training or fine-tuning. Among them, SAM uses a Transformer-based architecture, combined with large-scale visual data and pre-training technology, which can efficiently perform object segmentation in different scenarios. SAM includes Image-encoder, prompt-encoder, and mask-decoder. The image encoder converts the image into a vector, and the prompt encoder processes the input prompt information, that is, processes the prompt, and generates a segmentation mask through the mask decoder to obtain the image segmentation result.

[0041] In terms of application scenarios, SAM has broad application potential in computer vision, medical imaging, autonomous driving, and content creation. In the field of computer vision, SAM can be used for tasks such as image processing, object recognition, and target tracking. In medical imaging, it can help segment organs or lesion areas, assisting doctors in diagnosis and treatment planning. In autonomous driving systems, SAM can be used to identify and segment obstacles such as pedestrians and vehicles on the road, providing decision support for autonomous driving systems. In addition, in picture and video editing, SAM can also facilitate further editing and processing by quickly segmenting objects.

[0042] Indication information is also called prompt information, which will be represented by prompt later. It is used to identify the entity instance object of the image to be processed. After obtaining the denoised image, the image needs to be segmented to extract the key elements in the image for subsequent image analysis, understanding and processing. SAM supports a variety of prompt methods, including point prompts, box prompts and text prompts. Users can gradually optimize the segmentation results in an interactive way. For example, users can gradually correct the segmentation results output by the model by adding new prompts to obtain more accurate segmentation. The architecture of SAM consists of an encoder, a prompt encoder and a segmentation head. The encoder is used to extract the visual features of the image, the prompt encoder processes the prompt information input by the user, and the segmentation head generates the final segmentation results based on the visual features and prompts.

[0043] In this embodiment, during the image segmentation process, in order to reduce the need to train the image segmentation model and occupy excessive computing resources, the denoised image is directly segmented based on the SAM, a segmentation model that does not require training. There is no need to collect a large amount of training data and spend a lot of resources to train the model, thereby improving the resource utilization of image processing.

[0044] Specifically, when the denoised image is segmented based on SAM, the operation information set by the user can be received through the interactive interface, and the operation information can be converted into an indication prompt that can be recognized by SAM. Then SAM processes the denoised image according to the indication information to identify and segment multiple different entity instance objects, including the whole, part and sub-part of the instance object.

[0045] Furthermore, in addition to being obtained through the interactive interface, the prompt can also be obtained through other neural network models. That is, after obtaining the denoised image, it is first analyzed through the neural network model to obtain a prompt that can be an entity instance object of the image to be processed, and then the prompt and the denoised image are input into SAM for processing.

[0046] The untrained SwinIR and SAM models are used to perform denoising and image segmentation on the processed images respectively, which improves the accuracy of image segmentation while reducing the resource waste caused by model training and improving the utilization of computing resources.

[0047] This embodiment provides a semantic segmentation method based on image denoising. During the image segmentation process, feature information of the image to be processed is extracted based on a SwinIR model that does not require training, and denoising is performed on the image based on the extracted feature information, thereby improving the image segmentation accuracy without the need to train the denoising model. Then, the denoised image is segmented using a SAM model that does not require pre-training, thereby avoiding the need to train the image segmentation model, thereby improving resource utilization during image processing.

[0048] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the first embodiment can be referred to the above description, and will not be repeated in the following. Figure 2 Before step S30, the semantic segmentation method based on image denoising further includes steps S40 to S50: Step S40, obtaining prompt information received by the prompt editing interface of the object segmentation model, and determining a prompt type of the prompt information.

[0049] In this embodiment, the R&D personnel can develop a prompt editing interface based on SAM so that the user can give corresponding prompts in the interface, that is, the user can input the prompt information when performing image segmentation through the prompt editing interface of SAM, so as to generate a prompt that SAM can recognize, thereby improving the image segmentation effect. Among them, the prompt types include at least voice, text, point selection, box selection, mask, gesture and template input.

[0050] In this embodiment, the user can interact with the model during the segmentation process and optimize the segmentation result by adjusting the prompt. If the initial segmentation result is inaccurate, the user can correct the prompt by adding additional points, boxes or masks to obtain a more accurate segmentation result.

[0051] Step S50: generating the indication information corresponding to the prompt information according to the prompt type, and / or updating the indication information according to the prompt type.

[0052] In this embodiment, in the process of generating a prompt corresponding to the prompt information based on the prompt type, the content of the voice input can be converted into text by the voice recognition system, and then the text is processed as text input, or a prompt matching the voice content is directly generated according to the context. The text input may directly contain segmentation instructions (such as "segment out the cat in the image"), or a descriptive text (such as "the flowers in the image"). In this case, the processing module can parse the text and generate the corresponding prompt through NLP technology. Inputs such as point selection and box selection directly indicate a specific area or object in the image. The processing module can generate a prompt representing the area or object based on these inputs, and then pass it to the SAM model. The mask input is a binary image of the same size as the image, in which the white area represents the object of interest and the black area represents the background. The processing module can directly use this mask as part of the prompt, or generate a more concise prompt based on the mask. Gesture input is usually performed on a touch screen device. The processing module needs to be able to recognize and parse these gestures, and then generate a prompt matching the gestures. The template input is a preset segmentation template. The processing module can generate a corresponding prompt based on the template selected by the user, and then pass it to the SAM model for segmentation.

[0053] In the process of updating the prompt according to the prompt information, for example, if there are currently prompts corresponding to the click and box selection content, when the user selects new click and box selection content and updates these prompts accordingly, the existing prompts are updated based on the latest prompt.

[0054] It can be understood that when the prompt corresponding to the prompt type is empty, a prompt can be generated based on the prompt type. If it is not empty, the prompt is updated based on the prompt type. Therefore, when there are multiple prompt types, the prompt set associated with multiple prompt types can be determined. If the number of subsets of the prompt set does not match the number of prompt types, it means that some prompt types are associated with prompts, while others are not. For example, among the multiple prompt types, including box selection, text input, and gestures, the text input type corresponds to the prompt when the image is previously segmented. Therefore, the prompt is updated through the text input type, and new prompts corresponding to these types are generated based on gestures and box selection. Therefore, if the number of subsets of the prompt set does not match the number of prompt types, a prompt with empty status information can be generated according to the prompt type, and an existing prompt can be updated according to the prompt type.

[0055] Optionally, in another optional implementation of obtaining prompt, before step S30, steps S60 to S70 are also included: Step S60: identifying the entity instance object of the image to be processed based on a target detection algorithm, and acquiring position information and / or shape information of the entity instance object.

[0056] Step S70: Generate the indication information according to the position information and / or the shape information and send the indication information to the object segmentation model.

[0057] In this embodiment, in addition to receiving prompts through the prompt editing interface set by SAM, other computer vision technologies such as target detection algorithms can also be used to automatically identify entity instance objects in images, and convert the location, shape and other information of these objects into prompts and input them into the SAM model.

[0058] For example, the target detection model can be used to detect all cars in the image, and then the location information of these cars is input as prompt to the SAM model for segmentation.

[0059] Combined with the target detection algorithm, the image to be processed is automatically identified to obtain the prompt for segmentation processing, which reduces the manual operation process, improves the degree of automation of the image to be processed, and thus improves the segmentation efficiency of the image to be processed.

[0060] This embodiment provides a semantic segmentation method based on image denoising. Before SAM segments the denoised image, it generates a corresponding prompt or updates an existing prompt based on the prompt information received through the prompt editing interface of SAM, or recognizes the image to be processed through a target detection algorithm, and generates a prompt based on the position information and shape information of the recognized entity instance, so that SAM can accurately segment the denoised image through the prompt to improve the image segmentation effect.

[0061] Based on the first embodiment of the present application, in the third embodiment of the present application, the same or similar contents as those in the first embodiment can refer to the above introduction, and will not be repeated later. Figure 3 , step S30 also includes steps S31 to S34: Step S31, scaling the denoised image according to a preset image processing size of the object segmentation model, and discretizing the scaled image to obtain a processing vector.

[0062] In this embodiment, after obtaining the denoised image, it is necessary to scale the denoised image to the input size required by the SAM model (e.g., 1024x1024), and then use a convolution operation to discretize the image into a series of vectors, which will serve as input information for the Image-encoder. Therefore, it is necessary to scale the denoised image based on the preset image processing size of SAM, and then discretize the scaled image to obtain a processing vector.

[0063] The image segmentation efficiency is improved by scaling and discretizing the denoised image.

[0064] Step S32: input the processing vector into an image encoder to obtain image embedding information, and input the indication information into an indication information encoder to obtain indication information embedding information.

[0065] In this embodiment, after the processing vector is obtained, it is input into the image encoder to obtain the image embedding representation (image-embedding), that is, the image embedding information. At the same time, the prepared prompt information is input into the prompt-encoder indication information encoder to obtain the prompt information embedding representation (prompt-embedding), that is, prompt embedding information. By obtaining the embedded representation information, the accuracy of the segmentation process is improved.

[0066] Step S33: input the image embedding information and the indication information embedding information into a mask encoder to obtain a segmentation mask.

[0067] Step S34: using the segmentation mask as the graphic segmentation result.

[0068] In this embodiment, after the embedding information is obtained, the image-embedding and prompt-embedding are simultaneously input into the mask-decoder, a segmentation mask is generated through the decoding process, and the generated segmentation mask is optimized, such as removing small areas and smoothing edges, to improve the accuracy and consistency of the segmentation result. Finally, the optimized segmentation mask is output as the final image segmentation result.

[0069] Based on this, the pre-trained SAM model is used to segment the denoised image based on the received instruction prompt, so as to improve the resource utilization during image processing and the image segmentation effect.

[0070] Based on the first embodiment of the present application, in the fourth embodiment of the present application, the same or similar contents as those in the first embodiment can refer to the above description and will not be described in detail later. Figure 4 , step S10 also includes steps S11 to S13: Step S11, after receiving the image to be processed, extracting shallow features of the image to be processed based on the shallow feature extraction module, and inputting the shallow feature map into the deep feature extraction module.

[0071] In this embodiment, the shallow feature extraction module is mainly responsible for extracting the initial, relatively shallow feature information from the input low-quality image. A convolution layer is usually used for feature extraction. The function of this convolution layer is to map the input image from the image space to the high-dimensional feature space, so as to extract the basic features of the image. Among them, in the process of extracting the feature information of the image to be processed, the shallow features of the image to be processed can be first extracted by the shallow feature extraction module, and the shallow features at least include the edges and contours of the image to be processed, and then the shallow feature map after the information extraction is completed is input into the deep feature extraction module. Through the operation of the convolution layer, the shallow feature extraction module can extract low-frequency information in the image, such as basic features such as edges and contours, and these shallow features will then be passed to the deep feature extraction module for further feature extraction and processing.

[0072] Step S12: extracting features of the shallow feature map according to the deep feature extraction module to obtain deep features.

[0073] In this embodiment, the deep feature extraction module is the core part of the SwinIR model, which is responsible for extracting deep features in the image. The deep feature module is mainly composed of multiple residual Swin Transformer blocks (RSTBs), each of which contains multiple Swin Transformer layers and a residual connection. The Swin Transformer layer has a local attention mechanism that can capture local features in the image and pass the feature information to the next RSTB through the residual connection.

[0074] In the deep feature extraction process, the input image is processed by multiple RSTBs to gradually extract deeper feature information. These deep features include more complex features such as texture and details in the image. The extracted deep features will then be passed to a high-quality image reconstruction module for image reconstruction and restoration.

[0075] It should be noted that the deep feature extraction module does not directly process the original image. Its input is the feature map (or feature tensor) output by the shallow feature extraction module. These feature maps already contain some basic features of the image, such as edges, textures, etc. The task of the deep feature extraction module is to further extract deeper features based on these.

[0076] Step S13, performing feature fusion processing based on the shallow features and the deep features to obtain the fused feature information.

[0077] After obtaining the shallow features and deep features, in order to enable SwinIR to reconstruct and repair the image more accurately, it is usually necessary to fuse the shallow features and the deep features to improve the subsequent image denoising effect.

[0078] This embodiment provides a semantic segmentation method based on image denoising, which extracts low-frequency information and deep-level features in the image through a shallow feature extraction module and a deep feature extraction module, providing an important basis for subsequent denoising processing such as image reconstruction and restoration, and further improving the denoising effect of the image.

[0079] Based on the first embodiment of the present application, in the fifth embodiment of the present application, the same or similar contents as those in the first embodiment can be referred to the above description, and will not be described in detail later. On this basis, the SwinIR model can also be trained by the processed images to further improve the image denoising effect of SwinIR. Therefore, before step S10, a set of images to be processed can also be obtained, and the SwinIR model can be pre-trained based on the set of images to be processed, wherein the set of images to be processed is the denoised images.

[0080] In this embodiment, each image denoising is recorded after it is completed. Therefore, before denoising a new image to be processed, the model optimization process can be performed on the existing denoised image to improve the denoising effect.

[0081] It is understandable that the SwinIR model does not require a training set, and the model has acquired super strong generalization ability during the pre-training process. The SwinIR image denoising model can be summarized as the following function:

[0082] Among them, f represents the neural network, θ represents the network parameters (obtained by random initialization at the beginning), and z represents a fixed random noise initially input into the network. represents an image with noise, represents the output of the neural network, is the optimal solution of parameters obtained through training. The optimal output of the neural network is: .

[0083] During the pre-training process, you can use the images processed by SwinIR denoising to create a training data set for the image segmentation model, and then use the data annotation tool to annotate the images and sort out the original images, category labels and other information. In addition, you can also perform data enhancement on the images, including changing the brightness, adding random points, translation, flipping, etc. Optionally, you can also divide the data set. Divide it into 75% training set, 15% validation set, and 15% test set.

[0084] The present application provides an image processing device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the semantic segmentation method based on image denoising in the above-mentioned first embodiment.

[0085] Reference below Figure 5 , which shows a schematic diagram of the structure of an image processing device suitable for implementing an embodiment of the present application. The image processing device in the embodiment of the present application may include but is not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (PADs), portable multimedia players (PMPs), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 5The image processing device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0086] like Figure 5 As shown, the image processing device may include a processing device 1001 (such as a central processing unit, a graphics processor, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 to a random access memory (RAM) 1004. In the random access memory 1004, various programs and data required for the operation of the image processing device are also stored. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems can be connected to the input / output interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the image processing device to communicate with other devices wirelessly or by wire to exchange data. Although the image processing device with various systems is shown in the figure, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be implemented or have alternatively.

[0087] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.

[0088] The image processing device provided by the present application adopts the semantic segmentation method based on image denoising in the above embodiment, which can solve the technical problem of low resource utilization caused by the need for model training when processing images. Compared with the prior art, the beneficial effects of the image processing device provided by the present application are the same as the beneficial effects of the semantic segmentation method based on image denoising provided in the above embodiment, and the other technical features in the image processing device are the same as the features disclosed in the method of the previous embodiment, which will not be repeated here.

[0089] It should be understood that the various parts disclosed in this application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0090] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

[0091] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, computer programs) stored thereon, and the computer-readable program instructions are used to execute the semantic segmentation method based on image denoising in the above-mentioned embodiment.

[0092] The computer-readable storage medium provided in the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared or semiconductor systems, systems or devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM: Read Only Memory), an erasable programmable read-only memory (EPROM: Erasable Programmable Read Only Memory or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM, CD-Read Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, radio frequencies (RF, Radio Frequency), etc., or any suitable combination of the above.

[0093] The computer-readable storage medium may be included in the image processing device; or may exist independently without being assembled into the image processing device.

[0094] The computer-readable storage medium carries one or more programs. When the one or more programs are executed by the image processing device, the image processing device: after receiving the image to be processed, extracts feature information of the image to be processed based on the shallow feature extraction module and the deep feature extraction module of the image restoration model SwinIR; Based on the feature information, the image to be processed is subjected to denoising processing according to the image reconstruction module of SwinIR to obtain a denoised image; According to the everything segmentation model SAM and the indication information prompt of the denoised image, the denoised image is segmented to obtain an image segmentation result, and the indication information is used to identify the entity instance object of the image to be processed.

[0095] Computer program code for performing the operations of the present application may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0096] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of a code, and the module, a program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0097] The modules involved in the embodiments described in this application may be implemented by software or hardware, wherein the name of the module does not constitute a limitation on the unit itself in some cases.

[0098] The readable storage medium provided by the present application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned semantic segmentation method based on image denoising, and can solve the technical problem of low resource utilization caused by the need for model training when processing images. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by the present application are the same as the beneficial effects of the semantic segmentation method based on image denoising provided by the above-mentioned embodiment, and will not be elaborated here.

[0099] The above descriptions are only some embodiments of the present application, and are not intended to limit the patent scope of the present application. All equivalent structural changes made using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect applications in other related technical fields are included in the patent protection scope of the present application.

Claims

1. A semantic segmentation method based on image denoising, characterized in that: The semantic segmentation method based on image denoising includes: After receiving the image to be processed, extracting feature information of the image to be processed based on the shallow feature extraction module and the deep feature extraction module of the image restoration model; Based on the feature information, performing denoising processing on the image to be processed according to the image reconstruction module of the image restoration model to obtain a denoised image; The denoised image is segmented according to the object segmentation model and the indication information of the denoised image to obtain an image segmentation result, wherein the indication information is used to identify the entity instance object of the image to be processed.

2. The semantic segmentation method based on image denoising according to claim 1, characterized in that: Before the step of performing segmentation processing on the denoised image according to the object segmentation model and the indication information of the denoised image to obtain an image segmentation result, wherein the indication information is used to identify the entity instance object of the image to be processed, the step further includes: Acquire prompt information received by the prompt editing interface of the object segmentation model, and determine a prompt type of the prompt information, wherein the prompt type includes at least voice, text, point selection, box selection, mask, gesture and template input; The indication information corresponding to the prompt information is generated according to the prompt type, and / or the indication information is updated according to the prompt type.

3. The semantic segmentation method based on image denoising as claimed in claim 2, characterized in that: There are multiple prompt types, and the step of generating the indication information corresponding to the prompt information according to the prompt type, and / or updating the indication information according to the prompt type includes: Determining a set of indication information associated with a plurality of the prompt types; If the number of subsets of the indication information set does not match the number of the prompt types, the indication information with empty state information is generated according to the prompt type, and the existing indication information is updated according to the prompt type.

4. The semantic segmentation method based on image denoising according to claim 1, characterized in that: Before the step of performing segmentation processing on the denoised image according to the object segmentation model and the indication information of the denoised image to obtain an image segmentation result, wherein the indication information is used to identify the entity instance object of the image to be processed, the step further includes: Identify the entity instance object of the image to be processed based on a target detection algorithm, and obtain position information and / or shape information of the entity instance object; The indication information is generated according to the position information and / or the shape information and sent to the object segmentation model.

5. The semantic segmentation method based on image denoising according to claim 1, characterized in that: The step of performing segmentation processing on the denoised image according to the object segmentation model and the indication information of the denoised image to obtain an image segmentation result, wherein the indication information is used to identify the entity instance object of the image to be processed comprises: Scaling the denoised image according to a preset image processing size of the object segmentation model, and discretizing the scaled image to obtain a processing vector; Inputting the processing vector into an image encoder to obtain image embedding information, and inputting the indication information into an indication information encoder to obtain indication information embedding information; Inputting the image embedding information and the indication information embedding information into a mask encoder to obtain a segmentation mask; The segmentation mask is used as the image segmentation result.

6. The semantic segmentation method based on image denoising according to claim 1, characterized in that: After receiving the image to be processed, the step of extracting feature information of the image to be processed based on the shallow feature extraction module and the deep feature extraction module of the image restoration model includes: After receiving the image to be processed, extracting shallow features of the image to be processed based on the shallow feature extraction module, and inputting the shallow feature map into the deep feature extraction module, wherein the shallow features at least include the edge and contour of the image to be processed; Extracting features of the shallow feature map according to the deep feature extraction module to obtain deep features, wherein the deep features at least include texture and detail information of the image to be processed; Feature fusion processing is performed based on the shallow features and the deep features to obtain the fused feature information.

7. The semantic segmentation method based on image denoising according to claim 6, characterized in that: The step of performing denoising on the image to be processed based on the feature information according to the image reconstruction module of the image restoration model to obtain the denoised image comprises: Inputting the fused feature information into the image reconstruction module of the image restoration model, so as to convert the fused feature map into an image space through the reconstruction module; The reconstructed image is denoised according to preset network parameters to obtain the denoised image.

8. The semantic segmentation method based on image denoising according to claim 1, characterized in that: Before the step of performing denoising on the image to be processed based on the feature information according to the image reconstruction module of the image restoration model to obtain the denoised image, the step further includes: A set of images to be processed is obtained, and the image restoration model is pre-trained based on the set of images to be processed, wherein the set of images to be processed is a denoised image.

9. An image processing device, characterized in that: The image processing device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the semantic segmentation method based on image denoising as described in any one of claims 1 to 8.

10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the semantic segmentation method based on image denoising are implemented as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Detection method of coal rock infrared thermography damaged area

    CN117078931A

  • Image segmentation model training method, handheld object identification method, equipment and medium

    CN117893756A

  • Damaged luggage case damage assessment method and system based on SAM segmentation large model

    CN118397277A

  • Model construction method and apparatus, image segmentation method and apparatus, and device and medium

    WO2024131406A1

Cited By

  • Method for performing object detection by using prompt-based object detector and computing device using the same

    KR102987277B1