Image processing method and apparatus, and device, medium and program product

By performing semantic analysis on the first image and performing image denoising processing based on semantic analysis data, the problem of denoising errors in the prior art is solved, and the recognition accuracy and efficiency of image objects are improved.

WO2025112906A1PCT designated stage expired Publication Date: 2025-06-05TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/123313
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-29
Filing Date
2024-10-08
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

The prior art is prone to denoising errors when denoising product images, resulting in inconsistent reconstructed images and product images, which reduces the recognition accuracy of image objects.

Method used

By performing semantic analysis on the first image, semantic analysis data is obtained, and then image denoising the noise-added image is performed based on these data to obtain a second image, and finally, the defect recognition result is determined by comparing the feature similarity between the first image and the second image.

Benefits of technology

It improves the recognition accuracy and recognition efficiency of image objects, ensures the consistency between the denoised image objects and the original image objects, and avoids recognition errors caused by denoising errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024123313_05062025_PF_FP_ABST
    Figure CN2024123313_05062025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates the field of machine learning. Disclosed are an image processing method and apparatus, and a device, a medium and a program product. The method comprises: acquiring a noise-added image corresponding to a first image; performing semantic analysis on the first image, so as to obtain semantic analysis data; performing image denoising processing on the noise-added image on the basis of the semantic analysis data, so as to obtain a second image; and on the basis of the feature similarity between the first image and the second image, determining a defect identification result corresponding to the first image. That is, the semantic analysis data is guided to perform image denoising, so that an image object in the second image after denoising being consistent with an image object in the first image is ensured, and then the defect identification result is determined by means of the feature similarity that is obtained by means of a comparison, so that the accuracy and identification efficiency of image object identification can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Image processing method, device, equipment, medium and program product

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on November 29, 2023, with application number 2023116316198 and application name “Image processing method, device, equipment, medium and program product”, all contents of which are incorporated by reference into this application. Technical Field

[0002] The embodiments of the present application relate to the field of machine learning, and in particular to image processing. Background Art

[0003] Defect detection is commonly used in industrial production processes to identify manufacturing defects in products by inspecting product images.

[0004] In related technologies, defect detection methods usually involve pre-training a denoising model using normal sample images, adding noise to the product image to be identified, and then inputting the denoising model for denoising to obtain a denoised reconstructed image. By comparing the differences between the reconstructed image and the product image, the image object with defects in the product image can be identified.

[0005] However, in the related art, denoising errors may occur in the process of denoising the product image after noise processing through the denoising model, resulting in inconsistency between the image objects in the reconstructed image and the product image, making it impossible to determine the defective image object by comparing the reconstructed image and the product image, thereby reducing the recognition accuracy of the image object.

[0006] Summary of the Invention

[0007] The present invention provides an image processing method, apparatus, device, medium, and program product, which can improve the recognition accuracy of image objects. The technical solution is as follows:

[0008] In one aspect, an image processing method is provided, the method comprising:

[0009] Acquire a noisy image corresponding to a first image, where the noisy image is obtained by adding first noise data to the first image, and the first image includes an image object;

[0010] Performing semantic analysis on the first image to obtain semantic analysis data, where the semantic analysis data is used to represent information of the image object in the first image;

[0011] performing image denoising processing on the noisy image based on the semantic analysis data to obtain a second image;

[0012] Based on the feature similarity between the first image and the second image, a defect recognition result corresponding to the first image is determined, where the defect recognition result includes a defect condition of the image object.

[0013] In another aspect, an image processing apparatus is provided, the apparatus comprising:

[0014] an acquisition module, configured to acquire a noisy image corresponding to a first image, wherein the noisy image is obtained by adding first noise data to the first image, and the first image includes an image object;

[0015] an analysis module, configured to perform semantic analysis on the first image to obtain semantic analysis data, wherein the semantic analysis data is used to represent information of the image object in the first image;

[0016] a denoising module, configured to perform image denoising on the noisy image based on the semantic analysis data to obtain a second image;

[0017] A determination module is used to determine a defect recognition result corresponding to the first image based on feature similarity between the first image and the second image, where the defect recognition result includes the defect status of the image object.

[0018] On the other hand, a computer device is provided. The computer device includes a processor and a memory. The memory stores a computer program. The computer program is loaded and executed by the processor to implement the above method.

[0019] On the other hand, an embodiment of the present application provides a storage medium, which is used to store a computer program, and the computer program is used to execute the method of the above aspect.

[0020] On the other hand, an embodiment of the present application provides a computer program product including a computer program, which, when executed on a computer, enables the computer to execute the above method.

[0021] The beneficial effects of the technical solutions provided in the embodiments of the present application include at least:

[0022] After adding first noise data to the first image to obtain a noisy image, semantic analysis is performed on the first image to obtain semantic analysis data corresponding to information representing the image object in the first image. Image denoising is then performed on the noisy image based on the semantic analysis data to obtain a second image. Finally, the defect corresponding to the first image is determined by comparing the feature similarity between the first and second images. That is, by performing semantic analysis on the first image to obtain semantic analysis data, the semantic analysis data is used to guide image denoising. Since the semantic analysis object can reflect information about the image object in the first image, the image object in the noisy image is clearly indicated during image denoising, avoiding removing partial components of the image object as noise and preserving the complete image object in the second image to the greatest extent possible. Thus, when performing subsequent similarity matching, the lack of similarity will only be due to the defective portion of the image object in the first image, rather than any display issues of the image object in the second image caused by denoising. This improves the accuracy and efficiency of image object defect recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] FIG1 is a schematic diagram of an implementation environment provided by an exemplary embodiment of the present application;

[0024] FIG2 is a flow chart of an image processing method provided by an exemplary embodiment of the present application;

[0025] FIG3 is a flowchart of an image processing method provided by another exemplary embodiment of the present application;

[0026] FIG4 is a schematic diagram of a feature fusion process provided by another exemplary embodiment of the present application;

[0027] FIG5 is a flowchart of an image processing method provided by another exemplary embodiment of the present application;

[0028] FIG6 is a flowchart of an image processing method training process provided by another exemplary embodiment of the present application;

[0029] FIG7 is a schematic diagram of an image processing method provided by yet another exemplary embodiment of the present application;

[0030] FIG8 is a schematic diagram showing the effect of an image processing method provided by another exemplary embodiment of the present application;

[0031] FIG9 is a structural block diagram of an image processing apparatus provided by an exemplary embodiment of the present application;

[0032] FIG10 is a structural block diagram of an image processing apparatus provided by another exemplary embodiment of the present application;

[0033] FIG11 is a structural block diagram of a computer device provided by an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0034] To make the objectives, technical solutions, and advantages of this application more clear, the following will further describe the embodiments of this application in detail with reference to the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this application.

[0035] In this application, the terms "first", "second", etc. are used to distinguish identical or similar items with substantially the same effects and functions. It should be understood that there is no logical or temporal dependency between "first" and "second", nor is there any limitation on the quantity and execution order.

[0036] FIG1 is a schematic diagram of an implementation environment provided by an exemplary embodiment of the present application. As shown in FIG1 , the implementation environment includes a terminal 110, a server 120, and a communication network 130. Terminal 110 and server 120 are connected via communication network 130. Communication network 130 may be a wired network or a wireless network, which is not limited herein.

[0037] In the embodiment of the present application, terminal 110 is used to send data to server 120. Optionally, a target application with image recognition functionality is installed in terminal 110, which is not limited in this embodiment. Illustratively, the target application can be a traditional application, a cloud application, a mini-program or application module implemented in a host application, or a web platform, which is not limited in this embodiment.

[0038] After receiving the first image, the terminal 110 generates an object recognition request based on the first image and sends it to the server 120 . The object recognition request is used to request recognition of an image object in the first image.

[0039] After receiving the object recognition request, server 120 adds first noise data to the first image, thereby generating a noisy image corresponding to the first image. Furthermore, server 120 performs semantic analysis on the first image to obtain semantic analysis data. Based on the semantic analysis data, server 120 performs image denoising on the noisy image to obtain a second image. Based on the feature similarity between the first and second images, server 120 determines a defect recognition result corresponding to the first image. The defect recognition result is fed back to terminal 110 for display.

[0040] The above description uses the terminal 110 and the server 120 as examples of computer devices. The computer device is the execution subject of the embodiments of the present application. In some optional embodiments, when the computer device is the terminal 110, the terminal 110 is a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart home appliance, a smart car terminal, a smart speaker, an intelligent voice interaction device, an aircraft, etc., but is not limited to this.

[0041] It is worth noting that when the computer device is a server 120, the server 120 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0042] Cloud technology refers to a hosting technology that unifies hardware, software, and network resources within a wide area network (WAN) or local area network (LAN) to enable data computing, storage, processing, and sharing. Cloud technology is a general term for network technology, information technology, integration technology, management platform technology, and application technology, all based on the cloud computing business model. It can form a resource pool for on-demand, flexible, and convenient use. Cloud computing technology will become a crucial support. Backend services in technical network systems, such as those for video websites, image websites, and more portals, require significant computing and storage resources. With the rapid development and application of the internet industry, every item will likely have its own unique identifier, requiring transmission to backend systems for logical processing. Data of varying levels will be processed separately, requiring robust system support for all types of industry data, which can only be achieved through cloud computing. Alternatively, server 120 can also be implemented as a node in a blockchain system.

[0043] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.) and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards in the relevant regions.

[0044] In conjunction with the above introduction and implementation environment, FIG2 is a flowchart of an image processing method provided in an embodiment of the present application. The method can be executed by a computer device. The computer device can be, for example, the aforementioned terminal, a server, or both a terminal and a server. The embodiment of the present application takes the execution of the method by a server as an example for explanation. The method includes:

[0045] Step 210: Acquire a noisy image corresponding to the first image.

[0046] The noisy image is a result of adding first noise data to the first image, and the first image includes an image object.

[0047] Illustratively, the first image is an image containing at least one image object.

[0048] Optionally, the first image is an R (Red) G (Green) B (Blue) image; or, the first image is a grayscale image.

[0049] In some embodiments, the image object in the first image is used for defect recognition. Defect recognition refers to identifying whether the image object has a defective portion. For example, when the image object in the first image is a screw, embodiments of the present application can identify whether there is a concave portion on the screw cap.

[0050] Optionally, the image object in the first image is displayed in a two-dimensional form; or, the image in the first image is displayed in a three-dimensional form.

[0051] In some embodiments, noise data refers to data that interferes with image content in an image. For example, noise may cause the image to become blurred, distorted, or contain random brightness or color variations.

[0052] Optionally, the noise data includes at least one of salt and pepper noise (black and white pixels that appear randomly in an image), Gaussian noise (pixel values ​​in an image are affected by normally distributed random noise), analog signal noise (noise introduced during image transmission and acquisition due to interference such as noise from electronic components), compression noise (noise introduced during image compression), blur noise (noise corresponding to image blur caused by vibration, motion blur, or optical blur during shooting or transmission), color noise (noise differences between different color channels in an RGB image, such as color aberration), or environmental noise (noise data generated from the environment in the image, such as light changes, shadows, and reflections).

[0053] Optionally, a method for obtaining the first noise data includes at least one of the following methods:

[0054] First, a noise data set is obtained in advance, and at least one noise data is randomly adjusted from the noise data set as the first noise data;

[0055] Second, perform noise analysis on the first image. If noise data already exists in the first image, after obtaining the noise data, replicate the noise data multiple times and use the replicated results as the first noise data. For example, the first image contains 10 pixels, of which pixel 2 corresponds to Gaussian noise. Gaussian noise is introduced as the first noise data into other pixels to change their respective pixel values.

[0056] It is worth noting that the above-mentioned method for obtaining the first noise data is only an example for adaptation and is not limited to this embodiment of the present application.

[0057] Optionally, the first noise data includes at least one type of noise data. When the first noise data includes multiple types of noise data, the multiple types of noise data are noise data of the same noise type, or are noise data of different noise types.

[0058] In some embodiments, the process of adding the first noise data to the first image is referred to as an image noise addition process.

[0059] Optionally, the image noise addition process is performed once or multiple times. When the image noise addition process is performed multiple times, the noise addition result of the previous adding of the first noise data is used as the data to be processed for the current noise addition process. For example: the first noise data is added to image a to obtain noisy image 1, which is the first image noise addition process; the first noise data is added to the noisy image 1 to obtain noisy image 2, which is the second image noise addition process, until the image noise addition process is completed, and the noisy image obtained by the last noise addition process is used as the noisy image corresponding to the first image.

[0060] The noise data applied in each execution of the image noise addition process may be the same or different, which is not limited.

[0061] In an illustrative manner, a specified number of steps is obtained in advance, and noise addition processing is performed on the first image according to the number of noise addition times corresponding to the specified number of steps.

[0062] Step 220: Perform semantic analysis on the first image to obtain semantic analysis data.

[0063] The semantic analysis data is used to represent information about the image object in the first image.

[0064] Optionally, the information of the image object includes at least one of information types such as object direction, position distribution, placement posture, object category, etc. of the image object in the first image.

[0065] The object direction refers to the direction of the image object in the first image. For example, the image object in the first image includes a screw, wherein the screw head faces the left side of the first image.

[0066] The position distribution refers to the position of the image object in the first image. For example, the image object in the first image includes a gear, and the gear is located in the upper left area of ​​the first image.

[0067] Among them, the placement posture refers to the posture angle of the image object in the first image. For example, the image object in the first image includes a gear, wherein the lower edge line of the first image is used as the horizontal line, and the gear is displayed in the first image at an angle of 45 degrees rotated counterclockwise from the horizontal line.

[0068] Among them, the object category refers to the classification corresponding to the image object. For example, the image object in image a includes gears, and the image object in image b includes nuts. Since gears and nuts are workpieces with different functions, their corresponding object categories are different.

[0069] Illustratively, the semantic analysis method for obtaining semantic analysis data includes at least one of the following methods:

[0070] First, a semantic analysis model is pre-trained, the first image is input into the semantic analysis model, and semantic analysis data corresponding to the first image is output;

[0071] Second, the pixel values ​​corresponding to each pixel point in the first image are obtained, and a pixel value threshold is set. The pixel points that reach the pixel value threshold are used as the pixel points corresponding to the image object in the first image, thereby obtaining the pixel point distribution of the image object in the first image. According to the pixel point distribution, the position distribution of the image object in the first image can be obtained. According to the pixel point distribution, the object contour of the image object can also be obtained, and the object category, object direction and placement posture corresponding to the image object can be obtained in turn.

[0072] It is worth noting that the above-mentioned semantic analysis method is only an illustrative example and is not limited to this embodiment of the present application.

[0073] In the case of the semantic analysis model obtained by the above-mentioned pre-training, optionally, the semantic analysis model includes at least one of the model types such as a word vector model, a bag-of-words model, a document vector model, a recurrent neural network (RNN) model (for example, a long short-term memory network, a gated recurrent unit, etc.), a convolutional neural network (CNN) model, an attention mechanism model (for example, a Transformer model), and a pre-trained language model (for example, a BERT model).

[0074] Optionally, the semantic analysis data refers to state information expressed by a single image object; or, the semantic analysis data includes state information corresponding to multiple image objects.

[0075] Step 230 : performing image denoising processing on the noisy image based on the semantic analysis data to obtain a second image.

[0076] In some embodiments, image denoising processing refers to suppressing or eliminating noise data in a noisy image.

[0077] Optionally, the image denoising process includes at least one of the following methods:

[0078] First, mean filtering: replace each pixel in the image with the mean of its surrounding pixels;

[0079] Second, median filtering: replace each pixel in the image with the median value of its surrounding pixels;

[0080] Third, Gaussian filtering: use Gaussian function to blur the image;

[0081] Fourth, wavelet transform: convert the image into the wavelet domain and remove noise by thresholding the wavelet coefficients;

[0082] Fifth, total variation denoising: Optimize the image using the total variation regularization model to reduce noise by minimizing the total variation of the image;

[0083] Sixth, non-local mean denoising: using non-local similarity in the image to estimate noise and perform denoising;

[0084] Seventh, deep learning-based methods: pre-train a neural network model, and use the neural network model to learn and denoise the image.

[0085] It is worth noting that the above-mentioned method for image denoising is only an illustrative example and is not limited to this embodiment of the present application.

[0086] In some embodiments, the second image refers to an image generated by performing image denoising processing on the noisy image.

[0087] Optionally, the first image and the second image are the same, or there are certain differences between the first image and the second image.

[0088] Illustratively, when the first image is an image in which a defective portion exists in the image object, the second image is a corresponding image in the first image in which the defective portion does not exist in the image object.

[0089] In some embodiments, guided by the semantic analysis data, image denoising is performed on the noisy image. This allows the complete image object from the first image to be preserved in the second image. Furthermore, if the image object has a defective portion in the first image, the defective portion is repaired in the second image after denoising. Furthermore, guided by the semantic analysis data, the image object in the second image is generally indistinguishable from the image object in the first image, for example, in terms of object category, orientation, and position. Consequently, when calculating similarity in step 240, the sole cause of insufficient similarity is the defective portion of the image object in the first image, effectively improving the accuracy of defective portion recognition.

[0090] Step 240 : Determine a defect recognition result corresponding to the first image based on feature similarity between the first image and the second image.

[0091] The defect recognition result includes the defect status of the image object.

[0092] Schematically, the image similarity between the first image and the second image is determined by comparing the feature similarity between the first image and the second image. Since the second image is a defect-free image after denoising, the higher the similarity, the higher the possibility that the image object in the first image does not have defective parts. Conversely, if the similarity is lower than a pre-set similarity threshold, it indicates that the image object in the first image has defective parts.

[0093] In some embodiments, a first image feature representation corresponding to the first image is extracted, and a second image feature representation corresponding to the second image is extracted, so that the defect situation of the image object in the first image is obtained by calculating the feature similarity between the first image feature representation and the second image feature representation.

[0094] Optionally, the feature similarity may be calculated by calculating Euclidean distance, cosine similarity, correlation coefficient (calculating the correlation coefficient of two feature representations, i.e., the oblique variance of the two feature representations divided by the product of their standard deviations. The correlation coefficient ranges from [-1, 1], and the closer to 1, the higher the similarity), Hamming distance, and Pearson correlation coefficient (used to measure the linear relationship between two feature representations, calculating the covariance of two feature vectors divided by the product of their standard deviations).

[0095] In some embodiments, the defect condition includes the position distribution of defective parts of the image object in the first image, and the defect category corresponding to the defective parts (for example, dents, cracks, flaws, etc.).

[0096] Optionally, the defect situation is represented by coordinate points and numerical results, wherein the coordinate points are used to indicate the coordinate points corresponding to the defective part, and when different defect categories are pre-defined as different numerical values, the numerical results are used to indicate the defect category corresponding to the defective part; or, the defect situation is represented by a score map, wherein the score map is an image with the same pixel size as the first image, and the pixel channel value range is between 0 and 1, wherein 0 indicates that the pixel value of the pixel point of the first image and the second image is the same, and 1 indicates that the pixel values ​​of the pixel point of the first image and the second image are completely different, therefore, the position corresponding to the pixel channel value 1 is the defective part.

[0097] In summary, the image processing method provided by the embodiment of the present application, after adding first noise data to the first image to obtain a noisy image, performs semantic analysis on the first image to obtain semantic analysis data corresponding to the information used to characterize the image object in the first image, and then performs image denoising on the noisy image based on the semantic analysis data to obtain a second image, and finally determines the defect corresponding to the first image by comparing the feature similarity between the first image and the second image. That is, by performing semantic analysis on the first image to obtain semantic analysis data, the semantic analysis data is used to guide image denoising. Since the semantic analysis object can reflect the information of the image object in the first image, the image object in the noisy image will be clearly indicated during image denoising, avoiding removing part of the image object as noise, and retaining the complete image object in the second image to the greatest extent possible. In this way, when similarity matching is performed subsequently, the reason for the lack of similarity will only be the defective part of the image object in the first image, and will not be the display problem of the image object in the second image caused by denoising, which can improve the accuracy and efficiency of image object defect recognition.

[0098] In an optional embodiment, the semantic analysis process includes encoding processing and decoding processing. Schematically, please refer to Figure 3, which shows a flow chart of an image processing method provided by an exemplary embodiment of the present application, that is, step 220 also includes steps 221 to 224, and step 230 also includes steps 231 and 232. Schematically, as shown in Figure 4, the method includes the following steps.

[0099] Step 221 : downsample the first image to obtain a downsampling result.

[0100] Illustratively, image downsampling refers to reducing the image resolution or image size.

[0101] In this embodiment, the first image is in, represents channel, H represents height, and W represents width.

[0102] Step 222: perform iterative feature encoding processing on the downsampling result to obtain a plurality of semantic feature encoding data.

[0103] Among them, the i+1th semantic feature coding data is obtained by performing coding feature processing on the i-th semantic feature coding data, multiple semantic feature coding data correspond to different feature dimensions, and i is a positive integer.

[0104] Schematically, feature encoding is an image processing method used to extract key point information from an image.

[0105] In some embodiments, feature encoding is performed on the downsampling result to output semantic feature encoding data, which is used as input data for the next feature encoding process, thereby treating multiple feature encoding processes as iterative feature encoding processes.

[0106] Schematically, as the number of feature coding processes increases, the feature dimension corresponding to the semantic feature coding data gradually decreases. For example, the feature dimension of the first layer of semantic feature coding data obtained after the first feature coding process is 32×32, and the feature dimension of the second layer of semantic feature coding data obtained after the feature coding process of the first layer of semantic feature coding data is 16×16, and so on, until the last semantic feature coding data is obtained, and its corresponding feature dimension is the smallest.

[0107] In this embodiment, the iterative feature encoding process is described by taking four times of feature encoding as an example.

[0108] In some embodiments, after the fourth semantic feature encoding data is obtained, it is stored in a semantic data memory.

[0109] Step 223 : performing feature decoding processing on at least one semantic feature encoding data among the plurality of semantic feature encoding data to obtain at least one semantic feature decoding data as semantic analysis data.

[0110] Schematically, feature decoding processing refers to an image processing process of restoring data obtained after feature encoding processing to original data.

[0111] In this embodiment, after a plurality of semantic feature coded data are obtained through semantic feature coding processing, an iterative feature decoding process is performed on one or more of the semantic feature coded data to finally obtain one or more semantic feature decoded data.

[0112] Through the above-mentioned semantic analysis process of encoding and decoding the first image, the obtained semantic feature coding data can reflect clearer semantic information. For the feature dimensions of the semantic feature coding data, feature decoding processing that is more suitable for reflecting the semantics of the image object can be selected to obtain more accurate semantic analysis data.

[0113] Regarding step 223, in one possible implementation, at least two semantic feature encoding data among the plurality of semantic encoding feature data are subjected to feature fusion to obtain a first fused feature representation, and feature decoding processing is performed on the first fused feature representation to obtain semantic feature decoding data.

[0114] Illustratively, after obtaining a plurality of semantic feature encoding data, at least two semantic feature encodings in the plurality of semantic feature encoding data are subjected to feature fusion, thereby obtaining a first fused feature representation.

[0115] Optionally, the at least two semantic feature encoding data are at least two adjacent semantic feature encoding data; or, they are at least two non-adjacent semantic feature encoding feature data.

[0116] In some embodiments, at least two semantic feature encoding data include nth semantic feature encoding data and mth semantic feature encoding data, the nth semantic feature encoding data includes k-layer first semantic sub-data, and the mth semantic feature encoding data includes k-layer second semantic sub-data, where n, m, and k are positive integers; the k-layer second semantic sub-data are respectively subjected to convolution processing to obtain k convolution processing results; the k convolution processing results are feature fused with the j-layer first semantic sub-data in the k-layer first semantic sub-data to obtain the j-layer first fused sub-feature, where 0<j≤k and j is an integer; and a first fused feature representation is obtained based on the k-layer first fused sub-feature.

[0117] Schematically, for a single semantic feature encoding data, the single semantic feature encoding contains multiple layers of semantic sub-data, and the first fused feature representation is obtained by performing feature fusion on the semantic sub-data corresponding to at least two semantic feature encoding data.

[0118] In this embodiment, description is made by taking the case where the at least two semantic feature coding data are implemented as the third semantic feature coding data and the fourth semantic feature coding data as an example, that is, n=3, m=4.

[0119] In this embodiment, the third semantic feature coded data includes three layers of first semantic sub-data, and the fourth semantic feature coded data includes three layers of second semantic sub-data. That is, k=3.

[0120] In the process of feature fusion of the third semantic feature coding data and the fourth semantic feature coding data, the three layers of second semantic sub-data in the fourth semantic feature coding data are convolutionally processed to obtain convolution processing results corresponding to the three layers of second semantic sub-data respectively.

[0121] After obtaining the three convolution processing results, the first convolution processing result is concatenated with the third semantic feature encoding data to obtain the first fusion sub-feature of the first layer. The second convolution processing result is concatenated with the first fusion sub-feature of the first layer to obtain the first fusion sub-feature of the second layer. This process is repeated until the feature fusion is completed to obtain the three-layer first fusion sub-feature. The three-layer first fusion sub-feature is represented as the first fusion feature, which is recorded as J=3 means that there are three layers in the third semantic feature encoding data. Represents the low-scale output feature map in the third semantic feature encoding data (such as the fusion sub-features of each layer), Represents the high-scale output feature map in the fourth feature encoding data (such as the convolution processing result), Represents a convolutional module consisting of a 3×3 convolutional layer, a normalization layer, and an activation layer.

[0122] Schematically, please refer to Figure 4, which shows a schematic diagram of the feature fusion process provided by an exemplary embodiment of the present application. As shown in Figure 4, in the process of feature fusing the third semantic feature encoding data 410 and the fourth semantic feature encoding data 420, taking the first-layer second semantic sub-data 421 in the fourth semantic feature encoding data 420 as an example, the first-layer second semantic sub-data 421 in the fourth semantic feature encoding data 420 is convolved, and the convolution result obtained is feature spliced ​​with the first-layer first semantic sub-data 411 to obtain a first fusion result, the convolution result obtained by convolving the second-layer second semantic sub-data 422 is feature spliced ​​with the first fusion result to obtain a second fusion result, the convolution result obtained by convolving the third-layer second semantic sub-data 423 is feature spliced ​​with the second fusion result, and finally the first-layer first fusion sub-feature 431 is obtained. Finally, the first-layer first fusion sub-feature 431, the second-layer first fusion sub-feature 432 and the third-layer first fusion sub-feature 433 corresponding to the three layers of second semantic sub-data are represented as the first fusion feature.

[0123] Since different semantic feature coding data have different feature dimensions, a first fused feature representation is obtained by fusing multiple semantic feature coding data. This first fused feature representation can carry the characteristics of image objects under multiple feature dimensions and has richer useful information, thereby improving the expression ability and accuracy of semantic feature coding data for image objects.

[0124] In some embodiments, semantic segmentation is performed on target semantic feature encoding data among multiple semantic feature encoding data to obtain a first semantic segmentation feature, and the first semantic segmentation feature is used to represent a pixel classification result in the first image.

[0125] Schematically, semantic segmentation refers to classifying the pixels in the input image.

[0126] In this embodiment, after obtaining four semantic feature encoding data, the fourth semantic feature encoding data is input into the pre-trained first semantic segmentation model for semantic segmentation, thereby outputting the first semantic segmentation feature, which is used to characterize the classification results of each pixel point in the first image, for example: pixel point a and pixel point b belong to the pixel points corresponding to object 1 in the first image, and pixel point c, pixel point d and pixel point e belong to the pixel points corresponding to object 2 in the first image.

[0127] In this embodiment, the first semantic segmentation model is implemented as a spatial gridding module (SGM), and its corresponding module structure has three layers, consisting of a residual module (ResnetBlock), a spatial transformation module (Spatial Transformer) and a residual module. Among them, ResnetBlock is the basic building block in the residual network and is used to learn image features. It consists of two convolutional layers, which contain residual connections. SpatialTransformer is a module for learning image geometric transformations. It improves the robustness of the network to geometric changes in the image by learning how to perform geometric transformations such as translation, rotation, and scaling on the input image.

[0128] In one possible implementation, step 230 of performing image denoising on the noisy image based on the semantic analysis data to obtain a second image includes:

[0129] Step 231 : performing iterative feature coding processing on the noisy image to obtain a plurality of denoised feature coding data.

[0130] Schematically, the process of performing image denoising on a noisy image also includes two processes: feature encoding processing and feature decoding processing.

[0131] In this embodiment, the feature coding process is iterated four times as an example. First, the noisy image is feature coded to obtain the first denoised feature coded data. Secondly, the first denoised feature coded data is feature coded to obtain the second denoised feature coded data. Then, the second denoised feature coded data is feature coded to obtain the third denoised feature coded data. Finally, the third denoised feature coded data is feature coded to obtain the fourth denoised feature coded data.

[0132] In some embodiments, the qth denoised feature coding data is feature fused with the downsampling result to obtain a second fused feature representation; the second fused feature representation is iteratively feature coded to obtain multiple semantic feature coding data, where q<p and p,q are positive integers.

[0133] In this embodiment, after obtaining the first denoised feature coding data by performing feature coding processing on the noisy image, the first denoised feature coding data is feature fused with the downsampling result corresponding to the first image to obtain a second fused feature representation, and the above-mentioned iterative feature coding processing is performed on the second fused feature representation to obtain the above-mentioned multiple semantic feature coding data.

[0134] Step 232 : performing iterative feature decoding processing on the pth denoised feature coded data among the plurality of denoised feature coded data based on the semantic feature decoded data and the semantic feature coded data to obtain a second image.

[0135] Schematically, after obtaining the pth denoised feature coding data, the semantic feature coding data and / or semantic feature decoding data obtained in the above steps are iteratively feature decoded with the pth denoised feature coding data to obtain denoised feature decoding data. After iterative feature decoding of the denoised feature decoding data, a second image is finally obtained.

[0136] In this embodiment, after obtaining the last denoised feature coding data, it is stored together with the fourth semantic feature coding data to obtain fused data, and the semantic feature decoding data and the de-fused data are feature decoded to obtain the first denoised feature decoding data, and the semantic feature decoding data, the fourth denoised feature coding data and the first denoised feature decoding data are feature decoded to obtain the second denoised feature decoding data, and iterative feature decoding is performed until the second image is obtained.

[0137] In some embodiments, the pth denoising feature encoding data is iteratively feature decoded p times based on the semantic feature decoding data to obtain second noise data, and the second noise data is used to indicate the prediction result of the first noise data; the noisy image is denoised based on the second noise data to obtain data denoising features, and the data denoising features are used to characterize the image feature representation corresponding to the image after the second noise data is removed from the noisy image; the data denoising features are image decoded to obtain a second image.

[0138] In some embodiments, semantic segmentation is performed on the pth denoised feature encoding data to obtain a second semantic segmentation feature, and the second semantic segmentation feature is used to characterize the classification result of the pixel points in the noisy image; the first semantic segmentation feature and the second semantic segmentation feature are feature spliced ​​to obtain a semantic splicing feature; and the semantic splicing feature is iteratively feature decoded based on the semantic feature decoding data and the pth denoised feature encoding data to obtain a second noise data.

[0139] In this embodiment, taking p as 4 as an example, after obtaining the fourth denoised feature coded data, the fourth denoised feature coded data is input into the pre-trained second semantic segmentation model for semantic segmentation, and the second semantic segmentation feature is output, wherein the second semantic segmentation feature is used to characterize the classification result corresponding to each pixel in the noisy image. The second semantic segmentation model has the same model structure as the first semantic segmentation model, that is, the second semantic segmentation model (Spatial Gridding Module, SDM) also includes a layer of residual module (ResnetBlock), a layer of spatial transformation module (Spatial Transformer) and a layer of residual module (ResnetBlock).

[0140] In this embodiment, after obtaining the second semantic segmentation feature, the first semantic segmentation feature and the second semantic segmentation feature are concatenated to obtain a semantic concatenation feature.

[0141] In this embodiment, the semantic splicing feature is subjected to four iterative feature decoding processes in combination with the semantic feature decoding data and the fourth denoising feature encoding data, thereby obtaining the second noise data.

[0142] In this embodiment, taking the feature decoding process iteratively performed four times as an example, after performing four iterative feature decoding processes on the fourth denoised feature coded data, the second noise data ε is obtained. θ , perform data denoising on it, and thus obtain data denoising features

[0143] Optionally, the first noise data is the same as the second noise data; or, the first noise data is different from the second noise data.

[0144] In this embodiment, the denoising formula obtained in advance is used to perform data denoising on the noisy image based on the second noise data to obtain a data denoising feature, which is expressed as p θ (x0|x t-1 ), where x0 represents the first image, x t-1 It represents the noisy sub-image obtained after the image is denoised in step t-1.

[0145] In this embodiment, after obtaining the data denoising feature, it is output to the pre-trained image decoder for image decoding processing, and the second image is output. The second image is recorded as

[0146] It can be seen that when performing image denoising on a noisy image, the semantic feature decoding data as semantic analysis data is used as a guide for image denoising, so that the image object is taken into consideration during image denoising, avoiding the image object itself from being mistakenly removed as noise, thereby improving the image denoising accuracy.

[0147] In one possible implementation, step 240, determining a defect recognition result corresponding to the first image based on feature similarity between the first image and the second image, includes:

[0148] Step 241 : extracting a first image feature representation corresponding to the first image, and extracting a second image feature representation corresponding to the second image.

[0149] Schematically, the defect condition of the image object includes the position distribution of defects in the image object and the defect category.

[0150] Schematically, the first image and the second image are input into a common feature space of a same pre-trained feature extraction model to extract a first image feature representation and a second image feature representation corresponding to the first image.

[0151] In some embodiments, the first image feature representation includes feature representations of multiple different feature dimensions, and the second image feature representation also includes feature representations of multiple different feature dimensions, wherein the distribution of the feature dimensions in the first image feature representation is consistent with the distribution of the feature dimensions in the second image feature representation.

[0152] In this embodiment, the pre-obtained feature extraction model is implemented as a convolutional neural network resnet50.

[0153] Step 242 : Obtain a defect score distribution image based on the cosine similarity between the first image feature representation and the second image feature representation.

[0154] The pixel values ​​in the defect score distribution image are used to indicate the difference between the pixel values ​​in the first image and the pixel values ​​in the second image.

[0155] In this embodiment, the defect scores of different feature dimensions in the first image feature representation and the second image feature representation are calculated according to the cosine similarity formula, thereby obtaining a defect score distribution image. The cosine similarity formula can refer to the following formula 1.

[0156] Formula 1:

[0157] Where n represents the image feature representation corresponding to the nth feature dimension.

[0158] Schematically, the defect score partial image is a grayscale image with the same pixel size as the first image, and the corresponding pixel channel value range includes 0 to 1. When the channel value is 0, it means that the pixel value in the first image is the same as the pixel value at the same position in the second image. When the channel value is 1, it means that the pixel value in the first image is different from the pixel value at the same position in the second image.

[0159] Step 243 : perform feature upsampling on the defect score distribution image to obtain a defect position distribution image.

[0160] The defect location distribution image is used to indicate the location distribution of defects in the image object.

[0161] Schematically, after obtaining the defect score distribution image, feature upsampling is performed on it and the number of feature dimensions to be applied is selected to obtain the defect location distribution image. The feature upsampling process can be referred to the following formula 2.

[0162] Formula 2:

[0163] Among them, σ n represents the upsampling rate and N represents the number of feature dimensions used.

[0164] In step 244 , average pooling is performed on the defect position distribution image to obtain a defect category indicating a defect in the image object.

[0165] Schematically, after obtaining the defect position distribution image, a global average pooling process is performed on it, and the maximum value obtained by the process is used as the defect category result.

[0166] In step 245 , the defect location distribution image and the defect category result are used as the defect situation.

[0167] Finally, the defect location distribution image and defect category results are used as the defect recognition results.

[0168] In some embodiments, the defect scores of the first image and the second image at the RGB level may also be calculated and combined with the defect location distribution image and the defect category results to serve as the defect situation.

[0169] In summary, the image processing method provided in the embodiment of the present application, after adding first noise data to a first image to obtain a noisy image, performs semantic analysis on the first image to obtain semantic analysis data corresponding to information used to characterize image objects in the first image, and then performs image denoising on the noisy image based on the semantic analysis data to obtain a second image. Finally, the defect corresponding to the first image is determined by comparing the feature similarity between the first image and the second image. That is, by performing semantic analysis on the first image to obtain semantic analysis data, the semantic analysis data is used to guide image denoising, ensuring that the image objects in the denoised second image are consistent with the image objects in the first image, and then comparing the feature similarity to determine the defect of the image objects in the first image, which can improve the accuracy and efficiency of image object defect recognition.

[0170] In this embodiment, by performing feature encoding processing and feature decoding processing on the first image, the output result can include state information of the image object in the first image, thereby improving the accuracy of semantic analysis.

[0171] In this embodiment, after performing feature fusion on at least two semantic feature encoded data, feature decoding processing is performed on the data, which can improve the information richness of the semantic feature decoded data finally obtained.

[0172] In this embodiment, after the two semantic feature encoding data are layered, feature fusion is performed through convolution processing, which can improve the decoding accuracy of the semantic feature decoding data.

[0173] In this embodiment, the second image obtained by iteratively feature encoding the noisy image and iteratively decoding the denoised feature encoded data based on the semantic feature decoding data can improve the consistency of the second image and the first image.

[0174] In this embodiment, feature encoding processing is performed on the denoised feature coded data and the down-sampling result after feature fusion, thereby improving the accuracy of semantic analysis.

[0175] In an optional embodiment, the above-mentioned image object recognition process is implemented through a deep learning-based approach. For illustration, please refer to FIG5 , which shows a flow chart of an image processing method provided by an exemplary embodiment of the present application. As shown in FIG5 , the method includes the following steps.

[0176] Step 2201: Input the first image into a pre-trained semantic analysis model, and output semantic analysis data.

[0177] Schematically, the semantic analysis model includes an encoder and a decoder, which are respectively used to perform the above-mentioned feature encoding processing and feature decoding processing, thereby obtaining semantic analysis data.

[0178] Among them, the semantic analysis model includes multiple semantic encoding modules arranged in sequence, a first semantic analysis model and a semantic decoding module. The image denoising data is input into the first semantic encoding module, and the first semantic feature encoding data is output. The first semantic feature encoding data is input into the second semantic encoding module, and the second semantic feature encoding data is output. In this way, iterative feature encoding processing is performed to finally obtain multiple semantic feature encoding data.

[0179] The fourth semantic feature is encoded and input into the pre-trained first semantic segmentation model, and the first semantic segmentation feature is output.

[0180] A first fused feature representation obtained by performing feature fusion on at least two semantic feature encoding data is input into a decoding module, and semantic feature decoding data is output.

[0181] Among them, the semantic decoding module is implemented as a semantic guided decoding module (Semantic Guided Encoder Blocks, SGEB), and its network structure includes a three-layer residual module (ResnetBlock).

[0182] The semantic feature encoding data and the semantic feature decoding data are used as semantic analysis data.

[0183] Step 2301: Input the semantic analysis data and the noisy image into a pre-trained image denoising model to obtain a second image.

[0184] Illustratively, the image denoising model includes a plurality of denoising encoding modules and a plurality of denoising decoding modules arranged in sequence, which are respectively used to perform the above-mentioned denoising feature encoding processing and denoising feature decoding processing, thereby obtaining a second image.

[0185] The image denoising result is input into the first denoising coding module, and the first denoising feature coding data is output. The first denoising feature coding data is input into the second denoising coding module, and the second denoising feature coding data is output. The iterative coding process is performed to obtain the final denoising feature coding data.

[0186] The last denoised feature encoding data, the fourth semantic feature encoding data, and the semantic feature decoding data are input into the first denoising decoding module, and the first denoised feature decoding data is output. The first denoised feature decoding data is input into the second denoising decoding module, and the second denoised feature decoding data is output. The iterative feature decoding process is performed to obtain the second noise data. The data denoising process is performed based on the second noise data to finally obtain the second image. The second noise data can be referred to Formula 3.

[0187] Formula 3:

[0188] in, Represents the decoding module in the image denoising model, Represents the denoising data storage module in the image denoising model, Represents the encoding module in the image denoising model, Semantic data storage module representing the semantic analysis model, represents the convolutional neural network layer, Represents the decoding module of the semantic analysis model.

[0189] Among them, the semantic encoding module is implemented as a semantic guided decoding module (Semantic Guided Decoder Blocks, SGDB) which also includes three layers of ResnetBlock, and the denoising decoding module includes a layer of residual module (ResnetBlock), a layer of spatial transformation module (SpatialTransformer), a layer of residual module (ResnetBlock) and an upsampling module (Upsample). Therefore, the three-layer output results of the fourth semantic encoding module (the fourth semantic feature encoding data) and the three-layer output results of the semantic decoding module (semantic feature decoding data) are added together, and then connected to the corresponding residual module in the first denoising decoding module, and the fourth denoising feature encoding data is input into the first denoising decoding module, so as to perform denoising decoding processing and obtain the first denoising feature decoding data.

[0190] Below, the training process of the semantic analysis model and the image denoising model is explained. As shown in Figure 6, a flow chart of the image recognition method training process provided by an exemplary embodiment of the present application is shown. As shown in Figure 6, the method includes the following steps.

[0191] Step 610: Obtain a sample noisy image corresponding to the first sample image.

[0192] The sample noise-added image is a result obtained by adding the first sample noise data to the first sample image.

[0193] Illustratively, a first sample image for training is first obtained, the first sample image is input into a pre-trained encoder to obtain a latent variable representation, and noise data is randomly added to the latent variable representation to obtain a sample noisy image.

[0194] Illustratively, the first sample image is a sample image without defects.

[0195] Step 620: Input the first sample image into the sample analysis model, and output sample semantic analysis data.

[0196] The first sample image is input into the sample analysis model for feature encoding processing, feature decoding processing and feature fusion to obtain semantic analysis data.

[0197] Step 630: Input the sample semantic analysis data and the sample noisy image into a sample denoising model, and output first predicted noise data.

[0198] In the process of inputting the noisy image into the sample denoising model for image denoising processing, the semantic analysis data is input into the sample denoising model, and the first predicted noise data is obtained as an output.

[0199] Step 640: Based on the difference between the first sample noise data and the first predicted noise data, the sample analysis model and the sample denoising model are trained to obtain a semantic analysis model and an image denoising model.

[0200] By calculating the L2 distance loss value between the first sample noise data and the first predicted noise data, the sample analysis model and the sample denoising model are trained to obtain the semantic analysis model and the image denoising model.

[0201] In summary, the image processing method provided in the embodiment of the present application, after adding first noise data to a first image to obtain a noisy image, performs semantic analysis on the first image to obtain semantic analysis data corresponding to information used to characterize image objects in the first image, and then performs image denoising on the noisy image based on the semantic analysis data to obtain a second image. Finally, the defect corresponding to the first image is determined by comparing the feature similarity between the first image and the second image. That is, by performing semantic analysis on the first image to obtain semantic analysis data, the semantic analysis data is used to guide image denoising, ensuring that the image objects in the denoised second image are consistent with the image objects in the first image, and then comparing the feature similarity to determine the defect of the image objects in the first image, which can improve the accuracy and efficiency of image object defect recognition.

[0202] Schematically, please refer to FIG7 , which shows a schematic diagram of an image processing method provided by an exemplary embodiment of the present application. As shown in FIG7 , the method includes the following process.

[0203] A first image 701 is acquired, and the first image 701 is input into an encoder 702 to obtain a latent variable representation 703 . Image denoising is performed on the latent variable representation 703 to obtain a noisy image 704 .

[0204] The noisy image 704 is input into the image denoising model 710 for image denoising. Meanwhile, the first image 701 is input into the semantic analysis model 720 for semantic analysis.

[0205] During the semantic analysis process, the first image 701 is feature downsampled to obtain a downsampling result. At this time, during the image denoising process, the first denoised feature coding data is obtained by feature coding through the coding module 1 in the encoder in the image denoising model 710, and the data and the downsampling result are input into the coding module a in the semantic analysis model 720 for feature coding. The semantic feature coding data output by the coding module c and the coding module d are respectively input into the feature fusion module 721 for feature fusion after passing through the coding module b, the coding module c, and the coding module d to obtain a first fused feature representation. The first fused feature representation is input into the decoding module a to obtain semantic feature decoding data. In addition, the fourth semantic feature coding data output by the coding module d is input into the semantic segmentation model a to obtain a first semantic segmentation feature.

[0206] Among them, encoding module a, encoding module b, encoding module c, and encoding module d are implemented as four semantically guided encoding modules (SGEBs). Therefore, encoding module a corresponds to SGEB1, encoding module b corresponds to SGEB2, encoding module c corresponds to SGEB3, and encoding module d corresponds to SGEB4. The network structure of SGEB includes a residual module (ResnetBlock), a spatial transformer module (Spatial Transformer), a ResnetBlock, and a downsampling module (Downsample).

[0207] Among them, the semantic segmentation model a is implemented as an SGM model.

[0208] Among them, the decoding module a is implemented as a semantic guided decoding module (Semantic Guided Decoder Blocks, SGDB).

[0209] During the image denoising process, iterative feature coding processing is performed through encoding module 1, encoding module 2, encoding module 3 and encoding module 4, and the fourth denoising feature coding data obtained by encoding module 4 is input into the semantic segmentation model b to obtain the second semantic segmentation feature. The first semantic segmentation feature and the second semantic segmentation feature are feature fused and input into the decoding module 1 for decoding to obtain the first denoising feature coding data, and the first denoising feature coding data and the first fused feature representation are jointly input into the decoding module 2 for feature decoding, and then pass through the decoding module 3 and the decoding module 4 in sequence to obtain the second noise data.

[0210] Among them, encoding module 1, encoding module 2, encoding module 3 and encoding module 4 are implemented as four stable diffusion encoding modules (Stable Diffusion Encoder Blocks, SDEB). Therefore, encoding module 1 corresponds to SDEB1, encoding module 2 corresponds to SDEB2, encoding module 3 corresponds to SDEB3, and encoding module 4 corresponds to SDEB4.

[0211] Among them, decoding module 1, decoding module 2, decoding module 3 and decoding module 4 are implemented as four stable diffusion decoding modules (SDDB). Therefore, decoding module 1 corresponds to SDDB1, decoding module 2 corresponds to SDDB2, decoding module 3 corresponds to SDDB3, and decoding module 4 corresponds to SSDDB4.

[0212] After performing data denoising on the second noise data, a denoised feature representation 705 is obtained. After inputting the denoised feature representation 705 into the decoder 706, a second image 707 is obtained. The first image 701 and the second image 707 are input into the feature space 730, and the first image feature representation 731 and the second image feature representation 732 are extracted. Finally, the defect situation is obtained based on the cosine similarity between the first image feature representation 731 and the second image feature representation 732.

[0213] For illustration, please refer to Figure 8, which shows a comparison diagram of the effects of the image processing method provided by an exemplary embodiment of the present application. As shown in Figure 8, the first image 801 is the input image to be detected, image 802 is the image obtained by denoising after adding noise using related technologies, image 803 is the image obtained after processing using the solution of the present application, and image 804 is the contour reference image. It can be seen from Figure 8 that the accuracy of the image processing results of the image processing solution provided by the present application is higher than the accuracy corresponding to the images obtained by processing using related technologies.

[0214] In summary, the image processing method provided in the embodiments of the present application, after adding first noise data to a first image to obtain a noisy image, performs semantic analysis on the first image to obtain semantic analysis data corresponding to information used to characterize image objects in the first image, and then performs image denoising on the noisy image based on the semantic analysis data to obtain a second image. Ultimately, the defect corresponding to the first image is determined by comparing the feature similarity between the first image and the second image. That is, by performing semantic analysis on the first image to obtain semantic analysis data, the semantic analysis data is used to guide image denoising, ensuring that the image objects in the denoised second image are consistent with the image objects in the first image, and then comparing the feature similarity to determine the defect of the image objects in the first image, thereby improving the accuracy and efficiency of image object defect recognition.

[0215] This technical solution proposes a multi-class defect detection method based on the diffusion model framework, improves the existing diffusion model denoising network framework, adds a semantic guidance network, and improves the category error and semantic error problems of the existing diffusion model when dealing with multi-class defect detection tasks. It maintains the consistency of the semantic information of the input image and the reconstructed image while reconstructing large-area defect areas, and can effectively reconstruct defects of different categories and types into normal samples. It also extracts the input image and the reconstructed image through the feature extraction network to effectively detect and locate defects, and can cope with the detection and positioning of multi-category defects in actual industrial scenarios.

[0216] FIG9 is a block diagram of an image processing apparatus provided by an exemplary embodiment of the present application. As shown in FIG9 , the apparatus includes the following parts:

[0217] An acquisition module 910 is configured to acquire a noisy image corresponding to a first image, where the noisy image is obtained by adding first noise data to the first image, and the first image includes an image object.

[0218] An analysis module 920 is configured to perform semantic analysis on the first image to obtain semantic analysis data, where the semantic analysis data is used to represent information about the image object in the first image.

[0219] a denoising module 930 configured to perform image denoising on the noisy image based on the semantic analysis data to obtain a second image;

[0220] The determination module 940 is configured to determine a defect recognition result corresponding to the first image based on feature similarity between the first image and the second image, where the defect recognition result includes a defect condition of the image object.

[0221] In some embodiments, as shown in FIG10 , the analysis module 920 includes:

[0222] The sampling unit 921 is configured to perform image downsampling on the first image to obtain a downsampling result;

[0223] an encoding unit 922 configured to perform iterative feature encoding processing on the downsampling result to obtain a plurality of semantic feature encoding data, wherein the i+1th semantic feature encoding data is obtained by performing feature encoding processing on the i-th semantic feature encoding data, the plurality of semantic feature encoding data respectively corresponding to different feature dimensions, where i is a positive integer;

[0224] The decoding unit 923 is configured to perform feature decoding processing on at least one semantic feature encoding data among the plurality of semantic feature encoding data to obtain at least one semantic feature decoded data as semantic analysis data.

[0225] In some embodiments, the analysis module 920 further includes:

[0226] a fusion unit 924 configured to fuse at least two semantic feature encoding data among the plurality of semantic feature encoding data to obtain a first fused feature representation;

[0227] The decoding unit 923 is configured to perform feature decoding processing on the first fused feature representation to obtain the semantic feature decoding data as semantic analysis data.

[0228] In some embodiments, the at least two semantic feature encoding data include nth semantic feature encoding data and mth semantic feature encoding data, the nth semantic feature encoding data includes k layers of first semantic sub-data, the mth semantic feature encoding data includes k layers of second semantic sub-data, and n, m, and k are positive integers;

[0229] The fusion unit 924 is used to perform convolution processing on the k layers of second semantic sub-data respectively to obtain k convolution processing results; perform feature fusion on the k convolution processing results and the j-th layer of first semantic sub-data to obtain the j-th layer first fusion sub-feature, 0<j≤k and j is an integer; and obtain the first fusion feature representation based on the k-th layer first fusion sub-feature.

[0230] In some embodiments, the denoising module 930 is further used to perform iterative feature coding processing on the noisy image to obtain multiple denoised feature coding data; based on the semantic feature decoding data and the semantic feature coding data, iterative feature decoding processing is performed on the pth denoised feature coding data in the multiple denoised feature coding data to obtain the second image, where p is a positive integer.

[0231] In some embodiments, the denoising module 930 is further used to perform feature fusion on the qth denoised feature coding data and the downsampling result to obtain a second fused feature representation, where q<p and q is a positive integer; and perform iterative feature coding processing on the second fused feature representation to obtain multiple semantic feature coding data.

[0232] In some embodiments, the denoising module 930 is also used to perform feature decoding processing on the pth denoised feature encoding data based on the semantic feature decoding data and the semantic feature encoding data to obtain second noise data, and the second noise data is used to indicate the prediction result of the first noise data; based on the second noise data, the noisy image is subjected to data denoising processing to obtain data denoising features, and the data denoising features are used to characterize the image feature representation corresponding to the image after the noisy image is freed from the second noise data; and the data denoising features are subjected to image decoding processing to obtain the second image.

[0233] In some embodiments, the denoising module 930 is also used to perform semantic segmentation on the target semantic feature encoding data among the multiple semantic feature encoding data to obtain a first semantic segmentation feature, and the first semantic segmentation feature is used to characterize the pixel classification result in the first image; perform semantic segmentation on the pth denoised feature encoding data to obtain a second semantic segmentation feature, and the second semantic segmentation feature is used to characterize the pixel classification result in the noisy image; perform feature splicing on the first semantic segmentation feature and the second semantic segmentation feature to obtain a semantic splicing feature; perform iterative feature decoding processing on the semantic splicing feature based on the semantic feature decoding data and the pth denoised feature encoding data to obtain the second noise data.

[0234] In some embodiments, the defect condition of the image object includes the position distribution of defects of the image object and the defect category;

[0235] The determination module 940 is further used to extract a first image feature representation corresponding to the first image, and to extract a second image feature representation corresponding to the second image; based on the cosine similarity between the first image feature representation and the second image feature representation, a defect score distribution image is obtained, and the pixel values ​​in the defect score distribution image are used to indicate the difference between the pixel values ​​in the first image and the pixel values ​​in the second image; feature upsampling is performed on the defect score distribution image to obtain a defect position distribution image, and the defect position distribution image is used to indicate the position distribution of defects in the image object; average pooling is performed on the defect position distribution image to obtain a defect category of the defect in the image object; and the defect position distribution image and the defect category are used as the defect recognition result.

[0236] In some embodiments, the analysis module 920 is further configured to input the first image into a pre-trained semantic analysis model and output the semantic analysis data;

[0237] The denoising module 930 is further configured to input the semantic analysis data and the noisy image into a pre-trained image denoising model to obtain the second image.

[0238] In some embodiments, the apparatus further comprises:

[0239] A training module 950 is configured to obtain a sample noisy image corresponding to a first sample image, the sample noisy image being the result of adding first sample noise data to the first sample image; inputting the first sample image into a sample analysis model, and outputting sample semantic analysis data; inputting the sample semantic analysis data and the sample noisy image into a sample denoising model, and outputting first predicted noise data; and training the sample analysis model and the sample denoising model based on the difference between the first sample noise data and the first predicted noise data to obtain the semantic analysis model and the image denoising model.

[0240] In summary, the image processing device provided in the embodiment of the present application, after adding first noise data to a first image to obtain a noisy image, performs semantic analysis on the first image to obtain semantic analysis data corresponding to information used to characterize image objects in the first image, and then performs image denoising on the noisy image based on the semantic analysis data to obtain a second image. Finally, the defect corresponding to the first image is determined by comparing the feature similarity between the first image and the second image. That is, by performing semantic analysis on the first image to obtain semantic analysis data, the semantic analysis data is used to guide image denoising, ensuring that the image objects in the denoised second image are consistent with the image objects in the first image, and then comparing the feature similarity to determine the defect of the image objects in the first image, which can improve the accuracy and efficiency of image object defect recognition.

[0241] It should be noted that the image processing device provided in the above embodiment is merely an example of the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the image processing device provided in the above embodiment and the image processing method embodiment are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0242] FIG11 illustrates a block diagram of a computer device 1100 according to an exemplary embodiment of the present application. Computer device 1100 may be a portable mobile terminal, such as a smartphone, tablet computer, MP3 player (Moving Picture Experts Group Audio Layer III), MP4 player (Moving Picture Experts Group Audio Layer IV), laptop computer, or desktop computer. Computer device 1100 may also be referred to as user equipment, portable terminal, laptop terminal, desktop terminal, or other similar names.

[0243] Typically, the computer device 1100 includes a processor 1101 and a memory 1102 .

[0244] The processor 1101 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 1101 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 1101 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 1101 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1101 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.

[0245] The memory 1102 may include one or more computer-readable storage media, which may be non-transitory. The memory 1102 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 1102 is used to store at least one instruction, which is used to be executed by the processor 1101 to implement the classification model training method and data classification method provided in the method embodiment of the present application.

[0246] In some embodiments, the computer device 1100 may optionally include other components. Those skilled in the art will understand that the structure shown in Figure 11 does not constitute a limitation on the computer device 1100, and it may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component arrangement.

[0247] In addition, an embodiment of the present application further provides a storage medium, which is used to store a computer program, and the computer program is used to execute the method provided by the above embodiment.

[0248] An embodiment of the present application further provides a computer program product including a computer program, which, when executed on a computer, enables the computer to execute the method provided in the above embodiment.

[0249] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program. The program can be stored in a computer-readable storage medium, which can be a computer-readable storage medium included in the memory in the above embodiments; or a separate computer-readable storage medium that is not installed in the terminal. The computer-readable storage medium stores at least one instruction, at least one program, code set, or instruction set. The at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the classification model training method and data classification method described in any of the above embodiments.

[0250] Optionally, the computer-readable storage medium may include: a read-only memory (ROM), a random access memory (RAM), a solid-state drive (SSD), or an optical disk. Among them, the random access memory may include a resistance random access memory (ReRAM) and a dynamic random access memory (DRAM). The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0251] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk, or an optical disk, etc.

[0252] The above description is merely an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. An image processing method, the method being executed by a computer device, the method comprising: Acquire a noisy image corresponding to a first image, wherein the noisy image is obtained by adding first noise data to the first image, wherein the first image includes an image object; Performing semantic analysis on the first image to obtain semantic analysis data, wherein the semantic analysis data is used to characterize the image object; Performing image denoising processing on the noisy image based on the semantic analysis data to obtain a second image; Based on the feature similarity between the first image and the second image, a defect recognition result corresponding to the first image is determined, and the defect recognition result includes the defect situation of the image object.

2. According to the method of claim 1, the performing semantic analysis on the first image to obtain semantic analysis data comprises: Performing image downsampling on the first image to obtain a downsampling result; Performing iterative feature coding processing on the downsampling result to obtain a plurality of semantic feature coding data, wherein the i+1th semantic feature coding data is obtained by performing feature coding processing on the ith semantic feature coding data, the plurality of semantic feature coding data respectively correspond to different feature dimensions, and i is a positive integer; A feature decoding process is performed on at least one semantic feature encoding data among the plurality of semantic feature encoding data to obtain at least one semantic feature decoding data as the semantic analysis data.

3. According to the method of claim 2, the step of performing feature decoding processing on at least one semantic feature encoding data among the plurality of semantic feature encoding data to obtain at least one semantic feature decoding data as the semantic analysis data comprises: Performing feature fusion on at least two of the plurality of semantic feature coded data to obtain a first fused feature representation; The first fused feature representation is subjected to feature decoding processing to obtain the semantic feature decoding data as the semantic analysis data.

4. The method according to claim 3, wherein the at least two semantic feature encoding data include an nth semantic feature encoding data and an mth semantic feature encoding data, the nth semantic feature encoding data includes k layers of first semantic sub-data, the mth semantic feature encoding data includes k layers of second semantic sub-data, and n, m and k are positive integers; The step of fusing at least two semantic feature encoding data among the plurality of semantic feature encoding data to obtain a first fused feature representation includes: The k layers of second semantic sub-data are respectively subjected to convolution processing to obtain k convolution processing results; Perform feature fusion on the k convolution processing results and the first semantic sub-data of the j-th layer to obtain the first fused sub-feature of the j-th layer, where 0<j≤k and j is an integer; The first fused feature representation is obtained based on the k-layer first fused sub-features.

5. The method according to claim 2, wherein the performing image denoising on the noisy image based on the semantic analysis data to obtain a second image comprises: Performing iterative feature coding processing on the noisy image to obtain a plurality of denoising feature coding data; Based on the semantic feature decoding data and the semantic feature encoding data, The pth denoised feature coded data is iteratively decoded to obtain the second image, where p is a positive integer.

6. The method according to claim 5, wherein the step of performing iterative feature encoding processing on the downsampling result to obtain a plurality of semantic feature encoding data comprises: Perform feature fusion on the qth denoised feature encoding data and the downsampling result to obtain a second fused feature representation, wherein q<p and q is a positive integer; An iterative feature encoding process is performed on the second fused feature representation to obtain a plurality of semantic feature encoding data.

7. The method according to claim 5, wherein the iterative feature decoding process is performed on the pth denoised feature coded data among the plurality of denoised feature coded data based on the semantic feature decoded data and the semantic feature coded data to obtain the second image, comprising: Performing iterative feature decoding processing on the p-th denoised feature encoding data based on the semantic feature decoding data and the semantic feature encoding data to obtain second noise data, where the second noise data is used to indicate a prediction result of the first noise data; Performing data denoising processing on the noisy image based on the second noise data to obtain a data denoising feature, wherein the data denoising feature is used to represent an image feature representation corresponding to the image after the second noise data is removed from the noisy image; Perform image decoding processing on the data denoising feature to obtain the second image.

8. The method according to claim 7, before performing feature decoding processing on at least one semantic feature encoding data among the plurality of semantic feature encoding data to obtain at least one semantic feature decoded data, further comprising: Performing semantic segmentation on target semantic feature encoding data among the plurality of semantic feature encoding data to obtain a first semantic segmentation feature, where the first semantic segmentation feature is used to characterize a classification result of pixels in the first image; The iterative feature decoding process is performed on the p-th denoised feature coded data based on the semantic feature decoded data and the semantic feature coded data to obtain second noise data, comprising: Performing semantic segmentation on the p-th denoising feature encoding data to obtain a second semantic segmentation feature, where the second semantic segmentation feature is used to characterize a classification result of pixels in the noisy image; Performing feature splicing on the first semantic segmentation feature and the second semantic segmentation feature to obtain a semantic splicing feature; The semantic concatenation feature is iteratively decoded based on the semantic feature decoding data and the pth denoising feature encoding data to obtain the second noise data.

9. The method according to any one of claims 1 to 8, wherein determining the defect recognition result corresponding to the first image based on the feature similarity between the first image and the second image comprises: Extracting a first image feature representation corresponding to the first image, and extracting a second image feature representation corresponding to the second image; Based on the cosine similarity between the first image feature representation and the second image feature representation, a defect score distribution image is obtained, wherein the pixel values ​​in the defect score distribution image are used to indicate the difference between the pixel values ​​in the first image and the pixel values ​​in the second image; Performing feature upsampling on the defect score distribution image to obtain a defect position distribution image, wherein the defect position distribution image is used to indicate the position distribution of defects in the image object; Performing average pooling processing on the defect position distribution image to obtain a defect category of defects in the image object; The defect position distribution image and the defect category are used as the defect recognition result.

10. The method according to any one of claims 1 to 8, wherein the performing semantic analysis on the first image to obtain semantic analysis data comprises: Inputting the first image into a pre-trained semantic analysis model, and outputting the semantic analysis data; The performing image denoising processing on the noisy image based on the semantic analysis data to obtain a second image includes: The semantic analysis data and the noisy image are input into a pre-trained image denoising model to obtain the second image.

11. The method according to claim 10, before inputting the first image into a pre-trained semantic analysis model to output the semantic analysis data, further comprising: Acquire a sample noisy image corresponding to the first sample image, where the sample noisy image is a result of adding first sample noise data to the first sample image; Inputting the first sample image into a sample analysis model, and outputting sample semantic analysis data; Inputting the sample semantic analysis data and the sample noisy image into a sample denoising model, and outputting first predicted noise data; Based on the difference between the first sample noise data and the first predicted noise data, the sample analysis model and the sample denoising model are trained to obtain the semantic analysis model and the image denoising model.

12. An image processing device, comprising: An acquisition module, configured to acquire a noisy image corresponding to a first image, wherein the noisy image is a result of adding first noise data to the first image, and the first image includes an image object; An analysis module, configured to perform semantic analysis on the first image to obtain semantic analysis data, wherein the semantic analysis data is used to characterize information of the image object in the first image; A denoising module, configured to perform image denoising processing on the noisy image based on the semantic analysis data to obtain a second image; A determination module is used to determine a defect recognition result corresponding to the first image based on feature similarity between the first image and the second image, wherein the defect recognition result includes a defect condition of the image object.

13. A computer device, comprising a processor and a memory, wherein a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the image object recognition method according to any one of claims 1 to 11.

14. A computer-readable storage medium, wherein a computer program is stored in the storage medium, and the computer program is loaded and executed by a processor to implement the image object recognition method according to any one of claims 1 to 11.

15. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the method for recognizing an image object according to any one of claims 1 to 11 is implemented.

Citation Information

Patent Citations

  • Image denoising method and device, equipment and storage medium

    CN112990215A

  • Image generation method and device, equipment and storage medium

    CN117058276A

  • Image processing method and device, equipment, medium and program product

    CN117710295A

  • Method, device, and computer readable storage medium for image processing

    US20220207866A1