Training method and device of fouling detection model, electronic equipment, medium and product
By generating the location of the defilement area, the defilement detection model is trained, which solves the problem of the problem of the time and labor cost of obtaining and labeling defilement banknote samples, and the effect of improving model performance and reducing costs is achieved.
Patent Information
- Application Number
- CN202510095422.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-06-06
AI Technical Summary
In the field of paper currency destruction detection, it is difficult for the prior art to effectively train the destruction detection model because it takes a lot of time and labor costs to obtain and mark the destruction banknote samples.
By acquiring multiple images to be processed, a deficient area generation algorithm (such as Perlin noise) is used to generate the position of the deficient area, and then the corresponding deficient images are synthesized, which are used to train the initial deficient detection model to obtain the target deficient detection model.
This method effectively increases the diversity of training data, reduces dependence on real defile banknote samples, reduces time and labor costs, and improves the performance of the model in practical applications.
Smart Images

Figure CN120107715A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a training method, device, electronic equipment, medium and product for a contamination detection model. Background Art
[0002] Money circulates frequently, and banknotes, as an important medium of exchange, will inevitably be contaminated during the circulation process, including oil stains, ink, and other forms of contamination such as anti-propaganda. The degree of contamination and integrity directly affect the currency circulation. Therefore, it is of great significance to identify contaminated banknotes from multiple banknotes.
[0003] At present, in the field of banknote contamination detection, a banknote contamination detection model can be used to detect contaminated banknotes. However, in training the banknote contamination detection model, a large number of different types of contaminated banknotes are needed as samples. However, compared with normal banknotes, contaminated banknotes are more difficult to obtain and the number is smaller, and it takes a lot of time and manpower costs to mark each contaminated banknote. Summary of the invention
[0004] The embodiments of the present application provide a training method, device, electronic device, medium and product for a contamination detection model, so as to reduce the time and labor costs in the process of acquiring and marking contaminated banknotes.
[0005] In a first aspect, an embodiment of the present application provides a method for training a contamination detection model, comprising:
[0006] Acquire a plurality of images to be processed, wherein the images to be processed are images without contamination;
[0007] For each image to be processed, based on a damaged region generation algorithm, obtaining positions of each damaged region in at least one damaged region in the image to be processed;
[0008] Based on the positions corresponding to the defaced areas, at least one defaced image corresponding to the image to be processed is generated;
[0009] Based on at least one damaged image corresponding to each image to be processed and the position corresponding to each damaged area, the initial damage detection model is trained to obtain a target damage detection model, wherein the damage detection model is used to determine whether the image has a damaged area and the position of the damaged area according to the input image.
[0010] Optionally, the damaged region generation algorithm is Perlin noise, and for each image to be processed, based on the damaged region generation algorithm, obtaining the positions of each damaged region in the image to be processed of at least one damaged region includes:
[0011] For each image to be processed, generating a noise image with the same size as the image to be processed based on Perlin noise;
[0012] The noisy image is binarized to obtain a binary mask; wherein the binary mask is used to indicate the positions of the respective contaminated regions in the image to be processed.
[0013] Optionally, generating at least one defaced image corresponding to the image to be processed based on the positions corresponding to the defaced areas respectively includes:
[0014] Selecting a first texture image from a preset texture image set, and scaling the first texture image to obtain a second texture image, wherein the second texture image has the same size as the image to be processed;
[0015] Using the formula P = M i ⊙E i Determine the image to be superimposed; where M i represents the binary mask, E i represents the second texture image, and P represents the image to be superimposed;
[0016] At least one defaced image is determined according to the image to be superimposed and the image to be processed.
[0017] Optionally, the size of the image to be processed is h*w, and determining at least one defaced image according to the image to be superimposed and the image to be processed includes:
[0018] Determine at least one defaced image according to the following formula: wherein the defaced images correspond one to one with the coefficient β;
[0019]
[0020] Among them, 1 h×w represents a matrix with h rows and w columns, represents a defaced image, O i represents the original image; β represents the coefficient.
[0021] Optionally, the contamination detection model includes: a first contamination detection module and an improved Unet network; the improved Unet network includes n downsampling modules and n upsampling modules, the n downsampling modules correspond to the n upsampling modules one by one, for each upsampling module;
[0022] The up-sampling module is used to calculate the i ] up_new =Concat([F i ] up ,[F i ]down ⊙A i ) determines the target sampling features and outputs the target sampling features to the next module; wherein, [F i ] up is the upsampled feature map output by the previous module, [F i ] down is the down-sampling feature map output by the down-sampling module corresponding to the up-sampling module, where A i =x+ε, where x is the first detection result output by the first contamination detection module; ε is a constant bias matrix having the same resolution as x;
[0023] Accordingly, based on at least one contaminated image corresponding to each image to be processed and the position corresponding to each contaminated area, the initial contamination detection model is trained to obtain a target contamination detection model, including:
[0024] Input the text "contaminated" and at least one contaminated image into the first contaminated detection module to obtain first detection results corresponding to the contaminated images, wherein the first detection results are binary matrices; the number of rows and columns of the binary matrices are the same as those of the image to be processed; and the binary matrix is used to indicate the position of each contaminated area;
[0025] Inputting at least one contaminated image into the improved Unet network to obtain second detection results corresponding to each contaminated image; the second detection results are also used to indicate the position of each contaminated area;
[0026] The total loss value is determined according to the second detection results corresponding to each contaminated image, and the parameters of the improved Unet network are adjusted according to the total loss value to obtain the target contamination detection model.
[0027] Optionally, determining the total loss value according to the second detection results respectively corresponding to the defaced images includes:
[0028] Determine the loss value corresponding to each second test result according to the following formula, add the loss values corresponding to the second test results, and obtain the total loss value;
[0029] L seg =L focal_loss (R i ,M i )
[0030] Among them, R i Represents the second test result, M i Indicates the true area of contamination.
[0031] In a second aspect, an embodiment of the present application provides a training device for a contamination detection model, comprising:
[0032] An acquisition module, used for acquiring a plurality of images to be processed, wherein the images to be processed are images without contamination;
[0033] An obtaining module is used for obtaining, for each image to be processed, positions of each of the at least one damaged area in the image to be processed based on a damaged area generation algorithm;
[0034] A generating module, configured to generate at least one defaced image corresponding to the image to be processed based on the positions corresponding to the defaced areas;
[0035] A training module is used to train an initial contamination detection model based on at least one contaminated image corresponding to each image to be processed and the position corresponding to each contaminated area, so as to obtain a target contamination detection model, wherein the contamination detection model is used to determine whether the image has a contaminated area and the position of the contaminated area according to the input image.
[0036] In a third aspect, an embodiment of the present application provides an electronic device, including: a memory, a processor;
[0037] The memory stores computer-executable instructions;
[0038] The processor executes the computer-executable instructions stored in the memory, so that the processor executes the above first aspect and / or various possible implementations of the first aspect.
[0039] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the first aspect above and / or various possible implementations of the first aspect.
[0040] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the above first aspect and / or various possible implementation methods of the first aspect.
[0041] The training method, device, electronic device, medium and product of the contamination detection model provided in the embodiments of the present application are obtained by acquiring multiple images to be processed, wherein the images to be processed are images without contamination; for each image to be processed, based on a contaminated area generation algorithm, the positions of each contaminated area in at least one contaminated area in the image to be processed are obtained; based on the positions corresponding to each contaminated area, at least one contaminated image corresponding to the image to be processed is generated; based on the at least one contaminated image corresponding to each image to be processed and the positions corresponding to each contaminated area, an initial contamination detection model is trained to obtain a target contamination detection model, wherein the contamination detection model is used to determine whether the image has a contaminated area and the position of the contaminated area according to the input image, so that the contaminated image synthesized based on the contaminated area generation algorithm can effectively increase the diversity of training data, make up for the deficiency that it is difficult to obtain real contaminated banknote samples, and improve the performance of the model in subsequent practical applications, and in the process of generating each contaminated image, the position of the contaminated area has been accurately determined, reducing the dependence on manual labeling, thereby saving time and labor costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0043] Figure 1 An application scenario diagram provided for an embodiment of the present application;
[0044] Figure 2 A schematic diagram of a flow chart of a method for training a contamination detection model provided in an embodiment of the present application;
[0045] Figure 3 A schematic diagram of the structure of an improved Unet network provided in an embodiment of the present application;
[0046] Figure 4 A system architecture diagram of a method for training a contamination detection model provided in an embodiment of the present application;
[0047] Figure 5 A schematic diagram of the structure of a training device for a contamination detection model provided in this application;
[0048] Figure 6 A schematic diagram of the structure of the electronic device provided in this application.
[0049] The above drawings have shown clear embodiments of the present application, which will be described in more detail later. These drawings and text descriptions are not intended to limit the scope of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION
[0050] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0051] Money circulates frequently, and banknotes, as an important medium of exchange, are inevitably contaminated during the circulation process, including oil stains, ink, and other forms of contamination such as anti-propaganda. The degree of contamination and integrity of banknotes directly affect their circulation and service life in the market. Therefore, it is of great significance to identify contaminated banknotes from multiple banknotes, which not only helps to maintain the circulation quality of currency, but also improves the security and efficiency of financial transactions.
[0052] In the field of banknote contamination detection, advanced banknote contamination detection models can currently be used to automatically detect contaminated banknotes. These models are usually based on machine learning and computer vision technology, and can quickly and accurately identify the contaminated areas on banknotes and assess their severity. However, to effectively train banknote contamination detection models, a large number of different types of contaminated banknotes are required as samples. This is because the accuracy and robustness of the model largely rely on the diversity and richness of the training data.
[0053] However, compared with normal banknotes, soiled banknotes are more difficult to obtain and are less in number, which poses a challenge to sample collection. In addition, it takes a lot of time and manpower to make detailed annotations for each soiled banknote, because the annotation process usually requires manual precise identification and classification of the soiled area. This high cost and time investment limits the construction of large-scale datasets, thus affecting the training effect of the model.
[0054] In view of this, the present application provides a training method for a contamination detection model, which can obtain multiple images to be processed. For each image to be processed, based on a contaminated area generation algorithm, the positions of each contaminated area in at least one contaminated area in the image to be processed are obtained; based on the positions corresponding to each contaminated area, at least one contaminated image corresponding to the image to be processed is generated; based on at least one contaminated image corresponding to each image to be processed and the positions corresponding to each contaminated area, an initial contamination detection model is trained to obtain a target contamination detection model. In this way, the contaminated images synthesized based on the contaminated area generation algorithm can effectively increase the diversity of training data, make up for the deficiency that real contaminated banknote samples are difficult to obtain, and improve the performance of the model in subsequent practical applications. In the process of generating each contaminated image, the position of the contaminated area has been accurately determined, reducing the dependence on manual labeling, thereby saving time and labor costs.
[0055] Figure 1 An application scenario diagram provided for an embodiment of the present application, in which a client sends multiple images to be processed to a server. After the server obtains the multiple images to be processed, for each image to be processed, based on a damaged area generation algorithm, obtains the position of each damaged area in the image to be processed in at least one damaged area; based on the positions corresponding to each damaged area, generates at least one damaged image corresponding to the image to be processed; based on at least one damaged image corresponding to each image to be processed and the positions corresponding to each damaged area, trains an initial damage detection model to obtain a target damage detection model.
[0056] The client sends multiple images to be detected to the server. After receiving the multiple images to be detected, the server inputs the multiple images to be detected into the target contamination detection model to obtain the detection results corresponding to each image to be detected.
[0057] The technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems are described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.
[0058] Figure 2 A flowchart of a method for training a contamination detection model provided in an embodiment of the present application is provided. The execution subject of the present embodiment may be any device having a data processing function. The present application is specifically described with the server as the execution subject. Figure 2 As shown, a method for training a contamination detection model provided in an embodiment of the present application may include:
[0059] Step 201: Acquire multiple images to be processed, wherein the images to be processed are images without contamination.
[0060] Specifically, the user sends a plurality of images to be processed to the server through the client. The images to be processed are obtained by photographing real banknotes without soiling. The images to be processed are images of banknotes without soiling or being soiled.
[0061] The server obtains multiple images to be processed sent by the client.
[0062] Step 202: for each image to be processed, based on a damaged region generation algorithm, obtain the position of each damaged region in at least one damaged region in the image to be processed.
[0063] The defaced region generating algorithm may be any algorithm that can generate a defaced region, which is not limited in the present application. Exemplarily, it may be Perlin noise.
[0064] Specifically, for each image to be processed, the server may generate one or more damaged regions based on a damaged region generation algorithm, and determine the position of each damaged region in the image to be processed.
[0065] Optionally, the damaged region generation algorithm is Perlin noise, and for each image to be processed, based on the damaged region generation algorithm, obtaining the positions of each damaged region in the image to be processed of at least one damaged region includes:
[0066] For each image to be processed, generating a noise image with the same size as the image to be processed based on Perlin noise;
[0067] The noisy image is binarized to obtain a binary mask; wherein the binary mask is used to indicate the positions of the respective contaminated regions in the image to be processed.
[0068] Specifically, for each image to be processed, the size of the image to be processed is input into Perlin noise to obtain a noise image with the same size as the image to be processed, and the obtained noise image is binarized to obtain a binary mask corresponding to the noise image.
[0069] The specific process of binarization is as follows: first, a threshold is selected. For each pixel in the noise image, if the pixel value of the pixel is greater than the threshold, the pixel value of the pixel is updated to 1; if the pixel value of the pixel is less than the threshold, the pixel value of the pixel is updated to 0. The image obtained after updating the pixel value of each pixel in the noise image is a binary mask. The pixel value of the pixel in the binary mask can only be 0 or 1. The binary mask can also be called a "0-1" mask.
[0070] In the binary mask, the location of the pixel with a pixel value of 1 is the location of the pollution, so the binary mask can indicate the location of each pollution area in the image to be processed.
[0071] In this way, the texture generated by Perlin noise has a natural smooth transition characteristic, which makes the generated damaged area look more realistic, similar to the common damage or wear patterns in nature, and the binarized mask can be flexibly applied to the image to be processed. By determining the position of the pixel with a pixel value of 1, the position of the damaged area can be quickly determined, which is simple and efficient.
[0072] Step 203: Generate at least one damaged image corresponding to the image to be processed based on the positions corresponding to the damaged areas.
[0073] Specifically, the server generates at least one damaged image corresponding to the image to be processed based on the positions corresponding to the damaged areas.
[0074] Optionally, generating at least one defaced image corresponding to the image to be processed based on the positions corresponding to the defaced areas respectively includes:
[0075] Selecting a first texture image from a preset texture image set, and scaling the first texture image to obtain a second texture image, wherein the second texture image has the same size as the image to be processed;
[0076] Using the formula P = M i ⊙E i Determine the image to be superimposed; where M i represents the binary mask, E i represents the second texture image, and P represents the image to be superimposed;
[0077] At least one defaced image is determined according to the image to be superimposed and the image to be processed.
[0078] Among them, the preset texture image set includes multiple texture images. The server randomly selects a first texture image from the preset texture image set, and scales the selected first texture image to obtain a second texture image. The second texture image has the same size as the image to be processed.
[0079] The server uses the formula P = M i ⊙E i Determine the image to be superimposed; where M i represents a binary mask, E i represents the second texture image, and P represents the image to be superimposed.
[0080] Where P = M i ⊙Ei The meaning of the formula is: for any target pixel in the image to be superimposed, the pixel value of the pixel is equal to the product of the pixel value of the first pixel in the binary mask and the pixel value of the second pixel in the second texture image, wherein the relative position of the target pixel in the image to be superimposed is the same as the relative position of the first pixel in the binary mask and the relative position of the second pixel in the second texture image.
[0081] Optionally, using the formula P = M i ⊙E i Before determining the image to be superimposed, the second texture image can be subjected to random data enhancement operation to obtain an enhanced second texture image, and then according to the formula P = M i ⊙F i Determine the image to be superimposed, where F i Represents the enhanced second texture image.
[0082] Among them, random data augmentation operation is a technology used in training machine learning models (especially deep learning models), which aims to increase the diversity of data sets by randomly transforming training data, thereby improving the generalization ability and robustness of the model. Exemplary random data augmentation operations can include: color jitter, random rotation, Gaussian blur, etc.
[0083] In this way, by selecting the first texture image from the preset texture image set and scaling it to match the size of the image to be processed, the method can flexibly adapt to images of different sizes and resolutions, and use the formula P = M i ⊙E i By determining the image to be superimposed, the details of the texture image can be effectively superimposed on a specific area of the image to be processed. In this way, the generated defaced image can more realistically simulate the actual defacement effect, thereby improving the authenticity and visual effect of image processing.
[0084] Optionally, the size of the image to be processed is h*w, and determining at least one defaced image according to the image to be superimposed and the image to be processed includes:
[0085] Determine at least one defaced image according to the following formula: wherein the defaced images correspond one to one with the coefficient β;
[0086]
[0087] Among them, 1 h×w represents a matrix with h rows and w columns, represents a defaced image, O i represents the original image; β represents the coefficient.
[0088] The formula is explained below. The meaning of , where the sizes of the two images on which the symbol “⊙” is operated are the same, that is, the number of pixels contained in each row and each column is the same, and the symbol “⊙” means multiplying the pixel values of two pixels located at the same position in the two images.
[0089] M i ⊙O i The pixel value of the pixel in the damaged area is the same as the pixel value of the pixel at the corresponding position in the image to be processed, and the pixel value of the pixel not in the damaged area is 0.
[0090] (1-β)(M i ⊙O i )+βP means that the pixel value of the pixel in the damaged area is equal to the pixel value of the pixel at the corresponding position in the image to be processed multiplied by (1-β) plus the pixel value of the pixel at the corresponding position in the image to be superimposed multiplied by β, and the pixel value of the pixel not in the damaged area is 0.
[0091] (1 h×w -M i )⊙O i The pixel value of the pixel not in the stained area is the same as the pixel value of the pixel at the corresponding position in the image to be processed, and the pixel value of the pixel in the stained area is 0.
[0092] therefore It means that the pixel value of the pixel not in the damaged area is the same as the pixel value of the pixel at the corresponding position in the image to be processed, and the pixel value of the pixel in the damaged area is equal to the pixel value of the pixel at the corresponding position in the image to be processed multiplied by (1-β) plus the pixel value of the pixel at the corresponding position in the image to be superimposed multiplied by β.
[0093] In this way, multiple defaced images can be obtained for each image to be processed, and the defaced images obtained also have pixel-level labels of the defaced areas, eliminating the tedious manual labeling and saving a lot of manpower and time costs. In addition, the multiple defaced images obtained in this way also have the advantages of diversity, authenticity, and accessibility.
[0094] Step 204: Based on at least one damaged image corresponding to each image to be processed and the position corresponding to each damaged area, an initial damage detection model is trained to obtain a target damage detection model, wherein the damage detection model is used to determine whether the image has a damaged area and the position of the damaged area according to the input image.
[0095] Among them, the damage detection module can be any model that determines whether the image has a damaged area and the location of the damaged area based on the input image. This application does not limit this. Exemplarily, it can be a SAM (Segment Anything Model) or a Unet (U-Net Convolutional Network) model.
[0096] Exemplarily, based on at least one defaced image corresponding to each image to be processed and the position corresponding to each defaced area, the initial SAM model is trained to obtain a target SAM model.
[0097] The training method of the contamination detection model provided in the embodiment of the present application is by acquiring multiple images to be processed, wherein the images to be processed are images without contamination; for each image to be processed, based on a contaminated area generation algorithm, the position of each contaminated area in at least one contaminated area in the image to be processed is obtained; based on the positions corresponding to each contaminated area, at least one contaminated image corresponding to the image to be processed is generated; based on the at least one contaminated image corresponding to each image to be processed and the positions corresponding to each contaminated area, an initial contamination detection model is trained to obtain a target contamination detection model, wherein the contamination detection model is used to determine whether the image has a contaminated area and the position of the contaminated area according to the input image, so that the contaminated image synthesized based on the contaminated area generation algorithm can effectively increase the diversity of training data, make up for the deficiency that it is difficult to obtain real contaminated banknote samples, and improve the performance of the model in subsequent practical applications, and in the process of generating each contaminated image, the position of the contaminated area has been accurately determined, reducing the dependence on manual labeling, thereby saving time and labor costs.
[0098] Optionally, the contamination detection model includes: a first contamination detection module and an improved Unet network; the improved Unet network includes n downsampling modules and n upsampling modules, the n downsampling modules correspond to the n upsampling modules one by one, for each upsampling module;
[0099] The up-sampling module is used to calculate the i ] up_new =Concat([F i ] up ,[F i ] down ⊙A i ) determines the target sampling features and outputs the target sampling features to the next module; wherein, [F i ] up is the upsampled feature map output by the previous module, [F i ]down is the down-sampling feature map output by the down-sampling module corresponding to the up-sampling module, where A i =x+ε, where x is the first detection result output by the first contamination detection module; ε is a constant bias matrix having the same resolution as x;
[0100] Accordingly, based on at least one contaminated image corresponding to each image to be processed and the position corresponding to each contaminated area, the initial contamination detection model is trained to obtain a target contamination detection model, including:
[0101] Input the text "contaminated" and at least one contaminated image into the first contaminated detection module to obtain first detection results corresponding to the contaminated images, wherein the first detection results are binary matrices; the number of rows and columns of the binary matrices are the same as those of the image to be processed; and the binary matrix is used to indicate the position of each contaminated area;
[0102] Inputting at least one contaminated image into the improved Unet network to obtain second detection results corresponding to each contaminated image; the second detection results are also used to indicate the position of each contaminated area;
[0103] The total loss value is determined according to the second detection results corresponding to each contaminated image, and the parameters of the improved Unet network are adjusted according to the total loss value to obtain the target contamination detection model.
[0104] Among them, Concat means channel connection, Concat([F i ] up ,[F i ] down ⊙A i ) is to convert [F i ] up With [F i ] down ⊙A i Make channel connections.
[0105] Figure 3 A schematic diagram of the structure of an improved Unet network provided in an embodiment of the present application is shown in FIG. Figure 3 As shown in the figure, the improved Unet network includes 4 downsampling modules and 4 upsampling modules. The green blocks in the figure are downsampling modules, and the blocks combined with green and pink are upsampling modules. The 4 upsampling modules and 4 downsampling modules are symmetrical about the gray blocks in the middle. Each upsampling module corresponds to a downsampling module symmetrical about the gray blocks. The downsampling module is used to downsample the input feature image and output the downsampled feature image to the next module. For each upsampling module, the upsampling module is used to calculate the feature image according to the formula [F i ] up_new=Concat([F i ] up ,[F i ] down ⊙A i ) determines the target sampling features and outputs the target sampling features to the next module, wherein [F i ] up is the upsampled feature map output by the previous module, [F i ] down is the down-sampling feature map output by the down-sampling module corresponding to the up-sampling module, where A i =x+ε, where x is the first detection result output by the first contamination detection module; ε is a constant offset matrix having the same resolution as x, and the element values of all elements in the constant offset matrix are the same.
[0106] Specifically, the first contamination detection module can be a SAM model. For example, the server inputs the text "contaminated" and at least one contaminated image into the first contamination detection module to obtain the first detection results corresponding to each contaminated image. The first detection result is a binary matrix. The element value of the element in the binary matrix can only be 0 or 1. The image to be processed contains multiple pixels. The number of rows and columns of the image to be processed refers to the number of rows and columns of the pixels contained. The number of rows and columns of the binary matrix are the same as the image to be processed. The position of the element with the element value of 1 in the binary matrix is the position of the contamination, so the binary matrix can indicate the position of each contaminated area.
[0107] The server inputs at least one defaced image into the improved Unet network, obtains the second detection results corresponding to each defaced image, determines the loss value according to the second detection results corresponding to each defaced image, adjusts the parameters of the improved Unet network according to the loss value, and obtains the target defacement detection model.
[0108] In this way, the first detection result output by the first corruption detection module is introduced into the coding feature map of each jump layer connection. The introduction of the first detection result can further improve the accuracy of the subsequent second detection result, and adding a constant bias matrix can avoid the existence of 0 to retain the low-level auxiliary information provided by the jump layer.
[0109] Optionally, determining the total loss value according to the second detection results respectively corresponding to the defaced images includes:
[0110] Determine the loss value corresponding to each second test result according to the following formula, add the loss values corresponding to the second test results, and obtain the total loss value;
[0111] L seg =L focal_loss (R i ,M i )
[0112] Among them, R i Represents the second test result, M i Indicates the true area of contamination.
[0113] Among them, Focal Loss is a loss function used to deal with the problem of class imbalance. It was originally proposed by researchers in the object detection task. The traditional cross entropy loss may cause the model to pay too much attention to the majority class and ignore the minority class when dealing with class imbalance. Focal Loss helps the model better learn the minority class by adjusting the form of the loss function.
[0114] Thus, one of the main advantages of Focal Loss is its effectiveness in dealing with class imbalance problems. In a defaced image, the defaced area may only occupy a small part of the image, while the background or non-defaced area occupies the majority. Focal Loss helps the model learn these minority classes better by increasing attention to difficult-to-classify samples such as defaced areas.
[0115] Figure 4 A system architecture diagram of a method for training a contamination detection model provided in an embodiment of the present application, such as Figure 4 As shown in the figure, it is mainly divided into three parts: sample generation module, SAM-based attention generation module, and attention-guided contamination detection module. Since contaminated samples are difficult to obtain and the contamination patterns are diverse and unpredictable, the scheme first designs a contaminated sample generation module to artificially synthesize a large number of diverse contaminated banknotes. Then, the artificially synthesized contaminated banknotes are input into the SAM segmentation network for contamination segmentation, and the segmented results are used as spatial attention to provide pixel-level coarse-grained position information for the Unet segmentation network, thereby assisting it in obtaining more refined contamination detection results.
[0116] The specific implementation is as follows:
[0117] 1. Generation of contaminated samples
[0118] Since it is difficult to obtain damaged banknotes and the styles of damage are diverse, it is difficult to collect enough damaged samples and label them. For this reason, this patent first artificially synthesizes abnormal images with pixel-level labels on normal images. By artificially creating enough labeled abnormal images and training the model using the paradigm of supervised learning, the model can have a better understanding of the image context.
[0119] The forms of anomalies are varied and difficult to predict, but we can uniformly regard them as local appearances that are different from normal areas. We introduce out-of-distribution patterns as the source of anomalies based on normal images. The method for synthesizing defaced samples is as follows:
[0120] 1. For each normal image Generate a '0-1' mask using binarized perlin noise As defaced outlines.
[0121] 2. Secondly, use M i Importing texture images from external datasets As the defiled area. In order to improve the richness of the data, we first introduce E i Random data enhancement including color jitter, random rotation, Gaussian blur, etc. is used, and the defacement introduced is:
[0122] P=M i ⊙Rand(E i )
[0123] 3. Finally, the stained image P and the normal image O are combined by a coefficient β i The corresponding areas in the graph are superimposed to obtain the synthetic contaminated sample. The formula is as follows:
[0124]
[0125] In this way, K corresponding images can be generated for each normal image to intentionally collect a large number of synthetic samples with pixel-level labels. More importantly, these synthetic samples also show the advantages of diversity and easy acquisition.
[0126] 2. SAM-based attention image generation
[0127] SAM is an interactive image segmentation model that interacts with users by combining the attention mechanism, allowing users to actively participate in the image segmentation process and provide more guiding information. At the same time, SAM is trained with the help of a very large-scale dataset, enabling it to achieve zero-sample transfer and achieve accurate and detailed results in downstream tasks in a variety of scenarios.
[0128] Since SAM is an interactive segmentation model, it is necessary to input a prompt before segmentation to assist the model in generating a prediction mask of the desired segmentation area. The prompts include text, points, or boxes, etc. These prompts have a direct and important effect on the segmentation results. Therefore, after we input the "defaced" text prompt into the SAM segmentation network, we can directly use the provided pre-trained weights to obtain the banknote image. The contamination detection result S i , without any training process. However, this result may not be accurate, so we do not use it as the final detection result, but use it as spatial attention to provide the segmentation network with pixel-level coarse-grained position information to obtain more precise defect localization.
[0129] 3. Attention-guided Unet banknote damage detection
[0130] Unet is an efficient image segmentation network, whose architecture consists of two parts: encoder and decoder. The encoder extracts abstract features of the image through a series of convolutional layers, pooling layers, and nonlinear activation functions. In the decoder part, the network restores the resolution of the image through upsampling or a combination of transposed convolution and convolutional layers. In each upsampling, Unet concatenates the feature map with the feature map of the corresponding layer of the encoder, which helps the network to restore details while maintaining position information. The concatenated feature map is further processed by the convolutional layer to generate the final segmentation mask.
[0131] In order to make the segmentation network pay more attention to the target area rather than the interfering background, we introduce spatial attention A in the encoding feature map of each skip layer connection i , where A i =g(S i ), and g(x) = x + ε, 0, to avoid the existence of 0 to retain the low-level auxiliary information provided by the skip layer. More specifically, we use A i After optimizing the down-sampled features, they are skipped and connected to the corresponding feature maps of the decoder. The new up-sampled features for:
[0132]
[0133] in, represents the features of the image layer m extracted by the decoder through the segmentation network, Represents the features of the image layer m extracted by the encoder through the segmentation network, It is through A i The attention of the correct resolution scale obtained by downsampling, where m∈{1,…,C}, C is the number of segmentation network layers. Concat(*,*) represents channel connection.
[0134] Since the size of the stained area is often smaller than the normal area, this imbalance will inevitably lead to deviations in the model training process, so we use focal loss to optimize the Unet model to solve this problem. The Unet segmentation loss is as follows:
[0135] L seg =L focal_loss (R i ,M i )
[0136] Among them, R i is the output result of the Unet segmentation network, M i is the generated defacement mask.
[0137] 4. Beneficial effects brought by the technical solution of the present invention
[0138] The solution does not need to collect and label any damaged samples. Instead, it uses artificial synthesis to obtain a large number of diverse damaged images for model training, which is simple and efficient. Secondly, the solution introduces cross-modal text data through the SAM model to generate spatial attention images, and introduces them into the normal Unet segmentation model so that the segmentation model can pay more attention to the target area rather than the interfering background, and obtain more accurate detection results. The solution can automatically perform efficient and accurate banknote damage detection, reducing manpower investment and costs.
[0139] The embodiment of the present application also provides a contamination detection method, comprising: acquiring at least one image to be detected;
[0140] At least one image to be detected is input into the target contamination detection model to predict whether each image to be detected has contamination and the location of the contaminated area, wherein the target contamination detection model is the target contamination detection model in any of the above embodiments.
[0141] Corresponding to the above-mentioned training method of the contamination detection model, the embodiment of the present application further provides a training device for the contamination detection model. Figure 5 A schematic diagram of the structure of the training device for the contamination detection model provided in this application, such as Figure 5 As shown, the training device of the contamination detection model provided in this embodiment includes:
[0142] An acquisition module 501 is used to acquire a plurality of images to be processed, wherein the images to be processed are images without contamination;
[0143] An obtaining module 502 is used for obtaining, for each image to be processed, positions of each of the at least one damaged area in the image to be processed based on a damaged area generation algorithm;
[0144] A generating module 503, configured to generate at least one defaced image corresponding to the image to be processed based on the positions corresponding to the defaced areas;
[0145] The training module 504 is used to train the initial contamination detection model based on at least one contaminated image corresponding to each image to be processed and the position corresponding to each contaminated area, so as to obtain a target contamination detection model, wherein the contamination detection model is used to determine whether the image has a contaminated area and the position of the contaminated area according to the input image.
[0146] Optionally, the defaced region generation algorithm is Perlin noise, and module 502 is obtained, which is specifically used for:
[0147] For each image to be processed, generating a noise image with the same size as the image to be processed based on Perlin noise;
[0148] The noisy image is binarized to obtain a binary mask; wherein the binary mask is used to indicate the positions of the respective contaminated regions in the image to be processed.
[0149] Optionally, the generating module 503 is specifically used for:
[0150] Selecting a first texture image from a preset texture image set, and scaling the first texture image to obtain a second texture image, wherein the second texture image has the same size as the image to be processed;
[0151] Using the formula P = M i ⊙E i Determine the image to be superimposed; where M i represents the binary mask, E i represents the second texture image, and P represents the image to be superimposed;
[0152] At least one defaced image is determined according to the image to be superimposed and the image to be processed.
[0153] Optionally, the size of the image to be processed is h*w, and when the generating module 503 determines at least one defaced image according to the image to be superimposed and the image to be processed, it is specifically configured to:
[0154] Determine at least one defaced image according to the following formula: wherein the defaced images correspond one to one with the coefficient β;
[0155]
[0156] Among them, 1 h×w represents a matrix with h rows and w columns, represents a defaced image, O i represents the original image; β represents the coefficient.
[0157] The optional contamination detection model includes: a first contamination detection module and an improved Unet network; the improved Unet network includes n downsampling modules and n upsampling modules, the n downsampling modules correspond to the n upsampling modules one by one, for each upsampling module;
[0158] The up-sampling module is used to calculate the i ] up_new =Concat([F i ] up ,[F i ] down ⊙Ai ) determines the target sampling features and outputs the target sampling features to the next module; wherein, [F i ] up is the upsampled feature map output by the previous module, [F i ] down is the down-sampling feature map output by the down-sampling module corresponding to the up-sampling module, where A i =x+ε, where x is the first detection result output by the first contamination detection module; ε is a constant bias matrix having the same resolution as x;
[0159] Accordingly, the training module 504 is specifically used for:
[0160] Input the text "contaminated" and at least one contaminated image into the first contaminated detection module to obtain first detection results corresponding to the contaminated images, wherein the first detection results are binary matrices; the number of rows and columns of the binary matrices are the same as those of the image to be processed; and the binary matrix is used to indicate the position of each contaminated area;
[0161] Inputting at least one contaminated image into the improved Unet network to obtain second detection results corresponding to each contaminated image; the second detection results are also used to indicate the position of each contaminated area;
[0162] The total loss value is determined according to the second detection results corresponding to each contaminated image, and the parameters of the improved Unet network are adjusted according to the total loss value to obtain the target contamination detection model.
[0163] Optionally, when determining the total loss value according to the second detection results respectively corresponding to the defaced images, the training module 504 is specifically configured to:
[0164] Determine the loss value corresponding to each second test result according to the following formula, add the loss values corresponding to the second test results, and obtain the total loss value;
[0165] L seg =L focal_loss (R i ,M i )
[0166] Among them, R i Represents the second test result, M i Indicates the true area of contamination.
[0167] The training device for the contamination detection model provided in this embodiment can execute the method provided in the above method embodiment. Its implementation principle and technical effect are similar, and will not be described in detail in this embodiment.
[0168] Figure 6 This is a schematic diagram of the structure of the electronic device provided in this application. Figure 6 As shown, the electronic device 60 provided in this embodiment includes: at least one processor 601 and a memory 602. Optionally, the device 60 also includes a communication component 603. The processor 601, the memory 602 and the communication component 603 are connected via a bus 604.
[0169] In a specific implementation process, at least one processor 601 executes the computer execution instructions stored in the memory 602, so that at least one processor 601 executes the above method.
[0170] The specific implementation process of the processor 601 can be found in the above method embodiment, and its implementation principle and technical effect are similar, so this embodiment will not be repeated here.
[0171] In the above embodiments, it should be understood that the processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the invention may be directly implemented as being executed by a hardware processor, or may be executed by a combination of hardware and software modules in the processor.
[0172] The memory may include a high-speed memory (Random Access Memory, RAM), and may also include a non-volatile memory (Non-volatile Memory, NVM), such as at least one disk memory.
[0173] The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the bus in the drawings of the present application is not limited to only one bus or one type of bus.
[0174] The present application also provides a computer program product, including a computer program, which implements the above method when executed by a processor.
[0175] The present application also provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the above method is implemented.
[0176] The above-mentioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The readable storage medium can be any available medium that can be accessed by a general or special-purpose computer.
[0177] An exemplary readable storage medium is coupled to a processor so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (Application Specific Integrated Circuits, referred to as: ASIC). Of course, the processor and the readable storage medium can also exist in the device as discrete components.
[0178] The division of units is only a logical function division, and there may be other divisions in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other forms.
[0179] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0180] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0181] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods of each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0182] Those skilled in the art can understand that all or part of the steps of implementing the above-mentioned method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, the steps of the above-mentioned method embodiments are executed; and the aforementioned storage medium includes: ROM, RAM, disk or optical disk and other media that can store program codes.
[0183] Finally, it should be noted that those skilled in the art will readily conceive of other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. The present invention is intended to cover any variations, uses or adaptations of the present invention, which follow the general principles of the present invention and include common knowledge or customary technical means in the art not disclosed by the present invention, are not limited to the precise structure described above and shown in the drawings, and may be modified and changed in various ways without departing from the scope thereof. The scope of the present invention is limited only by the appended claims.
Claims
1. A method for training a contamination detection model, characterized in that: include: Acquire a plurality of images to be processed, wherein the images to be processed are images without contamination; For each image to be processed, based on a damaged region generation algorithm, obtaining positions of each damaged region in at least one damaged region in the image to be processed; Based on the positions corresponding to the defaced areas, at least one defaced image corresponding to the image to be processed is generated; Based on at least one damaged image corresponding to each image to be processed and the position corresponding to each damaged area, the initial damage detection model is trained to obtain a target damage detection model, wherein the damage detection model is used to determine whether the image has a damaged area and the position of the damaged area according to the input image.
2. The method according to claim 1, characterized in that The damaged area generation algorithm is Perlin noise. For each image to be processed, based on the damaged area generation algorithm, the positions of each damaged area in at least one damaged area in the image to be processed are obtained, including: For each image to be processed, generating a noise image with the same size as the image to be processed based on Perlin noise; The noisy image is binarized to obtain a binary mask; wherein the binary mask is used to indicate the positions of the respective contaminated regions in the image to be processed.
3. The method according to claim 2, characterized in that Generating at least one defaced image corresponding to the image to be processed based on the positions corresponding to the defaced areas, including: Selecting a first texture image from a preset texture image set, and scaling the first texture image to obtain a second texture image, wherein the second texture image has the same size as the image to be processed; Using the formula P = M i ⊙E i Determine the image to be superimposed; where M i represents the binary mask, E i represents the second texture image, and P represents the image to be superimposed; At least one defaced image is determined according to the image to be superimposed and the image to be processed.
4. The method according to claim 3, characterized in that The size of the image to be processed is h*w, and at least one defaced image is determined according to the image to be superimposed and the image to be processed, including: Determine at least one defaced image according to the following formula: wherein the defaced images correspond one to one with the coefficient β; Among them, 1 h×w represents a matrix with h rows and w columns, represents a defaced image, O i represents the original image; β represents the coefficient.
5. The method according to any one of claims 1 to 4, characterized in that: The contamination detection model includes: a first contamination detection module and an improved Unet network; the improved Unet network includes n downsampling modules and n upsampling modules, the n downsampling modules correspond to the n upsampling modules one by one, for each upsampling module; The up-sampling module is used to calculate the i ] up_new =Concat([F i ] up ,[F i ] down ⊙A i ) determines the target sampling features and outputs the target sampling features to the next module; wherein, [F i ] up is the upsampled feature map output by the previous module, [F i ] down is the down-sampling feature map output by the down-sampling module corresponding to the up-sampling module, where A i =x+ε, where x is the first detection result output by the first contamination detection module; ε is a constant bias matrix having the same resolution as x; Accordingly, based on at least one contaminated image corresponding to each image to be processed and the position corresponding to each contaminated area, the initial contamination detection model is trained to obtain a target contamination detection model, including: Input the text "Damaged" and at least one damaged image into the first damage detection module to obtain first detection results corresponding to the damaged images, wherein the first detection results are binary matrices; the number of rows and columns of the binary matrices are the same as those of the image to be processed; and the binary matrix is used to indicate the position of each damaged area; Inputting at least one contaminated image into the improved Unet network to obtain second detection results corresponding to each contaminated image; the second detection results are also used to indicate the position of each contaminated area; The total loss value is determined according to the second detection results corresponding to each contaminated image, and the parameters of the improved Unet network are adjusted according to the total loss value to obtain the target contamination detection model.
6. The method according to claim 5, characterized in that Determining the total loss value according to the second detection results respectively corresponding to the defiled images includes: Determine the loss value corresponding to each second test result according to the following formula, add the loss values corresponding to the second test results, and obtain the total loss value; L seg =L focal_loss (R i ,M i ) Among them, R i Represents the second test result, M i Indicates the true area of contamination.
7. A training device for a contamination detection model, characterized in that: include: An acquisition module, used for acquiring a plurality of images to be processed, wherein the images to be processed are images without contamination; An obtaining module is used for obtaining, for each image to be processed, positions of each of the at least one damaged area in the image to be processed based on a damaged area generation algorithm; A generating module, configured to generate at least one defaced image corresponding to the image to be processed based on the positions corresponding to the defaced areas; A training module is used to train an initial contamination detection model based on at least one contaminated image corresponding to each image to be processed and the position corresponding to each contaminated area, so as to obtain a target contamination detection model, wherein the contamination detection model is used to determine whether the image has a contaminated area and the position of the contaminated area according to the input image.
8. An electronic device, characterized in that: include: Memory, processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor performs the method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 6 when executed by a processor.
10. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 6 when being executed by a processor.