A method for generating a benchmark dataset for microscopic image recognition of particulate pollutants
By generating a benchmark dataset for microscopic image recognition of particulate contaminants, the problem of being unable to evaluate the accuracy of automatic cleanliness detectors was solved, and performance evaluation and accuracy analysis of the detection method were achieved.
Patent Information
- Application Number
- CN202211589632.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-04-29
- Filing Date
- 2022-12-12
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2042-12-12
AI Technical Summary
The accuracy of existing automatic cleanliness detectors cannot be evaluated, and there is a lack of public benchmark datasets for evaluating the accuracy of particulate contaminant microscopic image recognition methods.
Generate a benchmark dataset for microscopic image recognition of particulate pollutants. By collecting, processing and annotating microscopic images, performing image reduction, rotation, particle reduction and lighting compensation, counting the number and type of particles, using the DeepLabV3+ network for model training and testing, and calculating accuracy indicators.
It provides detailed information on particle pollutants, which can evaluate the performance of particle pollutant detection and identification methods and provide a reference for the accuracy evaluation of the detector.
Smart Images

Figure CN116416614B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an evaluation method for particulate pollutant microscopic image recognition, and in particular to a benchmark data set generation method for particulate pollutant microscopic image recognition. Background Art
[0002] With the continuous improvement of social and economic levels, the number of automobiles has increased annually to meet people's travel needs. According to statistics from the Ministry of Public Security, by the end of 2021, the national vehicle ownership reached 302 million, a 7.5% increase from 2019. At the same time, user quality requirements for automobiles are constantly increasing. Cleanliness is a key indicator of automotive transmission assemblies and components, directly affecting the performance and lifespan of transmission assemblies. Component cleanliness testing is one of the important means to improve product reliability. Therefore, controlling and monitoring component cleanliness has become an essential step in the automotive production process. The identification of particulate contaminants helps improve the measurement accuracy of system component cleanliness, promoting the development of higher-reliability systems.
[0003] For the inspection of the cleanliness of particulate contaminants, the international standard ISO 16232:2018 provides two standard analysis methods: optical analysis method and gravimetric analysis method. In recent years, with the improvement of particle filtration technology, particles can be separated and not overlapped on the filter membrane, making optical analysis method a common method for cleanliness analysis. Currently, the automatic cleanliness detectors in industry mainly use optical analysis methods, which usually consist of two parts: microscopic image acquisition and microscopic image analysis. However, the accuracy of the automatic cleanliness detectors on the market is unknown. The main reason is that there is a lack of a public, inability to evaluate the accuracy of the proposed identification method. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a method for generating a benchmark data set for particulate pollutant microscopic image recognition that can evaluate the accuracy of the particulate pollutant microscopic image recognition method. The benchmark data set contains detailed information on the particulate pollutants and can provide a reference basis for evaluating the performance of the particulate pollutant detection and identification method.
[0005] The technical solution adopted by the present invention to solve the above technical problems is: a method for generating a benchmark data set for microscopic image recognition of particulate pollutants, characterized in that the method comprises:
[0006] Step S1) collecting a sequence of microscopic images of particulate pollutants to generate an image basic data set;
[0007] Step S2) performing image reduction on the images in the image basic dataset to obtain an image reduction dataset;
[0008] Step S3) Rotate the images in the image reduction dataset to obtain an image rotation dataset;
[0009] Step S4) Reduce the particles in the images of the image basic dataset to obtain a particle reduction dataset;
[0010] Step S5) Perform light compensation processing on the images in the image reduction dataset to obtain a uniform illumination dataset;
[0011] Step S6) Perform non-uniform processing on the uniform illumination dataset to obtain a non-uniform illumination dataset;
[0012] Step S7) Perform quantity statistical analysis on all datasets by category and size to obtain the number of particles contained in each dataset, the category and size of each particle;
[0013] Step S8) Perform accuracy analysis on all datasets to obtain accuracy comparison data, thereby generating a benchmark dataset for particle pollutant microscopic image recognition composed of the accuracy comparison data.
[0014] The method for generating a benchmark dataset for particle pollutant microscopic image recognition specifically includes:
[0015] Step S1-1) Obtain a sequence S of microscopic gray-scale images of particle pollutants according to the particle pollutant inspection method in the international standard ISO 16232;
[0016] Step S1-2) For each image A in the sequence S of microscopic gray-scale images of particle pollutants, mark the larger particle pollutants to obtain a marked image B;
[0017] Step S1-3) Smooth the unmarked particle pollutants into the background, divide the marked image B into blocks to obtain a gray-scale image dataset D11, convert the marked image B into a binary image C to obtain a label image dataset D12, and generate a particle pollutant image basic dataset D1 containing the gray-scale image dataset D11 and the label image dataset D12 according to the gray-scale image dataset D11 and the label image dataset D12;
[0018] Step S2-1) Reduce the gray-scale image dataset D11 by k times to obtain a reduced gray-scale image dataset D21 of the image, where k is a positive integer greater than 2;
[0019] Step S2-2) Convert the binary image C in the label image dataset D12 into a binary image E, and the two values it contains are denoted as {e1, e2}, e1 < e2, and the element mapping relationship from C to E is:
[0020]
[0021] Where c1, c2, c3 are the pixel values of the ternary image C respectively;
[0022] Step S2-3) reducing the binary image E by a factor of k to obtain image E1, and binarizing image E1 to obtain binary image E2 containing pixel values c1 and c3;
[0023] Step S2-4) performing outer contour detection on the binary image E2, changing the pixel value of the outer contour at the corresponding position in the binary image E2 to c2, and obtaining a reduced image label image dataset D22;
[0024] Step 2-5) obtaining an image reduction dataset D2 containing the image reduction grayscale image dataset D21 and the image reduction label image dataset D22 from the image reduction grayscale image dataset D21 and the image reduction label image dataset D22;
[0025] Step S3-1) rotating the reduced grayscale image dataset D21 obtained in step S2-1) by p degrees to obtain a rotated grayscale image dataset D31, where p is any value from 0 to 360 degrees;
[0026] Step S3-2) rotating the binary image E1 obtained in step S2-3) by p degrees to obtain image E3, and binarizing image E3 to obtain a binary rotated image E4 containing pixel values c1 and c3;
[0027] Step S3-3) performing outer contour detection on the binary rotated image E4, changing the pixel value of the outer contour at the corresponding position in the binary rotated image E4 to c2, and obtaining a rotated label image dataset D32;
[0028] Step 3-4) obtaining an image rotation dataset D3 containing the rotated grayscale image dataset D31 and the rotated label image dataset D32 from the rotated grayscale image dataset D31 and the rotated label image dataset D32;
[0029] Step S4-1) performing a connected domain analysis on the binary image E obtained in step S2-2) to obtain a particle pollutant connected domain index mark image F;
[0030] Step S4-2) The connected domain index of the i-th particle pollutant is marked as f i , the background connected domain index is marked as f0, and each index mark f is traversed i , take out f i The image blocks corresponding to the circumscribed rectangle in the grayscale image dataset D11, the binary image E, and the connected domain index marker image F are denoted as grayscale image block G1, binary image block G2, and marker image block G3, respectively;
[0031] Step S4-3) Divide the labeled image B into blocks, find the b pixel values with the largest number of occurrences of unlabeled pixels in each sub-block image, find the b pixel values with the largest number of occurrences of non-particle pollutant pixels in the grayscale image block G1, and form a pixel set V1, V1 = {v1, v2, ..., v b}, b is a positive integer from 5 to 8;
[0032] Step S4-4) For the image blocks G1, G2, and G3, i The region is reduced by a factor of k1 by column to obtain image blocks G11, G21, and G31 with reduced width, where k1 is 2.
[0033] Step S4-5) According to step S4-4), the f in the image blocks G11, G21, and G31 are i The region is reduced by a factor of k1 by row to obtain highly reduced image blocks G12, G22, and G32 respectively;
[0034] Step S4-6) replacing the image block at position G1 in the grayscale image dataset D11 with the image block G12 to obtain a grayscale image dataset D41 with reduced particles;
[0035] Step S4-7) Replace the image block at position G2 in image E with G21 to obtain a binary image E3;
[0036] Step S4-8) performing outer contour detection on the image E3, changing the pixel value of the position corresponding to the outer contour in the image E3 to c2, and obtaining a particle-reduced labeled image dataset D42;
[0037] Step S4-9) Obtaining a particle reduction dataset D4 containing the particle reduction grayscale image dataset D41 and the particle reduction label image dataset D42 from the grayscale image dataset D41 and the label image dataset D42;
[0038] Step S5-1) calculating a background image H based on the image reduction dataset D2 obtained in step S2-1);
[0039] Step S5-2) defines the image illumination fitting function in the particle reduction dataset D4 as:
[0040]
[0041] Where g(x,y) represents the pixel value at coordinates (x,y), (x0,y0) represents the coordinate value of the light source position, t0 is the pixel value at the light source position, k0 is a constant, and the four parameters x0, y0, k0 and t0 in the function g(x,y) are fitted from the background image H;
[0042] Step S5-3) Perform illumination compensation on the grayscale image dataset D21 in the image reduction dataset D2 according to the illumination fitting function g(x, y) to obtain a uniform illumination grayscale image dataset D51. The illumination compensation function is:
[0043]
[0044] Where g′(x,y) is the pixel value after compensation;
[0045] Step S5-4) obtaining a uniform illumination image dataset D5 containing the grayscale image dataset D51 and the label image dataset D22 from the grayscale image dataset D51 and the label image dataset D22;
[0046] Step S6-1) performs illumination non-uniformity processing on the uniform illumination grayscale image dataset D51 according to the illumination fitting function obtained in step S5-2), and the mapping function is:
[0047]
[0048] Adjust the values of x0, y0, k0, and t0 to obtain the non-uniform illumination grayscale image dataset D61;
[0049] Step S6-2) obtaining a non-uniform illumination image dataset D6 containing the non-uniform illumination grayscale image dataset D61 and the label dataset D22 from the non-uniform illumination grayscale image dataset D61 and the label dataset D22;
[0050] Step S7-1) Calculate the length and width of each particle based on the labeled images of all data sets, using the diameter of the minimum circumscribed circle of the outline as the length γ of the particle pollutant, and the width of the minimum circumscribed rectangle of the outline as the width η of the particle pollutant;
[0051] Step S7-2) The particle pollutants are divided into 8 levels with a numerical unit of micrometer according to the width η of the particle pollutants, namely L K (η≥1000), L J (600≤η<1000), L I (400≤η<600), L H (200≤η<400),L G (150≤η<200),L F (100≤η<150), L e (50≤η<100), L D (η<50);
[0052] Step S7-3) Each particle is divided into three types: non-metal, metal and fiber. If the length and width of the particle pollutant satisfy γ / η>λ1 or the contour pixel area in the opencv system function is called satisfy The particle contaminant is fiber, otherwise it is non-fiber, where λ1 and λ2 are thresholds, which are 10 and 0.3 respectively. For non-fiber particles, the proportion of brighter pixel values in the grayscale image is used to determine whether they are metal or non-metal. The brighter pixel value threshold is λ3 = 220, the proportion threshold is λ4 = 0.1, and the number of pixels with values greater than λ3 is if If it is metal, it is non-metal; otherwise, it is metal.
[0053] Step S7-4) counting the number of particles of different types and width levels in each data set;
[0054] Step S8-1) Each dataset is divided into a training set and a test set. Model training and testing are performed using one of the semantic segmentation models of the DeepLabV3+ network. The pixel-level intersection-over-union (IoU), recall, and precision are calculated, as well as the recall and precision of different types and widths of the granularity.
[0055] Step S8-2) Testing the test set in the uniform illumination dataset D5 using the steps in the ISO 16232 standard, calculating the three metrics of pixel-level intersection-over-union (IoU), recall, and precision, as well as the recall and precision of particles of different types and widths.
[0056] Step S8-3) obtains the accuracy comparison data through step S8-1) and step S8-2), thereby generating a benchmark data set for particulate pollutant microscopic image recognition composed of the accuracy comparison data.
[0057] The step S1-2) specifically includes:
[0058] Step S1-2-1) Using a drawing tool, mark the outlines of particle pollutants in image A that are longer or wider than 10 pixels;
[0059] Step S1-2-2) Fill the inner area of the outline to obtain the labeled image B.
[0060] The step S1-3) specifically includes:
[0061] Step S1-3-1) Divide the labeled image B into blocks, find the b pixel values with the largest number of occurrences of unlabeled pixels in each sub-block image to form a pixel set V, denoted as V = {v1, v2, ..., v b};
[0062] Step S1-3-2) Randomly replace the unlabeled pixels of the sub-block with the corresponding elements in V to obtain a grayscale image dataset D11 in VOC format;
[0063] Step S1-3-3) Convert the labeled image B into a ternary image C, with the three values included denoted as {c1, c2, c3}, where the unlabeled pixel value corresponds to c1, the labeled pixel value corresponds to c2, and the pixel value of the filled area corresponds to c3, and c1 < c3 < c2 is satisfied, thereby obtaining the label image dataset D12 in VOC (Visual Object Classes) format;
[0064] Step S1-3-4) The particulate matter pollutant image basic dataset D1 is composed of the grayscale image dataset D11 and the label image dataset D12.
[0065] The specific steps of step S4-4) include:
[0066] Step S4-4-1) Traverse each row of the marked image block G3 to find the starting column value and the ending column value of the consecutive connected rows marked as f i Record the starting column value as Record the ending column value as where (β j , ε j ) represents the j-th consecutive connected row;
[0067] Step S4-4-2) Traverse each consecutive connected row and reduce each consecutive connected row by k1 times;
[0068] Step S4-4-3) Traverse the reduced consecutive connected rows. If there are other non-background connected marked values in the column positions of the image G3 in this row from to , then the entire image block is not reduced for particles. Otherwise, move each consecutive connected row to the left by pixels;
[0069] Step S4-4-4) After modifying all rows of the grayscale image block G1, the binary image block G2, and the marked image block G3 respectively, obtain the width-reduced grayscale image block G11, binary image block G21, and marked image block G31.
[0070] The specific steps of step S5-1) include:
[0071] Step S5-1-1) Traverse the image dataset D2. Denote the pixel value of the i-th image in the grayscale image dataset D21 at (x, y) as u i (x, y), and the pixel value of the i-th image in the label image dataset D22 at (x, y) as w i (x, y);
[0072] Step S5-1-2) The set Z(x, y) is composed of the pixel values of the grayscale image at the (x, y) position:
[0073] Z(x,y)={u i (x,y)|1≤i≤m,w i (x,y)=c1}
[0074] Step S5-1-3) Sort the elements in Z(x,y), and use the median of the sorted elements as the pixel value of the background image H at the (x,y) position.
[0075] The step S1-3-1) specifically includes:
[0076] Step S1-3-1-1) Divide the labeled image B into d×d blocks, and count the histogram of the unlabeled pixels in each sub-block image, which is recorded as
[0077] Step S1-3-1-2) Loop b times and find The maximum value h in i , if h i = 0, then end the loop; otherwise, the maximum value h i As the pixel value ν i Store it in V and let h i =0 to enter the next loop, and after the loop is finished, a pixel set V with b maximum values is obtained = {v1, v2, ..., v b}.
[0078] The step S4-4-1) specifically includes:
[0079] Step S4-4-1-1) Traverse each row of the marked image block G3 and record the initial list of connected row starting point column values as The initial list of endpoint column values is Whether the connectivity identifier hF = 0;
[0080] Step S4-4-1-2) Traverse all columns of the marked image block G3 in the row, and use the value of the mth column Determine whether the column is marked as f i Connect the starting point or end point of the line and add it to the list. If And hF=0, then the column is f i Connect the starting point of the line, and (β j =m) added to In the example, let hF=1; if And hF=1, then the column is f i Connect the endpoints of the line and change ε j =m-1 added to In the example, let hF=0, and make another judgment after all columns are traversed. If ε j=m added to Get the starting column value End point column value
[0081] The step S4-4-2) specifically includes:
[0082] Step S4-4-2-1) traverse each connected row (β j ,ε j ), its width w j =ε j -β j +1, reduce the width by k1 times Modify the grayscale image block G1 in the row β j +l column pixel value for:
[0083]
[0084] where r l represents the random element in the reduced pixel set V1, l represents the random element from β j The lth column after the start, its range is 0≤l<ε j ;
[0085] Step S4-4-2-2) Modify the binary image block G2 in the row β j +l column pixel value for:
[0086]
[0087] Step S4-4-2-3) Modify the marked image block G3 in the row β j +l column pixel value for:
[0088]
[0089] Compared with the existing technology, the advantage of the present invention is that by performing specific processing on the collected microscopic images of particulate pollutants and performing accuracy analysis through specific standard methods, accuracy control data is obtained, thereby generating a benchmark data set for particulate pollutant microscopic image identification composed of accuracy control data. Since the benchmark data set contains detailed information on particulate pollutants, it can provide a reference basis for evaluating the performance of particulate pollutant detection and identification methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0090] Figure 1 It is a schematic flow chart of the method for generating and evaluating a microscopic image data set of particulate pollutants of the present invention;
[0091] Figure 2This is a sub-block image rendering of the labeling process of the present invention;
[0092] Figure 3 is a schematic diagram of the particles of the present invention reduced in width and height;
[0093] Figure 4 It is a schematic diagram of the process of performing illumination change processing in the present invention. DETAILED DESCRIPTION
[0094] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0095] The application principle of the present invention is described in detail below with reference to the accompanying drawings.
[0096] like Figure 1 As shown, a method for generating and evaluating a particle pollutant microscopic image dataset includes the following steps:
[0097] Step S1) collects a sequence of microscopic images of particulate pollutants to generate an image basic data set; Step S1) specifically includes:
[0098] Step S1-1) obtaining a microscopic grayscale image sequence S of particulate matter according to the particulate matter inspection method in the international standard ISO 16232;
[0099] Step S1-2) annotating larger particle pollutants in each image A with a resolution of 1600*1200 in the particle pollutant microscopic grayscale image sequence S to obtain an annotated image B; Step S1-2) specifically includes:
[0100] Step S1-2-1) Using a drawing tool, mark the outlines of particle pollutants in image A that are longer or wider than 10 pixels;
[0101] Step S1-2-2) Fill the inner area of the outline to obtain the labeled image B. The sub-blocks of image A are as follows: Figure 2 As shown in _a, the sub-block after the outline is marked is as follows Figure 2 _b shown;
[0102] Step S1-3) smoothing the unlabeled particulate pollutants as a background, dividing the labeled image B into blocks to obtain a grayscale image dataset D11, converting the labeled image B into a ternary image C to obtain a labeled image dataset D12, and generating a particulate pollutant image basic dataset D1 including the grayscale image dataset D11 and the labeled image dataset D12 based on the grayscale image dataset D11 and the labeled image dataset D12; Step S1-3) specifically includes:
[0103] Step S1-3-1) Divide the labeled image B into blocks, and find the b pixel values with the most occurrences of unlabeled pixels in each sub-block image to form a pixel set V, denoted as V = {v1, v2,..., v b}; The specific steps of step S1-3-1) include:
[0104] Step S1-3-1-1) Divide the labeled image B into blocks of size d×d, and count the histogram of the occurrences of unlabeled pixels in each sub-block image, denoted as
[0105] Step S1-3-1-2) Loop b times for , and find the maximum value h in it. i , if h i = 0, then end the loop; otherwise, take the maximum value h [[ID=?]] i as the pixel value ν i and store it in V, and set h i = 0 to enter the next loop. After the loop ends, obtain a pixel set V = {v1, v2,..., v b} with b maximum values;
[0106] Step S1-3-2) Randomly replace the unlabeled pixels of the sub-blocks with the elements in the corresponding V to obtain a grayscale image dataset D11 in VOC format, and some of its sub-blocks are as shown in Figure 2 _d;
[0107] Step S1-3-3) Convert the labeled image B into a ternary image C, and the three values it contains are denoted as {c1, c2, c3}, where the unlabeled pixel value corresponds to c1, the labeled pixel value corresponds to c2, and the filled area pixel value corresponds to c3, and c1 < c3 < c2 is satisfied, so as to obtain a label image dataset D12 in VOC format;
[0108] Step S1-3-4) The particulate matter pollutant image basic dataset D1 is composed of the grayscale image dataset D11 and the label image dataset D12;
[0109] Step S2) Resize the images in the image basic dataset to obtain a resized image dataset; The specific steps of step S2) include:
[0110] Step S2-1) Resize the grayscale image dataset D11 by k times, where k is a positive integer greater than 2, to obtain a resized grayscale image dataset D21;
[0111] It should be noted that there seems to be an unclear tag in the original text at line 19 ([[ID=?]] i ), which is maintained as it is in the translation for the sake of following the rules.Step S2-2) Convert the ternary image C in the label image dataset D12 into a binary image E, with the two values denoted as {e1, e2}, where e1 < e2. The element mapping relationship from C to E is as follows:
[0112]
[0113] where c1, c2, and c3 are the pixel values of the ternary image C respectively;
[0114] Step S2-3) Reduce the binary image E by k times to obtain an image E1, and then binarize the image E1 to obtain a binary image E2 containing pixel values c1 and c3;
[0115] Step S2-4) Perform external contour detection on the binary image E2, and change the pixel values at the corresponding positions of the external contour in the binary image E2 to c2 to obtain a reduced label image dataset D22;
[0116] Step 2-5) Obtain an image reduction dataset D2 containing the reduced grayscale image dataset D21 and the reduced label image dataset D22 from the reduced grayscale image dataset D21 and the reduced label image dataset D22;
[0117] Step S3) Rotate the images in the image reduction dataset to obtain an image rotation dataset; The specific steps of Step S3) are as follows:
[0118] Step S3-1) Rotate the reduced grayscale image dataset D21 obtained in Step S2-1) by p degrees to obtain a rotated grayscale image dataset D31, where p is any value from 0 to 360°;
[0119] Step S3-2) Rotate the binary image E1 obtained in Step S2-3) by p degrees to obtain an image E3, and then binarize the image E3 to obtain a binary rotated image E4 containing pixel values c1 and c3;
[0120] Step S3-3) Perform external contour detection on the binary rotated image E4, and change the pixel values at the corresponding positions of the external contour in the binary rotated image E4 to c2 to obtain a rotated label image dataset D32;
[0121] Step 3-4) Obtain an image rotation dataset D3 containing the rotated grayscale image dataset D31 and the rotated label image dataset D32 from the rotated grayscale image dataset D31 and the rotated label image dataset D;
[0122] Step S4) Reduce the images in the image basic dataset to obtain a granule reduction dataset; The specific steps of Step S4) are as follows:
[0123] Step S4-1) performing a connected domain analysis on the binary image E obtained in step S2-2) to obtain a particle pollutant connected domain index mark image F;
[0124] Step S4-2) The connected domain index of the i-th particle pollutant is marked as f i , the background connected domain index is marked as f0, and each index mark f is traversed i , take out f i The image blocks corresponding to the circumscribed rectangle in the grayscale image dataset D11, the binary image E, and the connected domain index marker image F are denoted as grayscale image block G1, binary image block G2, and marker image block G3, respectively;
[0125] Step S4-3) Divide the labeled image B into blocks, find the b pixel values with the largest number of occurrences of unlabeled pixels in each sub-block image, find the b pixel values with the largest number of occurrences of non-particle pollutant pixels in the grayscale image block G1, and form a pixel set V1, V1 = {v1, v2, ..., v b}, b is a positive integer from 5 to 8;
[0126] Step S4-4) For the image blocks G1, G2, and G3, i The region is reduced by a factor of k1 by columns to obtain width-reduced image blocks G11, G21, and G31, respectively, where k1 is 2. Step S4-4) specifically includes:
[0127] Step S4-4-1) traverse each row of the marked image block G3 and find each row marked as f i The starting column value and the ending column value of the connected row are recorded as The end point column value is recorded as Among them (β j ,ε j ) represents the jth connected row; the step S4-4-1) specifically includes:
[0128] Step S4-4-1-1) Traverse each row of the marked image block G3 and record the initial list of connected row starting point column values as The initial list of endpoint column values is Whether the connectivity identifier hF = 0;
[0129] Step S4-4-1-2) Traverse all columns of the marked image block G3 in the row, and use the value of the mth column Determine whether the column is marked as f i Connect the starting point or end point of the line and add it to the list. If And hF=0, then the column is f i Connect the starting point of the line, and (β j =m) added to In the example, let hF=1; if And hF=1, then the column is f i Connect the endpoints of the line and change ε j =m-1 added to In the example, let hF=0, and make another judgment after all columns are traversed. If ε j =m added to Get the starting column value End point column value
[0130] Step S4-4-2) traverses each connected line segment and reduces each connected line segment by a factor of k1; the step S4-4-2) specifically includes:
[0131] Step S4-4-2-1) traverse each connected row (β j ,ε j ), its width w j =ε j -β j +1, reduce the width by k1 times Modify the grayscale image block G1 in the row β j +l column pixel value for:
[0132]
[0133] where r l represents the random element in the reduced pixel set V1, l represents the random element from β j The lth column after the start, its range is 0≤l<ε j ;
[0134] Step S4-4-2-2) Modify the binary image block G2 in the row β j +l column pixel value for:
[0135]
[0136] Step S4-4-2-3) Modify the marked image block G3 in the row β j +l column pixel value for:
[0137]
[0138] Step S4-4-3) traverse the connected rows after the reduction, if the image G3 is in the column position of the row from arrive If there are other non-background connected marker values on the image block, the entire image block will not be reduced in size, otherwise each connected row will be moved left. pixels;
[0139] Step S4-4-4) Modify all rows of the grayscale image block G1, the binary image block G2, and the marked image block G3 to obtain a reduced-width grayscale image block G11, a binary image block G21, and a marked image block G31, respectively;
[0140] Step S4-5) According to step S4-4), the f in the image blocks G11, G21, and G31 are i The region is reduced by a factor of k1 by row to obtain highly reduced image blocks G12, G22, and G32 respectively;
[0141] Step S4-6) replacing the image block at position G1 in the grayscale image dataset D11 with the image block G12 to obtain a grayscale image dataset D41 with reduced particles;
[0142] Step S4-7) Replace the image block at position G2 in image E with G21 to obtain a binary image E3;
[0143] Step S4-8) performing outer contour detection on the image E3, changing the pixel value of the position corresponding to the outer contour in the image E3 to c2, and obtaining a particle-reduced labeled image dataset D42;
[0144] Step S4-9) Obtaining a particle reduction dataset D4 containing the particle reduction grayscale image dataset D41 and the particle reduction label image dataset D42 from the grayscale image dataset D41 and the label image dataset D42;
[0145] Step S5) performs illumination compensation processing on the images in the image reduction dataset to obtain a uniform illumination dataset; the step S5) specifically includes:
[0146] Step S5-1) calculates the background image H based on the image reduction dataset D2 obtained in step S2-1); a grayscale image in the dataset D2 is as follows Figure 4 As shown in _a, the background image H is obtained as Figure 4 _b, the resolution is 800*600; the step S5-1) specifically includes:
[0147] Step S5-1-1) Traverse the image dataset D2 and record the pixel value of the i-th image at (x, y) in the grayscale image dataset D21 as u i (x, y), the pixel value of the i-th image at (x, y) in the labeled image dataset D22 is w i (x,y);
[0148] Step S5-1-2) The pixel values of the grayscale image at the position (x, y) form a set Z(x, y):
[0149] Z(x,y)={u i (x,y)|1≤i≤m,w i (x,y)=c1}
[0150] Step S5-1-3) Sort the elements in Z(x,y) and use the median of the sort as the pixel value of the background image H at the (x,y) position
[0151] Step S5-2) defines the image illumination fitting function in the particle reduction dataset D4 as:
[0152]
[0153] Where g(x,y) represents the pixel value at coordinates (x,y), (x0,y0) represents the coordinate value of the light source position, t0 is the pixel value at the light source position, k0 is a constant, and the four parameters x0, y0, k0 and t0 in the function g(x,y) are fitted from the background image H. The results are x0=334, y0=389, k0=1159000, and t0=192;
[0154] Step S5-3) Perform illumination compensation on the grayscale image dataset D21 in the image reduction dataset D2 according to the illumination fitting function g(x,y) to obtain a uniform illumination grayscale image dataset D51, in which an image such as Figure 4 As shown in _c, the illumination compensation function is:
[0155]
[0156] Where g′(x,y) is the pixel value after compensation;
[0157] Step S5-4) obtaining a uniform illumination image dataset D5 containing the grayscale image dataset D51 and the label image dataset D22 from the grayscale image dataset D51 and the label image dataset D22;
[0158] Step S6) performs non-uniform processing on the uniform illumination dataset to obtain a non-uniform illumination dataset; the step S6) specifically includes:
[0159] Step S6-1) performs illumination non-uniformity processing on the uniform illumination grayscale image dataset D51 according to the illumination fitting function obtained in step S5-2), and the mapping function is:
[0160]
[0161] Adjust the values of x0, y0, k0 and t0, set x0 = 334, y0 = 389, k0 = 115900, t0 = 192, and obtain the non-uniform illumination grayscale image dataset D61. Figure 4 _d shown;
[0162] Step S6-2) obtaining a non-uniform illumination image dataset D6 containing the non-uniform illumination grayscale image dataset D61 and the label dataset D22 from the non-uniform illumination grayscale image dataset D61 and the label dataset D22;
[0163] Step S7) performs quantitative statistical analysis on all data sets by category and size to obtain the number of particles contained in each data set, and the category and size of each particle; Step S7) specifically includes:
[0164] Step S7-1) Calculate the length and width of each particle based on the labeled images of all data sets, using the diameter of the minimum circumscribed circle of the outline as the length γ of the particle pollutant, and the width of the minimum circumscribed rectangle of the outline as the width η of the particle pollutant;
[0165] Step S7-2) The particle pollutants are divided into 8 levels with a numerical unit of micrometer according to the width η of the particle pollutants, namely L K (η≥1000), L J (600≤η<1000), L I (400≤η<600), L H (200≤η<400),L G (150≤η<200),L F (100≤η<150), L e (50≤η<100), L D (η<50);
[0166] Step S7-3) Each particle is divided into three types: non-metal, metal and fiber. If the length and width of the particle pollutant satisfy γ / η>λ1 or the contour pixel area in the opencv system function is called satisfy The particle contaminant is fiber, otherwise it is non-fiber, where λ1 and λ2 are thresholds, which are 10 and 0.3 respectively. For non-fiber particles, the proportion of brighter pixel values in the grayscale image is used to determine whether they are metal or non-metal. The brighter pixel value threshold is λ3 = 220, the proportion threshold is λ4 = 0.1, and the number of pixels with values greater than λ3 is if If it is metal, it is non-metal; otherwise, it is metal.
[0167] Step S7-4) counting the number of particles of different types and width levels in each data set;
[0168] Step S8) performs accuracy analysis on all data sets to obtain accuracy comparison data, thereby generating a benchmark data set for particulate pollutant microscopic image recognition consisting of the accuracy comparison data. Step S8-1) specifically includes:
[0169] Step S8-1) Each data set is divided into a training set and a test set, with 200 images in the training set and 85 images in the test set. The DeepLabV3+ network is used for model training and testing. The three indicators of intersection-over-union (IoU), recall, and precision are calculated at the pixel level, as well as the recall and precision of different types and widths at the particle level. The results at the pixel level are shown in Table 2, and the results at the particle level are shown in Tables 3-1 to 3-6.
[0170] Step S8-2) The test set in the uniform illumination dataset D5 is tested using the steps in the ISO 16232 standard. The three metrics of IoU, recall, and precision are calculated at the pixel level. The recall and precision of the granules at different types and widths are also calculated. The results are shown in Table 4. The recall and precision of the granules at different types and widths are also calculated. The results are shown in Table 5.
[0171] Step S8-3) obtains the accuracy comparison data through step S8-1) and step S8-2), thereby generating a benchmark data set for particulate pollutant microscopic image recognition composed of the accuracy comparison data.
[0172] Table 1 Classification statistics
[0173] Table 1-1 Particle count statistics for dataset D1
[0174]
[0175] Table 1-2 Statistics of the number of particles in the dataset D2
[0176]
[0177] Table 1-3 Statistics of the number of particles in the D3 dataset
[0178]
[0179]
[0180] Table 1-4 Dataset D4 particle count statistics
[0181]
[0182] Table 1-5 Statistics of the number of particles in the dataset D5
[0183]
[0184]
[0185] Table 1-6 Dataset D6 particle count statistics
[0186]
[0187] Table 2 Pixel-level test results of DeepLabV3+ datasets
[0188]
[0189]
[0190] Table 3-1 DeepLabV3+ dataset D1 granularity test results
[0191]
[0192]
[0193] Table 3-2 DeepLabV3+ dataset D2 granularity test results
[0194]
[0195]
[0196] Table 3-3 DeepLabV3+ dataset D3 granularity test results
[0197]
[0198]
[0199] Table 3-4 DeepLabV3+ dataset D4 granularity test results
[0200]
[0201]
[0202] Table 3-5 DeepLabV3+ dataset D5 granularity test results
[0203]
[0204]
[0205] Table 3-6 DeepLabV3+ dataset D6 granularity test results
[0206]
[0207]
[0208] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for generating a benchmark dataset for microscopic image recognition of particulate pollutants, characterized in that: The method includes: Step S1) Collect a microscopic image sequence of particulate pollutants to generate an image basic data set, specifically: Step S1-1) Obtain a microscopic gray-scale image sequence S of particulate pollutants according to the particulate pollutant inspection method in the international standard ISO 16232; Step S1-2) For each image A in the microscopic gray-scale image sequence S of particulate pollutants, mark the larger particulate pollutants to obtain a marked image B; Step S1-3) Smooth the unmarked particulate pollutants into the background, divide the marked image B into blocks to obtain a gray-scale image data set D11, convert the marked image B into a ternary image C to obtain a label image data set D12, and generate a particulate pollutant image basic data set D1 containing the gray-scale image data set D11 and the label image data set D12; Step S2) Reduce the images in the image basic data set to obtain an image reduction data set, specifically: Step S2-1) Reduce the gray-scale image data set D11 by k times to obtain a reduced gray-scale image data set D21 of the image, where k is a positive integer greater than 2; Step S2-2) Convert the ternary image C in the label image data set D12 into a binary image E, and the two values included are denoted as {e1, e2}, e1 < e2, and the element mapping relationship from C to E is: where c1, c2, and c3 are the pixel values of the ternary image C respectively; Step S2-3) Reduce the binary image E by k times to obtain an image E1, and perform binarization on the image E1 to obtain a binary image E2 containing pixel values c1 and c3; Step S2-4) Perform outer contour detection on the binary image E2, and change the pixel values at the corresponding positions of the outer contour in the binary image E2 to c2 to obtain a reduced label image data set D22 of the image; Step S2-5) Obtain an image reduction data set D2 containing the reduced gray-scale image data set D21 and the reduced label image data set D22 from the reduced gray-scale image data set D21 and the reduced label image data set D22 of the image; Step S3) Rotate the images in the image reduction data set to obtain an image rotation data set, specifically: Step S3-!) Rotate the reduced gray-scale image data set D21 obtained in step S2-1) by p degrees to obtain a rotated gray-scale image data set D31, where p is any value from 0 to 360°; Step S3-2) Rotate the binary image E1 obtained in step S2-3) by p degrees to obtain an image E3, and perform binarization on the image E3 to obtain a binary rotated image E4 containing pixel values c1 and c3; Step S3-3) Perform outer contour detection on the binary rotated image E4, and change the pixel values at the corresponding positions of the outer contour in the binary rotated image E4 to c2 to obtain a rotated label image data set D32; Step S3-4) Obtain an image rotation data set D3 containing the rotated gray-scale image data set D31 and the rotated label image data set D32 from the rotated gray-scale image data set D31 and the rotated label image data set D32; Step S4) performs particle reduction on the images in the image basic dataset to obtain a particle reduction dataset, specifically: Step S4-1) performing a connected domain analysis on the binary image E obtained in step S2-2) to obtain a particle pollutant connected domain index mark image F; Step S4-2) The connected domain index of the i-th particle pollutant is marked as f i , the background connected domain index is marked as f0, and each index mark f is traversed i , take out f i The image blocks corresponding to the circumscribed rectangle in the grayscale image dataset D11, the binary image E, and the connected domain index marker image F are denoted as grayscale image block G1, binary image block G2, and marker image block G3, respectively; Step S4-3) Divide the labeled image B into blocks, find the b pixel values of the grayscale image block G1 with the largest number of non-particle pollutant pixels to form a pixel set V1, V1 = {v1, v2, ..., v b }, b is a positive integer from 5 to 8; Step S4-4) For the image blocks G1, G2, and G3, i The region is reduced by a factor of k1 by column to obtain image blocks G11, G21, and G31 with reduced width, where k1 is 2. Step S4-5) According to step S4-4), the f in the image blocks G11, G21, and G31 are i The region is reduced by a factor of k1 by row to obtain highly reduced image blocks G12, G22, and G32 respectively; Step S4-6) replacing the image block at position G1 in the grayscale image dataset D11 with the image block G12 to obtain a grayscale image dataset D41 with reduced particles; Step S4-7) Replace the image block at position G2 in image E with G21 to obtain a binary image E3; Step S4-8) performing outer contour detection on the image E3, changing the pixel value of the position corresponding to the outer contour in the image E3 to c2, and obtaining a particle-reduced labeled image dataset D42; Step S4-9) Obtaining a particle reduction dataset D4 containing the particle reduction grayscale image dataset D41 and the particle reduction label image dataset D42 from the grayscale image dataset D41 and the label image dataset D42; Step S5) performs illumination compensation processing on the images in the image reduction dataset to obtain a uniform illumination dataset, specifically: Step S5-1) calculating a background image H based on the image reduction dataset D2 obtained in step S2-1); Step S5-2) defines the image illumination fitting function in the particle reduction dataset D4 as: Where g(x,y) represents the pixel value at coordinates (x,y), (x0,y0) represents the coordinate value of the light source position, t0 is the pixel value at the light source position, k0 is a constant, and the four parameters x0, y0, k0 and t0 in the function g(x,y) are fitted from the background image H; Step S5-3) Perform illumination compensation on the grayscale image dataset D21 in the image reduction dataset D2 according to the illumination fitting function g(x, y) to obtain a uniform illumination grayscale image dataset D51. The illumination compensation function is: Where g′(x,y) is the pixel value after compensation; Step S5-4) obtaining a uniform illumination image dataset D5 containing the grayscale image dataset D51 and the label image dataset D22 from the grayscale image dataset D51 and the label image dataset D22; Step S6) performs non-uniform processing on the uniform illumination data set to obtain a non-uniform illumination data set, specifically: Step S6-1) performs illumination non-uniformity processing on the uniform illumination grayscale image dataset D51 according to the illumination fitting function obtained in step S5-2), and the mapping function is: Adjust the values of x0, y0, k0, and t0 to obtain the non-uniform illumination grayscale image dataset D61; Step S6-2) obtaining a non-uniform illumination image dataset D6 from the non-uniform illumination grayscale image dataset D61 and the label dataset D22; Step S7) Perform quantitative statistical analysis on all data sets by category and size to obtain the number of particles contained in each data set, the category and size of each particle, specifically: Step S7-1) Calculate the length and width of each particle based on the labeled images of all data sets, using the diameter of the minimum circumscribed circle of the outline as the length γ of the particle pollutant, and the width of the minimum circumscribed rectangle of the outline as the width η of the particle pollutant; Step S7-2) The particle pollutants are divided into 8 levels with a numerical unit of micrometer according to the width η of the particle pollutants, namely L K (η≥1000), L J (600≤η<1000), L I (400≤η<600), L H (200≤η<400),L G (150≤η<200),L F (100≤η<150), L e (50≤η<100), L D (η<50); Step S7-3) Each particle is divided into three types: non-metal, metal and fiber. If the length and width of the particle pollutant satisfy γ / η>λ1 or the contour pixel area in the opencv system function is called satisfy The particle contaminant is fiber, otherwise it is non-fiber, where λ1 and λ2 are thresholds, which are 10 and 0.3 respectively. For non-fiber particles, the proportion of brighter pixel values in the grayscale image is used to determine whether they are metal or non-metal. The brighter pixel value threshold is λ3 = 220, the proportion threshold is λ4 = 0.1, and the number of pixels with values greater than λ3 is if If it is metal, it is non-metal; otherwise, it is metal. Step S7-4) Count the number of particles of each type and each width level in each dataset; Step S8) Perform accuracy analysis on all datasets to obtain accuracy comparison data, and thus generate a benchmark dataset for particulate pollutant microscopic image recognition composed of the accuracy comparison data, specifically: Step S8-1) Divide each dataset into a training set and a test set, use the DeepLabV3+ network for model training and testing, calculate three metrics: intersection over union, recall, and precision at the pixel level, as well as recall and precision at the particulate level under different types and different width levels; Step S8-2) Test the test set in the uniform illumination dataset D5 using the steps in ISO 16232 standard, and calculate three metrics: intersection over union, recall, and precision at the pixel level, as well as recall and precision at the particulate level under different types and different width levels; Step S8-3) Generate a benchmark dataset for particulate pollutant microscopic image recognition composed of the accuracy comparison data from Steps S8-1) and S8-2).
2. The method for generating a benchmark dataset for microscopic image recognition of particulate pollutants according to claim 1, characterized in that: The specific content of Step S1-2) includes: Step S1-2-1) Use a drawing tool to mark the contours of particulate pollutants in Image A with a length or width greater than 10 pixels; Step S1-2-2) Fill the area inside the contour to obtain a labeled image B.
3. The method for generating a benchmark dataset for microscopic image recognition of particulate pollutants according to claim 1, characterized in that: The specific content of Step S1-3) includes: Step S1-3-1) Divide the labeled image B into blocks, find the b pixel values with the largest number of occurrences of unlabeled pixels in each sub-block image to form a pixel set V, denoted as V = {v1, v2, ..., v b }; Step S1-3-2) Randomly replace the unlabeled pixels in the sub-blocks with the elements in the corresponding V to obtain a grayscale image dataset D11 in VOC format; Step S1-3-3) Convert the labeled image B into a ternary image C, and the three values it contains are denoted as {c1, c2, c3}, where the unlabeled pixel value corresponds to c1, the labeled pixel value corresponds to c2, and the filled area pixel value corresponds to c3, and c1 < c3 < c2 is satisfied, so as to obtain a label image dataset D12 in VOC format; Step S1-3-4) The particulate pollutant image basic dataset D1 is composed of the grayscale image dataset D11 and the label image dataset D12.
4. The method for generating a benchmark dataset for microscopic image recognition of particulate pollutants according to claim 1, characterized in that The specific content of Step S4-4) includes: Step S4-4-1) traverse each row of the marked image block G3 and find each row marked as f i The starting column value and the ending column value of the connected row are recorded as The end point column value is recorded as Among them (β j ,ε j ) represents the jth connected row; Step S4-4-2) Traverse each connected line segment and reduce each connected line segment by k1 times; Step S4-4-3) traverse the connected rows after the reduction, if the image G3 is in the column position of the row from arrive If there are other non-background connected marker values on the image block, the entire image block will not be reduced in size, otherwise each connected row will be moved left. pixels; Step S4-4-4) After modifying all rows of the grayscale image block G1, binary image block G2, and marker image block G3, obtain the width-reduced grayscale image block G11, binary image block G21, and marker image block G31 respectively.
5. The method for generating a benchmark data set for microscopic image recognition of particulate pollutants according to claim 1, characterized in that: The specific content of Step S5-1) includes: Step S5-1-1) Traverse the image dataset D2 and record the pixel value of the i-th image at (x, y) in the grayscale image dataset D21 as u i (x, y), the pixel value of the i-th image at (x, y) in the labeled image dataset D22 is w i (x,y); Step S5-1-2) The pixel values of the grayscale image at the (x, y) position form a set Z(x, y): Z(x,y)={u i (x,y)|1≤i≤m,w i (x,y)=c1} Step S5-1-3) Sort the elements in Z(x, y), and use the median after sorting as the pixel value of the background image H at the (x, y) position.
6. The method for generating a benchmark data set for microscopic image recognition of particulate pollutants according to claim 3, characterized in that: The specific content of Step S1-3-1) includes: Step S1-3-1-1) Divide the labeled image B into d×d blocks, and count the histogram of the unlabeled pixels in each sub-block image, which is recorded as Step S1-3-1-2) Loop b times and find The maximum value h in i , if h i = 0, then the loop ends; otherwise, the maximum value h i As the pixel value ν i Store it in V and let h i =0 to enter the next loop, and after the loop is finished, a pixel set V with b maximum values is obtained = {v1, v2, ..., v b }.
7. The method for generating a benchmark data set for microscopic image recognition of particulate pollutants according to claim 4, characterized in that: The specific content of Step S4-4-1) includes: Step S4-4-1-1) Traverse each row of the marked image block G3 and record the initial list of connected row starting point column values as The initial list of endpoint column values is Whether the connectivity identifier hF = 0; Step S4-4-1-2) Traverse all columns of the marked image block G3 in the row, and use the value of the mth column Determine whether the column is marked as f i Connect the starting point or end point of the line and add it to the list. If And hF=0, then the column is f i Connect the starting point of the line, and (β j =m) added to In the example, let hF=1; if And hF=1, then the column is f i Connect the endpoints of the line and change ε j =m-1 added to In the example, let hF=0, and make another judgment after all columns are traversed. If ε j =m added to Get the starting column value End point column value 8. The method for generating a benchmark data set for microscopic image recognition of particulate pollutants according to claim 4, characterized in that: The specific content of Step S4-4-2) includes: Step S4-4-2-1) traverse each connected row (β j ,ε j ), its width w j =ε j -β j +1, reduce the width by k1 times Modify the grayscale image block G1 in the row β j +l column pixel value for: where r l represents the random element in the reduced pixel set V1, l represents the random element from β j The lth column after the start, its range is 0≤l<ε j ; Step S4-4-2-2) Modify the binary image block G2 in the row β j +l column pixel value for: Step S4-4-2-3) Modify the marked image block G3 in the row β j +l column pixel value for:
Citation Information
Patent Citations
An automatic segmentation method for thin section microscopic image of sandstone
CN109523566A
Method for detecting and identifying particulate pollutants in image and computer readable storage medium
CN113139953A