Detection method, detection device, detection apparatus, and computer-readable storage medium
By performing data enhancement and multi-angle detection on printed images, generating multiple intermediate images and combining the detection results, the problem of insufficient accuracy in extracting regions of interest by printed material inspection machines is solved, achieving higher detection accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING LUSTER LIGHTTECH
- Filing Date
- 2022-12-29
- Publication Date
- 2026-07-31
AI Technical Summary
When inspecting printed materials for defects, printing inspection machines often struggle to accurately extract regions of interest from images, resulting in insufficient inspection accuracy.
Multiple intermediate images are generated by performing data augmentation on the image to be tested, and then input into a preset detection model for multiple detections. The detection results from different angles are comprehensively considered to determine the final detection result.
It improves the accuracy of printed matter inspection, especially the extraction accuracy of regions of interest (such as text regions).
Smart Images

Figure CN116188765B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of testing technology, and more specifically, to a testing method, a testing device, a testing equipment, and a non-volatile computer-readable storage medium. Background Technology
[0002] When inspecting printed materials for defects, the inspection machine needs to accurately extract the region of interest (ROI) of the printed material to complete the defect detection. Therefore, there is an urgent need for a solution that can accurately extract the region of interest (ROI) in an image. Summary of the Invention
[0003] This application provides a detection method, a detection apparatus, a detection device, and a non-volatile computer-readable storage medium.
[0004] The detection method of this application includes performing data augmentation on the image to be tested to generate multiple intermediate images; inputting the intermediate images into a preset detection model to determine the detection result of each intermediate image; and determining the detection result of the image to be tested based on the multiple detection results.
[0005] The detection apparatus of this application includes an enhancement module, an input module, and a determination module. The enhancement module is used to perform data enhancement on the image to be tested to generate multiple intermediate images; the input module is used to input the intermediate images into a preset detection model to determine the detection result of each intermediate image; the determination module is used to determine the detection result of the image to be tested based on the multiple detection results.
[0006] The detection device according to the embodiments of this application includes a processor, which is used to perform data augmentation on the image to be tested to generate a plurality of intermediate images; input the intermediate images to a preset detection model to determine the detection result of each intermediate image; and determine the detection result of the image to be tested based on the plurality of detection results.
[0007] The non-volatile computer-readable storage medium of this application includes a computer program that, when executed by a processor, causes the processor to perform the detection method. The detection method includes performing data augmentation on an image to be tested to generate a plurality of intermediate images; inputting the intermediate images into a preset detection model to determine a detection result for each intermediate image; and determining a detection result for the image to be tested based on the plurality of detection results.
[0008] The detection method, detection apparatus, detection equipment, and non-volatile computer-readable storage medium of this application generate multiple intermediate images by performing data enhancement on the image to be tested; then, the multiple intermediate images are input into a preset detection model to obtain the detection result of each intermediate image, thereby obtaining the detection result of the image to be tested at different angles (such as rotation angles); finally, the detection result of the image to be tested is determined based on the detection results of the multiple intermediate images. Compared with detecting the detection result of the image to be tested at only a single angle, performing multiple detections on intermediate images of the image to be tested at different angles and comprehensively considering the detection results of intermediate images at multiple different angles to determine the final detection result of the image to be tested can improve the detection accuracy of the detection result of the image to be tested, thereby achieving accurate extraction of the region of interest (such as text region) in the image to be tested.
[0009] Additional aspects and advantages of embodiments of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of embodiments of this application. Attached Figure Description
[0010] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, wherein:
[0011] Figure 1 This is a flowchart illustrating the detection method of some embodiments of this application;
[0012] Figure 2 This is a flowchart illustrating the detection method of some embodiments of this application;
[0013] Figure 3 This is a flowchart illustrating the detection method of some embodiments of this application;
[0014] Figure 4 This is a schematic diagram illustrating a usage scenario of the detection method according to certain embodiments of this application;
[0015] Figure 5 This is a flowchart illustrating the detection method of some embodiments of this application;
[0016] Figure 6 This is a flowchart illustrating the detection method of some embodiments of this application;
[0017] Figure 7 This is a flowchart illustrating the detection method of some embodiments of this application;
[0018] Figure 8 This is a flowchart illustrating the detection method of some embodiments of this application;
[0019] Figure 9This is a flowchart illustrating the detection method of some embodiments of this application;
[0020] Figure 10 This is a flowchart illustrating the detection method of some embodiments of this application;
[0021] Figure 11 This is a flowchart illustrating the detection method of some embodiments of this application;
[0022] Figure 12 This is a flowchart illustrating the detection method of some embodiments of this application;
[0023] Figure 13 This is a flowchart illustrating the detection method of some embodiments of this application;
[0024] Figure 14 This is a schematic diagram of the detection device according to some embodiments of this application;
[0025] Figure 15 This is a plan view of the detection device according to some embodiments of this application;
[0026] Figure 16 This is a schematic diagram illustrating the connection state of a non-volatile computer-readable storage medium and a processor in certain embodiments of this application. Detailed Implementation
[0027] The embodiments of this application are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the embodiments of this application, and should not be construed as limiting the embodiments of this application.
[0028] Please see Figure 1 This application provides a detection method, which includes:
[0029] Step 011: Perform data augmentation on the image to be tested to generate multiple intermediate images;
[0030] Specifically, before detecting the image to be tested, data augmentation is required to obtain multiple intermediate images, thereby increasing the number of detection samples for the image to be tested.
[0031] Furthermore, data augmentation includes at least one of rotation, reflection, mirroring, flipping, and scaling. After acquiring the image to be tested, several data augmentation methods can be randomly selected to perform different flipping transformations on the image to be tested, thereby generating intermediate images. Alternatively, the data augmentation methods can be specified, such as rotation, reflection, or mirroring of the image to be tested. In this case, each time the image to be tested is acquired, the generated intermediate images are intermediate images after rotation, reflection, and mirroring. The rotation angle of the image to be tested can be different, such as 90 degrees or 150 degrees. In this way, after data augmentation of the image to be tested, different forms of the image to be tested are obtained, i.e., intermediate images. This allows the detection model to detect the image to be tested in different forms and determine the detection result of the image to be tested based on the different intermediate images. This avoids insufficient detection accuracy of pixels in a certain form when the detection model detects the image to be tested, which would affect the detection accuracy of the detection model.
[0032] Step 012: Input the intermediate images into the preset detection model to determine the detection result for each intermediate image;
[0033] Specifically, after generating multiple intermediate images, these images can be input into a pre-defined detection model to determine the detection result for each intermediate image. Considering that some regions of interest (such as text regions) in the test image are pixel-level, the detection model chooses to detect the test image based on a segmentation network. The detection result is either whether each pixel in the test image is a pixel of a region of interest, such as whether a pixel is a text pixel, or the confidence level that each pixel in the test image is a text pixel.
[0034] Step 013: Determine the detection result of the image to be tested based on multiple detection results.
[0035] Specifically, after obtaining the detection results of each intermediate image, the detection results of the image to be tested can be determined based on the detection results of each intermediate image. Compared with the method of only performing one detection on the image to be tested, this application can detect the detection results of the image to be tested from different angles, and the impact of errors in a single detection on the detection results can be reduced through multiple detections.
[0036] Furthermore, when determining the detection result of the image to be tested based on the detection results of each intermediate image, the detection results of multiple intermediate images can be weighted and averaged, and the average result after weighted averaging can be used as the detection result. For example, data augmentation can be performed by rotating the image to be tested by 90 degrees, rotating it by 150 degrees, and flipping it horizontally. In this case, the image to be tested after rotating it by 90 degrees can be used as an intermediate image, the detection image after rotating it by 150 degrees can be used as an intermediate image, and the detection image after flipping it horizontally can be used as an intermediate image. Then, the three intermediate images are input into a preset detection model to obtain the corresponding detection results. Then, the three detection results are weighted and averaged to determine the detection result of the image to be tested. For example, the confidence scores of pixels at the same position in the detection results of the three intermediate images can be weighted and averaged, and the weighted average confidence score can be used as the confidence score of the pixel at the corresponding position in the image to be tested. Alternatively, when inputting intermediate images into the preset detection model, the image to be tested itself can also be used as an intermediate image input into the detection model to determine the corresponding detection result, and when determining the detection result of the image to be tested, the detection results of the intermediate images corresponding to the image to be tested itself can also be included in the calculation of the detection result of the image to be tested. For example, the detection results of the intermediate image corresponding to the image under test, the detection results of the intermediate image obtained after rotating the image under test by 90 degrees, the detection results of the intermediate image obtained after rotating the image under test by 150 degrees, and the detection results of the intermediate image obtained after horizontally flipping the image under test are weighted and averaged, and the detection result of the image under test is determined based on the weighted average result. That is, the detection result of the image under test can be calculated according to the formula R = Mean(M(x), M(rot90(x)), M(rot150(x)), M(flip(x))), where Mean is the method for calculating the weighted average of the detection results of multiple intermediate images (this embodiment is described with each item having a weight of 1 as an example), M is the preset detection model, x is the image under test, rot90(x) is the intermediate image obtained after rotating the image under test by 90 degrees, rot150(x) is the intermediate image obtained after rotating the image under test by 150 degrees, and flip(x) is the intermediate image obtained after horizontally flipping the image under test. In this way, the detection results of the intermediate images from four different angles can be comprehensively considered to determine the final detection result of the image under test, thereby improving the detection accuracy of the image under test.
[0037] The detection method of this application generates multiple intermediate images by performing data augmentation on the image to be tested; then, the multiple intermediate images are input into a preset detection model to obtain the detection result of each intermediate image, thereby obtaining the detection result of the image to be tested at different angles (such as rotation angles). Finally, the detection result of the image to be tested is determined based on the detection results of the multiple intermediate images. Compared with detecting the detection result of the image to be tested at only a single angle, performing multiple detections on intermediate images of the image to be tested at different angles and comprehensively considering the detection results of intermediate images at multiple different angles to determine the final detection result of the image to be tested can improve the detection accuracy of the detection result of the image to be tested, thereby achieving accurate extraction of the region of interest (such as text region) in the image to be tested.
[0038] Please see Figure 2 In some implementations, the detection results include the confidence score of each pixel as a preset type. Step 012: Input the intermediate image into a preset detection model to determine the detection result for each intermediate image, including:
[0039] Step 0121: Input the intermediate images into the preset detection model to determine the confidence level of each pixel in each intermediate image;
[0040] Step 013: Based on multiple detection results, determine the detection results of the image to be tested, including;
[0041] Step 0131: Align multiple intermediate images;
[0042] Step 0132: Determine the confidence level of the pixels at the target location in the image under test based on the confidence levels of the pixels at the target location in the multiple aligned intermediate images.
[0043] Specifically, the detection results include the confidence score of each pixel as belonging to a preset type. The confidence score can be understood as the probability that each pixel belongs to a preset type, which can be categorized into text type, face type, etc. After inputting the intermediate images into the preset detection model, the confidence score of each pixel in each intermediate image can be determined, thus facilitating the determination of the confidence score of each pixel in the target image. Since the intermediate images are the target images after flipping transformation, when determining the detection results of the target image based on the detection results of the intermediate images, multiple intermediate images need to be aligned to match the pixels in each intermediate image. For example, if the target image is rotated 90 degrees clockwise to obtain intermediate image A, and rotated 150 degrees clockwise to obtain intermediate image B, then during alignment, intermediate image A needs to be rotated 90 degrees counterclockwise, and intermediate image B needs to be rotated 150 degrees counterclockwise, so that the target positions of intermediate images A and B are aligned.
[0044] After aligning multiple intermediate images, the confidence level of the target location pixels in the test image can be determined based on the confidence levels of the pixels at the target location in the aligned intermediate images. The target location is the same location in the multiple intermediate images. For example, if the target location's coordinates in one intermediate image are (1, 1, 1), then its coordinates in another intermediate image are also (1, 1, 1). Therefore, the confidence level of the pixel at coordinates (1, 1, 1) in the test image can be determined based on the confidence levels of the pixels at coordinates (1, 1, 1) in the two intermediate images. In this way, determining the confidence level of the target location pixels in the test image based on the aligned intermediate images improves the accuracy of the target location confidence level determination. By determining the confidence levels of pixels at different target locations separately, the accuracy of the pixel confidence level determination in the test image can be guaranteed, thereby improving the detection accuracy of the detection model.
[0045] For example, the confidence level of the image under test can be determined by taking a weighted average of the confidence levels of multiple intermediate images. If the confidence level of a pixel at the target location in one intermediate image is 0.6 and the confidence level of a pixel at the target location in another intermediate image is 0.8, then the confidence level of the image under test can be determined to be 0.7.
[0046] Please see Figure 3 In some implementations, the detection method further includes:
[0047] Step 014: Determine the target region based on the confidence level of each pixel in the image to be tested and the preset confidence threshold.
[0048] Specifically, to determine the pixel type based on its confidence level, a preset confidence threshold is set. By comparing the confidence level of each pixel in the test image with the preset confidence threshold, the type of each pixel is determined, thus identifying the target region and improving the accuracy of target region recognition. For example, if the preset type is text, the target region is a text region, and the preset confidence threshold is 0.6, pixels in the test image with a confidence level below 0.6 are identified as non-text, while pixels with a confidence level of 0.6 or higher are identified as text. After determining the type of each pixel in the test image, text and non-text regions in the prediction image can be distinguished.
[0049] Optionally, the preset type can also be other types, such as the preset target type (e.g., face, vehicle, etc.), which is not limited here.
[0050] Please see Figure 4 and Figure 5 In some implementations, step 012: inputting intermediate images into a preset detection model to determine the detection result for each intermediate image, including:
[0051] Step 0122: Downsample the intermediate image to obtain the first feature vector;
[0052] Step 0123: Upsample the first feature vector to generate the second feature vector;
[0053] Step 0124: Output the confidence level of each pixel in the intermediate image based on the first feature vector and the second feature vector.
[0054] Specifically, after the intermediate image is input into the preset detection model, the preset detection model needs to determine the confidence level of each pixel in the intermediate image. First, the downsampling module M1 in the detection model downsamples the intermediate image to generate a first feature vector. Then, the upsampling module M2 in the detection model upsamples the first feature vector to generate a second feature vector. After obtaining the first and second feature vectors, the first and second feature vectors are fused to output the confidence level of each pixel in the intermediate image, thereby determining the detection result of the intermediate image.
[0055] Downsampling can be understood as reducing the resolution of the intermediate image. During downsampling of the intermediate image, the detection model contains multiple downsampling modules M1. These modules downsample the intermediate image multiple times, for example, six times, to gradually reduce its size and obtain a low-resolution first feature map. The first feature vector is the feature vector corresponding to this first feature map. The detection model also includes multiple upsampling modules M2, which gradually restore the low-resolution first feature map to a high-resolution second feature map of the same size as the intermediate image through multiple upsampling steps. The second feature vector is the feature vector corresponding to this second feature map.
[0056] In this way, by sequentially upsampling and downsampling the intermediate image, the confidence level of each pixel in the intermediate image as a pixel of a preset type is determined, thereby realizing the detection of each pixel in the intermediate image.
[0057] Please see Figure 4 , Figure 6 and Figure 7 In some implementations, the detection method further includes:
[0058] Step 015: Obtain the detection model.
[0059] Step 015: Obtain the detection model, including:
[0060] Step 0151: Label pixels of a preset type in multiple original images to generate multiple training samples:
[0061] Step 0152: Input training samples into the initial model to output detection information;
[0062] Step 0153: Calculate the loss value based on the detection information and the preset loss function;
[0063] Step 0154: Adjust the parameters of the initial model based on the loss value to generate a converged detection model.
[0064] Specifically, to improve the detection accuracy of the detection model, it needs to be trained. During training, pixels of a preset type in multiple original images are first labeled to generate multiple training samples. These training samples are then input into the initial model to output detection information. Next, the loss value is calculated based on the detection information and a preset loss function. The loss value is a commonly used metric for evaluating the segmentation performance of a model; it can be understood as the ratio of the number of overlapping pixels between the detected information and the training samples to the total number of pixels in the training samples. Therefore, the loss value can be used to evaluate the similarity between the detected information and the training samples. The smaller the loss value, the closer the detection information obtained after the initial model detects the training samples is to the training samples; that is, the higher the similarity, the more accurate the detection of the training samples. Therefore, after calculating the loss value, the corresponding parameters in the initial model can be adjusted based on the loss value for each training sample to generate a converged detection model. This ensures that the detection model has high detection accuracy for each training sample, and that the loss value between each training sample and the detection information obtained after detection by the model is less than the preset loss value. Before performing detection on the image to be tested, a converged detection model needs to be obtained to improve the detection accuracy of the image. Thus, by training the initial model and selectively adjusting its parameters, a converged detection model is generated, thereby improving the detection accuracy of the detection model on different training samples, and ultimately improving the detection accuracy of the image to be tested.
[0065] Furthermore, when labeling pixels of a preset type in multiple original images, manual labeling can be used to improve the flexibility of labeling. Alternatively, algorithmic labeling can be used. When using algorithms for labeling, the color levels, contrast, brightness, and thresholding of the original images can be adjusted to label the targets in the original images as pixels of a preset type, thereby improving the efficiency of labeling.
[0066] Furthermore, a test set is specifically set up to verify the training effect of the detection model, and a validation set is set up to assist in the initial model training. After labeling multiple original images, a test set for testing the detection model and a validation set for verifying the detection model are generated. Multiple training samples form the training set, while the samples in the test set and validation set are different from the training samples in the training set. After obtaining the converged detection model, it can be tested against the test set to verify the training effect of the detection model. The test set is used to test the accuracy of the model. Applying the test set to the detection model trained on the training set will yield a score for the detection model. The role of the test set is reflected in the testing process, used to evaluate the generalization ability of the detection model, but it cannot be used as the basis for algorithm-related selections such as parameter tuning and feature selection. Generalization ability refers to the detection model's ability to adapt to new samples. During the training process, the validation set can be used to assist in parameter tuning, feature selection, and other algorithm-related selections. The role of the validation set is reflected in the training process and can be used to check whether the initial model training effect is deteriorating. For example, by observing the changes in the loss values of the training and validation sets, it can be seen whether the initial model is overfitting. If it is overfitting, training can be stopped in time, and the initial model structure and hyperparameters can be adjusted, which greatly saves time.
[0067] Please see Figure 8 In some implementations, step 0152: inputting training samples into the initial model to output detection information includes:
[0068] Step 01521: Perform multiple downsampling operations on the training samples to obtain multiple third feature vectors at different scales;
[0069] Step 01522: Perform multiple upsamplings on the third feature vector obtained from the last downsampling to obtain multiple fourth feature vectors at different scales;
[0070] Step 01523: Determine the detection information based on the third and fourth feature vectors of the same scale.
[0071] Specifically, the detection information includes the confidence score of each pixel in the training samples that is of a preset type. After inputting the training samples into the initial model, the initial model also needs to determine the confidence score of each pixel in the training samples. Since pixels of the preset type can have different scales, if the detection model can only adapt to a few scales of pixels, or even just one scale, it will have significant detection limitations. Therefore, during training, it is also necessary to determine the confidence score of each pixel at different scales. When training the initial model, it performs multiple downsampling operations on the training samples to obtain multiple third feature vectors at different scales. Then, it performs multiple upsampling operations on the third feature vector obtained from the last downsampling to obtain multiple fourth feature vectors at different scales. Finally, the detection information, i.e., the confidence score of pixels at multiple scales, is determined based on the third and fourth feature vectors at the same scale. In this way, by obtaining the third and fourth feature vectors at different scales, the confidence score of pixels at different scales is determined, allowing the parameters of the initial model to be adjusted according to the confidence scores at different scales, thus enabling the generated detection model to adapt to pixels at different scales.
[0072] Please see Figure 9 In some implementations, the loss function includes multiple functions, each corresponding to a scale. Step 0153: Calculate the loss value based on the detection information and the preset loss function, including:
[0073] Step 01531: Calculate the loss value of each loss function based on the detection information and loss function corresponding to the same scale;
[0074] The parameters include the weights of each loss function. Step 0154: Adjust the parameters of the initial model based on the loss values, including:
[0075] Step 01541: Adjust the weight of each loss function according to the loss value of each loss function.
[0076] Specifically, there are multiple loss functions, each corresponding to a specific scale. The parameters of the initial model include the weights of each loss function. After acquiring detection information at different scales, the loss value of each loss function is calculated based on the detection information and loss function corresponding to the same scale. The weights of each loss function in the initial model are then adjusted according to its corresponding loss value. Therefore, the formula can be used... To calculate the loss value for each loss function, where L is the loss function, n is the number of scales, and w c T represents the weights of the loss function at the c-th scale. c F represents the type of the pixel in the third feature vector at the c-th scale. c Let be the type of the pixel in the fourth feature vector at the c-th scale.
[0077] Furthermore, to ensure the detection accuracy of the generated detection model, a requirement is placed on the loss value obtained during the initial model training. Only initial models with loss values less than or equal to a preset loss value will be output as detection models. Therefore, after adjusting the weights of each loss function, the training samples are input again into the initial model with adjusted loss function weights to obtain detection information at multiple scales, and the loss value of each loss function is calculated. If the loss values of all loss functions are less than or equal to the preset loss value, there is no need to adjust the weights of the loss functions, and the initial model can be output as a detection model. If the loss value of any loss function is greater than the preset loss value, the corresponding weights are adjusted again, and the initial model training process is repeated until the loss values of all loss functions are less than or equal to the preset loss value.
[0078] In this way, by calculating the loss values of loss functions at different scales and adjusting the weights of each loss function in a targeted manner, the loss value of the final generated detection model is small for pixels at each scale. That is, the final generated detection model can adapt to pixels at multiple scales, thereby improving the detection accuracy of the detection model, providing a more accurate region of interest for the detection of the image to be tested, and enabling the detection model to meet the detection requirements of most images to be tested.
[0079] Please see Figure 4 and Figure 10 In some implementations, step 01521 involves downsampling the training samples multiple times to obtain multiple third feature vectors at different scales, including:
[0080] Step 015211: Perform convolution processing on the third feature vector obtained from the current downsampling to generate the fifth feature vector;
[0081] Step 015212: Fuse the fifth feature vector and the third feature vector to generate the input feature vector;
[0082] Step 015213: Perform the next downsampling on the input feature vector to obtain the third feature vector corresponding to the next downsampling.
[0083] Specifically, when the initial model downsamples the training samples, the third feature vector after downsampling suffers information loss compared to the training samples. This loss occurs with each downsampling iteration and is carried over to the next, resulting in a significant information loss in the final third feature vector after multiple downsampling iterations, thus affecting the detection accuracy of the model. To minimize information loss during downsampling, a residual module S1 is added to the initial model. The residual module S1 performs convolution on the currently downsampled third feature vector to generate a fifth feature vector. Then, the fifth and third feature vectors are fused to generate the input feature vector, which now includes both the fifth and third feature vectors, resulting in a more complete information set. For example, by adding a residual module S1 to downsampling modules M11 and M12, when the third feature vector of downsampling module M11 is input to downsampling module M12, the residual module S1 convolves the third feature vector of downsampling module M11 to obtain the fifth feature vector of downsampling module M11. Then, the third and fifth feature vectors of downsampling module M11 are fused to form the input feature vector of downsampling module M11, which is then input into downsampling module M12. In this way, the residual module S1 can ensure the accuracy of the obtained third feature vector and the input feature vector input to downsampling module M1 as much as possible. Furthermore, to ensure the integrity of the pixels in the feature vector after each downsampling, the residual module S1 is used to process the third feature vector after each generation.
[0084] Specifically, this application uses the U-Net network to detect training samples. After extracting the detection portion of the network, it becomes clear that this portion is highly similar to the VGG-16 network. Therefore, a residual module S1 from the VGG-16 network can be added to the U-Net network to reduce the impact of increasing downsampling depth on detection accuracy. Alternatively, other networks or their residual modules S1 can be used to detect training samples, as long as the goal is to maintain detection accuracy while increasing downsampling depth.
[0085] Thus, by adding the residual module S1, after obtaining the third feature vector, a fifth feature vector with the same depth as the third feature vector can be generated. The third and fifth feature vectors are then fused to generate a more complete input feature vector for the pixel, ensuring the accuracy of the third feature vector obtained by the downconvolution. This improves the accuracy of the corresponding fourth feature vector and detection information, thereby ensuring the accuracy of the initial model's detection while increasing the depth of downsampling.
[0086] Please see Figure 4 and Figure 11 In some implementations, the detection method further includes:
[0087] Step 016: Randomly modify one or more dimensions of the third feature vector and / or the fourth feature vector to preset values.
[0088] Specifically, the types of images to be tested are diverse, thus requiring a high level of generalization ability from the detection model. However, the initial model has too many parameters and too few training samples, making it prone to overfitting. Overfitting manifests as a small loss value and high prediction accuracy on the training set, but a large loss function and low prediction accuracy on the validation set, thereby reducing the generalization ability of the trained detection model. To enhance the generalization ability of the initial model during training, a dropout layer S2 is added to the segmentation network. The dropout layer S2 is a structure used to reduce overfitting in the segmentation network. It randomly modifies one or more dimensions of the third and / or fourth feature vectors to preset values. This can be understood as randomly and temporarily removing one or more dimensions of the third and / or fourth feature vectors. For example, by adding a dropout layer S2 to the upsampling modules M21 and M22, the dropout layer S2 randomly modifies one or more dimensions of the fourth feature vector of the upsampling module M21 to preset values and inputs the modified fourth feature vector to the upsampling module M22. This allows the third and fourth feature vectors of the initial model to be randomized in each test, increasing the diversity of the test samples of the initial model and improving the generalization ability of the initial model, thereby ensuring that the detection model has a high generalization ability.
[0089] Please see Figure 12 In some implementations, before inputting training samples into the initial model to output detection information, step 015: obtaining the detection model includes:
[0090] Step 0155: Calculate the minimum distance between each first pixel and other second pixels in the training sample, where the first pixel is a pixel other than the pixel of the preset type, and the second pixel is a pixel of the preset type;
[0091] Step 0156: Determine the first pixel whose minimum distance is less than the preset distance threshold as the second pixel.
[0092] Specifically, some pixels in the training samples corresponding to the detection target may be incorrectly labeled as pixels other than the preset type, causing the generated detection model to be unable to correctly identify the type of the pixel, thus affecting the detection accuracy of the detection model. Therefore, it is necessary to ensure the integrity of pixels of the preset type in the training samples when training the initial model.
[0093] First, the first pixel in the training sample is determined to be a pixel other than a pixel of a preset type, and the second pixel is a pixel of a preset type. After acquiring the training sample, the distance between each first pixel and other second pixels in the training sample is calculated to determine the minimum distance between the first pixel and other second pixels. Then, the first pixel whose minimum distance is less than a preset distance threshold (such as a preset distance threshold of one or two pixels) is determined as the second pixel, so that all pixels in the region of interest of the training sample are accurately labeled as the second pixel.
[0094] For example, the minimum distance between the first pixel and other second pixels can be obtained using the formula D(p) = in(dist(p,q)), thus obtaining the distance-transformed training samples. Here, Min is the minimum value determination algorithm, dist is the algorithm for calculating the distance between the first pixel and other second pixels, p is the first pixel, and q is the second pixel. The formula... The pixels are processed to relabel the first pixel whose minimum distance is less than a preset distance threshold as the second pixel (i.e., all pixels with a minimum distance of 0 are considered second pixels). Here, T is the preset distance threshold, and D is the minimum distance between the first and second pixels. This allows the relabeling of the second pixel based on the distance between the first and second pixels, ensuring that all pixels within the contour of the region of interest (ROI) are of the preset type. This enables the generated detection model to better detect pixels of the preset type in the test image, resulting in more accurate ROI detection and indirectly improving the detection performance.
[0095] Furthermore, images are often affected by noise interference from imaging equipment and the external environment during digitization and transmission. This noise can interfere with the image and thus affect the detection accuracy of the detection model. Therefore, before inputting the original image into the image processing module, it is necessary to perform denoising on the original image. First, a Gaussian function is used to transform the distance-transformed training samples into a probability map. Then, a Gaussian filter is applied to the probability map to obtain the final probability image, and the training samples are then denoised based on the final probability image. The Gaussian function is... G represents the transformed probability, and σ represents the preset standard deviation. When performing Gaussian filtering on the probability plot, the formula is used. To obtain the final probability image after Gaussian filtering, where x and y are the coordinates of pixels in the probability image, σx and σ y The standard deviations of the x-axis and y-axis are preset. Thus, after two probability transformations, the training samples can be denoised using the final probability image after Gaussian filtering, reducing the impact of noise on the image and improving the training accuracy of the initial model.
[0096] Please see Figure 13 In some implementations, the detection method further includes:
[0097] Step 017: Perform adaptive spatial sampling on the third feature vector obtained from the last downsampling of the training samples to obtain the sixth feature vector;
[0098] Step 018: Perform dimensionality reduction on the sixth eigenvector and the third eigenvector obtained from the last downsampling to obtain the first probability matrix and the second probability matrix, respectively;
[0099] Step 019: Calculate the similarity between the first probability matrix and the second probability matrix, and use it as the similarity between the training samples;
[0100] Step 020: Adjust the number of training samples within multiple preset similarity ranges according to the similarity corresponding to each training sample, so that the difference between the number of training samples within any two preset similarity ranges is less than a preset difference.
[0101] Step 021: Retrain the detection model using the adjusted number of training samples until convergence.
[0102] Specifically, to further improve the training accuracy of the initial model, the number of training samples will be adjusted to ensure a more uniform distribution. Therefore, adaptive spatial sampling is first performed on the third feature vector obtained from the last downsampling of the training samples to obtain the sixth feature vector. For example, let F be the third feature vector obtained from the last downsampling. e The sixth feature vector after adaptive spatial sampling is F d The sixth feature vector after adaptive spatial sampling can be obtained as F according to the following formula. d :F d (nx,ny)=βF e (nx′,ny′)+(1-)βF e (nx ′ +1,ny ′ )+α(1-)F e (nx ′ ,y ′ +1)+(1-)(1-)F e (nx ′ +1,ny ′+1), where the number of horizontal samples is nx, the number of vertical samples is ny, α and β are the corresponding weights, and (nx′, ny′) is the coordinate in the third feature vector during sampling. In the formula, (nx, ny) is the coordinate in the sixth feature vector corresponding to (nx′, ny′). Thus, during sampling, interpolation is performed based on the features sampled in the third feature vector and the features near the sampled features to obtain the corresponding features in the sixth feature vector. For example, interpolation can be performed based on the features to the right of the sampled feature in the third feature vector, the features above the sampled pixel, and the features at the upper right corner of the sampled feature. The interpolated features are then used as the corresponding features in the sixth feature vector.
[0103] After obtaining the sixth eigenvector, both the sixth eigenvector and the third eigenvector obtained from the last downsampling are subjected to dimensionality reduction processing to obtain the first probability matrix and the second probability matrix. In this application, the Euclidean distance from the j-th pixel to the i-th pixel of the third eigenvector is converted into probability using a Gaussian kernel, i.e., using the formula... To confirm the first probability matrix, where k is the total number of pixels in the third eigenvector, σ i Let x be the preset standard deviation of the i-th pixel, and x be the coordinates of the corresponding pixel. Then, the Euclidean distance from the j-th pixel to the i-th pixel in the dimensionality-reduced sixth feature vector is converted into a probability using a Gaussian kernel. That is, the formula is used... To confirm the second probability matrix, where k is the total number of pixels in the sixth feature vector and y is the coordinate of the corresponding pixel.
[0104] After obtaining two probability matrices, the similarity between the first and second probability matrices is calculated to represent the similarity of corresponding samples. In this application, the similarity is evaluated by applying KL divergence to the first and second probability matrices. The formula is then used. The similarity between the first probability matrix and the second probability matrix is calculated, where C is the similarity score and KL is a preset threshold. The greater the difference between the first probability matrix and the second probability matrix, the larger the C value will be. Thus, the similarity between the two can be determined based on the C value.
[0105] After obtaining the similarity between the two training samples, the number of training samples within multiple preset similarity ranges is adjusted based on the similarity of each training sample, ensuring that the difference in the number of training samples within any two preset similarity ranges is less than a preset difference. Finally, the adjusted number of training samples is input into the trained initial model to allow the initial model to converge. At this point, the initial model can be confirmed as the detection model.
[0106] Thus, by reducing the dimensionality of the third feature vector obtained from the last downsampling and the sixth feature vector obtained by adaptively space-adapting the third feature vector, the corresponding probability matrices are obtained. Then, the similarity between the two probability matrices is calculated, and the number of training samples is adjusted according to the similarity to make the number of training samples evenly distributed. After inputting the evenly distributed training samples into the initial model, a convergent detection model can be generated, which is beneficial to improving the generalization ability of the detection model.
[0107] Please see Figure 14 To facilitate better implementation of the detection method described in this application, this application also provides a detection apparatus 10. The detection apparatus 10 includes an enhancement module 11, an input module 12, and a determination module 13. The enhancement module 11 performs data enhancement on the image to be tested to generate multiple intermediate images. The input module 12 inputs the intermediate images into a preset detection model to determine the detection result of each intermediate image. The determination module 13 determines the detection result of the image to be tested based on the multiple detection results.
[0108] The input module 12 is specifically used to input intermediate images into a preset detection model in order to determine the confidence level of each pixel in each intermediate image.
[0109] The determination module 13 is specifically used to align multiple intermediate objects and determine the confidence level of the pixels at the target position in the image under test based on the confidence level of the pixels at the target position in the multiple aligned intermediate images.
[0110] The determination module 13 is specifically used to determine the target region based on the confidence level of each pixel in the image to be tested and a preset confidence threshold.
[0111] The input module 12 is specifically used to downsample the intermediate image to obtain a first feature vector; upsample the first feature vector to generate a second feature vector; and output the confidence score of each pixel of the intermediate image based on the first feature vector and the second feature vector.
[0112] The detection device 10 also includes an acquisition module 14, a labeling module 15, a calculation module 16, and an adjustment module 17.
[0113] The acquisition module 14 is used to acquire the detection model.
[0114] The annotation module 15 is used to annotate pixels of a preset type in multiple original images to generate multiple training samples.
[0115] The input module 12 is specifically used to input training samples into the initial model in order to output detection information.
[0116] The calculation module 16 is used to calculate the loss value based on the detection information and the preset loss function.
[0117] The adjustment module 17 is used to adjust the parameters of the initial model based on the loss value to generate a converged detection model.
[0118] The input module 12 is specifically used to downsample the training samples multiple times to obtain multiple third feature vectors at different scales; to upsample the third feature vector obtained from the last downsampling multiple times to obtain multiple fourth feature vectors at different scales; and to determine the detection information based on the third and fourth feature vectors at the same scale.
[0119] The calculation module 16 is specifically used to calculate the loss value of each loss function based on the detection information and loss function corresponding to the same scale.
[0120] The adjustment module 17 is specifically used to adjust the weight of each loss function according to the loss value of each loss function.
[0121] The input module 12 is specifically used to perform convolution processing on the third feature vector obtained by the current downsampling to generate the fifth feature vector; to fuse the fifth feature vector and the third feature vector to generate the input feature vector; and to perform the next downsampling on the input feature vector to obtain the third feature vector corresponding to the next downsampling.
[0122] The detection device 10 also includes a modification module 18.
[0123] Modification module 18 is specifically used to randomly modify one or more dimensions of the third feature vector and / or the fourth feature vector to preset values.
[0124] The calculation module 16 is specifically used to calculate the minimum distance between each first pixel and other second pixels in the training sample. The first pixel is a pixel other than a pixel of a preset type, and the second pixel is a pixel of a preset type.
[0125] The determination module 13 is specifically used to determine the first pixel whose minimum distance is less than a preset distance threshold as the second pixel.
[0126] The input module 12 is specifically used to perform adaptive spatial sampling on the third feature vector obtained by the last downsampling of the training sample to obtain the sixth feature vector; and to perform dimensionality reduction processing on the sixth feature vector and the third feature vector obtained by the last downsampling to obtain the first probability matrix and the second probability matrix respectively.
[0127] The calculation module 16 is specifically used to calculate the similarity between the first probability matrix and the second probability matrix, so as to use the similarity between the training samples.
[0128] The adjustment module 17 is specifically used to adjust the number of training samples within multiple preset similarity ranges according to the similarity corresponding to each training sample, so that the difference between the number of training samples within any two preset similarity ranges is less than a preset difference.
[0129] The detection device 10 also includes a training module 19.
[0130] Training module 19 is used to retrain the detection model until convergence based on multiple training samples with adjusted numbers.
[0131] Please see Figure 15 The detection device 100 of this application includes a processor 20, which is used to perform data enhancement on the image to be tested to generate multiple intermediate images; input the intermediate images to a preset detection model to determine the detection result of each intermediate image; and determine the detection result of the image to be tested based on the multiple detection results.
[0132] Please see Figure 16 This application also provides a non-volatile computer-readable storage medium 200 storing a computer program 210. When the computer program 210 is executed by the processor 20, it implements the steps of the detection method of any of the above embodiments. For the sake of brevity, these steps will not be repeated here.
[0133] In the description of this specification, the references to terms such as "some embodiments," "in one example," "exemplarily," etc., indicate that a specific feature, structure, material, or characteristic described in connection with an embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0134] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the function involved, as will be understood by those skilled in the art to which embodiments of this application pertain.
[0135] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.
Claims
1. A method of detection, characterized in that, include: Data augmentation is performed on the image to be tested to generate multiple intermediate images; The intermediate images are input into a preset detection model to determine the detection result for each intermediate image; Based on the multiple detection results, the detection result of the image to be tested is determined; It also includes acquiring the detection model; The step of obtaining the detection model includes: labeling pixels of a preset type in multiple original images to generate multiple training samples; inputting the training samples into an initial model to output detection information; calculating a loss value based on the detection information and a preset loss function; and adjusting the parameters of the initial model based on the loss value to generate the converged detection model. The input of the training samples into the initial model to output detection information includes: The training samples are downsampled multiple times to obtain multiple third feature vectors at different scales; the third feature vectors obtained from the last downsampling are upsampled multiple times to obtain multiple fourth feature vectors at different scales; the detection information corresponding to each scale is determined based on the third feature vectors and fourth feature vectors corresponding to each scale. It also includes: sampling the third feature vector obtained by the last downsampling of the training sample, and performing weighted interpolation based on the features corresponding to the sampling position, the features corresponding to the right position of the sampling position, the features corresponding to the upper position of the sampling position, and the features corresponding to the upper right position of the sampling position to obtain the sixth feature vector; The sixth feature vector and the third feature vector obtained from the last downsampling are respectively subjected to dimensionality reduction processing to obtain the first probability matrix and the second probability matrix. Calculate the similarity between the first probability matrix and the second probability matrix, and use it as the similarity between the training samples. The number of training samples within multiple preset similarity ranges is adjusted according to the similarity corresponding to each training sample, so that the difference between the number of training samples within any two preset similarity ranges is less than a preset difference. The detection model is retrained using multiple training samples with adjusted quantities until convergence.
2. The detection method according to claim 1, characterized in that, The data augmentation includes at least one of rotation, reflection, mirroring, flipping, and scaling.
3. The detection method according to claim 1, characterized in that, The detection result includes the confidence score for each pixel as a preset type. The input of the intermediate image to a preset detection model to determine the detection result for each intermediate image includes: The intermediate image is input into a preset detection model to determine the confidence level of each pixel in each intermediate image; Determining the detection result of the image to be tested based on multiple detection results includes: Align the multiple intermediate images; The confidence level of the pixel at the target location in the image to be tested is determined based on the confidence level of the pixel at the target location in the multiple aligned intermediate images.
4. The detection method according to claim 3, characterized in that, Also includes: The target region is determined based on the confidence level of each pixel in the image under test and a preset confidence threshold.
5. The detection method according to claim 1, characterized in that, The detection result includes the confidence score for each pixel as a preset type. The input of the intermediate image to a preset detection model to determine the detection result for each intermediate image includes: The intermediate image is downsampled to obtain a first feature vector; The first feature vector is upsampled to generate a second feature vector; The confidence level of each pixel in the intermediate image is output based on the first feature vector and the second feature vector.
6. The detection method according to claim 1, characterized in that, The loss function includes multiple functions, each corresponding to a scale. The step of calculating the loss value based on the detection information and the preset loss function includes: Based on the detection information and the loss function corresponding to the same scale, calculate the loss value of each loss function; The parameters include the weights of each loss function, and adjusting the parameters of the initial model based on the loss values includes: The weights of each loss function are adjusted based on the loss value of each loss function.
7. The detection method according to claim 1, characterized in that, The step of downsampling the training samples multiple times to obtain multiple third feature vectors at different scales includes: The third feature vector obtained by the current downsampling is convolved to generate the fifth feature vector; The fifth feature vector and the third feature vector are fused to generate the input feature vector; The input feature vector is downsampled again to obtain the third feature vector corresponding to the next downsampling.
8. The detection method according to claim 1, characterized in that, Also includes: Randomly modify one or more dimensions of the third feature vector and / or the fourth feature vector to preset values.
9. The detection method according to claim 1, characterized in that, Before inputting training samples into the initial model to output detection information, obtaining the detection model further includes: Calculate the minimum distance between each first pixel and other second pixels in the training sample, where the first pixel is a pixel other than a pixel of a preset type, and the second pixel is a pixel of a preset type; The first pixel whose minimum distance is less than a preset distance threshold is determined as the second pixel.
10. A detection device, characterized in that, For implementing the method as described in any one of claims 1-9, comprising: The enhancement module is used to perform data enhancement on the image under test to generate multiple intermediate images; An input module is used to input the intermediate images into a preset detection model to determine the detection result for each intermediate image; and The determination module is used to determine the detection result of the image to be tested based on multiple detection results.
11. A testing device, characterized in that, For implementing the method as described in any one of claims 1-9, a processor is included, the processor being configured to perform data augmentation on the image to be tested to generate a plurality of intermediate images; The intermediate images are input into a preset detection model to determine the detection result for each intermediate image; Based on the multiple detection results, the detection result of the image to be tested is determined.
12. A non-volatile computer-readable storage medium comprising a computer program, wherein when executed by a processor, the computer program causes the processor to perform the detection method according to any one of claims 1-9.