Training image processing method, identification method, related equipment and storage medium
By performing specific processing on the training samples of the semantic segmentation model, ensuring the balance of positive samples and the diversity of negative samples, the problem of limited diversification ability of training samples in the existing technology is solved, and the training accuracy and robustness of the model are improved.
Patent Information
- Application Number
- CN202510093907.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-06-13
AI Technical Summary
In the prior art, the training sample diversification ability of the semantic segmentation model is limited, which affects the training accuracy and robustness of the model.
By processing positive and negative samples separately, and using methods such as sampling frequency calculation and defect proportion calculation, we ensure the balance of positive samples and the diversity of negative samples in the training data.
The training accuracy and robustness of the semantic segmentation model are improved, and the recognition accuracy of semantic segmentation tasks is enhanced.
Smart Images

Figure CN120147629A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technologies, and in particular, to a method and device for processing training images, a method and device for recognition, an electronic device, and a storage medium. Background Art
[0002] In the implementation of semantic segmentation tasks using semantic segmentation models, the training robustness and accuracy of the semantic segmentation models will affect the recognition accuracy of semantic segmentation tasks. It can be understood that the more precise and robust the semantic segmentation model is trained, the more accurate the recognition of semantic segmentation tasks will be. The diversification of training samples will ensure the training precision and robustness of the semantic segmentation model. In practical applications, the collection of training samples, such as positive samples and negative samples, is limited. Under limited training samples, the expansion of training samples is achieved by at least one of the following methods: data augmentation, data normalization and standardization, image stitching, and hybrid augmentation, so as to make the training samples diverse. The above methods are conventional means to make training samples diverse, and the diversification ability is limited. There is an urgent need for a solution that can make the diversification ability of training samples stronger. Summary of the Invention
[0003] This application provides a method and device for processing training images, a method and device for recognition, an electronic device, and a storage medium, so as to at least solve the above technical problems existing in the prior art.
[0004] According to the first aspect of this application, a method for processing training images is provided. The method includes:
[0005] Obtain positive sample images and negative sample images in the training images; perform a first process on the positive sample images to obtain target positive sample images; perform a second process on the negative sample images to obtain target negative sample images; based on the target positive sample images and the target negative sample images, obtain target training images, where the target training images are used to train a semantic segmentation model to be trained, and the trained semantic segmentation model is used to identify whether an image to be recognized is abnormal.
[0006] According to the second aspect of this application, a recognition method is provided, including: obtaining an image to be recognized; inputting the image to be recognized into the semantic segmentation model to obtain a recognition result of whether the image to be recognized is abnormal; where the semantic segmentation model is obtained by training the semantic segmentation model to be trained using the foregoing target training images.
[0007] According to the third aspect of this application, a device for processing training images is provided, including:
[0008] A first acquisition unit for acquiring positive sample images and negative sample images in training images; a first processing unit for performing a first processing on the positive sample images to obtain target positive sample images; a second processing unit for performing a second processing on the negative sample images to obtain target negative sample images; a second acquisition unit for obtaining target training images based on the target positive sample images and the target negative sample images, where the target training images are used to train a semantic segmentation model to be trained, and the trained semantic segmentation model is used to identify whether an abnormal situation exists in an image to be identified.
[0009] According to a fourth aspect of the present application, there is provided an identification device, the device including: a first obtaining unit for obtaining an image to be identified; a second obtaining unit for inputting the image to be identified into a semantic segmentation model to obtain an identification result of whether an abnormal situation exists in the image to be identified; wherein, the semantic segmentation model is obtained by training the semantic segmentation model to be trained with the foregoing target training images.
[0010] According to a fifth aspect of the present application, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method of the present application.
[0011] According to a sixth aspect of the present application, there is provided a non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are used to cause a computer to execute the method of the present application.
[0012] In the present application, the positive and negative samples are processed separately. In this way, the balance of the positive samples and the diversity of the negative samples in the training data used for training the semantic segmentation model to be trained can be ensured. Using such training data with balanced positive samples and diverse negative samples to train the semantic segmentation model can improve the accuracy and robustness of the semantic segmentation model. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present application will become easily understandable. In the drawings, several embodiments of the present application are shown in an exemplary rather than restrictive manner, where:
[0014] In the drawings, the same or corresponding reference numerals represent the same or corresponding parts.
[0015] Figure 1 Shows the implementation process schematic of the processing method of the training image in the embodiment of the present application Figure 1 ;
[0016] Figure 2 Shows the implementation effect diagram of the random cropping method in the embodiment of the present application;
[0017] Figure 3 Shows the schematic diagram of the implementation process of the random cropping method in the embodiment of the present application;
[0018] Figure 4 Shows the implementation effect diagram of the cropping method using a sliding window in the embodiment of the present application;
[0019] Figure 5 Shows the schematic diagrams of two processing methods for the image edge in the embodiment of the present application;
[0020] Figure 6 Shows the implementation effect diagram of the letterbox scheme in the embodiment of the present application;
[0021] Figure 7 Shows the schematic diagram of sample ratio cropping in the embodiment of the present application;
[0022] Figure 8 Shows the first schematic diagram of region expansion outside the region in the embodiment of the present application;
[0023] Figure 9 Shows the second schematic diagram of region expansion outside the region in the embodiment of the present application;
[0024] Figure 10 Shows the schematic diagram of the small defect image in the embodiment of the present application Figure 1 ;
[0025] Figure 11 Shows the schematic diagram of the small defect image in the embodiment of the present application Figure 2 ;
[0026] Figure 12 Shows the schematic diagram of the implementation process of the training method of the semantic segmentation model in the embodiment of the present application;
[0027] Figure 13 Shows the schematic diagram of the implementation process of the recognition method in the embodiment of the present application;
[0028] Figure 14 Shows the schematic diagram of the composition structure of the processing device for training images in the embodiment of the present application;
[0029] Figure 15 Shows the schematic diagram of the composition structure of the recognition device in the embodiment of the present application;
[0030] Figure 16 Shows the schematic diagram of the composition structure of the electronic device in the embodiment of the present application. Detailed implementation manners
[0031] To make the objectives, features, and advantages of this application more obvious and understandable, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of this application.
[0032] In this application, for the positive samples and negative samples in the training samples, by adopting different treatments for the two types of samples, such as performing a first treatment on the positive samples and a second treatment on the negative samples, the diversity of the training samples is improved. This addresses the problem in related technologies where the diversity of training samples is not strong due to not distinguishing between positive and negative samples and not performing targeted treatments on them separately.
[0033] The following further elaborates on the technical solutions of this application.
[0034] Figure 1 Shows the implementation process schematic of the processing method for training images in the embodiments of this application Figure 1 As Figure 1 shown, the method includes:
[0035] S101: Obtain the positive sample images and negative sample images in the training image.
[0036] In this application, the training sample is an image. For convenience of description, the training sample can be referred to as a training image. The training image (training sample or sample image) includes defective images and non-defective images. The defective images (abnormal images) can be used as negative samples (images), and the non-defective images (normal images) can be used as positive samples (images), and vice versa. Unless otherwise specified, the positive samples in this application refer to normal samples or normal image samples, and the negative samples refer to defective samples or defective image samples.
[0037] In practical applications, the training image can be an image that appears in industrial production, manufacturing, etc. For example, a glass image, a camera image. Glass and cameras may have defects such as scratches, dents, protrusions, and stains during the production and manufacturing processes. The training samples can be sample images collected for glass and cameras. The collected training images are divided into defective and non-defective images to obtain positive sample images and negative sample images.
[0038] S102: Perform a first treatment on the positive sample images to obtain target positive sample images.
[0039] S103: Perform a second treatment on the negative sample images to obtain target negative sample images.
[0040] In this application, considering the respective characteristics of positive samples and negative samples: for example, compared with the number of negative sample images, positive sample images are easier to collect. The number of positive samples in practical applications is much larger than that of negative samples. Due to the limited number of negative samples, the various defect situations reflected by the limited number of collected negative samples are particularly important. Based on this, according to the respective characteristics of positive and negative samples, in this application, positive and negative samples are processed separately. For example, the sampling frequency of positive samples is calculated so that each positive sample can be evenly sampled as the final training sample for the semantic segmentation model to be trained. The defect ratio of negative samples is calculated to ensure defect diversity. Among them, the first processing can be the calculation of the sampling frequency of positive samples. The second processing can be the calculation of the defect ratio of negative samples.
[0041] S104: Based on the target positive sample image and the target negative sample image, obtain a target training image, which is used to train the semantic segmentation model to be trained, and the trained semantic segmentation model is used to identify whether there is an abnormality in the image to be identified.
[0042] In this step, the target training image is obtained through two types of target samples (target positive samples and target negative samples) obtained by separately processing positive and negative samples. The target training image is the training data for the semantic segmentation model to be trained. The target training image is input into the semantic segmentation model to be trained, and the model is trained through iterative calculation of the semantic segmentation model to be trained.
[0043] In S101 - S104, because this application takes into account the respective characteristics of positive and negative samples, so according to their respective characteristics, positive and negative samples are processed separately. In this way, the balance of positive samples and the diversity of negative samples in the training data used for training the semantic segmentation model to be trained can be ensured. Using this training data with balanced positive samples and diverse negative samples to train the semantic segmentation model can improve the accuracy and robustness of the semantic segmentation model.
[0044] In practical applications, the training data used for training the semantic segmentation model to be trained has a certain image size, and this image size is used as the expected image size. In this application, positive and negative samples and the target positive samples themselves also have a certain image size. Based on this, in some embodiments, if the image size of the target positive sample image and / or the target negative sample image matches the expected image size of the semantic segmentation model, that is, the width and height of the target positive sample image and / or the target negative sample image are the same as or consistent with the width and height of the image required by the semantic segmentation model, then the target positive sample image and the target negative sample image are used as the target training image.
[0045] In some embodiments, if the image size of the target positive sample image and / or the target negative sample image does not match the expected image size of the semantic segmentation model, that is, the width of the image of at least one of the two target samples is different from the width of the image required by the semantic segmentation model, and / or the height of the image of at least one of the target samples is different from the height of the image required by the semantic segmentation model, then the target positive sample image and / or the target negative sample image is cropped to obtain a target training image. For example, according to the height and width of the image required by the semantic segmentation model, the target positive sample image and / or the target negative sample is cropped, and the cropped image is used as the training data of the semantic segmentation model. In this way, using the training data of the image size required by the semantic segmentation model to train the semantic segmentation model can meet the actual training needs and provide guarantee for the accurate training of the semantic segmentation model.
[0046] In this application, the situation where the image size of the target positive sample image and / or the target negative sample image does not match the expected image size of the semantic segmentation model includes two situations: the sample image size is greater than the expected image size (Situation 1) and the sample image size is less than the expected image size (Situation 2).
[0047] In Situation 1, one of the two methods of random cropping and sliding window cropping can be used to crop the image.
[0048] Random cropping method: If the sample image size of the target positive sample image and / or the target negative sample image is greater than the expected image size, for example, the width of the target sample image is greater than the expected image width (the image width required by the semantic segmentation model) and / or the height of the target sample image is greater than the expected image height (the image height required by the semantic segmentation model), then based on the random coordinate points generated in the target positive sample image and / or the target negative sample image, the target positive sample image and / or the target negative sample image is cropped to the size of the expected image to obtain a new cropped image, and the target training image is obtained based on the overlap degree between the new cropped image and the historical cropped image.
[0049] In this application, for the convenience of description, in the aforementioned random cropping method, the target positive sample image and the target negative sample image are collectively referred to as target samples, and the width and height of the image required by the semantic segmentation model are regarded as the network input size. Combining Figure 2 and Figure 3 as shown, a coordinate point is randomly generated in the current target sample image, and starting from this coordinate point, according to the width and height of the network input size, the image is cropped from the current target sample image to obtain Figure 2The cropped image in it can be used as a new cropped image. In practical applications, the current cropped image can be used as a training image for training a semantic segmentation model. Alternatively, a previously stored cropped image can be obtained, which was cropped randomly in the past. Calculate the overlap degree between the current newly cropped image and the historical cropped image. If the overlap degree between the current newly cropped image and the historical cropped image is greater than or equal to the overlap degree threshold, it means that the overlap degree between the current newly cropped image and the training images that the semantic segmentation model already had in the past is too high, so the current newly cropped image is considered a duplicate image and is not used as a training image for the semantic segmentation model. Another coordinate point is randomly generated again and the above-mentioned scheme is continued. If the overlap degree between the current newly cropped image and the historical cropped image is less than the overlap degree threshold, it means that the current newly cropped image does not overlap with the training images that the semantic segmentation model already had in the past, and the current newly cropped image can be used as a training image for the semantic segmentation model.
[0050] In this application, the calculation of the overlap degree (IOU) between two images can be obtained according to formula (1).
[0051] IOU = Area of the intersection of the regions of the two images / Area of the union of the regions of the two images (1)
[0052] If the overlap degree between the current newly cropped image and the historical cropped image is less than the overlap degree threshold, the current newly cropped image is saved or stored as a historical cropped image.
[0053] Sliding window cropping method: If the sample image size of the target positive sample image and / or the target negative sample image is larger than the desired image size, for example, the width of the sample image is greater than the desired image width (the image width required by the semantic segmentation model) and / or the height of the sample image is greater than the desired image height, then the sliding window cropping method is used to crop the target positive sample image and / or the target negative sample image to obtain the target training image; among them, the sliding step of the sliding window in the sliding window cropping method is less than the size of the sliding window.
[0054] In the sliding window cropping method, set the size of the sliding window to be the same as the network input size and set the sliding step of the sliding window. As shown in Figure 4 In the target sample shown on the left in Figure 4 The sliding window slides sequentially in the order from left to right and from top to bottom of the image, and slides with the set sliding step. Starting from the upper left corner of the target sample image, the sliding window is slid on the image with the set sliding step, and the image area covered by the sliding window is cropped. If the image area covered by the sliding window in the target sample during each slide is regarded as a small image, then multiple small images of the target sample can be cropped by sliding multiple times in the target sample image, as shown in Figure 4 On the right.
[0055] In the sliding window cropping method, the sliding step size can be set to be smaller than the size of the sliding window. In this way, there will be an overlap in two adjacent croppings, that is, there will be the same content in the small images obtained from two adjacent croppings. In this way, the integrity of the features (such as texture, edge, color) of the target sample can be ensured, and the error that may be caused when the feature is cropped to the edge of the sliding window can be reduced. This provides a more accurate training image for the semantic segmentation model.
[0056] It should be noted that no matter whether the random cropping or the sliding window cropping method is adopted, processing is required when cropping to the edge of the image. Refer to Figure 5 As shown, the edge processing method can be: filling the edge with fixed pixels to make its size reach the network input size. Or, translating the edge area towards the image direction so that an image can be exactly and completely cropped and the edge part is discarded.
[0057] In Case 2, if the sample image size of the target positive sample image and / or the target negative sample image is smaller than the expected image size, for example, the width of the image of the target sample is smaller than the expected image width, and / or the height of the image of the target sample is smaller than the expected image height (the image height required by the semantic segmentation model), then based on the sample image size of the target positive sample image and / or the target negative sample image and the expected image size, the first scaling ratio and the second scaling ratio are calculated, where the first scaling ratio is the scaling ratio in the width direction of the target positive sample image and / or the target negative sample image, and the second scaling ratio is the scaling ratio in the height direction of the target positive sample image and / or the target negative sample image; based on the first scaling ratio and the second scaling ratio, the target scaling ratio is obtained; according to the target scaling ratio, the sample image of the target positive sample image and / or the target negative sample image is scaled to obtain a scaled image; based on the scaled image, the target training image is obtained.
[0058] In implementation, Case 2 can be specifically: calculating the ratio between the width of the target sample and the width in the network input size as the width scaling ratio (the first scaling ratio). Calculating the ratio between the height of the target sample and the height in the network input size as the height scaling ratio (the second scaling ratio). Selecting the larger value from the width scaling ratio and the height scaling ratio as the target scaling ratio. According to the target scaling ratio, the target sample image is scaled to obtain a scaled image. The width and height of the scaled image are different from the width and height of the target sample. In the scaled image, subtracting the scaled image size from the network input size to obtain a difference value. For example, the network input size is 256 * 256 pixel points, and the scaled image size is 256 * 230 pixel points, and the added difference value between the two is 26 pixel points. Along the width direction or the height direction of the scaled image, the 26 pixel points are evenly divided. As Figure 6The first and last rows in the right figure each have 13 pixels, and the pixels in these two rows are the pixels to be filled. These pixels are filled with black, white, or gray to obtain the target training image. The foregoing solution is the letterbox solution in this application.
[0059] The above solution is to make the image size of the target sample whose image size does not match the network size conform to the size of the network input, so as to ensure the accurate training of the semantic segmentation model to be trained. The foregoing solution can be regarded as a solution to unify the image size of the target sample to the network input size, which is simply called the unified size cropping solution.
[0060] In this application, since the number of positive samples collected or gathered is usually more than the number of negative samples, and the balance between positive and negative samples in the training data of the semantic segmentation model can ensure the accurate training of the model. In this application, to ensure the relative balance of the number of positive and negative samples input to the semantic segmentation model to be trained, the negative sample image ratio parameter can be obtained in advance by configuring the ratio parameter of the positive and negative sample images. Based on the positive and negative sample image ratio parameter, the target positive sample image, and the target negative sample image, the target training image is obtained. Exemplarily, 200 positive sample images and 100 negative sample images are collected, and the ratio parameter of the positive and negative sample images is configured as 1:2. Sampling of the target positive and negative sample images is performed according to this ratio parameter to ensure that the number of positive and negative samples input to the semantic segmentation model to be trained is 1:1, providing positive and negative samples with a balanced number for the training of the model.
[0061] As Figure 7 shown, 2 small images can be cropped from the defective samples as the positive sample data of the semantic segmentation model, and 1 small image can be cropped from the normal samples as the negative sample data of the semantic segmentation model, so that the ratio of positive and negative data in the training dataset of the semantic segmentation model can reach 1:1 as much as possible, making the positive and negative samples in the training data as balanced as possible. This solution is the sample ratio parameter configuration or setting solution in this application. Through the sample ratio parameter configuration or setting solution, the balance of positive and negative training samples of the semantic segmentation model can be achieved.
[0062] The calculation process of the sampling frequency of positive samples is described below. It can be understood that the sampling frequency calculation only involves positive samples. By calculating the sampling frequency, it is ensured that each positive sample can be sampled and used as a positive training sample that can be input to the semantic segmentation model to be trained.
[0063] In this application, the target positive sample image is obtained by performing M rounds of first processing on the positive sample image. For example, the target positive sample image includes the target positive sample image obtained by calculating the sampling frequency in the first round - the target positive sample obtained by calculating the sampling frequency in the Mth round. Wherein, M is a positive integer greater than or equal to 2. Based on this, the solution of performing the first processing on the positive sample image in the aforementioned S102 to obtain the target positive sample image may include the solution shown in S201 to S204.
[0064] S201: Obtain the initial selection probability of each positive sample image in the ith round.
[0065] In this step, based on the number of times each positive sample image is selected in the (i - 1)th round and the maximum number of times a positive sample image is selected, the initial selection probability of each positive sample image in the ith round is obtained. Wherein, i is a positive integer greater than or equal to 2, and i ≤ M.
[0066] In specific implementation, the calculation of the initial selection probability sample_prob_1 of the positive sample in the ith round can be achieved according to formula (2).
[0067]
[0068] In formula (2), neg_freq is an array used to store the number of times each positive sample is selected in the (i - 1)th round. The max function is used to obtain the maximum number of times a positive sample is selected in the (i - 1)th round. Dividing the number of times each positive sample is selected by the maximum number of selections is equivalent to normalizing the sampling frequency to [0, 1]. Subtracting the normalization result from 1 is equivalent to reversing the normalization result, so that the positive sample with a lower selection frequency in the (i - 1)th round obtains a higher initial selection probability in the ith round, and the positive sample with a higher selection frequency obtains a lower initial selection probability in the ith round. sample_prob_1 is an array used to store the initial selection probability of each positive sample in the ith round.
[0069] S202: Process the initial selection probability of each positive sample image in the ith round to obtain the target selection probability of each positive sample image in the ith round, where the balance of each positive sample image being selected under the target selection probability is stronger than the balance of the positive sample image being selected under the initial selection probability.
[0070] In this step, first, the initial selection probability can be smoothed to obtain the first intermediate probability sample_prob_2 of each positive sample image. In implementation, the calculation of the first intermediate probability sample_prob_2 can be performed according to formula (3).
[0071] sample_prob_2 = sample_prob_1 - median(sample_prob_1) + 1 (3)
[0072] In formula (3), the median function is used to find the median of the array sample_prob_1. The probability distribution obtained through formula (2) is smoothed. To prevent extreme values in the distribution and make the distribution more uniform, the probability distribution obtained through formula (2) is subtracted by the median of the distribution and then added by one. This addition operation not only prevents negative values but also ensures that each positive sample in the entire probability distribution has at least a minimum base probability. sample_prob_2 is an array used to store the first intermediate probability of each positive sample in the i-th round.
[0073] Next, the probability distribution of the first intermediate probability of each positive sample image is adjusted to obtain the second intermediate probability sample_prob_3 of each positive sample image. The second intermediate probability can be calculated according to formula (4).
[0074] sample_prob_4 = sample_prob_3 log(len(sample_prob_3)*b) (4)
[0075] In formula (4), log() represents taking the logarithm of (), and the len function is used to obtain the length of the sample_prob_3 array; b is a constant that can be set according to actual needs. Using formula (4), the probability distribution is further adjusted, and the non-linear transformation log function is applied to raise the probability to the exponent, which makes the probability distribution sharper, such that the positive samples with a low (high) selection frequency in the (i - 1)-th round have an increased (decreased) selection probability in the i-th round. sample_prob_3 is an array used to store the second intermediate probability of each positive sample in the i-th round.
[0076] Subsequently, the second intermediate probability sample_prob_3 of each positive sample image is normalized to obtain the target selection probability sample_prob_4. The target selection probability sample_prob_4 can be calculated according to formula (5).
[0077]
[0078] In formula (5), the sum function is used to calculate the sum of all elements in the array sample_prob_3. The probability distributions of all positive samples are normalized so that their sum equals one. This can increase the probabilities of samples that have not been selected, and decrease the probabilities of samples that have been selected more times, thus enabling each normal sample to be evenly selected. sample_prob_4 is an array used to store the target selection probabilities of each positive sample in the i-th round.
[0079] S203: Based on the target selection probabilities, obtain the selected positive sample images in the i-th round.
[0080] In this step, from all positive sample images, the positive sample images with the top target selection probabilities can be selected as the selected positive sample images in the i-th round.
[0081] S204: Based on the selected positive sample images in the i-th round, obtain the target positive sample images in the i-th round.
[0082] In this step, the selected positive sample images in the i-th round can be used as the target positive sample images in the i-th round.
[0083] It should be noted that in the sampling frequency calculation scheme for the i = 1 round, the array neg_freq is initialized, that is, the initial selection times of each positive sample are initially set, and the calculations of formulas (2) - (5) are performed in sequence according to the initially set selection times of each positive sample to obtain the target selection probabilities of each positive sample in the 1st round, and based on these target selection probabilities, some positive samples are selected as the target positive sample images in the 1st round.
[0084] In this application, i = i + 1 can be used to repeatedly execute the scheme of S201 - S204 to obtain the target positive sample images in each round. The target positive sample images in all rounds are used as the training samples for the semantic segmentation model to be trained. That is, as shown in S201 - S204, through the sequential calculations of formulas (2) - (5), the probabilities of each positive sample being selected as the training positive sample of the semantic segmentation model can be obtained in each round. This scheme can be regarded as a scheme for calculating the sampling frequency of positive samples to ensure that all positive samples have similar or the same selection probabilities to be used as the training positive samples of the semantic segmentation model.
[0085] From the perspective of the semantic segmentation model, the selected positive samples are balanced to avoid the problem of insufficient precision in model training caused by the lack of diversity in training samples due to some positive samples being selected as model training samples multiple times while some positive samples cannot be selected as model training samples or are selected as model training samples too few times.
[0086] The following describes the calculation process of the defect proportion of negative samples. It can be understood that the calculation of the defect proportion only involves negative samples. By calculating the defect proportion of negative samples, the diversity of defects reflected in the negative training samples of the semantic segmentation model to be trained is ensured.
[0087] In this application, the target negative sample image is obtained through N rounds of first processing of the negative sample image. For example, the negative-labeled positive sample image includes the target positive sample image obtained through the defect proportion calculation in the first round - the target negative sample obtained through the defect proportion calculation in the Nth round. Where N is a positive integer greater than or equal to 2. The aforementioned S103 performs a second processing on the negative sample image to obtain the target negative sample image, which may include the solutions shown in S301 to S305.
[0088] S301: Based on the defect annotation information of each negative sample image in the jth round, determine the defect contour in the negative sample image.
[0089] In this step, the defect annotation information is used to annotate the position or boundary of the defect in the negative sample. Along the defect annotation information, the defect in the negative sample image is traced, and thus the defect contour in the negative sample can be determined. The area enclosed by the defect contour in the negative sample is the defect area. Where j is a positive integer greater than or equal to 1, and j ≤ N.
[0090] S302: Expand the defect contour in the negative sample image to obtain an expanded area.
[0091] In this step, the negative sample image is an image with defects. The defects in this image may be one or multiple. When there is one or more defects in the image, for a single defect in the image, based on the maximum point coordinates and minimum point coordinates of the defect contour in the negative sample image, obtain the circumscribed rectangle area of the defect contour; expand the circumscribed rectangle area to the expected image size to obtain the expanded area.
[0092] Generally speaking, in this application, first determine the defect contour in the negative sample image, use the defect contour to make the circumscribed rectangle of the defect, and use the circumscribed rectangle to expand the defect area in the negative sample. The expanded area is the expanded area.
[0093] Taking the upper left corner of the negative sample as the coordinate origin, the right side direction of the negative sample as the positive X-axis direction, and the lower side direction of the negative sample as the positive Y-axis direction, establish a reference coordinate system. The defect contour is composed of multiple points, and each point has its own coordinates (x, y) in the reference coordinate system. Such as Figure 8 and Figure 9The white area is the defect. From all the coordinate points that make up the defect contour, find the maximum and minimum values of x and the maximum and minimum values of y. The point formed by the minimum value of x and the minimum value of y serves as the upper left corner point of the defect circumscribed rectangle, and the point formed by the maximum value of x and the maximum value of y serves as the upper right corner point of the defect circumscribed rectangle. Take the difference between the maximum value of x and the minimum value of x as the width of the defect circumscribed rectangle, and take the difference between the maximum value of y and the minimum value of y as the height of the defect circumscribed rectangle. Based on the upper left corner point, upper right corner point of the rectangle, the width and height of the rectangle, the defect circumscribed rectangle can be drawn, as shown in Figure 8 and Figure 9 the defect circumscribed rectangle shown in the left figure in
[0094] Next, expand the defect circumscribed rectangle to the size of the network input dimension to obtain the expanded area. It can be understood that in the x-axis direction, increase the width of the defect circumscribed rectangle according to the width input_width in the network input dimension. In the y-axis direction, increase the height of the defect circumscribed rectangle according to the height input_height in the network input dimension, to obtain the expanded area as shown in Figure 8 and Figure 9 the left figure in
[0095] Compared with the size of the negative sample image, Figure 8 the left figure shows that the expanded area extends beyond the original negative sample image, Figure 9 the left figure shows that the expanded area does not extend beyond the original image but is located within the original image. The reason for these two expansion situations is, first, the size of the network input dimension, and second, the position of the defect in the original image. The larger the network input dimension, the easier it is to exceed the original image when expanding the defect circumscribed rectangle. The closer the defect is to the edge of the original image, the easier it is to exceed the original image.
[0096] The coordinates of the expanded area obtained by expanding the defect circumscribed rectangle to the size of the network input dimension can be calculated according to the following formula.
[0097] ROI_left = ma(0, x_min - input_width)
[0098] ROI_right = min(x_max + input_width, width)
[0099] ROI_top = max(0, y_min - input_height)
[0100] ROI_bottom = min(y_max - input_height, height)
[0101] Among them, ROI_left and ROI_top represent the x and y coordinates of the upper left corner point of the expanded region; ROI_left and ROI_bottom represent the x and y coordinates of the lower right corner point of the expanded region; x_min and y_min represent the x and y coordinates of the upper left corner point of the bounding rectangle of the defect; x_max and y_max represent the x and y coordinates of the lower right corner point of the bounding rectangle of the defect; input_width and input_height represent the width and height of the network input size; width and height represent the width and height of the original image.
[0102] Considering that the origin of the reference coordinate system is the upper left corner point of the original image, in order to make the expanded region as much as possible in the first quadrant of the reference coordinate point (both x and y are positive), the max function and min function are used to avoid the situation where the coordinate points in the expanded region are negative, and also to limit the size of the expanded region exceeding the original image.
[0103] S303: Cut out the target region in the expanded region based on the positional relationship of the expanded region relative to the negative sample image and the boundary of the expanded region. Among them, both the expanded region and the target region include defects, and the accuracy of the defect being in the middle position in the target region is higher than that in the expanded region.
[0104] In this step, for Figure 8 the case where the expanded region shown in the left figure exceeds the original image, cut out the image region corresponding to the original image in the expanded region as the target region, as Figure 8 the region composed of ROI_width width and ROI_height height in the right figure.
[0105] Generally speaking, as can be seen from Figure 8 the right figure, for the case where the expanded region exceeds the original image, it is equivalent to the case where the width (ROI_width) or height (ROI_height) of the target region is less than the width or height in the network input size. The excess regions on the left and right (or up and down) parts of the expanded region relative to the original image can be cut off, and the other region parts in the expanded region except the cut-off parts are the target region.
[0106] For Figure 9 the case where the expanded region shown in the left figure does not exceed the original image and is located within the original image, cut out the image region within the boundary of the expanded region in the original image as the target region, as Figure 9 shown in the right figure, which is the case where the width (ROI_width) or height (ROI_height) of the target region is greater than the width or height in the network input size.
[0107] It can be understood that, relative to the position of the defect in the original image, in the cropped target region, the defect is more located in the middle of the region. In this way, it is greatly convenient for subsequent calculation of the defect ratio.
[0108] S304: Based on the target region and the desired image size, obtain a small defect image in the target region.
[0109] In this step, crop the small defect image from the target region to facilitate subsequent calculation of the defect ratio in the small defect image.
[0110] During implementation, the following formula can be used to crop the small defect image from the target region.
[0111]
[0112] In the above formula, the random function represents randomly selecting a number random within the parentheses. x_start and x_end represent the starting x coordinate and the ending x coordinate of the small defect image at the cropping position. y_start and y_end represent the starting y coordinate and the ending y coordinate of the small defect image at the cropping position. That is, the cropped small defect image is a rectangle, the coordinates of the upper left corner point are (x_start, y_start), and the coordinates of the lower right corner point are (x_end, y_end). ROI_width and input_width represent the width of the target region and the width of the network input size respectively. ROI_height and input_height represent the height of the target region and the height of the network input size respectively.
[0113] S305: Based on the size of the defect in the small defect image and the size of the original defect in the negative sample image, obtain the target negative sample image in the j-th round.
[0114] In this step, the size of the defect in the small defect image such as the area and the size of the original defect in the negative sample image such as the area can be obtained. The circumscribed rectangle of the defect in the small defect image can be implemented using a similar scheme to the relevant content in the aforementioned S302, which will not be elaborated.
[0115] If the size of the circumscribed rectangle of the defect in the small defect image is smaller than the desired image size, then divide the area of the defect in the small defect image by the area of the original defect to obtain the defect ratio of this defect. As Figure 10 shown, the defect that appears in the cropped small image is a part of the original defect. Divide the area of this part by the area of the original defect to obtain the defect ratio of this defect.
[0116] If the size of the circumscribed rectangle of the defect in the defect sub-image is greater than the expected image size, based on the minimum point coordinates, maximum point coordinates of the circumscribed rectangle, and the expected image size, obtain the adjustment ratio; based on the adjustment ratio, the size of the original defect, and the size of the defect in the defect sub-image, obtain the defect ratio.
[0117] As Figure 11 shown, the white area is the defect. The width of the circumscribed rectangle of the defect in the defect sub-image is greater than the width of the network input size. If the ratio of the area of the defect in the defect sub-image to the area of the original defect is directly calculated, the calculated defect ratio will not be accurate enough. Therefore, the adjustment ratio can be calculated first to calculate the size of the defect ratio under the input network size, and then the defect ratio can be calculated based on the area of the defect in the defect sub-image, the area of the original defect, and the adjustment ratio.
[0118] During implementation, the adjustment ratio can be calculated through the following formula.
[0119]
[0120] In the above formula, area represents the area of the original defect. input_width and input_height represent the width and height of the network input size. x_min and y_min represent the x and y of the upper left corner coordinates of the circumscribed rectangle of the defect in the defect sub-image; x_max and y_max represent the x and y of the lower right corner coordinates of the circumscribed rectangle of the defect in the defect sub-image. cut_area represents the area of the defect in the defect sub-image. scale_height is the adjustment ratio in the height direction, and scale_width is the adjustment ratio in the width direction. The adjustment of the defect ratio can only involve the adjustment in the height direction, can also only involve the adjustment in the width direction, or can also involve the adjustment in both the height and width directions. As Figure 11 shown, for the case where the width of the circumscribed rectangle of the defect in the defect sub-image is greater than the width of the network input size, only the adjustment in the width direction is involved.
[0121] Next, the defect ratio defect_ratio can be calculated through the following formula.
[0122] defect_ratio = cut_area / (area * scale_width * scale_height)
[0123] Subsequently, based on the judgment result of whether the defect ratio satisfies the defect ratio condition, the target negative sample image in the j-th round is obtained. Further, the size relationship between the defect ratio and the preset defect ratio threshold is judged. If the defect ratio is greater than or equal to the preset defect ratio threshold, it is considered that the defect ratio condition is satisfied, and the small defect image is retained as the target negative sample image in the j-th round. If the defect ratio is less than the preset defect ratio threshold, the small defect image is discarded. Among them, the defect ratio condition is that the defect ratio is greater than or equal to the preset defect ratio threshold.
[0124] In this application, j = j + 1 can be used to repeatedly execute the solutions of S301 to S304 to obtain the target negative sample images in each round. The target negative sample images in all rounds are used as the training samples for the semantic segmentation model to be trained. That is, as shown in S301 to S304, by calculating the circumscribed rectangle of the defect in the original image, expanding the region, cropping the target region, cropping the small defect image, and calculating the defect ratio, the images with a large defect ratio are retained, and the images with a small defect ratio are deleted, increasing the diversity of defects, enabling the semantic segmentation model to focus more on learning defect features, improving the training accuracy of the semantic segmentation model, and thus enhancing the accuracy of the semantic segmentation model for defect recognition or detection.
[0125] Different from the original image where the target positive sample image is a positive sample, the target negative sample image in this application is the retained small defect image. This small defect image can be used as the training negative sample of the semantic segmentation model. By using the defect ratio, the deletion of images with a small defect ratio is realized, avoiding the influence of useless data on model training and improving the training accuracy of the model.
[0126] From the perspective of the semantic segmentation model, the defects of the retained negative samples are diverse, and it also avoids the useless influence of training negative samples with a small defect ratio on model training, improving the training accuracy of the model.
[0127] In this application, for defect samples, the screening of defect samples for the semantic segmentation model needs to be realized through the calculation of the defect ratio. The main reason is that: in industrial scenarios, it is difficult to collect defect samples and the collection cycle is long. Deep learning models rely on a large number of positive and negative sample data for training. In order to enable the semantic segmentation model, which is a deep learning model, to better focus on defects, in this application, the images with a large defect ratio are retained as the negative samples for model training through the calculation of the defect ratio, so that the model can focus more on learning the features of defects and improve the accuracy of the model for defect detection or recognition.
[0128] As can be seen from the above content, for the original training set of training images, by setting the positive-negative ratio parameter, the number of positive and negative samples can be constrained, and an attempt is made to ensure that the semantic segmentation model is trained with positive and negative samples having a balanced number of samples. For the positive samples in the original training set, by calculating the sampling frequency of each positive sample, the selection probability of each positive sample can be constrained, so that each positive sample is evenly selected. For the negative samples in the original training set, by calculating the defect proportion of the negative samples, it is ensured that the training negative samples of the semantic segmentation model are all images with a large defect proportion, thereby improving the detection accuracy of the model for defects.
[0129] Overall, for positive and negative samples, the sampling frequency and defect proportion can be calculated respectively in the same round. The total number of rounds M for calculating the sampling frequency of positive samples can be the same as the total number of rounds N for calculating the defect proportion of negative samples.
[0130] Thus, this application can obtain the original training set, and input the defect proportion threshold and the positive-negative sample image ratio parameter. Among them, according to the number of positive and negative samples in the original training set and this ratio parameter, the number of normal samples and the number of defective samples required for training the semantic segmentation model can be calculated. The sampling frequency of positive samples is calculated in a multi-round loop, and the positive samples are cropped to a unified size. The defect proportion of negative samples is calculated in a multi-round loop, and the small defective images retained by the defect proportion are cropped to a unified size.
[0131] After the multi-round loop of positive and negative samples ends, it is judged whether the number of normal images obtained through processes such as calculating the sampling frequency of positive samples and cropping to a unified size meets the number of normal samples required for the semantic segmentation model. If it meets, these normal images can be used as the training positive samples of the semantic segmentation model and input into the semantic segmentation model for training. If it does not meet, the loop continues until the number is satisfied. Also, it is judged whether the number of defective images obtained through processes such as calculating the defect proportion of negative samples and cropping to a unified size meets the number of defective samples required for the semantic segmentation model. If it meets, these defective images can be used as the training negative samples of the semantic segmentation model and input into the semantic segmentation model for training. If it does not meet, the loop continues until the number is satisfied.
[0132] In this application, the configuration or setting scheme of the above-mentioned positive and negative sample image ratio parameters, the calculation scheme of the sampling frequency of positive samples, the calculation scheme of the defect proportion of negative samples, the unified size cropping scheme, etc. are used to preprocess the original training set before training the semantic segmentation model to obtain the training data required by the semantic segmentation model. Based on this, the processing scheme of the training images in this application can be regarded as a preprocessing strategy scheme before semantic segmentation. Using this preprocessing strategy scheme, the original training set is preprocessed to obtain training data that meets the requirements of the semantic segmentation model.
[0133] Based on this, this application also provides a training scheme for the semantic segmentation model. As Figure 12 shown, the original training set is obtained through collection, and a preprocessing strategy scheme before semantic segmentation is performed on the original training set. The two target samples obtained through this scheme are used as the training data of the semantic segmentation model and input into the semantic segmentation model for training. When the training is completed, a semantic segmentation model that can be used to identify or detect defects or normality is obtained.
[0134] It can be understood that after the semantic segmentation model is trained, this application also provides an identification or detection scheme for implementing the semantic segmentation task using the semantic segmentation model. As Figure 13 shown, the identification method or detection method implemented using the semantic segmentation model includes:
[0135] S1301: Obtain the image to be identified.
[0136] In this step, the image to be identified can be obtained by collecting the image to be identified or by reading the already collected image to be identified. The image to be identified can be a glass image expecting to identify whether there are defects in the glass image, or a camera image expecting to identify whether there are defects in the camera image.
[0137] S1302: Input the image to be identified into the semantic segmentation model to obtain an identification result on whether there are abnormalities in the image to be identified; wherein, the semantic segmentation model is obtained by training the semantic segmentation model to be trained using the target training images processed by the above-mentioned training image processing method.
[0138] In this step, input the image to be identified into the semantic segmentation model to obtain an identification result on whether the image to be identified is an image with defects or an image without defects output by the semantic segmentation model. In addition, the semantic segmentation model can also inform whether there are images with defects in the identified image by outputting the position of the defect in the image to be identified. If the semantic segmentation model does not output the defect position, it means that the image to be identified is a normal image without defects.
[0139] In this application, the aforementioned sliding window type cropping method can also be used to crop the image to be recognized into two or more sub-images. The two or more sub-images are input into a semantic segmentation model trained using a target training image to obtain recognition results on whether each sub-image has an abnormality; based on the recognition results on whether each sub-image has an abnormality, a recognition result on whether the image to be recognized has an abnormality is obtained.
[0140] In practical applications, the trained semantic segmentation model can recognize or detect whether each sub-image is a defective image. If there is a defective image among the sub-images, it is considered that the image to be recognized is a defective image. If there is no defect in each sub-image, it is considered that the image to be recognized is a non-defective image, that is, a normal image. In addition, for a defective sub-image, the semantic segmentation model can also output the position of the defect in the sub-image. If there are two or more defective images among the sub-images, the defects can be stitched together according to the positions of the defects in these two or more defective images and the positional relationship between the two defective images, so as to obtain the position of the defect in the image to be recognized.
[0141] It can be understood that since the semantic segmentation model is trained by using the aforementioned preprocessing strategy solution before semantic segmentation, rich training samples are provided for training. Training the semantic segmentation model with rich and diverse training samples can greatly improve the training accuracy and robustness of the semantic segmentation model. Using a semantic segmentation model with strong accuracy and robustness to recognize or detect the image to be recognized can effectively improve the recognition or detection accuracy and avoid missed recognition and misrecognition.
[0142] This application provides a processing device for training images, such as Figure 14 shown, including:
[0143] A first obtaining unit 1301, configured to obtain positive sample images and negative sample images in the training image; a first processing unit 1302, configured to perform a first process on the positive sample images to obtain target positive sample images; a second processing unit 1303, configured to perform a second process on the negative sample images to obtain target negative sample images; a second obtaining unit 1304, configured to obtain a target training image based on the target positive sample images and the target negative sample images, where the target training image is used to train a semantic segmentation model to be trained, and the trained semantic segmentation model is used to recognize whether the image to be recognized has an abnormality.
[0144] In some embodiments, the second obtaining unit 1304 is configured to:
[0145] If the image sizes of the target positive sample image and / or the target negative sample image do not match the expected image size of the semantic segmentation model, then crop the target positive sample image and / or the target negative sample image to obtain the target training image; if the image sizes of the target positive sample image and / or the target negative sample image match the expected image size of the semantic segmentation model, then use the target positive sample image and the target negative sample image as the target training image.
[0146] In some embodiments, the second obtaining unit 1304 is configured to: obtain a positive and negative sample image ratio parameter; based on the positive and negative sample image ratio parameter, the target positive sample image, and the target negative sample image, obtain the target training image.
[0147] In some embodiments, the target positive sample image is obtained by performing M rounds of first processing on the positive sample image; M is a positive integer greater than or equal to 2;
[0148] The first processing unit 1302 is configured to: obtain the initial selection probability of each positive sample image in the i-th round, where i is a positive integer greater than or equal to 2 and i ≤ M; process the initial selection probability of each positive sample image in the i-th round to obtain the target selection probability of each positive sample image in the i-th round, wherein the balance of the selection of each positive sample image under the target selection probability is stronger than the balance of the selection of the positive sample image under the initial selection probability; based on the target selection probability, obtain the selected positive sample images in the i-th round; based on the selected positive sample images in the i-th round, obtain the target positive sample image in the i-th round.
[0149] In some embodiments, the first processing unit 1302 is configured to: based on the number of times each positive sample image in the (i - 1)-th round is selected and determine the maximum number of times a positive sample image is selected, obtain the initial selection probability of each positive sample image in the i-th round; perform a smoothing process on the initial selection probability to obtain the first intermediate probability of each positive sample image; adjust the probability distribution of the first intermediate probability of each positive sample image to obtain the second intermediate probability of each positive sample image; perform normalization on the second intermediate probability of each positive sample image to obtain the target selection probability.
[0150] In some embodiments, the target negative sample image is obtained by performing N rounds of second processing on the negative sample image; N is a positive integer greater than or equal to 2; the second processing unit 1303 is configured to: based on the defect annotation information of each negative sample image in the j-th round, determine the defect contour in the negative sample image, where j is a positive integer greater than or equal to 1 and j ≤ N; expand the defect contour in the negative sample image to obtain an expanded region; based on the positional relationship between the expanded region and the negative sample image and the boundary of the expanded region, cut out the target region in the expanded region, where both the expanded region and the target region include defects, and the accuracy of the defect being in the middle position in the target region is higher than that in the expanded region; based on the target region and the expected image size, obtain a defect thumbnail in the target region; based on the size of the defect in the defect thumbnail and the size of the original defect in the negative sample image, obtain the target negative sample image in the j-th round.
[0151] In some embodiments, the second processing unit 1303 is configured to: based on the maximum point coordinates and minimum point coordinates of the defect contour in the negative sample image, obtain the circumscribed rectangle region of the defect contour; expand the circumscribed rectangle region to the expected image size to obtain an expanded region. In some embodiments, the second processing unit 1303 is configured to: if the size of the circumscribed rectangle of the defect in the defect thumbnail is greater than the expected image size, based on the minimum point coordinates, maximum point coordinates of the circumscribed rectangle and the expected image size, obtain an adjustment ratio; based on the adjustment ratio, the size of the original defect and the size of the defect in the defect thumbnail, obtain a defect ratio; based on the judgment result of whether the defect ratio meets the defect ratio condition, obtain the target negative sample image in the j-th round.
[0152] In some embodiments, the second obtaining unit 1304 is configured to:
[0153] If the sample image size of the target positive sample image and / or the target negative sample image is greater than the expected image size, then based on the random coordinate points generated in the target positive sample image and / or the target negative sample image, perform cropping of the target positive sample image and / or the target negative sample image to the size of the expected image size to obtain a new cropped image, and based on the overlap degree between the new cropped image and the historical cropped image, obtain the target training image.
[0154] In some embodiments, the second obtaining unit 1304 is configured to:
[0155] If the sample image size of the target positive sample image and / or the target negative sample image is greater than the expected image size, then adopt a sliding window cropping method to perform image cropping on the target positive sample image and / or the target negative sample image to obtain the target training image; wherein, the sliding step of the sliding window in the sliding window cropping method is less than the size of the sliding window.
[0156] In some embodiments, the second obtaining unit 1304 is configured to: if the sample image size of the target positive sample image and / or the target negative sample image is smaller than the desired image size, calculate a first scaling ratio and a second scaling ratio based on the sample image size and the desired image size of the target positive sample image and / or the target negative sample image, where the first scaling ratio is the scaling ratio in the width direction of the target positive sample image and / or the target negative sample image, and the second scaling ratio is the scaling ratio in the height direction of the target positive sample image and / or the target negative sample image; obtain a target scaling ratio based on the first scaling ratio and the second scaling ratio; scale the sample image of the target positive sample image and / or the target negative sample image according to the target scaling ratio to obtain a scaled image; and obtain a target training image based on the scaled image.
[0157] In some embodiments, the third obtaining unit 1305 is configured to obtain an image to be recognized; and the fourth obtaining unit 1306 is configured to input the image to be recognized into a semantic segmentation model trained with the target training image to obtain a recognition result on whether there is an abnormality in the image to be recognized.
[0158] In some embodiments, the fourth obtaining unit 1305 is configured to: input two or more sub-images into a semantic segmentation model trained with the target training image to obtain a recognition result on whether there is an abnormality in each sub-image; and obtain a recognition result on whether there is an abnormality in the image to be recognized based on the recognition result on whether there is an abnormality in each sub-image.
[0159] This application further provides a recognition device, as Figure 15 shown, the device includes:
[0160] A first obtaining unit 1401, configured to obtain an image to be recognized; a second obtaining unit 1402, configured to input the image to be recognized into a semantic segmentation model to obtain a recognition result on whether there is an abnormality in the image to be recognized; where the semantic segmentation model is obtained by training a semantic segmentation model to be trained with the target training image according to any one of claims 1-11.
[0161] In some embodiments, the second obtaining unit 1402 is configured to: crop the image to be recognized into two or more sub-images; input the two or more sub-images into a semantic segmentation model trained with the target training image to obtain a recognition result on whether there is an abnormality in each sub-image; and obtain a recognition result on whether there is an abnormality in the image to be recognized based on the recognition result on whether there is an abnormality in each sub-image.
[0162] It should be noted that for the processing device and recognition device of the training images in the embodiments of the present application, since the principles of the processing device and recognition device of the training images are similar to the aforementioned processing method and recognition method of the training images, therefore, the implementation process and implementation principle of the processing device and recognition device of the training images can refer to the description of the implementation process and implementation principle of the aforementioned method, and the repeated parts will not be elaborated.
[0163] According to an embodiment of the present application, the present application also provides an electronic device and a readable storage medium.
[0164] Wherein, the electronic device includes at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the processing method and recognition method of the training images of the present application. The computer instructions are used to cause the computer to execute the processing method and recognition method of the training images of the present application.
[0165] The present application also provides a computer program product, including computer programs / instructions, and when the computer programs / instructions are executed by a processor, the processing method and recognition method of the training images of the present application are implemented.
[0166] Figure 16 FIG. shows a schematic block diagram of an exemplary electronic device 800 that can be used to implement the embodiments of the present application. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present application described and / or claimed herein.
[0167] As Figure 16 shown, the device 800 includes a computing unit 801, which can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 802 or the computer program loaded from the storage unit 808 into the random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. The input / output (I / O) interface 805 is also connected to the bus 804.
[0168] Multiple components in device 800 are connected to I / O interface 805, including: input unit 806, such as a keyboard, mouse, etc.; output unit 807, such as various types of displays, speakers, etc.; storage unit 808, such as a disk, optical disc, etc.; and communication unit 809, such as a network card, modem, wireless communication transceiver, etc. Communication unit 809 allows device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0169] Computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Computing unit 801 executes the various methods and processes described above, such as the processing method and recognition method of training images. For example, in some embodiments, the processing method and recognition method of training images can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by computing unit 801, one or more steps of the processing method and recognition method of training images described above can be executed. Alternatively, in other embodiments, computing unit 801 can be configured to execute the processing method and recognition method of training images in any other suitable way (e.g., by means of firmware).
[0170] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0171] The program code for implementing the method of the present application can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.
[0172] In the context of the present application, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0173] As described above, the foregoing are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily conceive of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A method for processing a training image, characterized in that: The method comprises: Obtain positive sample images and negative sample images in training images; Performing a first processing on the positive sample image to obtain a target positive sample image; Performing a second processing on the negative sample image to obtain a target negative sample image; Based on the target positive sample image and the target negative sample image, a target training image is obtained, and the target training image is used to train the semantic segmentation model to be trained. The trained semantic segmentation model is used to identify whether the image to be identified has an abnormality.
2. The method according to claim 1, characterized in that The step of obtaining a target training image based on the target positive sample image and the target negative sample image includes: If the image size of the target positive sample image and / or the target negative sample image does not match the expected image size of the semantic segmentation model, cropping the target positive sample image and / or the target negative sample image to obtain a target training image; If the image size of the target positive sample image and / or the target negative sample image matches the expected image size of the semantic segmentation model, the target positive sample image and the target negative sample image are used as target training images.
3. The method according to claim 1, characterized in that Also includes: Get the positive and negative sample image ratio parameters; Based on the positive-negative sample image ratio parameter, the target positive sample image and the target negative sample image, a target training image is obtained.
4. The method according to claim 1, characterized in that: The target positive sample image is obtained by performing M rounds of first processing on the positive sample image; M is a positive integer greater than or equal to 2; The first processing is performed on the positive sample image to obtain a target positive sample image, including: Obtain the initial probability of being selected for each positive sample image in the i-th round, where i is a positive integer greater than or equal to 2, i≤M; The initial selected probability of each positive sample image in the i-th round is processed to obtain the target selected probability of each positive sample image in the i-th round, wherein the balance of the positive sample images selected under the target selected probability is stronger than the balance of the positive sample images selected under the initial selected probability; Based on the probability of the target being selected, the selected positive sample image in the i-th round is obtained; Based on the selected positive sample images in the i-th round, the target positive sample images in the i-th round are obtained.
5. The method according to claim 4, characterized in that The processing of the initial selected probability of each positive sample image in the i-th round to obtain the target selected probability of each positive sample image in the i-th round includes: Based on the number of times each positive sample image is selected in the i-1th round and the maximum number of times among the positive sample images, the initial probability of selection of each positive sample image in the i-th round is obtained; Smoothing the initial selected probability to obtain the first intermediate probability of each positive sample image; Adjusting the probability distribution of the first intermediate probability of each positive sample image to obtain a second intermediate probability of each positive sample image; The second intermediate probability of each positive sample image is normalized to obtain the probability of the target being selected.
6. A recognition method, characterized in that: include: Obtaining an image to be recognized; The image to be identified is input into a semantic segmentation model to obtain an identification result of whether the image to be identified has an abnormality; wherein the semantic segmentation model is obtained by training the semantic segmentation model to be trained using the target training image described in any one of claims 1 to 5.
7. A training image processing device, characterized in that: include: A first obtaining unit, used to obtain positive sample images and negative sample images in training images; A first processing unit, configured to perform a first processing on the positive sample image to obtain a target positive sample image; A second processing unit, configured to perform a second processing on the negative sample image to obtain a target negative sample image; The second obtaining unit is used to obtain a target training image based on a target positive sample image and a target negative sample image. The target training image is used to train a semantic segmentation model to be trained. The trained semantic segmentation model is used to identify whether the image to be identified has an abnormality.
8. An identification device, characterized in that: The device comprises: A first obtaining unit is used to obtain an image to be recognized; The second obtaining unit is used to input the image to be identified into the semantic segmentation model to obtain a recognition result of whether the image to be identified has an abnormality; wherein the semantic segmentation model is obtained by training the semantic segmentation model to be trained using the target training image described in any one of claims 1 to 5.
9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 5 and / or the method of claim 6.
10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to enable a computer to execute the method according to any one of claims 1 to 5 and / or the method according to claim 6.