A training method of a carrier roller missing detection model and related equipment
By generating high-quality images of missing rollers and combining them with various evaluation metrics to enrich the training dataset, the problem of insufficient training effect and accuracy of deep learning network models in roller missing detection is solved, and efficient detection is achieved with limited samples.
Patent Information
- Application Number
- CN202511631069.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-10
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2045-11-10
AI Technical Summary
With limited real-world samples, existing technologies struggle to guarantee the training effectiveness and accuracy of deep learning network models, particularly in detecting missing rollers.
By processing images where no idler rollers are missing, a target mask is generated and then erased to produce images with missing rollers. High-quality images are selected by combining boundary tracelessness, color distribution matching metric, and illumination consistency evaluation metrics to enrich the training dataset. A deep learning network model is jointly trained using real samples and high-quality generated images.
Under limited real sample conditions, the training effect of deep learning network models and the accuracy of detection results are improved, alleviating the problem of scarce real missing roll samples and ensuring the stability and robustness of detection.
Smart Images

Figure CN121074562B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of equipment monitoring, in particular to a training method of a roller missing detection model and related equipment. BACKGROUND
[0002] Rollers are important components of belt conveyors, used to support the weight of the conveyor belt and the material. Therefore, as the basic components of belt conveyors, the health status of rollers is crucial to ensure the safe and smooth operation of belt conveyors. In the operation of belt conveyors, roller missing is an abnormal state that may occur, so roller missing detection is needed.
[0003] Currently, the method for detecting roller missing includes using a deep learning network model to detect roller missing. However, the training of the deep learning network model brings great difficulties to engineers. How to ensure the training effect of the deep learning network model and the accuracy of its detection results based on limited real samples has become a difficult problem for technical personnel in the field. SUMMARY
[0004] The present application aims to provide a training method of a roller missing detection model and related equipment to improve the above problems.
[0005] To achieve the above purpose, the technical solutions adopted by the embodiments of the present application are as follows:
[0006] In a first aspect, the embodiments of the present application provide a training method of a roller missing detection model, which comprises:
[0007] processing the first type of roller images that do not appear roller missing to obtain the target mask corresponding to the rollers in the first type of roller images;
[0008] performing erasing processing on the first type of roller images and their corresponding target masks to generate corresponding missing roller images;
[0009] screening the generated missing roller images according to a preset evaluation index, thereby obtaining target missing roller images, wherein the evaluation index includes boundary traceless measurement, color distribution matching measurement, and illumination consistency evaluation measurement;
[0010] training the roller missing detection model based on a training data set, wherein the training data set includes real samples and the obtained target missing roller images, and the real samples are missing roller image samples carrying artificial annotations.
[0011] In a second aspect, the embodiments of the present application provide a training device of a roller missing detection model, which comprises:
[0012] The first processing unit is configured to process the first type of roller image without roller missing to obtain a target mask corresponding to the roller in the first type of roller image.
[0013] The first processing unit is further configured to perform erasing processing on the first type of roller image and the target mask corresponding thereto to generate a corresponding missing roller image.
[0014] The first processing unit is further configured to screen the generated missing roller image according to a preset evaluation index, so as to obtain a target missing roller image, wherein the evaluation index includes a boundary traceless metric, a color distribution matching metric, and an illumination consistency evaluation metric.
[0015] The second processing unit is configured to train a roller missing detection model based on a training data set, wherein the training data set includes real samples and the obtained target missing roller image, and the real samples are missing roller image samples carrying artificial annotations.
[0016] In a third aspect, an embodiment of the present application provides a storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the method described above.
[0017] In a fourth aspect, an embodiment of the present application provides an electronic device, which includes a processor and a memory configured to store one or more programs, and when the one or more programs are executed by the processor, the method described above is implemented.
[0018] Compared with the prior art, the roller missing detection model training method and related device provided by the embodiment of the present application process the first type of roller image without roller missing to obtain a target mask corresponding to the roller in the first type of roller image, perform erasing processing on the first type of roller image and the target mask corresponding thereto to generate a corresponding missing roller image, screen the generated missing roller image according to a preset evaluation index, so as to obtain a target missing roller image, wherein the evaluation index includes a boundary traceless metric, a color distribution matching metric, and an illumination consistency evaluation metric, and train a roller missing detection model based on a training data set, wherein the training data set includes real samples and the obtained target missing roller image, and the real samples are missing roller image samples carrying artificial annotations. The roller body is positioned by the target mask, and then the first type of roller image is erased to generate a missing roller image. In combination with the three indexes of boundary traceless, color distribution, and illumination consistency evaluation, a high-quality target missing roller image is obtained through automatic screening. The high-quality target missing roller image is jointly trained with a small amount of real data samples, which effectively alleviates the scarcity of real missing roller samples. On the basis of limited real samples, the sample amount in the training data set is enriched by generating a target missing roller image, which guarantees the training effect of the deep learning network model and the accuracy of the detection result.
[0019] To make the above objectives, features and advantages of the present application more obvious and easy to understand, the following preferred embodiments are specifically described below, and the accompanying drawings are referred to for a detailed description. BRIEF DESCRIPTION OF DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can also be obtained without creative labor on the basis of these drawings.
[0021] Figure 1 The structural schematic diagram of the electronic device provided by the embodiments of the present application.
[0022] Figure 2 The flowchart of the training method of the carrier roller absence detection model provided by the embodiments of the present application.
[0023] Figure 3 The flowchart of the training method of the carrier roller absence detection model provided by the embodiments of the present application.
[0024] Figure 4 The unit schematic diagram of the carrier roller absence detection model training device provided by the embodiments of the present application.
[0025] In the figure: 10-processor; 11-memory; 12-bus; 13-communication interface; 501-first processing unit; 502-second processing unit. DETAILED DESCRIPTION
[0026] In order to make the objectives, technical solutions and advantages of the embodiments of the present application more clear, the following will combine the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are some of the embodiments of the present application, but not all the embodiments. The components of the embodiments of the present application described and shown in the drawings can be arranged and designed in various different configurations.
[0027] Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed present application, but only represents selected embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0028] It should be noted that like reference numerals and characters refer to like elements throughout the following description with like reference numerals and characters referring to like elements throughout the following description and across all figures. It should be noted that as used herein, the terms "first", "second", and the like, do not imply any relative importance of the described elements and do not imply any particular order of the described elements.
[0029] It should be noted that, in the description of the present application, the terms "first", "second", and the like, are used only to distinguish one entity or operation from another, and do not necessarily require or imply any actual relationship or order between the entities or operations. Also, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. Without further limitation, an element preceded by "comprises a" does not, without more constraints, foreclose the existence of additional identical elements in the process, method, article, or apparatus that comprises the recited element.
[0030] In the description of the present application, it should be noted that the terms "upper", "lower", "inner", "outer", and the like, indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the product of the present application is usually placed, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the device or element referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation on the present application.
[0031] In the description of the present application, it should also be noted that, unless otherwise explicitly specified and limited, the terms "provided", "connected" should be understood broadly, for example, it can be fixedly connected, or detachably connected, or integrally connected; it can be mechanically connected, or electrically connected; it can be directly connected, or indirectly connected through an intermediate medium, or it can be a communication between two elements. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.
[0032] Some embodiments of the present application will be described in detail below with reference to the accompanying drawings. The following examples and features in the examples can be combined with each other without conflict.
[0033] The electronic device provided by the embodiments of the present application can be a computer device, a mobile phone device, a server device, etc. Please refer to Figure 1, a structural schematic diagram of an electronic device. The electronic device includes a processor 10, a memory 11, and a bus 12. The processor 10 and the memory 11 are connected through the bus 12. The processor 10 is configured to execute an executable module stored in the memory 11, such as a computer program.
[0034] The processor 10 can be an integrated circuit chip having a signal processing capability. In the implementation process, each step of the training method of the missing roller detection model can be completed by the integrated logic circuit of the hardware or the instruction in the form of software in the processor 10. The processor 10 described above can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; or a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0035] The memory 11 can include a high-speed random access memory (RAM) and can also include a non-volatile memory, such as at least one disk memory.
[0036] The bus 12 can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. Figure 1 Only one bidirectional arrow is used to represent the bus 12, but it does not mean that there is only one bus 12 or only one type of bus 12.
[0037] The memory 11 is configured to store a program, such as a program corresponding to the training device of the missing roller detection model. The training device of the missing roller detection model includes at least one software functional module that can be stored in the memory 11 in the form of software or firmware or solidified in the operating system (OS) of the electronic device. After receiving an execution instruction, the processor 10 executes the program to implement the training method of the missing roller detection model.
[0038] Possibly, the electronic device provided by the embodiment of the application further comprises a communication interface 13. The communication interface 13 is connected with the processor 10 through the bus.
[0039] It should be understood that, Figure 1 The structure shown is only a schematic structure of part of the electronic device, and the electronic device can further comprise more or fewer components than those shown in the Figure 1 embodiment, or have a different configuration from that shown in the Figure 1 embodiment. Figure 1 The components shown in the embodiment can be realized by hardware, software or a combination thereof.
[0040] The roller absence detection model training method provided by the embodiment of the application can be applied to, but is not limited to, the electronic device shown in the embodiment, and the specific process is described in detail in the embodiment. Figure 1 The roller absence detection model training method provided by the embodiment of the application can be applied to, but is not limited to, the electronic device shown in the embodiment, and the specific process is described in detail in the embodiment. Figure 2 The roller absence detection model training method comprises S1, S2, S3 and S4, which are described as follows.
[0041] S1, processing the first type of roller image without roller absence to obtain a target mask corresponding to the roller in the first type of roller image.
[0042] S2, performing erasing processing on the first type of roller image and the target mask corresponding thereto to generate a corresponding roller absence image.
[0043] In an optional embodiment, the first type of roller image without roller absence can be erased according to the target mask by using an image erasing model (which can be, but is not limited to, a PowerPaint algorithm model) to generate a corresponding roller absence image.
[0044] S3, screening the generated roller absence image according to a preset evaluation index to obtain a target roller absence image.
[0045] The evaluation index comprises a boundary traceless metric, a color distribution matching metric and an illumination consistency evaluation metric.
[0046] S4, training a roller absence detection model based on a training data set.
[0047] The roller absence detection model (which can be, but is not limited to, a yolov11 model) is trained based on a training data set comprising real samples and the obtained target roller absence image (the ratio of the two can be, but is not limited to, 3:7), and the real samples are roller absence image samples carrying artificial annotations.
[0048] In the training method of the missing roller detection model provided in the embodiment of the application, the target mask is used to position the roller body, and then the first type of roller image is erased to generate a roller missing image. The three indexes of boundary traceless, color distribution and illumination consistency are combined to automatically screen, and a high-quality target roller missing image is obtained. The high-quality target roller missing image is combined with a small amount of real data samples for joint training, which effectively alleviates the scarcity of real roller missing samples. On the basis of limited real samples, the target roller missing image is generated to enrich the sample amount in the training data set, so as to ensure the training effect of the deep learning network model and the accuracy of the detection result.
[0049] On the basis of the foregoing, with regard to the content in S1, the embodiment of the application further provides an alternative implementation, please refer to the following. S1, the first type of roller image without roller missing is processed to obtain the target mask corresponding to the roller in the first type of roller image, including S11 and S12, which are specifically described as follows.
[0050] S11, the first type of roller image without roller missing is segmented to obtain the primary mask corresponding to the roller in the first type of roller image.
[0051] The first type of roller image without roller missing can be segmented based on an instance segmentation model (including but not limited to YOLOv11-Seg) to obtain the primary mask corresponding to the roller in the first type of roller image.
[0052] Optionally, in the training process of the instance segmentation model, a boundary traceless constraint condition is introduced to constrain the edge smoothness of the primary mask output by the model.
[0053] Specifically, the loss function of the boundary traceless constraint condition is:
[0054]
[0055] wherein, L denotes the boundary loss function, As the boundary traceless constraint condition, it is used to constrain the gradient (edge shape and intensity) of the predicted mask and the real mask at the boundary to be consistent, and the smaller the value is, the more natural the boundary transition is and the fewer the seams are, Y denotes the real mask (the roller instance binary mask (1 indicates the roller area and 0 indicates the background) obtained by labeling), X denotes the predicted mask (a probability graph or a continuous value after Sigmoid) output by the instance segmentation model, B denotes the pixel-level boundary set of the real mask (which can be obtained from a 3-pixel-wide ring is extracted, N denotes the number of boundary pixels of the real mask, denotes a boundary pixel, denotes an image gradient operator (first derivative of the mask in x, y direction, resulting in a two-dimensional vector ), denotes a gradient vector of the predicted mask at pixel point p, denotes a gradient vector of the real mask at pixel point p, denotes an L1 norm, i.e., the sum of absolute values of the difference between two vectors.
[0056] The boundary loss The intuitive meaning is to compare whether the direction and intensity of the predicted edge at the real boundary are consistent with the label; if the predicted boundary is too rough, misplaced, or jagged, the gradient difference at this place will be large, increasing; training minimizes this term, which makes the predicted boundary more consistent and the transition smoother, which is conducive to subsequent erasing and seamless connection.
[0057]
[0058] wherein, denotes a segmentation loss function, denotes the kth predicted mask, denotes the kth real mask, denotes the loss value between the kth real mask and the kth predicted mask, denotes the loss value between the real mask and the predicted mask of all instances;
[0059]
[0060] wherein, denotes a detection item loss, 、 、 denotes a loss weight coefficient, denotes the ith predicted box output by the instance segmentation model, denotes the real box corresponding to the ith predicted box, denotes an IoU loss function, which is used to measure the degree of overlap between the predicted box and the real box, denotes a Distribution FocalLoss, which is a new loss function for bounding box regression, denotes a binary cross-entropy loss (Binary Cross-Entropy Loss), which is used for classification tasks;
[0061]
[0062] wherein, denotes the total loss corresponding to the instance segmentation model, a weight coefficient of a boundary loss function, a weight coefficient of a boundary loss function.
[0063] S12, morphological expansion and smoothing processing are performed on the primary mask to obtain a target mask.
[0064] The target mask is used as a reference for the erasing processing in S2.
[0065] On the basis of the foregoing, regarding the content in S12, the embodiments of the present application further provide an alternative implementation, please refer to the following. S12, morphological expansion and smoothing processing are performed on the primary mask to obtain a target mask, comprising:
[0066] S121, geometric smoothing processing is performed on the primary mask to generate a corresponding hard mask.
[0067] Among them, the geometric smoothing uses the commonly used open-close combination to remove sharp corners and small holes.
[0068] Optionally, the formula of the geometric smoothing processing is:
[0069]
[0070] Among them, the hard mask, the prediction mask output by the instance segmentation model, i.e. the primary mask, the structural element, the closing operation, the opening operation.
[0071] S122, the hard mask is processed using a Gaussian blur algorithm to obtain a soft mask after edge blur processing as the target mask.
[0072] Optionally, the Gaussian blur edge band formula used by the Gaussian blur algorithm is as follows:
[0073]
[0074] Among them, the soft mask, the Gaussian kernel, the Gaussian kernel is used to convolve the hard mask to realize the gradual transition of the edge, and the standard deviation of the Gaussian kernel is , the convolution operation, the Gaussian kernel is convolved with the hard mask to generate a continuous gray value mask, the clipping operation, to limit the convolution result in the range of [0, 1], to ensure that the value range of the soft mask meets the requirements.
[0075] On the basis of the foregoing, with regard to the content in S3, the embodiment of the present application further provides an alternative implementation, please refer to the following. S3, the generated missing roller image is screened according to the preset evaluation index, so as to obtain the target missing roller image, comprising: S31, S32, S33 and S34, which are specifically described as follows.
[0076] S31, the boundary traceless measure of the missing roller image, the color distribution matching measure and the illumination consistency evaluation measure are obtained.
[0077] With regard to the specific formula of the boundary traceless measure, the color distribution matching measure and the illumination consistency evaluation measure, the embodiment of the present application further provides an alternative implementation, please refer to the following.
[0078] The boundary traceless measure is used to measure whether the boundary of the erasing area is smooth, and the specific formula is as follows:
[0079]
[0080] Among them, The boundary traceless measure is used to measure whether the boundary of the erasing area is smooth, and the specific formula is as follows: The boundary set of the mask M of the erasing area in the missing roller image, that is, the boundary line of the erasing area and the non-erasing area in the missing roller image, The boundary pixel number of the mask M, The single pixel point on the boundary of the mask M, The normal gradient operator is used to calculate the gradient of the image in the normal direction of the boundary (which can accurately reflect the visual continuity at the boundary), The Lab space brightness component (L channel) value of the pixel p1 inside the erasing area, The Lab space brightness component value of the adjacent pixel p1 outside the erasing area;
[0081] The boundary traceless measure formula quantifies the gradient discontinuity degree of the erasing area and the surrounding environment at the boundary. Ideally, if the erasing is perfect, the gradient at the boundary should be smooth transition, Should be close to 0. Higher Value indicates that there is obvious boundary trace or "joint". In the roller detection, since the roller is usually located below the conveying belt, its edge may have obvious contrast with the background. Therefore, reducing It is essential for generating natural missing roller image, especially when the roller part is blocked.
[0082] The color distribution matching measure is used to measure the evaluation of color distribution consistency, and the formula is as follows:
[0083]
[0084] where, represents a color distribution matching metric, measuring the difference in color distribution inside the erasure region , ) and outside the erasure region , , represents the mean vector inside the erasure region in Lab color space (each vector contains the average of L, a, b components), represents the mean vector outside the erasure region in Lab color space (each vector contains the average of L, a, b components), represents the covariance matrix inside the erasure region in Lab color space, represents the covariance matrix outside the erasure region in Lab color space (these matrices describe the variance and covariance relationships of color distribution), represents the trace of a matrix, which is the sum of diagonal elements, represents the square root of a matrix, calculated through eigenvalue decomposition. This step ensures the positive definiteness of the distance.
[0085] This formula is specifically designed for comparing two multi-dimensional Gaussian distributions. It takes into account both the mean difference (location difference) and the covariance difference (shape difference), allowing a comprehensive assessment of the consistency of color distribution. In industrial environments, lighting conditions can vary greatly, leading to differences in color distribution between different regions. Therefore, helps to ensure that the color of the erasure region not only matches in mean, but also in texture and color diversity with the surrounding environment.
[0086] Illumination consistency metric, used to measure the consistency of the illuminance field inside and outside the erasure region, the formula is as follows:
[0087]
[0088] where, represents the illumination consistency metric, used to measure the consistency of the illuminance field inside and outside the erasure region, represents the illuminance field inside the erasure region, represents the illuminance field outside the erasure region (in Retinex theory, the illuminance field reflects the intensity of the scene's light), represents the mean of the illuminance field inside the erasure region, represents the mean of the illuminance field outside the erasure region, represents the standard deviation of the illuminance field inside the erasure region, a standard deviation of the luminance field outside the erasing region, denotes a weight coefficient for balancing the importance of mean difference and standard deviation difference. The weight coefficient can be but not limited to set as 0.7.
[0089] The formula is based on the Retinex theory, which believes that the human eye perceives reflectance rather than direct illumination intensity. Therefore, even under different lighting conditions, as long as the reflectance pattern is similar, the human eye will consider that the object looks consistent. Ensures that the luminance distribution of the erasing region remains consistent with the surrounding environment in statistical properties. In industrial sites, the lighting system can be uneven, causing some areas to be brighter. Helps ensure that the area after erasing does not appear conspicuous due to inconsistent lighting, especially in areas with shadows or reflections.
[0090] S32, normalize the obtained boundary traceless measure, color distribution matching measure and lighting consistency evaluation measure to obtain boundary traceless normalized score, color distribution matching normalized score and lighting consistency normalized score.
[0091] S33, according to the boundary traceless normalized score, color distribution matching normalized score and lighting consistency normalized score, weighted operation is carried out to obtain the comprehensive quality score of the missing roll image.
[0092] Optionally, the formula of the comprehensive quality score is:
[0093]
[0094] wherein, denotes the comprehensive quality score, ranging from 0 to 1, the larger the value, the higher the quality of the generated data, , , denotes a preset weight coefficient, corresponding to the weights of boundary traceless, color distribution matching and lighting consistency respectively (these weights can be adjusted according to specific application scenarios), , , denotes , , The normalized single score (boundary traceless normalized score, color distribution matching normalized score and lighting consistency normalized score respectively). The original scores such as boundary traceless measure, color distribution matching measure and lighting consistency evaluation measure are scaled to the interval [0, 1] for easy weighted summation. The formula combines the three independent quality indicators into a single score, which is convenient for subsequent screening and sorting. By adjusting the weight, the importance of a particular aspect can be emphasized.
[0095] S34. Select the roll-missing images with a comprehensive quality score greater than the comprehensive score threshold as the target roll-missing images.
[0096] The roller missing detection model training method provided in this invention has the characteristics of interpretable and controllable data quality assurance, no reliance on manual selection, and high efficiency. It quantifies the boundary gradient smoothness, color distribution divergence and illumination consistency as criteria to form an auditable sample screening process, avoids "illusion" and artifact samples from entering the training set, and stabilizes the training convergence and online effect.
[0097] In one alternative implementation, the overall scoring threshold is reduced when the number of training iterations of the idler roller missing detection model reaches a set value.
[0098] Based on the comprehensive scoring threshold High-quality generated datasets were obtained through filtering. Overall scoring threshold Only when S> Only when the time is right is the sample considered high quality. A relatively strict threshold is set in the initial stage. =0.8 to ensure data quality. As the model training progresses and the number of high-quality training samples reaches the set value, the threshold can be appropriately relaxed to increase the amount of data.
[0099] Regarding S4, during the training of the idler missing detection model based on the training dataset, how should the loss function of the idler missing detection model be set? This embodiment of the invention also provides an optional implementation method, please refer to the following.
[0100] The loss function of the idler roller missing detection model is a sample-weighted loss function, calculated as follows:
[0101]
[0102] in, This represents the overall training loss function of the missing idler roller detection model, used to optimize its performance. , , These represent the loss weight coefficients (which control the importance of IoU loss, classification loss, and DFL loss, respectively). This represents the bounding box regression loss for the i-th sample, measuring the positional difference between the pre-defined bounding box and the ground truth bounding box. This represents the classification loss for the i-th sample, measuring the difference between the predicted and true classes. This represents the DFL loss of the i-th sample (DFL models the bounding box coordinates as a discrete distribution, rather than directly regressing continuous values, thus improving localization accuracy). This represents the weight of the i-th sample, determined by its overall quality score. determining, denotes the comprehensive quality score of the i-th sample (ranging between [0, 1], the larger the value, the higher the quality of the sample), denotes the weight adjustment coefficient, which controls the degree of influence of the quality score on the maximum weight.
[0103] The value of can be but is not limited to 0.5. This formula realizes sample weighted training, that is, high-quality generated samples are given higher weights, so as to play a greater role in the training process. This helps to prevent the model from overfitting low-quality samples and improve overall detection performance. In the roller detection, since the generated samples may have certain deviations, using the quality score as the weight can effectively alleviate this problem, so that the model focuses more on learning the features in high-quality samples. In the roller detection, since the generated samples (i.e. the generated missing roller images) may have certain deviations, using the quality score as the weight can effectively alleviate this problem, so that the model focuses more on learning the features in high-quality samples. By weighting the samples for learning and introducing a mask area negative constraint suppressing the false positive of "regarding the erased area as a roller". Effect: stabilize training, reduce systematic false detection, and improve robustness to complex background / strong light reflection scenes.
[0104] It should be understood that no roller should be predicted in the erased area. Therefore, when training the roller missing detection model, a negative sample consistency regularization term is introduced, and the formula is as follows:
[0105]
[0106] wherein, denotes the negative sample consistency regularization term, which is used to punish the roller missing detection model for the wrong prediction in the erased area, denotes the j-th prediction box, denotes the mask M of the erased area, that is, the area where the roller originally exists but has been erased, denotes the confidence of the j-th prediction box, which represents the probability that the roller missing detection model considers that the box contains "roller", denotes the prediction box the Intersection over Union between the prediction box and the mask M of the erased area, denotes the total loss function of the roller missing detection model, denotes the regularization weight coefficient, which controls the importance of the consistency regularization term.
[0107] If the missing idler roller detection model gives a high confidence score in a prediction box that completely belongs to the erased area, it will be penalized. This mechanism helps improve the model's robustness and avoids false alarms in areas where there are no idler rollers. In practical applications, there may be some residual features or noise that are difficult to completely remove, causing the missing idler roller detection model to misclassify. Through consistency regularization, the model can be forced to learn the correct semantic information within the erased area, thereby reducing the false alarm rate.
[0108] Please refer to Figure 3 In one optional implementation, S2, the erasure process is performed based on the first type of idler image and its corresponding target mask to generate the corresponding missing roll image, including: S21.
[0109] S21, erase the model of the first type of idler roller image and its corresponding target mask input image to obtain the corresponding missing roller image.
[0110] The image erasure model can be a large model for multimodal image editing and can be used for image erasure.
[0111] Based on this, regarding how to optimize the image erasure model, this embodiment of the invention also provides an optional implementation method, please refer to [link / reference needed]. Figure 3 After training the idler roller missing detection model based on the training dataset, the training methods for the idler roller missing detection model include S5 and S6, which are described in detail below.
[0112] S5 uses the trained idler roller missing detection model to score the detection consistency of the generated missing roller images in order to obtain the detection consistency score corresponding to the missing roller images.
[0113] Among them, the consistency score measures the credibility of the generated missing roll image within the erased area. The closer the value is to 1, the more "credible" the missing roll image (sample) is.
[0114] Optionally, the formula for calculating the consistency score is as follows:
[0115]
[0116] in, Indicates the consistency score of the detection. Represents the j-th prediction box Intersection over Union (IoU) between the mask M and the erasing area. The degree of overlap between the prediction box and the erased area was quantified. Let represent the confidence level of the j-th predicted box, and represent the probability that the missing roller detection model considers the box to contain "idler". denotes taking the maximum value in all the predicted boxes, ensuring that even a single high-confidence false positive will be captured.
[0117] This equation evaluates the quality of the generated data by calculating the likelihood of the missing roller detection model producing a "roller-like" structure in the erasing region. If ≈1, it means that the missing roller detection model did not produce a high-confidence "roller" prediction in the erasing region, i.e., the region does not look like a roller, and thus is more "trustworthy". Conversely, if a lower value indicates that the missing roller detection model is likely to identify a roller-like structure in the erasing region, which could be due to low-quality generation.
[0118] Rollers typically have specific geometric shapes and texture features. High-quality generated data should cause these features to disappear after erasing, thus avoiding false positives for the missing roller detection model. Therefore, helps to filter out generated samples that will not cause the missing roller detection model to be confused.
[0119] S6, using the missing roller images with detection consistency scores greater than the score threshold, fine-tuning the image erasing model in a reverse optimization manner.
[0120] Optionally, the missing roller images with detection consistency scores greater than the score threshold are used as selected samples to fine-tune the image erasing model in a reverse optimization manner until the detection performance converges.
[0121] Optionally, the corresponding loss function equation for the image erasing model training is as follows:
[0122]
[0123]
[0124] wherein, denotes the total loss function of the image erasing large model in the fine-tuning stage, denotes the perception prior loss, which measures the difference in high-level semantic features between the predicted image and the original image, denotes the unmasked consistency loss, which is used to penalize the pixel value and gradient change in the unmasked region, M1 denotes a binary mask of the erasing region, wherein M1=1 indicates the region to be erased, and M1=0 indicates the background region to be kept unchanged, denotes the complement mask (denoting the region to be kept), denotes the generated missing roller image (i.e., the image after PowerPaint processing), denotes the first type of roller image (original input image), and denotes element-wise multiplication (Hadamard product), denotes the gradient operator, and the Sobel filter is often used for calculation, denotes the L1 norm, that is, the sum of absolute values.
[0125] The unmasked consistency loss includes a pixel-level consistency loss and a gradient consistency loss. When fine-tuning the image erasing model, adding the unmasked consistency loss (Context-Preservation Loss) is an innovation, which aims to ensure that the erasing operation only affects the target area, while the original information of the surrounding environment is preserved to the greatest extent.
[0126] In the roller absence detection model training method provided in the embodiments of the present application, continuous self-adaptation and rapid updating are maintained, the trained roller absence detection model is used to fine-tune PowerPaint, and the large model is continuously aligned in the industrial field; based on new field data, more realistic roller absence samples are automatically generated, which are used to automatically fine-tune the detection model and shorten the time from online to stable.
[0127] Please continue to refer to Figure 3 , regarding the iteration termination condition of the image erasing model, the embodiments of the present application further provide an optional implementation, please refer to the following. After the image erasing model is fine-tuned by reverse optimization, the roller absence detection model training method further includes S7, S8 and S9, which are specifically described as follows.
[0128] S7, obtain the feature value change amount between the current iteration round and the previous round.
[0129] wherein the feature value (mAP50) is the value of the average precision mean at a set IoU threshold (which can be but is not limited to 0.5).
[0130] Optionally, the formula of the feature value change amount is:
[0131]
[0132] wherein, denotes the feature value change amount between the current iteration round t and the previous iteration round t-1. is the value of the average precision mean at a set IoU threshold, which is often used to evaluate the performance of a target detection model. denotes the change tolerance, which is usually set to a small positive number (such as 0.001), indicating that when the change of mAP50 is less than this value, the model performance is considered to be stable.
[0133] S8, obtaining an average quality score change between the current iteration round and the previous round.
[0134] Optionally, the average quality score change is wherein, denotes the average quality score of the current iteration round t, i.e., the average of the quality scores of all generated missing roller images, denotes the average quality score of the previous round t-1, i.e., the average of the quality scores of all generated missing roller images. The tolerance of the average quality score change is usually set to a small value (such as 0.01), indicating that when the change in the average quality score is less than this value, it is considered that the quality of the generated data has stabilized.
[0135] S9, determining whether to stop fine-tuning the image erasing model according to the feature value change and the average quality score change.
[0136] The formula of the termination condition of the iteration process is:
[0137]
[0138]
[0139] When the above two conditions are met for K consecutive rounds (such as K = 2-3), the iteration is stopped. This means that the performance of the detection model and the quality of the generated data have reached a stable state, and further training may not bring significant improvement.
[0140] In the roller missing detection model training method provided by the embodiment of the present application, a closed-loop optimization mechanism is proposed, which can simultaneously improve the quality of the generated data and the performance of the detection model, and finally realize efficient and accurate roller missing detection in an industrial scene. This iterative optimization strategy is particularly suitable for industrial application scenarios where data is scarce but demand is urgent.
[0141] Please refer to Figure 4 , Figure 4 A roller missing detection model training device is provided in the embodiment of the present application. Optionally, the roller missing detection model training device is applied to the electronic device described above.
[0142] The roller missing detection model training device comprises a first processing unit 501 and a second processing unit 502.
[0143] The first processing unit 501 is configured to process the first type of roller image in which no roller is missing to obtain a target mask corresponding to the roller in the first type of roller image.
[0144] The first processing unit 501 is further configured to perform erasing processing on the first type of roller image and the corresponding target mask to generate a corresponding missing roller image.
[0145] The first processing unit 501 is further configured to perform screening on the generated missing roller image according to a preset evaluation index, so as to obtain a target missing roller image, wherein the evaluation index includes a boundary traceless metric, a color distribution matching metric, and an illumination consistency evaluation metric.
[0146] The second processing unit 502 is configured to train the roller missing detection model based on a training data set, wherein the training data set includes real samples and the obtained target missing roller image, and the real samples are missing roller image samples carrying artificial annotations.
[0147] Optionally, the second processing unit 502 can perform the S4 described above, and the first processing unit 501 can perform other steps in the above method embodiments.
[0148] It should be noted that the roller missing detection model training apparatus provided in the present embodiment can perform the method processes shown in the above method process embodiments to achieve the corresponding technical effects. For brevity, some parts of the present embodiment are not mentioned, and the corresponding contents can be referred to the above embodiments.
[0149] The present embodiment further provides a storage medium storing computer instructions and programs, which perform the roller missing detection model training method of the above embodiments when read and run. The storage medium can include memory, flash memory, register, or a combination thereof.
[0150] The following provides an electronic device, which can be a computer device, a mobile phone device, a server device, etc. The electronic device can implement the above roller missing detection model training method, as shown in the following Figure 1 The electronic device includes a processor 10, a memory 11, and a bus 12. The processor 10 can be a CPU. The memory 11 is used to store one or more programs, which perform the roller missing detection model training method of the above embodiments when executed by the processor 10.
[0151] In summary, the roller absence detection model training method and related device provided by the embodiment of the present application processes the first type of roller image without roller absence to obtain a target mask corresponding to the roller in the first type of roller image; performs erasing processing on the first type of roller image and the target mask corresponding thereto to generate a corresponding roller absence image; and filters the generated roller absence image according to a preset evaluation index, thereby obtaining a target roller absence image, wherein the evaluation index includes a boundary traceless metric, a color distribution matching metric, and an illumination consistency evaluation metric; and the roller absence detection model is trained based on a training data set, wherein the training data set includes real samples and the obtained target roller absence image, and the real samples are roller absence image samples carrying artificial annotations. The roller body is positioned through the target mask, and then the first type of roller image is erased to generate a roller absence image, and the three indexes of boundary traceless, color distribution, and illumination consistency evaluation are combined to automatically filter, thereby obtaining a high-quality target roller absence image. The high-quality target roller absence image is jointly trained with a small amount of real data samples, effectively alleviating the scarcity of real roller absence samples. On the basis of limited real samples, the sample amount in the training data set is enriched by generating a target roller absence image, thereby ensuring the training effect of the deep learning network model and the accuracy of the detection result.
[0152] The above only describes preferred embodiments of the present application and is not intended to limit the present application. For those skilled in the art, the present application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
[0153] It is obvious for those skilled in the art that the present application is not limited to the details of the above exemplary embodiments, and the present application can be implemented in other specific forms without departing from the spirit or essential characteristics of the present application. Therefore, the embodiments should be regarded as exemplary and non-limiting, and the scope of the present application is defined by the appended claims rather than the above description, and it is intended to include all changes falling within the meaning and scope of the equivalent elements of the claims. Any reference signs in the claims should not be regarded as limiting the claims.
Claims
1. A method for training a missing idler roller detection model, characterized in that, The method includes: The first type of idler image without missing idlers is processed to obtain the target mask corresponding to the idler in the first type of idler image; The first type of idler roller image and its corresponding target mask are used for erasure processing to generate the corresponding missing roller image; The generated roll-missing images are filtered according to preset evaluation indicators to obtain target roll-missing images. This includes: acquiring the boundary flawlessness metric, color distribution matching metric, and illumination consistency evaluation metric of the roll-missing image; normalizing the obtained boundary flawlessness metric, color distribution matching metric, and illumination consistency evaluation metric to obtain boundary flawlessness normalization score, color distribution matching normalization score, and illumination consistency normalization score; performing a weighted calculation based on the boundary flawlessness normalization score, color distribution matching normalization score, and illumination consistency normalization score to obtain a comprehensive quality score for the roll-missing image; and selecting roll-missing images with a comprehensive quality score greater than a comprehensive score threshold as target roll-missing images. The evaluation indicators include the boundary flawlessness metric, color distribution matching metric, and illumination consistency evaluation metric. The idler roller missing detection model is trained based on the training dataset, wherein the training dataset includes real samples and the obtained target missing roller images, and the real samples are missing roller image samples with manual annotations.
2. The method for training a missing roller detection model as described in claim 1, characterized in that, The process of processing the first type of idler image where no idler is missing to obtain the target mask corresponding to the idler in the first type of idler image includes: The first type of idler image without missing idlers is segmented to obtain the initial mask corresponding to the idler in the first type of idler image; The initial mask is morphologically expanded and smoothed to obtain the target mask.
3. The method for training a missing roller detection model as described in claim 2, characterized in that, The step of performing morphological expansion and smoothing on the initial mask to obtain the target mask includes: The initial mask is geometrically smoothed to generate the corresponding hard mask; The hard mask is processed using a Gaussian blur algorithm to obtain a soft mask with blurred edges, which serves as the target mask.
4. The method for training a missing idler roller detection model as described in claim 1, characterized in that, When the number of training iterations of the idler roller missing detection model reaches a set value, the comprehensive scoring threshold is reduced accordingly.
5. The method for training a missing idler roller detection model as described in claim 1, characterized in that, The loss function of the missing roller detection model is a sample-weighted loss function, calculated as follows: in, This represents the first training loss function of the idler roller missing detection model, used to optimize its performance. , , This represents the loss weighting coefficient. This represents the bounding box regression loss for the i-th sample, measuring the positional difference between the pre-defined bounding box and the ground truth bounding box. This represents the classification loss for the i-th sample, measuring the difference between the predicted and true classes. This represents the DFL loss for the i-th sample. This represents the weight of the i-th sample, determined by its overall quality score. Decide, This represents the overall quality score of the i-th sample. This represents the weight adjustment coefficient, which controls the degree of influence of the quality score on the final weight.
6. The method for training a missing idler roller detection model as described in claim 5, characterized in that, When training the idler roller missing detection model, a negative sample consistency regularization term is introduced, with the following formula: in, This represents a negative sample consistency regularization term, used to penalize incorrect predictions by the idler roller missing detection model within the erasure area. This represents the j-th prediction box. The mask M represents the area to be erased. Let represent the confidence level of the j-th predicted box, and represent the probability that the missing roller detection model considers the box to contain "idler". Represents the prediction box The cross-union ratio between the mask M and the erased area, This represents the overall training loss function of the entire idler roller missing detection model. This represents the regularization weight coefficient, which controls the importance of the consistency regularization term.
7. The method for training a missing roller detection model as described in claim 1, characterized in that, The step of performing erasure processing based on the first type of idler image and its corresponding target mask to generate the corresponding missing roll image includes: inputting the first type of idler image and its corresponding target mask into the image erasure model to obtain the corresponding missing roll image; The method for training the idler roller missing detection model based on the training dataset includes: The trained idler roll missing detection model is used to score the detection consistency of the generated missing roll images to obtain the detection consistency score corresponding to the missing roll images. By using images with missing rolls whose consistency scores are greater than the score threshold, the image erasure model is fine-tuned through reverse optimization.
8. The method for training a missing roller detection model as described in claim 7, characterized in that, After performing reverse optimization and fine-tuning on the image erasure model, the method further includes: Obtain the change in feature value between the current iteration and the previous iteration, where the feature value is the average precision mean value at a set IoU threshold; Obtain the change in average quality score between the current iteration and the previous iteration; Based on the change in the feature value and the change in the average quality score, determine whether to stop the reverse optimization fine-tuning of the image erasure model.
9. A training device for a missing idler roller detection model, characterized in that, The device includes: The first processing unit is used to process the first type of idler image where no idler is missing, so as to obtain the target mask corresponding to the idler in the first type of idler image; The first processing unit is further configured to perform erasure processing based on the first type of idler roller image and its corresponding target mask to generate a corresponding missing roller image; The first processing unit is further configured to filter the generated missing roll images according to preset evaluation indicators to obtain target missing roll images, including: acquiring the boundary flawlessness measure, color distribution matching measure, and illumination consistency evaluation measure of the missing roll image; normalizing the obtained boundary flawlessness measure, color distribution matching measure, and illumination consistency evaluation measure to obtain the boundary flawlessness normalization score, color distribution matching normalization score, and illumination consistency normalization score; performing a weighted calculation based on the boundary flawlessness normalization score, color distribution matching normalization score, and illumination consistency normalization score to obtain the comprehensive quality score of the missing roll image; and taking the missing roll image with a comprehensive quality score greater than the comprehensive score threshold as the target missing roll image, wherein the evaluation indicators include the boundary flawlessness measure, color distribution matching measure, and illumination consistency evaluation measure. The second processing unit is used to train the idler roller missing detection model based on the training dataset, wherein the training dataset includes real samples and the obtained target missing roller images, and the real samples are missing roller image samples with manual annotations.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-8.
11. An electronic device, characterized in that, include: Processor and memory, the memory being used to store one or more programs; When the one or more programs are executed by the processor, the method as described in any one of claims 1-8 is implemented.
Citation Information
Patent Citations
Steel structure welding quality detection method and storage medium
CN120219374A
Image processing system and image processing method
JP2025026384A