A highly robust mixed data augmentation synthesis method

By employing a highly robust hybrid data enhancement synthesis method, high-quality and highly diverse synthetic detection target data is generated, solving the problems of data scarcity and insufficient diversity in the detection of cracks in chemical plant transport chain plates, and improving the model's detection performance in complex environments.

CN122116029APending Publication Date: 2026-05-29HUAIYIN INSTITUTE OF TECHNOLOGY

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUAIYIN INSTITUTE OF TECHNOLOGY
Filing Date
2026-01-23
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

In the detection of cracks in conveyor chain plates in chemical plants, there are problems such as scarce, low-quality and poor diversity of training data, which leads to insufficient robustness and generalization ability of deep learning models in complex environments.

Method used

A robust hybrid data augmentation synthesis method is adopted to generate high-quality and highly diverse synthetic detection target data through morphological and geometric transformations. By combining stress concentration areas and perspective transformations, the cracks are ensured to be visually and logically consistent with the real scene. Boundary overflow processing and IoU anti-overlap mechanism are introduced to generate efficient training data.

Benefits of technology

It significantly improves the robustness and generalization ability of the detection model in real industrial environments. The generated dataset can effectively cover complex and ever-changing chemical plant environments, improve the accuracy and recall of crack detection, and reduce data acquisition costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122116029A_ABST
    Figure CN122116029A_ABST
Patent Text Reader

Abstract

A high-robustness mixed data enhancement synthesis method, comprising the steps of: 1. preprocessing the input original image to generate an accurate binary mask; 2. performing transformation enhancement on the synthesized image, including but not limited to rotation, shearing, erosion and dilation operations; 3. through a probability-weighted random allocation mechanism, randomly generate high-quality irregular detection targets in stress-concentrated positions in terms of the number of detection objects and the physical layer, simulate real fault scenarios; 4. using a matching algorithm to paste the samples of the detection targets onto the background image, and introducing a boundary overflow processing mechanism to ensure that the detection targets do not exceed the ROI or image boundary, and to ensure that the position of the detection target corresponds to the image, and to ensure that the label does not change. This method is applied to application scenarios with scarce data sets in industrial scenarios, significantly improving the diversity and complexity of the data set, thereby greatly improving the generalization ability and detection robustness of the subsequent target detection model in complex environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of target detection technology, and more specifically to a highly robust hybrid data augmentation synthesis method. Background Technology

[0002] Conveyor chains in chemical plants are critical components for carrying and transporting materials. Operating under harsh environments such as high temperature, high pressure, and corrosive media for extended periods, they are prone to cracking due to factors like metal fatigue and stress concentration. If these cracks are not detected and addressed promptly, they can lead to chain breakage, causing production interruptions, equipment damage, and even serious safety accidents and environmental pollution incidents. Therefore, regular, efficient, and accurate crack detection of conveyor chains is a crucial step in ensuring safe chemical production and implementing preventative maintenance. Currently, crack detection methods in the industrial field mainly include manual visual inspection, traditional machine vision methods, and deep learning-based methods.

[0003] Manual visual inspection: This is the most traditional method, relying on experienced inspectors. The disadvantages of this method are obvious: it is highly subjective, has inconsistent inspection standards, is inefficient, and personnel fatigue can lead to missed or false positives. Furthermore, in certain high-risk or inaccessible chemical environments, the feasibility and safety of manual inspection are severely limited.

[0004] Traditional machine vision methods typically utilize image processing techniques such as edge detection (e.g., Canny and Sobel operators), filtering, and thresholding to identify cracks. However, the actual environment of chemical plants is extremely complex, with chain conveyor surfaces often covered in dirt, rust, water stains, scratches, and uneven lighting and reflections. Traditional methods rely on manually designed underlying image features, making them highly sensitive to these complex background noises and environmental changes. This results in poor algorithm robustness, weak generalization ability, and difficulty in stable application in variable field environments.

[0005] Deep learning-based methods: In recent years, deep learning object detection algorithms, represented by convolutional neural networks (CNNs) (such as YOLO and Faster R-CNN), have shown great potential in crack detection tasks due to their powerful feature self-learning capabilities. However, these data-driven methods heavily rely on large amounts of diverse and accurately labeled training data. In the specific scenario of crack detection on a chemical plant transport chain, obtaining high-quality training data faces the following serious challenges:

[0006] Data scarcity and class imbalance: Real-world crack samples are low-probability events, making it difficult to collect a sufficient number of samples in a short period. This results in a small training dataset and a severe imbalance between positive and negative samples (cracked / non-cracked), making it difficult for the model to fully learn crack features, prone to overfitting, and poor generalization ability to unseen crack types.

[0007] Insufficient diversity in crack morphology and background: Real-world cracks vary greatly in shape, size, direction, and location. Furthermore, the background of the chain conveyor varies at different workstations and with different cleanliness levels. Limited real-world data cannot cover all possible scenarios, causing a sharp decline in the model's detection performance when faced with cracks of rare morphology or against complex backgrounds.

[0008] Limitations of Data Synthesis Methods: To address the problem of data scarcity, existing technologies often employ data augmentation or synthesis techniques. Simple geometric transformations (such as rotation and flipping) can only increase data diversity to a limited extent. Furthermore, some "copy-and-paste" synthesis methods often neglect perspective relationships between cracks and the background, lighting consistency, and boundary blending issues, resulting in images with noticeable edge artifacts. When models learn from this low-quality synthesized data, they may learn artifact features rather than genuine crack features, thus reducing detection performance in real-world scenes.

[0009] In summary, existing technologies for solving the problem of crack detection in chemical plant transport chain plates generally face the bottleneck of insufficient robustness and generalization ability of deep learning models due to the "limited quantity, low quality, and poor diversity" of effective training data. Summary of the Invention

[0010] To address the aforementioned technical issues, this technical solution provides a highly robust hybrid data augmentation synthesis method that can generate a large amount of high-quality, highly diverse, and highly realistic synthetic detection target data at low cost and high efficiency. It can realistically simulate the random morphology, spatial distribution, and seamless integration with complex industrial backgrounds of the detection targets, thereby effectively improving the robustness and generalization ability of the detection model in real industrial environments; and effectively solving the technical problems.

[0011] This invention is achieved through the following technical solution:

[0012] A robust hybrid data augmentation synthesis method includes the following steps:

[0013] Step 1: Preprocess the input raw image to generate an accurate binary mask;

[0014] Step 2: Apply morphological transformations to process the original detection target material generated in Step 1 to simulate the different widths and shapes of detection targets in the real world due to factors such as formation time and stress magnitude;

[0015] Step 3: Perform geometric transformations on the original target materials extracted in Step 1. Geometric transformations include, but are not limited to, rotation and shearing operations to generate a large number of visually different but semantically identical target samples, enriching the diversity of the training dataset.

[0016] Step 4: On the background image, click on the four vertices in sequence through the interactive interface. Define an arbitrary quadrilateral as the stress concentration region; the coordinates of these four vertices... This data is recorded by the system and saved for easy retrieval.

[0017] Step 5: Realistically embed the detection target onto a single background image: Determine the image-level detection target density through a probability-weighted random allocation mechanism; use a matching algorithm to paste the detection target samples onto the background image, randomly generating high-quality irregular detection targets at stress concentration locations to simulate real fault scenarios; and introduce a boundary overflow handling mechanism to ensure that the detection target does not exceed the ROI or image boundary, while ensuring that the position of the detection target corresponds to the image and that the label does not change; within the specified stress concentration area, the detection targets are placed one by one through geometric and logical constraints, including but not limited to setting constraints for specifying the placement direction and boundary overflow truncation, generating a synthetic image that is visually realistic, geometrically conforms to perspective, and has a reasonable spatial distribution;

[0018] Step 6: Perform the above steps 1 to 5 in batches on the specified background set and the detection target set respectively to generate a set of synthetic images and corresponding labels;

[0019] Step 7: Train the target detection prediction model based on the set of synthetic images and corresponding labels generated in Step 6, and output the visualization graphs and index files required for performance evaluation.

[0020] Furthermore, the specific operation method of step one is as follows:

[0021] Step 1.1: Perform grayscale processing on the original three-channel color detection target image to convert it into a single-channel grayscale image;

[0022] Step 1.2: Apply binarization thresholding to the grayscale image to generate a binary mask that defines the effective region for detecting the target. ;

[0023] Step 1.3: Segment the original image based on the binary mask of the effective region to obtain the original detection target material. The threshold segmentation process mentioned above follows the mathematical definition:

[0024] ;

[0025] in: Represents the pixel coordinates in the image; It is a grayscale image in coordinates Pixel intensity value at; It is the output binary mask in coordinates The value at; It is a preset intensity threshold; It is the maximum value assigned when the pixel intensity is greater than the threshold.

[0026] Furthermore, the morphological transformation operations described in step two include two basic pixel-level operations: erosion and dilation; among them, the erosion operation... and expansion operations Used to generate target samples for connected width contraction and expansion respectively; their mathematical expressions are as follows:

[0027] Corrosion-related formulas:

[0028] ;

[0029] in, Represents a set Perform corrosion operation; This represents the coordinates of a pixel in the image; Indicates that the structural element Translate the origin to the coordinate system The meaning of the entire formula is: the output image is at... The point is a foreground pixel (white) if and only if it is a foreground pixel. Structural elements placed at the center All pixels covered in the original image The pixels in the middle are all foreground pixels;

[0030] Expansion-related formulas:

[0031] ;

[0032] in, Represents a set Perform an expansion operation; Indicates will Reflect it relative to its origin, then translate its origin to a coordinate system. The meaning of the entire formula is: the output image is at... The point is a foreground pixel (white) if and only if it is a foreground pixel. Structural elements placed at the center Within the covered area, at least one pixel is in the original image. The middle pixel is also a foreground pixel.

[0033] Furthermore, the geometric transformation described in step three also includes a random lateral truncation shearing transformation operation. This operation simulates the non-perpendicular morphology that the target may exhibit under different stress fields or viewpoints, enhancing the morphological diversity and realism of the synthesized target sample. The expression for the random lateral truncation shearing transformation is:

[0034] ;

[0035] in, The homogeneous transformation matrix represents the horizontal shearing (lateral truncation shearing) transformation, which is used to perform horizontal shearing deformation on the detected target in two-dimensional image space; This represents the horizontal shear factor, i.e., the slope.

[0036] Furthermore, the rotation described in step three simulates the morphology of the target object in nature under arbitrary growth directions and different observation angles, enhancing the directional diversity and realism of the synthesized target object samples; the expression for the rotation transformation is:

[0037] ;

[0038] in, It is a homogeneous rotation transformation matrix in a two-dimensional plane, used to rotate the target being detected; This represents the input parameter, i.e., the rotation angle.

[0039] Furthermore, the expression for the probability-weighted random allocation mechanism described in step five is:

[0040] ;

[0041] in, This indicates that the final generated image contains The probability of detecting a target; Indicates the preset "appearance" The relative probability of "one detection target"; This represents the sum of all prior weights and is used as a normalization factor.

[0042] Furthermore, the specific operation method of step five is as follows:

[0043] Step 5.1: Using the size of the bounding box aligned with the stress concentration region as the target, sample the scaling factor according to the truncated normal distribution to complete the scale transformation of the detection target image;

[0044] Step 5.2: Within the bounding box of the stress concentration area, randomly select a rectangular region as the target placement area, and calculate the perspective mapping matrix from the source image to the target placement area. Then, the image is transformed to obtain a perspective transformation result with the same size as the background image;

[0045] Step 5.3: Generate polygonal location labels for stress concentration areas, and perform a bitwise AND operation with the detected target mask after perspective to obtain the effective area of ​​the detected target after cropping, and then overlay it onto the background image to obtain the composite image;

[0046] Step 5.4: Calculate the axis-aligned bounding box of the currently placed detection target and perform IoU judgment with the bounding boxes of previously placed detection targets in the same stress concentration area. If the IoU exceeds the threshold, reject the current placement.

[0047] Step 5.5: Output and store the four coordinates of the placed target, the bounding box, and the corresponding stress concentration area identifier.

[0048] Furthermore, the scaling factor sampling based on the truncated normal distribution in step 5.1 to complete the scaling transformation of the target image is based on the scaling factor of the truncated normal distribution. Scaling is performed to complete the scale transformation of the detected target image; its expression is:

[0049] ;

[0050] in, , This indicates the width and height of the original target template. Indicates the scaling factor; , This indicates the final width and height of the detected target after scaling.

[0051] Furthermore, the perspective transformation described in step 5.2 is based on a multi-scale geometric matching algorithm, which calculates the source... Figure 4 point To the target placement area at four points Perspective mapping matrix This yields a perspective transformation result with the same size as the background image; the expression for the operation is:

[0052] Coordinate transformation formula:

[0053] ;

[0054] Formula for converting homogeneous coordinates to Cartesian coordinates:

[0055] ;

[0056] in, Represents the two-dimensional coordinates of any point in the source image; , This represents the new two-dimensional coordinates of the source point in the target image after perspective transformation; , Represents Cartesian coordinates , An intermediate representation; This represents a scaling factor; This represents the perspective transformation matrix, which fully defines the mapping relationship from the source plane to the target plane.

[0057] Furthermore, the synthesized image described in step 5.3 is obtained by acquiring a polygonal mask of the stress concentration region, performing a bitwise AND operation with the detected target mask after perspective viewing to obtain the effective area of ​​the detected target after cropping, and finally superimposing it onto the background image to obtain the synthesized image; its expression is:

[0058] ;

[0059] ;

[0060] ;

[0061] in, This indicates the final generated valid region mask in coordinates. Pixel value at; This indicates the target mask after perspective transformation in coordinates. Pixel value at; The stress concentration region mask is represented in coordinates Pixel value at; : indicates coordinates Normalized weights (alpha values) at each location; This indicates the final synthesized image in coordinates. The pixel color value at that location; Indicates the background image in coordinates The pixel color value at that location; : Represents the detected target texture image in coordinates after perspective transformation. The pixel color value at that location; This represents the perspective transformation matrix, which defines the mapping relationship from the source plane to the target plane.

[0062] Furthermore, the IoU determination in step 5.4 involves rejecting the placement if the IoU exceeds a threshold, calculated according to the following formula:

[0063] ;

[0064] in, This represents the intersection area of ​​the old and new boundary frames; The area of ​​the union;

[0065] Furthermore, the performance evaluation metrics output in step seven include: recall R, F1-score, and mAP; wherein the formula for calculating recall is:

[0066] ;

[0067] Wherein, TP indicates that the model correctly detected the target (detected target); FN indicates that the model failed to detect the actual target (detected target).

[0068] The formula for calculating the F1 score is:

[0069] ;

[0070] in, For accuracy, Recall rate;

[0071] The approximate formula for mAP calculated using the trapezoidal rule is as follows:

[0072] ;

[0073] in, For accuracy, This refers to the recall rate.

[0074] Beneficial effects

[0075] The highly robust hybrid data augmentation synthesis method proposed in this invention has the following advantages compared with existing technologies:

[0076] (1) This technical solution innovatively introduces a perspective transformation matching and boundary overflow processing mechanism based on stress concentration areas. The stress concentration areas interactively marked on the existing chain plate background image are randomly attached to the transformed cracks. Then, truncation and scaling rules are used to precisely constrain them within the specified physical area. This can accurately delineate the boundaries of the cracks without introducing other backgrounds. Visually and morphologically, they are indistinguishable from the original cracks. Compared with the traditional copy and paste method, which may cause edge artifacts and visual abruptness, and where crack truncation may even include the original background, this invention significantly improves the realism and quality of the synthesized data.

[0077] (2) This technical solution greatly enriches the diversity and complexity of the dataset: The original crack portions extracted from the external crack dataset based on mask features are geometrically (rotation, shearing) or morphologically (erosion, dilation) transformed to expand the subsequent merged samples. Then, the crack attachment in stress concentration areas is controlled through probability weighting to ensure it conforms to the real-world scenario. The significance of weighting lies in controlling the density of crack numbers, ensuring that the number of cracks attached to different stress areas conforms to a normal distribution based on different weights. This method can generate a massive number of crack samples with varying shapes, sizes, orientations, and quantities. Simultaneously, by combining these diverse cracks with real industrial backgrounds of different lighting and cleanliness levels, it effectively simulates the complex and ever-changing environment of a chemical plant, comprehensively covering various situations that may occur in real-world scenarios.

[0078] (3) This technical solution ensures the physical and logical rationality of the synthesized data: by placing cracks in stress concentration areas and introducing an IoU anti-overlap mechanism, specifically, during the random attachment process, the IoU of two different crack samples is calculated to ensure that the IoU is always less than 0, that is, it will not violate the normal state and the abnormal situation of two cracks in the same location will not occur. This invention ensures that the spatial distribution of the synthesized cracks is more in line with engineering reality and physical laws, making the synthesized data not only visually realistic but also logically more credible, thereby improving the effectiveness of the training data.

[0079] (4) This technical solution significantly improves the robustness and generalization ability of the detection model: Due to the use of high-quality, highly diverse, and highly realistic synthetic data generated by this method for training, the deep learning model can learn more essential and generalized crack features, rather than synthetic artifacts or limited sample features. At the same time, ablation experiments were conducted, and the basic model was used to predict corrosion expansion test set images, sheared dataset images, and rotated dataset images, respectively. The prediction model effect was compared with the dataset trained by the corresponding enhancement method. As shown in Table 1, the external model could not identify the cracks on the chain plate at all. On the contrary, the basic model had a recall of 0.6844, a mean average precision (MAP) of 0.8301, and an F1 score of 0.7986. The model evaluation results for the corrosion expansion test set showed that the corrosion expansion model improved the recall by 10.9% and the mean average precision (MAP) of the basic model. The model achieved a 6.9% improvement in recall and a 7.8% improvement in F1 score compared to the basic model for the rotating test set. For the rotating test set, the model evaluation results showed a 6.9% improvement in recall, a 3.7% improvement in mean precision (MAP), and a 6.7% improvement in F1 score compared to the basic model. For the shear test set, the model evaluation results showed a 15.6% improvement in recall, an 8.9% improvement in mean precision (MAP), and a 10.4% improvement in F1 score compared to the basic model. Therefore, when facing real, complex, and variable chemical field environments, the model's accuracy, recall, and overall robustness are significantly improved.

[0080] (5) This technical solution achieves efficient and low-cost data augmentation: This invention automates the entire synthesis process, which can quickly and cost-effectively generate tens of thousands of training samples with precise annotations, effectively solving the problems of difficult, costly and time-consuming real data collection in industrial scenarios, and providing strong data support for the application of deep learning in specific industrial fields. Attached Figure Description

[0081] Figure 1 This is a schematic diagram of the overall process of the present invention.

[0082] Figure 2This is a before-and-after comparison of mask extraction of crack material in an embodiment of the present invention.

[0083] Figure 3 The image shows crack material for morphological and geometric transformations in an embodiment of the present invention.

[0084] Figure 4 This is a diagram showing the crack synthesis result within the stress concentration area in an embodiment of the present invention.

[0085] Figures 5 to 8 This is a comparison chart showing the evaluation of models trained on the original dataset, the external dataset, and the enhanced dataset in this embodiment of the invention.

[0086] Figure 9 This is a visualization example of the predictions of various models in the embodiments of the present invention. Detailed Implementation

[0087] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments. The described embodiments are merely some embodiments of the present invention, and not all embodiments. Various modifications and improvements made to the technical solutions of the present invention by those skilled in the art without departing from the design concept of the present invention should fall within the protection scope of the present invention.

[0088] Example 1:

[0089] like Figure 1 As shown, a highly robust hybrid data augmentation synthesis method is disclosed, which is a method for enhancing the training data of target detection models. More specifically, this embodiment discloses a highly robust hybrid data augmentation synthesis method for a specific industrial scenario, namely, crack detection of conveyor chain plates in chemical plants. This method can generate a large amount of high-quality synthetic crack data that is visually realistic and credible, highly diverse in morphology, and spatially consistent with physical logic at low cost and high efficiency. This solves the problem of insufficient robustness and generalization ability of deep learning models caused by the "small quantity, low quality, and poor diversity" of training data; and solves the key technical bottlenecks faced by existing technologies in the crack detection task of conveyor chain plates in chemical plants.

[0090] The data augmentation synthesis method includes the following steps:

[0091] Step 1: Preprocess the input raw image to generate an accurate binary mask. This step aims to extract a clean crack template from the raw crack image. Prepare a set of raw crack images as input, such as... Figure 2 As shown, these images typically have the following characteristics: a relatively simple background and clear crack features.

[0092] The input original color crack image is converted to grayscale, and then binarized thresholding is used to generate a binary mask that accurately identifies the crack region. This step provides a clean crack base template for subsequent morphological and geometric transformations.

[0093] Step 1.1: Perform grayscale processing on the input original three-channel color crack image to convert it into a single-channel grayscale image.

[0094] Step 1.2: Apply binarization thresholding to the grayscale image to generate a binary mask that accurately identifies the crack region. This provides a clean crack-based template for subsequent morphological and geometric transformations.

[0095] Step 1.3: Segment the original image based on the binary mask of the effective region to obtain the original detection target material. The threshold segmentation process mentioned above follows the mathematical definition:

[0096] ;

[0097] in: Represents the pixel coordinates in the image; It is a grayscale image in coordinates Pixel intensity value at; It is the output binary mask in coordinates The value at; It is a preset intensity threshold; It is the maximum value assigned when the pixel intensity is greater than the threshold.

[0098] Step Two: Enhanced Crack Morphology Diversity; This step aims to simulate the natural variations in crack width, such as... Figure 3 As shown.

[0099] Morphological transformations are applied to the original crack material generated in step 1 to simulate the different widths and morphologies of cracks in the real world due to factors such as formation time and stress magnitude; this mainly includes corrosion and expansion operations. By randomly selecting different structural elements and iteration numbers, the contraction or expansion of crack width caused by factors such as formation time and stress magnitude in the real world can be simulated, generating crack samples with diverse morphologies. The specific operation method is as follows:

[0100] Input: The binary mask image generated in step one.

[0101] Randomized morphological operations: Define a set of operations {erosion, dilation, no operation}. For each input mask, randomly select one of these operations.

[0102] Erosion: If erosion is selected, a structuring element B is defined, and the erosion operation is applied, where the number of iterations n can be randomly selected within a range to generate finer cracks.

[0103] Corrosion operation Used to generate crack samples with shrinking connectivity; their mathematical expressions are as follows: Corrosion-related formulas:

[0104] ;

[0105] in, Represents a set Perform corrosion operation; This represents the coordinates of a pixel in the image; Indicates that the structural element Translate the origin to the coordinate system The meaning of the entire formula is: the output image is at... The point is a foreground pixel (white) if and only if it is a foreground pixel. Structural elements placed at the center All pixels covered in the original image The pixels in the middle are all foreground pixels.

[0106] Dilation: If dilation is chosen, a structuring element B is also used, and the dilation operation is applied, where the number of iterations m is also randomly chosen to generate coarser cracks. Dilation operation Used to generate expanded crack samples; their mathematical expressions are as follows:

[0107] Expansion-related formulas:

[0108] ;

[0109] in, Represents a set Perform an expansion operation; Indicates will Reflect it relative to its origin, then translate its origin to a coordinate system. The meaning of the entire formula is: the output image is at... The point is a foreground pixel (white) if and only if it is a foreground pixel. Structural elements placed at the center Within the covered area, at least one pixel is in the original image. The middle pixel is also a foreground pixel.

[0110] Output: After this step, a raw crack mask can be used to generate multiple masks with different widths, which together form a morphologically enhanced crack template library.

[0111] Step 3: Enhancing Crack Geometric Diversity. This step aims to simulate different observation angles and crack growth directions, such as... Figure 3 As shown.

[0112] Geometric transformations are performed on the original target materials extracted in step 1. These transformations include, but are not limited to, rotation and shearing operations, generating a large number of visually different but semantically identical target samples, thus enriching the diversity of the training dataset. The specific operation method is as follows:

[0113] Input: The binary mask image generated in step 1.

[0114] Random geometric transformation: Apply the following transformations sequentially or randomly to the input mask:

[0115] Rotation: Calculate the rotation matrix and apply it to the mask to generate cracks in arbitrary directions. To prevent image content from being cropped after rotation, the output image size usually needs to be adjusted to fit the rotated boundaries.

[0116] Rotation simulates the morphology of a target object in nature under arbitrary growth directions and different observation angles, enhancing the directional diversity and realism of the synthesized target object samples; the expression for the rotation transformation is:

[0117] ;

[0118] in, It is a homogeneous rotation transformation matrix in a two-dimensional plane, used to rotate the target being detected; This represents the input parameter, i.e., the rotation angle.

[0119] Shear: Set a random transverse shear factor sh_x. Apply a shear transformation to make the crack appear tilted, simulating the observation effect from a non-perpendicular viewpoint.

[0120] Shearing is a random lateral truncation shearing transformation operation. This operation simulates the non-vertical shape that the target may exhibit under different stress fields or viewpoints, enhancing the morphological diversity and realism of the synthesized target samples. The expression for the random lateral truncation shearing transformation is:

[0121] ;

[0122] in, The homogeneous transformation matrix represents the horizontal shearing (lateral truncation shearing) transformation, which is used to perform horizontal shearing deformation on the detected target in two-dimensional image space; This represents the horizontal shear factor, i.e., the slope.

[0123] Output: This step greatly expands the size and diversity of the crack template library, containing crack samples of various shapes, widths, directions and angles.

[0124] Step 4: Calibrating the stress concentration region; this step provides spatial constraints for subsequent precise synthesis, such as... Figure 4 As shown. The specific operation method is as follows:

[0125] Input: A background image to be composited, which is a real photograph of a chemical plant transport chain plate without cracks.

[0126] Interactive calibration: Develop a simple graphical user interface (GUI). The operator opens a background image in this interface and uses the mouse to click sequentially on the four key vertices of stress concentration on the chain plate. This allows us to define a stress concentration region of an arbitrary quadrilateral.

[0127] Data storage: The system records the coordinates of these four vertices. Then associate it with the background image file name and save it.

[0128] Step 5: Probability-weighted realistic crack synthesis and embedding; aiming to realistically embed the processed crack sample into the stress concentration region of the background image; such as Figure 4 As shown.

[0129] Determine the number of cracks: Based on the probability-weighted random distribution formula, determine the total number of cracks generated on a single background image. The probability-weighted random distribution formula for cracks is as follows:

[0130] ;

[0131] In the above formula, This indicates that the final generated image contains The probability of detecting a target; Indicates the preset "appearance" The relative probability of "one detection target"; This represents the sum of all prior weights and is used as a normalization factor.

[0132] This mechanism allows for the simulation of different damage levels, ranging from crack-free to multi-cracked. Then, for each crack to be placed, random sampling is performed according to a probability distribution to obtain the number of cracks k to be synthesized from the image. Each crack is synthesized iteratively (if k > 0): for i from 1 to k, the following sub-steps are executed:

[0133] Step 5.1: Multi-scale transformation: Using the size of the bounding box aligned with the stress concentration region as the target, the scaling factor is sampled according to the truncated normal distribution to complete the scale transformation of the detection target map.

[0134] Randomly select a crack template and its corresponding original texture map from the enhanced crack library generated in step three.

[0135] Calculate the width of the axis-aligned bounding box of the stress concentration region. and height Set the baseline scaling factor Sampling scaling factor from truncated normal distribution The selected crack template and texture map are scaled by a factor s; the scaling factor is based on a truncated normal distribution. Scaling is performed to complete the scale transformation of the detected target image. The formulas involved are as follows:

[0136] ;

[0137] in, , This indicates the width and height of the original target template. Indicates the scaling factor; , This indicates the final width and height of the detected target after scaling.

[0138] Step 5.2: Perspective Transformation: Perspective transformation is based on a multi-scale geometric matching algorithm, which calculates the source... Figure 4 point To the target placement area at four points Perspective mapping matrix Within the bounding box of the stress concentration area, a rectangular region is randomly selected as the target placement area, and the perspective mapping matrix from the source image to the target placement area is calculated. Then, a perspective transformation result with the same size as the background image is obtained.

[0139] Source Quadrilateral The four corner points of the scaled crack template are {(0, 0), (w, 0), (w, h), (0, h)}.

[0140] A target rectangular region is randomly generated within the bounding box of the stress concentration region, and its vertices are then slightly randomly perturbed to obtain the target quadrilateral. Then calculate the perspective transformation matrix. The scaled crack template and texture map are then combined according to the matrix. Perform a perspective transformation, outputting a perspective-transformed image and a perspective mask with the same dimensions as the background image. The relevant formulas are as follows:

[0141] Coordinate transformation formula:

[0142] ;

[0143] Formula for converting homogeneous coordinates to Cartesian coordinates:

[0144] ;

[0145] in, Represents the two-dimensional coordinates of any point in the source image; , This represents the new two-dimensional coordinates of the source point in the target image after perspective transformation; , Represents Cartesian coordinates , An intermediate representation; This represents a scaling factor; This represents the perspective transformation matrix, which fully defines the mapping relationship from the source plane to the target plane.

[0146] Step 5.3: Boundary Constraints and Blending: Create a completely black image of the same size as the background image. Using the stress concentration region vertices (P1, P2, P3, P4) saved in Step 4, draw a solid white polygon and generate a mask. .

[0147] The final effective crack region mask is obtained through a bitwise AND operation. This operation ensures that the crack after perspective transformation does not extend beyond the boundary of the stress concentration region.

[0148] Use As an alpha channel, it reveals the crack texture after perspective. blend into background image .

[0149] The composite image is obtained by acquiring a polygonal mask of the stress concentration region, performing a bitwise AND operation with the mask of the detected target after perspective mapping to obtain the effective area of ​​the detected target after cropping, and finally superimposing it onto the background image; its expression is:

[0150] ;

[0151] ;

[0152] ;

[0153] in, This indicates the final generated valid region mask in coordinates. Pixel value at; This indicates the target mask after perspective transformation in coordinates. Pixel value at; The stress concentration region mask is represented in coordinates Pixel value at; : indicates coordinates Normalized weights (alpha values) at each location; This indicates the final synthesized image in coordinates. The pixel color value at that location; Indicates the background image in coordinates The pixel color value at that location; : Represents the detected target texture image in coordinates after perspective transformation. The pixel color value at that location; This represents the perspective transformation matrix, which defines the mapping relationship from the source plane to the target plane.

[0154] Step 5.4: Anti-overlap mechanism: Calculate the axis-aligned bounding box of the currently placed detection target and perform IoU judgment with the bounding boxes of already placed detection targets in the same stress concentration area. If the IoU exceeds the threshold, the current placement is rejected.

[0155] calculate The smallest bounding rectangle is used as the bounding box for the current crack placement. The intersection-union ratio (IoU) is calculated by comparing it with the list of bounding boxes where cracks have been placed on the image.

[0156] Set an IoU threshold. If any IoU value exceeds this threshold, abandon the current placement attempt and repeat the above steps to find a new placement location. The IoU determination rejects the placement attempt when the IoU exceeds the threshold, and is calculated using the following formula:

[0157] ;

[0158] in, This represents the intersection area of ​​the old and new boundary frames; Let be the area of ​​the union.

[0159] Step 5.5: Label generation: If the placement is successful, record the label information of the crack placement box, output and store the four coordinates of the detected target, the bounding box, and the corresponding stress concentration area identifier; prepare for subsequent training of the target detection model.

[0160] Step 6: Batch generation: Perform the above steps 1 to 5 on the specified background set and the detection target set in batches to generate a set of composite images and corresponding labels;

[0161] Write an automated script that takes as input a directory containing multiple background images and their stress concentration region files, and a directory containing the original crack images. The script iterates through all the background images, performing step 5 for each image to generate a specified number of composite images and their corresponding label files.

[0162] Step 7: Train the target detection prediction model based on the set of synthetic images and corresponding labels generated in Step 6, and output the visualization graphs and index files required for performance evaluation.

[0163] like Figures 5 to 8 As shown, the performance of the model trained on the generated dataset is evaluated.

[0164] Dataset preparation: Divide the synthetic images generated in step 6 into training set, validation set and test set.

[0165] Model training: Select an object detection model and train the model on the prepared training set.

[0166] Performance evaluation: After training is complete, evaluate the model performance on the test set.

[0167] Calculate and record the defined metrics. The metrics output by the model performance evaluation include: recall (R), F1-score, and mAP.

[0168] Recall: Evaluates the model's ability to detect all true cracks. The formula for calculating recall is:

[0169] ;

[0170] In this context, TP indicates that the model correctly detected the target (detected target); FN indicates that the model failed to detect the actual target (detected target).

[0171] F1-score: The harmonic mean of precision and recall, used to comprehensively evaluate model performance. The formula for calculating F1-score is:

[0172] ;

[0173] in, For accuracy, This refers to the recall rate.

[0174] mAP: The core evaluation metric for object detection tasks, comprehensively measuring the model's accuracy in localization and classification at different confidence thresholds. The approximate formula for calculating mAP using the trapezoidal rule is:

[0175] ;

[0176] in, For accuracy, This refers to the recall rate.

[0177] Output results: Predict the images in the test set, visualize and save the detection results (images with bounding boxes), such as... Figure 6 As shown, this is used to intuitively analyze the detection performance of the model.

[0178] Ablation experiments were conducted on the resulting datasets. YOLOv8 model weight files were used for all experiments. The basic model was trained based on the original data, and the corresponding target detection models were trained based on the data after erosion, dilation, shearing, and rotation transformations.

[0179] Meanwhile, a corresponding target detection model is trained using an external crack dataset, and the trained external model is used to test the performance of the original data and the basic model for comparison.

[0180] Finally, the basic model was used to predict the images in the erosion and dilation test set, the shearing test set, and the rotation test set, respectively. The results were compared with the prediction models trained on the datasets enhanced by the corresponding enhancement methods. The results are shown in Table 1.

[0181] Table 1

[0182]

[0183] In this embodiment, as shown in Table 1: the external model completely failed to identify the cracks on the chain plate, while the basic model achieved a recall of 0.6844, a mean average precision (MAP) of 0.8301, and an F1 score of 0.7986. For the corrosion expansion test set, the model evaluation results showed that the corrosion expansion model improved recall by 10.9%, MAP by 6.9%, and F1 score by 7.8% compared to the basic model. For the rotation test set, the model evaluation results showed that the rotation model improved recall compared to the basic model. The mean precision (MAP) improved by 3.7%, and the F1 score improved by 6.7%. For the shear test set, the model evaluation results showed that the shear model improved recall by 15.6%, mean precision (MAP) by 8.9%, and the F1 score by 10.4% compared to the base model. This invention successfully generated a large amount of high-quality, highly diverse synthetic data and demonstrated that the model trained using this data significantly improved its robustness and generalization ability in crack detection when faced with real images of chemical plant transport chain plates with complex backgrounds and variable environments.

[0184] The above embodiments are only for illustrating the technical concept and features of the present invention, and are intended to enable those skilled in the art to understand the content of the present invention and implement it accordingly. They should not be construed as limiting the scope of protection of the present invention. All equivalent changes or modifications made in accordance with the spirit and essence of the present invention should be covered by the present invention.

Claims

1. A highly robust hybrid data augmentation synthesis method, characterized in that: Including the following steps: Step 1: Preprocess the input raw image to generate an accurate binary mask; Step 2: Apply morphological transformations to process the original detection target material generated in Step 1 to simulate the different widths and shapes of detection targets in the real world due to factors such as formation time and stress magnitude; Step 3: Perform geometric transformations on the original detection target materials extracted in Step 1. Geometric transformations include, but are not limited to, rotation and shearing operations to generate a large number of visually different but semantically identical detection target samples, enriching the diversity of the training dataset. Step 4: On the background image, click on the four vertices in sequence through the interactive interface. Define an arbitrary quadrilateral as the stress concentration region; the coordinates of these four vertices... This data is recorded by the system and saved for easy retrieval. Step 5: Realistically embed the detection target on a single background image: Determine the image-level detection target density through a probability-weighted random allocation mechanism; The matching algorithm is used to paste the sample of the detected target onto the background image, and high-quality irregular detected targets are randomly generated at the stress concentration location to simulate real fault scenarios. Furthermore, a boundary overflow handling mechanism is introduced to ensure that the detected target does not exceed the ROI or image boundary, while ensuring that the position of the detected target corresponds to the image and that the label does not change. In addition, within the specified stress concentration area, the detected targets are placed one by one through geometric and logical constraints. The constraints include, but are not limited to, the setting of constraint rules for specifying the placement direction and boundary overflow truncation, to generate a synthetic image that is visually realistic, geometrically conforms to perspective, and is spatially reasonable. Step 6: Perform the above steps 1 to 5 in batches on the specified background set and the detection target set respectively to generate a set of synthetic images and corresponding labels; Step 7: Train the target detection prediction model based on the set of synthetic images and corresponding labels generated in Step 6, and output the visualization graphs and index files required for performance evaluation.

2. The highly robust hybrid data augmentation synthesis method according to claim 1, characterized in that: The specific operation method for step one is as follows: Step 1.1: Perform grayscale processing on the original three-channel color detection target image to convert it into a single-channel grayscale image; Step 1.2: Apply binarization thresholding to the grayscale image to generate a binary mask that defines the effective region for detecting the target. ; Step 1.3: Segment the original image based on the binary mask of the effective region to obtain the original detection target material. The threshold segmentation process mentioned above follows the mathematical definition: ; in: Represents the pixel coordinates in the image; It is a grayscale image in coordinates Pixel intensity value at; It is the output binary mask in coordinates The value at; It is a preset intensity threshold; It is the maximum value assigned when the pixel intensity is greater than the threshold.

3. The highly robust hybrid data augmentation synthesis method according to claim 1, characterized in that: The morphological transformation operations described in step two include two basic pixel-level operations: erosion and dilation; among them, the erosion operation... and expansion operations Used to generate target samples for connected width contraction and expansion respectively; their mathematical expressions are as follows: Corrosion-related formulas: ; in, Represents a set Perform corrosion operation; This represents the coordinates of a pixel in the image; Indicates that the structural element Translate the origin to the coordinate system The meaning of the entire formula is: the output image is at... The point is a foreground pixel (white) if and only if it is a foreground pixel. Structural elements placed at the center All pixels covered in the original image The pixels in the middle are all foreground pixels; Expansion-related formulas: ; in, Represents a set Perform an expansion operation; Indicates will Reflect it relative to its origin, then translate its origin to a coordinate system. The meaning of the entire formula is: the output image is at... The point is a foreground pixel (white) if and only if it is a foreground pixel. Structural elements placed at the center Within the covered area, at least one pixel is in the original image. The middle pixel is also a foreground pixel.

4. The highly robust hybrid data augmentation synthesis method according to claim 1, characterized in that: Step 3's geometric transformation also includes a random lateral truncation shearing transformation, which simulates the non-perpendicular morphology that the target may exhibit under different stress fields or viewpoints, enhancing the morphological diversity and realism of the synthesized target samples; the expression for the random lateral truncation shearing transformation is: ; in, The homogeneous transformation matrix represents the horizontal shearing (lateral truncation shearing) transformation, which is used to perform horizontal shearing deformation on the detected target in two-dimensional image space; This represents the horizontal shear factor, i.e., the slope. The rotation described simulates the morphology of the target object in nature under arbitrary growth directions and different observation angles, enhancing the directional diversity and realism of the synthesized target object samples; the expression for the rotation transformation is: ; in, It is a homogeneous rotation transformation matrix in a two-dimensional plane, used to rotate the target being detected; This represents the input parameter, namely the rotation angle.

5. The highly robust hybrid data augmentation synthesis method according to claim 1, characterized in that: The expression for the probability-weighted random allocation mechanism described in step five is: ; in, This indicates that the final generated image contains The probability of detecting a target; Indicates the preset "appearance" The relative probability of "one detection target"; This represents the sum of all prior weights and is used as a normalization factor. The specific operation method for step five is as follows: Step 5.1: Using the size of the bounding box aligned with the stress concentration region as the target, sample the scaling factor according to the truncated normal distribution to complete the scale transformation of the detection target image; Step 5.2: Within the bounding box of the stress concentration area, randomly select a rectangular region as the target placement area, and calculate the perspective mapping matrix from the source image to the target placement area. Then, the image is transformed to obtain a perspective transformation result with the same size as the background image; Step 5.3: Generate polygonal location labels for stress concentration areas, and perform a bitwise AND operation with the detected target mask after perspective to obtain the effective area of ​​the detected target after cropping, and then overlay it onto the background image to obtain the composite image; Step 5.4: Calculate the axis-aligned bounding box of the currently placed detection target and perform IoU judgment with the bounding boxes of previously placed detection targets in the same stress concentration area. If the IoU exceeds the threshold, reject the current placement. Step 5.5: Output and store the four coordinates of the placed target, the bounding box, and the corresponding stress concentration area identifier.

6. The highly robust hybrid data augmentation synthesis method according to claim 5, characterized in that: Step 5.1, which samples the scaling factor based on a truncated normal distribution to complete the scaling transformation of the target image, is a scaling factor based on a truncated normal distribution. Scaling is performed to complete the scale transformation of the detected target image; Its expression is: ; in, , This indicates the width and height of the original target template. Indicates the scaling factor; , This indicates the final width and height of the detected target after scaling.

7. The highly robust hybrid data augmentation synthesis method according to claim 5, characterized in that: The perspective transformation described in step 5.2 is based on a multi-scale geometric matching algorithm, which calculates four points in the source image. To the target placement area at four points Perspective mapping matrix This yields a perspective transformation result with the same size as the background image; the expression for the operation is: Coordinate transformation formula: ; The formula for converting homogeneous coordinates to Cartesian coordinates: ; in, Represents the two-dimensional coordinates of any point in the source image; , This represents the new two-dimensional coordinates of the source point in the target image after perspective transformation; , Represents Cartesian coordinates , An intermediate representation; This represents a scaling factor; This represents the perspective transformation matrix, which fully defines the mapping relationship from the source plane to the target plane.

8. The highly robust hybrid data augmentation synthesis method according to claim 5, characterized in that: The synthesized image described in step 5.3 is obtained by acquiring a polygonal mask of the stress concentration region, performing a bitwise AND operation with the detected target mask after perspective viewing to obtain the effective area of ​​the detected target after cropping, and finally superimposing it onto the background image to obtain the synthesized image; its expression is: ; ; ; in, This indicates the final generated valid region mask in coordinates. Pixel value at; This indicates the target mask after perspective transformation in coordinates. Pixel value at; The stress concentration region mask is represented in coordinates Pixel value at; : indicates coordinates Normalized weights (alpha values) at each location; This indicates the final synthesized image in coordinates. The pixel color value at that location; Indicates the background image in coordinates The pixel color value at that location; : Represents the detected target texture image in coordinates after perspective transformation. The pixel color value at that location; This represents the perspective transformation matrix, which defines the mapping relationship from the source plane to the target plane.

9. The highly robust hybrid data augmentation synthesis method according to claim 5, characterized in that: The IoU determination in step 5.4 involves rejecting the placement if the IoU exceeds a threshold, calculated according to the following formula: ; in, This represents the intersection area of ​​the old and new boundary frames; Let be the area of ​​the union.

10. The highly robust hybrid data augmentation synthesis method according to claim 1, characterized in that: The performance evaluation metrics output in step seven include: recall (R), F1-score, and mAP; the formula for calculating recall is: ; Wherein, TP indicates that the model correctly detected the target (detected target); FN indicates that the model failed to detect the actual target (detected target). The formula for calculating the F1 score is: ; in, For accuracy, Recall rate; The approximate formula for mAP calculated using the trapezoidal rule is as follows: ; in, For accuracy, This refers to the recall rate.