Rotating shaft simulation data enhancement method and system combined with generative adversarial network
By determining the required parameters for shaft data and configuring a generative adversarial network, simulation data with consistent features is generated and filtered, solving the problems of scarce shaft data and inconsistent features in the existing technology, and improving the accuracy of shaft quality inspection and fault diagnosis.
Patent Information
- Application Number
- CN202511769183.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-28
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2045-11-28
AI Technical Summary
Existing technologies struggle to generate simulation data that is highly consistent with real shaft data and has rich features, failing to effectively meet the data requirements for shaft quality inspection and control. Furthermore, acquiring data on rare defect types and shafts made of various materials is costly and time-consuming.
By determining the set of parameters required for the pivot data, configuring the adaptation parameters of the generative adversarial network, and using a real pivot data sample set to drive the generative adversarial network, simulation data with consistent features are generated and selected, and integrated to form a pivot enhancement data set.
The generated simulation data is highly consistent with the real data, expanding the scale and diversity of shaft data, improving the accuracy of shaft quality inspection and fault diagnosis, and enhancing the quality control capabilities of industrial production.
Smart Images

Figure CN121389804A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, in particular to a shaft simulation data enhancement method and system combining a generative adversarial network. BACKGROUND
[0002] In industrial production, as a key moving part, the performance and quality of the shaft directly affect the running stability and reliability of the entire equipment. Precise detection and analysis of the shaft are important links to ensure its quality, which requires a large amount of and diversified data support. At present, shaft data is mainly obtained by detection and sampling in actual production, but the above-mentioned method has many limitations. On the one hand, the occurrence of various defect types of the shaft in actual production is uncertain, and it is difficult to collect sufficient and comprehensive defect sample data in a short time, especially the data of some rare but dangerous defect types are even more scarce. On the other hand, shafts of different materials have large differences in performance during production and use, and it takes a lot of time and cost to obtain sufficient data covering shafts of various materials. In addition, most of the existing data enhancement methods are simple and cannot generate simulation data that is highly consistent with real shaft data and has rich features, which cannot effectively meet the data needs of in-depth analysis and precise detection of the shaft, thereby limiting the development of shaft quality detection and control technology. SUMMARY
[0003] In view of the above-mentioned problems, in combination with the first aspect of the present application, the embodiments of the present application provide a shaft simulation data enhancement method combining a generative adversarial network, which comprises: determining a set of shaft data requirement parameters to be enhanced, the set of shaft data requirement parameters to be enhanced comprising a shaft surface defect type, a shaft material representation parameter, and a data sample quantity requirement; configuring shaft adaptation parameters of the generative adversarial network, the shaft adaptation parameters of the generative adversarial network comprising a number of shaft feature mapping layers of a generator, a shaft data discrimination threshold of a discriminator, and an iteration termination condition of network training; obtaining a real shaft data sample set, driving the generative adversarial network to run according to the shaft adaptation parameters of the generative adversarial network based on the real shaft data sample set, and generating a preliminary shaft simulation data set; performing feature consistency comparison processing on the preliminary shaft simulation data set and the real shaft data sample set, eliminating preliminary shaft simulation data with inconsistent features, and obtaining an effective shaft simulation data set with consistent features; integrating the effective shaft simulation data set and the real shaft data sample set to form a final shaft enhancement data set, the final shaft enhancement data set comprising real shaft data samples and effective shaft simulation data.
[0004] In still another aspect, the embodiments of the present application also provide a rotating shaft simulation data enhancement system combined with a generative adversarial network, comprising a processor, a machine readable storage medium, the machine readable storage medium being connected with the processor, the machine readable storage medium being used for storing programs, instructions or codes, and the processor being used for running the programs, instructions or codes in the machine readable storage medium to realize the above-mentioned method.
[0005] Based on the above aspects, the embodiments of the present application determine the rotating shaft data requirement parameter set to be enhanced, explicitly determine key elements such as rotating shaft surface defect type, material characterization parameter and data sample quantity requirement, and ensure that the generated simulation data can meet the requirements of rotating shaft data diversity and pertinence in actual application. The rotating shaft adaptation parameters of the generative adversarial network are configured, including the rotating shaft feature mapping layer number of the generator, the rotating shaft data discrimination threshold of the discriminator and the iteration termination condition of network training, so that the generative adversarial network can better adapt to the characteristics of the rotating shaft data, and improve the quality and efficiency of the generated simulation data. The generative adversarial network is driven based on the real rotating shaft data sample set to generate a preliminary rotating shaft simulation data set, which fully utilizes the information of the real data and lays a foundation for subsequent data enhancement. The preliminary rotating shaft simulation data set and the real rotating shaft data sample set are subjected to feature consistency comparison processing, and the data with inconsistent features are removed, so as to ensure the high consistency of the generated simulation data and the real data in features, and enhance the credibility and usability of the simulation data. Finally, the effective rotating shaft simulation data set and the real rotating shaft data sample set are integrated to form a rotating shaft enhancement data set, which effectively expands the scale and diversity of the rotating shaft data, provides more abundant and more accurate training data for quality detection, fault diagnosis and performance analysis of the rotating shaft, and helps to improve the research level of rotating shaft related technology and the quality control ability of industrial production. BRIEF DESCRIPTION OF DRAWINGS
[0006] Figure 1 is the execution flow diagram of the rotating shaft simulation data enhancement method combined with the generative adversarial network provided by the embodiments of the present application.
[0007] Figure 2 is the hardware architecture diagram of the rotating shaft simulation data enhancement system combined with the generative adversarial network provided by the embodiments of the present application. DETAILED DESCRIPTION
[0008] The present application will be specifically described below in combination with the drawings of the specification, Figure 1 is the flow diagram of the rotating shaft simulation data enhancement method combined with the generative adversarial network provided by an embodiment of the present application, and the rotating shaft simulation data enhancement method combined with the generative adversarial network will be described in detail below.
[0009] Step S110: Determine the set of to-be-enhanced shaft data requirement parameters, which includes shaft surface defect type, shaft material characterization parameter, and data sample quantity requirement.
[0010] In this embodiment, the training data enhancement of the detection model for complex surface defects of industrial motor shafts is taken as the scenario, and the complex surface defects include composite defects (such as scratch + rust composite, depression + coating shedding composite), micro-crack defects (small width and fuzzy edge), and fuzzy edge defects (no obvious boundary between the defect and the normal region). The determination of the set of requirement parameters needs to go through the following steps: existing data evaluation, defect type analysis, material parameter statistics, and sample quantity calculation, and the details are as follows.
[0011] First, the existing real shaft data sample set used to train the defect detection model is called, and the sample set includes shaft surface images collected under different working conditions and corresponding labeled data. The metadata information of the sample set is read through the data management module to clearly indicate the labeled defect types, material parameters, and total sample quantity in the sample set, which provides a basis for subsequent determination of requirement parameters.
[0012] Step S111: Analyze the coverage of shaft surface defect types in the existing real shaft data sample set to obtain a defect type coverage result, which includes the covered defect types and the sample quantity of each covered defect type.
[0013] The defect type coverage analysis is completed through the following sub-steps: Step S1111: Obtain all data samples in the existing real shaft data sample set, and read the defect labeling information of each data sample one by one, which includes the shaft surface defect type corresponding to the sample.
[0014] Each sample in the existing real shaft data sample set is traversed, and the XML annotation file of each sample is read through the annotation analysis module to extract the “defect type” field information. For example, the “defect type” field in the annotation file of a sample is labeled as “scratch + rust composite”, and another sample is labeled as “micro-crack”. The defect type name corresponding to each sample is recorded to form a sample-defect type mapping table.
[0015] For samples that are not labeled or have ambiguous labels, mark them as “to be confirmed” and include them in the analysis after supplementing the labels through manual review to ensure that no unlabeled missing samples affect the accuracy of the coverage result.
[0016] Step S1112: Establish a defect type statistical table, and the columns of the defect type statistical table include the shaft surface defect type name, and the rows of the defect type statistical table are used to record the sample quantity of each defect type.
[0017] A two-dimensional statistical table is created in the data statistics module, the first column of the table is set as "rotating shaft surface defect type name", and all non-repeated defect types extracted from the annotation information are listed in turn, including single defect types (scratch, rust, indentation, coating peeling) and complex defect types (scratch + rust composite, indentation + coating peeling composite, micro-crack, edge blur defect). The second column of the table is set as "sample number", and the initial value is set to zero for subsequent counting statistics.
[0018] Step S1113: Traverse the defect annotation information of each data sample, find the corresponding defect type name column in the defect type statistical table, and accumulate the sample number of the corresponding defect type name column in the defect type statistical table.
[0019] According to the order of the sample-defect type mapping table, the defect type of each sample is matched with the type name in the statistical table one by one. If the matching is successful, the "sample number" value corresponding to the type is increased by one; if it is a composite defect type, it needs to be increased by one in the "sample number" column of the corresponding single defect type and composite defect type at the same time, for example, the "scratch + rust composite" sample needs to be increased by one in the "scratch", "rust" and "scratch + rust composite" columns respectively, to ensure the coverage of single defects and composite defects.
[0020] After the traversal is completed, it is checked whether the total sample number of the statistical table is consistent with the total sample number of the existing real rotating shaft data sample set. If it is not consistent, the matching and counting process is rechecked to correct the missing or repeated counting problems.
[0021] Step S1114: After completing the traversal of all data samples, the defect types with a sample number greater than 0 in the defect type statistical table are counted, and the defect types are determined as covered defect types.
[0022] The rows with a "sample number" column value greater than 0 in the statistical table are screened, and the corresponding "defect type name" is extracted to form a covered defect type list. For example, the list may include "scratch", "rust", "scratch + rust composite", "micro-crack", and the sample number of "indentation + coating peeling composite" and "edge blur defect" is 0, which is not included in the covered defect type list, indicating that these complex defect types are not covered in the existing sample set.
[0023] Step S1115: The covered defect types and the sample number corresponding to each covered defect type are arranged into structured data, which is the defect type coverage result.
[0024] The defect type coverage list and the corresponding sample quantity are combined to generate structured data in JSON format, including two fields: "defect type" and "sample quantity". For example, a structured data fragment is: [{"defect type": "scratch", "sample quantity":...}, {"defect type": "scratch + rust complex", "sample quantity":...}, {"defect type": "micro-crack", "sample quantity":...}].
[0025] The structured data is stored in a local database, and a visual statistical chart is generated to intuitively display the sample quantity distribution of each defect type, facilitating subsequent analysis of sample defect conditions.
[0026] Step S112: Based on the defect type coverage result, identify the defect type with a sample quantity less than a preset training requirement threshold, and determine the defect type as the target defect type to be supplemented. The target defect type to be supplemented belongs to the shaft surface defect type in the set of enhanced shaft data requirement parameters.
[0027] The preset training requirement threshold is retrieved, which is determined based on the recognition accuracy requirement of the defect detection model for each defect type. The threshold for different defect types can be adjusted according to their recognition difficulty. The threshold for complex defect types is usually higher than that for single defect types, as more samples are needed to support feature learning for complex defects.
[0028] Compare the sample quantity of each defect type in the defect type coverage result with the corresponding threshold: if the sample quantity of a certain defect type is less than its corresponding training requirement threshold, and the type is a key defect type that the detection model must recognize, then it is determined as the target defect type to be supplemented. For example, the sample quantity of "micro-crack" is less than its threshold, and the sample quantity of "scratch + rust complex" is also less than its threshold, and both are common fault defects of motor shafts, so both are determined as the target defect type to be supplemented.
[0029] The determined target defect type name and corresponding sample gap quantity (threshold minus existing sample quantity) are arranged into a target defect list, which is an important part of the requirement parameter set.
[0030] Step S113: Collect the shaft material representation parameters of all samples in the existing real shaft data sample set, and count the distribution frequency of each shaft material representation parameter to obtain the material parameter distribution statistical result, which includes the occurrence times and proportion of each material representation parameter.
[0031] The material quality characterization parameters of the rotating shaft include the material type (such as 45 steel, stainless steel, and alloy steel), the surface treatment method (such as turning, grinding, and chrome plating), the surface roughness parameter, the hardness parameter, and the like. These parameters directly affect the texture characteristics and defect performance of the surface of the rotating shaft. For example, the light reflection characteristics of stainless steel are different from those of 45 steel, and the defect morphology of the chrome-plated surface is different from that of the surface without chrome plating.
[0032] The material quality parameter statistics are completed through the following sub-steps: Step S1131: Each sample in the existing real rotating shaft data sample set is traversed, the metadata file of the sample is read, and the rotating shaft material quality characterization parameter field information in the metadata file is extracted.
[0033] The metadata file of each sample is stored in the JSON format and includes the fields of “material type”, “surface treatment method”, “surface roughness”, and “hardness”. For example, in the metadata of a certain sample, the “material type” is “45 steel”, the “surface treatment method” is “turning”, the “surface roughness” is a certain level, and the “hardness” is a certain level. The field values are extracted and associated with the sample ID to form a sample-material parameter mapping table.
[0034] Step S1132: For each material quality characterization parameter category, a parameter value statistics sub-table is established, and the number of occurrences of each parameter value in the sample set is counted.
[0035] The statistics sub-tables are respectively established according to the material quality characterization parameter categories. The “material type statistics sub-table” includes the columns of “material name” and “number of occurrences”; the “surface treatment method statistics sub-table” includes the columns of “treatment method name” and “number of occurrences”; the “surface roughness statistics sub-table” includes the columns of “roughness level” and “number of occurrences”; and the “hardness statistics sub-table” includes the columns of “hardness level” and “number of occurrences”.
[0036] According to the sample-material parameter mapping table, the “number of occurrences” of the corresponding parameter values in each sub-table is accumulated and counted. For example, if the “material type” of 100 samples is “45 steel”, the “number of occurrences” of “45 steel” in the “material type statistics sub-table” is recorded as 100.
[0037] Step S1133: The occurrence proportion of each material quality characterization parameter value is calculated, which is the ratio of the number of occurrences of the parameter value to the total number of samples in the sample set.
[0038] For each parameter value in each statistics sub-table, the “number of occurrences” is divided by the total number of samples in the existing real rotating shaft data sample set to obtain the occurrence proportion of the parameter value. For example, the total number of samples in the sample set is 500, and the number of occurrences of “stainless steel” is 50. Therefore, the occurrence proportion of “stainless steel” is 50 / 500, and the calculation result is rounded to two decimal places and filled in the newly added “occurrence proportion” column of each sub-table.
[0039] Step S1134: integrate the statistical sub-tables of all material characterization parameter categories to form a material parameter distribution statistical result.
[0040] The "material category statistical sub-table", "surface treatment method statistical sub-table", "surface roughness statistical sub-table", and "hardness statistical sub-table" are integrated into a complete structured document. The document is divided into chapters according to parameter categories, and each chapter contains corresponding sub-table data. The statistical result clearly presents the distribution of each material parameter, for example, the appearance ratio of "turning" surface treatment method is the highest, the appearance ratio of "chromium plating" treatment is the lowest, and the appearance ratio of "45 steel" material is more than half.
[0041] Step S114: based on the material parameter distribution statistical result, determine the material characterization parameter category that needs to be strengthened and supplemented, determine the material characterization parameter of the material characterization parameter category as the target material characterization parameter, and the target material characterization parameter belongs to the shaft material characterization parameter in the shaft data requirement parameter set to be enhanced.
[0042] Analyze the uniformity of the distribution of each material characterization parameter category: if the appearance ratio of each parameter value in a certain parameter category is too different, for example, the appearance ratio of "turning" in "surface treatment method" is too high, and the appearance ratio of "chromium plating" and "grinding" is too low, it indicates that the distribution of this parameter category is not balanced, which may lead to insufficient recognition accuracy of the detection model for the shaft surface defects corresponding to the low-occurrence-ratio parameters, and the low-occurrence-ratio parameters in this parameter category need to be strengthened and supplemented.
[0043] Determine the material characterization parameter category with uneven distribution as the category that needs to be strengthened and supplemented, and the low-occurrence-ratio parameter values under this category are the target material characterization parameters. For example, the appearance ratio of "chromium plating" and "grinding" in the "surface treatment method" category is too low, and "chromium plating" and "grinding" are determined as the target material characterization parameters; the appearance ratio of "alloy steel" in the "material category" category is too low, and "alloy steel" is determined as the target material characterization parameter.
[0044] Record the name, category, and current appearance ratio of the target material characterization parameter, and clearly indicate the target ratio to be achieved after supplementation (such as the ratio of each parameter value tends to be balanced), which provides a basis for generating simulation data later.
[0045] Step S115: according to the total amount requirement of the training sample of the shaft data driven model, combined with the sample quantity of the existing real shaft data sample set, calculate the number of simulation data samples that need to be supplemented, and determine the number of simulation data samples as the target sample quantity, which belongs to the data sample quantity requirement in the shaft data requirement parameter set to be enhanced.
[0046] Firstly, the total training sample requirement of the shaft defect detection model (such as a detection model based on a convolutional neural network) is determined, which is based on the number of network layers, the number of parameters, and the expected recognition accuracy of the model. For example, the model needs to input a certain number of samples to reach the preset accuracy.
[0047] The amount of simulation data needed to be supplemented is calculated by the formula "target sample number = total training sample requirement - sample number of existing real shaft data sample set". If the result is negative, it means that the existing sample number has met the total requirement, but the target defect type sample gap determined in step S112 and the target material characteristic parameter supplement demand determined in step S114 need to be combined to adjust the target sample number to ensure that the defect type gap is filled and the material parameter distribution is balanced.
[0048] For example, the total training sample requirement is a certain value, the existing sample number is a certain value, and the initial target sample number is calculated to be a certain value. However, the sample gap of the target defect type "micro crack" is a certain value, the sample gap of the target defect type "scratch + rust complex" is a certain value, and the sample gap of the target material characteristic parameter is a certain value. After comprehensive adjustment, the final target sample number needs to cover all gaps to ensure that the sample number of each target defect type after supplementation meets the training requirement threshold and the proportion of each target material characteristic parameter tends to be balanced.
[0049] Step S116: Integrate the target defect type to be supplemented, the target material characteristic parameter, and the target sample number to form a shaft data demand parameter set to be enhanced.
[0050] The target defect type to be supplemented (such as "micro crack" and "scratch + rust complex") determined in step S112, the target material characteristic parameter (such as "chromium plating", "grinding", and "alloy steel") determined in step S114, and the target sample number determined in step S115 are arranged into a JSON format demand parameter set.
[0051] The demand parameter set includes three first-level fields: "target defect type to be supplemented" (including defect type name, sample gap number, and training requirement threshold), "target material characteristic parameter" (including parameter name, category, current proportion, and target proportion), and "target sample number" (including total target number, defect type allocation number, and material parameter allocation number).
[0052] The demand parameter set is stored in the project database and pushed to the configuration module of the generative adversarial network as the core basis for network parameter configuration and simulation data generation.
[0053] Step S120: configuring the rotation axis adaptation parameter of the generative adversarial network, the rotation axis adaptation parameter of the generative adversarial network including the rotation axis feature mapping layer number of the generator, the rotation axis data discrimination threshold of the discriminator, and the iteration termination condition of network training.
[0054] The generative adversarial network (GAN) is composed of a generator and a discriminator. In order to meet the data enhancement requirements of the rotation axis complex defect, the adaptive network parameters need to be configured to ensure that the generated simulation data can accurately capture the surface characteristics corresponding to the feature and material parameter of the complex defect, as follows.
[0055] Step S121: based on the target defect type to be supplemented in the rotation axis data requirement parameter set to be enhanced, analyzing the rotation axis surface feature complexity corresponding to different defect types, if the parameter value corresponding to the rotation axis surface feature complexity is greater than the preset complexity threshold, the feature mapping layer number required by the generator is greater than the preset basic layer number, and the initial rotation axis feature mapping layer number of the generator is determined accordingly.
[0056] First, define the evaluation index of the rotation axis surface feature complexity, including defect edge sharpness, feature dimension number, and difficulty in distinguishing from normal area: micro-crack defects have high feature complexity due to their small width and blurred edges; composite defects have high complexity due to the superposition of multiple defect features; single defects such as ordinary scratches have low complexity due to their clear edges.
[0057] Calculate the complexity parameter value of each target defect type through the feature analysis module, and compare the value with the preset complexity threshold: if the complexity parameter value is greater than the threshold (such as micro-crack and composite defect), it indicates that the generator needs more feature mapping layers to capture subtle and complex features, and the initial feature mapping layer number of the generator needs to be greater than the preset basic layer number (the preset basic layer number is set based on the generation requirements of single defects); if the complexity parameter value is less than or equal to the threshold (such as ordinary scratch), the initial feature mapping layer number can adopt the preset basic layer number.
[0058] For example, the preset basic layer number is a certain value, the complexity parameter value of micro-crack is greater than the threshold, and the initial feature mapping layer number is set to the basic layer number plus several layers; the complexity parameter value of scratch + rust composite defect is also greater than the threshold, and the same initial layer number is set; the complexity parameter value of ordinary rust is less than the threshold, and the basic layer number is adopted.
[0059] Record the initial feature mapping layer number corresponding to each target defect type to form a generator layer number configuration table, which provides a reference for subsequent test adjustment.
[0060] Step S122: Test adjustment is performed on the initial number of shaft feature mapping layers. The feature similarity of the simulation data generated by the generator under different initial numbers of shaft feature mapping layers is compared with the real shaft data sample. The initial number of shaft feature mapping layers with feature similarity greater than that of other initial numbers of layers is selected as the final number of shaft feature mapping layers of the generator, and the final number of shaft feature mapping layers of the generator belongs to the shaft adaptation parameter of the generative adversarial network.
[0061] A test generator model is constructed. Different initial numbers of feature mapping layers (such as the basic number of layers, the basic number of layers + 1, and the basic number of layers + 2) determined in step S121 are used to input the same random noise and real shaft data feature vector to generate simulation data samples corresponding to the number of layers.
[0062] Real samples matching the target defect type and target material characterization parameter are selected from the real shaft data sample set. The surface features (texture features and defect morphology features) of these samples are compared with the surface features of the simulation data samples generated by different numbers of layers.
[0063] The similarity of the two is calculated by a feature similarity calculation module. The similarity calculation is based on multi-dimensional features such as the arrangement rule of the texture, the edge profile of the defect, and the gray distribution mode. If the feature similarity of the simulation data generated by a certain initial number of layers is higher than that of other numbers of layers, then the number of layers is more suitable for capturing the complex features of the corresponding defect, and it is determined as the final number of shaft feature mapping layers.
[0064] For example, the simulation data of microcracks generated by the basic number of layers + 2 has the highest feature similarity with the real microcrack samples, and the basic number of layers + 2 is determined as the final feature mapping layer number corresponding to microcracks. The simulation data of scratches + corrosion generated by the basic number of layers + 1 has the highest similarity, and it is determined as the final layer number corresponding to the defect, forming a generator layer number configuration scheme distinguished by defect type.
[0065] Step S123: Based on the target material characterization parameter in the set of shaft data enhancement requirement parameters, the pixel distribution features of the real shaft data samples corresponding to the target material characterization parameter are extracted. The mean value of the pixel distribution features is used as a reference value, and an allowed fluctuation range around the reference value is set. The boundary values of the allowed fluctuation range are determined as the initial shaft data discrimination threshold of the discriminator.
[0066] The initial discrimination threshold of the discriminator is configured through the following sub-steps: Step S1231: The target material characterization parameter is extracted from the set of shaft data enhancement requirement parameters, and the specific category and value range of the target material characterization parameter are determined.
[0067] Read the "target material characterization parameter" field in the demand parameter set, and determine the category and specific value of each parameter. For example, "chromium plating" belongs to the "surface treatment method" category, and the value is "chromium plating treatment"; "alloy steel" belongs to the "material category" category, and the value is "alloy steel material".
[0068] For each target material characterization parameter, determine its corresponding pixel distribution feature dimension. For example, the pixel distribution feature of a chromium-plated surface includes the gray value range of the reflective area and the pixel arrangement density of the texture; the pixel distribution feature of an alloy steel material includes the overall gray mean value and noise distribution pattern.
[0069] Step S1232: Filter all samples in the real rotating shaft data sample set whose material characterization parameters belong to the target material characterization parameter category to form a target material sample subset.
[0070] Based on the sample metadata annotation, filter the samples in the real rotating shaft data sample set whose "material category" is "alloy steel" and "surface treatment method" is "chromium plating" or "grinding", and integrate the above samples to form a target material sample subset. If a sample meets both "material category is alloy steel" and "surface treatment method is chromium plating", it is also included in the subset to ensure that the subset covers all samples of the target material characterization parameter combination.
[0071] After filtering, count the number of target material sample subsets. If the number is too small (e.g., less than a certain number), the filtering range needs to be expanded, and samples with similar material parameters (e.g., "high alloy steel" and "chromium plating + polishing") are included in the subset as supplementary samples to avoid bias in subsequent pixel distribution feature extraction due to insufficient sample size.
[0072] Step S1233: Extract pixel values for each sample in the target material sample subset to obtain the pixel values of all pixel points in each sample and form a pixel value set.
[0073] For each image sample in the target material sample subset, the image processing module traverses each pixel point of the image in row-major order, extracts the gray value of each pixel point (if it is a color image, it needs to be converted to a grayscale image first), and records the coordinates (X, Y) of the pixel point in the image.
[0074] Store the pixel values of all samples by sample ID to form a pixel value set. Each element in the pixel value set contains the sample ID, pixel coordinates, and corresponding gray value. For example, the pixel value set of sample ID S001 contains (X1, Y1, gray value 1), (X2, Y2, gray value 2), …, (Xn, Yn, gray value n), ensuring that all pixel points of each sample are completely extracted.
[0075] Step S1234: Calculate the mean value of the pixel value set, which is the reference value of the pixel distribution feature of the real rotation axis data sample corresponding to the target material representation parameter.
[0076] The gray values of all pixel points in the pixel value set are summarized, and the arithmetic mean value is calculated. The formula is "reference value = sum of gray values of all pixel points / total number of pixel points". For example, the pixel value set contains m samples, each sample has n pixel points, and the total number of pixel points is m x n. After adding all n x m gray values and dividing by the total number, the reference value is obtained.
[0077] During the calculation process, abnormal pixel values (such as extreme values with gray values of 0 or 255) need to be excluded. By setting a gray value range (such as 10-245), valid pixel values are filtered, and the mean value is calculated to ensure that the reference value can truly reflect the pixel distribution feature of the target material.
[0078] Step S1235: Analyze the dispersion degree of the pixel value set, and determine the size of the allowed fluctuation range according to the dispersion degree. When the parameter value corresponding to the dispersion degree is greater than the preset dispersion threshold, the allowed fluctuation range is greater than the preset basic range; when the parameter value corresponding to the dispersion degree is less than or equal to the preset dispersion threshold, the allowed fluctuation range is less than or equal to the preset basic range.
[0079] The dispersion degree is measured by calculating the standard deviation of the pixel value set. The larger the standard deviation, the more dispersed the pixel value distribution, and the higher the dispersion degree. The smaller the standard deviation, the more concentrated the distribution, and the lower the dispersion degree. Compare the calculated standard deviation with the preset dispersion threshold: if the standard deviation is greater than the preset dispersion threshold, it means that the pixel value of the target material sample fluctuates greatly, and a larger allowed fluctuation range needs to be set to accommodate the natural differences of real data; if the standard deviation is less than or equal to the preset dispersion threshold, the pixel value distribution is concentrated, and the allowed fluctuation range can use the preset basic range to avoid excessive relaxation leading to a decrease in discrimination accuracy.
[0080] The preset basic range is set based on the pixel fluctuation of the conventional material sample, for example, a certain basic range is reference value ± k, which can be adjusted to ± 2k when the dispersion degree is high, to ensure that the range size is adapted to the data dispersion degree.
[0081] Step S1236: Based on the reference value and the determined size of the allowed fluctuation range, calculate the upper limit value and the lower limit value of the allowed fluctuation range. The upper limit value is the reference value plus the size of the allowed fluctuation range, and the lower limit value is the reference value minus the size of the allowed fluctuation range.
[0082] According to the size of the allowable fluctuation range determined in step S1235, the upper limit value is calculated as upper limit value = reference value + size of allowable fluctuation range, and the lower limit value is calculated as lower limit value = reference value - size of allowable fluctuation range. For example, if the reference value is G and the size of the allowable fluctuation range is ΔG, the upper limit value is G + ΔG, and the lower limit value is G - ΔG.
[0083] If the calculated lower limit value is less than 0 (the minimum value of the gray value), the lower limit value is adjusted to 0; if the upper limit value is greater than 255 (the maximum value of the gray value), the upper limit value is adjusted to 255, so as to ensure that the upper and lower limit values meet the value range of the image pixel value, and avoid invalid threshold values.
[0084] Step S1237: The upper limit value and the lower limit value of the allowable fluctuation range are determined as the initial threshold values of the discriminator. The upper limit value is used as the upper limit standard of the pixel value of the real data determined by the discriminator, and the lower limit value is used as the lower limit standard of the pixel value of the real data determined by the discriminator.
[0085] The calculated upper limit value and lower limit value are stored as the initial threshold values of the discriminator. In the subsequent identification process, the discriminator can compare the pixel value of the input data with the threshold value: if the pixel value of the data is within the range of [lower limit value, upper limit value], it is preliminarily determined as real data meeting the target material characteristic; if it is out of the range, it is determined as abnormal data (may be simulation data or data of other materials).
[0086] For different combinations of target material characterization parameters (such as "alloy steel + chromium plating" and "alloy steel + grinding"), the corresponding reference values and upper and lower limit values need to be calculated respectively to form multiple sets of initial threshold values, so as to ensure that the discriminator can accurately identify different material combinations.
[0087] Step S1238: The pixel values of the target material sample subset are compared with the initial threshold values of the discriminator, the proportion of the number of pixel values within the allowable fluctuation range to the total number of pixel values of the target material sample subset is calculated, and if the proportion is less than a preset proportion threshold, the size of the allowable fluctuation range is adjusted, and the initial threshold values of the discriminator are recalculated until the proportion is greater than or equal to the preset proportion threshold.
[0088] The pixel value set of the target material sample subset is traversed, the number of pixels with gray values within the range of [lower limit value, upper limit value] in each sample is counted, and the average proportion of all samples is calculated, which is "proportion = (sum of valid pixel numbers of each sample / sum of total pixel numbers of each sample) x 100%".
[0089] The ratio is compared with a preset ratio threshold (such as 90%): if the ratio is less than the threshold, it indicates that the initial threshold range is too narrow, resulting in a large number of real pixel values being misjudged as abnormal, and the allowed fluctuation range size needs to be increased (such as from ΔG to 1.2ΔG), and the upper and lower limit values are recalculated; if the ratio is greater than or equal to the threshold, it indicates that the initial threshold range is reasonable and can be used as the final initial discrimination threshold.
[0090] The steps S1236-S1238 need to be repeated during the adjustment process until the ratio meets the standard, ensuring that the initial threshold can cover most of the pixel values of the real material, while not being too loose to cause discrimination failure.
[0091] Step S124: Test the discrimination accuracy of the discriminator on real shaft data samples and noise data under different initial shaft data discrimination thresholds, and retain the initial shaft data discrimination thresholds with discrimination accuracy greater than a preset accuracy threshold. Select the threshold that can maximize the discrimination between real data and noise data as the final shaft data discrimination threshold of the discriminator, which belongs to the shaft adaptation parameter of the generative adversarial network.
[0092] First, prepare the test data set, which includes two parts: one part is real shaft data samples outside the target material sample subset (consistent with the target material parameters), and the other part is randomly generated noise data (such as Gaussian noise, salt and pepper noise images). The number of samples of the two categories is configured in a 1:1 ratio, and neither has participated in the initial threshold calculation.
[0093] Input different groups of initial discrimination thresholds (such as thresholds for "alloy steel + chrome plating" and "alloy steel + grinding") into the discriminator, and the discriminator discriminates each sample in the test data set: calculate the proportion of pixel values within the threshold range in the sample. If the proportion is greater than a preset judgment ratio (such as 80%), it is judged as real data; otherwise, it is judged as noise data.
[0094] Calculate the discrimination accuracy of each initial threshold, the formula is "accuracy = (correctly judged real sample number + correctly judged noise sample number) / total number of test samples x 100%". Retain the initial threshold with accuracy greater than a preset accuracy threshold (such as 95%), and select the threshold with the lowest real sample misjudgment rate and the lowest noise sample misjudgment rate from these thresholds as the final discrimination threshold.
[0095] For example, the accuracy of threshold A is 96%, the real sample misjudgment rate is 2%, and the noise sample misjudgment rate is 2%; the accuracy of threshold B is 95%, the real sample misjudgment rate is 3%, and the noise sample misjudgment rate is 2%, then select threshold A as the final shaft data discrimination threshold of the discriminator, to ensure that the discriminator can accurately distinguish real data from noise and simulation data.
[0096] Step S125: Based on the target sample quantity in the set of data requirement parameters of the rotation axis to be enhanced, the sample quantity generated in a single iteration of the generative adversarial network, the minimum iteration quantity required is calculated, and the minimum iteration quantity is taken as the initial iteration termination condition of the generative adversarial network training.
[0097] Firstly, the sample quantity generated by the generator in a single iteration of the generative adversarial network is determined, which depends on the batch size (BatchSize) of the generator and the hardware computing capability (such as GPU memory), for example, a certain quantity of samples can be generated in a single iteration.
[0098] The initial iteration termination condition is calculated by the formula "minimum iteration quantity = ceiling (target sample quantity / sample quantity generated in a single iteration)". For example, the target sample quantity is a certain value, the sample quantity generated in a single iteration is a certain value, the minimum iteration quantity is calculated to be a certain value, and if there is a remainder, an additional iteration is required to ensure that the total quantity of generated samples is not less than the target sample quantity.
[0099] If the target sample quantity is split according to the defect type and material parameter (for example, "micro crack + alloy steel" requires a certain quantity of samples, and "scratch + rust composite + chrome plating" requires a certain quantity of samples), the minimum iteration quantity of each combination needs to be calculated respectively, and the maximum value is taken as the initial iteration termination condition to ensure that the sample requirement of all combinations can be met.
[0100] Step S126: Monitor the training stability of the generative adversarial network under different initial iteration termination conditions. When the generation loss value and the discrimination loss value of the generative adversarial network tend to be stable and no longer decrease significantly, the iteration quantity at this time is recorded, the iteration quantity is compared with the minimum iteration quantity, and the iteration quantity greater than the minimum iteration quantity is selected as the final iteration termination condition of the generative adversarial network training. The final iteration termination condition of the generative adversarial network training belongs to the rotation axis adaptation parameter of the generative adversarial network.
[0101] A training monitoring module of the generative adversarial network is built to record the generator loss value (GeneratorLoss) and the discriminator loss value (DiscriminatorLoss) of each iteration in real time. The generator loss value reflects the difference between the generated samples and the real samples, and the smaller the value is, the better it is. The discriminator loss value reflects the ability of the discriminator to distinguish between real and generated samples, and the smaller the value is, the more accurate the discrimination is.
[0102] A plurality of training experiments are performed for different initial iteration termination conditions (e.g., iteration number 1, iteration number 2, iteration number 3). Each experiment uses the same training subset and network parameters, and only the iteration number is adjusted. The trend of the loss value in each experiment is monitored: when the fluctuation amplitude of the generation loss value and the discrimination loss value of continuous multiple iterations (e.g., 20 rounds) is less than a preset fluctuation threshold (e.g., 0.001), and there is no obvious downward trend, it is determined that the network training has reached a stable state, and the iteration number at this time is recorded.
[0103] The stable iteration number is compared with the minimum iteration number calculated in step S125: if the stable iteration number is greater than the minimum iteration number, the stable iteration number is taken as the final termination condition; if the stable iteration number is less than the minimum iteration number, the iteration number needs to be increased until the minimum iteration number is reached and the loss value remains stable, to avoid insufficient iteration leading to poor quality of generated samples.
[0104] For example, the minimum iteration number is a certain value, and the loss value of a certain experiment is stable when the iteration number reaches a certain value, so the final termination condition is set to the iteration number; if the loss value of another experiment is stable when the iteration number is a certain value but less than the minimum iteration number, the iteration needs to be continued to the minimum iteration number to ensure that the number of generated samples meets the standard.
[0105] Step S127: Integrate the final generator feature mapping layer number of the shaft, the final discriminator data discrimination threshold of the shaft, and the iteration termination condition of the final generative adversarial network training to form the shaft adaptation parameter of the generative adversarial network.
[0106] The generator feature mapping layer number determined in step S122 according to the defect type (e.g., "microcrack" corresponds to layer number 1, "scratch + rust composite" corresponds to layer number 2), the discriminator threshold value determined in step S124 according to the material combination (e.g., "alloy steel + chrome plating" corresponds to threshold value 1, "alloy steel + grinding" corresponds to threshold value 2), and the final iteration termination condition determined in step S126 are integrated into a JSON format shaft adaptation parameter file.
[0107] The file contains three core modules: "generator configuration" (defect type-layer number mapping table, activation function type, optimizer parameter), "discriminator configuration" (material combination-threshold value mapping table, discrimination ratio, loss function type), "training configuration" (iteration termination number, batch size, learning rate).
[0108] The adaptation parameter file is loaded into the core training module of the generative adversarial network, and is also backed up to the project database, to facilitate parameter tuning and tracing in the subsequent training process.
[0109] Step S130: Obtain a real rotating shaft data sample set, and drive the generative adversarial network to run according to the rotating shaft adaptation parameters of the generative adversarial network based on the real rotating shaft data sample set, to generate a preliminary rotating shaft simulation data set.
[0110] The real rotating shaft data is collected through multiple channels, and after cleaning and labeling, a training subset is formed. The generative adversarial network is driven to train according to the adaptation parameters and generate simulation data. The specific process is as follows.
[0111] Step S131: Collect rotating shaft surface image data and corresponding material parameter records under different working conditions, and remove invalid data from the collected image data to obtain cleaned rotating shaft data. The different working conditions include different running time, different load intensity working conditions.
[0112] Industrial vision acquisition system is used to collect real rotating shaft data: the system is composed of multiple industrial cameras, ring light sources and rotating stages. The camera resolution and frame rate are adjusted according to the size of the rotating shaft defects (such as high-resolution cameras for micro-cracks). Collect rotating shaft samples under different working conditions: divide into new shaft (not running), short-time running shaft and long-time running shaft according to running time; divide into light-load running shaft, medium-load running shaft and heavy-load running shaft according to load intensity, to ensure that the defect evolution state of the rotating shaft throughout its life cycle is covered.
[0113] The material parameter records (material type, surface treatment method, heat treatment process, etc.) of each rotating shaft are recorded synchronously during collection and stored as metadata files associated with image data. Invalid data is removed from the collected image data: delete blurred images (with a clarity lower than a preset clarity threshold), overexposed / underexposed images (with a gray mean value outside the normal range), and duplicate images (identified by image hash value comparison), to obtain cleaned rotating shaft data, ensuring that the data quality meets the training requirements.
[0114] Step S132: Label the cleaned rotating shaft data to form a real rotating shaft data sample set. The labeling content includes rotating shaft surface defect type and rotating shaft material characterization parameter.
[0115] An artificial labeling platform is built, and professional detection personnel are organized to label the cleaned rotating shaft images: first, label the defect type. For single defects, directly label the type name (such as "scratch" and "rust"), and for complex defects, label the composite type (such as "scratch + rust composite" and "micro-crack + indentation"). The position of the defect in the image is marked by a bounding box. Second, label the material characterization parameters. Extract the material type, surface treatment method and other information from the metadata file, and supplement the label to the attribute field of the image sample.
[0116] After labeling, cross-validation is performed: multiple annotators independently label the same batch of samples, and the annotation consistency (such as IOU value, type consistency ratio) is calculated. If the consistency is lower than the preset consistency threshold (such as 90%), the annotators need to review and modify the annotation differences together to ensure the accuracy of the annotation. The labeled samples are stored in a unified format (such as VOC format, COCO format) to form a real rotating shaft data sample set.
[0117] Step S133: The real rotating shaft data sample set is divided into a training subset and a validation subset according to a preset ratio. The training subset is used to drive the training of the generative adversarial network, and the validation subset is used to monitor the generation effect in the network training process.
[0118] The data set division is completed through the following sub-steps: Step S1331: The functional requirements of the generative adversarial network training for the training subset and the validation subset are clearly defined. Based on the functional requirements, the division ratio of the training subset and the validation subset is set. The division ratio needs to ensure that the number of samples in the training subset is greater than the minimum sample threshold for network training, and the number of samples in the validation subset is greater than the minimum sample threshold for effect monitoring.
[0119] The function of the training subset is to provide basic data for network learning, which needs to include a variety of defect types and material parameter combinations, and the number of samples needs to meet the needs of network parameter optimization (i.e. greater than the minimum sample threshold for training). The function of the validation subset is to evaluate the training effect in real time, and the number of samples needs to be sufficient to statistically evaluate the quality indicators of the generated samples (i.e. greater than the minimum sample threshold for effect monitoring).
[0120] Based on the functional requirements, the division ratio (such as 7:3, 8:2) is set. For example, when the ratio is 8:2, 80% of the samples are assigned to the training subset and 20% to the validation subset. If the total number of samples is large, the proportion of the validation subset can be appropriately reduced (such as 9:1), but the number of samples in the validation subset needs to be ensured not to be less than the minimum sample threshold for effect monitoring.
[0121] Step S1332: The real rotating shaft data sample set is divided into layers according to the rotating shaft surface defect type and the rotating shaft material parameter representation. In each layer, samples are randomly selected according to the set division ratio. The larger part of the selected samples is assigned to the training subset, and the smaller part is assigned to the validation subset.
[0122] The sample set is divided into layers according to the combination of "defect type + material parameter", such as "micro-crack + alloy steel", "scratch + rust composite + chrome plating", "ordinary scratch + 45 steel", etc., to ensure that each layer contains a certain number of samples.
[0123] In each layer, a random number generator is used to select samples: if the division ratio is 8:2, 80% of the samples are randomly selected into the training subset, and the remaining 20% are selected into the validation subset. For example, the "micro-crack + alloy steel" layer contains 100 samples, 80 of which are randomly selected into the training subset, and 20 of which are randomly selected into the validation subset, ensuring that the sample distribution in each layer remains consistent between the training and validation subsets.
[0124] Step S1333: The number of samples of different shaft surface defect types and the number of samples of different shaft material characterization parameters included in the training subset and the validation subset are counted, and the sample distribution of the two is made to deviate from the overall distribution of the real shaft data sample set by less than a preset distribution deviation threshold.
[0125] The proportion of each defect type (such as the proportion of "micro-crack" and the proportion of "scratch + corrosion complex") and the proportion of each material parameter (such as the proportion of "alloy steel" and the proportion of "chromium plating") in the training subset and the validation subset are calculated, and the deviation value is calculated by comparing the corresponding proportion of the original sample set. The formula is "deviation value = |proportion in subset - proportion in original set |".
[0126] If the deviation values of all defect types and material parameters are less than the preset distribution deviation threshold (such as 5%), it is determined that the division is qualified; if there is a deviation value that exceeds the standard, the sample selection in the corresponding layer needs to be adjusted, for example, if the proportion of a certain defect type in the training subset is too low, some samples are transferred from the validation subset of the layer to the training subset until the deviation value meets the requirements.
[0127] Step S1334: If the number of samples in the training subset is less than the network training minimum sample threshold, adjust the division ratio to increase the number of samples in the training subset and reduce the number of samples in the validation subset until the number of samples in the training subset is greater than the network training minimum sample threshold.
[0128] Compare the total number of samples in the training subset with the network training minimum sample threshold: if the total number is insufficient, the division ratio of the training subset needs to be increased (such as from 8:2 to 9:1), and samples are selected from each layer again to increase the number of training subsets and reduce the number of validation subsets.
[0129] After adjustment, the sample distribution deviation needs to be verified again to ensure that while increasing the number of training samples, the distribution of each defect type and material parameter still deviates from the overall distribution of the original sample set by less than the preset distribution deviation threshold. If the adjusted ratio still cannot meet the sample quantity requirement of the training subset, it is necessary to consider supplementing the collection of real shaft data or including some similar material samples with accurate labeling to ensure that the sample quantity of the training subset is sufficient and the distribution is reasonable.
[0130] Step S1335: If the number of samples in the validation subset is less than the effect monitoring minimum sample threshold, adjust the division ratio, increase the number of samples in the validation subset, and reduce the number of samples in the training subset until the number of samples in the validation subset is greater than the effect monitoring minimum sample threshold.
[0131] Compare the total number of samples in the validation subset with the effect monitoring minimum sample threshold: if the total number is insufficient, the division ratio of the training subset needs to be reduced (e.g., from 8:2 to 7:3), and some samples are transferred from the training subsets of each layer to the validation subset to increase the number of samples in the validation subset.
[0132] After adjustment, the distribution deviation of each defect type and material parameter also needs to be recalculated. If the deviation is out of tolerance, the sample allocation in the layer is fine-tuned (e.g., more samples are transferred from layers with sufficient distribution), to ensure that the validation subset can meet the quantity requirement and accurately reflect the distribution characteristics of the original sample set, providing a reliable basis for training effect monitoring.
[0133] Step S1336: After adjustment, the final selected samples are integrated to form the training subset and the validation subset.
[0134] The samples in each layer that are determined to belong to the training subset are classified and stored according to defect type and material parameter, generating a label file and a metadata file for the training subset. The label file contains the defect bounding box and defect type label for each sample, and the metadata file contains information such as material representation parameters and acquisition conditions.
[0135] Similarly, the samples in the validation subset are integrated to generate corresponding label files and metadata files, ensuring that the file formats of the training subset and the validation subset are uniform, facilitating the generation of adversarial networks for reading and processing. The two subsets are stored in separate folders with clear labels "training subset" and "validation subset", and accompanied by a division explanation document recording the division ratio, the number of samples in each layer, and the distribution deviation.
[0136] Step S134: Input the training subset into the generator of the generative adversarial network, and perform multi-layer mapping processing on the shaft features of the training subset according to the number of generator shaft feature mapping layers in the shaft adaptation parameters of the generative adversarial network, to generate initial simulation data.
[0137] Start the generator module of the generative adversarial network, convert the image data of the training subset into a tensor format recognizable by the generator, and extract the feature vector (including texture features, defect morphology features, and material parameter features) of each sample as input to the generator.
[0138] According to the number of feature mapping layers set in the pivot adaptation parameters according to the defect type, the generator performs multi-layer mapping processing on the input features: the first layer of mapping extracts the basic texture and material characteristics through convolution operation; the middle layer of mapping captures complex defect characteristics (such as the edge blur feature of micro-cracks and the superimposed feature of composite defects) in detail by increasing the number of convolution kernels and the size of the convolution window to enhance the extraction ability of subtle features; the last layer of mapping maps the low-dimensional features to high-dimensional image data through deconvolution operation to generate initial simulation data consistent with the sample size and resolution of the training subset.
[0139] For example, for the “micro-crack + alloy steel” sample, the generator uses the set high-layer number feature mapping to gradually extract the micro-width features of the micro-crack and the gray distribution features of the alloy steel through multi-layer convolution, and then generates initial simulation data containing clear micro-crack morphology and matching material texture through deconvolution; for the “scratch + rust composite + chrome plating” sample, the generator captures the superimposed relationship between the linear profile of the scratch and the patch-like feature of the rust through multi-layer mapping, while also restoring the reflective properties of the chrome-plated surface to ensure the feature authenticity of the initial simulation data.
[0140] Step S135: input the initial simulation data and the real data of the training subset into the discriminator of the generative adversarial network, and the discriminator discriminates the authenticity of the initial simulation data according to the pivot data discrimination threshold of the discriminator in the pivot adaptation parameters of the generative adversarial network, and outputs the discrimination result.
[0141] The initial simulation data and the real data of the training subset are mixed in a 1:1 ratio, and then randomly shuffled and input into the discriminator. The discriminator first preprocesses the input data (such as normalization processing, converting the pixel value to the range of 0-1), and then extracts the pixel distribution features, texture features and defect morphology features of the data.
[0142] According to the discrimination threshold set in the pivot adaptation parameters according to the material combination, the discriminator calculates the matching degree of the pixel value of the input data and the corresponding material threshold: if the pixel distribution is within the threshold range, and the texture features (such as the reflective texture of the chrome-plated surface) and the defect morphology features (such as the edge gradient of the micro-crack) have a similarity to the real data that exceeds the preset similarity threshold, it is determined to be “real data”; if the pixel distribution exceeds the threshold range, or the feature similarity is lower than the threshold, it is determined to be “simulation data”.
[0143] The discriminator outputs the discrimination result (“real” or “simulation”) and the corresponding confidence value (0-1 range, the value closer to 1 indicates a higher confidence in determining the real data) of each input data, compares the discrimination result with the real label of the data (the training subset data is labeled as “real”, and the initial simulation data is labeled as “simulation”), and calculates the discrimination loss value of the discriminator.
[0144] Step S136: According to the identification result, the network parameters of the generator and the discriminator are adjusted by back propagation, and the process of generating initial simulation data by the generator, identifying by the discriminator and adjusting the network parameters is repeatedly performed until the iteration termination condition of network training in the pivot adaptation parameter of the generative adversarial network is reached. All initial simulation data output by the generator at the iteration termination is integrated to form a preliminary pivot simulation data set.
[0145] According to the identification loss value output by the discriminator, the network parameters of the generator and the discriminator are adjusted by the back propagation algorithm: for the generator, if the proportion of the generated simulation data misjudged by the discriminator as "real data" is low, the weight and bias parameters of the convolution layer need to be adjusted to strengthen the learning of real features (such as increasing the extraction weight of micro-crack edge features); for the discriminator, if the misjudgment rate of real data is high, the parameters of the fully connected layer need to be adjusted to optimize the feature matching logic and improve the identification accuracy.
[0146] After completing a round of parameter adjustment, the generator generates a batch of initial simulation data again, the discriminator identifies and calculates the loss value again, and the above process is repeated. During the training process, the generation effect is monitored in real time by the validation subset: every certain number of iterations, the simulation data output by the generator is compared with the real data of the validation subset in terms of features, the feature similarity is calculated, and if the similarity continues to improve and tends to be stable, it indicates that the training effect is good.
[0147] When the number of iterations reaches the iteration termination condition set in the pivot adaptation parameter, and the generation loss value of the generator and the identification loss value of the discriminator are stable within the preset range, the training is stopped. All initial simulation data output by the generator at the iteration termination are classified according to defect types and material parameters, and integrated to form a preliminary pivot simulation data set, each simulation data sample is attached with a generation label (including defect type, material parameter, and generation iteration number).
[0148] Step S140: Perform feature consistency comparison processing on the preliminary pivot simulation data set and the real pivot data sample set, eliminate the preliminary pivot simulation data with inconsistent features, and obtain an effective pivot simulation data set with consistent features.
[0149] The authenticity of the simulation data is verified by multi-dimensional feature comparison to ensure that it matches the features of the real data. The specific process is as follows.
[0150] Step S141: Extract the pivot surface texture features and defect morphology features of each sample from the real pivot data sample set, and integrate to form a real pivot feature set, which includes a real texture feature subset and a real defect morphology feature subset.
[0151] The real pivot features are extracted by the following sub-steps: Step S1411: Obtain each data sample in the real shaft data sample set, read the image data and the corresponding label information of each data sample, and the label information includes the shaft surface defect type and the shaft material characterization parameter.
[0152] Traverse the real shaft data sample set, load the image data (gray image or color image) of each sample through the image reading module, read the corresponding XML label file and JSON metadata file at the same time, extract the key information such as “defect type”, “material type” and “surface treatment method”, and establish the mapping relationship between the sample ID and the label information.
[0153] Step S1412: Perform texture feature extraction on the image data of each data sample, capture the arrangement rule and gray level change mode of the pixels in the image data by using the texture feature extraction logic, and obtain the shaft surface texture feature of the data sample.
[0154] For the texture features of different materials, the corresponding extraction logic is adopted: for mechanical processing textures such as turning and grinding, four parameters of energy, entropy, contrast and correlation of the texture are calculated through the gray level co-occurrence matrix, the energy reflects the uniformity of the texture, the entropy reflects the complexity of the texture, the contrast reflects the clarity of the texture, and the correlation reflects the directionality of the texture; for coating textures such as chrome plating, the microstructure features of the texture are extracted through the local binary pattern, and the particle distribution and light reflection mode of the coating surface are captured.
[0155] Combine the texture feature parameters (four parameters of the gray level co-occurrence matrix and the local binary pattern feature vector) of each sample to form the shaft surface texture feature of the sample, and the dimension of the feature vector is determined according to the number of extracted parameters (such as four parameters of the gray level co-occurrence matrix + 256-dimensional local binary pattern feature, forming a 260-dimensional texture feature vector).
[0156] Step S1413: Perform defect morphology feature extraction on the image data of each data sample, locate the defect area in the image data according to the shaft surface defect type in the data sample label information, extract the shape, size and edge contour features of the defect area, and obtain the defect morphology feature of the data sample.
[0157] According to the defect boundary box coordinates in the annotation file, the defect area in the image is located: for micro-crack defects, the edge profile of the crack is extracted by the Canny edge detection algorithm, the length, width and bending degree parameters of the profile are calculated, and the bending degree is measured by the average value of the curvature change of each point on the profile; for composite defects (such as scratches + rust), the linear features (length, width, direction) of the scratches and the patch features (area, circularity, gray mean value) of the rust are extracted respectively, and then the two types of features are combined; for edge fuzzy defects, the fuzzy boundary of the defect area is determined by adaptive threshold segmentation, and the fuzzy degree parameter (gray gradient standard deviation of edge pixels) of the boundary and the area and shape parameters of the region are calculated.
[0158] The shape, size and edge profile parameters of the defect are combined to form the defect morphology feature vector of the sample (such as the length, width and bending degree of micro-crack + edge gradient feature, forming a multi-dimensional feature vector).
[0159] Step S1414: Establish a real texture feature sub-set, and store the surface texture features of each data sample according to the surface texture features of the data sample. The texture features of data samples with the same material representation parameters are classified into the same category.
[0160] A dictionary structure with “material type + surface treatment method” as the key value is created as the real texture feature sub-set. For example, the key value “alloy steel + chrome plating” corresponds to the texture feature vector list of all samples of this material combination; the key value “45 steel + turning” corresponds to the texture feature vector list of samples of this material combination.
[0161] The texture feature vector of each sample is classified into the corresponding dictionary key value according to its material representation parameter. If the number of samples of a certain material combination is large, further storage is performed according to the surface roughness parameter to ensure that the classification of the real texture feature sub-set is refined and ordered, which facilitates subsequent comparison with the texture features of simulation data.
[0162] Step S1415: Establish a real defect morphology feature sub-set, and store the defect morphology features of each data sample according to the surface defect type of the data sample. The defect morphology features of data samples with the same defect type are classified into the same category.
[0163] A dictionary structure with defect type as the key value is created as the real defect morphology feature sub-set. For example, the key value “micro-crack” corresponds to the defect morphology feature vector list of all micro-crack samples; the key value “scratch + rust composite” corresponds to the defect morphology feature vector list of all samples of this composite defect; the key value “edge fuzzy defect” corresponds to the defect morphology feature vector list of all samples of this type of defect.
[0164] The defect morphology feature vector of each sample is classified into the corresponding dictionary key value according to the defect type. For a composite defect, the corresponding single defect feature of the sample is stored in the single defect type key value contained in the composite defect (for example, the scratch feature of the "scratch + rust composite" sample is stored in the "scratch" key value), facilitating multi-dimensional comparison.
[0165] Step S1416: The number of texture features in each category in the real texture feature subset is counted. If the number of texture features is less than the number of data samples corresponding to the material representation parameter, the data samples without extracted texture features are reprocessed to supplement the missing texture features.
[0166] Each dictionary key value of the real texture feature subset is traversed, and the number of texture feature vectors under the key value is counted and compared with the number of data samples of the material combination. If the numbers do not match, it indicates that there are samples without extracted texture features. These samples are reloaded through the image reading module, the texture feature extraction is supplemented according to the extraction logic of step S1412, and is classified into the corresponding category to ensure that each real sample has corresponding texture features stored in the subset.
[0167] Step S1417: The number of defect morphology features in each category in the real defect morphology feature subset is counted. If the number of defect morphology features is less than the number of data samples corresponding to the defect type, the data samples without extracted defect morphology features are reprocessed to supplement the missing defect morphology features.
[0168] Using the same logic as step S1416, the matching of the number of feature vectors under each key value of the real defect morphology feature subset and the number of samples of the corresponding defect type is counted. The samples with missing features are re-executed the defect region positioning and feature extraction process to supplement the defect morphology feature vectors, ensuring the integrity of the real defect morphology feature subset.
[0169] Step S1418: The real texture feature subset and the real defect morphology feature subset are integrated to form a real shaft feature set.
[0170] The real texture feature subset and the real defect morphology feature subset are associated according to the sample ID to form a structured real shaft feature set containing "sample ID-material parameter-texture feature-defect morphology feature", which is stored as a JSON format file. Each entry corresponds to the complete feature information of a real sample, facilitating one-to-one comparison with simulation data features.
[0171] Step S142: The shaft surface texture features and defect morphology features of each sample are extracted from the preliminary shaft simulation data set to form a simulation shaft feature set, which includes a simulation texture feature subset and a simulation defect morphology feature subset.
[0172] For each simulation sample in the preliminary rotation axis simulation data set, the same texture feature extraction logic as step S1412 (the same gray level co-occurrence matrix parameters, local binary pattern dimension) is used to extract the texture feature vector, ensuring consistency in extraction standards; the same defect morphology feature extraction logic as step S1413 (the same edge detection algorithm, parameter calculation method) is used to extract the defect morphology feature vector, including the shape, size, and edge profile parameters of the defect.
[0173] According to the classification method of step S1414, the texture feature vector of the simulation sample is classified into the corresponding category of the simulation texture feature sub-set according to the "material type + surface treatment method" in its generated label; according to the classification method of step S1415, the defect morphology feature vector of the simulation sample is classified into the corresponding category of the simulation defect morphology feature sub-set according to the "defect type" in its generated label.
[0174] The matching of the number of features in the simulation texture feature sub-set and the simulation defect morphology feature sub-set with the number of preliminary simulation data samples is counted, and after the missing features are supplemented, the simulation rotation axis feature set containing "sample ID - generated material label - generated defect label - simulation texture feature - simulation defect morphology feature" is formed according to the sample ID, ensuring that its structure is completely consistent with the real rotation axis feature set, laying a foundation for subsequent consistency comparison.
[0175] Step S143: Compare each simulation texture feature in the simulation texture feature sub-set with the real texture feature under the corresponding material representation parameter in the real texture feature sub-set, calculate the feature agreement degree of the two, and the real texture feature under the corresponding material representation parameter refers to the texture feature of the real data sample with the same material representation parameter as the simulation data sample.
[0176] Select a simulation texture feature vector and the corresponding material representation parameter (such as "alloy steel + chrome plating") from the simulation texture feature sub-set, and find all real texture feature vectors in the real texture feature sub-set under the same material representation parameter category to form a comparison reference set.
[0177] Calculate the cosine similarity of the simulation texture feature vector and each real texture feature vector in the comparison reference set. The cosine similarity measures the degree of feature similarity by calculating the cosine of the angle between two vectors. The closer the value is to 1, the higher the agreement degree. Take the average of all cosine similarities calculated as the feature agreement degree of the simulation texture feature and the real texture feature. If the number of real texture features in the comparison reference set is zero (corresponding to no real sample for the material), then select real texture features with similar material parameters to form a reference set (such as "alloy steel + polishing" instead of "alloy steel + chrome plating"), and mark "approximate material comparison" in the agreement calculation result.
[0178] Step S144: Compare each simulated defect morphology feature in the simulated defect morphology feature subset with the real defect morphology feature under the corresponding defect type in the real defect morphology feature subset, calculate the morphology coincidence degree of the two, and the real defect morphology feature under the corresponding defect type refers to the defect morphology feature of the real data sample of the same defect type as the simulation data sample.
[0179] Select a simulated defect morphology feature vector and the corresponding defect type (such as "micro-crack") from the simulated defect morphology feature subset, and find all real defect morphology feature vectors under the same defect type category in the real defect morphology feature subset to form a defect comparison reference set.
[0180] Calculate the Euclidean distance between the simulated defect morphology feature vector and each real feature vector in the defect comparison reference set. The Euclidean distance is obtained by calculating the square sum of the difference between the corresponding dimensions of the two vectors and taking the square root. The smaller the value, the more similar the morphology. Convert the Euclidean distance to a morphology coincidence degree (such as converting to a value in the range of 0-1 by "morphology coincidence degree = 1 / (1+Euclidean distance)"), and take the average of all morphology coincidence degrees as the final morphology coincidence degree of the simulated defect morphology feature and the real defect morphology feature.
[0181] For complex defects (such as "scratch + rust complex"), the coincidence degree of the scratch sub-feature in the simulation feature and the real scratch feature, and the coincidence degree of the rust sub-feature and the real rust feature need to be calculated respectively, and then weighted and summed according to the preset weight (such as scratch 0.4, rust 0.6) to obtain the morphology coincidence degree of the complex defect, ensuring that the comparison covers all complex features.
[0182] Step S145: Compare the feature coincidence degree of each preliminary simulated data sample of the rotating shaft with the feature coincidence degree threshold, and compare the morphology coincidence degree with the morphology coincidence degree threshold. If the feature coincidence degree of the preliminary simulated data sample of the rotating shaft is greater than or equal to the feature coincidence degree threshold and the morphology coincidence degree is greater than or equal to the morphology coincidence degree threshold, it is determined that the sample is a feature consistent simulation data; if the feature coincidence degree of the preliminary simulated data sample of the rotating shaft is less than the feature coincidence degree threshold or the morphology coincidence degree is less than the morphology coincidence degree threshold, it is determined that the sample is a feature inconsistent simulation data, which is excluded.
[0183] Set the feature coincidence degree threshold and the morphology coincidence degree threshold (based on a large number of real-simulation comparison experiments, the feature coincidence degree threshold is usually not less than 0.7, the morphology coincidence degree threshold is usually not less than 0.65, and the threshold of complex defects can be appropriately reduced by 0.05-0.1).
[0184] For each preliminary simulation sample, compare its feature fitness and shape fitness simultaneously: if both reach or exceed the corresponding threshold, it indicates that the texture features and defect morphology of the simulation sample are highly matched with the real data, and it is determined as feature-consistent simulation data; if either fitness is below the threshold (such as the texture feature fitness meets the standard but the defect morphology fitness is insufficient, or vice versa), it indicates that the sample has feature distortion (such as the edge of micro-cracks is too clear, the chrome texture does not conform to the real reflection pattern), and it is determined as feature-inconsistent simulation data, marked as “to be removed” and the reason for not meeting the standard is recorded (such as “defect morphology fitness is low-micro-crack width is abnormal”).
[0185] Step S146: Collect all simulation data determined as feature-consistent, and integrate to form a feature-consistent effective shaft simulation data set.
[0186] Screen all simulation samples determined as “feature-consistent” in the preliminary shaft simulation data set, classify and organize them according to defect type and material parameter, and generate a label file and a metadata file of effective simulation data. The label file format of effective simulation data is consistent with that of the real shaft data sample set, containing the defect bounding box coordinates of each simulation sample, defect type label (such as “micro-crack” “scratch + rust composite”); the metadata file contains generated label (labeled “simulation data”), corresponding material characterization parameters (such as “alloy steel + chrome plating”), generation iteration number, generator feature mapping layer number, etc. information, which is convenient for subsequent tracing of the generation process of simulation data.
[0187] Integrate the classified and organized effective simulation samples and the corresponding label files and metadata files, create independent folders according to defect type (such as “micro-crack simulation data” “scratch + rust composite simulation data”), and further subdivide sub-folders under each folder according to material parameters (such as “micro-crack simulation data / alloy steel + chrome plating” “micro-crack simulation data / 45 steel + grinding”), forming a structured effective shaft simulation data set. At the same time, generate a collection list document to record the number of simulation samples corresponding to each defect type and each material parameter, ensuring the integrity and manageability of the collection.
[0188] Step S150: Integrate the effective shaft simulation data set and the real shaft data sample set to form a final shaft augmented data set, which contains real shaft data samples and effective shaft simulation data.
[0189] Through data fusion, classification optimization, gap filling, etc., real data and effective simulation data are integrated into an augmented data set that meets the training requirements, and the specific process is as follows.
[0190] For example, step S151: obtain all valid shaft simulation data samples in the valid shaft simulation data set, add simulation data identification to each valid shaft simulation data sample, and the simulation data identification is used to distinguish real data from simulation data.
[0191] Traverse each simulation sample in the valid shaft simulation data set, add "simulation data" identification in the "data type" field of its metadata file, and generate a unique simulation data ID (format "SIM-year-month-day-sequence number"), which is associated with the sample file name (such as "SIM-2025-09-01-001.jpg").
[0192] At the same time, add the "simulation identification" field in the annotation file, and take the value as "yes", which is clearly distinguished from the "simulation identification" field "no" of the real data sample, so that the data type can be selectively called during subsequent model training, or the difference in the influence of real data and simulation data on model performance can be analyzed.
[0193] Step S152: obtain all real shaft data samples in the real shaft data sample set, and add real data identification to each real shaft data sample.
[0194] For each sample in the real shaft data sample set, clearly mark "real data" in the "data type" field of its metadata file, and continue to use the original real sample ID (format "REAL-sequence number"), to ensure that the naming rules of the simulation data ID are clearly distinguished.
[0195] Check the annotation file of all real samples, supplement the "simulation identification" field and take the value as "no", to ensure that real data and simulation data are completely independent in identification, and to avoid data type confusion during integration.
[0196] Step S153: classify the valid shaft simulation data samples according to the shaft surface defect type and shaft material characterization parameter, form a simulation data classification subset, and each simulation data classification subset contains valid shaft simulation data samples of the same defect type and the same material parameter.
[0197] Classify the samples in the valid shaft simulation data set according to "defect type + material characterization parameter" as classification dimensions: for example, all simulation samples with "micro-crack" defects and "alloy steel + chrome plating" material are classified into one classification subset, and all simulation samples with "scratch + rust composite" defects and "45 steel + grinding" material are classified into another classification subset.
[0198] Each simulation data classification subset contains the image file, annotation file and metadata file of the corresponding sample, and a classification description document is generated to record the defect type, material parameter, sample quantity, generator configuration parameter (such as the number of feature mapping layers) and other information of the subset, providing a basis for subsequent merging with the real data classification subset.
[0199] Step S154: The real shaft data samples are classified according to the shaft surface defect type and shaft material representation parameter to form real data classification subsets, and each real data classification subset contains real shaft data samples of the same defect type and the same material parameter.
[0200] The real shaft data sample set is classified using the same classification dimensions ("defect type + material representation parameter") as step S153: for example, the "micro crack + alloy steel + chrome plating" real sample forms a real data classification subset, and the "scratch + rust complex + 45 steel + grinding" real sample forms another real data classification subset.
[0201] Each real data classification subset also contains sample images, annotation files, metadata files and classification description documents, ensuring that its classification logic is completely consistent with the simulation data classification subset, and laying the foundation for the accurate merging of the two types of data.
[0202] Step S155: The simulation data classification subsets and the real data classification subsets of the same category are merged, that is, the simulation data and the real data of the same defect type and the same material parameter are classified into the same merged subset, to obtain multiple merged subsets.
[0203] All simulation data classification subsets are traversed, and the corresponding real data classification subsets are found according to their "defect type + material parameter" category, and the two are merged into a unified merged subset. For example, the "micro crack + alloy steel + chrome plating" simulation data classification subset and the "micro crack + alloy steel + chrome plating" real data classification subset are merged to form the "micro crack + alloy steel + chrome plating" merged subset.
[0204] During the merging process, the original identification (simulation / real), ID, annotation and metadata information of the two types of data are retained, and only the sample files are stored in folders in the merged subset according to the data type (such as "merged subset / micro crack + alloy steel + chrome plating / real data" and "merged subset / micro crack + alloy steel + chrome plating / simulation data"), which not only realizes data integration, but also facilitates tracing the data source.
[0205] For the categories with only simulation data classification subset without corresponding real data classification subset (such as some rare material and complex defect combinations), directly take the simulation data classification subset as an independent merged subset, and mark “no corresponding real data” in the classification description document; for the categories with only real data classification subset without corresponding simulation data classification subset, take the real data classification subset as an independent merged subset, and mark “no corresponding simulation data”.
[0206] Step S156: Count the sample quantity of each merged subset, and if the sample quantity is less than the data sample quantity requirement in the to-be-enhanced shaft data requirement parameter set, record the sample gap quantity of the merged subset.
[0207] Retrieve the “sample quantity requirement allocated according to defect type + material parameter” in the to-be-enhanced shaft data requirement parameter set, and compare the sample total quantity (real sample quantity + simulation sample quantity) of each merged subset with the corresponding requirement quantity one by one: if the sample total quantity of the merged subset is less than the requirement quantity, calculate “sample gap quantity = requirement quantity - sample total quantity of merged subset”, and record the gap quantity and the corresponding merged subset category (such as “micro-crack + alloy steel + chrome plating” merged subset gap quantity is a certain value).
[0208] For the merged subset with “no corresponding real data”, the sample total quantity is the simulation sample quantity, and if it is less than the requirement quantity, the gap quantity is also calculated; for the merged subset with “no corresponding simulation data”, if the sample total quantity is less than the requirement quantity, mark it as “need to supplement simulation data” and include it in the gap statistics range.
[0209] Step S157: For the merged subsets with sample gap quantity, screen the preliminary shaft simulation data with similar features but not included in the effective shaft simulation data set, compare the features of the preliminary shaft simulation data, and supplement the preliminary shaft simulation data with feature coincidence degree greater than or equal to feature coincidence degree threshold and shape coincidence degree greater than or equal to shape coincidence degree threshold to the corresponding merged subset until the sample quantity of the merged subset is greater than or equal to the data sample quantity requirement in the to-be-enhanced shaft data requirement parameter set.
[0210] From the preliminary shaft simulation data set, screen out the simulation samples marked as “to be removed” but with similar features (such as feature coincidence degree slightly lower than the threshold but shape coincidence degree meets the standard, or vice versa, and the difference is small), match them according to the “defect type + material parameter” category of the merged subset, and get the candidate supplement samples corresponding to each gap merged subset.
[0211] For the candidate supplementary samples, repeat steps S143-S145 for feature consistency comparison: using the same subset of real features as before as a reference, recalculate the feature fit and morphological fit of the candidate samples. If both recalculated fits reach the corresponding threshold (for complex defects, the reduced threshold can be used), then the candidate sample is included in the scope of valid simulation data, supplemented to the corresponding merged subset, and the sample count of the merged subset is updated.
[0212] If gaps still exist in the merged subset after supplementing candidate samples (e.g., insufficient number of candidate samples or failure to meet the standard after re-comparison), the generative adversarial network is restarted. For the defect type and material parameters of the gap category, the number of generator feature mapping layers is adjusted (e.g., increasing the number of layers to improve feature capture accuracy) and the discriminator discrimination threshold is adjusted (e.g., fine-tuning the threshold to improve the quality of generated samples). A new batch of preliminary simulation data is generated. After feature consistency comparison, qualified samples are added to the merged subset until the number of samples in all merged subsets meets the quantity requirements in the set of required parameters.
[0213] Step S158: Integrate all merged subsets and sort them according to the type of surface defect on the shaft to form a structured final shaft enhancement data set.
[0214] All merged subsets are sorted according to the importance of the defect type (such as failure rate and detection difficulty). For example, the sorting order is "microcrack", "scratch + rust composite", "blurred edge defect", "ordinary scratch", and "ordinary rust". The merged subsets under each defect type are then sorted according to the alphabetical order of the material parameters (such as "alloy steel + chrome plating", "alloy steel + grinding", and "45 steel + turning").
[0215] The sorted and merged subsets are integrated to form the final axle enhancement data set. Under the root directory of the set, a first-level folder is created according to the defect type. Under each first-level folder, a second-level folder (i.e., merged subset folder) is created according to "defect type + material parameter". The second-level folders contain "real data" and "simulation data" subfolders and classification description documents.
[0216] Simultaneously, a master directory document for the enhanced dataset is generated, which records in detail the number of real samples, simulation samples, and total samples corresponding to each defect type and material parameter, as well as key information such as the dataset creation time, generative adversarial network configuration parameters, and feature consistency comparison threshold, making it easy for users to quickly understand the overall structure and attributes of the dataset.
[0217] Step S159: Count the total number of real shaft data samples and effective shaft simulation data samples in the final shaft enhancement data set, and make the total number greater than or equal to the number of data samples required in the set of required parameters for the shaft data to be enhanced.
[0218] The total number of all real samples in the final augmented data set and the total number of all valid simulation samples are counted respectively, and the sum of the two is calculated to obtain the total number of samples of the augmented data set. The total number is compared with the total target sample number in the shaft data requirement parameter set to be augmented: if the total number is greater than or equal to the total target number, and the sample number of each defect type and each material parameter meets the subdivision requirement, the data set integration is completed; if the total number meets the requirement but there is still a gap in some categories (such as the sample number of a certain material parameter is insufficient), steps S157-S158 are repeated to supplement qualified simulation samples for the gap category until all categories meet the requirement.
[0219] The finally generated data set verification report contains statistical information such as total sample number, real and simulation sample number ratio, defect type sample distribution, material parameter sample distribution, and comparison results with the requirement parameter set, ensuring that the final shaft augmented data set completely meets the training data requirement of the industrial motor shaft complex defect detection model, and can be directly used for model training, verification and testing.
[0220] Figure 2 The hardware structure of the shaft simulation data augmentation system 100 for implementing the above-mentioned shaft simulation data augmentation method based on a generative adversarial network provided by the embodiments of the present application is shown in FIG. 1. Figure 2 As shown in FIG. 1, the shaft simulation data augmentation system 100 based on a generative adversarial network can include a processor 110, a machine readable storage medium 120, a bus 130 and a communication unit 140.
[0221] In the specific implementation process, one or more processors 110 execute computer executable instructions stored in the machine readable storage medium 120, so that the processor 110 can execute the shaft simulation data augmentation method based on a generative adversarial network of the above method embodiments. The processor 110, the machine readable storage medium 120 and the communication unit 140 are connected through the bus 130, and the processor 110 can be used to control the transceiving action of the communication unit 140.
[0222] The specific implementation process of the processor 110 can refer to the above-mentioned various method embodiments executed by the shaft simulation data augmentation system 100 based on a generative adversarial network, which has similar implementation principles and technical effects, and will not be described here.
[0223] In addition, the present embodiment also provides a readable storage medium, wherein computer executable instructions are set in the readable storage medium, and when the processor runs the computer executable instructions, the above-mentioned shaft simulation data augmentation method based on a generative adversarial network is realized.
[0224] It should be noted that the foregoing description of embodiments of the application has been presented for the purposes of simplicity and explanation, and is not intended to limit the description of embodiments of the application to the forms described. As should be appreciated by those skilled in the art, the specific form described above is intended to be illustrative only and not limiting of the embodiments of the application.
Claims
1. A method for data augmentation of rotation-invariant data by combining a generative adversarial network, characterized in that, The method comprises: determining a set of shaft data demand parameters to be enhanced, the set of shaft data demand parameters to be enhanced comprising a shaft surface defect type, a shaft material characterization parameter, and a data sample quantity demand; configuring shaft adaptation parameters of a generative adversarial network, the shaft adaptation parameters of the generative adversarial network comprising a number of shaft feature mapping layers of a generator, a shaft data discrimination threshold of a discriminator, and an iteration termination condition of network training; obtaining a real shaft data sample set, driving the generative adversarial network to operate according to the shaft adaptation parameters of the generative adversarial network based on the real shaft data sample set, and generating a preliminary shaft simulation data set; performing feature consistency comparison processing on the preliminary shaft simulation data set and the real shaft data sample set, eliminating preliminary shaft simulation data that is inconsistent in features, and obtaining a set of effective shaft simulation data that is consistent in features; integrating the set of effective shaft simulation data and the real shaft data sample set to form a final shaft enhanced data set, the final shaft enhanced data set comprising real shaft data samples and effective shaft simulation data.
2. The data augmentation method of claim 1, wherein, The determination of the set of shaft data demand parameters to be enhanced comprises: analyzing a shaft surface defect type coverage in an existing real shaft data sample set to obtain a defect type coverage result, the defect type coverage result comprising covered defect types and sample quantities of each covered defect type; based on the defect type coverage result, identifying a defect type with a sample quantity less than a preset training demand threshold, determining the defect type as a target defect type to be supplemented, and the target defect type to be supplemented belonging to the shaft surface defect type in the set of shaft data demand parameters to be enhanced; collecting shaft material characterization parameters of all samples in the existing real shaft data sample set, counting distribution frequencies of each shaft material characterization parameter, and obtaining a material parameter distribution statistical result, the material parameter distribution statistical result comprising occurrence times and proportions of each material characterization parameter; based on the material parameter distribution statistical result, determining a material characterization parameter category that needs to be supplemented and strengthened, determining the material characterization parameter of the material characterization parameter category as a target material characterization parameter, and the target material characterization parameter belonging to the shaft material characterization parameter in the set of shaft data demand parameters to be enhanced; according to a total training sample quantity requirement of a shaft data driving model, combining a sample quantity of the existing real shaft data sample set, calculating a simulation data sample quantity that needs to be supplemented, and determining the simulation data sample quantity as a target sample quantity, the target sample quantity belonging to the data sample quantity demand in the set of shaft data demand parameters to be enhanced; integrating the target defect type to be supplemented, the target material characterization parameter, and the target sample quantity to form the set of shaft data demand parameters to be enhanced.
3. The method of claim 2, wherein the data augmentation is performed by rotating the image by a random angle. The analysis of the shaft surface defect type coverage in the existing real shaft data sample set to obtain the defect type coverage result comprises: Obtain all data samples in the existing real shaft data sample set, read the defect annotation information of each data sample one by one, and the defect annotation information contains the shaft surface defect type corresponding to the sample; A defect type statistical table is established, the columns of the defect type statistical table contain the shaft surface defect type name, and the rows of the defect type statistical table are used to record the sample quantity of each defect type; Iterate through the defect annotation information of each data sample, find the corresponding defect type name column in the defect type statistical table, and accumulate the sample quantity of the corresponding defect type name column in the defect type statistical table; After completing the iteration of all data samples, count the defect types with sample quantity greater than 0 in the defect type statistical table, and determine the defect types as covered defect types; The covered defect types and the sample quantity corresponding to each covered defect type are arranged into structured data, and the structured data is the defect type coverage result.
4. The data augmentation method of claim 1, wherein, The configuration of the shaft adaptation parameter of the generative adversarial network comprises: Based on the target defect type to be supplemented in the shaft data demand parameter set to be enhanced, the shaft surface feature complexity corresponding to different defect types is analyzed, and when the parameter value corresponding to the shaft surface feature complexity is greater than a preset complexity threshold, the feature mapping layer number required by the generator is greater than a preset basic layer number, and the initial shaft feature mapping layer number of the generator is determined accordingly; Test and adjust the initial shaft feature mapping layer number, compare the feature similarity of the simulation data output by the generator under different initial shaft feature mapping layer numbers with the real shaft data sample, select the initial shaft feature mapping layer number with the feature similarity greater than the feature similarity corresponding to other initial layer numbers as the final shaft feature mapping layer number of the generator, and the final shaft feature mapping layer number of the generator belongs to the shaft adaptation parameter of the generative adversarial network; Based on the target material representation parameter in the shaft data demand parameter set to be enhanced, the pixel distribution feature of the real shaft data sample corresponding to the target material representation parameter is extracted, the mean value of the pixel distribution feature is taken as a reference value, an allowed fluctuation range around the reference value is set, and the boundary value of the allowed fluctuation range is determined as the initial shaft data discrimination threshold of the discriminator; Test the discrimination accuracy of the discriminator on real shaft data samples and noise data under different initial shaft data discrimination thresholds, retain the initial shaft data discrimination threshold with discrimination accuracy greater than a preset accuracy threshold, and select the threshold that can maximize the discrimination between real data and noise data as the final shaft data discrimination threshold of the discriminator, which belongs to the shaft adaptation parameter of the generative adversarial network; Based on the target sample quantity in the shaft data demand parameter set to be enhanced, combine the generated sample quantity of a single iteration of the generative adversarial network, calculate the required minimum iteration number, and take the minimum iteration number as the initial iteration termination condition for training the generative adversarial network; The training stability of the generative adversarial network under different initial iteration termination conditions is monitored, when the generation loss value and the discrimination loss value of the generative adversarial network both tend to be stable and no longer significantly decrease, the iteration number at this time is recorded, the iteration number is compared with the minimum iteration number, and the iteration number greater than the minimum iteration number is selected as the final iteration termination condition of the generative adversarial network training, and the final iteration termination condition of the generative adversarial network training belongs to the pivot adaptation parameter of the generative adversarial network; The pivot feature mapping layer number of the final generator, the pivot data discrimination threshold value of the final discriminator and the iteration termination condition of the final generative adversarial network training are integrated to form the pivot adaptation parameter of the generative adversarial network.
5. The data augmentation method of claim 4, wherein, Based on the target material characteristic parameter in the to-be-enhanced pivot data demand parameter set, the pixel distribution feature of the real pivot data sample corresponding to the target material characteristic parameter is extracted, the mean value of the pixel distribution feature is taken as a reference value, an allowable fluctuation range around the reference value is set, and the boundary value of the allowable fluctuation range is determined as the initial pivot data discrimination threshold value of the discriminator, including: The target material characteristic parameter is extracted from the to-be-enhanced pivot data demand parameter set, and the specific category and value range of the target material characteristic parameter are determined; All samples in which the material characteristic parameter belongs to the target material characteristic parameter category are filtered out from the real pivot data sample set to form a target material sample subset; The pixel value of each sample in the target material sample subset is extracted to obtain the pixel value of each sample to form a pixel value set; The mean value of the pixel value set is calculated, and the mean value is the reference value of the pixel distribution feature of the real pivot data sample corresponding to the target material characteristic parameter; The dispersion degree of the pixel value set is analyzed, and the size of the allowable fluctuation range is determined according to the dispersion degree; when the parameter value corresponding to the dispersion degree is greater than the preset dispersion threshold value, the allowable fluctuation range is greater than the preset basic range; when the parameter value corresponding to the dispersion degree is less than or equal to the preset dispersion threshold value, the allowable fluctuation range is less than or equal to the preset basic range; Based on the reference value and the determined size of the allowable fluctuation range, the upper limit value and the lower limit value of the allowable fluctuation range are calculated, the upper limit value is the reference value plus the size of the allowable fluctuation range, and the lower limit value is the reference value minus the size of the allowable fluctuation range; The upper limit value and the lower limit value of the allowable fluctuation range are determined as the initial pivot data discrimination threshold value of the discriminator, wherein the upper limit value is taken as the pixel value upper limit standard of the discriminator for determining the data as real data, and the lower limit value is taken as the pixel value lower limit standard of the discriminator for determining the data as real data; The pixel value of the target material sample subset is compared with the initial pivot data discrimination threshold value, the proportion of the pixel value in the allowable fluctuation range to the total pixel value of the target material sample subset is calculated, and if the proportion is less than a preset proportion threshold value, the size of the allowable fluctuation range is adjusted, the initial pivot data discrimination threshold value is recalculated, and the process is repeated until the proportion is greater than or equal to the preset proportion threshold value.
6. The data augmentation method of rotational simulation of a GAN according to claim 1, wherein, The real rotating shaft data sample set is obtained, and the generative adversarial network is driven to run according to the rotating shaft adaptation parameters of the generative adversarial network based on the real rotating shaft data sample set, and a preliminary rotating shaft simulation data set is generated, including: Collect rotating shaft surface image data and corresponding material parameter records under different working conditions, and remove invalid data from the collected image data to obtain cleaned rotating shaft data. The different working conditions include working conditions under different running times and different load intensities. Label the cleaned rotating shaft data to form a real rotating shaft data sample set. The labeling content includes rotating shaft surface defect types and rotating shaft material representation parameters. The real rotating shaft data sample set is divided into a training subset and a verification subset according to a preset proportion. The training subset is used to drive the generative adversarial network training, and the verification subset is used to monitor the generation effect in the network training process. The training subset is input into the generator of the generative adversarial network, and the rotating shaft features of the training subset are processed by multi-layer mapping according to the rotating shaft feature mapping layer number of the generator in the rotating shaft adaptation parameters of the generative adversarial network to generate initial simulation data. The initial simulation data and the real data of the training subset are jointly input into the discriminator of the generative adversarial network. The discriminator discriminates the authenticity of the initial simulation data according to the rotating shaft data discrimination threshold of the discriminator in the rotating shaft adaptation parameters of the generative adversarial network, and outputs a discrimination result. According to the discrimination result, the network parameters of the generator and the discriminator are adjusted through back propagation. The process of generating initial simulation data by the generator, discriminating by the discriminator, and adjusting the network parameters is repeated until the iteration termination condition of the network training in the rotating shaft adaptation parameters of the generative adversarial network is reached. All initial simulation data output by the generator at the iteration termination are integrated to form a preliminary rotating shaft simulation data set.
7. The method of claim 6, wherein the data augmentation is performed by rotating the image by a random angle. The real rotating shaft data sample set is divided into a training subset and a verification subset according to a preset proportion, including: Clearly define the functional requirements of the generative adversarial network training on the training subset and the verification subset. Based on the functional requirements, set the division proportion of the training subset and the verification subset. The division proportion needs to ensure that the number of samples in the training subset is greater than the minimum sample threshold for network training, and the number of samples in the verification subset is greater than the minimum sample threshold for effect monitoring. According to the rotating shaft surface defect types and the rotating shaft material representation parameters, the real rotating shaft data sample set is layered. In each layer, samples are randomly selected according to the set division proportion. The part with a larger division proportion in the selected samples is assigned to the training subset, and the part with a smaller division proportion is assigned to the verification subset. The number of samples of different rotating shaft surface defect types and the number of samples of different rotating shaft material representation parameters included in the training subset and the verification subset are counted, and the sample distribution of the two is made to deviate from the overall distribution of the real rotating shaft data sample set by less than a preset distribution deviation threshold. If the number of samples in the training subset is less than the minimum sample threshold for network training, adjust the division proportion to increase the number of samples in the training subset and reduce the number of samples in the verification subset until the number of samples in the training subset is greater than the minimum sample threshold for network training. If the number of samples in the verification subset is less than the minimum sample threshold for effect monitoring, the division ratio is adjusted to increase the number of samples in the verification subset and reduce the number of samples in the training subset until the number of samples in the verification subset is greater than the minimum sample threshold for effect monitoring; After the adjustment is completed, the finally selected samples are integrated to form the training subset and the verification subset.
8. The data augmentation method of rotational simulation of a GAN according to claim 1, wherein, The feature consistency comparison processing is performed on the preliminary shaft simulation data set and the real shaft data sample set, and the preliminary shaft simulation data with inconsistent features is removed to obtain the effective shaft simulation data set with consistent features, including: Extracting the shaft surface texture features and defect morphology features of each sample from the real shaft data sample set, and integrating to form a real shaft feature set, the real shaft feature set containing a real texture feature sub-set and a real defect morphology feature sub-set; Extracting the shaft surface texture features and defect morphology features of each sample from the preliminary shaft simulation data set, and integrating to form a simulation shaft feature set, the simulation shaft feature set containing a simulation texture feature sub-set and a simulation defect morphology feature sub-set; Comparing each simulation texture feature in the simulation texture feature sub-set with the real texture feature under the corresponding material representation parameter in the real texture feature sub-set, calculating the feature coincidence degree of the two, the real texture feature under the corresponding material representation parameter being the texture feature of the real data sample with the same material representation parameter as the simulation data sample; Comparing each simulation defect morphology feature in the simulation defect morphology feature sub-set with the real defect morphology feature under the corresponding defect type in the real defect morphology feature sub-set, calculating the morphology coincidence degree of the two, the real defect morphology feature under the corresponding defect type being the defect morphology feature of the real data sample with the same defect type as the simulation data sample; Comparing the feature coincidence degree of each preliminary shaft simulation data sample with the feature coincidence degree threshold and the morphology coincidence degree with the morphology coincidence degree threshold, if the feature coincidence degree of the preliminary shaft simulation data sample is greater than or equal to the feature coincidence degree threshold and the morphology coincidence degree is greater than or equal to the morphology coincidence degree threshold, it is determined that the sample is the simulation data with consistent features; if the feature coincidence degree of the preliminary shaft simulation data sample is less than the feature coincidence degree threshold or the morphology coincidence degree is less than the morphology coincidence degree threshold, it is determined that the sample is the simulation data with inconsistent features, which is removed; Collecting all the simulation data determined to be consistent in features, and integrating to form the effective shaft simulation data set with consistent features.
9. The data augmentation method of claim 8, wherein, The extracting the shaft surface texture features and defect morphology features of each sample from the real shaft data sample set, and integrating to form a real shaft feature set, includes: Obtaining each data sample in the real shaft data sample set, reading the image data and corresponding annotation information of each data sample, the annotation information including the shaft surface defect type and the shaft material representation parameter; Performing texture feature extraction on the image data of each data sample, using texture feature extraction logic to capture the arrangement rule and gray scale change mode of pixels in the image data to obtain the shaft surface texture features of the data sample; The image data of each data sample is subjected to defect morphology feature extraction, and according to the shaft surface defect type in the data sample annotation information, the defect area in the image data is located, the shape, size and edge contour features of the defect area are extracted, and the defect morphology features of the data sample are obtained. A real texture feature sub-set is established, and the shaft surface texture features of each data sample are classified and stored according to the shaft material representation parameters of the data sample. The texture features of data samples with the same material representation parameters are classified into the same category. A real defect morphology feature sub-set is established, and the defect morphology features of each data sample are classified and stored according to the shaft surface defect type of the data sample. The defect morphology features of data samples with the same defect type are classified into the same category. The number of texture features in each category in the real texture feature sub-set is counted. If the number of texture features is less than the number of data samples corresponding to the material representation parameters, the data samples without extracted texture features are processed again to supplement the missing texture features. The number of defect morphology features in each category in the real defect morphology feature sub-set is counted. If the number of defect morphology features is less than the number of data samples corresponding to the defect type, the data samples without extracted defect morphology features are processed again to supplement the missing defect morphology features. The real texture feature sub-set and the real defect morphology feature sub-set are integrated to form a real shaft feature set.
10. A data augmentation system for rotationally symmetric data, which combines a generative adversarial network, characterized in that, The shaft simulation data enhancement system combined with the generative adversarial network includes a processor and a memory, the memory and the processor are connected, the memory is used to store programs, instructions or codes, and the processor is used to run the programs, instructions or codes in the memory to realize the shaft simulation data enhancement method combined with the generative adversarial network in any one of claims 1-9.
Citation Information
Patent Citations
Electric cooker liner image data enhancement method based on mask generative adversarial network
CN115908379A
Lipstick product surface defect data augmentation method based on small sample feature migration
CN116229205A
Strip steel surface defect data enhancement method based on Wasserstein GAN
CN118658023A
Steel surface defect data enhancement method and system based on defect perception generative adversarial network
CN120431423A
Chemical material detection method and system based on deep learning
CN120852319A