Metal defect visual detection method and system for gearbox shell

Through the combination of multiple convolutional neural network models and the self-attention mechanism, the accuracy and uniformity problems of metal defect detection on the gearbox housing were solved, and efficient metal defect detection of the gearbox housing was achieved.

CN120673156AInactive Publication Date: 2025-09-19XUZHOU CHAOJIE ELECTRIC VEHICLE PARTS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510787522.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-09-19
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing deep learning computer vision inspection models have difficulty accurately detecting irregularly shaped metal defects on gearbox housings, and lack a unified inspection method for various types of gearbox housings.

Method used

Multiple convolutional neural network models are used to form a heterogeneous model combination, and the self-attention mechanism is combined to integrate features. The metal defect confidence is generated through image preprocessing, convolution calculation, self-attention mechanism and feedforward neural network.

Benefits of technology

It realizes accurate detection of irregularly shaped metal defects on the gearbox housing and can uniformly detect metal defects of various types of gearbox housings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673156A_ABST
    Figure CN120673156A_ABST
Patent Text Reader

Abstract

The invention discloses a metal defect visual detection method and system for a gearbox shell, and the method comprises the steps: obtaining a plane image of the gearbox shell through an image collection device, carrying out the graying and normalization processing, and generating a preprocessing image; respectively inputting the preprocessed image into N convolutional neural network defect detection models, and carrying out parallel calculation to generate N feature vectors; the N feature vectors are sequentially combined into a sequence form, self-attention mechanism calculation without position coding is carried out, and an attention sequence is generated; the first representation vector in the attention sequence is input into the feedforward neural network model, and the confidence coefficient of the metal defect existing on the surface of the current gearbox shell is generated. And metal defect detection under a unified normal form can be carried out on various gearbox shells.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision detection, and in particular to a method and system for visually detecting metal defects of a gearbox housing. Background Art

[0002] During the production of gearbox housings, various reasons can lead to surface metal defects such as pores, cold shuts, shrinkage cavities, porosity, cracks, mechanical damage, and inclusions. To prevent these defective gearbox housings from entering the production line and impacting product quality, metal quality inspection is required. Currently, computer vision inspection technology is commonly used to map and detect metal defects on gearbox housings using deep learning computer vision inspection models.

[0003] When using computer vision inspection technology to map and detect metal defects on gearbox housings, conventional deep learning computer vision inspection models have difficulty accurately detecting the irregular shapes of metal defects on gearbox housings. Furthermore, various gearbox housings have different shapes, and currently there is no method that can uniformly detect metal defects across all types of gearbox housings. Summary of the Invention

[0004] The purpose of the present invention is to provide a method and system for visual detection of metal defects in gearbox housings, aiming to solve the problem that in the process of using computer vision detection technology to detect metal defects in gearbox housings, the shapes of metal defects on gearbox housings are irregular, and conventional deep learning computer vision detection models are difficult to perform accurate detection. In addition, various types of gearbox housings have different shapes, and at this stage there is no relevant method that can perform unified metal defect detection on various types of gearbox housings.

[0005] In view of the above problems, the present application provides a method and system for visually detecting metal defects of a gearbox housing.

[0006] A first aspect disclosed in the present application provides a method for visually inspecting metal defects in a gearbox housing, the method comprising: Step 1: Obtain a plane image of the gearbox housing through an image acquisition device, perform grayscale and normalization processing, and generate a preprocessed image; Step 2: Input the preprocessed image into N convolutional neural network defect detection models respectively, perform calculations in parallel, and generate N feature vectors; Step 3: Combine the N feature vectors into a sequence, perform self-attention calculation without position encoding, and generate an attention sequence; Step 4: Input the first representation vector in the attention sequence into the feedforward neural network model to generate the confidence that there are metal defects on the surface of the current gearbox housing.

[0007] Preferably, the step 1 specifically includes: Step 1.1: Use an industrial camera to obtain an RGB three-channel planar image of the gearbox housing to generate an initial image; Step 1.2: Iteratively perform anisotropic diffusion filtering on the initial image to retain edge features while suppressing Gaussian noise to generate a denoised image; Step 1.3: Convert the RGB image to the YCrCb color space and extract the brightness component to generate a grayscale image; Step 1.4: Calculate the mean and standard deviation of the brightness values ​​of all pixels in the grayscale image, perform Z-score normalization on the brightness value of each pixel in the grayscale image, then apply dynamic range compression and map the brightness value of each pixel to the (0, 1) interval through the softmax function to generate a preprocessed image.

[0008] Preferably, the step 2 specifically includes: Step 2.1: Construct a heterogeneous model combination consisting of a ResNet34 neural network model, a DenseNet121 neural network model, and an EfficientNet-B3 neural network model. The output results of the ResNet34 neural network model, the DenseNet121 neural network model, and the EfficientNet-B3 neural network model are of dimension 1 after the full connection layer adaptation. d The eigenvector of Step 2.2: Input the preprocessed image into the ResNet34 neural network model, the DenseNet121 neural network model, and the EfficientNet-B3 neural network model respectively. Use different processes to perform parallel calculations for each neural network model. Use CUDA streams to isolate the calculation tasks of different neural network models and generate corresponding feature vectors. Step 2.3: Perform layer normalization on each feature vector, and then map each feature vector to a unified feature space through affine transformation.

[0009] Preferably, the step 3 specifically includes: Step 3.1: Combine the N eigenvectors into a sequence to form a sequence matrix. Perform matrix multiplication on the sequence matrix with the query projection matrix, key projection matrix, and value projection matrix to generate the query matrix. Q , key matrix K Sum Matrix V ; Step 3.2: Pass The self-attention mechanism is calculated in the same way, and the attention sequence is generated through the residual connection, where for K The transposed matrix of .

[0010] Preferably, the step 4 specifically includes: Step 4.1: Obtain the first representation vector in the attention sequence and input it into the GeLU activation function to enhance the nonlinear expression ability, and perform layer normalization to stabilize the feature distribution and generate a global vector; Step 4.2: Input the global vector into the feedforward neural network model, and input the output result into the sigmoid function to generate the confidence level of the presence of metal defects on the surface of the current gearbox housing, wherein the feedforward neural network model includes an input layer, a hidden layer, and an output layer.

[0011] A second aspect disclosed in the present application provides a metal defect visual detection system for a gearbox housing, the system being used in the above-mentioned metal defect visual detection method for a gearbox housing, the system comprising: A preprocessing module, configured to acquire a planar image of the gearbox housing through an image acquisition device, perform grayscale and normalization processing on the image, and generate a preprocessed image; A convolution module, which is used to input the preprocessed images into N convolutional neural network defect detection models respectively, perform calculations in parallel, and generate N feature vectors; An attention module, which is used to sequentially combine N feature vectors into a sequence, perform a self-attention mechanism calculation without positional encoding, and generate an attention sequence; A confidence module is used to input the first representation vector in the attention sequence into the feedforward neural network model to generate a confidence that there is a metal defect on the surface of the current gearbox housing.

[0012] The third aspect disclosed in the present application provides a computer device including a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above-mentioned method for visual detection of metal defects of a gearbox housing when executing the computer program.

[0013] The fourth aspect disclosed in the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-mentioned method for visual inspection of metal defects of a gearbox housing.

[0014] The fifth aspect disclosed in the present application provides a computer program product, including a computer program or instructions, which, when executed by a processor, implement the steps of the above-mentioned method for visual inspection of metal defects of a gearbox housing.

[0015] The beneficial effects of the present invention are: (1) Using multiple convolutional neural network models to form a heterogeneous model combination, the metal defect detection of the gearbox housing is carried out from multiple angles and levels, and the computer vision detection of irregularly shaped metal defects on the gearbox housing can be performed accurately; (2) The self-attention mechanism is used to integrate the inference results of multiple convolutional neural network models, and the instance information of the gearbox housing itself is used as a blueprint for feature capture, which can perform metal defect detection on various types of gearbox housings under a unified paradigm. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0017] Figure 1 The figure is an overall flow chart of a metal defect visual inspection method for a gearbox housing.

[0018] Figure 2 This is an overall structural diagram of a metal defect visual inspection system for a gearbox housing. DETAILED DESCRIPTION

[0019] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0020] Example 1: like Figure 1 As shown, an embodiment of the present application provides a method for visually detecting metal defects of a gearbox housing, the method comprising: Step 1: Obtain a plane image of the gearbox housing through an image acquisition device, perform grayscale and normalization processing, and generate a preprocessed image.

[0021] Step 1 specifically includes the following steps: Step 1.1: Use a ring-shaped LED array light source for illumination compensation and use an industrial camera to acquire a three-channel RGB image of the gearbox housing to generate an initial image. In this initial image, the gearbox housing occupies the main part of the frame, and the background color is a solid color. Step 1.2: The initial image is iteratively subjected to anisotropic diffusion filtering. The neighborhood gradient calculation uses the Sobel operator with 3 iterations and a transmission coefficient parameter of 25 to retain edge features while suppressing Gaussian noise to generate a denoised image. Step 1.3: Convert the RGB image to the YCrCb color space. Since YCrCb is a color space that separates brightness and chrominance, the color information and brightness information of the image are separated in the YCrCb color space. Extract the brightness component to generate a grayscale image. Step 1.4: Calculate the mean and standard deviation of the brightness values ​​of all pixels in the grayscale image, perform Z-score normalization on the brightness value of each pixel in the grayscale image, then apply dynamic range compression and map the brightness value of each pixel to the (0, 1) interval through the softmax function to generate a preprocessed image.

[0022] Step 2: Input the preprocessed image into N convolutional neural network defect detection models respectively, perform calculations in parallel, and generate N feature vectors.

[0023] Step 2 specifically includes the following steps: Step 2.1: Construct a heterogeneous model combination consisting of a ResNet34 neural network model that is sensitive to shallow image features, a DenseNet121 neural network model that implements dense feature reuse, and an EfficientNet-B3 neural network model that implements multi-scale perception. The output results of the ResNet34 neural network model, the DenseNet121 neural network model, and the EfficientNet-B3 neural network model are adapted to the dimension of the fully connected layer. d The eigenvector of Step 2.2: Input the preprocessed image into the ResNet34 neural network model, the DenseNet121 neural network model, and the EfficientNet-B3 neural network model respectively. Use different processes to perform parallel calculations for each neural network model. Use CUDA streams to isolate the calculation tasks of different neural network models and generate corresponding feature vectors. Step 2.3: Perform layer normalization on each eigenvector, and then map each eigenvector to a unified feature space through an affine transformation consisting of a projection matrix and an offset.

[0024] Step 3: Combine the N feature vectors into a sequence in sequence, perform self-attention calculation without position encoding, and generate an attention sequence.

[0025] Step 3 specifically includes the following steps: Step 3.1: Combine the N eigenvectors into a sequence to form a sequence matrix. Perform matrix multiplication on the sequence matrix with the query projection matrix, key projection matrix, and value projection matrix to generate the query matrix. Q , key matrix K Sum Matrix V , where the dimension of the projection matrix is d × d ; Step 3.2: Pass The self-attention mechanism is calculated in the same way, and the attention sequence is generated through the residual connection, where, for K The purpose of the residual connection is to prevent the gradient from disappearing during the training process.

[0026] Step 4: Input the first representation vector in the attention sequence into the feedforward neural network model to generate the confidence that there are metal defects on the surface of the current gearbox housing.

[0027] Step 4 specifically includes: Step 4.1: Obtain the first representation vector in the attention sequence and input it into the GeLU activation function to enhance the nonlinear expression ability, and perform layer normalization to stabilize the feature distribution and generate a global vector; Step 4.2: Input the global vector into the feedforward neural network model, and input the output result into the sigmoid function to generate the confidence level of the presence of metal defects on the surface of the current gearbox housing, wherein the feedforward neural network model includes an input layer, a hidden layer, and an output layer.

[0028] Specifically, all neural network models and projection matrices in this method need to be trained using the conventional deep learning training method. The dataset is divided into training set, validation set, and test set in an 8:1:1 ratio. The loss function uses the binary cross-entropy loss function. Dropout and early stopping methods are used during training to prevent overfitting. Accuracy and recall are used to evaluate its accuracy.

[0029] In summary, the metal defect visual inspection method for a gearbox housing provided by the embodiments of the present application has the following technical effects: (1) Using multiple convolutional neural network models to form a heterogeneous model combination, the metal defect detection of the gearbox housing is carried out from multiple angles and levels, and the computer vision detection of irregularly shaped metal defects on the gearbox housing can be performed accurately; (2) The self-attention mechanism is used to integrate the inference results of multiple convolutional neural network models, and the instance information of the gearbox housing itself is used as a blueprint for feature capture, which can perform metal defect detection on various types of gearbox housings under a unified paradigm.

[0030] Example 2: Based on the same inventive concept as the method for visually detecting metal defects of a gearbox housing in Example 1, Figure 2 As shown, the present application provides a metal defect visual inspection system for a gearbox housing, the system comprising: A preprocessing module, configured to acquire a planar image of the gearbox housing through an image acquisition device, perform grayscale and normalization processing on the image, and generate a preprocessed image; A convolution module, which is used to input the preprocessed images into N convolutional neural network defect detection models respectively, perform calculations in parallel, and generate N feature vectors; An attention module, which is used to sequentially combine N feature vectors into a sequence, perform a self-attention mechanism calculation without positional encoding, and generate an attention sequence; A confidence module is used to input the first representation vector in the attention sequence into the feedforward neural network model to generate a confidence that there is a metal defect on the surface of the current gearbox housing.

[0031] Through the detailed description of a metal defect visual detection method for a gearbox housing in the foregoing description, those skilled in the art can clearly understand a metal defect visual detection system for a gearbox housing in this embodiment. Since it corresponds to the method disclosed in the embodiment, the description is relatively simple. For relevant matters, please refer to the method section.

[0032] Example 3: In a third embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps of the above-mentioned method for visual inspection of metal defects of a gearbox housing when executing the computer program.

[0033] Example 4: In a fourth embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned method for visual inspection of metal defects of a gearbox housing are implemented.

[0034] Embodiment 5: In the fifth embodiment, a computer program product is provided, including a computer program or instructions, which implement the steps of the above-mentioned method for visual inspection of metal defects of a gearbox housing when the computer program or instructions are executed by a processor.

[0035] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0036] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for visually inspecting metal defects of a gearbox housing, characterized in that: The method comprises: Step 1: Obtain a plane image of the gearbox housing through an image acquisition device, perform grayscale and normalization processing, and generate a preprocessed image; Step 2: Input the preprocessed image into N convolutional neural network defect detection models respectively, perform calculations in parallel, and generate N feature vectors; Step 3: Combine the N feature vectors into a sequence, perform self-attention calculation without position encoding, and generate an attention sequence; Step 4: Input the first representation vector in the attention sequence into the feedforward neural network model to generate the confidence that there are metal defects on the surface of the current gearbox housing.

2. A method for visually inspecting metal defects of a gearbox housing according to claim 1, characterized in that: The step 1 specifically includes: Step 1.1: Use an industrial camera to obtain an RGB three-channel planar image of the gearbox housing to generate an initial image; Step 1.2: Iteratively perform anisotropic diffusion filtering on the initial image to generate a denoised image; Step 1.3: Convert the RGB image to the YCrCb color space and extract the brightness component to generate a grayscale image; Step 1.4: Calculate the mean and standard deviation of the brightness values ​​of all pixels in the grayscale image, perform Z-score normalization on the brightness value of each pixel in the grayscale image, and then use the softmax function to map the brightness value of each pixel to the (0, 1) interval to generate a preprocessed image.

3. A method for visually inspecting metal defects of a gearbox housing according to claim 2, characterized in that: The step 2 specifically includes: Step 2.1: Construct a heterogeneous model combination consisting of a ResNet34 neural network model, a DenseNet121 neural network model, and an EfficientNet-B3 neural network model. The output results of the ResNet34 neural network model, the DenseNet121 neural network model, and the EfficientNet-B3 neural network model are of dimension 1 after the full connection layer adaptation. d The eigenvector of Step 2.2: Input the preprocessed image into the ResNet34 neural network model, the DenseNet121 neural network model, and the EfficientNet-B3 neural network model respectively. For each neural network model, different processes are used to perform parallel calculations to generate the corresponding feature vectors. Step 2.3: Perform layer normalization on each feature vector, and then map each feature vector to a unified feature space through affine transformation.

4. A method for visually inspecting metal defects of a gearbox housing according to claim 3, characterized in that: The step 3 specifically includes: Step 3.1: Combine the N eigenvectors into a sequence to form a sequence matrix. Perform matrix multiplication on the sequence matrix with the query projection matrix, key projection matrix, and value projection matrix to generate the query matrix. Q , key matrix K Sum Matrix V ; Step 3.2: Pass The self-attention mechanism is calculated in the same way, and the attention sequence is generated through the residual connection, where for K The transposed matrix of .

5. A method for visually inspecting metal defects of a gearbox housing according to claim 4, characterized in that: The step 4 specifically includes: Step 4.1: Get the first representation vector in the attention sequence and input it into the GeLU activation function, and perform layer normalization to generate a global vector; Step 4.2: Input the global vector into the feedforward neural network model, and input the output result into the sigmoid function to generate the confidence level of the presence of metal defects on the surface of the current gearbox housing, wherein the feedforward neural network model includes an input layer, a hidden layer, and an output layer.

6. A metal defect visual inspection system for a gearbox housing, characterized in that: The system comprises: A preprocessing module, configured to acquire a planar image of the gearbox housing through an image acquisition device, perform grayscale and normalization processing on the image, and generate a preprocessed image; A convolution module, which is used to input the preprocessed images into N convolutional neural network defect detection models respectively, perform calculations in parallel, and generate N feature vectors; An attention module, which is used to sequentially combine N feature vectors into a sequence, perform a self-attention mechanism calculation without positional encoding, and generate an attention sequence; A confidence module is used to input the first representation vector in the attention sequence into the feedforward neural network model to generate a confidence that there is a metal defect on the surface of the current gearbox housing.

7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of a method for visually inspecting metal defects of a gearbox housing according to any one of claims 1 to 5 are implemented.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of a method for visually inspecting metal defects of a gearbox housing according to any one of claims 1 to 5 are implemented.

9. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the steps of a method for visually inspecting metal defects of a gearbox housing according to any one of claims 1 to 5 are implemented.