A method and system for image recognition of semi-trailer frame weld defects
Patent Information
- Application Number
- CN202610989345.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-03
- Publication Date
- 2026-09-25
AI Technical Summary
[0004]为解决模型易受环境扰动且对低频缺陷学习不足的技术问题,本发明提出了一种半挂车车架焊缝缺陷的图像识别方法及系统,能够增强抗扰动能力并提升稀有缺陷识别精度
[0022]本发明通过融合神经网络预设层通道激活统计特征与梯度范数构建激活状态向量,利用主成分分析提取低方差子空间并基于投影增量生成扰动增广图像,增强模型对弱响应特征方向的训练覆盖,同时将激活状态向量表征为激活哈希编码并根据出现频率分配稀有度权重,联合优化加权分类识别损失与激活模式正交化损失,从而提升半挂车车架焊缝缺陷识别的准确性和抗扰动能力。
Smart Images

Figure CN122821272A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image recognition technology. More specifically, this invention relates to an image recognition method and system for weld defects in semi-trailer frames. Background Technology
[0002] The chassis is the main load-bearing structure of a semi-trailer, and its welding quality directly affects the overall structural strength and operational safety of the vehicle. Due to factors such as fluctuations in production processes, changes in the operating environment, and deviations in welding parameters, defects such as porosity, cracks, slag inclusions, and incomplete penetration are prone to occur at the chassis welds. These defects may gradually expand during long-term heavy-load operation, eventually leading to structural failure or even fracture. Current non-destructive testing processes for welds still rely to some extent on manual visual judgment and experience-based identification. This is not only inefficient and labor-intensive, but also easily affected by the subjective experience, fatigue level, and environmental conditions of the inspectors, making it difficult to meet the testing requirements of modern large-scale, high-paced industrial production.
[0003] Currently, with the development of deep learning and computer vision technologies, using neural networks for automatic identification of weld defect images has become an important direction for improving the efficiency and consistency of weld inspection. However, in the actual production inspection site of semi-trailer frame welds, factors such as changes in lighting, dust splashes, oil stains, and sensor noise can introduce unpredictable micro-perturbations into the original defect images. Existing neural network models are still somewhat vulnerable to such perturbation inputs, easily leading to fluctuations in prediction results, resulting in missed or false detections. Furthermore, the actual weld defect image data collected typically exhibits a long-tail distribution, with a large number of common defect samples and a limited number of high-risk but low-frequency defect samples. Conventional training mechanisms tend to bias the model towards high-frequency samples, resulting in insufficient learning of rare defect samples. This leads to insufficient feature responses related to rare defects in deep networks, affecting the model's ability to express and identify rare defects. Therefore, how to improve the model's resistance to environmental perturbations while enhancing its learning ability for low-frequency defect samples and weak response feature patterns, and balancing the training contributions between high-frequency and low-frequency samples, has become a pressing technical problem to be solved in the field of intelligent inspection of semi-trailer frame welds. Summary of the Invention
[0004] To address the technical problems of models being susceptible to environmental disturbances and insufficient learning of low-frequency defects, this invention proposes an image recognition method and system for weld defects in semi-trailer frames, which can enhance the anti-disturbance capability and improve the recognition accuracy of rare defects.
[0005] In a first aspect, the present invention provides an image recognition method for weld defects in semi-trailer frames, comprising: S1, acquiring a dataset of semi-trailer weld defect images and initializing a neural network model; in training iterations, calculating an initial loss for the current batch of original images, backpropagating to obtain the gradient tensor of the initial loss with respect to a preset layer feature map, and concatenating the statistical features of the activation values of each channel of the preset layer with the L2 norm of the corresponding channel's gradient tensor to form an activation state vector corresponding to the image; S2, performing principal component analysis on the set of activation state vectors of the current batch, extracting principal component vectors of a preset proportion after ranking the explained variance to form a low-variance subspace; and maximizing the activation state vectors... The generator function of the projection increment onto the low-variance subspace is used to calculate the adversarial perturbation and applied to the input image to generate a perturbation augmented image. The activation state vector of the original image is projected onto the periodically updated global principal component basis to generate an activation hash code and store it in the signature library. Images with frequencies below a preset threshold are assigned sparse weights calculated based on the reciprocal of the frequency. In step S3, the classification loss of the perturbation augmented image is calculated. The activation mode orthogonalization loss is calculated by penalizing the projection components of the activation state vector onto the global low-variance principal components. The classification loss weighted by the sparse weights is added to the orthogonalization loss to obtain the total loss. The model weights are updated until the preset termination condition is met.
[0006] By adopting the above technical solution, the activation state vector is constructed by integrating the pre-set layer channel activation statistical features and gradient norm. The perturbation augmented image is generated by using low variance subspace projection increment to enhance the training coverage of the model in the direction of weak response features. Combined with the sparse weight allocation mechanism based on activation hash coding frequency and the activation mode orthogonalization loss constraint, the model's ability to resist environmental disturbances is improved while enhancing its learning ability for low-frequency defect samples and weak response feature patterns. This effectively reduces the false negative rate and false positive rate of semi-trailer frame weld defect identification.
[0007] Preferably, the step of concatenating the statistical features of the activation values of each channel in the preset layer with the L2 norm of the gradient tensor of the corresponding channel to form the activation state vector corresponding to the image includes: for each image in the current batch, extracting the feature map output of each channel in the preset layer, and calculating the spatial global average pooling value of the feature map as the statistical feature of the activation value; simultaneously, extracting the backpropagation gradient tensor corresponding to the channel in the feature map of the preset layer from the initial loss, and calculating the square root of the sum of squares of the gradient tensor in all dimensions as the L2 norm of the gradient tensor; concatenating the global average pooling value of all channels in the preset layer with the L2 norm of the gradient tensor to construct the activation state vector corresponding to the image.
[0008] By adopting the above technical solution, the spatial global average pooling value is used as the activation statistical feature and concatenated with the L2 norm of the gradient tensor in the channel dimension. This can simultaneously characterize the activation intensity of neurons in each channel and their sensitivity to the loss function. This allows the activation state vector to more comprehensively reflect the model's response state and vulnerability characteristics under the current sample, providing a more discriminative basic representation for subsequent perturbation generation and sparseness evaluation.
[0009] Preferably, the step of performing principal component analysis on the set of activation state vectors in the current batch, extracting principal component vectors of a predetermined proportion after ranking the explained variance, and spanning a low-variance subspace includes: arranging the activation state vectors corresponding to all images in the current batch by row to construct the state matrix of the current batch; calculating the feature covariance matrix after performing mean-neutralization on the state matrix; performing eigenvalue decomposition on the covariance matrix to obtain each eigenvalue and its corresponding orthogonal eigenvector; sorting the eigenvectors in descending order of their corresponding eigenvalues, extracting the eigenvectors of a predetermined proportion that are ranked last, and using the extracted eigenvectors to form a basis matrix, spanning a low-variance subspace.
[0010] By adopting the above technical solution, principal component analysis is performed on the current batch of activated state vectors, and the principal component vectors with lower explanatory variance ranking are extracted to form a low variance subspace. This can accurately locate the feature directions in which the model has a weak response and insufficient coverage on the current batch of data, thereby providing a targeted projection space for the generation of subsequent adversarial perturbations. This allows the augmented samples to effectively expose the model's weak links in the recognition of boundary transition classes and non-mainstream visual features.
[0011] Preferably, the step of calculating the adversarial perturbation by maximizing the generating function of the projection increment of the activation state vector onto the low variance subspace includes: assuming the basis matrix of the low variance subspace is P, the original image is x, the adversarial perturbation is δ, the activation state vector corresponding to the original image is the first vector, and the activation state vector corresponding to the image after the perturbation is the second vector; constructing a target generating function, the value of which is the squared L2 norm of the product of the transpose of P and the second vector, minus the squared L2 norm of the product of the transpose of P and the first vector; using a projection gradient ascent algorithm based on zero-order optimization, with the goal of maximizing the target generating function, iteratively solving for the optimal adversarial perturbation within a preset perturbation norm constraint range.
[0012] By adopting the above technical solution, a generating function is constructed with the projection increment of the activation state vector onto the low variance subspace as the optimization objective. The zero-order optimized projection gradient ascent algorithm is used to iteratively solve the adversarial perturbation within the perturbation norm constraint. This can efficiently generate targeted augmented samples without relying on the calculation of higher-order derivatives, induce the model features to shift towards the low variance direction, thereby improving the learning strength and stability of low-response defect features.
[0013] Preferably, the step of applying the perturbation to the input image to generate the perturbation augmented image includes: superimposing the acquired adversarial perturbation onto the original image and limiting it to the effective pixel value range through a pixel value cropping operation to generate the perturbation augmented image.
[0014] By adopting the above technical solution, the adversarial perturbation obtained by solving is superimposed on the original image and the perturbation augmented image is generated by limiting the pixel value to the effective pixel value range through pixel value cropping. While maintaining the visual consistency of the image, adversarial changes targeting the model's vulnerable subspace can be introduced at the feature level, providing high-quality data augmentation samples for subsequent training and avoiding information distortion caused by pixel overflow.
[0015] Preferably, the step of calculating the activation mode orthogonalization loss by penalizing the projection components of the activation state vector onto the global low-variance principal components includes: obtaining an orthogonal basis matrix composed of global low-variance principal component vectors calculated from historical batch activation state vectors; for the activation state vector induced by the current perturbation augmented image at a preset layer, obtaining the gradient L2 norm component based on the classification loss of the perturbation augmented image, and performing gradient stopping processing on the gradient L2 norm component, calculating the projection sum of squares of the vector onto the orthogonal basis matrix, the projection sum of squares being the sum of the squares of the inner products of the activation state vector and each column vector of the orthogonal basis matrix, multiplied by the orthogonalization regularization coefficient; and using the projection sum of squares as the activation mode orthogonalization loss.
[0016] Preferably, the step of obtaining the orthogonal basis matrix composed of global low-variance principal component vectors calculated from historical batch activation state vectors includes: maintaining a global covariance matrix calculated from historical batch activation state vectors, periodically performing eigenvalue decomposition on the global covariance matrix, obtaining the global low-variance principal component vectors corresponding to the last K smallest eigenvalues, and forming an orthogonal basis matrix.
[0017] Preferably, the step of generating an activation hash code and storing it in the signature library includes: extracting the first M principal component vectors of the global historical activation state vector to form the principal component basis that is fixed for use in the current period; performing an inner product operation on the activation state vector of each original image in the current batch with the fixed principal component basis to obtain an M-dimensional projection coordinate vector; for each dimension of the projection coordinate vector, if the value is greater than a preset zero threshold, it is represented as 1, otherwise it is represented as 0; converting the M-dimensional projection coordinate vector into an M-bit binary sequence composed of 0 and 1; using the binary sequence as the activation hash code of the image, registering it in the global activation signature library, and updating the cumulative occurrence count of the hash code in the library.
[0018] Preferably, assigning rarity weights based on the reciprocal of frequency to images with frequencies below a preset threshold includes: counting the cumulative occurrences of the activation hash code of the current image in the global activation signature database, dividing by the total number of records in the signature database to calculate the occurrence frequency of the hash code; determining whether the occurrence frequency is below a preset rarity threshold, and if so, using the reciprocal of the occurrence frequency as a base amplification factor; multiplying the base amplification factor by a preset smoothing coefficient and adding a constant 1, and then truncating it by a preset upper limit to obtain the rarity weights corresponding to the original image.
[0019] Secondly, the present invention provides an image recognition system for weld defects in a semi-trailer frame, comprising a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the above-mentioned image recognition method for weld defects in a semi-trailer frame is implemented.
[0020] By adopting the above technical solution, the image recognition method for weld defects in a semi-trailer frame is generated into a computer program and stored in a memory for loading and execution by a processor. This allows for the creation of a terminal device based on the memory and processor, facilitating its use.
[0021] The present invention has the following beneficial effects:
[0022] This invention constructs an activation state vector by fusing the activation statistical features of a preset layer channel in a neural network with the gradient norm. It then uses principal component analysis to extract a low-variance subspace and generates a perturbation augmented image based on the projection increment, thereby enhancing the model's training coverage of weak response feature directions. Simultaneously, it represents the activation state vector as an activation hash code and assigns sparse weights according to the frequency of occurrence. It jointly optimizes the weighted classification recognition loss and the activation pattern orthogonalization loss, thereby improving the accuracy and anti-perturbation capability of semi-trailer frame weld defect identification.
[0023] Furthermore, by assigning sparse weights to low-frequency activation mode samples and introducing activation mode orthogonalization loss constraints, this invention reduces the model's excessive dependence on vulnerable feature directions, effectively enhances the learning ability of rare defect samples in long-tail distribution, and improves the reliability of rare defect detection in semi-trailer frame weld defect detection. Attached Figure Description
[0024] Figure 1 This is a flowchart of an image recognition method for weld defects in a semi-trailer frame. Figure 2 This is a diagram of the optimization process to counteract disturbances; Figure 3 This is a diagram of the optimization process to counteract disturbances. Detailed Implementation
[0025] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments.
[0026] This invention discloses an image recognition method for weld defects in semi-trailer frames, referring to... Figure 1 This includes steps S1-S3: S1. Obtain the image dataset of weld defects in semi-trailers and initialize the neural network model. During the training iteration, calculate the initial loss for the original images of the current batch, backpropagate to obtain the gradient tensor of the initial loss with respect to the feature map of the preset layer, and concatenate the statistical features of the activation values of each channel of the preset layer with the L2 norm of the gradient tensor of the corresponding channel to form the activation state vector corresponding to the image.
[0027] Visual images of the weld seams on the semi-trailer frame are acquired using an industrial camera. The image pixel matrix is read, and all images are uniformly scaled to a preset fixed resolution size, such as 224×224. Tensor transformation and pixel value normalization operations are performed to construct a tensor-formatted dataset.
[0028] The ResNet network structure is instantiated as a neural network model, and the parameters of the convolutional and fully connected layers are randomly initialized using the Kaiming normal distribution method.
[0029] After entering the training iteration loop, original images of the current batch are randomly selected from the dataset according to a set batch size, such as 32, and fed into the model for forward propagation. The predicted probability distribution of weld defect categories is output. The cross-entropy loss function is then called to calculate the initial loss between the predicted probability and the true label.
[0030] Perform backpropagation calculations by using the register_hook mechanism to attach hook functions during model forward propagation, intercept and save the gradient tensor of the preset layer feature map, and set the size of the feature map tensor to "batch number × channel number × height × width".
[0031] For the feature map output by the preset layer, the mean of the activation value of each channel is calculated in the height and width spatial dimensions as a statistical feature. At the same time, the L2 norm of the captured gradient tensor of the corresponding channel is calculated in the above spatial dimensions to obtain the gradient norm of each channel.
[0032] In terms of channel features, the mean of the activation values and the gradient norm are concatenated to generate a vector with twice the number of channels. This vector is then flattened and used as the activation state vector for each original image in the current batch.
[0033] In an optional embodiment, the statistical features of the activation values of each channel in the preset layer are concatenated with the L2 norm of the gradient tensor of the corresponding channel to form the activation state vector corresponding to the image. This includes: for each image in the current batch, extracting the feature map output of each channel in the preset layer, and calculating the spatial global average pooling value of the feature map as the statistical feature of the activation value; simultaneously, extracting the backpropagation gradient tensor of the channel corresponding to the initial loss in the feature map of the preset layer, and calculating the square root of the sum of squares of the gradient tensor in all dimensions as the L2 norm of the gradient tensor; concatenating the global average pooling value of all channels in the preset layer with the L2 norm of the gradient tensor to construct the activation state vector corresponding to the image.
[0034] The neural network model used in the preset layer is an image classification model based on a residual network architecture. This model's network structure sequentially includes a 7×7 initial convolutional layer, a max-pooling layer, four cascaded residual convolutional modules, a global average pooling layer, and a fully connected layer. The input to this neural network model is a three-channel, 224×224 original image of a semi-trailer weld. The output of this neural network model is a classification probability distribution vector indicating various weld defect categories.
[0035] During the training of the aforementioned neural network model, the output of the last residual block of the fourth residual convolutional module group is selected as the preset layer. For an image input with a batch size of B=32, the dimension of the output feature map of this preset layer is usually [B,C,H,W], for example [32,2048,7,7], where C represents the number of channels 2048, and H and W are the spatial height and width, respectively.
[0036] For the c-th channel of a single image, the feature map output is a 7×7 two-dimensional matrix. A spatial global average pooling operation is then performed on it, which involves summing the activation values of these 49 pixels and dividing by 49 to obtain a single scalar value. , representing the statistical characteristics of the channel.
[0037] Calculate the initial cross-entropy classification loss using the backpropagation mechanism. gradient tensor of the feature map of this channel This tensor is also a 7×7 matrix. The square root of the sum of the squares of these 49 gradient values is then taken to obtain the L2 norm of the gradient. The pooling scalar value sequence of all 2048 channels obtained in the above steps. With the corresponding gradient L2 norm sequence The vectors are concatenated along the dimensional axis to generate a one-dimensional vector of length 4096, which is the activation state vector S corresponding to the image. For example, if the pooling value of the 10th channel of a specific image is 0.45 and the gradient L2 norm is 0.015, then the 10th element of its activation state vector is 0.45 and the 2058th element is 0.015.
[0038] The current batch will form a state matrix of size [32, 4096]. This state representation, which integrates the activation output with its sensitivity, i.e., gradient, can represent the neuron activation patterns and vulnerability characteristics of the network for defects to be detected, such as porosity and cracks in semi-trailer welds.
[0039] S2: Perform principal component analysis on the set of active state vectors in the current batch, extract the principal component vectors after explanatory variance ranking and span a low-variance subspace; calculate the adversarial perturbation by maximizing the generation function of the incremental projection of the active state vectors onto the low-variance subspace, and apply it to the input image to generate a perturbation augmented image; project the active state vectors of the original image onto the periodically updated global principal component basis, generate an activation hash code and store it in the signature library, and assign sparse weights based on the reciprocal of frequency to images with frequencies below a preset threshold.
[0040] Collect the activation state vectors of all images in the current batch to form a feature matrix. Then, use the PCA class from a machine learning library, or utilize singular value decomposition (SVD) to perform principal component analysis on this feature matrix, extracting the eigenvalues and corresponding eigenvectors of all principal components.
[0041] The principal components are sorted in descending order of their eigenvalues, and the set of eigenvectors corresponding to the lowest-ranked proportions is selected. These vectors are then linearly combined in the feature space to form a low-variance subspace.
[0042] In the adversarial perturbation computation stage, a target generating function is defined. This function is composed of the norm of the projection vector of the new activation state vector of the perturbated image onto the aforementioned low-variance subspace, and is measured by multiplying the projection matrix by the activation vector and calculating the squared L2 norm. The adversarial perturbation is initialized as a small noise matrix conforming to a preset distribution, and is superimposed on the input image to calculate the output of the generating function. Using a projection gradient ascent algorithm based on zero-order optimization, the approximate gradient of the target generating function with respect to the adversarial perturbation is estimated through finite difference. The perturbation matrix is updated according to a preset step size and truncated within the infinite norm constraint boundary. After multiple iterations to maximize this projection increment, the adversarial perturbation is obtained. This adversarial perturbation is element-wise summed and applied to the pixel tensor of the original input image to generate the perturbation augmented image.
[0043] A global principal component basis matrix is established in external memory, and the principal component basis is updated every fixed period, such as 5 batches, using an exponential moving average of historical sliding window data. During each periodic update of the global principal component basis, the inner product of each principal component vector obtained in the current period and the corresponding principal component vector in the previous period is calculated. If the inner product is less than 0, the current principal component vector is multiplied by -1 to perform a sign flip, aligning its direction with the corresponding principal component vector in the previous period; if the inner product is not less than 0, the current principal component vector remains unchanged. This sign alignment process reduces the impact of principal component direction uncertainty on the stability of activation hash encoding. The activation state vector of the current original image is multiplied by the global principal component basis matrix to achieve projection, obtaining the dimensionality-reduced projection vector.
[0044] A hash encoding method based on principal component projection coordinate symbol representation is adopted. The coordinates of each dimension in the projection vector are compared with a preset zero threshold. If the coordinates are greater than the preset zero threshold, they are represented as 1; otherwise, they are represented as 0. A binary sequence of 0 and 1 is obtained, which generates the activation hash code. This code is then stored as a key in the signature library of the Python dictionary structure to count the frequency.
[0045] During the query phase, the cumulative occurrence count of the current image's hash code is retrieved from the signature database, and the occurrence frequency of the hash code is calculated by combining this with the total number of records in the signature database. A preset rarity threshold is set, for example, 0.01. For images with an occurrence frequency lower than the preset rarity threshold, a rarity weight proportional to the reciprocal of the occurrence frequency is calculated and assigned. For images with an occurrence frequency not lower than the preset rarity threshold, a base weight of 1 is assigned.
[0046] In an optional embodiment, principal component analysis is performed on the set of activation state vectors in the current batch to extract principal component vectors of a predetermined proportion after ranking the explained variance, spanning a low-variance subspace. This includes: arranging the activation state vectors corresponding to all images in the current batch by row to construct the state matrix of the current batch; performing mean-neutralization on the state matrix and calculating the feature covariance matrix; performing eigenvalue decomposition on the covariance matrix to obtain each eigenvalue and its corresponding orthogonal eigenvector; sorting the eigenvectors in descending order of their corresponding eigenvalues, extracting the eigenvectors of a predetermined proportion that are ranked last, and using the extracted eigenvectors to form a basis matrix, spanning a low-variance subspace.
[0047] Suppose the current training batch contains 32 images of semi-trailer weld seams. The corresponding 4096-dimensional activation state vectors are stacked row-wise to obtain a state matrix M of size [32, 4096]. Mean-neutralization is performed on the state matrix M along the column direction, i.e., the mean of each column is calculated. This mean represents the batch average level of the neuron's feature or gradient. Subtracting this mean from the corresponding column elements in matrix M yields the centered matrix. For example, if the batch average in column 1 is 0.21, then subtract 0.21 from all elements in column 1. Using the formula... Calculate the eigencovariance matrix C. Since the 4096×4096 matrix is quite large, in practical calculations it can be done through singular value decomposition. By decomposing the data, equivalent eigenvalues and orthogonal eigenvectors can be obtained, thereby improving computational efficiency and reducing memory overflow issues.
[0048] After eigenvalue decomposition, extract the eigenvalues and sort them in descending order of value. In practical calculations, singular value decomposition can be used to truncate the centered matrix, obtaining no more than B-1 effective principal components; let the number of effective principal components be R, and R≤B-1, then the corresponding eigenvalues are sorted in descending order of value. Arrangement. Within this sequence, the lower the eigenvalue, the smaller the variance of the data distribution in the corresponding eigenvector dimension. This type of dimension carries fragile feature patterns that are under-covered by neurons and susceptible to adversarial perturbations.
[0049] Low-variance principal components are selected according to a preset ratio; for example, when the preset ratio is 50%, the top half of the principal components are discarded, and the bottom half are selected. R / 2 The sequence of eigenvectors corresponding to each eigenvalue. When the current batch size B is 32, the maximum number of effective principal components R of the centered matrix is 31; for example, when R=31, the sequence of eigenvectors corresponding to the last 16 eigenvalues can be extracted. The extracted column vectors are merged along the column direction to construct an orthogonal basis matrix P. The low-variance subspace spanned by this matrix is used to characterize the relatively inactive neuron response dimension in the deep feature manifold of the current batch of weld defect images. When generating adversarial perturbations, it can induce image features to project onto this low-variance space.
[0050] In an optional embodiment, the adversarial perturbation is calculated by maximizing the generating function of the projection increment of the activation state vector onto the low-variance subspace, including: assuming the basis matrix of the low-variance subspace is P, the original image is x, the adversarial perturbation is δ, the activation state vector corresponding to the original image is the first vector, and the activation state vector corresponding to the image after the perturbation is the second vector; constructing a target generating function, the value of which is the squared L2 norm of the product of the transpose of P and the second vector, minus the squared L2 norm of the product of the transpose of P and the first vector; using a projection gradient ascent algorithm based on zero-order optimization, with the goal of maximizing the target generating function, iteratively solving for the optimal adversarial perturbation within a preset perturbation norm constraint.
[0051] For a given original image of a semi-trailer weld. The image after superimposing the anti-perturbation δ is set as Let the 4096-dimensional activation state vector extracted from the original image in the model be the first vector S(x), and the activation state vector corresponding to the perturbed image be the second vector S(x+δ). The constructed low-variance subspace basis matrix is denoted as... The square of the projection of the first vector onto the low-variance subspace is... , representing the energy distribution of the original activation features in the low variance space. Similarly, the square of the projection of the second vector is... .
[0052] Based on this, a target generation function is constructed. Since calculating the L2 norm of the feature map and the second derivative of the corresponding backpropagation gradient is computationally expensive, a projection gradient ascent algorithm based on zero-order optimization is chosen to avoid this high-order derivative process. This algorithm estimates the approximate gradient with respect to the perturbation through multiple objective function value evaluations; each objective function value evaluation includes forward propagation and a backpropagation for extracting the gradient tensor of the preset layer, but does not require further calculation of the higher-order derivative of the gradient tensor's norm.
[0053] In the optimization process, the goal is to maximize the generating function F(δ), and the range of the infinite norm perturbation constraint is set as follows. ,in The preferred value is typically 0.0314, which refers to pixel values normalized to the 0-1 range.
[0054] In the single-step optimization, a unit probe vector u following a Gaussian distribution is randomly generated, and the probe step size μ is preferably 0.001. A zero-order stochastic gradient estimation method based on bilateral finite difference is adopted to estimate the approximate gradient with respect to δ based on the difference in function values between F(δ+μu) and F(δ-μu). Using the sign direction of this approximate gradient, δ is iteratively updated with a step size of α = 2 / 255, i.e. , where Π is the projection clipping operator. The trend of the target generating function value during the iterative optimization process is as follows: Figure 2 As shown.
[0055] After 10 iterations, the optimal adversarial perturbation is found, which can induce a more obvious projection change of the model's activation state vector into the low variance subspace within the preset perturbation constraint range. This helps to expose the model's weak areas in recognizing low-response, non-mainstream, or boundary transition visual features. Visual features in semi-trailer weld seam images can be represented by local features such as defect edges, blurred weld seam areas, or fine crack textures, and provide targeted augmentation samples for subsequent robust training.
[0056] In an optional embodiment, applying an adversarial perturbation to the input image to generate a perturbation augmented image includes: superimposing the acquired adversarial perturbation onto the original image and limiting it to the effective pixel value range through a pixel value cropping operation to generate a perturbation augmented image.
[0057] The optimal perturbation is added element-wise to the pixel matrix of the original weld defect image. To ensure the pixel values of the synthesized image still conform to normal digital image representation rules, a global pixel value domain cropping operation is then performed, resulting in the superimposed image. The constraint is within the normalized effective value range of 0 to 1. For example, if the sum of a pixel value is 1.02, then it is set to 1.
[0058] This generates a perturbation augmented image for fine-tuning model stability. The augmented sample is visually almost identical to the original image, but contains adversarial features targeting the model's vulnerable subspaces at the feature level.
[0059] In an optional embodiment, the activation state vector of the original image is projected onto a periodically updated global principal component basis to generate an activation hash code and store it in a signature library. This includes: extracting the first M principal component vectors of the global historical activation state vector to form the principal component basis that is fixed for use in the current period; performing an inner product operation on the activation state vector of each original image in the current batch with the fixed principal component basis to obtain an M-dimensional projection coordinate vector; for each dimension of the projection coordinate vector, if the value is greater than a preset zero threshold, it is represented as 1, otherwise it is represented as 0; converting the M-dimensional projection coordinate vector into an M-bit binary sequence composed of 0 and 1; using the binary sequence as the activation hash code of the image, registering it in the global activation signature library, and updating the cumulative occurrence count of the hash code in the library.
[0060] For all orthogonal eigenvectors obtained by periodically decomposing the global historical feature covariance matrix, select the top M principal component vectors that represent the main feature distributions and have the highest variance explanation rate. For example, take M=64, and construct a fixed principal component basis matrix of size 4096×64. For a single original image of a semi-trailer weld in the current batch of input, its 4096-dimensional activation state vector is... Performing matrix multiplication (inner product) with this fixed principal component basis yields a projected coordinate vector of length 64. Each component of this coordinate vector represents the response intensity of the original image in the corresponding mainstream feature space direction.
[0061] Discretize the projected coordinate vector composed of continuous numerical values. Set a specific zero threshold τ, which is usually set as the mean or median of all historical projected values, for example, it can be set to 0.
[0062] Iterate through each element of the 64-dimensional coordinate vector. Where j ranges from 1 to 64, the representation determination is performed using the step function: if If >0, then the dimension is represented as 1. If ≤0, the output is 0. Through this fast binarization operation, the original high-dimensional floating-point state vector is compressed into a compact 64-bit binary sequence of 0s and 1s, such as a sequence of 10110.
[0063] This highly spatially specific binary sequence is set as the exclusive activation hash code for this image, and it is inserted as a unique index key into the global activation signature library maintained in the memory dictionary structure. If the code already exists, its cumulative occurrence counter is incremented by 1.
[0064] In an optional embodiment, images with frequencies below a preset threshold are assigned rarity weights calculated based on the reciprocal of the frequency. This includes: counting the cumulative occurrences of the activation hash code of the current image in the global activation signature database, dividing by the total number of records in the signature database, and calculating the occurrence frequency of the hash code; determining whether the occurrence frequency is below a preset rarity threshold, and if so, using the reciprocal of the occurrence frequency as a base amplification factor; multiplying the base amplification factor by a preset smoothing coefficient and adding a constant 1, and then truncating it by a preset upper limit to obtain the rarity weights corresponding to the original image.
[0065] After obtaining the 64-bit activation hash code generated from the current weld seam's original image, a hash table lookup algorithm is used to retrieve the cumulative occurrence count of this code from the global activation signature database. Simultaneously, the total number of records for all image samples processed in the signature database at this point is obtained. Through formula Calculate the current global frequency f of the activation pattern represented by the hash code. For example, if a hash code generated by a specific porosity defect image appears only 15 times, and the total number of samples is 10,000, then the frequency is 0.0015.
[0066] A rarity threshold is preset. The preferred value is 0.01, or 1%. Determine if the frequency f is lower than 0.01. If the condition is met, it indicates that the neuronal response pattern induced by the image is rare in global training. Then, calculate the reciprocal of its occurrence frequency, 1 / f, as the basic amplification factor. For the above example, this amplification factor is approximately 666.7.
[0067] Since using the reciprocal of the frequency might cause a sudden change in the loss function, leading to gradient explosion, a preset smoothing coefficient γ, preferably 0.005, is used for damped scaling. The intermediate weights are calculated according to the formula W = 1 + γ·(1 / f), and a maximum upper limit threshold is applied. The value is typically set to 5 for truncation. If the frequency exceeds the preset rarity threshold, the default setting is used. =1.
[0068] S3, calculate the classification loss of the perturbation augmented image, calculate the activation mode orthogonalization loss by penalizing the projection components of the activation state vector on the global low variance principal components; add the classification loss after sparse weighting to the orthogonalization loss to obtain the total loss, update the model weights, until the preset termination condition is met.
[0069] The perturbation-enhanced image generated in the previous step is input back into the neural network model, and forward propagation is performed. The output is a tensor of the defect classification prediction probability under adversarial example conditions. The cross-entropy loss between this predicted probability and the original true label distribution is calculated and used as the classification loss.
[0070] Based on the classification loss of the perturbation-augmented image, the activation state vector corresponding to the perturbation-augmented image is calculated. Specifically, after inputting the perturbation-augmented image into the neural network model, the classification loss is calculated using its prediction results and the corresponding ground truth labels of the original image. The gradient tensor of this classification loss with respect to the feature map of a preset layer is extracted, the L2 norm of the gradient of the corresponding channel is calculated, and it is concatenated with the activation statistical features of the preset layer channels to obtain the activation state vector corresponding to the perturbation-augmented image. The global low-variance principal component vectors with the lowest explained variance in the global principal component basis are obtained to construct the projection basis matrix. The inner product of the activation state vector of the augmented image on the projection basis matrix is calculated to obtain the projection components. The sum of squares of all projection components is calculated and multiplied by a preset orthogonalization regularization coefficient that controls the strength, as the activation mode orthogonalization loss. This constrains the model to learn more stable feature representations, thereby reducing over-response in vulnerable directions.
[0071] When calculating the orthogonalization loss of the activation mode, the gradient L2 norm component in the activation state vector is used as a state statistic in the projection calculation, and the gradient L2 norm component is subjected to a stopping gradient process; the activation statistical feature component in the activation state vector retains the computation graph used for model parameter updates, so that the back propagation of the orthogonalization loss mainly acts on the model parameters through the activation statistical feature component, thereby avoiding the need to continue to calculate higher-order derivatives of the gradient L2 norm.
[0072] Using basic arithmetic multiplication, the previously calculated image-specific rarity weights are multiplied by the classification loss to perform a weighted sum. The weighted classification loss is then added to the activation pattern orthogonalization loss, and the sum is used to obtain the scalar total loss.
[0073] Based on the total loss, the backpropagation function is called to calculate the gradients of all learnable parameters of the model. By calling the pre-instantiated optimizer object, combined with the set learning rate and momentum parameters, the model weights are updated along the negative direction of the gradient, and the gradient zeroing function is called to clear the historical gradients accumulated in the previous iteration, while updating the current training iteration number.
[0074] The entire process from "acquiring batch images" to "updating model weights" is executed iteratively. After each iteration, the classification accuracy of the model on the validation set and the number of iterations are evaluated. When the validation set loss no longer decreases for a set number of consecutive iterations, or when the set maximum number of training iterations is reached, the preset termination condition is met. At this point, the training process ends, and the converged neural network model weight file is saved.
[0075] In an optional embodiment, the activation mode orthogonalization loss is calculated by penalizing the projection components of the activation state vector onto the global low-variance principal components. This includes: maintaining a global covariance matrix calculated from historical batches of activation state vectors; periodically performing eigenvalue decomposition to obtain the global low-variance principal component vectors corresponding to the K smallest eigenvalues, forming an orthogonal basis matrix; for the activation state vector induced by the current perturbation augmented image at a preset layer, obtaining the gradient L2 norm component based on the classification loss of the perturbation augmented image, and performing gradient stopping processing on the gradient L2 norm component, calculating the projection sum of squares of the vector onto the orthogonal basis matrix, where the projection sum of squares is the sum of the squares of the inner products of the activation state vector and each column vector of the orthogonal basis matrix, multiplied by the orthogonalization regularization coefficient; and using the projection sum of squares as the activation mode orthogonalization loss.
[0076] Throughout the model training process, a global feature covariance matrix of size 4096×4096 is maintained using a moving average algorithm or a queue caching mechanism. The updated formula is as follows The momentum coefficient β is preferably set to 0.05. Eigenvalue decomposition is performed on the global covariance matrix every 50 training batches to obtain all eigenvalues, which are then sorted from largest to smallest. The eigenvectors corresponding to the K=32 smallest eigenvalues are extracted.
[0077] By concatenating the columns of 32 mutually orthogonal 4096-dimensional global low-variance principal component vectors, a global orthogonal basis matrix of size 4096×32 is constructed. It encompasses the fragile feature space orientations that the model rarely activates or produces stable gradient responses throughout the entire training history.
[0078] During the forward propagation of the currently generated perturbation augmented image, the 4096-dimensional activation state vector induced in the preset layer is extracted. Calculate the sum of squares of the projections of the state vector onto the global orthogonal basis matrix, i.e., calculate the sum of squares of the projections of the state vector onto the global orthogonal basis matrix. and Each column vector Take the inner product and square it, where i belongs to 1 to 32. Summate these 32 squared values to get... .
[0079] Multiply the sum of squares of the projection by a preset orthogonalization regularization coefficient. This value is typically set to between 0.01 and 0.05, for example, 0.02, to calculate the activation mode orthogonalization loss. The setting of this loss term can constrain the tendency of augmented samples to shift to the global vulnerable space when updating weights during backpropagation, making the extracted features more stable and reducing the dependence on low-variance vulnerable directions. This helps to improve the generalization and anti-interference ability of the detection results of low-contrast defects, fine crack textures or boundary transition regions on weld surfaces.
[0080] The ablation experiments were conducted using a dataset of images of weld defects in semi-trailers for testing and verification. The network employed an image classification model based on a residual network architecture. Four sets of comparative conditions were set up in the experiments: the first set was the baseline model trained using only standard cross-entropy loss; the second set incorporated a scheme based on low-variance subspace adversarial perturbation augmentation into the baseline model; the third set added an activation mode orthogonalization loss constraint; and the fourth set integrated the complete inventive scheme based on an activation hash coding sparse weight allocation mechanism.
[0081] The first baseline model achieved an average classification accuracy of 82.5% on the test set, but its recall rate for rare crack defects was only 65.3%. The second scheme improved the average classification accuracy to 86.2%, and the recall rate for rare crack defects to 73.4%. The third scheme achieved an average classification accuracy of 88.7%, and the recall rate for rare crack defects increased to 78.9%. The fourth complete scheme achieved the highest average classification accuracy on the test set, reaching 92.3%, and its recall rate for rare crack defects rose to 89.6%. The comparison results of the average classification accuracy of each ablation experiment are as follows: Figure 3 As shown.
[0082] Analysis of the experimental results revealed that adversarial perturbation augmentation increases the model's training coverage of low-response and boundary transition visual features, leading to an initial improvement in basic recognition performance. Superimposing activation mode orthogonalization loss constrains the excessive shift of network features towards the global vulnerable space, enhancing the anti-interference ability for detecting small defects. Applying a rarity weight allocation mechanism suppresses overfitting on the majority of normal samples, assigning greater gradient weights to rare pores and long-tailed cracks, thus improving overall classification accuracy and the success rate of rare defect detection.
[0083] This invention also discloses an image recognition system for weld defects in semi-trailer frames, including a processor and a memory. The memory stores computer program instructions, and when the computer program instructions are executed by the processor, an image recognition method for weld defects in semi-trailer frames according to the present invention is implemented.
[0084] The system also includes other components well known to those skilled in the art, such as communication buses and communication interfaces, the settings and functions of which are known in the art and will not be described in detail here.
[0085] In the description of this specification, "multiple" or "several" means at least two, such as two, three or more, unless otherwise expressly and specifically defined.
Claims
1. An image recognition method for weld defects in a semi-trailer frame, characterized in that, include: S1, acquire the image dataset of weld defects in semi-trailers and initialize the neural network model; In the training iteration, the initial loss is calculated for the original images of the current batch. Backpropagation is used to obtain the gradient tensor of the initial loss with respect to the feature map of the preset layer. The statistical features of the activation values of each channel of the preset layer are concatenated with the L2 norm of the gradient tensor of the corresponding channel to form the activation state vector corresponding to the image. In S2, principal component analysis is performed on the set of activation state vectors of the current batch. After extracting the principal component vectors with a preset proportion after explaining the variance ranking, a low-variance subspace is formed. The adversarial perturbation is calculated by maximizing the generation function of the projection increment of the activation state vector onto the low-variance subspace and applied to the input image to generate a perturbation augmented image. The activation state vector of the original image is projected onto the periodically updated global principal component basis to generate an activation hash code and store it in the signature library. A sparse weight calculated based on the reciprocal of the frequency is assigned to images with a frequency lower than a preset threshold. In S3, the classification loss of the perturbation augmented image is calculated. The activation mode orthogonalization loss is calculated by penalizing the projection component of the activation state vector onto the global low-variance principal component. The classification loss after weighting the sparse weight is added to the orthogonalization loss to obtain the total loss. The model weights are updated until the preset termination condition is met.
2. The method according to claim 1, characterized in that, The step of concatenating the statistical features of the activation values of each channel in the preset layer with the L2 norm of the gradient tensor of the corresponding channel to form the activation state vector corresponding to the image includes: For each image in the current batch, extract the feature map output of each channel in the preset layer, and calculate the spatial global average pooling value of the feature map as the statistical feature of the activation value; At the same time, the backpropagation gradient tensor corresponding to the channel in the feature map of the preset layer is extracted from the initial loss, and the square root of the sum of squares of the gradient tensor in all dimensions is calculated as the L2 norm of the gradient tensor. The global average pooling value of all channels in the image within the preset layer is concatenated with the L2 norm of the gradient tensor to construct the activation state vector corresponding to the image.
3. The method according to claim 1, characterized in that, The step of performing principal component analysis on the current batch of active state vector sets, extracting and ranking the principal component vectors at a predetermined ratio to span a low-variance subspace includes: Arrange the activation state vectors corresponding to all images in the current batch by row to construct the state matrix of the current batch; After the state matrix is centered by removing the mean, the feature covariance matrix is calculated. Eigenvalue decomposition is performed on the covariance matrix to obtain each eigenvalue and the corresponding orthogonal eigenvectors; The eigenvectors are sorted in descending order of their corresponding eigenvalues. The eigenvectors at the bottom of the list are truncated at a predetermined ratio. The truncated eigenvectors form a basis matrix and span a low-variance subspace.
4. The method according to claim 2 or 3, characterized in that, The calculation of adversarial perturbations by maximizing the generation function that projects the increment of the activated state vector onto the low-variance subspace includes: Let P be the basis matrix of the low variance subspace, x be the original image, δ be the adversarial perturbation, the activation state vector corresponding to the original image be the first vector, and the activation state vector corresponding to the image after the perturbation is the second vector. Construct a target generating function, the value of which is the squared L2 norm of the product of the transpose of P and the second vector, minus the squared L2 norm of the product of the transpose of P and the first vector. By using the projection gradient ascent algorithm based on zero-order optimization, the optimal adversarial perturbation is iteratively solved within the preset perturbation norm constraint range with the goal of maximizing the target generating function.
5. The method according to claim 1, characterized in that, The process of applying a perturbation to the input image to generate the augmented image includes: The acquired adversarial perturbation is superimposed on the original image, and then restricted to the effective pixel value range through pixel value cropping to generate a perturbation augmented image.
6. The method according to claim 1, characterized in that, The calculation of activation mode orthogonalization loss by penalizing the projection components of the activation state vector onto the global low-variance principal components includes: Obtain the orthogonal basis matrix composed of global low-variance principal component vectors calculated from the historical batch activation state vectors; For the activation state vector induced by the current perturbation augmented image in the preset layer, the gradient L2 norm component is obtained based on the classification loss of the perturbation augmented image. After performing the stopping gradient processing on the gradient L2 norm component, the projection sum of squares of the vector on the orthogonal basis matrix is calculated. The projection sum of squares is the sum of the squares of the inner product of the activation state vector and each column vector of the orthogonal basis matrix, and then multiplied by the orthogonalization regularization coefficient. The sum of squared projections is used as the orthogonalization loss of the activation modes.
7. The method according to claim 6, characterized in that, The step of obtaining the orthogonal basis matrix composed of global low-variance principal component vectors calculated from historical batch activation state vectors includes: Maintain a global covariance matrix calculated from the activation state vectors of historical batches. Periodically perform eigenvalue decomposition on the global covariance matrix to obtain the global low-variance principal component vectors corresponding to the last K smallest eigenvalues, which form an orthogonal basis matrix.
8. The method according to claim 1, characterized in that, The process of generating and storing the activation hash code in the signature database includes: The first M principal component vectors of the global historical activation state vector are extracted to form the principal component basis that is fixed to be used in the current period. The activation state vector of each original image in the current batch is subjected to inner product operation with the fixed principal component basis to obtain an M-dimensional projected coordinate vector. For each dimension of the projected coordinate vector, if the value is greater than a preset zero threshold, it is represented as 1; otherwise, it is represented as 0. The M-dimensional projected coordinate vector is converted into an M-bit binary sequence consisting of 0 and 1. The binary sequence is used as the activation hash code of the image, registered in the global activation signature library, and the cumulative occurrence count of the hash code in the library is updated.
9. The method according to claim 1, characterized in that, The process of assigning rarity weights based on the reciprocal of frequency to images with frequencies below a preset threshold includes: The frequency of the hash code is calculated by counting the cumulative number of occurrences of the activation hash code of the current image in the global activation signature database and dividing it by the total number of records in the signature database. Determine whether the frequency of occurrence is lower than a preset rarity threshold. If it is lower, use the reciprocal of the frequency of occurrence as the basic amplification factor. The sparse weights corresponding to the original image are obtained by multiplying the base magnification factor by a preset smoothing coefficient and adding a constant of 1, and then truncating the image by a preset upper limit.
10. An image recognition system for weld defects in a semi-trailer frame, characterized in that, include: A processor and a memory, the memory storing computer program instructions that, when executed by the processor, implement the image recognition method for weld defects in a semi-trailer frame according to any one of claims 1-9.