Defect identification method and device for aluminum matrix composite agricultural machinery parts

CN122820546APending Publication Date: 2026-09-25CHINA ELECTRONIC INFORMATION IND DEV RES INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610733771.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-26
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0005]本发明的目的在于提供一种铝基复合材料农机零部件缺陷识别方法及装置,旨在解决现有技术中的上述问题

Benefits of technology

[0009]根据本发明实施例的第四方面,提供一种计算机可读存储介质,其上存储有信息传递的实现程序,该程序被处理器执行时实现本公开第一方面所提供的铝基复合材料农机零部件缺陷识别方法的步骤。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122820546A_ABST
    Figure CN122820546A_ABST
Patent Text Reader

Abstract

The application provides an aluminum matrix composite agricultural machinery part defect identification method and device, wherein the method comprises: acquiring SEM microscopic images and EDX spectral images corresponding to aluminum matrix composite agricultural machinery parts respectively; forming an embedded vector sequence based on the SEM microscopic images, forming a spectral sequence based on the EDX spectral images, and constructing the embedded vector sequence and the spectral sequence into a joint sequence; inputting the joint sequence into a visual transformation encoder, interacting through the visual transformation encoder using self-attention and cross-modal attention, using global labels to converge multi-modal information and outputting global feature representation, and outputting a defect identification result according to the global feature representation. The application fully integrates SEM images and EDX spectral multi-modal information, improves strong global feature modeling capability, and realizes accurate, stable and intelligent detection of defects such as cracks, holes, inclusions and delamination of aluminum matrix composite agricultural machinery parts.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent inspection and non-destructive testing technology for agricultural machinery and equipment, and in particular to a method and device for identifying defects in agricultural machinery parts made of aluminum-based composite materials. Background Technology

[0002] Aluminum-based composites, with their outstanding advantages of lightweight, high strength, wear resistance, and corrosion resistance, have been widely used in the manufacturing of agricultural machinery parts. However, during production, processing, assembly, debugging, and long-term field operation, these parts are highly susceptible to microscopic defects such as cracks, holes, inclusions, and delamination. These defects not only weaken the mechanical properties of the parts and shorten their service life, but may also cause malfunctions in agricultural machinery, and in severe cases, even lead to equipment damage and production safety accidents. Therefore, accurate and efficient defect detection of aluminum-based composite agricultural machinery parts is of great practical significance for ensuring the stable operation of agricultural machinery, extending equipment life, and improving the safety of production operations.

[0003] Currently available defect detection methods mainly include manual visual inspection, ultrasonic inspection, X-ray inspection, and automated inspection methods based on single-modal image analysis. However, these methods all have significant shortcomings. Manual inspection is time-consuming and labor-intensive, making it difficult to meet the inspection needs of large batches of parts. For small or dispersed defects, omissions or misjudgments are prone to occur. Single-modal inspection methods typically utilize only one of the morphological or compositional information, failing to simultaneously reflect the microstructural features and chemical composition of the parts, making it difficult to achieve comprehensive and accurate identification of complex defects. For example, scanning electron microscopy (SEM) can clearly present the microscopic morphology of the surface and near-surface layers of parts, but it cannot provide information on the chemical composition of defect areas. Energy dispersive X-ray spectroscopy (EDX) can reveal the elemental composition and distribution patterns of defect areas, but it lacks spatial morphology information. Using SEM or EDX technology alone is insufficient to comprehensively and completely characterize the features of defects such as cracks, pores, inclusions, and delamination. Furthermore, with the rapid development of deep learning technology, methods based on Convolutional Neural Networks (CNNs) have achieved certain results in the field of image recognition. However, traditional CNNs mainly rely on local receptive fields to extract features, which limits their ability to model global defects or subtle texture changes that are distributed across regions, making it difficult to fully integrate multimodal information for collaborative analysis.

[0004] Based on the above analysis of the development status of this technology field, the existing technologies lack efficient defect identification schemes that fully integrate SEM images and EDX spectral multimodal information. Summary of the Invention

[0005] The purpose of this invention is to provide a method and apparatus for identifying defects in agricultural machinery parts made of aluminum-based composite materials, in order to solve the aforementioned problems in the prior art.

[0006] According to a first aspect of the present invention, a method for identifying defects in aluminum-based composite agricultural machinery parts is provided, comprising: SEM microscopic images and EDX spectral images of aluminum-based composite agricultural machinery parts were obtained respectively. An embedding vector sequence is generated based on SEM microscopic images, and a spectral sequence is generated based on EDX spectral images. The embedding vector sequence and the spectral sequence are then combined to construct a joint sequence. The joint sequence is input into the visual transformation encoder, which interacts with self-attention and cross-modal attention. Global labels are used to aggregate multimodal information and output a global feature representation. Based on the global feature representation, the defect identification result is output.

[0007] According to a second aspect of the present invention, a defect identification device for aluminum-based composite agricultural machinery parts is provided, comprising: The initial acquisition module is used to acquire SEM microscopic images and EDX spectral images of aluminum-based composite agricultural machinery parts, respectively. The joint construction module is used to generate an embedding vector sequence based on SEM microscopic images, a spectral sequence based on EDX spectral images, and to construct a joint sequence from the embedding vector sequence and the spectral sequence. The defect identification module is used to input the joint sequence into the visual transformation encoder. The visual transformation encoder interacts with self-attention and cross-modal attention, uses global labels to aggregate multimodal information and outputs a global feature representation, and outputs the defect identification result based on the global feature representation.

[0008] According to a third aspect of the present invention, an electronic device is provided, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the method for identifying defects in aluminum-based composite agricultural machinery parts as provided in the first aspect of the present disclosure.

[0009] According to a fourth aspect of the present invention, a computer-readable storage medium is provided, on which an information transmission implementation program is stored, which, when executed by a processor, implements the steps of the method for identifying defects in aluminum-based composite agricultural machinery parts provided in the first aspect of the present disclosure.

[0010] The technical solution provided by the embodiments of the present invention has the following beneficial effects: it fully integrates multimodal information such as SEM images and EDX spectra, enhances the powerful global feature modeling capability, so as to achieve accurate, stable and intelligent detection of defects such as cracks, holes, inclusions and delamination in aluminum-based composite agricultural machinery parts.

[0011] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0012] To more clearly illustrate the technical solutions in one or more embodiments of this specification or in the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 This is a flowchart of a method for identifying defects in agricultural machinery parts made of aluminum-based composite materials according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the multimodal Transformer model structure according to an embodiment of the present invention; Figure 3 This is a schematic diagram of an aluminum-based composite material agricultural machinery parts defect identification device according to an embodiment of the present invention; Figure 4 This is a schematic diagram of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0014] To enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the technical solutions in one or more embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, and not all of the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of this document.

[0015] Method Implementation Examples According to embodiments of the present invention, a method for identifying defects in agricultural machinery parts made of aluminum-based composite materials is provided. Figure 1 This is a flowchart of a method for identifying defects in aluminum-based composite agricultural machinery parts according to an embodiment of the present invention, as shown below. Figure 1 As shown, the method for identifying defects in aluminum-based composite agricultural machinery parts according to an embodiment of the present invention specifically includes: In step S110, SEM microscopic images and EDX spectral images of the aluminum-based composite agricultural machinery parts are acquired, specifically including: SEM images and EDX spectra were used as multimodal input data; SEM microscopic images were obtained by scanning with an electron microscope, and EDX spectral images were obtained by energy dispersive X-ray diffraction.

[0016] In step S120, an embedding vector sequence is formed based on the SEM microscopic image, and a spectral sequence is formed based on the EDX spectral image. The embedding vector sequence and the spectral sequence are then combined to construct a joint sequence, specifically including: The dimensions of an SEM image are represented as C×H×N, where C represents the number of channels, and H and N represent the height and width of the image, respectively. First, the SEM image is divided into several non-overlapping image blocks, each block being [size missing]. The corresponding number of image patches is Each image patch is then flattened into a one-dimensional vector and projected onto a fixed-dimensional vector through a linear mapping layer. The embedding space is used to obtain the image patch sequence representation; at the same time, the EDX spectral data is preprocessed and feature encoded, and then concatenated with the image patch sequence to form a unified input sequence, so as to realize the joint representation and modeling of multimodal features.

[0017] It is mainly used to standardize the collected multimodal detection data and convert it into a unified feature representation form suitable for subsequent analysis and defect identification. Through this process, the surface and internal defects of aluminum-based composite agricultural machinery parts can be initially quantified and structured, providing basic data support for subsequent feature fusion and defect discrimination.

[0018] SEM images are two-dimensional high-resolution images that can effectively characterize the microscopic morphology of defects on the surface and near the surface of parts, including key information such as crack morphology, pore boundaries, particle distribution, and interface structure, providing a solid foundation for subsequent multimodal fusion.

[0019] First, the original SEM image is resized and standardized. Let the original SEM image be represented as ;in, These represent the image's height, width, and number of channels, respectively. The SEM microscopic images were sequentially resized, normalized, and flattened. To ensure consistency in the model input size, bilinear interpolation was used to uniformly adjust the images. Its calculation form is ,in, This indicates the position of the target pixel in the resized image. Pixel value at; Represents the original SEM micrograph; and These represent the positions of the target pixels in the original image, respectively. The corresponding adjacent horizontal and vertical coordinates; and These represent the interpolation weights for the horizontal and vertical directions, respectively. and The adjacent pixel index is 0 or 1. After adjustment, the image satisfies the following: To reduce imaging noise and brightness differences between different samples, the image pixel values ​​are normalized and mapped to the [0, 1] interval. The normalization formula is as follows: When the SEM image is multi-channel data, perform the above normalization operation on each channel separately. Subsequently, the standardized SEM image is divided into several non-overlapping image blocks of size G×G; in this embodiment, G=16, so each SEM image is divided into... ; Each image patch is flattened to form a one-dimensional vector. Equation 1 is used to flatten the image patch into a one-dimensional vector, and Equation 2 is used to map the one-dimensional vector to a dimension of through a linear projection layer. L Embedded space: Formula 1; Formula 2; in, Indicates the first Each image block is flattened. Represents the set of real numbers. Indicates the side length of the non-overlapping image patch. Indicates the number of channels. Indicates the number of image divisions. This indicates block embedding representation. Represents the weight matrix. Represents the bias vector; To achieve the overall ability to identify defects in aluminum-based composite agricultural machinery parts, a learnable classification tag vector is introduced. In this embodiment of the invention, the first classification tag vector, the second classification tag vector, and the global tag CLS (Class Token) are special tag vectors specifically used for classification tasks. Location encoding is represented as It is used to characterize the spatial location information between different image blocks, which helps the model understand the crack propagation direction, defect aggregation area and various spatially related features, thereby improving the expressive ability of defect recognition; Add learnable first-class label vectors and positional codes, and use Formula 3 to form the embedding vector sequence, which is the complete input sequence of the SEM modality: Formula 3; in, Represents an embedded vector sequence. Represents the first classification label vector. Representing an image Location encoding.

[0020] In aluminum-based composite agricultural machinery parts, defects such as cracks, pores, inclusions, and delamination are often accompanied by localized compositional anomalies, such as enrichment of reinforcing phases, concentration of impurity elements, or uneven element distribution at interfaces. Energy-dispersive X-ray spectroscopy can reflect the elemental composition and corresponding energy distribution characteristics within the defect region, providing an important basis for distinguishing and identifying different defect types. Therefore, EDX spectroscopy is introduced as an important modal feature for characterizing the causes and types of defects in aluminum-based composite agricultural machinery parts. The EDX spectral images were then normalized and segmented sequentially. The raw EDX spectral data acquired within the defect candidate region are represented as follows: ,in, This indicates the number of sampling points on the energy axis; spectral intensity reflects the characteristic peak information of corresponding elements at different energy positions. To eliminate differences in acquisition conditions and measurement scales among different samples, the original EDX spectrum is normalized, and the calculation formula is as follows: ,in, and These represent the maximum and minimum values ​​of the spectral intensity, respectively. In aluminum matrix composites, different defect types often correspond to significantly different compositional characteristics. For example, crack and delamination defect regions typically exhibit discontinuous distribution of interfacial elements; pore defect regions may show an overall decrease in spectral intensity or the absence of characteristic element peaks; while inclusion defects are often accompanied by anomalous element peaks or local element enrichment. To highlight the compositional variation characteristics within local energy ranges and reduce the computational complexity caused by long sequences, this embodiment segments the normalized EDX spectrum into segments with a fixed length J=64. The spectral data is divided into... One spectral segment; Formula 4 is used to represent the segmented spectral segments, each corresponding to a local interval on the energy axis, used to describe the distribution characteristics of elemental peaks within that interval. Formula 5 is used to perform feature embedding processing on the spectral segments using a multilayer perceptron to extract compositional variation patterns related to defect types. Formula 4; Formula 5; in, Indicates the first A spectral segment, Indicates a fixed segment length. Indicates the starting index spectrum. Indicates the end of the index spectrum. Indicates the number of segments. This represents a multilayer perceptron. This represents fragment embedding representation; MLP consists of multiple layers of linear mappings and nonlinear activation functions, used to learn elemental peak intensity features, inter-peak correlations, and local component anomalies. Its output embedding vector has a dimension of [missing information]. L ; Introducing position encoding for each spectral segment embedding vector This is used to characterize the relative order of different energy ranges in the spectral sequence, enabling the model to distinguish the elemental composition variation characteristics corresponding to each energy range, thereby enhancing the ability to identify defect types such as inclusions, component segregation, and interface anomalies; and a learnable classification label vector corresponding to the EDX mode is added. ; Add learnable second-class label vectors and positional codes, and use Formula 6 to form a spectral sequence: Formula 6; in, Represents a spectral sequence. Represents the second classification label vector. Representing fragments Location encoding.

[0021] By concatenating the sequences of SEM and EDX modalities along the sequence dimension, a joint input sequence is constructed. The embedding vector sequence and the spectral sequence are then concatenated to form the initial sequence. ,in, ; Preferably, to distinguish the feature sources of different modalities, modal embedding is introduced into the concatenated sequence, defining a learnable modal embedding vector. Use Equations 7 and 8 to add the learnable modality embedding vectors to the features corresponding to the initial sequence: Formula 7; Formula 8; in, This represents the feature representation after adding modal embedding vectors to the embedding vector sequence. Represents the first in the embedding vector sequence Each value, excluding the marker. Equivalent to , Represents the first modality embedding vector. Indicates the number of image divisions. This represents the feature representation of the spectral sequence after adding modal embedding vectors. Indicates the first spectral sequence One value, Represents the second modality embedding vector. Indicates the number of segments; The final joint sequence can be represented using Equation 9: Formula 9; in, This indicates a joint sequence.

[0022] In step S130, the joint sequence is input into the visual transformation encoder. The visual transformation encoder interacts with self-attention and cross-modal attention, uses global labels to aggregate multimodal information and outputs a global feature representation. Based on the global feature representation, the defect identification result is output, specifically including: joint sequence The input is processed by the ViT encoder; in the first stage, the input is processed by the Transformer encoder to achieve preliminary feature fusion; in the second stage, a cross-modal attention mechanism is introduced to further enhance the explicit interaction between the two modalities and generate a comprehensive feature representation for defect classification.

[0023] Input the joint sequence into a visual transform encoder in the form of a ViT encoder; The ViT encoder is composed of stacked Transformer layers. In this embodiment of the invention, the encoder consists of... L = It consists of 6 stacked Transformer layers. Each Transformer layer includes a multi-head attention layer (MHSA) for information interaction and a feedforward neural network (FFN) for feature transformation and mapping. Except for the limitations of the following embodiments of the present invention, the other processing can be carried out in the existing manner.

[0024] The self-attention mechanism is executed in the first 3 Transformer layers of a predetermined number. By modeling the correlation between the elements of the sequence, the self-attention mechanism is used to achieve global information fusion. The self-attention of the first 3 layers is shown in Equation 10: Formula 10; in, , ; These represent the learnable linear projection matrices corresponding to the query feature, key feature, and value feature, respectively. They represent the joint sequences respectively. The query features, key features, and value features obtained after linear mapping Represents the dimension of the embedded vector. This represents the feature dimension of each attention head. This indicates the number of attention heads.

[0025] The output sequence dimension after processing by the first 3 encoder layers is still (F+E+2)×L. Learnable global labels are introduced after the global information fusion of the first 3 layers. , forming intermediate feature sequences ,in, This represents the initial fused features obtained after the previous Transformer layer has been executed; To further enhance cross-modal feature interaction capabilities, this embodiment of the invention adds a cross-modal attention module after the third encoder and before the fourth encoder. This module performs a bidirectional query based on the initial fused features, using Equations 11 and 12 to represent the cross-modal attention mechanism, which optimizes bidirectional queries between the SEM and EDX modalities. Formula 11; Formula 12; in, , ;in, This represents the feature subsequence corresponding to the SEM modality in the initial fused features output by the first 3 Transformer encoder layers. This represents the feature subsequence corresponding to the EDX mode in the initial fused features output by the first 3 Transformer encoder layers; , , These represent the learnable linear projection matrices corresponding to the query feature, key feature, and value feature in the SEM modality, respectively. , , These represent the learnable linear projection matrices corresponding to the query feature, key feature, and value feature in the EDX modality, respectively. and These represent the query features corresponding to the SEM modality and the EDX modality, respectively. and These represent the key features corresponding to the SEM mode and the EDX mode, respectively. and These represent the value features corresponding to the SEM mode and the EDX mode, respectively. This represents the feature dimension of each attention head.

[0026] The enhanced features are obtained by concatenating the features after cross-modal attention enhancement, which means re-concatenating the SEM and EDX feature sequences after cross-modal attention enhancement. The enhanced features are input into a predetermined number of subsequent Transformer layers for deep modeling, i.e., the subsequent 3 Transformer encoder layers, using global labels. It aggregates multimodal information, that is, it aggregates the features output during the deep modeling process and outputs a global feature representation. .

[0027] The global feature vector integrates morphological and compositional information and can be used for subsequent defect type discrimination and recognition tasks. The global feature representation is input into the fully connected layer, and the softmax function in the fully connected layer is used to predict the defect recognition result.

[0028] Figure 2 This is a schematic diagram of the multimodal Transformer model structure according to an embodiment of the present invention, as shown below. Figure 2 As shown, an architecture for identifying defects in agricultural machinery parts made of aluminum-based composite materials is presented. The feedforward neural network in the Transformer layer can be set as a multilayer perceptron.

[0029] The specific settings for the ViT encoder during training include: For the multimodal identification of defects in aluminum-based composite agricultural machinery parts, the data itself has significant multimodal characteristics and also suffers from class imbalance, i.e., the number of normal samples is significantly greater than the number of defective samples, and there is an asymmetry in the information representation capabilities between different modalities. Therefore, during the model training process, it is necessary to ensure that the model has a fast convergence speed, good generalization ability, and can fully explore and utilize the complementary information between SEM images and EDX spectra.

[0030] Since Transformer models typically rely on large-scale data for pre-training to obtain good parameter initialization, while the sample size of the aluminum-based composite agricultural machinery parts defect dataset is relatively limited, if random initialization is used directly for training, it is easy to cause the model to converge slowly or even overfit. Therefore, this invention employs a training strategy that combines pre-training and fine-tuning.

[0031] In the specific implementation, the encoder of the SEM branch is initialized with the weights of the ViT-B / 16 model pre-trained on the ImageNet dataset. The low-level features (such as edges and textures) learned during ImageNet pre-training have strong generality and can be transferred to the SEM image domain to a certain extent, providing reasonable initial parameters for the Patch Embedding layer and the first few Transformer encoder layers, thereby significantly accelerating the convergence speed of the model.

[0032] Since the EDX branch lacks mature pre-trained models that can be directly transferred, its network weights are randomly initialized and follow a Xavier normal distribution. At the same time, a low learning rate is set for the EDX branch during training (about 1 / 10 of the learning rate of the SEM branch) to ensure that the two branches can maintain a stable and coordinated parameter update rhythm during training.

[0033] To further mitigate the risk of overfitting due to the limited number of samples, this invention introduces data augmentation strategies to expand the training sample space; for SEM images, random horizontal flipping (probability 0.5) and random rotation (angle range of...) are employed. Enhancement is achieved through methods such as 10° to 10° and brightness perturbation (±10%); for EDX spectral data, random Gaussian noise (standard deviation of 0.01) and slight shifts in peak positions (±2 sampling points) are introduced to enhance the model's robustness to noise and measurement fluctuations.

[0034] During the fine-tuning phase, pre-trained weights are loaded into the entire multimodal model, including the SEM branch, EDX branch, and feature fusion module, and end-to-end training is performed based on a dataset of defects in aluminum-based composite agricultural machinery parts. Specifically, the first three layers of the encoder are updated with a low learning rate (1 / 5 of the initial learning rate) to retain general feature representation capabilities; the last three layers of the encoder and the classification head are trained with a normal learning rate, enabling the model to better adapt to specific defect recognition tasks.

[0035] In selecting the optimization algorithm, considering the large parameter scale and high training stability requirements of the Transformer architecture, AdamW was used as the optimizer during training. While retaining the advantages of Adam's adaptive learning rate, AdamW introduces an independent weight decay term, making it more suitable for the training needs of Vision Transformer-like models compared to standard SGD or traditional Adam. AdamW achieves adaptive gradient adjustment through first-order and second-order momentum estimation, which helps to accelerate the convergence speed and alleviate the parameter sparsity problem. At the same time, the weight decay mechanism effectively regularizes the ViT model with approximately 25 million parameters, reducing the risk of overfitting.

[0036] The parameter update process of AdamW can be represented by formulas 13 to 16: Formula 13; Formula 14; Formula 15; Formula 16 in, and These represent the first-order and second-order momentum estimates, respectively. Indicates the first The gradient of the loss function at each step. Indicates model parameters, For learning rate, This is the weight decay coefficient. It is the numerical stability constant. and The momentum decay coefficients are set to 0.9 and 0.999 respectively in this embodiment.

[0037] Regarding hyperparameter settings, this invention uses an initial maximum learning rate. Combined with a cosine annealing learning rate scheduling strategy; the weight decay coefficient is set to... Numerical stability term The batch size is set to 32; the total number of training epochs is set to 200; and an early stopping mechanism is introduced, which terminates training when the validation set loss does not decrease for 10 consecutive epochs.

[0038] Through the above training strategies and parameter configurations, this embodiment of the invention can train a multimodal Transformer model with fast convergence speed and excellent generalization performance under conditions of limited sample size, imbalanced multimodal data, and unequal modal information, thereby achieving high-precision identification and classification of various defects such as cracks, holes, inclusions, and delamination in aluminum-based composite agricultural machinery parts.

[0039] Preferably, in the embodiments of the present invention, for the multimodal recognition task of defects in agricultural machinery parts made of aluminum-based composite materials, the data has the characteristics of class imbalance and modal information asymmetry; therefore, the design of the loss function needs to solve the problem of training difficulties of rare classes at the same time, and balance the contributions of SEM and EDX modes in the feature fusion process. To this end, the present invention proposes a composite loss function that combines Focal Loss and modal balance regularization term. Focal Loss is defined as shown in Formula 17: Formula 17; in, This represents the true label (one-hot encoding). This represents the predicted probability after model fusion. For focusing parameters.

[0040] To balance the influence of SEM and EDX modes in fusion prediction, this embodiment of the invention introduces a mode balance regularization term. Specifically, the method involves calculating independent prediction results for the SEM branch and the EDX branch separately. and The CLSTokens are obtained by inputting each branch's CLSToken into an independent fully connected layer, and then fused together with the prediction. The KL divergence is shown as a regularization term in Formula 18: Formula 18; Wherein, KL divergence is defined as ; The final composite loss function is defined as ; in, The mode balance regularization term ensures that the prediction results of the SEM and EDX branches are consistent with the fusion results, avoiding the dominant role of a single mode in the final prediction.

[0041] Compared to traditional cross-entropy loss, Focal Loss focuses more on hard-to-classify samples and is suitable for solving minority class defect identification problems. Meanwhile, the modal balance regularization term enhances the complementary use of multimodal information, ensuring that EDX chemical features, such as the abnormal distribution of reinforcing phase elements in aluminum matrix composites, play a full role in the classification process, thereby improving the identification accuracy of defects such as cracks, pores, inclusions, and delamination.

[0042] The main difference between the composite loss function of this invention and the prior art is that, considering the special characteristics of modal information of aluminum-based composites, the morphological features of SEM and the compositional features of EDX are highly complementary but have large distribution differences. By using a modal balance regularization term to force the prediction consistency of the two modes, rather than a simple weighted fusion, the deep collaborative utilization of multimodal information can be ensured.

[0043] In summary, addressing the existing problems, this invention presents a defect identification method for aluminum-based composite agricultural machinery parts. It fully integrates SEM images and EDX spectral multimodal information to enhance powerful global feature modeling capabilities. Furthermore, leveraging the self-attention and cross-modal attention mechanisms of the Transformer, the first few layers of self-attention capture long-distance dependencies, while the later layers of cross-modal attention learn the correspondences between different modalities, achieving explicit interaction and bidirectional fusion of morphological and compositional information. The introduction of learnable global classification labels aggregates multimodal information, enabling the model to comprehensively assess defects based on global context. Simultaneously, the use of the AdamW optimizer, hierarchical learning rate strategy, and early stopping mechanism effectively accelerates model convergence and improves generalization ability. Overall, this method achieves accurate, stable, and intelligent detection of defects such as cracks, holes, inclusions, and delamination in aluminum-based composite agricultural machinery parts.

[0044] Device Examples According to an embodiment of the present invention, a defect identification device for agricultural machinery parts made of aluminum-based composite materials is provided. Figure 3 This is a schematic diagram of an aluminum-based composite material agricultural machinery parts defect identification device according to an embodiment of the present invention, as shown below. Figure 3 As shown, the aluminum-based composite material agricultural machinery parts defect identification device according to an embodiment of the present invention specifically includes: The initial acquisition module 30 is used to acquire SEM microscopic images and EDX spectral images corresponding to aluminum-based composite agricultural machinery parts, specifically: to obtain SEM microscopic images by scanning with an electron microscope and to obtain EDX spectral images by energy dispersive X-rays.

[0045] The joint construction module 32 is used to generate an embedding vector sequence based on SEM microscopic images and a spectral sequence based on EDX spectral images, and to construct a joint sequence from the embedding vector sequence and the spectral sequence. Specifically, it is used for: The SEM micrographs were sequentially resized, normalized, and flattened. Formula 1 was used to flatten the image blocks into a one-dimensional vector, and Formula 2 was used to map the one-dimensional vector to a dimension of [dimensional value missing] using a linear projection layer. L Embedded space: Formula 1; Formula 2; in, Indicates the first Each image block is flattened. Represents the set of real numbers. Indicates the side length of the non-overlapping image patch. Indicates the number of channels. Indicates the number of image divisions. This indicates block embedding representation. Represents the weight matrix. Represents the bias vector; Add learnable first-class label vectors and positional codes, and use Formula 3 to form an embedding vector sequence: Formula 3; in, Represents an embedded vector sequence. Represents the first classification label vector. Representing an image Location encoding; The EDX spectral image is then normalized and segmented sequentially. Equation 4 represents the segmented spectral segments, and Equation 5 is used to embed features into the spectral segments using a multilayer perceptron. Formula 4; Formula 5; in, Indicates the first A spectral segment, Indicates a fixed segment length. Indicates the starting index spectrum. Indicates the end of the index spectrum. Indicates the number of segments. This represents a multilayer perceptron. This represents fragment embedding representation; Add learnable second-class label vectors and positional codes, and use Formula 6 to form a spectral sequence: Formula 6; in, Represents a spectral sequence. Represents the second classification label vector. Representing fragments Location encoding.

[0046] The embedded vector sequence and the spectral sequence are concatenated to form the initial sequence; Use Equations 7 and 8 to add the learnable modality embedding vectors to the features corresponding to the initial sequence: Formula 7; Formula 8; in, This represents the feature representation after adding modal embedding vectors to the embedding vector sequence. Represents the first in the embedding vector sequence One value, Represents the first modality embedding vector. Indicates the number of image divisions. This represents the feature representation of the spectral sequence after adding modal embedding vectors. Indicates the first spectral sequence One value, Represents the second modality embedding vector. Indicates the number of segments; The final joint sequence can be represented using Equation 9: Formula 9; in, This indicates a joint sequence.

[0047] The defect identification module 34 is used to input the joint sequence into the visual transformation encoder, through which the visual transformation encoder interacts using self-attention and cross-modal attention, uses global labels to aggregate multimodal information and outputs a global feature representation, and outputs the defect identification result based on the global feature representation. Specifically, it is used for: The joint sequence is input into a visual transform encoder in the form of a ViT encoder, wherein the ViT encoder is composed of stacked Transformer layers, including a multi-head attention layer for information interaction and a feedforward neural network for feature transformation and mapping. A self-attention mechanism is executed in the first few Transformer layers according to a preset number of conditions, and a learnable global label is introduced after global information fusion. , forming intermediate feature sequences ,in, This represents the initial fused features obtained after the previous Transformer layer has been executed; A cross-modal attention mechanism that performs bidirectional queries based on initial fusion features forms enhanced features by concatenating the features obtained after cross-modal attention enhancement. Deep modeling is performed by inputting a preset number of enhanced features into subsequent Transformer layers, using global labeling. It aggregates multimodal information and outputs a global feature representation.

[0048] The global feature representation is input into the fully connected layer, and the defect identification result is predicted by the softmax function in the fully connected layer.

[0049] The ViT encoder specifically includes: AdamW was used as the optimizer during training.

[0050] In summary, addressing the existing problems, this invention presents a defect identification device for aluminum-based composite agricultural machinery parts. It fully integrates SEM images and EDX spectral multimodal information to enhance powerful global feature modeling capabilities. Furthermore, leveraging the self-attention and cross-modal attention mechanisms of the Transformer, the first few layers of self-attention capture long-distance dependencies, while the later layers of cross-modal attention learn the correspondences between different modalities, achieving explicit interaction and bidirectional fusion of morphological and compositional information. The introduction of learnable global classification labels aggregates multimodal information, enabling the model to comprehensively assess defects based on global context. Simultaneously, the use of the AdamW optimizer, hierarchical learning rate strategy, and early stopping mechanism effectively accelerates model convergence and improves generalization ability. Overall, it achieves accurate, stable, and intelligent detection of defects such as cracks, holes, inclusions, and delamination in aluminum-based composite agricultural machinery parts.

[0051] Electronic device examples Figure 4 This is a schematic diagram of an electronic device according to an embodiment of the present invention. The electronic device 400 may include at least one processor 410 and a memory 420. The processor 410 can execute instructions stored in the memory 420. The processor 410 is communicatively connected to the memory 420 via a data bus. In addition to the memory 420, the processor 410 can also be communicatively connected to an input device 430, an output device 440, and a communication device 450 via the data bus.

[0052] Processor 410 can be any conventional processor, such as a commercially available CPU. Processors may also include graphics processing units (GPUs), field-programmable gate arrays (FPGAs), systems-on-chips (SoCs), application-specific integrated circuits (ASICs), or combinations thereof.

[0053] The memory 420 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.

[0054] In this embodiment of the present disclosure, the memory 420 stores executable instructions, and the processor 410 can read the executable instructions from the memory 420 and execute the instructions to implement all or part of the steps of the aluminum-based composite material agricultural machinery component defect identification method in any of the above exemplary embodiments.

[0055] Computer-readable storage medium embodiments In addition to the methods and apparatus described above, exemplary embodiments of this disclosure may also be a computer program product or a computer-readable storage medium storing the computer program product, wherein the computer program product includes computer program instructions that can be executed by a processor to implement all or part of the steps described in any of the aluminum-based composite material agricultural machinery component defect identification methods in the exemplary embodiments described above.

[0056] Computer program products can be written in any combination of one or more programming languages ​​to perform the operations of the embodiments of this application. Programming languages ​​include object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar languages, and scripting languages ​​(e.g., Python). The program code can be executed entirely on the user's computing device, partially on the user's device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0057] Computer-readable storage media may be any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example,, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media include: static random access memory (SRAM) having one or more electrically connected wires; electrically erasable programmable read-only memory (EEPROM); erasable programmable read-only memory (EPROM); programmable read-only memory (PROM); read-only memory (ROM); magnetic storage; flash memory; magnetic disk or optical disk; or any suitable combination thereof.

[0058] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for identifying defects in aluminum-based composite agricultural machinery parts, characterized in that, include: SEM microscopic images and EDX spectral images of aluminum-based composite agricultural machinery parts were obtained respectively. An embedding vector sequence is formed based on the SEM microscopic image, and a spectral sequence is formed based on the EDX spectral image. The embedding vector sequence and the spectral sequence are then combined to form a joint sequence. The joint sequence is input into a visual transformation encoder, which interacts with self-attention and cross-modal attention. Global labels are used to aggregate multimodal information and output a global feature representation. Based on the global feature representation, the defect identification result is output.

2. The method according to claim 1, characterized in that, The specific steps of obtaining SEM microscopic images and EDX spectral images corresponding to aluminum-based composite agricultural machinery parts include: obtaining SEM microscopic images by scanning with an electron microscope and obtaining EDX spectral images by energy dispersive X-ray diffraction.

3. The method according to claim 1, characterized in that, The process of forming an embedding vector sequence based on the SEM microscopic image and forming a spectral sequence based on the EDX spectral image specifically includes: The SEM micrographs were sequentially resized, normalized, and flattened. Formula 1 was used to flatten the image blocks into a one-dimensional vector, and Formula 2 was used to map the one-dimensional vector to a linear projection layer with dimension [0, 1]. L Embedded space: Official 1; Official 2; in, Indicates the first Each image block is flattened. Represents the set of real numbers. Indicates the side length of the non-overlapping image patch. Indicates the number of channels. Indicates the number of image divisions. This indicates block embedding representation. Represents the weight matrix. Represents the bias vector; Add learnable first-class label vectors and positional codes, and use Formula 3 to form an embedding vector sequence: Official 3; in, Represents an embedded vector sequence. Represents the first classification label vector. Representing an image Location encoding; The EDX spectral image is sequentially normalized and segmented. Equation 4 represents the segmented spectral segments, and Equation 5 is used to perform feature embedding processing on the spectral segments using a multilayer perceptron. Official 4; Official 5; in, Indicates the first A spectral segment, Indicates a fixed segment length. Indicates the starting index spectrum. Indicates the end of the index spectrum. Indicates the number of segments. This represents a multilayer perceptron. This represents fragment embedding representation; Add learnable second-class label vectors and positional codes, and use Formula 6 to form a spectral sequence: Official 6; in, Represents a spectral sequence. Represents the second classification label vector. Representing fragments Location encoding.

4. The method according to claim 1, characterized in that, The step of constructing a joint sequence from the embedded vector sequence and the spectral sequence specifically includes: The embedded vector sequence is concatenated with the spectral sequence to form an initial sequence; Use Equations 7 and 8 to add the learnable modality embedding vectors to the features corresponding to the initial sequence: Official 7; Official 8; in, This represents the feature representation after adding modal embedding vectors to the embedding vector sequence. Represents the first in the embedding vector sequence One value, Represents the first modality embedding vector. Indicates the number of image divisions. This represents the feature representation of the spectral sequence after adding modal embedding vectors. Indicates the first spectral sequence One value, Represents the second modality embedding vector. Indicates the number of segments; The final joint sequence can be represented using Equation 9: Official 9; in, This indicates a joint sequence.

5. The method according to claim 1, characterized in that, The step of using a visual transformation encoder to interact with self-attention and cross-modal attention, and using global labels to aggregate multimodal information and output a global feature representation specifically includes: The joint sequence is input into a visual transform encoder in the form of a ViT encoder, wherein the ViT encoder is composed of stacked Transformer layers, the Transformer layers including a multi-head attention layer for information interaction and a feedforward neural network for feature transformation and mapping; A self-attention mechanism is executed in the first few Transformer layers according to a preset number of conditions, and a learnable global label is introduced after global information fusion. , forming intermediate feature sequences ,in, This represents the initial fused features obtained after the previous Transformer layer has been executed; A cross-modal attention mechanism that performs bidirectional queries based on the initial fusion features is used to form enhanced features by concatenating the features obtained after cross-modal attention enhancement. The enhanced features are input into a predetermined number of subsequent Transformer layers for depth modeling, using global labeling. It aggregates multimodal information and outputs a global feature representation.

6. The method according to claim 5, characterized in that, The ViT encoder specifically includes: AdamW was used as the optimizer during training.

7. The method according to claim 1, characterized in that, The step of outputting the defect identification result based on the global feature representation specifically includes: inputting the global feature representation into a fully connected layer, and predicting the defect identification result through the softmax function in the fully connected layer.

8. A defect identification device for agricultural machinery parts made of aluminum-based composite materials, characterized in that, include: The initial acquisition module is used to acquire SEM microscopic images and EDX spectral images of aluminum-based composite agricultural machinery parts, respectively. A joint construction module is used to form an embedding vector sequence based on the SEM microscopic image, form a spectral sequence based on the EDX spectral image, and construct a joint sequence by the embedding vector sequence and the spectral sequence. The defect identification module is used to input the joint sequence into the visual transformation encoder, and the visual transformation encoder interacts with self-attention and cross-modal attention. It uses global labels to aggregate multimodal information and outputs a global feature representation, and outputs the defect identification result based on the global feature representation.

9. An electronic device, characterized in that, include: The memory, the processor, and the computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the method for identifying defects in aluminum-based composite agricultural machinery parts as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores an information transmission implementation program, which, when executed by a processor, implements the steps of the method for identifying defects in aluminum-based composite agricultural machinery parts as described in any one of claims 1 to 7.