A multimodal fusion learning system for alloy fracture surfaces

By using a multimodal fusion learning system that combines image recognition and machine learning, the ratio of dimple fracture to brittle fracture in the fracture surface of 7050 aluminum alloy can be analyzed quickly and accurately. This solves the problems of long time consumption and high cost of traditional methods, and realizes an efficient way to optimize the plasticity and toughness of materials. It is applicable to aerospace, military engineering and high-end manufacturing fields.

CN120411056BActive Publication Date: 2025-12-02CHONGQING UNIV OF ARTS & SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510561058.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-24
Publication Date
2025-12-02
Estimated Expiration
2044-10-24

AI Technical Summary

Technical Problem

Existing technologies make it difficult to quickly and accurately quantify the ratio of dimple fracture to brittle fracture in the fracture surface of 7050 aluminum alloy, which leads to difficulties in optimizing the material's plasticity and toughness. Traditional models are time-consuming and costly, and are unable to describe the complex relationship between the material's microstructure parameters and the fracture ratio.

Method used

A multimodal fusion learning system was adopted, combining image recognition and machine learning. The improved YOLOv5 model was used for semantic segmentation and object detection. The XGBoost model was used to construct a regression prediction model for the proportion of dimple fracture and brittle fracture. The influence of key features was verified through feature importance analysis and SHAP analysis.

Benefits of technology

It enables rapid and accurate identification and prediction of the proportion of dimple fracture and brittle fracture in fracture surfaces, improves the efficiency of obtaining quantitative data on fracture morphology, provides accurate data support for material optimization, and meets the material needs of aerospace, military engineering, and high-end manufacturing fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120411056B_ABST
    Figure CN120411056B_ABST
Patent Text Reader

Abstract

This invention provides a multimodal fusion learning system for alloy fracture surfaces, relating to the field of visual inspection of metal fracture surfaces. The system includes: a data collection module, an image collection module, a target detection module, a prediction module, and a feature analysis module. The data collection module is used to obtain key feature parameters; the image collection module is used to collect images of different alloy fracture surfaces; the target detection module performs semantic segmentation and target detection on the fracture surface images; the prediction module performs prediction and output; and the feature analysis module includes a feature importance analysis module and a SHAP analysis module, used for interpretive analysis of various key features. This system utilizes the combination of image recognition and machine learning to accurately predict dimple fracture and brittle fracture regions in materials, providing data support for material performance optimization.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This invention is a divisional application of patent application number 202411489077X, entitled "Automatic Recognition and Learning Method for Fracture Morphology Based on Multimodal Fusion Learning". Technical Field

[0002] This invention relates to the field of visual inspection technology for metal fracture surfaces, and more specifically to a multimodal fusion learning system for alloy fracture surfaces. Background Technology

[0003] 7050 aluminum alloy is a high-performance Al-Zn-Mg-Cu superalloy widely used in aerospace, military engineering, and high-end manufacturing due to its excellent properties such as high strength, high toughness, and corrosion resistance. However, with the increasing demands for lightweight, high reliability, and long service life in these fields, optimizing the strength, plasticity, and toughness of 7050 aluminum alloy is crucial and remains a challenge. It requires finding a suitable balance among various material properties and achieving a systematic balance between load-bearing capacity, formability, and failure resistance. Research has found a correlation between the proportion of dimple fracture and its plasticity and toughness (i.e., a higher proportion of dimple fracture corresponds to better plasticity and toughness, and vice versa). Therefore, studying the ratio of dimple fracture to brittle fracture regions in the fracture surface of 7050 aluminum alloy can effectively optimize its plasticity and toughness. However, in actual fracture conditions, dimple fracture and brittle fracture usually occur simultaneously or alternately. Quantitative analysis of dimple fracture and brittle fracture not only requires a large amount of human and material resources and is time-consuming, but also makes it impossible to accurately quantify the ratio of dimple fracture to brittle fracture due to the diversity and complexity of the material's microstructure. Currently, the prediction of material strength is mainly achieved through multi-dimensional methods such as experimental testing, theoretical modeling, numerical simulation, and data-driven methods. These methods primarily establish mathematical and physical models including solid solution strengthening, precipitation strengthening, dislocation strengthening, and grain refinement. However, these traditional models are not only time-consuming and costly to collect data, but also struggle to accurately describe the complex relationship between the material's microstructure parameters and the ratio of dimple fracture to brittle fracture, thus increasing the difficulty of optimizing the plasticity and toughness of the material (7050 aluminum alloy). Summary of the Invention

[0004] To address the problems existing in the prior art, the present invention aims to provide a multimodal fusion learning system for alloy fracture surfaces. This system utilizes the combination of image recognition and machine learning to accurately predict the dimple fracture and brittle fracture regions of materials.

[0005] The objective of this invention is achieved through the following technical solution:

[0006] A multimodal fusion learning system for alloy fracture surfaces includes:

[0007] Step S1: Constructing the dataset: The samples undergo heat treatment, tensile testing, and morphology analysis sequentially to construct a dataset containing six key features: first aging time, second aging time, dislocation density, and grain size in three directions. Step S2: Capturing and annotating fracture surface images. After annotation, the dataset is divided into training, validation, and test sets. Step S3: Multimodal fusion learning: The improved YOLOv5 model is used for semantic segmentation of the fracture surface images, and target detection is performed based on the segmented images. Subsequently, a machine learning model is used, employing the six key features as input, to construct a regression prediction model between the proportion of dimple fracture and brittle fracture. Step S4: Feature importance analysis and SHAP analysis: Feature importance analysis is used to calculate the importance of each key feature in predicting the proportion of dimple fracture and brittle fracture. SHAP analysis is used to calculate the influence of each key feature on crack morphology. Combining feature importance analysis and SHAP analysis verifies the accuracy of the key features' influence on the model.

[0008] Based on further optimization of the above scheme, the sample in step S1 is made of 7050 aluminum alloy forging, and the 7050 aluminum alloy forging is cut along the L, T and S directions respectively.

[0009] Based on further optimization of the above scheme, the heat treatment in step S1 is as follows: First, the sample is solution treated at 477℃ for 4 hours; then, it is water quenched at 66℃; subsequently, a two-stage aging process is performed sequentially: the sample is first aged at 121℃ for 4-6 hours, then the temperature is increased to 177℃ at a rate of 5℃ / min, and then a second aging process is performed for 4-7 hours; the tensile test is performed using an electronic universal mechanical testing machine, the cross-sectional area of ​​the tensile specimen is 2×2mm², the gauge length is 20mm, the tensile rate is set to 1.0mm / min, and a stretch gauge with a gauge length of 12.5mm is used; immediately after the heat treatment is completed, the tensile test is repeated 3 times.

[0010] Based on further optimization of the above scheme, the morphology analysis in step S1 is specifically as follows: After the tensile test, X-ray diffractometer analysis is first performed. The instrument is set with a voltage of 45 kV, a current of 40 mA, a step size of 0.01313°, a scanning range of 2θ = 35°~140°, and a scanning speed of 0.05° / s. Each diffraction peak is fitted using a pseudo-voigt function to obtain the full width at half maximum (FWHM). Then, the dislocation density is calculated using a modified Williamson-Hall plot. Specifically:

[0011]

[0012] In the formula: FWHM represents strain broadening in reciprocal space; D g denoted by , where A represents the average grain size; , where A represents the grain area; and b represents the Burgers vector. Represents dislocation density; K Indicates the diffraction vector; Average contrast factor of different diffraction peaks;

[0013]

[0014] In the formula: Indicates the diffraction angle. This represents the full width at half maximum (FWHM) of the diffraction peak. Indicates the wavelength of X-rays;

[0015] Then, the grain size of different surfaces of the sample was analyzed using an electron backscatter diffractometer, employing the equivalent circle diameter. D Granularity representation:

[0016]

[0017] In the formula: A represents the area of ​​the grain;

[0018] The grain size was measured in the L, T, and S directions respectively. These measurements were designated as grain size 1, grain size 2, and grain size 3, respectively, corresponding to the actual tensile direction of the sample in the right-hand Cartesian coordinate system, thus obtaining the grain size in the three directions.

[0019] Based on further optimization of the above scheme, step S2 specifically includes: Image acquisition: Fracture surface imaging is performed using a scanning electron microscope. The image acquisition process follows standard material characterization procedures: imaging is performed under a scanning electron microscope at a fixed magnification of 1000x, and the contrast and clarity of the image are optimized using an accelerating voltage of 3kV and a working distance; for each sample, 10 images of non-overlapping areas are randomly captured. To ensure high-quality image data, imaging is performed under stable vacuum conditions to minimize the impact of external factors on image quality; Image annotation: The dimple and brittle fracture areas are accurately annotated using Labelme software, with red and green polygons used to mark dimples and brittle fractures, respectively; to ensure the consistency and accuracy of annotation, detailed annotation guidelines are developed, clearly defining the microscopic characteristics of dimples and brittle fractures: dimples are generally circular or elliptical pits with smooth edges and uniform distribution, while brittle fractures present as irregular linear or angular network structures; before final confirmation, each image undergoes multiple rounds of review to ensure no mislabeling or omissions.

[0020] Based on further optimization of the above scheme, the ratio of training set, validation set and test set in step S2 is 7:1:2.

[0021] Based on further optimization of the above scheme, the specific steps of step S3, which involves semantic segmentation of the fracture image using the improved YOLOv5 model and target detection based on the segmented image, are as follows:

[0022] Step S31: Process the input image using mosaic data augmentation and normalization to optimize the data input; Step S32: The improved YOLOv5 model includes a backbone, neck, and head: The backbone uses the CSPDarknet53 architecture to extract spatial and semantic information from the image. During the backbone optimization process, more front convolutional layers are added to capture fine texture details, and the use of back convolutional layers is reduced to prevent over-abstraction of these details; The neck uses FPN and PANet structures for feature fusion. During the optimization process, lateral connections are added, and an adaptive weighting mechanism is introduced to ensure that detailed information is not lost. In addition, the bottom-up feature propagation and convolutional operations of key nodes in PANet are enhanced, improving the sensitivity and accuracy of identifying subtle structural changes; The head integrates bounding box regression, object classification, and confidence scoring functions. To optimize performance, a dynamic anchor mechanism is implemented to adaptively adjust the size of the anchor boxes; Step S33: First, for each anchor box, softmax is used... The activation function predicts the probability of belonging to each category, thus completing the category prediction. Then, the model predicts the bounding box regression parameters for each detection box, and the IoU loss function is introduced to improve the accuracy of the bounding box, thereby accurately fitting the target location. Subsequently, the prediction output is post-processed with Non-Maximum Suppression (NMS) to select the optimal bounding box (i.e., removing overlapping low-confidence boxes and retaining high-confidence boxes), thereby achieving the purpose of predicting the location and category of the target box. Finally, a confidence threshold is set to filter low-confidence predictions, thereby reducing the false detection rate. Step S34: The image recognition results are evaluated using precision (P), recall (R), and F1 score.

[0023] Based on further optimization of the above scheme, the specific steps of step S3, which involve constructing a regression prediction model for the ratio of dimple fracture to brittle fracture using the machine learning model, are as follows: Six key features—first aging time, second aging time, dislocation density, and grain size in three directions—are used as inputs and randomly divided into training data and test data, with a training data:test data ratio of 7:3. An XGBoost model is constructed using the training data in a Python environment to establish the relationship between input and output features, with the percentage of dimple fracture and brittle fracture as the output. This analyzes and explores the key factors affecting the ratio of dimple fracture to brittle fracture on the fracture surface, thus providing a reference for optimizing the plasticity and toughness of the material. The constructed XGBoost model is then used to validate the test data, and the accuracy of the model is evaluated using the mean absolute error (MAE), mean square error (MSE), root mean square error (RMSE), and coefficient of determination (R²).

[0024] Based on further optimization of the above scheme, the feature importance analysis in step S4 calculates the key features that play an important role in predicting the proportion of dimple fracture and brittle fracture in the XGBoost model through three factors: Gain, Cover, and Frequency.

[0025] Based on further optimization of the above scheme, the SHAP analysis in step S4 is specifically as follows:

[0026] First, the Shapley value of each key feature is calculated using the SHAP analysis method in Python. :

[0027]

[0028] In the formula: Represents the set of all input features. p Indicates the number of input features; Indicates that the set is not included. The set of all input features; Representing a feature subset S The prediction;

[0029] In Chinese: denominator represents p The features, in any order, have The combination of cases; numerator meaning: in determining the subset S The set of features is subset S It itself contains A sequence of combinations, followed by feature j and the remaining features. contain Combinations, i.e., molecular representations determine subsets S After p Each feature has, under a specific ranking condition Various combinations;

[0030] Then, by summing the Shapley values ​​of all samples, the overall impact of the features is visualized; based on the average absolute Shapley value of each feature, the importance of each feature is measured.

[0031] .

[0032] The following are the technical effects of the present invention:

[0033] This invention utilizes an improved YOLOv5 model to quickly and accurately identify brittle fracture and dimple fracture in images, significantly improving the efficiency of acquiring quantitative fracture morphology data. The acquisition of quantitative fracture morphology data is faster, more efficient, and more accurate. Through the coordination of the XGBoost model's machine learning model with inputs of six key features—first aging time, second aging time, dislocation density, and grain size in three directions—it describes the relationship between microstructure parameters and crack morphology, predicts the special importance of the dimple fracture and brittle fracture ratio, and identifies characteristic influencing factors. Through feature importance analysis and SHAP analysis, it analyzes and explores the key factors affecting the dimple and brittle fracture ratio on the fracture surface, i.e., the fracture morphology. This allows for precise identification of key factors influencing the fracture morphology of metals, providing a fast and convenient new approach for optimizing material plasticity and toughness. The system employed in this invention can efficiently and accurately describe the complex relationship between material microstructure parameters and the proportion of dimple and brittle fracture, thus providing accurate and effective data support for material optimization and effectively meeting the material needs of aerospace, military engineering, and high-end manufacturing fields. Attached Figure Description

[0034] Figure 1 The images show a comparison of the identification of aluminum alloy fracture surfaces using the improved YOLOv5 model in this embodiment of the invention; wherein, Figure 1 (a) is the original image before recognition. Figure 1 (b) is the image after recognition.

[0035] Figure 2 This is a structural block diagram of the improved YOLOv5 model used for semantic segmentation of fracture images in an embodiment of the present invention.

[0036] Figure 3 This is a graph showing the image recognition evaluation results of the improved YOLOv5 model in an embodiment of the present invention.

[0037] Figure 4This is a graph showing the prediction results of the XGBoost model in an embodiment of the present invention.

[0038] Figure 5 This is a schematic diagram of the SHAP analysis results in an embodiment of the present invention.

[0039] Figure 6 This is a schematic diagram showing the results of feature importance analysis in an embodiment of the present invention. Detailed Implementation

[0040] The technical solutions in the embodiments of the present invention will be clearly and completely described below. In the following description, specific details such as specific system structures and technologies are presented for illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present invention.

[0041] Example 1:

[0042] A multimodal fusion learning method for alloy fracture surfaces uses 7050 aluminum alloy forgings as samples, and the 7050 aluminum alloy forgings are cut along the L, T, and S directions respectively, including:

[0043] Step S1: Constructing a dataset: Perform heat treatment, tensile testing, and morphology analysis on the sample in sequence to construct a dataset containing six key features: first aging time, second aging time, dislocation density, and grain size in three directions.

[0044] Step S11, Heat Treatment: First, the sample is solution treated at 477℃ for 4 hours; then, it is water quenched at 66℃; subsequently, a two-stage aging process is performed sequentially: the sample is first aged at 121℃ for 4-6 hours (to obtain the first aging time, in this embodiment, the first aging time can be 4 hours, 5 hours, or 6 hours), then the temperature is increased to 177℃ at a rate of 5℃ / min, and then a second aging is performed for 4-7 hours (to obtain the second aging time, in this embodiment, the second aging time can be 4 hours, 5 hours, 6 hours, or 7 hours). Step S12, Tensile testing: A universal mechanical testing machine (e.g., a 10 kN MTS E43 type) is used. The cross-sectional area of ​​the tensile specimen is 2×2 mm², the gauge length is 20 mm, the tensile rate is set to 1.0 mm / min, and a 12.5 mm long extensometer is used. The tensile test is repeated three times immediately after heat treatment. Step S13, Morphology analysis: After the tensile test, X-ray diffractometer (e.g., a Uranus X-ray diffractometer equipped with a Cu Kα12 radiation source) is used for analysis. The instrument settings are: voltage 45 kV, current 40 mA, step size 0.01313°, scanning range 2θ = 35°~140°, and scanning speed 0.05° / s. Each diffraction peak is fitted using a pseudo-voigt function to obtain the full width at half maximum (FWHM). Then, the dislocation density is calculated using a modified Williamson-Hall plot. Specifically:

[0045]

[0046] In the formula: FWHM represents strain broadening in reciprocal space; D g denoted by , where A represents the average grain size; , where A represents the grain area; and b represents the Burgers vector. Represents dislocation density; K Indicates the diffraction vector; Average contrast factor of different diffraction peaks;

[0047]

[0048] In the formula: Indicates the diffraction angle. This represents the full width at half maximum (FWHM) of the diffraction peak. Indicates the wavelength of X-rays;

[0049] Then, the grain size of different surfaces of the sample was analyzed using an electron backscatter diffractometer (e.g., a Zeiss 300 field emission scanning electron microscope equipped with an Oxford symmetric S1 detector), employing the equivalent circle diameter. D Granularity representation:

[0050]

[0051] In the formula: A represents the area of ​​the grain;

[0052] The grain size was measured in the L, T, and S directions respectively. These measurements were designated as grain size 1, grain size 2, and grain size 3, respectively, corresponding to the actual tensile direction of the sample in the right-hand Cartesian coordinate system, thus obtaining the grain size in the three directions.

[0053] Step S2: Take and annotate the fracture surface images.

[0054] Image Acquisition: Fracture surface imaging was performed using a scanning electron microscope (SEM). The image acquisition process followed standard material characterization procedures: imaging was conducted at a fixed magnification of 1000x under the SEM, and an accelerating voltage of 3kV and working distance were used to optimize image contrast and sharpness. For each sample, 10 images of non-overlapping areas were randomly captured. To ensure high-quality image data, imaging was performed under stable vacuum conditions to minimize the impact of external factors on image quality. Image Annotation: The dimple and brittle fracture regions were precisely annotated using Labelme software (an open-source image annotation tool developed by MIT). Red and green polygons were used to mark dimples and brittle fractures, respectively. To ensure consistency and accuracy of annotation, detailed annotation guidelines were developed, clearly defining the microscopic characteristics of dimples and brittle fractures: dimples are generally circular or elliptical pits with smooth edges and uniform distribution, while brittle fractures present as irregular linear or angular network structures. Each image underwent multiple rounds of review before final confirmation to ensure no mislabeling or omissions.

[0055] After labeling, the dataset is divided into training, validation, and test sets in a ratio of 7:1:2.

[0056] Step S3, Multimodal Fusion Learning: The improved YOLOv5 model is used to perform semantic segmentation on the fracture image, and target detection is performed based on the segmented image, such as... Figure 2 As shown, the specific steps are as follows:

[0057] Step S31: Process the input image using mosaic data augmentation and normalization to optimize the data input; specifically: First, using mosaic data augmentation technology, scale and crop the four original images (each 640×640), and then combine them into a 640×640 composite image, so that the model can see more diverse features in a single image; then, in order to effectively handle the differences in pixel values ​​between different images, perform normalization processing, scaling the pixel values ​​from 0 to 255 to between 0 and 1, that is: pixel value 255 becomes 1 after normalization processing, while 0 remains 0;

[0058] Image normalization aims to scale input pixel values ​​to a uniform range (typically 0 to 1) to eliminate differences in pixel value ranges between images and enhance the stability of model training. The specific formula is as follows:

[0059]

[0060] In the formula: I Represents the original image matrix. max(I), min(I) These represent the maximum and minimum values ​​of the original image matrix, respectively. This represents the image matrix after normalization.

[0061] Step S32: The improved YOLOv5 model includes the trunk, neck, and head:

[0062] The backbone uses the CSPDarknet53 architecture to extract spatial and semantic information from images. During backbone optimization, more front-layer convolutional layers are added to capture fine texture details, while the use of later convolutional layers is reduced to prevent over-abstraction of these details. (Refer to...) Figure 2As shown: In the input part of the backbone network, the initial convolutional block (32x320x320 CBS) is responsible for extracting low-level features, increasing the number of channels to 32 while maintaining a high spatial resolution (320x320). Here, CBS represents Convolution, Batch Normalization, and ReLU activation function. During feature extraction, the CSP (Cross Stage Partial) module optimizes computation through branching structure while maintaining feature expressiveness. The CSP1_1 and CSP1_2 modules enhance the network's feature learning ability by segmenting and merging features. These modules focus on optimizing gradient flow, helping the network learn complex features more efficiently. As the network deepens, the spatial size of the feature map gradually decreases, while the number of channels gradually increases (e.g., from 64x160x160). The feature map size was increased to 128x80x80 through downsampling (convolution or pooling with a stride of 2). Subsequent CBS and CSP modules (e.g., CSP1_3) further enhanced this, ultimately increasing the number of channels in the feature map to 512. The network also introduced the SSPF (Spatial Pyramid Pooling Fast) module for multi-scale feature fusion. This module enhances the network's ability to detect targets of different sizes by pooling and merging features at different scales, ensuring the generation of fixed-size feature maps and improving detection robustness. The neck region utilizes FPN and PANet structures for feature fusion. During optimization, lateral connections were added, and an adaptive weighting mechanism was introduced to ensure no loss of detailed information. Furthermore, bottom-up feature propagation and convolutional operations at key nodes in PANet were enhanced, improving the sensitivity and accuracy of identifying subtle structural changes. (Refer to...) Figure 2As shown, the Neck network utilizes upsampling and concatenation operations to process feature maps during feature fusion. Upsampling adjusts low-resolution feature maps to higher resolutions, facilitating fusion with high-resolution feature maps. Concatenation integrates the upsampled feature maps with the previous ones, enabling the model to utilize both high-level semantic information and low-level detail information, thereby enhancing feature descriptive capabilities. Furthermore, a small CSP module continues to play a role in the Neck network, maintaining the efficiency and diversity of the feature flow through further convolution and normalization operations. This design enriches the feature representation of the Neck network, contributing to improved object detection accuracy. By combining high-level semantics with low-level spatial information, the Neck network significantly enhances the model's multi-scale detection capabilities. This capability allows the model to exhibit better detection performance and robustness when facing targets of different sizes, especially small targets. The FPN structure is used to fuse features across different scales to enhance multi-scale target detection capabilities; its core idea is bottom-up multi-level feature aggregation. The PANet structure further improves the FPN structure by adding a top-down path to enhance information flow and better aggregate features. The head integrates bounding box regression, object classification, and confidence scoring functions. To optimize performance, a dynamic anchoring mechanism is implemented to adaptively adjust the size of the anchor boxes; (Refer to...) Figure 2 As shown: First, the head network typically contains multiple prediction layers, each operating on different feature scales. The head network has two main branches: a classification branch and a regression branch. The classification branch is responsible for determining the category of each candidate region in the feature map, generating predictions corresponding to the number of categories through a series of convolutional operations. The regression branch is responsible for predicting the location and size of the target, typically adjusting predefined anchor boxes (or anchors) using bounding box regression techniques to better fit the target's location. The head network uses an anchor box mechanism; anchor boxes are predefined and have different shapes and sizes to identify targets at specific locations, helping the model cope with targets of different scales and aspect ratios. For model training, the head network uses loss functions that simultaneously consider classification error and bounding box regression error, such as cross-entropy loss and smoothed L1 loss.

[0063] Step S33: First, for each anchor box, the softmax activation function is used to predict the probability of belonging to each category, thus completing the category prediction. Then, the bounding box regression parameters of each detection box are predicted by the model, and the IoU loss function is introduced to improve the accuracy of the bounding box, thereby accurately fitting the target position. Subsequently, Non-Maximum Suppression (NMS) post-processing is applied to the prediction output to select the optimal bounding box (i.e., removing overlapping low-confidence boxes and retaining high-confidence boxes), thereby achieving the purpose of predicting the position and category of the target box. Finally, a confidence threshold is set to filter low-confidence predictions, thereby reducing the false detection rate.

[0064] Step S34: Evaluate the image recognition results using precision (P), recall (R), and F1 score (where precision is used to evaluate the accuracy of the classification model, recall is used to evaluate the model's ability to identify positive samples, and F1 score is used to comprehensively evaluate the model's performance; precision, recall, and F1 score are calculated using methods conventional in the art); refer to Figure 3 As shown: the light blue line corresponds to the dimple fracture category, the yellow line corresponds to the brittle fracture category, and the thicker dark blue line represents the average value of all categories; for example... Figure 3 As shown in (a): when the confidence level is 0.497, the F1 score for each class is 0.98, indicating that the improved YOLOv5 model performs well in predicting dimple fracture and brittle fracture, with high accuracy and recall. The accuracy curve is shown below. Figure 3 As shown in (b), the overall trend indicates that accuracy increases with confidence; accuracy reaches 1 when confidence reaches 0.845. Figure 3 (c) The recall rates at different confidence levels are given. The recall rates for dimple fracture and brittle fracture initially remain unchanged, then decline rapidly. Initially, the average recall rate is slightly lower than that for brittle fracture and slightly higher than that for dimple fracture. Finally, the three rates are close, especially when the confidence index exceeds 0.8. Figure 3(d) shows the accuracy and recall of the classifier at different thresholds: the dimple fracture line remains stable at around 0.990, while the brittle fracture line remains at 0.995, and the mean precision-recall curve hovers around 0.993. These three lines indicate that as the recall increases, the accuracy decreases slightly, but overall remains at a satisfactory level, demonstrating that the improved YOLOv5 model performs very well in identifying dimple fracture and brittle fracture, exhibiting high stability and accuracy. Subsequently, a machine learning model was employed, using six key features as input, to construct a regression prediction model between the proportion of dimple fracture and brittle fracture. Specifically, the six key features—first aging time, second aging time, dislocation density, and grain size in three directions—were used as input, and randomly divided into training and testing data with a ratio of 7:3. An XGBoost model was then constructed using the training data in a Python environment to establish the relationship between the input and output features, with the percentage of dimple fracture and brittle fracture as the output. This analysis explored the key factors influencing the proportion of dimple fracture and brittle fracture on the fracture surface, thus providing a reference for optimizing the plasticity and toughness of the material. The constructed XGBoost model was then validated using test data, and the accuracy of the model was evaluated using mean absolute error (MAE), mean square error (MSE), root mean square error (RMSE), and coefficient of determination (R²).

[0065]

[0066] In the formula: x i This represents the true value of the sample; y i Represents the predicted value of the sample; The mean squared difference (MAE) represents the average of all true values. The MAE calculates the mean absolute difference between the predicted and true values: a smaller MAE value indicates that the model's predictions are closer to the actual values, and thus the model's accuracy is higher. The mean squared difference (MSE) calculates the average of the squared differences between the predicted and true values: a smaller MSE value indicates a smaller prediction error and higher prediction accuracy. The mean squared error (RMSE) is the square root of the MSE and measures the deviation between the predicted and true values: a smaller RMSE value indicates that the model's predictions are more consistent with the actual values. R² is used to evaluate the performance of a regression model: the closer the R² value is to 1, the stronger the model's ability to explain data variability, indicating better model performance.

[0067] The mean absolute error (MAE), mean squared error (MSE), root mean squared error (RMSE), and coefficient of determination (R²) thresholds are preset respectively. If the MAE, MSE, RMSE, and R² values of the XGBoost model are all within the preset range, the XGBoost model is determined to be qualified; otherwise, the training of the XGBoost model is restarted until it is qualified. In this embodiment, the MAE, MSE, RMSE, and R² for the prediction of the dimple fracture and brittle fracture ratios of the XGBoost model using the training data are 5.05, 2.25, 1.91, and 0.73 respectively; the prediction results and experimental results of the XGBoost model are as Figure 4 shown: In the prediction of dimple fracture, the R² values of the training data and the test data are 0.98 and 0.71 respectively; in the prediction of brittle fracture, the R² values of the training data and the test data are 0.99 and 0.75 respectively; the experimental data shows that the XGBoost model has high accuracy in predicting the crack morphology characteristics and can effectively capture the relationship between the crack morphology characteristics and the microstructural parameters.

[0068] Step S4, Feature Importance Analysis and SHAP Analysis: Calculate the importance of each key feature in predicting the dimple fracture and brittle fracture ratios through feature importance analysis. In this embodiment, the key features that play an important role in predicting the dimple fracture and brittle fracture ratios in the XGBoost model are calculated through three factors: Gain, Cover, and Frequency; among them, Gain (gain) is used to represent the contribution of a certain feature to the improvement of the model prediction performance when used for splitting, Cover (coverage) is used to represent the sum of the number of data samples used for splitting when constructing a tree, and Frequency (frequency) represents the number of times a feature is used as a splitting point. The feature with a higher number of usage times contributes more to the model prediction (Gain, Cover, and Frequency all adopt conventional means in the art). After performing the importance feature analysis, the results are visualized. Referring to Figure 6 shown, according to the importance feature analysis: the ratios of dimple fracture and brittle fracture are mainly affected by the dislocation density and the first aging time, and the influence of grain size on dimple fracture and brittle fracture is relatively small.

[0069] Calculate the influence degree of each key feature on the crack morphology through SHAP analysis (the input is the trained XGBoost and the training data, and the output is the SHAP value of each key factor). The specific SHAP analysis is as follows:

[0070] First, use the SHAP analysis in python to explain the model and calculate the Shapley value of each key feature respectively :

[0071]

[0072] In the formula: Represents the set of all input features. p Indicates the number of input features; Indicates that the set is not included. The set of all input features; Representing a feature subset S The prediction;

[0073] In Chinese: denominator represents p The features, in any order, have The combination of cases; numerator meaning: in determining the subset S The set of features is subset S It itself contains A sequence of combinations, followed by feature j and the remaining features. contain Combinations, i.e., molecular representations determine subsets S After p Each feature has, under a specific ranking condition Various combinations;

[0074] Then, by summing the Shapley values ​​of all samples, the overall impact of the features is visualized; based on the average absolute Shapley value of each feature, the importance of each feature is measured.

[0075] .

[0076] Figure 5 This is a summary graph of SHAP values ​​in this embodiment. The position on the x-axis represents the influence of the feature value on the prediction; red dots indicate a significant feature influence, and blue dots indicate a minor influence. According to SHAP analysis, the key features with the greatest impact on dimple fracture and brittle fracture are dislocation density and first aging time, which is consistent with the results of feature importance analysis. Combining feature importance analysis and SHAP analysis verifies the accuracy of the key feature's influence on the model, confirming that dislocation density and first aging time are the key features with the greatest impact on dimple fracture and brittle fracture.

[0077] Example 2:

[0078] As another preferred embodiment of the present invention, based on the scheme of Example 1, the bounding box parameters predicted in the head network include the center position offset (t) x ,t y ), width and height (t) w ,t h and confidence level .

[0079]

[0080] In the formula, This indicates the coordinates of the center point of the predicted bounding box in the feature map, along with its length and width. Indicates the coordinates of the top-left point of each unit; This represents the length and width of the prior bounding box relative to the feature map.

[0081] Example 3:

[0082] A multimodal fusion learning system for alloy fracture surfaces employs the automatic identification and learning method for fracture morphology based on multimodal fusion learning as described in either Example 1 or Example 2. The system includes: a data collection module, an image collection module, a target detection module, a prediction module, and a feature analysis module.

[0083] The data collection module is electrically connected to the heat treatment timer, electronic universal mechanical testing machine, X-ray diffractometer, and electron backscatter diffractometer to obtain key characteristic parameters such as the first aging time, the second aging time, dislocation density, and grain size in three directions. The image collection module is connected to an external camera device to collect images of different fracture surfaces of the alloy. The target detection module uses an improved YOLOv5 model to perform semantic segmentation on the fracture surface images and performs target detection at the fracture surface based on the segmented images. The prediction module uses the XGBoost model to predict the ratio of dimple fracture to brittle fracture at the fracture surface and outputs the results. The feature analysis module includes a feature importance analysis module and a SHAP analysis module, which perform interpretability analysis on each key feature in the XGBoost model to screen the most important key features affecting the ratio of dimple fracture to brittle fracture at the fracture surface, providing data support for the research of alloy materials.

Claims

1. A multimodal fusion learning system for alloy fracture surfaces, characterized in that: include: Data collection module, image collection module, object detection module, prediction module, feature analysis module. The data collection module is electrically connected to the heat treatment timer, electronic universal mechanical testing machine, X-ray diffractometer, and electron backscatter diffractometer to obtain key characteristic parameters such as the first aging time, the second aging time, dislocation density, and grain size in three directions. The image collection module is connected to an external camera device to collect images of different fracture surfaces of the alloy. The target detection module uses an improved YOLOv5 model to perform semantic segmentation on the fracture surface images and performs target detection at the fracture surface based on the segmented images. The prediction module uses the XGBoost model to predict the ratio of dimple fracture to brittle fracture at the fracture surface and outputs the prediction results. The feature analysis module includes a feature importance analysis module and a SHAP analysis module, which respectively perform interpretive analysis on each key feature in the XGBoost model; The system's multimodal fusion learning method involves sequentially performing heat treatment, tensile testing, and morphology analysis on the sample. Specifically, the morphology analysis is performed as follows: After the tensile test, X-ray diffraction is used for analysis. The instrument settings are: voltage 45 kV, current 40 mA, step size 0.01313°, scanning range 2θ = 35°–140°, and scanning speed 0.05° / s. Each diffraction peak is fitted using a pseudo-voigt function to obtain the full width at half maximum (FWHM). Then, the dislocation density is calculated using a modified Williamson-Hall plot. In the formula: FWHM represents strain broadening in reciprocal space; D g Indicates the average grain size; A represents the grain area; b represents the Burgers vector; Represents dislocation density; K Indicates the diffraction vector; Average contrast factor of different diffraction peaks; In the formula: Indicates the diffraction angle. This represents the full width at half maximum (FWHM) of the diffraction peak. Indicates the wavelength of X-rays; Then, the grain size of different surfaces of the sample was analyzed using an electron backscatter diffractometer, employing the equivalent circle diameter. D Granularity representation: In the formula: A represents the area of ​​the grain; The grain size was measured in the L, T, and S directions respectively. These measurements were designated as grain size 1, grain size 2, and grain size 3, respectively, corresponding to the actual tensile direction of the sample in the right-hand Cartesian coordinate system, thus obtaining the grain size in the three directions.

2. The multimodal fusion learning system for alloy fracture surfaces according to claim 1, characterized in that: The steps of this system's multimodal fusion learning method are as follows: Step S1: Constructing the dataset: Perform heat treatment, tensile testing, and morphology analysis on the samples sequentially to construct a dataset containing six key features: first aging time, second aging time, dislocation density, and grain size in three directions. Step S2: Capturing and annotating fracture surface images. After annotation, divide the dataset into training, validation, and test sets. Step S3: Multimodal fusion learning: Use an improved YOLOv5 model to perform semantic segmentation on the fracture surface images and perform target detection based on the segmented images. Subsequently, use a machine learning model, with the six key features as input, to construct a regression prediction model between the proportion of dimple fracture and brittle fracture. Step S4: Feature importance analysis and SHAP analysis: Calculate the importance of each key feature in predicting the proportion of dimple fracture and brittle fracture through feature importance analysis, and calculate the influence of each key feature on crack morphology through SHAP analysis.

3. The multimodal fusion learning system for alloy fracture surfaces according to claim 2, characterized in that: The sample in step S1 is made of 7050 aluminum alloy forging, and the 7050 aluminum alloy forging is cut along the L, T and S directions respectively; the heat treatment in step S1 is as follows: first, the sample is solution treated at 477°C for 4 hours; then, it is water quenched at 66°C. Subsequently, a two-stage aging process was performed sequentially: the sample was first aged at 121℃ for 4–6 hours, and then the temperature was increased to 177℃ at a rate of 5℃ / min for a second aging process of 4–7 hours; the tensile test in step S1 was specifically performed using an electronic universal mechanical testing machine, with a cross-sectional area of ​​2×2mm² for the tensile specimen, a gauge length of 20mm, a tensile rate of 1.0mm / min, and a stretch gauge with a gauge length of 12.5mm; immediately after the heat treatment was completed, the tensile test was repeated 3 times.

4. The multimodal fusion learning system for alloy fracture surfaces according to claim 3, characterized in that: Step S2 specifically involves: Image acquisition: Fracture surface imaging was performed using a scanning electron microscope. The image acquisition process followed standard material characterization procedures: imaging was conducted under a scanning electron microscope at a fixed magnification of 1000x, and the contrast and sharpness of the images were optimized using an accelerating voltage of 3kV and a working distance. For each sample, 10 images of non-overlapping areas were randomly captured, and imaging was performed under stable vacuum conditions. Image annotation: The dimple and brittle fracture areas were accurately annotated using Labelme software, with red and green polygons used to mark dimples and brittle fractures, respectively. Microscopic features of dimples and brittle fractures were defined: dimples are circular or elliptical pits with smooth edges and uniform distribution, while brittle fractures present as irregular linear or angular network structures.

5. The multimodal fusion learning system for alloy fracture surfaces according to claim 2, characterized in that: The specific steps of performing semantic segmentation on the fracture image using the improved YOLOv5 model in step S3, and then performing target detection based on the segmented image, are as follows: Step S31: Process the input image using mosaic data augmentation and normalization to optimize the data input; Step S32: The improved YOLOv5 model includes the trunk, neck, and head: The trunk uses the CSPDarknet53 architecture to extract spatial and semantic information from the image; The neck uses FPN and PANet structures for feature fusion. The head integrates bounding box regression, target classification, and confidence scoring functions; Step S33: First, for each anchor box, the softmax activation function is used to predict the probability of belonging to each category, completing the category prediction; then, the bounding box regression parameters of each detection box are predicted by the model, and the IoU loss function is introduced to improve the accuracy of the bounding box, thereby accurately fitting the target position; subsequently, the prediction output is post-processed with Non-Maximum Suppression to select the optimal bounding box; finally, a confidence threshold is set to filter low-confidence predictions, thereby reducing the false detection rate; Step S34: The image recognition results are evaluated using precision, recall, and F1 score.

6. The multimodal fusion learning system for alloy fracture surfaces according to claim 5, characterized in that: The specific steps of constructing a regression prediction model between the proportion of dimple fracture and brittle fracture using the machine learning model in step S3 are as follows: Six key features—first aging time, second aging time, dislocation density, and grain size in three directions—are used as inputs and randomly divided into training data and test data, with a training data:test data ratio of 7:

3. An XGBoost model is constructed using the training data in a Python environment to establish the relationship between the input and output features, with the percentage of dimple fracture and brittle fracture as the output. The key factors influencing the proportion of dimple fracture and brittle fracture on the fracture surface are analyzed and explored. The constructed XGBoost model is then used to validate the test data, and the accuracy of the model is evaluated using the mean absolute error (MAE), mean square error (MSE), root mean square error (RMSE), and coefficient of determination (R²).

7. The multimodal fusion learning system for alloy fracture surfaces according to claim 2, characterized in that: The feature importance analysis in step S4 calculates the key features that play an important role in predicting the proportion of dimple fracture and brittle fracture in the XGBoost model using three factors: Gain, Cover, and Frequency.

Citation Information

Patent Citations

  • Audience evaluation data-driven silent product video creation auxiliary method and device

    CN114005077A

  • Method for improving YOLOv8 network and application of method in strip steel surface defect detection

    CN118365599A