Construction method of acoustic time sequence image and application of acoustic time sequence image in metal plate crack detection
By constructing acoustic temporal images and combining them with deep learning models, the problems of low efficiency and poor recognition effect in metal plate crack detection in existing technologies are solved. High-precision, low-complexity micro-crack recognition is achieved, which is suitable for multi-frequency excitation and large-area non-contact measurement.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HEFEI UNIV OF TECH
- Filing Date
- 2026-01-27
- Publication Date
- 2026-05-12
Smart Images

Figure CN122017016A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of nondestructive testing and intelligent recognition technology, specifically to a method for identifying and locating cracks in metal plates by constructing acoustic time-series images based on signals acquired by an acoustic array and combining them with a deep learning model. Background Technology
[0002] Due to their excellent specific strength, machinability, and formability, sheet metal is widely used in the manufacturing and assembly of various components such as vehicle bodies, ship hulls, wings, shells, and tanks in the engineering and machinery manufacturing fields. During service, sheet metal is susceptible to micro-cracks due to manufacturing defects, external loads, impacts and vibrations, environmental corrosion, fatigue accumulation, and structural aging. These micro-cracks can gradually propagate under cyclic loading, eventually leading to fracture and failure. Therefore, timely and reliable identification of crack locations in sheet metal is of significant engineering importance for preventing safety accidents, reducing maintenance costs, and providing a basis for material selection and process optimization.
[0003] Among existing crack detection methods, manual inspection relies on experience, resulting in low efficiency and difficulty in detecting early, minute defects. Contact inspection methods are sensitive to surface conditions and coupling conditions, have limited detection efficiency, and are difficult to rapidly cover large-area structures. Some methods are not well-suited for curved surfaces, thin plates, or coated components. Acoustic inspection methods offer advantages such as non-contact operation and high sensitivity, enabling the identification of early cracks without damaging the structure.
[0004] With the development of data-driven methods, deep learning has been gradually introduced into non-destructive testing scenarios. Among them, convolutional neural networks (CNNs) can replace the traditional "manual feature extraction + classifier" process through end-to-end learning, demonstrating good feature extraction capabilities and robustness in 2D image recognition tasks. However, in metal plate crack detection tasks, sound field information not only includes spatial distribution but also changes continuously with time and excitation frequency. 2D convolution only extracts features on a plane and cannot effectively capture the dynamic changes of the sound field in the time dimension, nor can it characterize the temporal response features under different frequency excitations. Therefore, simply using 2D CNNs to process sound pressure cloud maps will lose key temporal correlation information, and the recognition effect will significantly decrease when the crack is small or the sound field difference is weak.
[0005] To simultaneously consider spatial and temporal features, 3D CNNs are used to jointly model dynamic features through convolutional operations in the spatiotemporal domain, effectively addressing the issue of 2D CNNs neglecting temporal information. However, 3D CNNs still fundamentally rely on local convolutional operations with a fixed receptive field, which limits their ability to handle long-term signals or cross-regional global dependencies. Specifically, 3D CNNs exhibit low sensitivity to global changes in the sound field at different excitation frequencies, making it difficult to effectively capture long-range relevant information during sound wave propagation. Summary of the Invention
[0006] To overcome the shortcomings of the prior art, this invention provides a method for constructing acoustic temporal images and its application in metal plate crack detection. It aims to achieve high-precision localization and classification of metal plate cracks under non-contact conditions, and is particularly suitable for the identification of micro-cracks under multi-frequency excitation conditions.
[0007] To achieve its objectives, the present invention employs the following technical solution: The method for constructing an acoustic temporal image in this invention is characterized by: applying excitation to a metal plate and using an acoustic array to non-contactly measure surface sound field information to obtain sound pressure sampling sequences at multiple spatial locations; segmenting the sound pressure sequence according to a preset time interval, and mapping the sound pressure data of each time interval into a two-dimensional sound pressure amplitude distribution map based on the spatial arrangement of the sensors; arranging the two-dimensional sound pressure amplitude distribution map in chronological order to construct an acoustic temporal image, which is used to reflect the spatiotemporal changes in sound pressure distribution.
[0008] The method for constructing the acoustic temporal image of the present invention is carried out according to the following steps: Step 1, Acoustic signal acquisition: Apply excitation to the metal plate and use an acoustic array to perform non-contact measurement of the sound field on the surface of the metal plate, and obtain the sound pressure sampling sequence at multiple spatial locations within the same measurement plane; Step 2, Time Segmentation Processing: The sound pressure sampling sequence is segmented according to a preset time interval to obtain sound pressure sampling signals within multiple time segments; Step 3, Spatial mapping processing: Combining the spatial arrangement relationship of each sensor in the acoustic array, the corresponding sound pressure sampling data in each time segment is mapped into a two-dimensional sound pressure amplitude distribution map; Step 4, Image Mapping and Temporal Arrangement: Arrange the two-dimensional sound pressure amplitude distribution maps in chronological order to construct an acoustic temporal image, which is composed of multiple two-dimensional sound pressure amplitude distribution maps arranged in chronological order.
[0009] The acoustic time-series image construction method of this invention is characterized by the following: time segmentation processing involves dividing the high temporal resolution sound pressure sampling sequence into segments according to a preset time interval, and recombining the original high sampling rate acoustic data into an acoustic time-series image composed of two-dimensional sound pressure amplitude distribution maps corresponding to multiple time segments. This reduces the data scale used for crack detection from time-series level data to image sequence level data. Thus, while retaining key spatiotemporal feature information of the sound field, it significantly reduces data storage requirements and improves the engineering feasibility and real-time processing capability of the crack detection method. The acoustic time-series image can be constructed in the form of a two-dimensional grayscale image, a multi-channel color image, or a multi-feature channel image.
[0010] The method for constructing acoustic time-series images in this invention is also characterized by the following: In step 2, time segmentation is performed as follows: Let s(t) represent the sound pressure sampling signal, which is the time function corresponding to the sound pressure sampling sequence, with a total duration of T; the total duration T is divided into N equal time intervals, with... The start time of the nth time interval is represented by... Represents the end time of the nth time interval. For the preset time range, Then we have: ; ; The sound pressure data sequence in the nth time period for: , ; in, .
[0011] The method for constructing acoustic temporal images in this invention is also characterized by the following: In step 3, spatial mapping processing is performed as follows: by The representation is located in spatial coordinates The sound pressure data is obtained from the sound pressure sensor at the location. The corresponding sound pressure amplitude characteristics Obtained from equation (1) or equation (2): (1); (2).
[0012] The method for constructing acoustic temporal images in this invention is also characterized by the following: in step 4, image mapping and temporal arrangement are performed as follows: sound pressure amplitude characteristics Normalization is performed, and the result is mapped to pixels in a two-dimensional image. Image channel intensity value : , where the mapping function The linear normalized form shown in equation (3) is: (3); in, and These represent the reference minimum and reference maximum values corresponding to the sound pressure amplitude characteristics, respectively. c For image channel indexing, This is a scaling factor used to map the normalized grayscale intensity to a preset image display range; The mapping function f(·) may be a nonlinear mapping, a piecewise mapping, or an adaptive mapping. The reference minimum and reference maximum values are determined based on a single two-dimensional sound pressure amplitude distribution map, multiple time segments, or a preset statistical range.
[0013] The crack detection model for metal plate crack detection in this invention is characterized by being based on acoustic time-series images. The crack detection model for metal plate crack detection includes: The local spatiotemporal feature extraction module is used to perform joint feature extraction on acoustic time-series images in the spatial and temporal dimensions to obtain feature representations that characterize the local sound field changes. The temporal feature modeling module is used to model the global dependencies of local spatiotemporal features in the time dimension in order to obtain a global temporal feature representation that characterizes the evolution of the sound field. The feature fusion and discrimination module is used to fuse local spatiotemporal features and global temporal features, and output the detection results of metal plate cracks based on the fused features; The crack detection model is trained using an acoustic time-series image dataset consisting of several acoustic time-series images. The acoustic time-series images are used as training samples to input into the crack detection model, and the model parameters are learned and optimized to obtain a parameterized crack detection model for metal plate crack detection. The parameterized crack detection model is used to characterize whether there is a crack in the metal plate and the mapping relationship between the crack location and the acoustic time-series images. The crack detection model performs feature mapping on the acoustic time-series image to obtain a feature representation for crack discrimination. The feature mapping relationship is expressed as follows: (4); In formula (4): This represents an acoustic time-series image composed of T-frame two-dimensional sound pressure amplitude distribution maps, with... This represents a nonlinear mapping function used for global spatiotemporal feature modeling and temporal feature extraction. This represents the spatiotemporal characteristics of the fused sound field.
[0014] The crack detection model for metal plate crack detection in this invention is also characterized by the following: the training process of the crack detection model includes: Acoustic time-series image samples containing known crack states and crack locations are obtained, and an acoustic time-series image training dataset with crack annotation information is constructed. The acoustic temporal image training dataset is input into the crack detection model, and the crack detection prediction result is obtained through forward feature mapping. Based on the crack detection prediction results and the corresponding crack annotation information, a model training loss function is constructed, and a parameter optimization algorithm is used to iteratively update the parameters of the crack detection model. When the model training loss function satisfies the preset convergence condition, a parameterized crack detection model for metal plate crack detection is obtained.
[0015] The characteristics of the metal plate crack detection method of the present invention are: An acoustic time-series image of the metal plate to be tested is constructed according to the method of claim 1, which is the acoustic time-series image of the metal plate to be tested; The acoustic time-series image of the metal plate to be detected is input into the trained parameterized crack detection model of the metal plate to be detected. The parameterized crack detection model of the metal plate under test is used to perform feature mapping and discrimination processing on the acoustic time series image of the metal plate under test, and output the detection result of whether there is a crack in the metal plate under test and the corresponding crack location information. by This represents the crack detection results of the metal plate output by the model: ; by This represents an acoustic time-series image input consisting of T-frame two-dimensional sound pressure amplitude distribution maps arranged in chronological order. by The parameter is The parameterized crack detection model includes a nonlinear mapping structure that combines local spatiotemporal feature extraction with global temporal dependency modeling.
[0016] Compared with existing technologies, the beneficial effects of this invention are reflected in: 1. Significantly Reduced Data Scale and Computational Complexity: This invention segments high-sampling-rate sound pressure time series data and extracts each segment into a two-dimensional sound pressure amplitude distribution map. These segments are then arranged chronologically to form an acoustic time series image, transforming the original "multi-channel long-time sequence-level data" into "image sequence-level data." This achieves an order-of-magnitude reduction in data scale (e.g., optimizing GB-level raw data to KB-level image data). Compared to directly processing the original time series signal, this significantly reduces the computational overhead of data storage, transmission, and training / inference, effectively improving the method's engineering feasibility and real-time processing capabilities.
[0017] 2. Structured Representation and Improved Learnability of Sound Field Information: This invention maps the sound pressure data of discrete sensor sampling points into a two-dimensional distribution map according to the spatial arrangement relationship, forming a structured representation consistent with the physical space; combined with the time dimension arrangement to form a time-series image, the key change laws of the sound field in the spatial and time dimensions are preserved, thus making it more suitable for end-to-end modeling using deep learning and improving the separability of crack features.
[0018] 3. Balancing local details and global temporal dependencies to improve recognition robustness: In the crack detection model, the local spatiotemporal feature extraction module is used to capture the spatiotemporal variation details of the local sound field, while the temporal feature modeling module is used to establish global dependencies across time periods. This enables the model to more effectively utilize multi-frame sound field evolution information and has better robustness and generalization ability for sound field changes under conditions of micro-cracks, weak difference scenes, and multi-frequency excitation.
[0019] 4. Non-contact measurement adapts to engineering scenarios: This invention uses an acoustic array for non-contact measurement, which reduces the dependence on the surface condition and coupling conditions of the measured structure, making it easy to apply in large-area, difficult-to-access or unsuitable sensor attachment conditions. Attached Figure Description
[0020] Figure 1 This is a flowchart illustrating the overall process of acoustic temporal image construction and crack detection in this invention. Figure 2 This is a schematic diagram of the acoustic time-series image construction principle in an embodiment of the present invention; wherein, (a) is a schematic diagram of the acquired multi-channel sound pressure time-series signal and time slice selection, and (b) is a two-dimensional sound pressure amplitude distribution map (pixel matrix) generated after feature extraction and normalization processing. Figure 3 This is a schematic diagram of the mesh division of the metal plate simulation model in this invention, including structural mesh (a), acoustic mesh (b), and reconstructed sound field mesh (c). Figure 4 This is a schematic diagram of acoustic time-series image samples constructed in an embodiment of the present invention, showing a sequence from 10... -4 20×10 seconds-4 A continuous 20-frame sound pressure amplitude cloud map per second; Figure 5 This is a block diagram of the crack detection model combining 3D convolutional neural network (3D CNN) and Transformer proposed in an embodiment of the present invention.
[0021] Figure 6 The T-SNE feature visualization analysis results show that the four types of crack states are clearly distinguishable in the feature space, verifying the effectiveness of this method. Detailed Implementation
[0022] The specific embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.
[0023] See Figure 1 In this embodiment, the method for constructing the acoustic time-series image involves applying excitation to a metal plate and using an acoustic array to non-contactly measure the surface sound field information to obtain sound pressure sampling sequences at multiple spatial locations. The sound pressure sequences are then segmented according to preset time intervals, and combined with the spatial arrangement of the sensors, the sound pressure data for each time interval is mapped into a two-dimensional sound pressure amplitude distribution map. These two-dimensional sound pressure amplitude distribution maps are arranged chronologically to construct an acoustic time-series image, which reflects the spatiotemporal changes in sound pressure distribution. This embodiment transforms a one-dimensional acoustic time series into a two-dimensional image sequence with a spatial topological structure, utilizing computer vision technology to solve the problem of jointly extracting spatiotemporal features in acoustic detection.
[0024] In this embodiment, the method for constructing the acoustic temporal image is performed according to the following steps: Step 1: Acoustic Signal Acquisition. An excitation is applied to a metal plate, and a non-contact measurement of the sound field on the metal plate surface is performed using an acoustic array to obtain sound pressure sampling sequences at multiple spatial locations within the same measurement plane. This embodiment uses finite element simulation for verification: a metal plate model is created using the 3D modeling software SOLIDWORKS, with a length and width of 450 mm and a thickness of 3 mm. Figure 3 The mesh generation of the metal plate simulation model is illustrated, including the structural mesh (a), acoustic mesh (b), and reconstructed sound field mesh (c). The model was meshed using HYPERMESH: the structural mesh used 10mm square meshes; both the acoustic and reconstructed sound field meshes used 10mm triangular meshes. The meshes were imported into LMS Virtual.lab software, and the material was set to steel with the following parameters: density ρ = 7.85 × 10⁻⁶. 3 kg / m 3 Young's modulus E = 2.1 × 10⁻⁶ 5MPa, Poisson's ratio υ=0.3. During the signal acquisition phase, transient excitation is applied to the metal plate, and sound pressure data is acquired using an acoustic array.
[0025] Step 2, Time Segmentation Processing: The sound pressure sampling sequence is segmented according to a preset time interval to obtain sound pressure sampling signals within multiple time segments.
[0026] Time-segmented processing divides the high-temporal-resolution sound pressure sampling sequence into segments according to preset time intervals. The original high-sampling-rate acoustic data is then reconstructed into an acoustic time-series image composed of two-dimensional sound pressure amplitude distribution maps corresponding to multiple time segments. This reduces the data scale for crack detection from time-series level data to image sequence level data. Thus, while preserving key spatiotemporal features of the sound field, it significantly reduces data storage requirements and improves the engineering feasibility and real-time processing capabilities of the crack detection method. The acoustic time-series image can be constructed using two-dimensional grayscale images, multi-channel color images, or multi-feature channel images. In the simulation, the sampling time step is set to 10. -4 Seconds. Traditional processing methods typically require storing tens of thousands of sampling points, resulting in data redundancy and high computational costs. This embodiment uses time segmentation, dividing the data into 10-second intervals. -4 Extract one feature frame per second. For example, select a frame from 10... -4 20×10 seconds -4 In a time span of seconds, only 20 key moments need to be retained to characterize a complete sound wave propagation process, greatly reducing the data size.
[0027] In this embodiment, time segmentation is performed as follows: Let s(t) represent the sound pressure sampling signal, which is the time function corresponding to the sound pressure sampling sequence, with a total duration of T; the total duration T is divided into N equal time intervals, with... The start time of the nth time interval is represented by... Represents the end time of the nth time interval. For the preset time range, Then we have: ; ; The sound pressure data sequence in the nth time period for: , ; in, ; Figure 2Figure (a) illustrates the multi-channel sound pressure level timing signal. The dashed lines in the figure, representing "time slices," correspond to a specific time period or moment. At this moment, data from all channels are synchronously captured as the basic data for constructing the current frame image. The time slice corresponds to a data window within the nth preset time interval. Amplitude features are extracted from the sound pressure sequence of each channel within this window to obtain channel amplitude vectors, which are used as pixel / node inputs for the two-dimensional sound pressure amplitude distribution map of this frame.
[0028] Step 3, Spatial Mapping Processing: Combining the spatial arrangement of each sensor in the acoustic array, the corresponding sound pressure sampling data in each time segment is mapped into a two-dimensional sound pressure amplitude distribution map.
[0029] Figure 2 Figure (b) is for illustrative purposes: when the array has 16 channels arranged in a 4×4 pattern, the channel amplitude features are directly constructed into a 4×4 pixel matrix. For implementations requiring higher spatial resolution, while maintaining the consistency of the sensor spatial topology, interpolation / reconstruction methods are used to map discrete channel features to a denser two-dimensional pixel grid, which is then arranged in chronological order to form a frame sequence. The number of sensor channels is not limited to 16, and the array shape is not limited to a regular 4×4 arrangement; when the number of channels is M and the sensor spatial arrangement is an arbitrary set of two-dimensional coordinates, rasterization mapping, interpolation mapping, or reconstruction-based mapping can be performed according to the coordinate relationship to generate a two-dimensional sound pressure amplitude distribution map consistent with the physical topology.
[0030] In this embodiment, spatial mapping is performed as follows: by The representation is located in spatial coordinates The sound pressure data is obtained from the sound pressure sensor at the location. The corresponding sound pressure amplitude characteristics Obtained from equation (1) or equation (2): (1); (2).
[0031] Step 4, Image Mapping and Temporal Arrangement: Arrange the two-dimensional sound pressure amplitude distribution maps in chronological order to construct an acoustic temporal image, which is composed of multiple two-dimensional sound pressure amplitude distribution maps arranged in chronological order.
[0032] In this embodiment, image mapping and temporal arrangement are performed as follows: sound pressure amplitude characteristics Normalization is performed, and the result is mapped to pixels in a two-dimensional image. Image channel intensity value : , where the mapping function The linear normalized form shown in equation (3) is: (3); in, and These represent the reference minimum and reference maximum values corresponding to the sound pressure amplitude characteristics, respectively. c For image channel indexing, This is a scaling factor used to map the normalized grayscale intensity to a preset image display range; The mapping function f(·) may be a nonlinear mapping, a piecewise mapping, or an adaptive mapping. The reference minimum and reference maximum values are determined based on a single two-dimensional sound pressure amplitude distribution map, multiple time segments, or a preset statistical range.
[0033] like Figure 2 As shown in (b), the normalized values range from 0.00 to 1.00 and are displayed visually in grayscale, with dark colors representing strong sound pressure levels and light colors representing weak sound pressure levels.
[0034] Figure 4 The image shown is a sample of the final constructed acoustic time-series image, illustrating the sequence from 10... -4 20×10 seconds -4 Twenty consecutive sound pressure level (SPL) amplitude contour images were generated per second. To construct the dataset for deep learning, this embodiment set four operating conditions: no cracks, and cracks at three different locations, with crack distances of 225mm, 175mm, and 125mm from the left edge, and a length of 10mm for each crack. The final generated dataset contains four classes of data, with 25 samples in each class. To adapt to the model input, the generated SPL images were randomly cropped to 188×188 pixels, forming a tensor data with dimensions of 3×20@188×188. In this embodiment, the three channels are only used for pseudo-color visualization; the data is mapped through pseudo-color mapping to the R, G, and B channels, without changing the numerical values.
[0035] In this embodiment, a crack detection model is constructed based on acoustic time-series images for crack detection in metal plates. The crack detection model includes: The local spatiotemporal feature extraction module is used to perform joint feature extraction on acoustic time-series images in the spatial and temporal dimensions to obtain feature representations that characterize the local sound field changes. The temporal feature modeling module is used to model the global dependencies of local spatiotemporal features in the time dimension in order to obtain a global temporal feature representation that characterizes the evolution of the sound field. The feature fusion and discrimination module is used to fuse local spatiotemporal features and global temporal features, and output the detection results of metal plate cracks based on the fused features; The crack detection model is trained using an acoustic time-series image dataset consisting of several acoustic time-series images. The acoustic time-series images are used as training samples to input into the crack detection model, and the model parameters are learned and optimized to obtain a parameterized crack detection model for metal plate crack detection. The parameterized crack detection model is used to characterize whether there is a crack in the metal plate and the mapping relationship between the crack location and the acoustic time-series images. The crack detection model performs feature mapping on the acoustic time-series image to obtain a feature representation for crack discrimination. The feature mapping relationship is expressed as follows: (4); In formula (4): This represents an acoustic time-series image composed of T-frame two-dimensional sound pressure amplitude distribution maps, with... This represents a nonlinear mapping function used for global spatiotemporal feature modeling and temporal feature extraction. This represents the spatiotemporal characteristics of the fused sound field.
[0036] Figure 5 The diagram shows the crack detection model, which combines a 3D convolutional neural network and a Transformer. (1) Local spatiotemporal feature extraction module: multi-layer 3D convolution is used. The input data dimension is 3×19@188×188. It should be noted that in order to enhance the convergence speed of the overall model, one frame is randomly discarded, so the number of channels becomes 19. The first layer p1 uses 64 convolutional kernels and outputs 64×19@94×94; the number of channels in subsequent layers p2 to p4 increases to 512, and the feature map size decreases layer by layer. This module can capture the local distortion features of the sound wave front. (2) Temporal feature modeling module: Transformer encoder is used. The feature map is input into the Transformer after adaptive pooling. This module contains 4 heads, a multi-head self-attention layer with a dimension of 128, and a feedforward layer. Its function is to capture the long-range dependence of the sound field evolution and improve the robustness of the identification of small cracks. (3) Feature fusion and discrimination module: The features output by Transformer are processed by global average pooling and fully connected layers, and finally output a 1×4 classification vector.
[0037] The training process of the crack detection model in this embodiment includes: Acoustic time-series image samples containing known crack states and crack locations are obtained, and an acoustic time-series image training dataset with crack annotation information is constructed. The acoustic temporal image training dataset is input into the crack detection model, and the crack detection prediction result is obtained through forward feature mapping. Based on the crack detection prediction results and the corresponding crack annotation information, a model training loss function is constructed, and a parameter optimization algorithm is used to iteratively update the parameters of the crack detection model. When the model training loss function satisfies the preset convergence condition, a parameterized crack detection model for metal plate crack detection is obtained.
[0038] In this embodiment, the dataset is divided into three sets: 100 samples as the training set, 30 as the validation set, and 40 as the test set. The Adam optimizer and cross-entropy loss function are used for training. Experimental results show that the method achieves 100% detection accuracy on the test set. Figure 6 The t-SNE feature visualization analysis shows that the four types of crack states are clearly distinguishable in the feature space, which verifies the effectiveness of the proposed method.
[0039] In this embodiment, the method for detecting cracks in a metal plate is as follows: An acoustic time-series image of the metal plate to be detected is constructed; this image is then input into a trained parametric crack detection model for the metal plate; the parametric crack detection model is used to perform feature mapping and discrimination processing on the acoustic time-series image of the metal plate to be detected, outputting the detection result of whether a crack exists in the metal plate and the corresponding crack location information; This represents the crack detection results of the metal plate output by the model: by This represents an acoustic time-series image input consisting of T-frame two-dimensional sound pressure amplitude distribution maps arranged in chronological order; with The parameter is The parameterized crack detection model includes a nonlinear mapping structure that combines local spatiotemporal feature extraction with global temporal dependency modeling.
[0040] This invention transforms acoustic radiation data acquired by an acoustic array into a time-series acoustic image, and combines 3D convolution and Transformer modules to model the local and global features of the sound field. This method enables high-precision identification of crack locations in metal plates under non-contact measurement conditions. Experimental results show that this embodiment can identify crack locations with a length of not less than 10 mm, exhibiting high generalization performance and robustness, especially under complex sound fields and multi-frequency excitation conditions.
Claims
1. A method for constructing an acoustic temporal image, characterized by: By applying excitation to a metal plate, the acoustic field information of the surface is measured non-contactly using an acoustic array to obtain sound pressure sampling sequences at multiple spatial locations. The sound pressure sequences are segmented according to a preset time interval, and the sound pressure data of each time interval is mapped into a two-dimensional sound pressure amplitude distribution map based on the spatial arrangement of the sensors. The two-dimensional sound pressure amplitude distribution map is arranged in chronological order to construct an acoustic time-series image, which is used to reflect the changes in sound pressure distribution over time and space.
2. The method for constructing an acoustic time-series image according to claim 1, characterized in that: Follow these steps: Step 1, Acoustic signal acquisition: Apply excitation to the metal plate and use an acoustic array to perform non-contact measurement of the sound field on the surface of the metal plate, and obtain the sound pressure sampling sequence at multiple spatial locations within the same measurement plane; Step 2, Time Segmentation Processing: The sound pressure sampling sequence is segmented according to a preset time interval to obtain sound pressure sampling signals within multiple time segments; Step 3, Spatial mapping processing: Combining the spatial arrangement relationship of each sensor in the acoustic array, the corresponding sound pressure sampling data in each time segment is mapped into a two-dimensional sound pressure amplitude distribution map; Step 4, Image Mapping and Temporal Arrangement: Arrange the two-dimensional sound pressure amplitude distribution maps in chronological order to construct an acoustic temporal image, which is composed of multiple two-dimensional sound pressure amplitude distribution maps arranged in chronological order.
3. The method for constructing an acoustic time-series image according to claim 2, characterized in that: Time-segmented processing involves dividing the high-temporal-resolution sound pressure sampling sequence into segments according to preset time intervals. The original high-sampling-rate acoustic data is then reconstructed into an acoustic time-series image composed of two-dimensional sound pressure amplitude distribution maps corresponding to multiple time segments. This reduces the data scale for crack detection from time-series level data to image sequence level data. Consequently, while preserving key spatiotemporal features of the sound field, the data storage requirements are significantly reduced, improving the engineering feasibility and real-time processing capabilities of the crack detection method. The acoustic time-series image can be constructed using two-dimensional grayscale images, multi-channel color images, or multi-feature channel images.
4. The method for constructing an acoustic time-series image according to claim 2, characterized in that: in In step 2, time segmentation is performed as follows: Let s(t) represent the sound pressure sampling signal, which is the time function corresponding to the sound pressure sampling sequence, with a total duration of T; the total duration T is divided into N equal time intervals, with... The start time of the nth time interval is represented by... Represents the end time of the nth time interval. For the preset time range, Then we have: ; ; The sound pressure data sequence in the nth time period for: , ; in, .
5. The method for constructing an acoustic time-series image according to claim 2, characterized in that: in In step 3, spatial mapping is performed as follows: by The representation is located in spatial coordinates The sound pressure data is obtained from the sound pressure sensor at the location. The corresponding sound pressure amplitude characteristics Obtained from equation (1) or equation (2): (1); (2)。 6. The method for constructing an acoustic time-series image according to claim 5, characterized in that: in In step 4, image mapping and temporal arrangement are performed as follows: for sound pressure amplitude features Normalization is performed, and the result is mapped to pixels in a two-dimensional image. Image channel intensity value : , where the mapping function The linear normalized form shown in equation (3) is: (3); in, and These represent the reference minimum and reference maximum values corresponding to the sound pressure amplitude characteristics, respectively. c For image channel indexing, This is a scaling factor used to map the normalized grayscale intensity to a preset image display range; The mapping function f(·) may be a nonlinear mapping, a piecewise mapping, or an adaptive mapping. The reference minimum and reference maximum values are determined based on a single two-dimensional sound pressure amplitude distribution map, multiple time segments, or a preset statistical range.
7. A crack detection model for detecting cracks in metal plates, characterized in that: A crack detection model is constructed based on the acoustic time-series image in claim 1 for crack detection in metal plates. The crack detection model includes: The local spatiotemporal feature extraction module is used to perform joint feature extraction on acoustic time-series images in the spatial and temporal dimensions to obtain feature representations that characterize the local sound field changes. The temporal feature modeling module is used to model the global dependencies of local spatiotemporal features in the time dimension in order to obtain a global temporal feature representation that characterizes the evolution of the sound field. The feature fusion and discrimination module is used to fuse local spatiotemporal features and global temporal features, and output the detection results of metal plate cracks based on the fused features; The crack detection model is trained using an acoustic time-series image dataset consisting of several acoustic time-series images. The acoustic time-series images are used as training samples to input into the crack detection model, and the model parameters are learned and optimized to obtain a parameterized crack detection model for metal plate crack detection. The parameterized crack detection model is used to characterize whether there is a crack in the metal plate and the mapping relationship between the crack location and the acoustic time-series images. The crack detection model performs feature mapping on the acoustic time-series image to obtain a feature representation for crack discrimination. The feature mapping relationship is expressed as follows: (4); In formula (4): This represents an acoustic time-series image composed of T-frame two-dimensional sound pressure amplitude distribution maps, with... This represents a nonlinear mapping function used for global spatiotemporal feature modeling and temporal feature extraction. This represents the spatiotemporal characteristics of the fused sound field.
8. The crack detection model for detecting cracks in metal plates according to claim 7, characterized in that: The training process for the crack detection model includes: Acoustic time-series image samples containing known crack states and crack locations are obtained, and an acoustic time-series image training dataset with crack annotation information is constructed. The acoustic temporal image training dataset is input into the crack detection model, and the crack detection prediction result is obtained through forward feature mapping. Based on the crack detection prediction results and the corresponding crack annotation information, a model training loss function is constructed, and a parameter optimization algorithm is used to iteratively update the parameters of the crack detection model. When the model training loss function satisfies the preset convergence condition, a parameterized crack detection model for metal plate crack detection is obtained.
9. A method for detecting cracks in a metal plate, characterized in that: An acoustic time-series image of the metal plate to be tested is constructed according to the method of claim 1, which is the acoustic time-series image of the metal plate to be tested; The acoustic time-series image of the metal plate to be detected is input into the trained parameterized crack detection model of the metal plate to be detected. The parameterized crack detection model of the metal plate under test is used to perform feature mapping and discrimination processing on the acoustic time series image of the metal plate under test, and output the detection result of whether there is a crack in the metal plate under test and the corresponding crack location information. by This represents the crack detection results of the metal plate output by the model: ; by This represents an acoustic time-series image input consisting of T-frame two-dimensional sound pressure amplitude distribution maps arranged in chronological order. by The parameter is The parameterized crack detection model includes a nonlinear mapping structure that combines local spatiotemporal feature extraction with global temporal dependency modeling.