A method and device for predicting acoustic quality in shipbuilding based on working condition perception and multi-scale attention autoencoder

By adopting an acoustic quality prediction method based on working condition perception and multi-scale attention autoencoder, the problems of unbalanced feature extraction and non-robust evaluation in acoustic quality detection during ship construction are solved, and high accuracy and high fault tolerance acoustic quality assessment are achieved.

CN122364631APending Publication Date: 2026-07-10HARBIN ENG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HARBIN ENG UNIV
Filing Date
2026-04-03
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

Existing acoustic quality testing methods suffer from problems such as single feature extraction scale, unbalanced multi-channel signal processing, lack of adaptability to operating conditions, and non-robust evaluation indicators during ship construction. These problems result in poor stability of test results and make it difficult to achieve high-accuracy non-destructive evaluation in complex environments.

Method used

An acoustic quality prediction method based on condition perception and multi-scale attention autoencoder is adopted. By introducing multi-scale parallel convolution and spatial-channel attention mechanism, combined with kernel density estimation, a highly robust assessment of acoustic signals is achieved.

Benefits of technology

It effectively avoids misjudgment due to fluctuations in operating conditions, improves detection accuracy, and achieves acoustic quality assessment with high fault tolerance and probabilistic statistical significance in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122364631A_ABST
    Figure CN122364631A_ABST
Patent Text Reader

Abstract

This invention discloses an acoustic quality prediction method and device based on working condition perception and multi-scale attention autoencoder in shipbuilding, belonging to the field of non-destructive testing and intelligent evaluation technology in shipbuilding. The method acquires acoustic signals and corresponding working condition parameters of different ship structures under normal operation, constructing a training dataset. Then, a multi-scale conditional attention autoencoder model is used to output the predicted acoustic signals. Under normal operation, the mean square error between the predicted and sample acoustic signals is calculated to form a reconstruction error set. A non-parametric kernel density estimation is introduced to fit the reconstruction error, obtaining a baseline probability density function. The acoustic signal to be evaluated is input into the trained model. The average mean square error of all the sound signals to be evaluated and predicted under this working condition is calculated, and the cumulative distribution function value is estimated based on the density function to obtain the acoustic quality probability score. This invention provides the output quality score with rigorous probabilistic support and higher error tolerance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of non-destructive testing and intelligent evaluation technology in shipbuilding, specifically relating to an acoustic quality prediction method and device based on working condition perception and multi-scale attention autoencoder in shipbuilding. Background Technology

[0002] During shipbuilding, welding, assembly, and outfitting processes generate a large amount of structural vibration and acoustic signals. These signals contain crucial information such as structural integrity, weld stability, and process consistency. By analyzing these acoustic signals, the quality of hull construction can be assessed without damaging the structure, making it an important non-destructive testing method.

[0003] However, traditional acoustic quality testing methods mainly rely on human experience or discrimination models based on statistical characteristics. For example, commonly used features such as root mean square (RMS), peak factor, spectral kurtosis, and spectral centroid are typically used to measure the overall change in acoustic energy or frequency distribution. These methods are computationally simple, but can only reflect the macroscopic statistical characteristics of the signal and are difficult to characterize complex temporal evolution. In addition, these features are highly sensitive to environmental noise and fluctuations under different operating conditions, resulting in poor stability of the test results and making it difficult to achieve robust acoustic quality assessment in complex shipbuilding environments.

[0004] To overcome the limitations of manual feature extraction, some studies have attempted to introduce dimensionality reduction methods such as principal component analysis and independent component analysis to extract the main variation patterns of acoustic signals. However, these linear dimensionality reduction models assume a linear relationship between signal features and cannot effectively capture the large number of nonlinear dynamic features present in acoustic signals, thus limiting their recognition capabilities in multi-condition and complex structural vibration scenarios.

[0005] In recent years, the rapid development of deep learning technology has provided new avenues for unsupervised modeling of acoustic signals. Reconstruction learning methods based on autoencoders (AEs) can automatically extract features through end-to-end training, eliminating the need for manual labeling. The basic idea is to train an encoder-decoder structure so that the model's output signal is as close as possible to the input signal. By analyzing the reconstruction errors of data under different operating conditions, label-free detection of acoustic pattern differences can be achieved. Although subsequent research introduced one-dimensional convolutional autoencoders (Conv1d-AEs) to process temporal data, existing convolutional autoencoder methods still have the following significant limitations in complex shipbuilding scenarios: First, the feature extraction scale is singular, making it difficult to comprehensively capture complex frequency band features. Acoustic signals generated during ship welding and assembly possess both high-frequency transient impacts (such as spatter and microcrack initiation) and low-frequency continuous resonance characteristics. Traditional Conv1d-AEs typically use a single convolutional kernel of fixed size, limiting the receptive field and making it difficult to simultaneously consider the multi-scale temporal evolution patterns across a wide frequency band. Second, they neglect the spatial topology and channel quality differences of multi-sensor arrays. In practical multi-channel acquisition (such as a 4-channel sensor array), the signal-to-noise ratio (SNR) of each channel varies greatly due to the influence of installation location, transmission path, and ambient background noise. Existing models typically treat multi-channel data equally, lacking the ability to adaptively focus on high-quality signal sources, making the reconstruction process susceptible to interference from noisy channels. Furthermore, they lack adaptability to complex operating condition fluctuations. Changes in shipbuilding process parameters (such as welding current, speed, and fixture type) can cause significant drift in the normal acoustic baseline distribution. Traditional unsupervised models mix all operating condition data within a single potential space, easily misjudging signal fluctuations caused by normal operating condition switching as structural quality anomalies, resulting in extremely high false alarm rates in multi-condition scenarios. Finally, quality assessment indicators lack statistical robustness. Existing acoustic quality assessments often use simple mean square error combined with linear normalization of maximum and minimum values ​​for scoring. This approach is highly sensitive to extremely rare transient outliers, and the scoring results lack rigorous probabilistic statistical significance, making it difficult to provide reliable confidence assessments in industrial settings. Summary of the Invention

[0006] To address the shortcomings of existing technologies, this invention proposes an acoustic quality prediction method based on condition perception and a multi-scale attention autoencoder. This method not only overcomes the bottleneck of feature extraction based on single scale and equal channels, but also achieves a new paradigm for label-free acoustic quality assessment with high robustness and probabilistic statistical significance through conditional constraints and kernel density estimation.

[0007] This invention provides an acoustic quality prediction method based on condition perception and multi-scale attention autoencoder, comprising the following steps:

[0008] Step 1: Obtain acoustic signals and corresponding operating parameters of different hull structures under normal operating conditions, and construct a training dataset;

[0009] Step 2: Construct and train a multi-scale conditional attention autoencoder model; the multi-scale conditional attention autoencoder model introduces a one-dimensional space-channel attention mechanism module on the basis of the encoder and decoder to perform global temporal feature weighting on the input feature tensor; the encoder uses a multi-scale parallel temporal convolution module for feature extraction, and uses a depth downsampling network for layer-by-layer convolution and compression of low-dimensional latent vectors;

[0010] Step 3: After training is completed, calculate the mean square error between the predicted acoustic signal output by the multi-scale conditional attention autoencoder model and the acoustic signal of the sample under normal operating conditions, forming a set of reconstruction errors under normal operating conditions; introduce non-parametric kernel density estimation to fit the reconstruction error and obtain the baseline probability density function.

[0011] Step 4: Input the acoustic signal to be evaluated into the trained multi-scale conditional attention autoencoder model and output the predicted acoustic signal; calculate the average root mean square error of all the sound signals to be evaluated and the predicted sound signals under this condition, and estimate its cumulative distribution function value according to the baseline probability density function to obtain the acoustic quality probability score.

[0012] Furthermore, the operating parameters include welding current amplitude, welding moving speed, workpiece fixture type, and ambient background noise level.

[0013] Further, in step 1, the acoustic signal data is preprocessed by standardization, and the standardized continuous signal is segmented using a sliding window. Features are extracted from each segment to obtain the corresponding feature tensor. The operating condition parameters are mapped to discrete classification labels or continuous parameter matrices, and a one-dimensional operating condition vector aligned with each acoustic window is constructed. The feature tensor and the operating condition vector are combined to construct a training dataset.

[0014] Furthermore, in step 2, constructing and training the multi-scale conditional attention autoencoder model includes the following steps:

[0015] Step 2.1: The input feature tensor is globally averaged through the spatial-channel attention mechanism module. Then, a multilayer perceptron and a sigmoid activation function are used to generate the weight coefficients for each channel. Finally, the weight coefficients are multiplied with the feature tensor channel by channel.

[0016] Step 2.2: Use three parallel one-dimensional convolution kernels to extract feature tensors through the multi-scale temporal convolution module to obtain multi-scale feature tensors;

[0017] Step 2.3: Input the multi-scale feature tensor into a depth downsampling network consisting of four one-dimensional convolutional layers and the ReLU nonlinear activation function;

[0018] Step 2.4: Flatten the feature tensor output by the deep downsampling network and map it to a low-dimensional latent space vector through a fully connected layer;

[0019] Step 2.5: The working condition vector and the latent space vector in the dataset are concatenated, mapped and reconstructed into a three-dimensional feature tensor through a fully connected layer, and used as the input of the decoder; the decoder is constructed with four symmetric one-dimensional deconvolution layers. Under the explicit constraint of the working condition vector, the time dimension of the signal is restored by linear upsampling layer by layer. The number of channels decreases in turn, and the spatial scale and structure of the signal are restored layer by layer. Finally, the reconstructed acoustic signal with the same size as the input signal is output.

[0020] Further, in step 2.1, the spatial-channel attention mechanism module operates as follows:

[0021]

[0022]

[0023] in, Input feature tensor The weights; The Sigmoid activation function is used; MLP is used for multilayer perceptron; AvgPool is used for average pooling. For the input feature tensor, This is the weighted feature map; This involves multiplying each channel sequentially.

[0024] Furthermore, in step 2, the multi-scale conditional attention autoencoder model is trained using the mean squared error loss function, and training is stopped and parameters are retained when the maximum number of training iterations is reached or the loss function converges.

[0025] Furthermore, in step 3, the baseline probability density function for:

[0026]

[0027] in, This represents the total number of samples under normal operating conditions. For the first The mean squared error of a normal sample; For smoothing kernel function; To smooth bandwidth.

[0028] Further, in step 4, the cumulative distribution function value is:

[0029]

[0030] in, This is the average of the standard deviations of all samples.

[0031] Furthermore, the acoustic quality probability score for:

[0032] .

[0033] Furthermore, in step 4, if the average root mean square error of the acoustic signal to be evaluated is less than or equal to the average root mean square error under normal operating conditions, then the acoustic quality probability score is directly assigned. A value of 1 corresponds to excellent processing quality; if the average mean square error of the evaluated acoustic signal is greater than the average mean square error under normal operating conditions, an acoustic quality probability score is calculated. , The smaller the value, the worse the corresponding processing quality.

[0034] The present invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the acoustic quality prediction method based on condition perception and multi-scale attention autoencoder described above.

[0035] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the acoustic quality prediction method based on condition perception and multi-scale attention autoencoder described above.

[0036] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the acoustic quality prediction method based on condition perception and multi-scale attention autoencoder described above.

[0037] The beneficial effects of this invention are as follows:

[0038] This invention injects operating condition parameters into the conditional autoencoder architecture, enabling the model to dynamically adjust and reconstruct expectations based on different environments. This effectively avoids misjudging normal operating condition fluctuations as acoustic anomalies, significantly improving detection accuracy in complex shipbuilding environments. The introduction of multi-scale parallel convolution overcomes the limitations of a single receptive field, achieving comprehensive capture of broadband acoustic features. Simultaneously, a spatial-channel attention mechanism adaptively allocates sensor weights, shielding edge channel interference noise from the model's underlying layers. Replacing the traditional linear extremum mapping formula with a probability confidence score based on kernel density estimation effectively eliminates the interference of local extreme outliers on the overall score, giving the output quality score rigorous probabilistic support and higher fault tolerance. This invention proposes an acoustic quality prediction method based on baseline operating conditions under normal conditions, capable of selecting superior processing quality under different demanding environments. Attached Figure Description

[0039] Figure 1 This is a flowchart of the acoustic quality prediction method based on working condition perception and multi-scale attention autoencoder of the present invention.

[0040] Figure 2 This is a probability distribution diagram of acoustic feature reconstruction error kernel density estimation under different operating conditions in the embodiment;

[0041] Figure 3 The above is a comparison diagram of the temporal waveform reconstruction of acoustic signals based on a multi-scale attention network in the embodiments.

[0042] Figure 4 This is a biaxial comparison chart of the average reconstruction error and the KDE acoustic quality probability score trend for each working condition in the embodiment. Detailed Implementation

[0043] The present invention will be further described below with reference to the accompanying drawings. The embodiments are only used to explain the present invention and are not intended to limit the scope of protection of the present invention.

[0044] This invention discloses an acoustic quality prediction method based on working condition perception and a multi-scale attention autoencoder, used for self-supervised learning and robust unlabeled quality assessment of acoustic signals under complex and variable working conditions during shipbuilding. By deeply fusing prior knowledge of working conditions, spatial topological relationships, and multi-band temporal features, this invention can achieve intelligent identification of structural acoustic states and statistical quantitative evaluation of processing quality, thereby replacing traditional detection methods that rely on experience and are susceptible to noise and fluctuations in working conditions.

[0045] The core idea of ​​this invention is to utilize a multi-scale conditional attention autoencoder for feature extraction and conditional reconstruction learning of acoustic signals. Addressing the shortcomings of traditional autoencoders, which are prone to misjudgment under various operating conditions, this invention injects operating condition parameters as conditional priors into the model's latent space. Considering the extremely wide bandwidth of acoustic signals and the varying spatial weights of sensors, a multi-scale parallel convolution and spatial attention mechanism is introduced. Finally, abandoning the traditional linear normalization scoring, kernel density estimation is used to transform the reconstruction error into a statistically significant probability confidence level, achieving continuous and highly fault-tolerant acoustic quality scoring.

[0046] The specific technical solution is as follows:

[0047] Step 1: Acquisition and preprocessing of multi-channel acoustic data and operating condition metadata.

[0048] Multi-channel acceleration or acoustic emission sensors are deployed at key locations on the hull structure to collect vibration or acoustic signals under different processing conditions, and simultaneously record corresponding operating parameters (such as welding current, speed, fixture status, etc.). The original signals are denoised, normalized, and segmented using a sliding window to obtain time-series samples of shape (C, L). Simultaneously, discrete and continuous operating parameters are vectorized and encoded to generate corresponding one-dimensional condition vectors. .

[0049] Step 2: Construction and training of the multi-scale conditional attention model.

[0050] A multi-scale conditional attention autoencoder model is constructed. First, a one-dimensional spatial-channel attention mechanism is introduced at the network front end to adaptively weight the input multi-channel signals, enhancing high-quality sensor signals and suppressing background noise. Second, the encoder uses multi-scale parallel one-dimensional convolutional kernels (e.g., sizes 3, 7, and 15) to simultaneously extract high-frequency transient impact (e.g., splashing, cracks) and low-frequency continuous resonance features, and compresses them into low-dimensional latent vectors. Subsequently, acoustic feature vectors are placed in the latent space. With operating condition vector The data is then stitched together and fused; finally, the decoder performs deconvolution reconstruction based on the fused features. The model uses the mean squared error between the input and the reconstructed output as the loss function, and automatically learns the normal acoustic distribution pattern under specific operating conditions from a large amount of data.

[0051] Step 3: Probabilistic quality assessment based on kernel density estimation (KDE).

[0052] The trained model is applied to various operating condition data to obtain reconstruction errors. The mean squared error set under normal baseline operating conditions is extracted, and its true probability density function is fitted using kernel density estimation to construct a baseline distribution of normal acoustic modes. For the operating condition sample to be evaluated, the cumulative distribution probability of its average reconstruction error under this baseline distribution is calculated, i.e., the probability that the error falls within the normal range, and a continuous acoustic quality probability score is defined accordingly. This score directly reflects the statistical confidence level that the signal belongs to a normal acoustic pattern.

[0053] Step 4: Outputting Results and Multidimensional Visualization.

[0054] The system automatically outputs the average reconstruction error and the corresponding acoustic quality probability score (0~1 range) for each working condition, and generates waveform time-domain reconstruction comparison chart, reconstruction error KDE probability density distribution chart and quality score trend chart, which intuitively and quantitatively show the differences in acoustic evolution and structural integrity under different working conditions.

[0055] Example 1

[0056] An acoustic quality prediction method based on condition perception and multi-scale attention autoencoder includes the following steps:

[0057] Step 1: Multi-channel acoustic signal acquisition and construction of operating condition metadata

[0058] During shipboard welding and assembly processes, structural vibrations and acoustic signals are significantly affected by physical space transmission attenuation and the complex environment. This step aims to construct a high-quality acoustic sample set containing rich spatial topological information and clear operating condition labels.

[0059] Step 1.1: Sensor Spatial Deployment and Multi-Channel Synchronous Acquisition. First, select representative locations at both ends of the hull structure under test (e.g., a typical weld), the welding table / clamp, and adjacent plates, and rigidly fix four accelerometers (denoted as AI1-01 to AI1-04). To ensure the subsequent model can learn the spatial importance differences of the sensors, the three-dimensional coordinates, relative position from the sound source center, installation height, and calibration parameters (static zero point, sensitivity gain, and bias) of each sensor must be recorded in detail as spatial metadata. All sampling devices are set to a uniform 1kHz sampling frequency for synchronous acquisition of multi-channel vibration and acoustic signals. A unified hardware trigger or network time synchronization protocol (PTP / NTP) is used to ensure accurate time alignment of multiple channels; if independent acquisition devices are used, timestamps must be recorded, and time delay and clock drift must be corrected using cross-correlation in the post-processing stage.

[0060] Step 1.2: Recording Working Condition Metadata and Acquiring Data from Multiple Scenarios. To reflect different structural processing states and achieve working condition awareness of the model, this embodiment sets four typical processing working conditions, denoted as A, B, C, and D. For each working condition, its detailed working condition parameters (including but not limited to: welding current amplitude, welding movement speed, workpiece fixture type, and ambient background noise level) must be clearly defined and quantified. This parameter information is written as the "prior metadata" of that working condition into the header information of each acquisition, ensuring data traceability and providing a basis for subsequent construction of condition vectors. In batch acquisition management, at least approximately 10,000 samples are collected for each working condition. Real-time integrity monitoring is performed during acquisition, and severely distorted data such as channel loss, sensor saturation, or abnormal offset (overflow) is promptly labeled or removed.

[0061] Step 1.3: Multimodal Data Alignment and File Organization. Each acquisition generates an independent data file, with the filename including the operating condition category, acquisition date, and batch number. The file synchronously stores the timestamp and the original four-channel calibrated acceleration physical quantities column-by-column. Simultaneously, the corresponding operating condition metadata is strictly aligned and bound to the time-series data in the form of an external configuration file or header file. This dual-track storage mode of "time-series signal + operating condition parameters" provides complete data support for the subsequent joint input of multi-scale conditional autoencoders.

[0062] Step 2: Data Preprocessing and Vectorization of Operating Conditions

[0063] To ensure the consistency of amplitudes between different sensor channels, suppress environmental noise, and provide standardized spatiotemporal and prior inputs for multi-scale conditional autoencoders, this embodiment performs unified joint preprocessing on the acquired multi-channel acoustic signals and operating condition metadata.

[0064] Step 2.1: Outlier Removal and Standardization. First, a batch signal integrity check is performed on all acquired raw acoustic sequences to remove outlier time periods containing NaN, infinite values, or prolonged sensor saturation. Then, the global mean and standard deviation are calculated for each channel on the training set, and the data is standardized to zero mean and unit variance. This not only eliminates inconsistencies in dimensions caused by differences in sensor installation location or sensitivity but also accelerates the convergence of subsequent deep learning networks. The standardized parameters from the training set are strictly reused during the inference phase to maintain absolute consistency in data scale.

[0065] Step 2.2: Sliding Window Segmentation and Condition Label Alignment. To improve modeling efficiency and the learnability of temporal features, the standardized continuous signal is segmented using a sliding window method. The window length is set to 512 sampling points (corresponding to 0.512s physical duration at a 1kHz sampling rate), and the sliding step size is 10 sampling points. Thus, the original one-dimensional long sequence is segmented into spatiotemporal tensor samples with dimensions of (C, L), where C=4 represents four sensor channels and L=512 represents the time step. Simultaneously, the prior metadata of the operating conditions recorded in Step 1 (operating conditions A, B, C, D, where operating condition A is the standard operating state) is extracted and mapped to discrete classification labels or continuous parameter matrices, constructing a one-dimensional condition vector strictly aligned with each acoustic window sample. These processed tensors, together with the conditional vectors, constitute a unified sample set for self-supervised learning of the model.

[0066] Step 3: Design of Multi-Scale Conditional Attention Autoencoder Network

[0067] This invention overcomes the limitation of traditional single-core one-dimensional convolution in handling wide-bandwidth features, and innovatively constructs a conditional autoencoder architecture that integrates spatial attention and multi-scale feature extraction. The network consists of four core modules: attention feature weighting, multi-scale encoding, conditional injection, and deconvolution reconstruction.

[0068] Step 3.1: One-Dimensional Spatial-Channel Attention Mechanism. To address the issue of inconsistent signal quality from multiple sensors in complex structures, the model first introduces a one-dimensional spatial-channel attention module at the input. This module receives a multi-channel tensor of shape (4, 512), extracts the global spatial energy distribution of each channel through global adaptive average pooling, and then inputs it into a weight generation network composed of a multilayer perceptron and a sigmoid activation function. This module can adaptively output four channels within a certain range. Attention weights between The input tensor is multiplied by the corresponding weight for each channel, thereby dynamically amplifying sensor signals with high signal-to-noise ratio that are close to abnormal sound sources at the bottom layer of the network, while suppressing edge or severely interfered channel signals.

[0069] Step 3.2: Multi-scale Parallel Temporal Convolution. The attention-weighted feature maps are fed into the multi-scale encoder. To simultaneously capture high-frequency transient impacts (such as splashes and fractures) and low-frequency sustained resonances (such as welding thermal deformation) in the acoustic signals of shipbuilding, the encoder's first layer uses three parallel one-dimensional convolutional kernels for feature extraction. In this embodiment, the convolutional kernel sizes are set to 3, 7, and 15, with corresponding padding of 1, 3, and 7 to maintain consistent sequence lengths. The three convolutional branches with different receptive fields each output feature maps with 16 channels, which are then concatenated along the channel dimension to generate a wideband multi-scale feature tensor with 48 channels, thereby significantly improving the model's ability to characterize complex nonlinear temporal evolution.

[0070] Step 3.3: Depth Downsampling Encoding and Latent Space Mapping. The multi-scale feature tensor is then fed into a depth downsampling network consisting of four layers of one-dimensional convolutions (kernel size 4, stride 2) and the ReLU nonlinear activation function. Through layer-by-layer convolution, the number of channels is successively expanded to 32, 64, 128, and 128, while the time series length is progressively compressed to 1 / 16 of the original length (i.e., 32 time steps). The downsampled feature tensor is flattened and mapped to a low-dimensional latent space vector through a fully connected layer. This vector highly condenses the pure acoustic mode distribution of the current input signal.

[0071] Step 3.4: Operating Condition Injection and Decoder Reconstruction. This module is crucial for achieving robust detection across operating conditions in this invention. During the decoding stage, the model does not rely solely on acoustic latent vectors. Instead, it uses the working condition vector constructed in step two. With latent vector The features are concatenated to generate a joint feature vector that incorporates prior knowledge. Subsequently, The initial 3D tensor shape is projected back to the decoder through a fully connected layer. The decoder consists of four symmetrical one-dimensional deconvolutional layers. The deconvolutional layers operate under specific conditions. Under explicit constraints, the time dimension is gradually recovered through linear upsampling (the number of channels in each layer decreases sequentially to 128, 64, 32, and then to 4), ultimately outputting a reconstructed acoustic signal with the same shape as the original input. Through this conditional reconstruction, the model can clearly identify whether the current signal fluctuation belongs to a change in normal operating conditions or an abnormality in processing quality, thus fundamentally solving the problem of high false alarm rate in traditional models under multiple operating conditions.

[0072] Step 4: Self-supervised training of the model and construction of the normal baseline distribution

[0073] This step aims to optimize the network parameters of the multi-scale conditional attention autoencoder through self-supervised reconstruction learning using a large amount of unlabeled data. After training, kernel density estimation (KDE) is used to extract the acoustic error probability density features under normal processing conditions, and a statistical benchmark is constructed for subsequent quality assessment.

[0074] Step 4.1: Network Parameter Optimization and Self-Supervised Training. First, the unified sample set constructed in Step 2 is divided into a training set and a validation set at a ratio of 9:1, and stratified sampling is used to ensure a balanced proportion of samples for each working condition. A data loader is constructed, and the batch size is set to 128 to fully utilize the hardware memory and ensure the smoothness of gradient descent.

[0075] In terms of training configuration, the model uses the Adam optimizer for parameter updates and sets an initial learning rate. And introduce a weight decay coefficient To prevent overfitting, the loss function uses mean squared error and is calculated based on the multi-channel raw input signal. With conditional reconstruction output signal Reconstruction loss between time steps and channels:

[0076]

[0077] The training process continuously sets the number of rounds and introduces an early stopping strategy: if the loss on the validation set does not show a significant decrease for 30 consecutive rounds, the training is terminated early, and the model weights with the minimum loss on the validation set are saved as the final inference model.

[0078] Step 4.2: Fitting the baseline probability density function based on kernel density estimation. After the model training converges, extract normal operating condition data known to be in optimal or standard processing conditions, input them into the trained model to obtain the reconstructed signal, and calculate the mean square error of each sample window to form a set of reconstruction errors for normal conditions. .

[0079] To overcome the limitations of traditional parameter estimation methods (such as assuming the error follows a normal distribution), this invention introduces non-parametric kernel density estimation to fit the data. The probability density function, denoted as... ;

[0080]

[0081] in, This represents the total number of samples under normal operating conditions. For the first The mean squared error of a normal sample; The kernel function is a smoothing kernel (preferably a Gaussian kernel in this embodiment); To smooth the bandwidth, this embodiment uses the Silverman rule of thumb to adaptively calculate the optimal bandwidth. The fitted bandwidth is... That is, the statistical benchmark distribution used to evaluate the acoustic quality under all operating conditions.

[0082] Step 5: Acoustic Quality Probabilistic Assessment and Multidimensional Visualization Output

[0083] This step utilizes the trained model and the fitted KDE baseline distribution to perform rigorous probabilistic quality quantification of acoustic signals under unknown or complex operating conditions, and generates intuitive visualization charts for engineers to make decisions.

[0084] Step 5.1: Test Set Reconstruction and Reconstruction Error Calculation. Input the time-series window samples of the work condition to be evaluated and their corresponding work condition vectors in batches into the multi-scale conditional attention autoencoder with frozen weights. Calculate the mean square error of each window sample and calculate the overall average error of all samples under that work condition. Thanks to the pre-implementation condition injection mechanism in the model, if the fluctuations in the input signal are normal process variations under this operating condition, the model can still reconstruct the data with high accuracy. Keep the signal low; if the signal contains hidden structural defects or deteriorated processing quality, the model cannot be reconstructed using normal priors. It will increase significantly.

[0085] Step 5.2: Calculation of Continuous Probability Quality Score. Abandoning the traditional maximum-minimum linear normalization formula, which is highly susceptible to interference from extreme outliers, this invention calculates the average reconstruction error of the target working condition. Substitute the probability density function constructed in step 4 In this process, calculate its cumulative distribution function value (which can be compared by calculating the probability density when the root mean square deviation is the same under the same operating conditions), that is, the probability that the test error falls within the normal distribution interval (the sample root mean square deviation is less than the mean root mean square deviation):

[0086]

[0087] Subsequently, a continuous acoustic quality probability score with rigorous statistical confidence was defined. for:

[0088]

[0089] To align with engineering intuition, when testing errors... When the average error is less than or equal to that of the normal reference, it is directly assigned This rating The range of values ​​is strictly defined within Interval. A higher value indicates a smaller signal reconstruction error and a better match to the expected acoustic pattern under specific operating conditions, corresponding to superior processing quality; conversely, a lower value indicates a lower signal reconstruction error and a lower signal quality. The lower the value, the higher the confidence level of its deviation from the normal distribution, indicating a substantial degradation in processing quality.

[0090] The average reconstruction error and acoustic quality score obtained under four operating conditions in this embodiment are shown in Table 1:

[0091]

[0092] The results show that, with condition A as the baseline, condition C has the highest acoustic quality, followed by condition B, while condition D has the lowest acoustic quality. The model has a good ability to distinguish the acoustic differences between different conditions.

[0093] In particular, in some preferred embodiments of the present invention, a computer device is also provided, including a memory and a processor and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the acoustic quality prediction method based on condition perception and multi-scale attention autoencoder described in any of the above embodiments.

[0094] In some other preferred embodiments of the present invention, a computer-readable storage medium is also provided, on which a computer program / instruction is stored, wherein when the computer program is executed by a processor, the steps of the acoustic quality prediction method based on condition perception and multi-scale attention autoencoder described in any of the above embodiments are implemented.

[0095] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above embodiments of the acoustic quality prediction method based on working condition perception and multi-scale attention autoencoder, which will not be repeated here.

[0096] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.

[0097] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0098] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0099] The above description is merely an embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural or procedural changes made using the content of this invention and its drawings, or any direct or indirect applications in related technical fields, are similarly included within the patent protection scope of the present invention.

Claims

1. A method for predicting acoustic quality in shipbuilding based on condition perception and multi-scale attention autoencoders, characterized in that, Includes the following steps: Step 1: Obtain acoustic signals and corresponding operating parameters of different hull structures under normal operating conditions, and construct a training dataset; Step 2: Construct and train a multi-scale conditional attention autoencoder model; the multi-scale conditional attention autoencoder model introduces a one-dimensional space-channel attention mechanism module on the basis of the encoder and decoder to perform global temporal feature weighting on the input feature tensor; The encoder uses a multi-scale parallel temporal convolution module for feature extraction, and employs a depth downsampling network for layer-by-layer convolution and compression of low-dimensional latent vectors. Step 3: After training is completed, calculate the mean square error between the predicted acoustic signal output by the multi-scale conditional attention autoencoder model and the acoustic signal of the sample under normal operating conditions, forming a set of reconstruction errors under normal operating conditions; introduce non-parametric kernel density estimation to fit the reconstruction error and obtain the baseline probability density function. Step 4: Input the acoustic signal to be evaluated into the trained multi-scale conditional attention autoencoder model and output the predicted acoustic signal; calculate the average root mean square error of all the sound signals to be evaluated and the predicted sound signals under this condition, and estimate its cumulative distribution function value according to the baseline probability density function to obtain the acoustic quality probability score.

2. The acoustic quality prediction method based on condition perception and multi-scale attention autoencoder according to claim 1, characterized in that, In step 1, the acoustic signal data is standardized and preprocessed, and the standardized continuous signal is segmented using a sliding window. Features are extracted from each segment to obtain the corresponding feature tensor. The operating condition parameters are mapped to discrete classification labels or continuous parameter matrices, and a one-dimensional operating condition vector aligned with each acoustic window is constructed. The feature tensor and the operating condition vector are combined to construct a training dataset.

3. The acoustic quality prediction method based on condition perception and multi-scale attention autoencoder according to claim 2, characterized in that, Step 2, in which the multi-scale conditional attention autoencoder model is constructed and trained, includes the following steps: Step 2.1: The input feature tensor is globally averaged through the spatial-channel attention mechanism module. Then, a multilayer perceptron and a sigmoid activation function are used to generate the weight coefficients for each channel. Finally, the weight coefficients are multiplied with the feature tensor channel by channel. Step 2.2: Use three parallel one-dimensional convolution kernels to extract feature tensors through the multi-scale temporal convolution module to obtain multi-scale feature tensors; Step 2.3: Input the multi-scale feature tensor into a depth downsampling network consisting of four one-dimensional convolutional layers and the ReLU nonlinear activation function; Step 2.4: Flatten the feature tensor output by the deep downsampling network and map it to a low-dimensional latent space vector through a fully connected layer; Step 2.5: The working condition vector and the latent space vector in the dataset are concatenated, mapped and reconstructed into a three-dimensional feature tensor through a fully connected layer, and used as the input of the decoder; the decoder is constructed with four symmetric one-dimensional deconvolution layers. Under the explicit constraint of the working condition vector, the time dimension of the signal is restored by linear upsampling layer by layer. The number of channels decreases in turn, and the spatial scale and structure of the signal are restored layer by layer. Finally, the reconstructed acoustic signal with the same size as the input signal is output.

4. The acoustic quality prediction method based on condition perception and multi-scale attention autoencoder according to claim 3, characterized in that, In step 2.1, the spatial-channel attention mechanism module operates as follows: in, Input feature tensor The weights; The Sigmoid activation function is used; MLP is used for multilayer perceptron; AvgPool is used for average pooling. For the input feature tensor, This is the weighted feature map; This involves multiplying each channel sequentially.

5. The acoustic quality prediction method based on condition perception and multi-scale attention autoencoder according to claim 1, characterized in that, In step 2, the multi-scale conditional attention autoencoder model is trained using the mean squared error loss function. Training stops and parameters are retained when the maximum number of training iterations is reached or the loss function converges.

6. The acoustic quality prediction method based on condition perception and multi-scale attention autoencoder according to claim 1, characterized in that, In step 3, the baseline probability density function for: in, This represents the total number of samples under normal operating conditions. For the first The mean squared error of a normal sample; For smoothing kernel function; To smooth bandwidth.

7. The acoustic quality prediction method based on condition perception and multi-scale attention autoencoder according to claim 1, characterized in that, In step 4, the cumulative distribution function value is: in, This is the average of the standard deviations of all samples.

8. The acoustic quality prediction method based on condition perception and multi-scale attention autoencoder according to claim 7, characterized in that, The acoustic quality probability score for: 。 9. The acoustic quality prediction method based on condition perception and multi-scale attention autoencoder according to claim 1, characterized in that, In step 4, if the mean square error of the acoustic signal to be evaluated is less than or equal to the mean square error under normal operating conditions, then the acoustic quality probability score is directly assigned. A value of 1 corresponds to excellent processing quality; if the average mean square error of the evaluated acoustic signal is greater than the average mean square error under normal operating conditions, an acoustic quality probability score is calculated. , The smaller the value, the worse the corresponding processing quality.

10. A computer device, comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method as described in any one of claims 1 to 9.