Ship engine vibration signal working condition classification method, system, medium, program and terminal based on domain adaptation
Patent Information
- Application Number
- CN202511067143.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2045-07-31
AI Technical Summary
[0016]本申请首先对样本进行分层特征提取;其中,所述分层特征提取的方法包括:首先将船舶发动机的源域数据集和目标域数据集输入至浅层特征提取器以分别得到对应的浅层特征,并将所述浅层特征输入至深层特征提取器以分别得到对应的深层特征。这种分层特征提取的方法能够充分利用源域和目标域的数据特性,进而使得模型能够学到更加丰富和具有区分度的特征表示,使得模型能够更好地适应多种数据分布,增强了模型在分析船舶发动机振动信号时的泛化能力。接下来,将所述源域和目标域的浅层特征经梯度反转后输入至域判别器,通过反向传播梯度抑制所述浅层特征的域可分性,从而确保了特征在域之间的迁移能力,在这个过程中,模型不仅能够保持源域的任务分类能力,还能够减小目标域与源域的数据分布差异,使得目标域数据也能被有效处理。并且,在深层特征提取器中加入多核最大均值差异(MK-MMD)正则函数,通过在每一层进行特征对齐,减少了由于域间差异导致的负面影响,提高了模型对目标域的多种型号的船舶发动机振动信号的适应性,提升了跨域学习的效果和精度。从而在船舶发动机振动信号的分析过程中,能够充分捕捉到发动机复杂机型间的振动信号多样性,提升了模型在该应用场景下的泛化能力。综上,本申请通过上述对传统模型架构的改进,提升了模型在通过船舶发动机振动信号来分析不同机型的发动机的工况类别时的泛化能力。
Smart Images

Figure CN120951129B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of marine engines, and in particular to a domain-adaptive method, system, medium, program, and terminal for classifying the operating conditions of marine engine vibration signals. Background Technology
[0002] Domain adaptation is a technique in transfer learning that aims to effectively transfer knowledge learned from a source domain to a target domain, reducing the distributional differences between the source and target domains and enabling models trained on the source domain to generalize better in the target domain. Adversarial domain adaptation is a type of domain adaptation that addresses the problem of weak model generalization when there are significant differences in the distributions of data between the source and target domains. Through adversarial training, it minimizes the distributional differences between the source and target domains, thereby improving the model's generalization ability in the target domain. Traditional domain adaptation methods (such as Domain Adversarial Neural Networks, DANNs) typically include three key modules: a feature extractor, a classifier, and a discriminator. The feature extractor is responsible for extracting high-level features from the input data; the classifier performs sample classification based on these features; and the discriminator distinguishes whether the input features belong to the source or target domain. The feature extractor and the domain discriminator are connected by a gradient inversion layer (GRL). During backpropagation, the GRL inverts the sign of the gradient of the domain discriminant loss (multiplying it by -1), forcing the feature extractor to learn domain-invariant features, thereby effectively reducing the difference between the source domain and the target domain and improving the model's target domain generalization ability.
[0003] In the application scenario of classifying the operating conditions of ship engine vibration signals, different engine models often exhibit significant differences in the time-frequency characteristics (e.g., spectral energy distribution, harmonic components) of their vibration signals due to variations in structure, speed, and load. Traditional domain adaptation methods have limitations in addressing these complex domain differences in ship engine vibration signals, failing to fully capture the diversity among complex engine models. This is because traditional methods rely on a single feature extractor to directly output mixed features, making it difficult for the model to adapt to complex domain differences. Furthermore, traditional methods only mine domain-invariant features through adversarial learning, failing to fully utilize domain discriminative information, resulting in unsatisfactory performance in application scenarios involving complex domain differences and multiple target domains. Consequently, when using traditional models for classifying the operating conditions of ship engine vibration signals, the model's generalization ability is poor, and its performance on target engine models is subpar. Summary of the Invention
[0004] In view of the shortcomings of the prior art described above, the purpose of this application is to provide a domain-adaptive method, system, medium, program and terminal for classifying the operating conditions of ship engine vibration signals, in order to solve the above problems.
[0005] To achieve the above and other related objectives, a first aspect of this application provides a domain-adaptive-based method for classifying the operating conditions of ship engine vibration signals, comprising: acquiring a source domain dataset and a target domain dataset of the ship engine vibration signals; wherein the source domain dataset includes source domain vibration signals x s and its corresponding working condition label y sc-0 and domain tag y sd-0 The target domain dataset includes target domain vibration signals x t and its corresponding working condition label y tc-0 and domain tag y td-0 ; for x s and x t Hierarchical feature extraction is performed separately; wherein, the method of hierarchical feature extraction includes: x s and x t The input is fed into a shallow feature extractor to obtain the first source domain shallow features z. ss-1 and the shallow features z of the first target domain ts-1 , will z ss-1 and z ts-1 Input to a deep feature extractor to obtain the first source domain deep features z sd-1 and the deep features of the first target domain z td-1 ; will z sd-1 and z td -1 Input the load condition category classifier to obtain the first source domain predicted load condition label y sc-1 And the first target domain predicted working condition label y tc -1 ; Calculate y sc-1 and y sc-0 The first source domain classification loss L1 and y tc-1 and y tc-0 The first target domain classification loss L2 between z; ss-1 and z ts-1 After gradient reversal, the input is fed into the domain discriminator to obtain the source domain predicted domain label y. tc-1 And target domain prediction domain label y td-1 Combined with y sd-0 and y td-0 Calculate the domain adversarial loss L3 of the source domain and the domain adversarial loss L4 of the target domain; use the optimizer to backpropagate and update the network parameters until the total loss of the weighted sum of L1, L2, L3 and L4 satisfies the convergence condition.
[0006] In one embodiment of the first aspect of this application, the method further includes: adding a multi-kernel maximum mean difference regularization calculation z to each hidden layer of the deep feature extractor. sd-1 and z td-1 MK-MMD loss L 5,The optimizer is used to backpropagate and update the network parameters until the total loss obtained by weighted summation of L5, L1, L2, L3, and L4 satisfies the convergence condition.
[0007] In one embodiment of the first aspect of this application, when x s and x t Before performing hierarchical feature extraction, x s and x t Preprocessing is performed; wherein, the preprocessing process includes: processing x s and x t Downsampling is performed to match the number of sampling points in each signal cycle for data with different sampling rates in the source and target domains; the downsampled x... s and x t Perform frame segmentation and windowing; wherein, frame segmentation and windowing refers to dividing x s or x t The signal is divided into fixed-length short frames, each with a frame length not less than the period length of the vibration signal; a window function is applied to each short frame; the x-axis after windowing is then processed. s and x t Perform a short-time Fourier transform to generate a time-frequency graph.
[0008] In one embodiment of the first aspect of this application, the shallow feature extractor includes a multi-layer convolutional network and a depth residual shrinking network; wherein, each layer of the multi-layer convolutional network has a different convolutional kernel size, and each layer is composed of a Conv layer, a BN layer, a PRelu unit, and a MaxPooling layer in sequence; the step of extracting x... s and x t The input is fed into a shallow feature extractor to obtain the first source domain shallow features z. ss-1 and the shallow features z of the first target domain ts-1 The method includes: processing x through a multi-layer convolutional network of the shallow feature extractor respectively. s and x t After extracting features at different scales, the features at different scales are fused and concatenated. The concatenated features are then input into a residual shrinkage network for noise reduction, and a soft thresholding function is used to reduce x. s and x t Interference information in the middle.
[0009] In one embodiment of the first aspect of this application, the deep feature extractor includes multiple fully connected layers.
[0010] In one embodiment of the first aspect of this application, the domain discriminator includes a single-output softmax binary classifier for outputting the probability value that the vibration signal originates from the source domain or the target domain.
[0011] To achieve the above and other related objectives, a second aspect of this application provides a domain-adaptive-based classification system for ship engine vibration signals, comprising: a data acquisition module for acquiring a source domain dataset and a target domain dataset; wherein the source domain dataset includes source domain vibration signals x s and its corresponding working condition label y sc-0 and domain tag y sd-0 The target domain dataset includes target domain vibration signals x t and its corresponding working condition label y tc-0 and domain tag y td-0 The feature extraction module is used to extract features from x. s and x t Hierarchical feature extraction is performed separately; wherein, the method of hierarchical feature extraction includes: x s and x t The input is fed into a shallow feature extractor to obtain the first source domain shallow features z. ss-1 and the shallow features z of the first target domain ts-1 , will z ss-1 and z ts-1 Input to a deep feature extractor to obtain the first source domain deep features z sd-1 and the deep features of the first target domain z td-1 The working condition classification module is used to classify z sd-1 and z td-1 Input the load condition category classifier to obtain the first source domain predicted load condition label y sc-1 And the first target domain predicted working condition label y tc-1 ; Calculate y sc-1 and y sc-0 The first source domain classification loss L1 and y tc-1 and y tc-0 The first target domain classification loss L2 is used to classify z. ss-1 and z ts-1 After gradient reversal, the input is fed into the domain discriminator to obtain the source domain predicted domain label y. tc-1 And target domain prediction domain label y td-1 Combined with y sd-0 and y td-0 The module calculates the domain adversarial loss L3 of the source domain and the domain adversarial loss L4 of the target domain; the parameter optimization module is used to update the network parameters through backpropagation using the optimizer until the total loss of the weighted sum of L1, L2, L3 and L4 satisfies the convergence condition.
[0012] To achieve the above and other related objectives, a third aspect of this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in any of the preceding claims.
[0013] To achieve the above and other related objectives, a fourth aspect of this application provides a computer program product comprising computer program code that, when executed on a computer, causes the computer to perform the method described in any of the preceding claims.
[0014] To achieve the above and other related objectives, a fifth aspect of this application provides an electronic terminal, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the method described in any of the preceding claims.
[0015] As described above, this application has the following beneficial effects:
[0016] This application first performs hierarchical feature extraction on the samples. The hierarchical feature extraction method includes: firstly, inputting the source domain dataset and target domain dataset of the ship engine into a shallow feature extractor to obtain corresponding shallow features, and then inputting the shallow features into a deep feature extractor to obtain corresponding deep features. This hierarchical feature extraction method can fully utilize the data characteristics of the source and target domains, enabling the model to learn richer and more discriminative feature representations. This allows the model to better adapt to various data distributions and enhances its generalization ability when analyzing ship engine vibration signals. Next, the shallow features of the source and target domains are input into a domain discriminator after gradient inversion. Backpropagation of gradients suppresses the domain separability of the shallow features, thereby ensuring the feature transferability between domains. In this process, the model not only maintains the task classification ability of the source domain but also reduces the data distribution difference between the target and source domains, allowing the target domain data to be effectively processed. Furthermore, a multi-kernel maximum mean difference (MK-MMD) regularization function is incorporated into the deep feature extractor. By aligning features at each layer, the negative impact of inter-domain differences is reduced, improving the model's adaptability to vibration signals from various types of ship engines in the target domain and enhancing the effectiveness and accuracy of cross-domain learning. This allows for the full capture of the diverse vibration signals across complex engine models during the analysis of ship engine vibration signals, improving the model's generalization ability in this application scenario. In summary, this application, through the aforementioned improvements to the traditional model architecture, enhances the model's generalization ability when analyzing the operating conditions of different engine types using ship engine vibration signals. Attached Figure Description
[0017] Figure 1 The diagram shown is a flowchart of a domain-adaptive-based method for classifying the operating conditions of ship engine vibration signals in one embodiment of this application.
[0018] Figure 2The diagram shown is a flowchart of a domain-adaptive-based method for classifying the operating conditions of ship engine vibration signals in one embodiment of this application (including the MK-MMD algorithm).
[0019] Figure 3 The diagram shown is a flowchart of a domain-adaptive-based method for classifying the operating conditions of ship engine vibration signals in one embodiment of this application (including data preprocessing and model pretraining).
[0020] Figure 4 The diagram shown is a schematic representation of the structure of a multi-layer convolutional network in a shallow feature classifier according to an embodiment of this application.
[0021] Figure 5 The diagram shown is a model architecture diagram of a domain-adaptive ship engine vibration signal condition classification method in one embodiment of this application.
[0022] Figure 6 The diagram shown is a schematic representation of a domain-adaptive ship engine vibration signal condition classification system according to an embodiment of this application.
[0023] Figure 7 The diagram shown is a structural schematic of an electronic terminal according to an embodiment of this application. Detailed Implementation
[0024] The following specific examples illustrate the implementation of this application. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. This application can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this application. It should be noted that, unless otherwise specified, the following embodiments and features in the embodiments can be combined with each other.
[0025] In the embodiments of this application, terms such as "first" and "second" are used to distinguish identical or similar items with essentially the same function and effect. For example, "first XX" and "second XX" are merely used to distinguish different XXs and do not limit their order. Those skilled in the art will understand that terms such as "first" and "second" do not limit the quantity or execution order, and that terms such as "first" and "second" do not necessarily imply that they are different.
[0026] It should be noted that, in the embodiments of this application, the words "exemplary" or "for example" indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0027] In this application embodiment, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.
[0028] like Figure 1 As shown, the second aspect of this application provides a domain-adaptive-based method for classifying the operating conditions of ship engine vibration signals, including:
[0029] S1: Obtain the source domain dataset and target domain dataset of the ship engine vibration signal; wherein, the source domain dataset includes the source domain vibration signal x s and its corresponding working condition label y sc-0 and domain tag y sd-0 The target domain dataset includes target domain vibration signals x t and its corresponding working condition label y tc-0 and domain tag y td-0 .
[0030] It should be understood that the input signals (i.e., vibration signals) in the source and target domains, along with their corresponding labels (operating condition labels and domain labels), provide rich feature information for subsequent training. Operating condition labels represent system behavior under different operating conditions, such as engine start-up, acceleration, and idling. Domain labels identify whether the data comes from the source or target domain, helping to distinguish the different feature distributions between the source and target domains during training. Distinguishing between source and target domain data through labels helps the model perform initial training in the source domain. By addressing the differences between the source and target domains, the model's generalization ability to target domain data is improved. Through the organic combination of vibration signals and their labels in the source and target domains, the system can effectively identify key features in the vibration signals, ensuring the accuracy and effectiveness of subsequent model training.
[0031] Shallow feature extractors primarily focus on low-level, cross-domain shared features. These features represent fundamental information shared between the source and target domains. By extracting shallow features, commonalities between the source and target domains in lower dimensions can be effectively captured, laying a solid foundation for subsequent deep feature extraction and domain adaptation. Deep feature extractors are responsible for abstracting high-level semantic information from shallow features. This information not only enhances task-related discriminative capabilities but also facilitates the alignment of feature distributions across domains. In the deep feature space, features from the source and target domains must not only maintain efficient task classification capabilities but also minimize inter-domain differences to achieve cross-domain knowledge transfer. The goal of the domain discriminator is to maximize the difference in feature distributions between the source and target domains, making it impossible for the discriminator to distinguish their origins. This pushes shallow features to be as domain-insensitive as possible. This process helps reduce the differences between the source and target domains, enabling the model to learn domain-independent shared features. This not only reduces losses during domain transfer but also improves the adaptability to target domain data. In summary, adversarial training can suppress the domain separability of shallow features, effectively enhance the stability and robustness of cross-domain learning, make the feature distribution between the source and target domains more consistent, reduce inter-domain bias, and improve the model's generalization ability in the target domain.
[0032] In one embodiment of the first aspect of this application, the shallow feature extractor includes a multi-layer convolutional network and a depth residual shrinking network; wherein, each layer of the multi-layer convolutional network has a different convolutional kernel size, and each layer is composed of a Conv layer, a BN layer, a PRelu unit, and a MaxPooling layer in sequence; the step of extracting x... s and x t The input is fed into a shallow feature extractor to obtain the first source domain shallow features z. ss-1 and the shallow features z of the first target domain ts-1 The method includes: processing x through a multi-layer convolutional network of the shallow feature extractor respectively. s and x t After extracting features at different scales, the features at different scales are fused and concatenated. The concatenated features are then input into a residual shrinkage network for noise reduction, and a soft thresholding function is used to reduce x. s and x t Interference information in the middle.
[0033] Preferably, the convolutional kernel is a wide convolutional kernel (e.g., 13 or more). It should be understood that in vibration signals, short-term features are often manifested in rapid changes or local fluctuations in the signal. Therefore, using a larger convolutional kernel can capture a wider signal window, helping to identify the instantaneous change patterns of the signal. Furthermore, smaller convolutional kernels tend to focus on details (e.g., subtle local fluctuations), while larger convolutional kernels can better extract the global context and overall pattern of the signal. Therefore, a wide convolutional kernel can comprehensively consider more time steps, reducing the impact of local changes on feature extraction, thereby enabling the network to extract short-term features more stably, especially in the analysis of complex ship engine vibration signals.
[0034] like Figure 4 As shown, for example, the multi-layer convolutional network consists of 8 convolutional layers with kernel sizes of 13, 15, 17, 19, 30, 50, and 70. Each convolutional layer consists of a Conv layer, a BN layer, a PReLU unit, and a MaxPooling layer in sequence. The PReLU unit (parametric ReLU) is placed between the BN layer (batch normalization layer) and the MaxPooling layer (maximum pooling layer). It should be understood that since the values of ship engine vibration signals are usually positive and negative, using the ReLU function as the activation function would directly set the negative values of the input data to 0, resulting in the loss of some information. In contrast, the negative output of PReLU is a learnable dynamic parameter, therefore PReLU has stronger fitting and generalization capabilities than ReLU.
[0035] When x s and x t After being input into the shallow feature extractor, the multi-layer convolutional network processes x respectively. s and x t Feature extraction is performed, with each network layer extracting features at different scales. These features are then fused and concatenated, and the concatenated features are fed into a Residual Shrinking Network (DRSN) to learn the base features. A soft thresholding function is then used to reduce x. s and x tInterference information in the signal. It should be understood that the spliced features contain signal information at different scales. By feeding them into the DRSN, the network can learn basic features from these spliced multi-scale features, namely features related to the essential pattern of the signal (such as vibration frequency, fluctuation amplitude, etc.). By introducing residual connections (the idea of ResNet), the DRSN can avoid information from disappearing or becoming redundant in deep networks. At the same time, it reduces irrelevant or noisy features through shrinkage operations, thereby highlighting useful basic features. Soft thresholding is to “threshold” the signal values, suppressing signal values below a certain threshold to 0 and retaining the larger signal portion. Since vibration signals are often accompanied by noise, which may come from sensor errors, external interference, etc., if noise is not removed during feature extraction, this interference information will affect subsequent analysis and classification tasks. Therefore, through soft thresholding, the network can effectively compress or suppress noise signals, retaining only the main features related to the operating conditions, thereby reducing the impact of interference information and improving the clarity and accuracy of features.
[0036] Preferably, by introducing an attention mechanism to adjust the threshold, noise interference is suppressed and working condition-related features are enhanced, while retaining more transferable key features. This not only improves the accuracy and stability of the features, but also provides a cleaner and more robust input for subsequent working condition classification tasks and MK-MMD alignment.
[0037] In one embodiment of the first aspect of this application, z ss-1 and z ts-1 This includes the frequency component characteristics, time domain characteristics, and local mode characteristics of the ship engine vibration signal; wherein, the local mode characteristics include the short-time characteristics corresponding to the ship engine under impact, fluctuation, or oscillation conditions.
[0038] The deep feature extractor incorporates a multi-kernel maximum mean difference (MK-MMD) regularization term into each hidden layer to align deep features between the source and target domains. This method precisely aligns the feature distributions of the source and target domains at each layer, gradually reducing the distributional discrepancy between them. MK-MMD, a statistical measure of the difference between two distributions, optimizes the objective to make the features of the source and target domains converge layer by layer in the high-dimensional space, achieving domain-independent feature representations in the deep feature space. Furthermore, this distribution alignment is not merely a single-level feature matching process; it eliminates the bias between the source and target domains at each layer in a hierarchical manner. This ensures better fusion of features between the source and target domains at each abstraction level, helping deep networks avoid inter-domain feature interference during task-related learning, improving the model's generalization ability to target domain data, enabling the model to better adapt to the target domain, enhancing cross-domain learning performance, and ultimately improving task classification accuracy and model stability.
[0039] In one embodiment of the first aspect of this application, the deep feature extractor includes multiple fully connected layers.
[0040] Preferably, the multilayer fully connected layer includes three fully connected layers.
[0041] It should be understood that multi-layer fully connected layers can abstract features from the data layer by layer. Each layer weights and combines the features of the previous layer to extract more advanced and complex features. For example, the first layer learns simple features (such as linear relationships), the second layer learns more complex feature combinations based on this, and the third layer further abstracts more discriminative features. Therefore, multi-layer fully connected layers can provide richer nonlinear transformations, thus capturing potential nonlinear patterns in the input features of vibration signals, making the network more capable of fitting, and further enhancing the model's generalization ability for vibration signals from different engine models with large differences in feature distribution. Reducing the number of fully connected layers may make the model lack sufficient capacity to handle more complex patterns and high-dimensional data, while too many layers may lead to overfitting and increased computational overhead. When there are three fully connected layers, compared to other numbers of layers, it ensures that the model has sufficient learning capacity while avoiding overfitting and excessive computational overhead, achieving better performance with limited data and computational resources. S2: For x s and x t Hierarchical feature extraction is performed separately; wherein, the method of hierarchical feature extraction includes: x s and x t The input is fed into a shallow feature extractor to obtain the first source domain shallow features z. ss-1 and the shallow features z of the first target domain ts-1 , will z ss-1 and z ts-1 Input to a deep feature extractor to obtain the first source domain deep features z sd-1 and the deep features of the first target domain z td-1 .
[0042] In one embodiment of the first aspect of this application, the domain-adaptive-based method for classifying the operating conditions of ship engine vibration signals further includes:
[0043] S11: In relation to x s and x t Before performing hierarchical feature extraction, x s and x t Preprocessing is performed; wherein, the preprocessing process includes: processing x s and x tDownsampling is performed to match the number of sampling points in each signal cycle for data with different sampling rates in the source and target domains; the downsampled x... s and x t Perform frame segmentation and windowing; wherein, frame segmentation and windowing refers to dividing x s or x t The signal is divided into fixed-length short frames, each with a frame length not less than the period length of the vibration signal; a window function is applied to each short frame; the x-axis after windowing is then processed. s and x t Perform a short-time Fourier transform to generate a time-frequency graph.
[0044] It should be understood that downsampling refers to reducing the sampling rate of data, i.e., reducing the number of data points collected per second. Since different sampling rates between the source and target domains can lead to signals from the two domains having different time resolutions and different numbers of sampling points within the same signal cycle, thus affecting subsequent processing and model training, downsampling ensures consistency in the sampling rate of data from different source and target domains. For example, assuming the source domain's data sampling frequency is 1000Hz (1000 sampling points per second) and the target domain's data sampling frequency is 500Hz (500 sampling points per second), downsampling can reduce the source domain's data to 500Hz, ensuring that both have the same number of sampling points in each signal cycle, thereby guaranteeing their alignment on the time axis. Furthermore, for vibration signals, downsampling can remove unnecessary high-frequency information from the signal, reducing the amount of data and computational burden, while accelerating subsequent processing without affecting the main characteristics of the signal. Moreover, downsampling can reduce the amount of data processed by the model and the computational complexity, making it suitable for real-time signal processing applications with large amounts of data.
[0045] Frame-based windowing is a method for segmenting time-series signals. It first divides the vibration signal into several short frames, each with a frame length no less than the period length of the vibration signal. Preferably, the length of each short frame is the period length of the vibration signal. It should be understood that if the frame length is shorter than the period length, the periodic pattern of the signal may be lost; conversely, if the frame length is longer than the period length, it may contain multiple periods, causing the periodic characteristics of the signal to be averaged. Therefore, by setting the length of each frame to a complete period, the model can better identify the features within each period, especially key features such as peak value, amplitude, and frequency.
[0046] After framing, without windowing, the connection points between frames will experience drastic changes ("jumping"), meaning there will be abrupt changes in signal values between adjacent frames, causing high-frequency interference in the frequency domain. Therefore, a window function needs to be applied to each short frame. The window function can smooth the boundaries of each frame, allowing the signal values at the beginning and end of the frame to gradually transition to zero, thereby eliminating jumps and making the signal smoother and more continuous between frames. This reduces unnecessary high-frequency noise in subsequent frequency domain analysis (such as Fourier transform). Framing and windowing helps transform long-term signal sequences into local analysis problems, allowing the frequency characteristics within a short time to be effectively captured. In vibration signal processing, framing and windowing ensures that the signal remains stable within a short time, avoiding distortion introduced by improper truncation, thus improving the accuracy and stability of frequency domain analysis. Preferably, the window function is a Hanning window. Since the values at both ends of the Hanning window are close to zero, the weight of the center of the window is higher, and the weight of the ends is lower, the signal transition is smoother, helping to reduce the edge effects caused by signal cutting and making the connection between frames more natural. Therefore, when performing frequency domain analysis on ship engine vibration signals, it can improve the resolution of the spectrum, especially in the higher frequency range, and more accurately capture the periodic characteristics of the signal.
[0047] The Short-Time Fourier Transform (SFT) is a commonly used tool in signal processing to convert signals from the time domain to the time-frequency domain. Its basic idea is to divide the signal into frames and window them, then apply the SFT to each short frame, simultaneously displaying the signal's changes in time and frequency. The SFT can extract the frequency components of a signal at different time points, making it suitable for analyzing non-stationary signals (such as vibration signals) because their frequency components change over time. The advantage of the SFT is that it provides a time-frequency representation of the signal, helping us identify the frequency characteristics of the signal over different time periods.
[0048] The aforementioned data preprocessing methods also include performing only downsampling, or performing frame-by-frame windowing after downsampling without short-time Fourier transform, or performing short-time Fourier transform only after frame-by-frame windowing without downsampling. Any feasible combination of downsampling, frame-by-frame windowing, and short-time Fourier transform falls within the protection scope of the preprocessing methods described in this invention, and the foregoing examples do not constitute a limitation on the protection scope of the preprocessing methods of this invention. Downsampling is a process of reducing the amount of data, typically used to reduce the sampling rate and data volume of a signal. Performing only downsampling may result in the loss of some high-frequency information, causing a loss of signal details, especially for vibration signals where high-frequency components are important. It is suitable for working scenarios where the high-frequency components of the signal are not important, or already contain sufficient low-frequency information (e.g., vibration signals under specific operating conditions), and can reduce the computational burden. Frame-by-frame windowing of the signal after downsampling reduces the amount of data and allows for localized signal processing. However, without Fourier transform, frame-by-frame windowing simply cuts the signal into short segments and does not involve frequency domain analysis. This approach is suitable for applications that analyze the time-domain characteristics of signals (such as amplitude variations) rather than frequency components. Performing a short-time Fourier transform only after frame windowing, without downsampling, yields relatively accurate time-frequency information. However, the absence of downsampling leads to a higher computational load. It is suitable for situations requiring high-precision time-frequency analysis where a higher computational load is acceptable.
[0049] like Figure 2 As shown, in one embodiment of the first aspect of this application, the domain-adaptive-based method for classifying the operating conditions of ship engine vibration signals further includes:
[0050] Step S12: On x s and x t Before performing hierarchical feature extraction, the model is pre-trained; wherein, the pre-training steps include: setting x s Input the shallow feature extractor to obtain the shallow features z of the second source domain. ss-2 ; will z ss-2 Input the deep feature extractor to obtain the deep features z of the second source domain. sd-2 , will z sd-2 Input the operating condition category classifier to obtain the second source domain predicted operating condition label y sc-2 ;Calculate y using the minimum cross-entropy function sc-0 and y sc-2 The second source domain condition classification loss L6 is used; the model parameters are iterated through the optimizer until L6 meets the convergence condition.
[0051] like Figure 3 As shown, it should be understood that when x s and x tBefore performing hierarchical feature extraction, the model is pre-trained and x is tested simultaneously. s and x t When performing preprocessing, the preprocessing sequence should be placed before the pretraining.
[0052] S3: z sd-1 and z td-1 Input the load condition category classifier to obtain the first source domain predicted load condition label y sc-1 And the first target domain prediction working condition label y tc-1 ; Calculate y sc-1 and y sc-0 The first source domain classification loss L1 and y tc-1 and y tc-0 The first target domain condition classification loss L2 is between them.
[0053] It should be understood that y sc-1 and y tc-1 The L1 loss function reflects the model's ability to identify working conditions in both the source and target domains. The L1 loss reflects the model's accuracy in classifying working conditions in the source domain, while the L2 loss evaluates the model's performance in classifying working conditions in the target domain. By continuously minimizing these loss functions, the model's parameters are optimized, further improving its ability to classify working conditions in both the source and target domains. In this process, the cross-entropy loss not only helps optimize the model's classification accuracy in its respective domain but also provides crucial optimization basis for subsequent adversarial learning and feature alignment, ensuring consistent model performance between the source and target domains.
[0054] S4: z ss-1 and z ts-1 After gradient reversal, the input is fed into the domain discriminator to obtain the source domain predicted domain label y. tc-1 And target domain prediction domain label y td-1 Combined with y sd-0 and y td-0 Calculate the domain adversarial loss L3 of the source domain and the domain adversarial loss L4 of the target domain.
[0055] It should be understood that the purpose of gradient inversion is to reverse the direction of the gradient during backpropagation, thereby forcing the model to generate domain-independent features during feature extraction. This operation facilitates feature alignment between the source and target domains, enabling them to share features and achieve inter-domain adaptation. Through the cross-entropy loss function, the model measures the difference between the predicted domain label and the true label, thus optimizing the domain discriminators for the source and target domains. By minimizing these losses, the shallow feature extractor is forced to learn domain-invariant features of the target and source domain data, while the deep network collaborates with MK-MMD to align semantic features, forming a two-layer cross-domain adaptation of "shallow domain invariance + deep semantic alignment." This demonstrates the role of hierarchical feature alignment, ensuring consistency between the source and target domains in the feature space, thereby improving the classification accuracy and model robustness of the target domain task and enhancing the model's generalization ability when analyzing ship engine vibration signals.
[0056] In one embodiment of the first aspect of this application, the domain discriminator includes a single-output softmax binary classifier for outputting the probability value that the vibration signal originates from the source domain or the target domain.
[0057] It should be understood that a single-output softmax binary classifier, by outputting the probability values of the two classes, clearly tells the network whether the signal comes from the source domain (class 0) or the target domain (class 1). This structure is relatively simple and straightforward. For example, if the output of a signal is [0.2, 0.8], it means that the probability of the signal belonging to the target domain is 80%, and the probability of it belonging to the source domain is 20%. Although sigmoid can also be used for binary classification tasks (outputting probabilities between 0 and 1), it predicts the probability of each class separately. Softmax, on the other hand, considers the comparison relationship between the two classes simultaneously in binary classification. Therefore, the output of softmax better reflects the mutual exclusivity of the two classes because it ensures that the total probability sums to 1, while sigmoid predicts the probability of each class independently and cannot directly obtain the comparison relationship between the two. Furthermore, a single-output softmax binary classifier can be directly trained using the cross-entropy loss function, which is efficient, stable, and relatively simple to calculate in binary classification problems.
[0058] In one embodiment of the first aspect of this application, the calculation algorithm for L1, L2, L3 and L4 is to minimize the cross-entropy function.
[0059] Because cross-entropy loss provides rich gradient information, it helps to update model parameters more quickly and effectively. Compared with other loss functions, cross-entropy can finely adjust the model's classification boundary, thereby improving training efficiency and model performance. Furthermore, the cross-entropy loss function is computationally stable during training, avoiding gradient vanishing or exploding problems that can occur with other loss functions such as mean squared error, making the training process more stable and easier to converge.
[0060] S5: Use the optimizer to backpropagate and update the network parameters until the total loss of the weighted sum of L1, L2, L3 and L4 satisfies the convergence condition.
[0061] It should be understood that the optimizer's role is to update network parameters through backpropagation to minimize the value of the loss function. The loss function consists of the condition classification loss for the source and target domains, as well as the domain adversarial loss. The optimization objective is the total loss obtained by weighted summation. Through iterative training, the optimizer continuously adjusts the network parameters to gradually reduce the total loss until a preset convergence condition is met. This process ensures that the feature extractor, condition classifier, and domain discriminator are trained to their optimal performance, further enhancing the model's feature alignment ability between the source and target domains and improving the model's performance and stability. In this way, the model can learn more robust and general features, thereby improving its adaptability and generalization ability across different domains, ultimately achieving better performance in the target domain.
[0062] In one embodiment of this application, the method further includes: adding a multi-kernel maximum mean difference regularization calculation z to each hidden layer of the deep feature extractor. sd-1 and z td-1 The MK-MMD loss L5 is used to backpropagate and update the network parameters using the optimizer until the total loss obtained by weighted summing L5 with L1, L2, L3 and L4 satisfies the convergence condition.
[0063] Preferably, the model uses the Adam optimizer and the initial learning rate Ir is set to 0.001.
[0064] Because the Adam optimizer's momentum method and gradient normalization properties help mitigate the vanishing and exploding gradient problems, it enhances the model's robustness to noise in vibration signals, especially in scenarios with environmental interference or equipment aging noise, such as ship engines. Furthermore, the Adam optimizer can adaptively adjust the learning rate, accelerating model convergence, particularly when dealing with complex vibration signals. This ensures that low-frequency and high-frequency features in the signal are effectively learned, guaranteeing the model's stability and high classification accuracy under different operating conditions and data distributions. However, an excessively high learning rate (Ir) can lead to gradient oscillations during training, preventing the model from converging stably and even causing training failure. For complex vibration signals, the model's training process is easily affected by data fluctuations and noise; an excessively high learning rate may cause the model to skip optimal solutions, resulting in unsatisfactory final results. Conversely, an excessively low learning rate (Ir) slows down the training process. While it may converge stably, it may take a long time to achieve a satisfactory result, especially on large vibration signal datasets. A low learning rate can increase training time and may cause the model to get stuck in local optima, failing to effectively mine deeper features in the data. Therefore, when the initial learning rate Ir is 0.001, it can maintain an appropriate training speed while ensuring stability, avoiding the oscillation caused by an excessively high learning rate and the problem of excessively long training time caused by an excessively low learning rate.
[0065] In summary, the model achieves a progressive transfer from shallow denoising, domain-invariant feature extraction, deep semantic alignment, to operational condition classification through hierarchical decoupling and mutual synergy of shallow features, deep features, and the loss function. The shallow layers effectively eliminate noise and inter-domain differences in the data through denoising and learning domain-invariant features. The deep layers further optimize the feature distribution of the source and target domains through semantic alignment, making high-level features more unified and transferable. The hierarchical design of the loss function allows each layer's objective task to be optimized independently, while synergistic effects ensure that the entire network's learning process moves towards a unified goal (improving classification accuracy in the target domain). Ultimately, this progressive transfer process effectively improves the model's classification accuracy and robustness in the target domain, enhancing its adaptability to diverse engine models and operational conditions.
[0066] like Figure 6 As shown, a third aspect of this application provides a domain-adaptive-based classification system for ship engine vibration signals, comprising: a data acquisition module for acquiring a source domain dataset and a target domain dataset; wherein the source domain dataset includes source domain vibration signals x s and its corresponding operating condition label y sc-0 and domain tag y sd-0 The target domain dataset includes target domain vibration signals x t and its corresponding operating condition label y tc-0and domain tag y td-0 The feature extraction module is used to extract features from x. s and x t Hierarchical feature extraction is performed separately; wherein, the method of hierarchical feature extraction includes: x s and x t The input is fed into a shallow feature extractor to obtain the first source domain shallow features z. ss-1 and the shallow features z of the first target domain ts-1 , will z ss-1 and z ts-1 Input to a deep feature extractor to obtain the first source domain deep features z sd-1 and the deep features of the first target domain z td-1 The working condition classification module is used to classify z sd-1 and z td-1 Input the load condition category classifier to obtain the first source domain predicted load condition label y sc-1 And the first target domain predicted working condition label y tc-1 ; Calculate y sc-1 and y sc-0 The first source domain classification loss L1 and y tc-1 and y tc-0 The first target domain classification loss L2 is used to classify z. ss-1 and z ts-1 After gradient inversion, the input is fed into the domain discriminator to obtain the source domain predicted domain label y. tc-1 And target domain prediction domain label y td-1 Combined with y sd-0 and y td-0 The module calculates the domain adversarial loss L3 of the source domain and the domain adversarial loss L4 of the target domain; the parameter optimization module is used to update the network parameters through backpropagation using the optimizer until the total loss of the weighted sum of L1, L2, L3 and L4 satisfies the convergence condition.
[0067] It should be understood that the specific process of each module performing the above-mentioned steps has been described in detail in the above method embodiments, and will not be repeated here for the sake of brevity.
[0068] It should also be understood that the module division in the embodiments of this application is illustrative and only represents a logical functional division; in actual implementation, there may be other division methods. Furthermore, the functional modules in the various embodiments of this application can be integrated into a single processor, exist as separate physical entities, or be integrated into a single module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0069] A fourth aspect of this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in any of the preceding claims.
[0070] The fifth aspect of this application provides a computer program product including computer program code that, when run on a computer, causes the computer to perform the method described in any of the preceding claims.
[0071] like Figure 7 As shown, a sixth aspect of this application provides an electronic terminal comprising: at least one processor 701, a memory 702, at least one network interface 703, and a user interface 705. The various components in the device are coupled together via a bus system 704. It is understood that the bus system 704 is used to implement communication between these components. In addition to a data bus, the bus system 704 also includes a power bus, a control bus, and a status signal bus.
[0072] The user interface 705 may include a monitor, keyboard, mouse, trackball, clicker, button, touchpad, or touch screen.
[0073] It is understood that memory 702 can be volatile memory or non-volatile memory, or both. Non-volatile memory can be read-only memory (ROM) or programmable read-only memory (PROM), used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM) and synchronous static random access memory (SSRAM). The memories described in the embodiments of this invention are intended to include, but are not limited to, these and any other suitable categories of memory.
[0074] In this embodiment of the invention, the memory 702 is used to store various types of data to support the operation of the electronic terminal 700. Examples of this data include: any executable program for operation on the electronic terminal 700, such as operating system 7021 and application program 7022; operating system 7021 includes various system programs, such as framework layer, core library layer, driver layer, etc., for implementing various basic services and handling hardware-based tasks. Application program 7022 may include various applications, such as media player, browser, etc., for implementing various application services. The methods provided in this embodiment of the invention can be included in application program 7022.
[0075] The methods disclosed in the above embodiments of the present invention can be applied to or implemented by processor 701. Processor 701 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in processor 701 or by instructions in software form. The processor 701 may be a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Processor 701 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. General-purpose processor 701 may be a microprocessor or any conventional processor, etc. The steps of the accessory optimization method provided in the embodiments of the present invention can be directly reflected as being executed by a hardware decoding processor, or being executed by a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium, which is located in memory. The processor reads the information in the memory and combines it with its hardware to complete the steps of the aforementioned method.
[0076] In an exemplary embodiment, the electronic terminal 700 may be used by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), or complex programmable logic devices (CPLDs) to execute the aforementioned method.
[0077] According to the method provided in the embodiments of this application, this application also provides a computer program product, which includes: computer program code, which, when run on a computer, causes the computer to perform the method of any embodiment in the embodiments of this application.
[0078] According to the method provided in the embodiments of this application, this application also provides a computer-readable storage medium storing program code, which, when run on a computer, causes the computer to perform the method of any embodiment in the embodiments of this application.
[0079] The terms “component,” “module,” “system,” etc., used in this specification are used to refer to computer-related entities, hardware, firmware, combinations of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, and / or a computer. As illustrated, applications running on computing devices and computing devices can both be components. One or more components may reside in a process and / or an execution thread, and components may be located on a single computer and / or distributed among two or more computers. Furthermore, these components can be executed from various computer-readable media on which various data structures are stored. Components can communicate, for example, via local and / or remote processes based on signals having one or more data packets (e.g., data from two components interacting with another component between a local system, a distributed system, and / or a network, such as the Internet interacting with other systems via signals).
[0080] Those skilled in the art will recognize that the various illustrative logical blocks and steps described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.
[0081] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0082] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0083] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0084] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0085] In the above embodiments, the functions of each functional unit can be implemented entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. A computer program product includes one or more computer instructions (programs). When the computer program instructions (programs) are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. Computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., high-density digital video discs, DVDs), or semiconductor media (e.g., solid-state disks, SSDs, etc.).
[0086] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0087] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0088] In summary, this application effectively overcomes the various shortcomings of the prior art and has high industrial application value.
[0089] The above embodiments are merely illustrative of the principles and effects of this application and are not intended to limit this application. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of this application. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in this application should still be covered by the claims of this application.
Claims
1. A domain-adaptive-based method for classifying the operating conditions of ship engine vibration signals, characterized in that, include: Obtain the source domain dataset and target domain dataset of the ship engine vibration signal; wherein, the source domain dataset includes the source domain vibration signal x. s and its corresponding operating condition label y sc-0 and domain tag y sd-0 The target domain dataset includes target domain vibration signals x t and its corresponding operating condition label y tc-0 and domain tag y td-0 ; For x s and x t Hierarchical feature extraction is performed separately; wherein, the method of hierarchical feature extraction includes: x s and x t The input is fed into a shallow feature extractor to obtain the first source domain shallow features z. ss-1 and the shallow features z of the first target domain ts-1 , will z ss-1 and z ts-1 Input to a deep feature extractor to obtain the first source domain deep features z sd-1 and the deep features z of the first target domain td-1 ; z sd-1 and z td-1 Input the load condition category classifier to obtain the first source domain predicted load condition label y sc-1 And the first target domain prediction working condition label y tc-1 ; Calculate y sc-1 and y sc-0 The first source domain classification loss L1 and y tc-1 and y tc-0 The first target domain condition classification loss L2; z ss-1 and z ts-1 After gradient reversal, the input is fed into the domain discriminator to obtain the source domain predicted domain label y. tc-1 And target domain prediction domain label y td-1 Combined with y sd-0 and y td-0 Calculate the domain adversarial loss L3 of the source domain and the domain adversarial loss L4 of the target domain; The network parameters are updated by backpropagation using the optimizer until the total loss of the weighted sum of L1, L2, L3 and L4 satisfies the convergence condition. The shallow feature extractor includes a multi-layer convolutional network and a depth residual shrinking network; each layer of the multi-layer convolutional network has a different kernel size, and each layer consists of a Conv layer, a BN layer, a PRelu unit, and a MaxPooling layer in sequence; the step of extracting x... s and x t The input is fed into a shallow feature extractor to obtain the first source domain shallow features z. ss -1 and the shallow features z of the first target domain ts-1 The method includes: in the multi-layer convolutional network of the shallow feature extractor, respectively processing x s and x t After extracting features at different scales, the features at different scales are fused and concatenated. The concatenated features are then input into a residual shrinkage network for noise reduction, and a soft thresholding function is used to reduce x. s and x t Interference information in; The deep feature extractor comprises multiple fully connected layers.
2. The method for classifying the operating conditions of ship engine vibration signals based on domain adaptation according to claim 1, characterized in that, Also includes: Multi-kernel maximum mean difference regularization calculation z is added to each hidden layer of the deep feature extractor. sd-1 and z td-1 MK-MMD loss L 5, The optimizer is used to backpropagate and update the network parameters until L5 matches L1 and L2. 2、 The total loss obtained by weighted summation of L3 and L4 satisfies the convergence condition.
3. The method for classifying the operating conditions of ship engine vibration signals based on domain adaptation according to claim 1, characterized in that, In x s and x t Before performing hierarchical feature extraction, x s and x t Preprocessing is performed; wherein the preprocessing process includes: For x s and x t Downsampling is performed to match the number of sampling points of data with different sampling rates in the source and target domains within each signal cycle; downsampled x s and x t Perform frame segmentation and windowing; wherein, frame segmentation and windowing refers to dividing x s or x t The signal is divided into short frames of fixed length, and the length of each short frame is not less than the period length of the vibration signal. Apply a window function to each short frame; apply the windowed x-axis to each frame. s and x t Perform a short-time Fourier transform to generate a time-frequency graph.
4. The method for classifying the operating conditions of ship engine vibration signals based on domain adaptation according to claim 1, characterized in that, The domain discriminator includes a single-output softmax binary classifier, which outputs the probability value of whether the vibration signal comes from the source domain or the target domain.
5. A domain-adaptive-based classification system for the operating conditions of marine engine vibration signals, characterized in that, include: The data acquisition module is used to acquire source domain datasets and target domain datasets; wherein, the source domain dataset includes source domain vibration signals x. s and its corresponding operating condition label y sc-0 and domain tag y sd-0 The target domain dataset includes target domain vibration signals x t and its corresponding operating condition label y tc-0 and domain tag y td-0 ; The feature extraction module is used to extract features from x. s and x t Hierarchical feature extraction is performed separately; wherein, the method of hierarchical feature extraction includes: x s and x t The input is fed into a shallow feature extractor to obtain the first source domain shallow features z. ss-1 and the shallow features z of the first target domain ts-1 , will z ss-1 and z ts-1 Input to a deep feature extractor to obtain the first source domain deep features z sd-1 and the deep features z of the first target domain td-1 ; The working condition classification module is used to classify z sd-1 and z td-1 Input the load condition category classifier to obtain the first source domain predicted load condition label y sc-1 And the first target domain prediction working condition label y tc-1 ; Calculate y sc-1 and y sc-0 The first source domain classification loss L1 and y tc-1 and y tc-0 The first target domain condition classification loss L2; The domain discrimination module is used to determine z ss-1 and z ts-1 After gradient reversal, the input is fed into the domain discriminator to obtain the source domain predicted domain label y. tc-1 And target domain prediction domain label y td-1 Combined with y sd-0 and y td-0 Calculate the domain adversarial loss L3 of the source domain and the domain adversarial loss L4 of the target domain; The parameter optimization module is used to update the network parameters through backpropagation using the optimizer until the total loss of the weighted sum of L1, L2, L3 and L4 satisfies the convergence condition. The shallow feature extractor includes a multi-layer convolutional network and a depth residual shrinking network; each layer of the multi-layer convolutional network has a different kernel size, and each layer consists of a Conv layer, a BN layer, a PRelu unit, and a MaxPooling layer in sequence; the step of extracting x... s and x t The input is fed into a shallow feature extractor to obtain the first source domain shallow features z. ss -1 and the shallow features z of the first target domain ts-1 The method includes: in the multi-layer convolutional network of the shallow feature extractor, respectively processing x s and x t After extracting features at different scales, the features at different scales are fused and concatenated. The concatenated features are then input into a residual shrinkage network for noise reduction, and a soft thresholding function is used to reduce x. s and x t Interference information in; The deep feature extractor comprises multiple fully connected layers.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1-4.
7. A computer program product, characterized in that, The computer program product includes computer program code that, when run on a computer, causes the computer to implement the method as described in any one of claims 1-4.
8. An electronic terminal, comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the method according to any one of claims 1-4.
Citation Information
Patent Citations
Elevator traction sheave fault judgment method based on improved DANN model
CN115358276A
Adaptive bearing fault diagnosis method based on depth discrimination confrontation domain
CN116451022A