Training methods, state recognition methods, devices and equipment for state recognition models
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-17
- Publication Date
- 2026-08-14
AI Technical Summary
然而,电力机车实际运行环境复杂多变,主压缩机在工作时不仅会产生强烈的机械噪声,而且还会受到轮轨噪声、气流湍流噪声、牵引电机电磁噪声及各类环境噪声的多重干扰
[0034]本申请实施例提供了一种状态识别模型的训练方法以及一种状态识别方法。本方案的状态识别模型,能够在强噪声环境下基于现场采集音频对机车压缩机进行高精度与高鲁棒性识别。
Smart Images

Figure CN122575397A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio processing technology, and in particular to a training method, state recognition method, apparatus and device for a state recognition model. Background Technology
[0002] As the core power source of the locomotive's braking system, the main compressor's operating status directly affects train safety. With the development of intelligent operation and maintenance technology, audio signal-based fault diagnosis technology has been widely applied to the condition monitoring and fault identification of the main compressor. However, the actual operating environment of electric locomotives is complex and variable. During operation, the main compressor not only generates strong mechanical noise but is also subject to multiple interferences from wheel-rail noise, airflow turbulence noise, traction motor electromagnetic noise, and various environmental noises.
[0003] In actual operating conditions, the aforementioned complex background noise often drowns out the weak characteristic signals reflecting equipment faults, leading to decreased accuracy and insufficient reliability in identifying the main compressor's operating status. Especially in the early stages of a fault, the fault characteristics themselves are relatively weak, making them highly susceptible to missed detections and misjudgments under strong noise interference. Therefore, how to achieve high-precision and robust identification of the main compressor's operating status based on audio signals in high-noise environments has become a pressing problem to be solved in this field. Summary of the Invention
[0004] This application provides a training method for a state recognition model, a state recognition method, an apparatus, and a device. The technical solution is shown below.
[0005] On the one hand, a training method for a state recognition model is provided, the method comprising: Multiple clean audio samples were acquired when the locomotive compressor was in operation. These clean audio samples were collected in a laboratory environment while the locomotive compressor was running. Based on the noise category and preset signal-to-noise ratio in the locomotive operating environment, multiple meta-learning tasks are constructed; wherein, each meta-learning task corresponds to a noise category and a signal-to-noise ratio; For any meta-learning task, according to the noise category and signal-to-noise ratio corresponding to the meta-learning task, multiple sample pairs are constructed for the meta-learning task based on the multiple clean audios, and each sample pair includes a clean audio and a corresponding noisy frequency; Obtain the frequency domain feature maps of the clean audio and the noisy audio in each sample pair; Based on the frequency domain feature map, the network parameters of the denoising network of the state recognition model are updated to obtain the denoising network with the first update.
[0006] In some embodiments, constructing multiple sample pairs for the meta-learning task based on the multiple clean audio tracks according to the noise category and signal-to-noise ratio corresponding to the meta-learning task includes: Select a preset number of clean audio tracks from the plurality of clean audio tracks; Based on the noise category and signal-to-noise ratio corresponding to the meta-learning task, noise is added to the preset number of clean audio segments to obtain the preset number of noise-banded frequencies. Based on the preset number of clean audio samples and the preset number of noisy audio samples, a preset number of sample pairs are constructed for the meta-learning task.
[0007] In other embodiments, obtaining the frequency domain feature maps of the clean audio and noisy frequencies in each sample pair includes: The Mel spectrogram of the pure audio in each sample pair is obtained, and the Mel spectrogram of the pure audio is processed into a single-channel grayscale image to obtain the frequency domain feature map of the pure audio. The Mel spectrum of the noisy frequency in each sample pair is obtained, and the Mel spectrum of the noisy frequency is processed into a single-channel grayscale image to obtain the frequency domain feature map of the noisy frequency.
[0008] In other embodiments, updating the network parameters of the denoising network of the state recognition model based on the frequency domain feature map to obtain the initially updated denoising network includes: Based on the frequency domain feature map and according to the selected meta-learning training method, the initial network parameters of the noise reduction network are updated to obtain the noise reduction network with the first update. The noise reduction network in the initial update has optimized initial network parameters.
[0009] In other embodiments, in response to the meta-learning training method being a model-independent meta-learning training method, the method further includes: For any meta-learning task, select a first number of sample pairs from the preset number of sample pairs corresponding to it to obtain the support set of the meta-learning task; The remaining second number of sample pairs are used as the query set for the meta-learning task; the sum of the first number and the second number is equal to the preset number.
[0010] In other embodiments, updating the initial network parameters of the denoising network based on the frequency domain feature map and according to a selected meta-learning training method includes: In each round of training, a target number of meta-learning tasks are selected from the multiple meta-learning tasks; For any selected meta-learning task, the initial network parameters are updated based on the support set of the meta-learning task to obtain the temporarily updated denoising network; wherein the temporarily updated denoising network has temporary network parameters adapted to the meta-learning task. Obtain the temporarily updated denoising loss of the denoising network on the query set of the meta-learning task; The meta-loss is obtained based on the denoising loss corresponding to each meta-learning task in the plurality of meta-learning tasks, and the initial network parameters are updated based on the meta-loss. Repeat the steps of selecting a target number of meta-learning tasks and updating the network parameters of the denoising network until the training termination condition is met.
[0011] In other embodiments, updating the initial network parameters based on the support set of the meta-learning task includes: Starting with the initial network parameters, multiple gradient update operations are performed on the support set of the meta-learning task.
[0012] In other embodiments, performing multiple gradient update operations on the support set of the meta-learning task includes: Based on the support set of the meta-learning task, the noise reduction loss of the noise reduction network on the support set is obtained; Based on the initial network parameters, the first learning rate, the gradient operator, and the denoising loss, multiple parameter update operations are performed on the denoising network; wherein, the gradient operator is an operator that calculates the gradient of the initial network parameters.
[0013] In other embodiments, obtaining the temporarily updated denoising loss of the denoising network on the query set of the meta-learning task includes: Obtain the frequency domain feature map of each sample pair with noise in the query set to obtain the frequency domain feature map before noise reduction; The noise reduction network is temporarily updated to perform noise reduction processing on the frequency domain feature map before noise reduction, so as to obtain the frequency domain feature map after noise reduction. Based on the denoised frequency domain feature map and the frequency domain feature map of the clean audio of each sample pair in the query set, the average denoising loss of the temporarily updated denoising network on the query set is obtained.
[0014] In other embodiments, obtaining the meta-loss based on the denoising loss corresponding to each of the plurality of meta-learning tasks, and updating the initial network parameters based on the meta-loss, includes: The sum of the average noise reduction losses corresponding to each of the multiple meta-learning tasks is taken as the meta-loss. The initial network parameters are updated based on the second learning rate and the meta-loss gradient; wherein the meta-loss gradient is the gradient of the meta-loss with respect to the initial network parameters.
[0015] On the other hand, a state recognition method is provided, the method comprising: Acquire audio to be processed collected at the operation site, wherein the operation site includes the target locomotive compressor for which status identification is to be performed, and the audio to be processed includes sound signals generated by the target locomotive compressor during operation; Based on the frequency domain feature map of the audio to be processed, the network parameters of the initially updated noise reduction network are updated again to obtain the noise reduction network updated again. The frequency domain feature map of the audio to be processed is divided into a training set and a test set; The noise reduction network is updated again, and noise reduction processing is performed on the frequency domain feature maps of the training set and the test set to obtain the noise-reduced training set and the noise-reduced test set. Based on the denoised training set, the network parameters of the classification network are updated to obtain the updated classification network; The noise-reduced test set is input into the updated classification network, and the updated classification network is used to identify the operating status of the target locomotive compressor to obtain the operating status identification result.
[0016] In some embodiments, the noise reduction network updated for the first time has optimized initial network parameters; the further updating of the network parameters of the noise reduction network based on the frequency domain feature map of the audio to be processed includes: Starting with the optimized initial network parameters, multiple gradient update operations are performed on the frequency domain feature map of the audio to be processed.
[0017] In other embodiments, updating the network parameters of the classification network based on the denoised training set includes: Since the frequency domain feature map in the training set after noise reduction is in single-channel form, channel expansion is performed on the single-channel form of the frequency domain feature map to obtain a three-channel form of the frequency domain feature map. The network parameters of the classification network are updated based on the frequency domain feature map in the form of the three channels.
[0018] On the other hand, a training device for a state recognition model is provided, the device comprising: The first acquisition module is configured to acquire multiple clean audio signals when the locomotive compressor is in operation, wherein the clean audio signals are collected in a laboratory environment during the operation of the locomotive compressor; The first construction module is configured to construct multiple meta-learning tasks based on the noise category and preset signal-to-noise ratio in the locomotive operating environment; wherein, each meta-learning task corresponds to a noise category and a signal-to-noise ratio. The second construction module is configured to, for any meta-learning task, construct multiple sample pairs for the meta-learning task based on the multiple clean audios according to the noise category and signal-to-noise ratio corresponding to the meta-learning task, wherein each sample pair includes a clean audio and a corresponding noisy audio. The second acquisition module is configured to acquire frequency domain feature maps of the clean audio and noisy frequencies in each sample pair; The first training module is configured to update the network parameters of the denoising network of the state recognition model based on the frequency domain feature map, so as to obtain the denoising network with the first update.
[0019] In some embodiments, the second building module is configured as follows: A preset number of clean audio tracks are selected from the plurality of clean audio tracks; based on the noise category and signal-to-noise ratio corresponding to the meta-learning task, noise is added to the preset number of clean audio tracks to obtain a preset number of noise-banded audio tracks; Based on the preset number of clean audio samples and the preset number of noisy audio samples, a preset number of sample pairs are constructed for the meta-learning task.
[0020] In other embodiments, the first acquisition module is configured to: The Mel spectrogram of the pure audio in each sample pair is obtained, and the Mel spectrogram of the pure audio is processed into a single-channel grayscale image to obtain the frequency domain feature map of the pure audio. The Mel spectrum of the noisy frequency in each sample pair is obtained, and the Mel spectrum of the noisy frequency is processed into a single-channel grayscale image to obtain the frequency domain feature map of the noisy frequency.
[0021] In other embodiments, the first training module is configured as follows: Based on the frequency domain feature map and according to the selected meta-learning training method, the initial network parameters of the noise reduction network are updated to obtain the noise reduction network with the first update. The noise reduction network in the initial update has optimized initial network parameters.
[0022] In other embodiments, in response to the meta-learning training method being a model-independent meta-learning training method, the construction module is further configured to: For any meta-learning task, select a first number of sample pairs from the preset number of sample pairs corresponding to it to obtain the support set of the meta-learning task; The remaining second number of sample pairs are used as the query set for the meta-learning task; the sum of the first number and the second number is equal to the preset number.
[0023] In other embodiments, the first training module is configured as follows: In each round of training, a target number of meta-learning tasks are selected from the multiple meta-learning tasks; For any selected meta-learning task, the initial network parameters are updated based on the support set of the meta-learning task to obtain the temporarily updated denoising network; wherein the temporarily updated denoising network has temporary network parameters adapted to the meta-learning task. Obtain the temporarily updated denoising loss of the denoising network on the query set of the meta-learning task; The meta-loss is obtained based on the denoising loss corresponding to each meta-learning task in the plurality of meta-learning tasks, and the initial network parameters are updated based on the meta-loss. Repeat the steps of selecting a target number of meta-learning tasks and updating the network parameters of the denoising network until the training termination condition is met.
[0024] In other embodiments, the first training module is configured as follows: Starting with the initial network parameters, multiple gradient update operations are performed on the support set of the meta-learning task.
[0025] In other embodiments, the first training module is configured as follows: Based on the support set of the meta-learning task, the noise reduction loss of the noise reduction network on the support set is obtained; Based on the initial network parameters, the first learning rate, the gradient operator, and the denoising loss, multiple parameter update operations are performed on the denoising network; wherein, the gradient operator is an operator that calculates the gradient of the initial network parameters.
[0026] In other embodiments, the first training module is configured as follows: Obtain the frequency domain feature map of each sample pair with noise in the query set to obtain the frequency domain feature map before noise reduction; The noise reduction network is temporarily updated to perform noise reduction processing on the frequency domain feature map before noise reduction, so as to obtain the frequency domain feature map after noise reduction. Based on the denoised frequency domain feature map and the frequency domain feature map of the clean audio of each sample pair in the query set, the average denoising loss of the temporarily updated denoising network on the query set is obtained.
[0027] In other embodiments, the first training module is configured as follows: The sum of the average noise reduction losses corresponding to each of the multiple meta-learning tasks is taken as the meta-loss. The initial network parameters are updated based on the second learning rate and the meta-loss gradient; wherein the meta-loss gradient is the gradient of the meta-loss with respect to the initial network parameters.
[0028] On the other hand, a state recognition device is provided, the device comprising: The third acquisition module is configured to acquire audio to be processed collected at the operating site, the operating site including the target locomotive compressor to be identified in terms of its status, and the audio to be processed including the sound signal generated by the target locomotive compressor during its operation. The second training module is configured to update the network parameters of the initially updated noise reduction network based on the frequency domain feature map of the audio to be processed, so as to obtain the noise reduction network updated again. The data partitioning module is configured to divide the frequency domain feature map of the audio to be processed into a training set and a test set; The noise reduction module is configured to perform noise reduction processing on the frequency domain feature maps in the training set through the noise reduction network that has been updated again, so as to obtain the noise-reduced training set; The third training module is configured to update the network parameters of the classification network based on the denoised training set, thereby obtaining the updated classification network. The noise reduction module is further configured to perform noise reduction processing on the frequency domain feature map of the test set through the updated noise reduction network to obtain the noise-reduced test set; The identification module is configured to input the noise-reduced test set into the updated classification network, and to identify the operating status of the target locomotive compressor through the updated classification network to obtain the operating status identification result.
[0029] In some embodiments, the noise reduction network updated for the first time has optimized initial network parameters; The second training module is configured as follows: Starting with the optimized initial network parameters, multiple gradient update operations are performed on the frequency domain feature map of the audio to be processed.
[0030] In other embodiments, the third training module is configured as follows: Since the frequency domain feature map in the training set after noise reduction is in single-channel form, channel expansion is performed on the single-channel form of the frequency domain feature map to obtain a three-channel form of the frequency domain feature map. The network parameters of the classification network are updated based on the frequency domain feature map in the form of the three channels.
[0031] On the other hand, a computer device is provided, the device including a processor and a memory, the memory storing computer program code, the computer program code being loaded and executed by the processor to implement the training method of the state recognition model described above, or the state recognition method described above.
[0032] On the other hand, a computer-readable storage medium is provided, wherein computer program code is stored in the storage medium, the computer program code being loaded and executed by a processor of a computer device to implement the training method of the state recognition model described above, or the state recognition method described above.
[0033] On the other hand, a computer program product is provided, the computer program product including computer program code stored in a computer-readable storage medium, a processor of a computer device reading the computer program code from the computer-readable storage medium, the processor executing the computer program code, causing the computer device to execute the training method of the state recognition model described above, or the state recognition method described above.
[0034] This application provides a training method for a state recognition model and a state recognition method. The state recognition model of this solution can perform high-precision and robust recognition of locomotive compressors based on on-site audio acquisition in noisy environments.
[0035] In detail, the denoising network of the state recognition model is trained using a meta-learning method. The optimization goal of meta-learning is to first obtain an optimal set of initial network parameters. This allows the denoising network to quickly adapt to the new noise environment (corresponding to the current noise environment) using only a small amount of on-site audio data after deployment, completing the transition from the denoising network with the initial parameter update to the denoising network with the next parameter update. This training strategy introduces meta-learning into the front-end signal preprocessing stage, enabling the generation of a denoising network adapted to the current noise environment in real time before state recognition, even when facing unknown and complex on-site noise environments like those of locomotive compressors, demonstrating strong robustness. Because the denoising network provides high signal-to-noise ratio input features to the back-end classification network, weak fault features that were originally masked by noise in strong noise environments can be recovered and effectively captured by the classification network, thus ensuring the accuracy of state recognition by the back-end classification network. Attached Figure Description
[0036] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0037] Figure 1 This is a schematic diagram of an implementation environment for training a state recognition model to perform state recognition on a locomotive compressor, provided in an embodiment of this application. Figure 2 This is a flowchart of an embodiment of the present application providing an adaptive noise reduction and state recognition process for audio collected from a locomotive compressor based on meta-learning; Figure 3 This is a flowchart of a training method for a state recognition model provided in an embodiment of this application; Figure 4 This is a flowchart of a state recognition method provided in an embodiment of this application; Figure 5 This is a network architecture diagram of a classification network provided in an embodiment of this application; Figure 6 This is a schematic diagram of the structure of a training device for a state recognition model provided in an embodiment of this application; Figure 7 This is a schematic diagram of the structure of a state recognition device provided in an embodiment of this application; Figure 8 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application; Figure 9 This is a schematic diagram of the structure of another computer device provided in an embodiment of this application. Detailed Implementation
[0038] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0039] In this application, the terms "first," "second," etc., are used to distinguish identical or similar items that have essentially the same function. It should be understood that there is no logical or temporal dependency between "first," "second," and "nth," nor does it limit the quantity or execution order. It should also be understood that although the following description uses the terms "first," "second," etc., to describe various elements, these elements should not be limited by the terms.
[0040] These terms are simply used to distinguish one element from another. For example, without departing from the various examples, the first element can be referred to as the second element, and similarly, the second element can be referred to as the first element. Both the first and second elements can be elements, and in some cases, they can be separate and distinct elements.
[0041] "At least one" refers to one or more elements. For example, at least one element can be one element, two elements, three elements, or any integer number of elements greater than or equal to one. "Multiple" refers to two or more elements. For example, multiple elements can be two elements, three elements, or any integer number of elements greater than or equal to two.
[0042] In this article, "and / or" indicates that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0043] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, data stored, data displayed, etc.) and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant regions.
[0044] Figure 1 This is a schematic diagram of an implementation environment for training a state recognition model to perform state recognition on a locomotive compressor, provided in an embodiment of this application.
[0045] In this embodiment, the state recognition model includes a noise reduction network for audio noise reduction and a classification network for state classification. The data processing object in this embodiment is audio, and non-contact acoustic sensors, such as microphone arrays, are deployed around the locomotive compressor (e.g., the locomotive's main compressor).
[0046] See Figure 1 The implementation environment includes a first computer device 110 and a second computer device 120. The first computer device 110 is used for training noise reduction and enhancement based on audio samples for complex noise environments. In other words, the noise reduction network abandons the single-task supervised training paradigm and instead adopts a gradient-based meta-learning strategy for updating network parameters. Its core mechanism is that it does not pursue obtaining the globally optimal solution under a specific noise distribution, but rather focuses on learning a set of optimal initial network parameters (enabling the network to have high generalization ability).
[0047] The second computer device 120 is used to quickly adjust the network parameters of the noise reduction network (starting from the optimal initial network parameters mentioned above) using a small amount of audio collected at the operating site, so that it can quickly adapt to the current noise environment, thereby providing high-quality clean audio for subsequent status classification. The operating site includes the locomotive compressor to be identified.
[0048] It should be noted that the first computer device 110 can be a terminal node, an edge computing node, or a server (belonging to a cloud computing center), etc., and this application does not limit this. The second computer device 120 can be a terminal node, such as a portable handheld device specifically used for status identification of locomotive compressors, a smartphone or tablet with a specific application installed, etc., and this application does not limit this. The specific application has the function of identifying the status of the locomotive compressor. Furthermore, the specific application can be a standalone application or a sub-application embedded in a parent application, such as a mini-program, and this application does not limit this.
[0049] Figure 2 This is a flowchart of an embodiment of the present application for adaptive noise reduction and state recognition of audio collected from a locomotive compressor based on meta-learning.
[0050] See Figure 2 The embodiments of this application include: a meta-learning noise reduction and enhancement training stage for complex noise environments and a stage for identifying the operating status of the locomotive compressor based on the noise-reduced frequency domain feature map.
[0051] During the training phase, the denoising network is not trained using a single-task supervised training method, but rather using a meta-learning training method. The optimization objective of meta-learning is to obtain a set of optimal initial network parameters. This allows the denoising network to quickly adapt to the current noise environment by utilizing data collected on-site and updating gradients in only a few steps when faced with new denoising tasks. This training strategy is designed for the complex and variable operating environment of electric locomotives, aiming to reduce the interference of background noise on the model's recognition accuracy while simultaneously improving the generalization performance of the neural network and meeting the requirement for broad data adaptation across multiple operating conditions. The technical logic of this training strategy is as follows.
[0052] 1. Task Construction and Simulation: During the training phase, pure audio of locomotive compressors during operation is collected in a laboratory environment, and noise-laden frequencies are simulated accordingly. This enables the construction of a multi-source noise library covering various typical locomotive internal noises, and the development of various meta-learning tasks (also known as noise adaptation tasks).
[0053] like Figure 2As shown on the left, this scheme pre-sets C noise categories and S signal-to-noise ratios for the operating environment of electric locomotives. It then collects clean audio data of the locomotive compressor during operation in a laboratory environment, thereby constructing multiple meta-learning tasks based on the C noise categories and S signal-to-noise ratios. Each meta-learning task corresponds to one noise category and one signal-to-noise ratio.
[0054] It should be noted that for any meta-learning task, this scheme constructs multiple sample pairs based on the collected clean audio, according to the noise category and signal-to-noise ratio corresponding to that meta-learning task. Each sample pair includes a clean audio segment and its corresponding noisy frequency; that is, this scheme simulates noisy frequencies. Furthermore, after constructing sample pairs for each meta-learning task, this scheme further divides the data into a support set and a query set. For any given meta-learning task, this scheme uses a subset of the constructed sample pairs as the support set (K audio pairs) and another subset as the query set (M audio pairs).
[0055] 2. Meta-optimization process: Utilizing meta-learning algorithms, inner-loop updates and outer-loop optimizations are performed on the support and query sets of each meta-learning task. The inner-loop update simulates the process by which the denoising network achieves rapid gradient adaptation and updates using a small number of samples when facing new noise environments. The outer-loop optimization iteratively optimizes the initial network parameters of the denoising network based on the model performance feedback after the inner-loop update. .
[0056] like Figure 2 As shown in the left part, the meta-optimization process involves steps such as extracting frequency domain feature maps from sample pairs and training a denoising network based on the extracted frequency domain feature maps, thereby obtaining an initially updated denoising network. The initially updated denoising network has a set of optimal initial network parameters. .
[0057] For detailed information on extracting frequency domain feature maps and training the denoising network, please refer to step 304 below.
[0058] 3. Rapid adaptation mechanism: The optimal initial network parameters are obtained through meta-learning training. It contains general structural prior knowledge of noise processing. When the noise reduction network is deployed at the operation site of an electric locomotive and faces an unknown real-time noise environment, it only needs to use a small amount of on-site audio and perform a very small number of gradient updates to enable the network parameters of the noise reduction network to quickly converge to the optimal state that fits the current noise distribution.
[0059] like Figure 2As shown on the right side, when the noise reduction network is deployed at the operating site of an electric locomotive, this solution collects audio from the locomotive compressor on-site and extracts frequency domain feature maps. Based on these extracted feature maps, the noise reduction network is fine-tuned to obtain an updated network. Furthermore, this solution divides the extracted frequency domain feature maps into training and testing sets, and then uses the updated noise reduction network to perform noise reduction processing on the frequency domain feature maps in both sets, resulting in denoised training and testing sets.
[0060] Furthermore, this scheme updates the network parameters of the classification network based on the denoised training set, thereby obtaining an updated classification network. This process involves multiple rounds of training and the final saving of the network parameters. Accordingly, this scheme loads the updated classification network and inputs the denoised test set into the updated classification network to identify the operating status of the locomotive compressor.
[0061] It should be noted that for detailed information on fine-tuning the noise reduction network and training the classification network, please refer to steps 401-404 below.
[0062] Based on the above technical logic, the training method and state recognition method of the state recognition model provided in this application embodiment at least solve the following problems.
[0063] 1. It can solve the mismatch problem between static noise reduction networks and dynamic noise environments.
[0064] This solution utilizes a meta-learning mechanism, enabling the denoising network to perceive the current noise distribution and update its parameters in real time. This allows the denoising network to match the current noise environment in any complex operating condition, eliminating feature distortion or noise residue caused by network mismatch at the source. It can cope with the rapidly changing noise spectrum characteristics during locomotive operation.
[0065] 2. The denoising network learns general structural prior knowledge of noise processing. This allows for the acquisition of network parameters specific to the current noise environment based on only a small amount of on-site audio data after actual model deployment. This solves the problem of how to quickly and adaptively adjust the parameters of the denoising network in blind source conditions.
[0066] 3. It can solve the problem of early faults being undiagnosable due to strong noise masking weak features.
[0067] Related technologies attempt to forcibly extract target features from noisy frequencies, which essentially involves probabilistic guessing under low signal-to-noise ratios, making it prone to missing early, subtle faults. Our solution, however, moves the meta-optimization process forward to the audio preprocessing stage, prioritizing high fidelity and high signal-to-noise ratio of the input signal. This provides the backend classifier with physically clear and identifiable target features, fundamentally eliminating the physical limitations of strong noise environments on early fault diagnosis.
[0068] Figure 3 This is a flowchart illustrating a training method for a state recognition model provided in an embodiment of this application. The method is executed by a computer device, such as... Figure 1 The first computer device 110 in the system. See also Figure 3 The method includes the following steps.
[0069] 301. The computer equipment acquires multiple clean audio signals when the locomotive compressor is in operation. The clean audio signals are collected in a laboratory environment while the locomotive compressor is running.
[0070] It should be noted that steps 301-304 correspond to Figure 2 The left side. These four steps aim to train the denoising network using a meta-learning training method, utilizing clean audio obtained in a laboratory environment and simulated noisy frequencies, to obtain an optimal set of initial network parameters. This is to prepare for the complex and unknown noise environment of the future. The laboratory environment is a quiet environment free from external noise interference.
[0071] This step is used to acquire clean audio. For example, in this embodiment of the application, N audio clips (e.g., N=200) of a locomotive compressor during operation are acquired in a laboratory environment, with each audio clip lasting 1 second. These audio clips must have a sufficiently high signal-to-noise ratio, and can be approximated as noise-free signals, serving as the basis for subsequent noise addition.
[0072] 302. The computer equipment constructs multiple meta-learning tasks based on the noise category and preset signal-to-noise ratio in the locomotive operating environment; wherein, each meta-learning task corresponds to a noise category and a signal-to-noise ratio.
[0073] In this embodiment, the typical background noise categories of an electric locomotive in its operating environment include: Gaussian white noise, motor electromagnetic noise, airflow turbulence noise, impact noise, and human voice, totaling C types. As an example, C is set to 10. Furthermore, for each noise category, this scheme sets S different signal-to-noise ratios. Taking S=6 as an example, these six signal-to-noise ratios can be -5dB, 0dB, 5dB, 10dB, 15dB, and 20dB, respectively; this application does not limit these specific values.
[0074] It should be noted that this scheme defines a combination of a noise category and a signal-to-noise ratio as an independent meta-learning task T. Therefore, the total number of meta-learning tasks... .
[0075] 303. For any meta-learning task, the computer device constructs multiple sample pairs for the meta-learning task based on the multiple clean audios according to the noise category and signal-to-noise ratio corresponding to the meta-learning task. Each sample pair includes a clean audio and a corresponding noisy frequency.
[0076] This step is used to generate sample pairs within the task.
[0077] In some embodiments, for any meta-learning task, a preset number of clean audio tracks are selected from multiple clean audio tracks; based on the noise category and signal-to-noise ratio corresponding to the meta-learning task, noise is added to the preset number of clean audio tracks to obtain a preset number of noise-banded frequencies; based on the preset number of clean audio tracks and the preset number of noise-banded frequencies, a preset number of sample pairs are constructed for the meta-learning task.
[0078] As an example, taking a total of 200 clean audio tracks N and a preset quantity Y=20, then for each task... Perform the following operations: First, 20 clean audio tracks are randomly selected from 200 clean audio tracks; Next, according to the task Based on the corresponding noise category and signal-to-noise ratio, noise is added to each of the 20 clean audio tracks to obtain 20 noisy frequencies. These 20 noisy frequencies are then paired with their corresponding 20 clean audio tracks to obtain 20 sample pairs of clean audio and noisy frequency combinations.
[0079] Where i is a positive integer, and the value of i ranges from 1 to... As mentioned above, This refers to the total number of meta-learning tasks. As an example, noise can be added to clean audio as follows: First, generate a pure noise signal with the same sampling rate and length as the clean audio (taking Gaussian white noise as an example, then the pure noise signal is Gaussian white noise); then, according to the corresponding signal-to-noise ratio (e.g., -5dB), perform power normalization scaling on the noise amplitude, and then superimpose the normalized noise onto the clean audio to obtain the corresponding noisy frequency.
[0080] In summary, the embodiments of this application, through the above steps 301-303, realize the construction of a multi-source noise library and the generation of mixed data.
[0081] 304. The computer device acquires the frequency domain feature maps of the clean audio and the noisy audio in each sample pair, and updates the network parameters of the denoising network of the state recognition model based on the acquired frequency domain feature maps to obtain the first updated denoising network.
[0082] In some embodiments, this solution uses the U-Net network as the basic architecture for the noise reduction network. The symmetric encoder-decoder structure of the U-Net network, combined with skip connections, can effectively suppress noise while preserving detailed information of the input feature map.
[0083] It should be noted that other lightweight convolutional networks that meet the requirements (such as convolutional autoencoders, residual autoencoders, etc.) can also be used to replace the U-Net network as the basic architecture of the denoising network, and all of the above fall within the protection scope of this application. The initial network parameters of the denoising network are as follows: Taking (random initialization) as an example, after the training phase, the denoising network will obtain a set of optimal initial network parameters. For example, the input to the initially updated denoising network is a noisy frequency domain feature map of size 224×224×1, and the output is a denoised frequency domain feature map of size 224×224×1. This application does not limit this.
[0084] In other embodiments, taking the Mel spectrogram as an example, obtaining the frequency domain feature map of each sample pair for both the clean audio and noisy frequencies means: The Mel spectrogram of the pure audio in each sample pair is obtained, and the Mel spectrogram of the pure audio is processed into a single-channel grayscale image to obtain the frequency domain feature map of the pure audio; and the Mel spectrogram of the noisy frequency in each sample pair is obtained, and the Mel spectrogram of the noisy frequency is processed into a single-channel grayscale image to obtain the frequency domain feature map of the noisy frequency.
[0085] For example, the frequency domain feature map of the clean audio and the frequency domain feature map of the noisy audio are both 224×224×1 in size, and this application does not limit this.
[0086] In other embodiments, this scheme adjusts the initial network parameters of the denoising network based on the acquired frequency domain feature map and according to the selected meta-learning training method. The network is updated to obtain the initially updated denoising network; the initially updated denoising network has optimized initial network parameters. .
[0087] It should be noted that the selected meta-learning training method can be MAML (Model-Agnostic Meta-Learning), Reptile, FOMAML (First-Order MAML), or ANIL (Almost No Inner Loop MAML), etc., and this application does not limit it.
[0088] In other embodiments, assuming the selected meta-learning training method is MAML, the scheme further includes the following steps: for any meta-learning task, select a first number of sample pairs from the preset number of sample pairs corresponding to it to obtain the support set of the meta-learning task; use the remaining second number of sample pairs as the query set of the meta-learning task; wherein the sum of the first number and the second number is equal to the preset number.
[0089] Taking the first quantity as K and the second quantity as M as an example, then Y = K + M. Assuming the total number of clean audio tracks is N = 200, Y = 20, K = 5, and M = 15, then for each task... Perform the following operations: From the mission Five sample pairs were randomly selected from the 20 constructed sample pairs as the task. The support set is used to simulate rapid fine-tuning in the field; the remaining 15 sample pairs serve as the task. The query set is used to evaluate the suitability of the denoising network for the task. The noise reduction effect afterward.
[0090] In other embodiments, taking the frequency domain feature map as a Mel spectrum as an example, for the task... For any sample pair in the support set and query set, this scheme will extract the Mel spectrograms of the clean audio and the noisy audio in the sample pair. After extracting the Mel spectrograms, this scheme will uniformly scale them into a single-channel grayscale image of size 224×224. That is, the input and output of the noise reduction network are both single-channel grayscale images of size 224×224. This application does not limit this.
[0091] It should be noted that this article introduces the scheme using the example where the support set and query set consist of sample pairs. For instance, the support set may contain K sample pairs, and the query set M sample pairs. In addition, the support set and query set can also include frequency domain feature maps of sample pairs, such as those found in the task... The support set includes 5 Mel spectrogram pairs, and the query set includes 15 Mel spectrogram pairs, all in the form of (Mel spectrogram of pure audio, Mel spectrogram of noisy frequency), and this application does not limit this.
[0092] In some embodiments, taking MAML as an example of the selected meta-learning training method, the initial network parameters of the denoising network are adjusted based on the acquired frequency domain feature map and according to the selected meta-learning training method. The initial update of the noise reduction network is performed, including the following steps.
[0093] 3041. In each training round, a target number of meta-learning tasks are selected from multiple meta-learning tasks; for any selected meta-learning task, the initial network parameters are adjusted based on the support set of that meta-learning task. An update is performed to obtain a temporarily updated noise reduction network.
[0094] The temporarily updated denoising network has temporary network parameters adapted to the meta-learning task.
[0095] This step is used to sample task batches and perform inner loop updates. For example, taking a total of 60 meta-learning tasks and a target number of 4 as an example, a batch of tasks is randomly selected from all 60 meta-learning tasks in each training round, with a batch size of B=4. This application does not limit this.
[0096] In some embodiments, the initial network parameters of the denoising network are determined based on the support set of the meta-learning task. Updating means updating the noise reduction network to its current initial parameters. Starting from this point, perform multiple gradient update operations on the support set of this meta-learning task.
[0097] It should be noted that the initial network parameters of the noise reduction network are as follows. Starting from this point, the meaning of this sentence is: when training the denoising network for the first time, directly use the current initial network parameters. These are used as initial values for training, rather than randomly initializing the network parameters of the denoising network; that is, the current initial network parameters... Based on this, the network parameters of the noise reduction network are iteratively updated, which can effectively improve training efficiency.
[0098] In other embodiments, performing multiple gradient update operations on the support set of the meta-learning task means: based on the support set of the meta-learning task, obtaining the denoising loss of the denoising network on that support set; and based on the current initial network parameters of the denoising network... First learning rate gradient operator The obtained noise reduction loss is used to perform multiple parameter update operations on the noise reduction network.
[0099] Among them, the first learning rate Also known as the inner loop learning rate, gradient operator The current initial network parameters Operators for finding gradients.
[0100] For example, with meta-learning tasks For example, this step is used for: task-based Support set From the initial network parameters Initially, in the support set Upward Subgradient descent update yields a result adapted to the task. Temporary network parameters This process can be represented by the following formula 1.
[0101] Formula 1:
[0102] In Formula 1, This refers to the loss function of the noise reduction network. This refers to the noise reduction network in the task. Support set The noise reduction loss on the surface This refers to a noise reduction network. For example, in Formula 1, The value of is 0.01, and the value of L is 5; this application does not impose any restrictions on these values. Furthermore, the loss function of the denoising network can be the L1 loss function; this application also does not impose any restrictions on this.
[0103] 3042. Obtain the temporary updated denoising loss of the denoising network on the query set of the meta-learning task.
[0104] This step is used to evaluate the effectiveness of the denoising network in adapting to the meta-learning task.
[0105] In some embodiments, this scheme obtains the denoising loss of the temporarily updated denoising network on the query set of the meta-learning task in the following manner: Obtain the frequency domain feature map of the noisy frequency for each sample pair in the query set of the meta-learning task to obtain the frequency domain feature map before denoising; perform denoising processing on the frequency domain feature map before denoising through a temporarily updated denoising network to obtain the frequency domain feature map after denoising; based on the frequency domain feature map after denoising and the frequency domain feature map of the clean audio for each sample pair in the query set, obtain the average denoising loss of the temporarily updated denoising network on the query set.
[0106] For example, with meta-learning tasks For example, this step is used for: a noise reduction network based on temporary updates (with features adapted to the task). Temporary network parameters ), for the task query set The frequency domain feature map of each sample pair with noisy frequencies is denoised, and the result is calculated in the query set. The average noise reduction loss. This process can be expressed as Equation 2 below.
[0107] Formula 2:
[0108] In Formula 2, This refers to the query set The average noise reduction loss is used to reflect the network's adaptation in the task. The actual performance on the platform. It refers to the frequency domain feature map of a single sample pair. It refers to the frequency domain characteristic map of the noisy frequency. It refers to the frequency domain characteristic map of pure audio. This refers to a temporarily updated noise reduction network. The noise-reduced frequency domain feature map output by the network, where M refers to the task... query set The number of sample pairs included.
[0109] 3043. Obtain the meta-loss based on the denoising loss corresponding to each meta-learning task in multiple meta-learning tasks, and update the initial network parameters based on the meta-loss.
[0110] This step is used to update the outer loop.
[0111] In some embodiments, a meta-loss is obtained based on the denoising loss corresponding to each meta-learning task in multiple meta-learning tasks, and the initial network parameters are adjusted based on the obtained meta-loss. Updating refers to using the sum of the average denoising losses corresponding to each of the multiple meta-learning tasks as the meta-loss. Based on the second learning rate and meta-loss gradient For the initial network parameters Update.
[0112] Among them, the second learning rate Also known as the outer loop learning rate, or the meta-loss gradient. It is a loss Regarding initial network parameters gradient, meta-loss It can be calculated using the following formula 3.
[0113] Formula 3:
[0114] In formula 3, This refers to the distribution of meta-learning tasks. This refers to having network parameters as The noise reduction network takes a frequency domain feature map with noise as input and outputs a noise-reduced frequency domain feature map.
[0115] In addition, this scheme is based on the outer loop learning rate. And the meta-loss gradient, which is calculated using the following formula 4 for the initial network parameters. Update: Formula 4:
[0116] For example, The value of is 0.001, and this application does not limit it. Additionally, Equation 4 involves the calculation of the second-order gradient to ensure the initial network parameters... It can accurately capture common features between tasks, thus enabling it to quickly adapt to new noise reduction tasks.
[0117] 3044. Repeat the steps of selecting a target number of meta-learning tasks and updating the network parameters of the denoising network until the training termination condition is met.
[0118] Repeat steps 3041-3043 above, processing a new batch of denoising tasks in each round, until the network parameters of the denoising network stabilize. After training, a set of optimal initial network parameters for the denoising network is obtained. This is the result of meta-learning training.
[0119] It should be noted that the above example, using MAML as the selected meta-learning training method, illustrates how to obtain a set of optimal initial network parameters for the denoising network. In addition to the above, other gradient-based meta-learning algorithms (such as Reptile, FOMAML, ANIL, etc.) can be used to achieve the same training objective, and this application does not limit this to any particular algorithm. Those skilled in the art can choose appropriate algorithms according to actual needs, and these all fall within the scope of protection of this application. For example, assuming the Reptile algorithm is used, the initial network parameters... Starting from the task Support set After performing multiple gradient update operations to obtain the fine-tuned network parameters, the initial network parameters are then directly applied. Move a small step in the direction of the fine-tuned network parameters. Since this method does not require calculating the second-order gradient, it reduces training complexity.
[0120] In summary, for the fault diagnosis scenario of locomotive compressors, the embodiments of this application introduce meta-learning technology into the front-end noise reduction stage, constructing a complete technical system from the laboratory simulated noise environment to the locomotive operation site. This system achieves collaborative operation of noise reduction enhancement based on meta-learning, rapid on-site adaptation, and fault diagnosis. This system solves problems such as the mismatch between static noise reduction networks and dynamic noise environments, the difficulty in constructing dedicated noise reduction networks, and the inability to diagnose early faults due to strong noise masking weak features.
[0121] This paper proposes a meta-learning-based training mechanism for denoising networks. Unlike single-task supervised training paradigms, this approach constructs diverse meta-learning tasks in a laboratory environment, covering various typical noise categories and signal-to-noise ratio combinations. Based on these tasks, a meta-learning strategy is employed to train the denoising network. It's important to note that the optimization objective of meta-learning is not to obtain the optimal solution for a specific noise distribution, but rather to obtain a set of optimal initial network parameters with high generalization ability. These initial network parameters contain prior knowledge of the general structure of noise processing. Thus, after actual model deployment, when facing unknown and dynamically changing on-site noise environments, the denoising network can be rapidly fine-tuned using only a small amount of on-site audio data, without relying on difficult-to-obtain clean reference audio, thereby ensuring accurate adaptation to the current noise distribution. This mechanism fundamentally solves the mismatch problem between static denoising networks and dynamic noise environments and enables the rapid construction of dedicated denoising networks in blind source states.
[0122] In summary, this scheme moves the optimization objective of meta-learning to the audio preprocessing stage, prioritizing the high signal-to-noise ratio of the input features. This allows weak fault features that were originally masked by noise in a noisy environment to be recovered and effectively captured by the classification network, solving the problem of undiagnosable early faults.
[0123] Figure 4 This is a flowchart of a state recognition method provided in an embodiment of this application. The execution subject of this method is a computer device, such as... Figure 1 The second computer device 120. See also Figure 4 The method includes the following steps.
[0124] 401. The computer equipment acquires the audio to be processed collected at the operation site, and based on the frequency domain feature map of the audio to be processed, the network parameters of the initially updated noise reduction network are updated again to obtain the updated noise reduction network.
[0125] This step is used to capture audio on-site (also known as the audio to be processed) and quickly fine-tune the noise reduction network.
[0126] It should be noted that the operating site includes the target locomotive compressor to be identified in terms of its status, and the audio to be processed includes the sound signals generated by the target locomotive compressor during its operation.
[0127] For example, 5-10 seconds of audio can be collected at the operation site, and the frequency domain feature map of the collected audio (e.g., with a size of 224×224×1) can be obtained as fine-tuning data for the noise reduction network.
[0128] In addition, the audio collected on site can be either a pure noise signal or a noisy signal that includes the operating sound of the target locomotive compressor; this application does not limit this.
[0129] It should be noted that, since the audio collected on-site is usually quite long, it is typically segmented during actual processing, and the frequency domain feature map of each segment is obtained.
[0130] As mentioned earlier, the initially updated denoising network has optimized initial network parameters. Accordingly, based on the frequency domain feature map of the audio to be processed, the network parameters of the initially updated noise reduction network are further updated. This means that the optimized initial network parameters are used as the basis for further updates. Starting from this point, multiple gradient update operations are performed on the frequency domain feature map of the audio collected on-site.
[0131] It should be noted that the optimized initial network parameters are used. Starting from this point, the meaning of this sentence is: when retraining the denoising network, directly use the optimized initial network parameters. The network parameters of the denoising network are used as initial values for training, rather than being randomly initialized; that is, the optimized initial network parameters are used. The network parameters of the noise reduction network are iteratively updated based on this.
[0132] As an example, taking the T-step gradient update (e.g., T=5 steps) on the frequency domain feature map of the audio collected on-site as an example, the final network parameters adapted to the current noise environment are obtained. The process is as follows: The frequency domain feature map of the on-site audio is input into the noise reduction network, and T rounds of forward propagation, noise reduction loss calculation, backpropagation, and gradient descent update are executed sequentially. In each round, the network parameters are optimized in a direction that better suits the current noise environment, and finally the network parameters are obtained. Since the fine-tuning process requires only a small amount of computation, it can be completed within milliseconds.
[0133] It should be noted that after the frequency domain feature map of the on-site audio is input into the noise reduction network, the input frequency domain feature map starts from the input end, goes through each network layer of the noise reduction network to complete linear and nonlinear operations, and finally outputs the noise-reduced frequency domain feature map. This is one forward propagation process.
[0134] For example, taking a self-supervised fine-tuning approach, for each round A forward propagation process can be represented by the following formula 5, which is not limited in this application.
[0135] Formula 5:
[0136] Where F refers to the mapping function of the classification network; , This refers to the frequency domain characteristic map of the audio collected on-site; Added noise; This refers to the The frequency domain feature map obtained after adding noise and then performing noise reduction processing.
[0137] Noise reduction loss calculation refers to calculating based on the loss function. and The deviation between them. Among them, the greater the noise reduction loss, the worse the noise reduction effect; the smaller the noise reduction loss, the closer the noise reduction result is to the ideal state.
[0138] For example, the noise reduction loss can be calculated using the self-supervised loss function shown in Formula 6 below, but this application does not limit it.
[0139] Formula 6:
[0140] in, This refers to the noise reduction loss, and N refers to the total number of frequency domain feature maps extracted from the audio collected on-site. It refers to the square of the L2 norm, where i is a positive integer. This refers to the i-th frequency domain feature map extracted from the audio collected on-site. The frequency domain feature map obtained after noise reduction corresponding to the i-th frequency domain feature map.
[0141] Backpropagation refers to working backward from the output to the input of the denoising network based on the calculated denoising loss, to determine how to adjust each network parameter to reduce the denoising loss. Gradient descent update refers to slightly modifying the network parameters according to the gradient direction (positive or negative) obtained from backpropagation, so that the denoising loss calculated in the next iteration is smaller.
[0142] In this context, backpropagation to find the gradient involves taking the partial derivative of the loss function with respect to the network parameters, thereby obtaining the gradient g shown in Equation 7 below. t .
[0143] Formula 7:
[0144] For example, the network parameters can be adjusted slightly according to the gradient direction and the learning rate to obtain the new network parameters used in the next round. If the gradient is greater than 0, it indicates that the current network parameters will increase the noise reduction loss, and the network parameters need to be slightly reduced, that is, the network parameters must be adjusted in the direction of decreasing; if the gradient is less than 0, it indicates that the current network parameters are too small, and the network parameters need to be slightly increased, that is, the network parameters must be adjusted in the direction of increasing.
[0145] In summary, the overall logic of the T-round update process is as follows: using the optimized initial network parameters... Using these as initial training values, the above process of forward propagation, noise reduction loss calculation, backpropagation, and gradient descent update is repeated T times. After completing T rounds of updates, the network parameters are obtained. .
[0146] 402. The computer device divides the frequency domain feature map of the audio to be processed into a training set and a test set; through the updated noise reduction network, noise reduction processing is performed on the frequency domain feature maps in the training set and the test set to obtain the noise-reduced training set and the noise-reduced test set.
[0147] For example, the training and test sets can be divided in an 8:2 ratio, but this application does not limit this to a specific ratio. The final trained denoising network has network parameters adapted to the current noise environment. Therefore, by inputting the frequency domain feature maps from the training and test sets into the finally trained denoising network, a frequency domain feature map with good denoising effect can be obtained.
[0148] 403. The computer equipment updates the network parameters of the classification network based on the noise-reduced training set, thus obtaining the updated classification network.
[0149] To adapt to the input format requirements of the classification network, and in response to the fact that the frequency domain feature map in the denoised training set is in single-channel form, it is necessary to perform channel expansion on the single-channel frequency domain feature map to obtain a three-channel frequency domain feature map; then, based on the three-channel frequency domain feature map, the network parameters of the classification network are updated.
[0150] For example, assuming a single-channel frequency domain feature map is 224×224×1 in size, it can be expanded into a three-channel feature map, forming a feature map of size 224×224×3, to match the input dimension of the classification network. Alternatively, during channel expansion, the single-channel frequency domain feature map can be copied to each channel, mapping it to all three channels, thus achieving channel expansion. This application does not limit the scope of this application.
[0151] For example, Figure 5 This is a network architecture diagram of a classification network provided in an embodiment of this application. For example... Figure 5 As shown, the classification network includes: Base-CNN (Base Convolutional Neural Network), Attention module, and classification module.
[0152] Base-CNN extracts deep information from the input frequency domain feature map by stacking convolutional layers, batch normalization layers, activation function (RELU) layers, and max pooling layers. The Attention module adaptively learns attention weights to weight and enhance the deep information extracted by Base-CNN, thereby strengthening weak fault features and suppressing complex background noise interference from the locomotive. The classification module performs classification mapping on the enhanced weighted features through adaptive average pooling, Dropout regularization, and linear layers, and finally outputs classification probabilities to indicate the current operating status of the locomotive compressor.
[0153] Batch normalization layers are used to standardize the distribution of features output by convolutional layers, accelerating the iterative update speed of the classification network and improving the identification accuracy of fault features. Max pooling layers are used for downsampling to reduce the length of feature sequences and decrease the overall computational cost of the network. Adaptive average pooling differs from ordinary average pooling; it does not require manual setting of the window and stride. Only the output feature length needs to be specified manually, and the classification network will automatically calculate the appropriate window and stride, averaging all features within each sliding window before outputting the degree. Regardless of the length of the input features, the output size is strictly equal to the preset value. Dropout regularization is used to prevent overfitting during the training process of the classification network.
[0154] In some embodiments, the training process of the classification network includes: For cases where the three-channel frequency domain feature map is labeled, the three-channel frequency domain feature map is input... Figure 5 After the classifier shown, the network parameters of Base-CNN, the Attention module, and the classifier are iteratively updated using the running state recognition loss as the optimization objective through the backpropagation algorithm, ultimately obtaining the optimal classification network adapted to the complex noise environment of the locomotive. For example, the frequency domain feature maps in the training set can be labeled by on-site technicians or algorithm developers; this application does not limit this.
[0155] The first point to clarify is that using motion state recognition loss as the optimization objective means that during the training process, the goal is to minimize the motion state recognition loss, thereby guiding the network to learn features and improve the accuracy of motion state recognition.
[0156] The second point to note is that the meaning of backpropagation can be referred to in step 401 above, and will not be repeated here. The backpropagation algorithm can be a stochastic gradient descent update algorithm, which is not limited in this application. In this application, iterative update refers to the process of gradually adjusting the network parameters through multiple iterations of optimization using the calculated gradient, so as to continuously reduce the classification loss and continuously improve the classification accuracy.
[0157] 404. The computer equipment inputs the noise-reduced test set into the updated classification network, and uses the updated classification network to identify the operating status of the target locomotive compressor, thereby obtaining the operating status identification result.
[0158] In this embodiment, the denoised frequency domain feature maps from the test set are input into an updated classification network to provide noise-free, highly recognizable data for the next step of operational status identification. Specifically, after the denoised frequency domain feature maps from the test set are input into the classification network, the network automatically performs operational status identification.
[0159] Using classification networks Figure 5 Taking the network architecture shown as an example, Base-CNN is used to extract deep information from the frequency domain feature map after noise reduction in the test set by stacking convolutional layers, batch normalization layers, activation function layers and max pooling layers; the Attention module is used to weight the deep information extracted by Base-CNN to enhance weak fault features and suppress the interference of complex background noise of the locomotive; the classification module is used to classify and map the weighted features through adaptive average pooling, Dropout regularization and linear layers, and finally output the running status category to indicate the current running status of the target locomotive's main compressor.
[0160] As an example, in addition to the current operating status category of the target locomotive's main compressor (such as normal, coupling failure, or fan blade failure), the operating status identification result may also include the confidence level (also known as classification probability) of the given operating status category and / or the corresponding maintenance suggestions. This application does not limit this.
[0161] Based on the above description, this solution proposes a two-stage architecture that coordinates front-end denoising and back-end state recognition. That is, this solution constructs a complete technical path from front-end denoising to back-end state recognition. First, meta-learning techniques are used to train the denoising network based on diverse meta-learning tasks, enabling the denoising network to obtain initial network parameters with rapid adaptability.
[0162] Secondly, at the locomotive operation site, only a small number of gradient updates are needed to rapidly fine-tune the noise reduction network based on a limited amount of on-site audio data, ensuring that the noise reduction network's capabilities precisely match the current noise environment. Finally, the frequency domain feature map of the on-site audio data is processed by the fine-tuned noise reduction network to obtain a denoised frequency domain feature map, which is then used as the input feature for the classification network to determine the operating status of the locomotive compressor.
[0163] In summary, this application provides a training method for a state recognition model and a state recognition method. The state recognition model of this solution can perform high-precision and robust recognition of locomotive compressors based on on-site audio acquisition in noisy environments.
[0164] In detail, the denoising network of the state recognition model is trained using a meta-learning training method. The optimization goal of meta-learning is to first obtain a set of optimal initial network parameters. This way, after the actual model is deployed, when faced with a new denoising task (corresponding to the current noise environment), the denoising network can quickly adapt to the current noise environment using only a small amount of on-site audio, completing the transition from the denoising network after the initial parameter update to the denoising network after the second parameter update.
[0165] This training strategy introduces a meta-learning strategy to train the noise reduction network in the front-end signal preprocessing stage. This enables the generation of a noise reduction network adapted to the current noise environment in real time before state recognition when facing unknown and complex on-site noise environments of locomotive compressors, demonstrating strong robustness.
[0166] Because the denoising network provides high signal-to-noise ratio input features to the backend classification network, it allows weak fault features that were originally masked by noise in noisy environments to be recovered and effectively captured by the classification network, thus ensuring the accuracy of state recognition in the backend classification network. In other words, because the pre-denoising stage effectively suppresses background noise and enhances fault features, the backend classification network can focus more on the key regions of the input features, thereby further improving the accuracy of state recognition.
[0167] Figure 6 This is a schematic diagram of the structure of a training device for a state recognition model provided in an embodiment of this application. See also... Figure 6 The device includes: The first acquisition module 601 is configured to acquire multiple clean audio signals when the locomotive compressor is in operation, wherein the clean audio signals are collected in a laboratory environment during the operation of the locomotive compressor. The first construction module 602 is configured to construct multiple meta-learning tasks based on the noise category and preset signal-to-noise ratio in the locomotive operating environment; wherein, one meta-learning task corresponds to a noise category and a signal-to-noise ratio. The second construction module 603 is configured to, for any meta-learning task, construct multiple sample pairs for the meta-learning task based on the multiple clean audios according to the noise category and signal-to-noise ratio corresponding to the meta-learning task, wherein each sample pair includes a clean audio and a corresponding noisy audio. The second acquisition module 604 is configured to acquire frequency domain feature maps of the clean audio and noisy frequencies in each sample pair; The first training module 605 is configured to update the network parameters of the denoising network of the state recognition model based on the frequency domain feature map, so as to obtain the denoising network with the first update.
[0168] The state recognition model provided in this application can perform high-precision and robust identification of locomotive compressors based on on-site audio in high-noise environments. Specifically, the denoising network of the state recognition model is trained using a meta-learning training method. The optimization objective of meta-learning is to first obtain a set of optimal initial network parameters. Thus, after actual model deployment, when facing a new denoising task (corresponding to the current noise environment), only a small amount of on-site audio is needed for the denoising network to quickly adapt to the current noise environment, completing the transition from the denoising network after the initial parameter update to the denoising network after a second parameter update. This training strategy introduces a meta-learning strategy to train the denoising network in the front-end signal preprocessing stage, enabling the generation of a denoising network adapted to the current noise environment in real time before state recognition, even when facing unknown and complex on-site noise environments of locomotive compressors, resulting in strong robustness. Since the denoising network provides high signal-to-noise ratio input features for the back-end classification network, weak fault features that were originally masked by noise in high-noise environments can be recovered and effectively captured by the classification network, thus ensuring the accuracy of state recognition by the back-end classification network.
[0169] In some embodiments, the second building module 603 is configured to: A preset number of clean audio tracks are selected from the plurality of clean audio tracks; based on the noise category and signal-to-noise ratio corresponding to the meta-learning task, noise is added to the preset number of clean audio tracks to obtain a preset number of noise-banded audio tracks; Based on the preset number of clean audio samples and the preset number of noisy audio samples, a preset number of sample pairs are constructed for the meta-learning task.
[0170] In other embodiments, the second acquisition module 604 is configured to: The Mel spectrogram of the pure audio in each sample pair is obtained, and the Mel spectrogram of the pure audio is processed into a single-channel grayscale image to obtain the frequency domain feature map of the pure audio. The Mel spectrum of the noisy frequency in each sample pair is obtained, and the Mel spectrum of the noisy frequency is processed into a single-channel grayscale image to obtain the frequency domain feature map of the noisy frequency.
[0171] In other embodiments, the first training module 605 is configured as follows: Based on the frequency domain feature map and according to the selected meta-learning training method, the initial network parameters of the noise reduction network are updated to obtain the noise reduction network with the first update. The noise reduction network in the initial update has optimized initial network parameters.
[0172] In other embodiments, in response to the meta-learning training method being a model-independent meta-learning training method, the second construction module 603 is further configured to: For any meta-learning task, select a first number of sample pairs from the preset number of sample pairs corresponding to it to obtain the support set of the meta-learning task; The remaining second number of sample pairs are used as the query set for the meta-learning task; the sum of the first number and the second number is equal to the preset number.
[0173] In other embodiments, the first training module 605 is configured as follows: In each round of training, a target number of meta-learning tasks are selected from the multiple meta-learning tasks; For any selected meta-learning task, the initial network parameters are updated based on the support set of the meta-learning task to obtain the temporarily updated denoising network; wherein the temporarily updated denoising network has temporary network parameters adapted to the meta-learning task. Obtain the temporarily updated denoising loss of the denoising network on the query set of the meta-learning task; The meta-loss is obtained based on the denoising loss corresponding to each meta-learning task in the plurality of meta-learning tasks, and the initial network parameters are updated based on the meta-loss. Repeat the steps of selecting a target number of meta-learning tasks and updating the network parameters of the denoising network until the training termination condition is met.
[0174] In other embodiments, the first training module 605 is configured as follows: Starting with the initial network parameters, multiple gradient update operations are performed on the support set of the meta-learning task.
[0175] In other embodiments, the first training module 605 is configured as follows: Based on the support set of the meta-learning task, the noise reduction loss of the noise reduction network on the support set is obtained; Based on the initial network parameters, the first learning rate, the gradient operator, and the denoising loss, multiple parameter update operations are performed on the denoising network; wherein, the gradient operator is an operator that calculates the gradient of the initial network parameters.
[0176] In other embodiments, the first training module 605 is configured as follows: Obtain the frequency domain feature map of each sample pair with noise in the query set to obtain the frequency domain feature map before noise reduction; The noise reduction network is temporarily updated to perform noise reduction processing on the frequency domain feature map before noise reduction, so as to obtain the frequency domain feature map after noise reduction. Based on the denoised frequency domain feature map and the frequency domain feature map of the clean audio of each sample pair in the query set, the average denoising loss of the temporarily updated denoising network on the query set is obtained.
[0177] In other embodiments, the first training module 605 is configured as follows: The sum of the average noise reduction losses corresponding to each of the multiple meta-learning tasks is taken as the meta-loss. The initial network parameters are updated based on the second learning rate and the meta-loss gradient; wherein the meta-loss gradient is the gradient of the meta-loss with respect to the initial network parameters.
[0178] All of the above-mentioned optional technical solutions can be combined in any way to form the optional embodiments of this application, and will not be described in detail here.
[0179] Figure 7 This is a schematic diagram of the structure of a state recognition device provided in an embodiment of this application. See also... Figure 7 The device includes: The third acquisition module 701 is configured to acquire audio to be processed collected at the operating site, the operating site including a target locomotive compressor to be identified in terms of its status, and the audio to be processed including sound signals generated by the target locomotive compressor during its operation. The second training module 702 is configured to update the network parameters of the initially updated noise reduction network based on the frequency domain feature map of the audio to be processed, so as to obtain the noise reduction network updated again. The data partitioning module 703 is configured to partition the frequency domain feature map of the audio to be processed into a training set and a test set; The noise reduction module 704 is configured to perform noise reduction processing on the frequency domain feature maps of the training set and the test set through the noise reduction network that has been updated again, so as to obtain the noise-reduced training set and the noise-reduced test set. The third training module 705 is configured to update the network parameters of the classification network based on the denoised training set, thereby obtaining the updated classification network. The identification module 706 is configured to input the noise-reduced test set into the updated classification network, and to identify the operating status of the target locomotive compressor through the updated classification network to obtain the operating status identification result.
[0180] The state recognition model provided in this application can perform high-precision and robust identification of locomotive compressors based on on-site audio in high-noise environments. Specifically, the denoising network of the state recognition model is trained using a meta-learning training method. The optimization objective of meta-learning is to first obtain a set of optimal initial network parameters. Thus, after actual model deployment, when facing a new denoising task (corresponding to the current noise environment), only a small amount of on-site audio is needed for the denoising network to quickly adapt to the current noise environment, completing the transition from the denoising network after the initial parameter update to the denoising network after a second parameter update. This training strategy introduces a meta-learning strategy to train the denoising network in the front-end signal preprocessing stage, enabling the generation of a denoising network adapted to the current noise environment in real time before state recognition, even when facing unknown and complex on-site noise environments of locomotive compressors, resulting in strong robustness. Since the denoising network provides high signal-to-noise ratio input features for the back-end classification network, weak fault features that were originally masked by noise in high-noise environments can be recovered and effectively captured by the classification network, thus ensuring the accuracy of state recognition by the back-end classification network.
[0181] In some embodiments, the noise reduction network updated for the first time has optimized initial network parameters; The second training module 702 is configured as follows: Starting with the optimized initial network parameters, multiple gradient update operations are performed on the frequency domain feature map of the audio to be processed.
[0182] In other embodiments, the third training module 705 is configured as follows: Since the frequency domain feature map in the training set after noise reduction is in single-channel form, channel expansion is performed on the single-channel form of the frequency domain feature map to obtain a three-channel form of the frequency domain feature map. The network parameters of the classification network are updated based on the frequency domain feature map in the form of the three channels.
[0183] All of the above-mentioned optional technical solutions can be combined in any way to form the optional embodiments of this application, and will not be described in detail here.
[0184] It should be noted that the state recognition model training device provided in the above embodiments, when training the state recognition model and when performing state recognition, is only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. Furthermore, the state recognition model training device and the state recognition model training method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process is detailed in the method embodiments, which will not be repeated here.
[0185] Figure 8 This is a schematic diagram of the structure of a computer device 800 provided in an embodiment of this application.
[0186] Typically, computer device 800 includes a processor 801 and a memory 802.
[0187] Processor 801 includes one or more processing cores, such as a quad-core processor or an octa-core processor. Processor 801 is implemented using at least one of the following hardware forms: DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Alternatively, processor 801 includes a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state.
[0188] In some embodiments, the processor 801 integrates a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content that the display screen needs to show.
[0189] In some embodiments, processor 801 further includes an AI (Artificial Intelligence) processor for processing computational operations related to machine learning.
[0190] The memory 802 includes one or more computer-readable storage media that are non-transitory. The memory 802 also includes high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices.
[0191] In some embodiments, the non-transitory computer-readable storage medium in memory 802 is used to store computer program code, which is executed by processor 801 to implement the training method or state recognition method of the state recognition model provided in the embodiments of this application.
[0192] In some embodiments, the computer device 800 further includes a peripheral device interface 803 and at least one peripheral device. The processor 801, memory 802, and peripheral device interface 803 are connected via a bus or signal line. Each peripheral device is connected to the peripheral device interface 803 via a bus, signal line, or circuit board. The peripheral device includes at least one of the following: a radio frequency circuit 804, a display screen 805, a camera assembly 806, an audio circuit 807, and a power supply 808.
[0193] Peripheral device interface 803 is used to connect at least one I / O (Input / Output) related peripheral device to processor 801 and memory 802. In some embodiments, processor 801, memory 802, and peripheral device interface 803 are integrated on the same chip or circuit board. In other embodiments, any one or two of processor 801, memory 802, and peripheral device interface 803 are implemented on separate chips or circuit boards, which is not limited in this application.
[0194] The radio frequency (RF) circuit 804 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 804 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 804 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals into electrical signals. The RF circuit 804 communicates with other computer devices through at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks, wireless local area networks, and / or WiFi (Wireless Fidelity) networks.
[0195] Display screen 805 is used to display a user interface (UI). This UI includes graphics, text, icons, videos, and any combination thereof. When display screen 805 is a touch display, it also has the ability to collect touch signals on or above its surface. These touch signals are input as control signals to processor 801 for processing. In this case, display screen 805 is also used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, display screen 805 is made of materials such as LCD (Liquid Crystal Display) or OLED (Organic Light-Emitting Diode).
[0196] The camera assembly 806 is used to acquire images or videos. In some embodiments, the camera assembly 806 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the computer device 800, and the rear-facing camera is located on the back of the computer device 800. In some embodiments, the camera assembly 806 also includes a flash. The flash is a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm-light flash and a cool-light flash, used for light compensation at different color temperatures.
[0197] Audio circuitry 807 includes a microphone and a speaker. Power supply 808 supplies power to the various components in computer device 800. Power supply 808 can be alternating current, direct current, a disposable battery, or a rechargeable battery. If power supply 808 includes a rechargeable battery, it can be a wired or wirelessly rechargeable battery. A wired rechargeable battery is charged via a wired connection, while a wirelessly rechargeable battery is charged via a wireless coil. The rechargeable battery also supports fast charging technology.
[0198] Those skilled in the art will understand that Figure 8 The structure shown does not constitute a limitation on the computer device 800, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.
[0199] Figure 9This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. The computer device 900 can be a server. The computer device 900 can vary significantly due to differences in configuration or performance, including one or more Central Processing Units (CPUs) 901 and one or more memories 902. The memories 902 store computer program code, which is loaded and executed by the processors 901 to implement the training method for the aforementioned state recognition model or the aforementioned state recognition method. Of course, the computer device 900 also has wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The computer device 900 also includes other components for implementing device functions, which will not be elaborated here.
[0200] In some embodiments, this application also provides a computer-readable storage medium, such as a memory including computer program code, which can be executed by a processor in a computer device to complete the training method of the state recognition model or the state recognition method described above. For example, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.
[0201] In some embodiments, this application also provides a computer program product, the computer program product including computer program code, the computer program code being stored in a computer-readable storage medium, a processor of a computer device reading the computer program code from the computer-readable storage medium, the processor executing the computer program code, causing the computer device to execute the training method of the state recognition model or the state recognition method described above.
[0202] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0203] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A training method for a state recognition model, characterized in that, The method includes: Multiple clean audio samples were acquired when the locomotive compressor was in operation. These clean audio samples were collected in a laboratory environment while the locomotive compressor was running. Based on the noise category and preset signal-to-noise ratio in the locomotive operating environment, multiple meta-learning tasks are constructed; wherein, each meta-learning task corresponds to a noise category and a signal-to-noise ratio; For any meta-learning task, according to the noise category and signal-to-noise ratio corresponding to the meta-learning task, multiple sample pairs are constructed for the meta-learning task based on the multiple clean audios, and each sample pair includes a clean audio and a corresponding noisy frequency; Obtain the frequency domain feature maps of the clean audio and the noisy audio in each sample pair; Based on the frequency domain feature map, the network parameters of the denoising network of the state recognition model are updated to obtain the denoising network with the first update.
2. The method according to claim 1, characterized in that, The step involves constructing multiple sample pairs for the meta-learning task based on the multiple clean audio clips, according to the noise category and signal-to-noise ratio corresponding to the meta-learning task. Select a preset number of clean audio tracks from the plurality of clean audio tracks; Based on the noise category and signal-to-noise ratio corresponding to the meta-learning task, noise is added to the preset number of clean audio segments to obtain the preset number of noise-banded frequencies. Based on the preset number of clean audio samples and the preset number of noisy audio samples, a preset number of sample pairs are constructed for the meta-learning task.
3. The method according to claim 1, characterized in that, The step of obtaining the frequency domain feature maps of the clean audio and noisy frequencies in each sample pair includes: The Mel spectrogram of the pure audio in each sample pair is obtained, and the Mel spectrogram of the pure audio is processed into a single-channel grayscale image to obtain the frequency domain feature map of the pure audio. The Mel spectrum of the noisy frequency in each sample pair is obtained, and the Mel spectrum of the noisy frequency is processed into a single-channel grayscale image to obtain the frequency domain feature map of the noisy frequency.
4. The method according to claim 2, characterized in that, The step of updating the network parameters of the denoising network of the state recognition model based on the frequency domain feature map to obtain the initially updated denoising network includes: Based on the frequency domain feature map and according to the selected meta-learning training method, the initial network parameters of the noise reduction network are updated to obtain the noise reduction network with the first update. The noise reduction network in the initial update has optimized initial network parameters.
5. The method according to claim 4, characterized in that, In response to the meta-learning training method being a model-independent meta-learning training method, the method further includes: For any meta-learning task, select a first number of sample pairs from the preset number of sample pairs corresponding to it to obtain the support set of the meta-learning task; The remaining second number of sample pairs are used as the query set for the meta-learning task; the sum of the first number and the second number is equal to the preset number.
6. The method according to claim 5, characterized in that, The step of updating the initial network parameters of the denoising network based on the frequency domain feature map and according to the selected meta-learning training method includes: In each round of training, a target number of meta-learning tasks are selected from the multiple meta-learning tasks; For any selected meta-learning task, the initial network parameters are updated based on the support set of the meta-learning task to obtain the temporarily updated denoising network; wherein the temporarily updated denoising network has temporary network parameters adapted to the meta-learning task. Obtain the temporarily updated denoising loss of the denoising network on the query set of the meta-learning task; The meta-loss is obtained based on the denoising loss corresponding to each meta-learning task in the plurality of meta-learning tasks, and the initial network parameters are updated based on the meta-loss. Repeat the steps of selecting a target number of meta-learning tasks and updating the network parameters of the denoising network until the training termination condition is met.
7. The method according to claim 6, characterized in that, The update of the initial network parameters based on the support set of the meta-learning task includes: Starting with the initial network parameters, multiple gradient update operations are performed on the support set of the meta-learning task.
8. The method according to claim 7, characterized in that, The step of performing multiple gradient update operations on the support set of the meta-learning task includes: Based on the support set of the meta-learning task, the noise reduction loss of the noise reduction network on the support set is obtained; Based on the initial network parameters, the first learning rate, the gradient operator, and the denoising loss, multiple parameter update operations are performed on the denoising network; wherein, the gradient operator is an operator that calculates the gradient of the initial network parameters.
9. The method according to claim 6, characterized in that, The step of obtaining the temporarily updated denoising loss of the denoising network on the query set of the meta-learning task includes: Obtain the frequency domain feature map of each sample pair with noise in the query set to obtain the frequency domain feature map before noise reduction; The noise reduction network is temporarily updated to perform noise reduction processing on the frequency domain feature map before noise reduction, so as to obtain the frequency domain feature map after noise reduction. Based on the denoised frequency domain feature map and the frequency domain feature map of the clean audio of each sample pair in the query set, the average denoising loss of the temporarily updated denoising network on the query set is obtained.
10. The method according to claim 6, characterized in that, The step of obtaining the meta-loss based on the denoising loss corresponding to each of the plurality of meta-learning tasks, and updating the initial network parameters based on the meta-loss, includes: The sum of the average noise reduction losses corresponding to each of the multiple meta-learning tasks is taken as the meta-loss. The initial network parameters are updated based on the second learning rate and the meta-loss gradient; wherein the meta-loss gradient is the gradient of the meta-loss with respect to the initial network parameters.
11. A state recognition method, characterized in that, The method further includes: Acquire audio to be processed collected at the operation site, wherein the operation site includes the target locomotive compressor for which status identification is to be performed, and the audio to be processed includes sound signals generated by the target locomotive compressor during operation; Based on the frequency domain feature map of the audio to be processed, the network parameters of the initially updated noise reduction network are updated again to obtain the noise reduction network updated again. The frequency domain feature map of the audio to be processed is divided into a training set and a test set; The noise reduction network is updated again, and noise reduction processing is performed on the frequency domain feature maps of the training set and the test set to obtain the noise-reduced training set and the noise-reduced test set. Based on the denoised training set, the network parameters of the classification network are updated to obtain the updated classification network; The noise-reduced test set is input into the updated classification network, and the updated classification network is used to identify the operating status of the target locomotive compressor to obtain the operating status identification result.
12. The method according to claim 11, characterized in that, The noise reduction network in the initial update has optimized initial network parameters; the subsequent update of the network parameters of the initially updated noise reduction network based on the frequency domain feature map of the audio to be processed includes: Starting with the optimized initial network parameters, multiple gradient update operations are performed on the frequency domain feature map of the audio to be processed.
13. The method according to claim 11, characterized in that, The process of updating the network parameters of the classification network based on the denoised training set includes: Since the frequency domain feature map in the training set after noise reduction is in single-channel form, channel expansion is performed on the single-channel form of the frequency domain feature map to obtain a three-channel form of the frequency domain feature map. The network parameters of the classification network are updated based on the frequency domain feature map in the form of the three channels.
14. A training device for a state recognition model, characterized in that, The device includes: The first acquisition module is configured to acquire multiple clean audio signals when the locomotive compressor is in operation, wherein the clean audio signals are collected in a laboratory environment during the operation of the locomotive compressor; The first construction module is configured to construct multiple meta-learning tasks based on the noise category and preset signal-to-noise ratio in the locomotive operating environment; wherein, each meta-learning task corresponds to a noise category and a signal-to-noise ratio. The second construction module is configured to, for any meta-learning task, construct multiple sample pairs for the meta-learning task based on the multiple clean audios according to the noise category and signal-to-noise ratio corresponding to the meta-learning task, wherein each sample pair includes a clean audio and a corresponding noisy audio. The second acquisition module is configured to acquire frequency domain feature maps of the clean audio and noisy frequencies in each sample pair; The first training module is configured to update the network parameters of the denoising network of the state recognition model based on the frequency domain feature map, so as to obtain the denoising network with the first update.
15. A state recognition device, characterized in that, The device further includes: The third acquisition module is configured to acquire audio to be processed collected at the operating site, the operating site including the target locomotive compressor to be identified in terms of its status, and the audio to be processed including the sound signal generated by the target locomotive compressor during its operation. The second training module is configured to update the network parameters of the initially updated noise reduction network based on the frequency domain feature map of the audio to be processed, so as to obtain the noise reduction network updated again. The data partitioning module is configured to divide the frequency domain feature map of the audio to be processed into a training set and a test set; The noise reduction module is configured to perform noise reduction processing on the frequency domain feature maps of the training set and the test set through the noise reduction network that has been updated again, so as to obtain the noise-reduced training set and the noise-reduced test set. The third training module is configured to update the network parameters of the classification network based on the denoised training set, thereby obtaining the updated classification network. The identification module is configured to input the noise-reduced test set into the updated classification network, and to identify the operating status of the target locomotive compressor through the updated classification network to obtain the operating status identification result.
16. A computer device, characterized in that, The device includes a processor and a memory, the memory storing computer program code, the computer program code being loaded and executed by the processor to implement the training method of the state recognition model as described in any one of claims 1 to 10, or the state recognition method as described in any one of claims 11 to 13.
17. A computer-readable storage medium, characterized in that, The storage medium stores computer program code, which is loaded and executed by a processor to implement the training method of the state recognition model as described in any one of claims 1 to 10, or the state recognition method as described in any one of claims 11 to 13.
18. A computer program product, characterized in that, The computer program product includes computer program code stored in a computer-readable storage medium. A processor of a computer device reads the computer program code from the computer-readable storage medium and executes the computer program code, causing the computer device to perform a training method for a state recognition model as described in any one of claims 1 to 10, or a state recognition method as described in any one of claims 11 to 13.