A method and system for constructing an enhanced model of server voiceprint data, an enhancement method, an electronic device, and a storage medium
Patent Information
- Application Number
- CN202610971671.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-01
- Publication Date
- 2026-09-25
AI Technical Summary
通常服务器故障包括风扇轴承磨损、叶片裂纹等,然而在服务器运行时的声纹数据分析过程中依赖大量带标签的故障声纹数据,在实际运维中故障声纹数据稀缺,占比<5%,导致服务器故障预警系统泛化能力差
[0024]通过从样本库中选取故障声纹数据及其对应的工况参数,其中每组故障声纹数据与对应的工况参数的映射关系,使通过增强模型生成的高保真伪时频谱图可精确匹配工况参数,极大增强了生成的高保真伪时频谱图与真实应用场景的适配性;对故障声纹数据进行傅里叶变换后可以得到初始时频谱图,对初始时频谱图进行归一化处理后得到可以使用的声纹时频谱图。首先通过声纹时频谱图和对应的工况参数可以对深度学习模型中的分类器进行预训练,将分类器训练好之后,再进行深度学习模型的增强训练,使得到的增强模型更加稳定,增强得到的数据更加准确可靠,便于后续使用和处理。解决了现有技术中未针对服务器噪声特点设计,生成的数据与真实数据存在域差异,工况参数不匹配的问题。
Smart Images

Figure CN122821986A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data augmentation, specifically to a method, system, augmentation method, electronic device, and storage medium for constructing an augmentation model for server voiceprint data. Background Technology
[0002] Analyzing the acoustic signature data of a server during operation can provide early warnings of server failures. Common server failures include fan bearing wear and blade cracks. However, acoustic signature data analysis relies on a large amount of tagged fault acoustic signature data, which is scarce in actual operation and maintenance, accounting for less than 5%, resulting in poor generalization ability of the server failure early warning system. To address this issue, methods such as adding noise, time stretching, and pitch transformation are commonly used. However, these methods only alter surface features and cannot generate true fault modes.
[0003] To address the aforementioned issues, a method using basic generative adversarial networks (GANs) to generate fault voiceprint data has been proposed. However, this method still suffers from the following problems: 1. It is primarily used for speech synthesis and is not designed specifically for server noise characteristics, resulting in domain differences between the generated data and real data, and mismatches in operating parameters; 2. It does not consider the aliasing characteristics of multi-source noise on servers. Using GANs to generate bearing fault signals is only suitable for ideal scenarios with a single fault (specifically, determining whether there is a fan bearing wear fault in a clean environment), and cannot meet the needs of real-world scenarios (server chassis with fan, hard drive, and power supply noise); 3. It lacks physical consistency, and the generated voiceprints may violate the laws of sound wave propagation, making the model unusable in real-world applications; 4. It suffers from insufficient sample diversity, making it impossible to generate samples under different combinations of operating parameters; 5. It has poor adaptability to small samples, and when the number of fault voiceprint data is extremely small, it is impossible to train a stable model, leading to a decrease in generation quality. Summary of the Invention
[0004] To address one of the aforementioned technical deficiencies, this application provides a method, system, enhancement method, electronic device, and storage medium for constructing an enhanced model of server voiceprint data.
[0005] According to a first aspect of this application, a method for constructing an enhanced model of server voiceprint data is provided, comprising:
[0006] Multiple fault acoustic fingerprint data of servers and their corresponding operating parameters were selected from the sample library; the operating parameters include fan speed, CPU load, ambient temperature, fault type, and server model.
[0007] Perform a Fourier transform on each fault voiceprint data to obtain the initial time spectrum; normalize the initial time spectrum to obtain the voiceprint time spectrum.
[0008] The classifier in the deep learning model is pre-trained by multiple time-spectral maps of voiceprints and corresponding operating parameters to obtain a pre-trained classifier and a pre-trained deep learning model.
[0009] The pre-trained deep learning model is enhanced by using multiple audioprint time-spectrum maps and corresponding operating parameters to obtain an enhanced model.
[0010] According to a second aspect of this application, a system for constructing an enhanced model of server voiceprint data is provided, including modules for implementing the method for constructing an enhanced model of server voiceprint data as described above.
[0011] According to a third aspect of this application, a method for enhancing server voiceprint data is provided, comprising:
[0012] Acquire the fault acoustic signature data to be enhanced, along with its corresponding operating parameters and the expected number of samples to be generated;
[0013] Perform Fourier transform on the enhanced fault acoustic data to obtain the initial fault time spectrum; normalize the initial fault time spectrum to obtain the fault acoustic data time spectrum.
[0014] The operating parameters and expected number of fault voiceprint data to be enhanced are input into the enhancement model, which is an enhancement model constructed using the aforementioned method for constructing enhancement models of server voiceprint data.
[0015] The model is enhanced to output the corresponding high-fidelity pseudo-time spectrum.
[0016] The faulty voiceprint time spectrum map and the corresponding high-fidelity pseudo-time spectrum map are used to construct the enhanced voiceprint dataset.
[0017] According to a fourth aspect of this application, an electronic device is provided, comprising:
[0018] Memory;
[0019] Processor; and
[0020] Computer programs;
[0021] The computer program is stored in the memory and is configured to be executed by the processor to implement the aforementioned method for constructing an enhanced model of server voiceprint data, or to be configured to be executed by the processor to implement the aforementioned method for enhancing server voiceprint data.
[0022] According to a fifth aspect of this application, a computer-readable storage medium is provided having a computer program stored thereon; the computer program is executed by a processor to implement a method for constructing an enhanced model of server voiceprint data as described above, or to implement a method for enhancing server voiceprint data as described above.
[0023] The beneficial effects of this application are as follows:
[0024] By selecting fault acoustic signature data and their corresponding operating parameters from a sample library, and establishing the mapping relationship between each set of fault acoustic signature data and its corresponding operating parameters, the high-fidelity pseudo-time spectrum generated by the enhancement model can accurately match the operating parameters, greatly enhancing the adaptability of the generated high-fidelity pseudo-time spectrum to real application scenarios. After performing a Fourier transform on the fault acoustic signature data, an initial time spectrum is obtained. Normalizing this initial time spectrum yields a usable acoustic signature time spectrum. Firstly, the classifier in the deep learning model can be pre-trained using the acoustic signature time spectrum and corresponding operating parameters. After the classifier is trained, the deep learning model undergoes enhancement training, resulting in a more stable enhancement model and more accurate and reliable enhanced data, facilitating subsequent use and processing. This addresses the problems in existing technologies where the design did not specifically address server noise characteristics, leading to domain differences between the generated data and real data, and mismatched operating parameters.
[0025] Other features and advantages of this application will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing the application. The objectives and other advantages of this application may be realized and obtained by means of what is pointed out in the written description and the accompanying drawings. Attached Figure Description
[0026] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments of this application and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0027] Figure 1 A flowchart illustrating a method for constructing an enhanced model of server voiceprint data provided in this application;
[0028] Figure 2 The pre-training flowchart of the classifier provided in this application;
[0029] Figure 3 The flowchart for the augmented training provided in this application;
[0030] Figure 4 Flowcharts of the alternating optimization discriminator and generator provided in this application;
[0031] Figure 5A flowchart illustrating the acquisition of the first pseudo-time spectrum provided in this application;
[0032] Figure 6 A flowchart for calculating the physical constraint loss value provided in this application;
[0033] Figure 7 A schematic diagram of the functional structure of a system for constructing an enhanced model of server voiceprint data provided in this application;
[0034] Figure 8 for Figure 7 A schematic diagram of the functional structure of the enhanced training module;
[0035] Figure 9 for Figure 8 Functional structure diagram of the alternating optimization unit;
[0036] Figure 10 for Figure 7 A schematic diagram of the functional structure of the classifier pre-training module;
[0037] In the picture:
[0038] 10 is the selection module, 20 is the transformation module, 30 is the normalization module, 40 is the classifier pre-training module, 50 is the augmentation training module, 401 is the partitioning unit, 402 is the pre-training unit, 403 is the validation unit, 404 is the return unit, 405 is the classifier freezing unit, 501 is the alternating optimization unit, 502 is the pseudo-time spectrogram generation unit, 503 is the classification unit, 504 is the weight update unit, 505 is the judgment unit, 506 is the training termination unit, 507 is the loop execution unit, and 5011 is the discriminator optimization unit. 5012 is the generator optimization unit, 50111 is the first construction unit, 50112 is the first generation unit, 50113 is the image definition unit, 50114 is the discriminator loss calculation unit, 50115 is the discriminator optimization unit, 50116 is the discriminator freezing unit, 50121 is the second construction unit, 50122 is the second generation unit, 50123 is the adversarial loss calculation unit, 50124 is the physical constraint loss calculation unit, 50125 is the generator loss calculation unit, and 50126 is the generator optimization unit. Detailed Implementation
[0039] To make the technical solutions and advantages of the embodiments of this application clearer, the exemplary embodiments of this application will be described in further detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not an exhaustive list of all embodiments. It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of this application can be combined with each other.
[0040] Example 1
[0041] In Example 1, the fault type is "fan bearing wear".
[0042] like Figure 1 As shown, in response to the above problems, the first aspect of Embodiment 1 of this application provides a method for constructing an enhanced model of server voiceprint data, including:
[0043] Select multiple fault voiceprint data of the server from the sample library (usually ≥20, except for extremely low sample scenarios, 50 in this case) and their corresponding operating parameters;
[0044] Fourier transform was performed on each fault voiceprint data (Hanning window, window length 2048 sampling points; frame shift: 512 sampling points, overlap rate 75%; FFT number of points: 2048, resulting in a frequency resolution of approximately 46.875Hz). 256 consecutive frames from the center were taken to obtain the initial time-frequency spectrum (initial time-frequency spectrum is 256×256 pixels). The initial time-frequency spectrum was then normalized (the amplitude values of all initial time-frequency spectra were linearly mapped to the [0,1] interval) to obtain the voiceprint time-frequency spectrum (256×256 pixels).
[0045] The classifier (ResNet-18) in the deep learning model is pre-trained using multiple time-spectrum maps of voiceprints and corresponding operating parameters to obtain the pre-trained classifier and the pre-trained deep learning model.
[0046] The pre-trained deep learning model is enhanced by using multiple audioprint time-spectrum maps and corresponding operating parameters to obtain an enhanced model.
[0047] Based on the above scheme, fault acoustic signature data of servers and their corresponding operating parameters are selected from the sample library. The mapping relationship between each set of fault acoustic signature data and the corresponding operating parameters ensures that the high-fidelity pseudo-time spectrum generated by the enhancement model can accurately match the operating parameters, greatly enhancing the adaptability of the generated high-fidelity pseudo-time spectrum to real application scenarios. After performing a Fourier transform on the fault acoustic signature data, an initial time spectrum is obtained. After normalizing the initial time spectrum, a usable acoustic signature time spectrum is obtained. First, the classifier in the deep learning model can be pre-trained using the acoustic signature time spectrum and the corresponding operating parameters. After the classifier is trained, the deep learning model is then enhanced, making the resulting enhancement model more stable and the enhanced data more accurate and reliable, facilitating subsequent use and processing. This solves the problems in existing technologies where the design does not take into account the characteristics of server noise, resulting in domain differences between the generated data and real data, and mismatches in operating parameters.
[0048] In some possible implementations of the first aspect, the sample library includes 50 fault acoustic fingerprint data points of fan bearing wear failure collected from rack servers (1U / 2U rack servers, here 2U rack servers). These data points are obtained by multi-channel fusion processing after being collected by four MEMS microphones (INMP441, sampling rate 96kHz, placed 10cm from the fan outlet). The acquisition time for each fault acoustic fingerprint data point and the corresponding operating condition parameters is 10 seconds. The fan speed, CPU load, and ambient temperature are recorded at a frequency of 1Hz, and the average values are taken respectively.
[0049] Specifically, the multi-channel fusion processing involves fusing the multi-channel signals acquired by the four MEMS microphones using a delay-and-sum beamforming method. This method first calculates the relative time delay of each channel signal based on the geometric position of the microphone array and the direction of the target sound source (fan exhaust). The relative time delay of each channel signal is then compensated to align its phase. Finally, the compensated channel signals are weighted and summed to enhance the sound source signal from the target direction while suppressing environmental noise and interference from other directions.
[0050] In some possible implementations of the first aspect, the operating parameters in the sample library include fan speed, CPU load, ambient temperature, fault type, and server model.
[0051] The operating parameters have 12 dimensions, including:
[0052] Fan speed (1-dimensional) is a normalized value. The specific normalization method is to normalize 0-10000RPM to [0,1].
[0053] CPU load (1-dimensional) is a normalized value. The normalization method is to normalize 0-100% to [0,1]. Specifically, it includes three cases: low load, medium load and high load.
[0054] Table 1 Classification of CPU Load
[0055]
[0056] Ambient temperature (1-dimensional) is a normalized value. The specific normalization method is to normalize 20-50℃ to [0,1].
[0057] Fault types (8 dimensions) are uniquely encoded (in a binary vector of length 8, only the dimension corresponding to the fault type is 1, and the other 7 dimensions are all 0), including: fan bearing wear encoding, blade crack encoding, magnetic head abnormal noise encoding, power supply whistling encoding, fan imbalance encoding, capacitor aging encoding, connector loosening encoding, and no fault encoding.
[0058] Table 2 Fault Type Description Table 1
[0059]
[0060] Server model identifier (1D), model code, set to an integer value.
[0061] Specifically, in this first embodiment, the fan speed is 5400 RPM, the CPU load is 80%, and the ambient temperature is 45℃; the processed operating parameters are: fan speed [0.54], CPU load [0.8], ambient temperature [0.67]; fault type [1, 0, 0, 0, 0, 0, 0], and server model identifier [0] (fixed value, representing model KunPeng920).
[0062] Based on the above scheme, fan speed, CPU load, and ambient temperature are all normalized values. Fault types are encoded using one-hot encoding, and server model identification uses model encoding. These methods facilitate the identification and further processing of operating parameters, making the construction of the augmented model more time-efficient. This solves the problem of insufficient sample diversity in existing technologies, which prevents the generation of samples under different combinations of operating parameters.
[0063] like Figure 2 As shown, in some possible implementations of the first aspect, the step of pre-training the classifier in the deep learning model using multiple voiceprint time-spectrum maps and corresponding operating parameters to obtain a pre-trained classifier specifically includes:
[0064] Multiple audioprint time-spectrum maps and corresponding operating parameters are divided into training set and validation set;
[0065] Pre-train the classifier in the deep learning model using the training set;
[0066] The classification results of the classifier are validated using a validation set;
[0067] If the accuracy of the validation result is less than 90%, return to the training set to pre-train the classifier in the deep learning model for the next round.
[0068] When the accuracy of the verification results is not less than 90%, the classifier parameters are frozen to obtain a pre-trained classifier.
[0069] In some possible implementations of the first aspect, the classifier in the deep learning model is pre-trained using a training set (this part is existing technology, where the time-spectrum image of the voiceprint and the corresponding operating parameters can be directly input), specifically including:
[0070] Based on the audioprint time spectrum, the prediction score is calculated through forward propagation.
[0071] The predicted scores are transformed using the Softmax function to obtain the predicted probabilities of the classifier.
[0072] Cross-entropy loss is calculated based on the predicted probability and the corresponding fault type in the operating parameters. The calculation formula is: ;
[0073] In the formula, A represents the number of spectrograms in the voiceprints in the training set. (Corresponding to eight fault types) This represents the true label (0 or 1) of the fault type b in the operating condition parameters corresponding to the spectrum of the a-th voiceprint. Let be the probability predicted by the classifier that the spectrum of the a-th voiceprint contains the b-th type of fault.
[0074] The parameter update gradient is obtained by backpropagation based on the cross-entropy loss, and the classifier parameters are updated based on the parameter update gradient.
[0075] Based on the above scheme, a training set and a validation set were created. The classifier was pre-trained using the training set, and then the classification results were validated using the validation set. When the accuracy of the validation results was not less than 90%, it indicated that the classifier's classification results were relatively accurate, and the classifier parameters were frozen, resulting in the pre-trained classifier. When the accuracy of the validation results was less than 90%, it indicated that the classifier's classification results were inaccurate and the classification performance was poor. In this case, a new round of pre-training was performed using the training set until the accuracy of the validation results was not less than 90%.
[0076] like Figure 3 As shown, in some possible implementations of the first aspect, the deep learning model is a conditional generative adversarial network (cGAN) model, comprising a generator (U-Net architecture) and a judge (PatchGAN).
[0077] The process involves enhancing the pre-trained deep learning model using multiple time-spectrum audiograms and corresponding operating parameters (GPU acceleration (NVIDIA V100 / A100 or equivalent computing power) is recommended during enhancement training, with a video memory requirement of ≥16GB) to obtain the enhanced model, specifically including:
[0078] The discriminator and generator are alternately optimized by using multiple time-spectrum diagrams of acoustic signatures and corresponding operating parameters.
[0079] After each k iterations (k is 10 in this embodiment) of alternating optimization, multiple operating parameters are selected and input into the generator to generate multiple pseudo-time spectrum diagrams.
[0080] The classification results are obtained by classifying multiple pseudo-time spectrograms using a trained classifier.
[0081] Update the input weights of the operating condition parameters in the input generator based on the classification results;
[0082] Determine whether the total number of optimization attempts of the generator has reached the preset training threshold or whether the total loss of the generator has converged.
[0083] If so, then training ends and a trained augmented model is obtained;
[0084] Otherwise, repeat the above steps.
[0085] Based on the above scheme, a conditional generative adversarial network (GAN) model is used as a deep learning model for training the augmentation model. The specific training process is as follows: the discriminator and generator are alternately optimized using the time-spectrum image of the voiceprint and corresponding operating parameters (alternating optimization is typically set to optimize the discriminator at least twice and the generator once); then periodic classification is performed. Every k cycles of alternating optimization, a hard example weight adjustment is triggered (the confidence value of the hard example is lower than a preset threshold; the hard example weight adjustment is based on the classification results of the trained classifier, and is set to increase the weight of the hard example). This process is repeated until the total number of optimizations for the generator (200 in this embodiment) reaches the preset training threshold or the generator's total loss converges. This method allows the generator to input hard examples more frequently, forcing it to focus on generating "difficult" samples that easily confuse the classifier. This solves the problem that generated samples are often concentrated in easily classified regions, improves the diversity and boundary quality of the generated data, and addresses the technical problem of insufficient boundary samples.
[0086] like Figure 4 As shown, in some possible implementations of the first aspect, the alternating optimization of the discriminator and generator using multiple time-spectrum maps of voiceprints and corresponding operating parameters specifically includes:
[0087] Optimize the discriminator, specifically including:
[0088] Multiple operating parameters (used to control the attributes of the generated sample data) are selected and constructed into the corresponding first condition vector (112-dimensional) by combining them with a real-time acquired random noise vector (100-dimensional, following a standard Gaussian distribution N(0,1), which is independently sampled using a deep learning architecture (PyTorch) or a symbolic mathematics system (TensorFlow) each time the generator is called, to provide the randomness required for the diversity of the generated sample data).
[0089] Multiple first condition vectors are input into the generator for processing to obtain multiple first pseudo-time spectrum maps;
[0090] The acoustic signature time spectrum and the corresponding first pseudo-time spectrum of the selected multiple operating parameters are used as the original image, and the image obtained by downsampling the original image is used as the preprocessed image.
[0091] Each original image and its corresponding preprocessed image are input into the discriminator to calculate the loss value for each original image, and the total loss value of the discriminator is calculated using the loss value of each original image.
[0092] Based on the total loss value of the discriminator, the discriminator parameters are optimized by the optimizer to obtain the optimized discriminator;
[0093] After optimizing the discriminator M times (twice in this embodiment), the discriminator parameters are frozen; and the generator is optimized according to the following steps:
[0094] Multiple operating parameters are selected and constructed into a corresponding second condition vector (112-dimensional) by combining them with a real-time acquired random noise vector (100-dimensional, following a standard Gaussian distribution N(0,1)).
[0095] Multiple second condition vectors are input into the generator for processing to obtain multiple second pseudo-time spectrum maps;
[0096] The second pseudo-time spectrum corresponding to the selected multiple operating parameters is input into the optimized discriminator to calculate the adversarial loss value (the discriminator can determine the probability of classifying the second time spectrum as real; the higher the probability, the smaller the adversarial loss value).
[0097] Based on the acoustic waveform time spectrum corresponding to the selected multiple operating parameters, the physical constraint loss value (i.e. the degree of deviation from the acoustic wave equation) is calculated according to the residual calculation formula of the wave equation.
[0098] The total loss value of the generator is calculated by combining the adversarial loss value and the physical constraint loss value.
[0099] Based on the total loss of the generator, the generator parameters are optimized using an optimizer (Adam optimizer, with a learning rate of 0.0002 and an exponential decay rate β1=0.5) to obtain the optimized generator.
[0100] Based on the above scheme, specific optimization methods for the discriminator and generator are provided. When optimizing the discriminator, multiple operating parameters are first selected and combined with real-time acquired random noise vectors to construct the corresponding first condition vector. Then, the generator generates the corresponding first pseudo-time spectrum. The original image (the voiceprint time spectrum corresponding to the operating parameters and the first pseudo-time spectrum) and the preprocessed image are input into the discriminator. The discriminator can determine the authenticity of the input original and preprocessed images, and can also calculate the loss value corresponding to each original image (including the loss value of the genuine voiceprint time spectrum and the loss value of the fake first pseudo-time spectrum). Thus, the total loss value of the discriminator can be calculated. The total loss value of the discriminator is used as the basis for optimizing the discriminator parameters (backpropagation). The optimized discriminator is obtained through optimization. After optimizing the discriminator M times, the discriminator parameters are frozen, and the generator is optimized according to the following steps. When optimizing the generator, operating parameters are selected and combined with real-time acquired random noise vectors to construct a second conditional vector. The generator then generates a corresponding second pseudo-time spectrum. The adversarial loss value is calculated using the optimized discriminator, and the physical constraint loss value is calculated according to the fluctuation loss formula. This yields the generator's total loss value, which is used as the basis for optimizing the generator parameters (backpropagation). The optimized generator is then obtained through optimization. During this alternating optimization process, the generator parameters are frozen when optimizing the discriminator, and vice versa. The real-time acquired random noise vector provides the generator with "seed" randomness, allowing for the generation of diverse first and second pseudo-sample time spectra even under the same operating parameters due to different random noise vectors. An optimization frequency of M times for the discriminator and once for the generator enables stable adversarial training, improving the stability of the trained augmented model. The influence of the physical constraint loss value is considered, addressing the problem in existing technologies where lack of physical consistency and the potential for generated soundprints to violate sound wave propagation laws, rendering the model unusable, are issues.
[0101] In some possible implementations of the first aspect, the random noise vector is acquired in real time by dynamically sampling it at each generation; the distribution type is a standard Gaussian distribution N(0, 1) (mean 0, variance 1); and the dimension is 100.
[0102] like Figure 5 As shown, in some possible implementations of the first aspect, the step of inputting multiple first condition vectors into a generator for processing to obtain multiple first pseudo-time spectrum diagrams specifically includes:
[0103] 1. Reshape the first conditional vector (112-dimensional) into an initial feature map (represented as 16×16×512, where the size is 16×16 and the number of channels is 512) through a fully connected mapping layer.
[0104] The initial feature map is processed sequentially by D (D is 4 in this embodiment) downsampling convolutional layers in the encoder module to obtain a high-dimensional feature map. Each downsampling convolutional layer has a kernel size of 4×4, a stride of 2, and padding of 1. The number of channels changes sequentially as follows: 512→1024→2048→4096→4096 (wherein, the final change in the number of channels from 4096 to 4096 is a "bottleneck truncation" strategy commonly used in large U-Net architectures (similar to the bottleneck design in ResNet), which aims to maintain model performance while ensuring that training can run smoothly on a conventional industrial-grade GPU (such as NVIDIA V100 32GB).
[0105] Specifically, batch normalization (BatchNorm) can be used to standardize the input of each downsampled convolutional layer, thereby stabilizing the data distribution, accelerating training convergence, and improving network performance.
[0106] The formula for batch normalization is: ;
[0107] In the formula, This represents the mean of the initial feature maps in the current batch. This represents the variance of the initial feature map in the current batch. Represents a small constant (default is ). ), to prevent division by zero, This represents the learnable scaling parameter (scale). This represents the learnable shift parameter. The batch division is determined by the batch dimension and spatial dimension, and each channel is calculated independently. x represents all initial feature maps for a given channel.
[0108] Furthermore, after batch normalization, it is determined whether the input of the downsampled convolutional layer is negative;
[0109] If not, then directly input it into the downsampling convolutional layer;
[0110] If so, the negative value is processed by the activation function (LeakyReLU(0.2), which is an improvement on the standard ReLU), and then the processed negative value is input into the downsampling convolutional layer. When the input is negative, a small gradient is allowed to pass through, thereby avoiding the "neuron death" problem.
[0111] The formula for processing the values in the input downsampling convolutional layer is: ;
[0112] In the formula, In this embodiment, the slope is represented as follows: This means that when the input is negative, the output = input × 0.2.
[0113] 2. The high-dimensional feature map is processed sequentially by D transposed convolutional layers in the decoder module to obtain a low-dimensional feature map. Each transposed convolutional layer concatenates the feature map of the corresponding downsampling convolutional layer in the encoder with the feature map of the corresponding transposed convolutional layer in the decoder through skip connections. Then, the dimensionality of the concatenated feature map is reduced. The number of channels changes from 4096 to 6144 to 4096 to 2048 to 1024.
[0114] The low-dimensional feature map is processed by the final upsampling module to obtain the first pseudo-time spectrum. The activation function of the final upsampling module (kernel size 3×3 pixels, stride 1 (pixel stride), padding 1 pixel) is Tanh, the number of channels of the first pseudo-time spectrum is 1 (channel number), and the output size is 256×256 pixels.
[0115] Similarly, the specific steps of inputting multiple second condition vectors into the generator for processing to obtain multiple second pseudo-time spectrum diagrams correspond to the steps described above. For the purpose of saving space and keeping it concise, they will not be repeated here.
[0116] In some possible implementations of the first aspect, the preprocessed image includes a first preprocessed image downsampled by 2 times and a second preprocessed image downsampled by 4 times;
[0117] The process of inputting each original image and its corresponding preprocessed image into the discriminator to calculate the loss value for each original image, and then calculating the total loss value of the discriminator using the loss value of each original image, specifically includes:
[0118] Each original image and its corresponding preprocessed image are input into the discriminator, which calculates the probability value of each original image.
[0119] The loss value of each original image is calculated by using the probability value of each original image and a preset label (in this embodiment, the preset label is 1 when the original image is a voiceprint spectrum and 0 when the original image is a first pseudo-time spectrum).
[0120] The total loss value of the discriminator is calculated using the loss value of each original image.
[0121] Specifically, the discriminator includes three sub-discriminators of different scales (D1: 16×16, D2: 32×32, D3: 64×64).
[0122] The step of inputting each original image and its corresponding preprocessed image into the discriminator, and calculating the probability value of each original image by the discriminator, specifically includes:
[0123] Step 1: Multi-scale input construction
[0124] The original image (256×256), the first preprocessed image (128×128), and the second preprocessed image (64×64) are used as the input images for the three sub-discriminators D1, D2, and D3, respectively.
[0125] Step 2: Parallel Convolution Processing
[0126] Each sub-discriminator is a 5-layer convolutional network that extracts features from the corresponding input image and generates probability maps of 30×30, 14×14, and 6×6, respectively. Each value on the probability map represents the probability that the corresponding input image is a real sample.
[0127] Step 3: Conditional Vector Injection
[0128] In the intermediate layer of each sub-discriminator (at the output of the 4th convolution layer), the operating condition parameters (12-dimensional) corresponding to the input image are linearly projected and added element-wise to the two-dimensional matrix (with channel number) processed by the intermediate layer of the sub-discriminator. This allows the discriminator to simultaneously verify the authenticity of the input image and whether it matches the operating condition parameters and the fault type (probability p→1 when the input image is a real sample and matches; probability p→0 when the input image is a generated sample and does not match).
[0129] Step 4: Probability Fusion
[0130] Global average pooling is performed on the generated probability maps to obtain three scalar probability values. The arithmetic mean of the three scalar probability values is then taken as the probability value of each original image.
[0131] The formula for calculating the probability value of the original image is:
[0132] ;
[0133] ;
[0134] In the formula, , These represent the probability values that the original image is the i-th real sample (voiceprint time spectrum) and the j-th generated sample (first pseudo-time spectrum), respectively. The values are both in the range of [0,1]. The closer to 1, the higher the probability that the input image is a real sample. , , Let D1, D2, and D3 represent the probability values corresponding to the i-th true sample output by sub-discriminators D1, D2, and D3, respectively. , , These represent the probability values corresponding to the j-th generated sample output by sub-discriminators D1, D2, and D3, respectively. In this embodiment, the weights of the three sub-discriminators are fixed and equal (here, 1 / 3) when performing the arithmetic average.
[0135] More specifically, each sub-discriminator is a 5-layer convolutional network. Taking one sub-discriminator as an example, its 5 convolutional layers are as follows:
[0136] First layer: The 3-channel input is convolved by Conv2D(3,64,4,2) (4×4 kernel, stride 2), and then directly connected to the LeakyReLU activation function, without a normalization layer;
[0137] The second layer: the 64-channel input is convolved by Conv2D(64,128,4,2) (4×4 kernel, stride 2), followed by the InstanceNorm normalization layer and the LeakyReLU activation function in sequence;
[0138] The third layer: The 128-channel input is convolved by Conv2D(128,256,4,2) (4×4 kernel, stride 2), followed by the InstanceNorm normalization layer and the LeakyReLU activation function in sequence;
[0139] Fourth layer: The 256-channel input is convolved by Conv2D(256,512,4,1) (4×4 kernel, stride 1), followed by the InstanceNorm normalization layer and the LeakyReLU activation function in sequence;
[0140] Fifth layer: The 512-channel input is convolved by Conv2D(512,1,4,1) (convolution kernel 4×4, stride 1), and then directly connected to the Sigmoid activation function to output a single-channel probability map.
[0141] Based on the above scheme, the number of channels in the entire sequence changes from 3→64→128→256→512→1, and the step size changes from 2→2→2→1→1. Except for the first and last layers, all other layers include the InstanceNorm normalization operation.
[0142] In some possible implementations of the first aspect, the loss value of each original image is calculated by using the probability value of each original image and a preset label (in this embodiment, the preset label is 1 when the original image is a voiceprint spectrogram, and 0 when the original image is a first pseudo-time spectrogram), specifically including:
[0143] The formula for calculating the loss value corresponding to the i-th real sample in the original image is: ;
[0144] In the formula, This represents the loss value of the sub-discriminator Dz corresponding to the i-th real sample. , This represents the probability value corresponding to the i-th real sample output by the sub-discriminator Dz;
[0145] The formula for calculating the loss value corresponding to the j-th generated sample in the original image is: ;
[0146] In the formula, This represents the loss value of the sub-discriminator Dz corresponding to the j-th generated sample. This represents the probability value corresponding to the j-th generated sample output by the sub-discriminator Dz.
[0147] In some possible implementations of the first aspect, the total loss value of the discriminator is calculated from the loss value of each original image, specifically including:
[0148] The loss value for each sub-discriminator is calculated using the LSGAN (least squares GAN) loss function. The formula for calculating the loss value of each sub-discriminator is as follows: ;
[0149] In the formula, The total loss value corresponding to the sub-discriminator Dz is represented by N, the number of time spectrum maps of the voiceprint is represented by J, and the number of first pseudo-time spectrum maps is represented by J. In this embodiment, the number of time spectrum maps of the voiceprint is the same as the number of first pseudo-time spectrum maps, and N=J.
[0150] Based on the loss value corresponding to each sub-discriminator, the total loss value of the discriminator is obtained by multi-scale fusion; the formula for calculating the total loss value of the discriminator is: ;
[0151] In the formula, , , These represent the total loss values corresponding to sub-discriminators D1, D2, and D3, respectively.
[0152] The following example illustrates the calculation process of the loss value corresponding to sub-discriminator D1:
[0153] With N=J=4 as the default, the loss values of the original image are calculated as shown in Table 3 below:
[0154] Table 3 Loss values corresponding to sub-discriminator D1
[0155]
[0156] The calculation process for the loss value corresponding to sub-discriminator D1 is as follows:
[0157] Total loss for the actual sample: 0.0032 + 0.00845 + 0.00125 + 0.0072 = 0.0201;
[0158] Total loss for generating samples: 0.0512 + 0.0392 + 0.10125 + 0.02205 = 0.2137;
[0159] The loss value corresponding to sub-discriminator D1:
[0160] .
[0161] In some possible implementations of the first aspect, the discriminator parameters are optimized by an optimizer based on the total loss value of the discriminator to obtain an optimized discriminator, specifically including:
[0162] The optimized formula is: ;
[0163] In the formula, This represents the optimized discriminator parameters. This represents the old discriminator parameters. The learning rate is represented in this first embodiment. This is used to control the step size of each discriminator parameter optimization; This represents the gradient (partial derivative vector) of the total loss of the discriminator with respect to the discriminator parameters, indicating the direction and magnitude in which the discriminator parameters should be adjusted.
[0164] Based on the above scheme, this multi-scale mechanism enables the discriminator to capture both local texture details (small receptive field) and global structural consistency (large receptive field), thereby effectively guiding the generator to produce high-fidelity and diverse temporal spectra.
[0165] Similarly, in the specific implementation process, the second pseudo-time spectrum corresponding to the selected multiple operating parameters is input into the optimized discriminator to calculate the adversarial loss value. The specific steps correspond to the steps described above for calculating the loss value of each original image. To save space and keep it concise, they will not be repeated here.
[0166] like Figure 6 As shown, in some possible implementations of the first aspect, the calculation of the physical constraint loss value based on the acoustic waveform time-spectrum diagrams corresponding to the selected multiple operating parameters, according to the fluctuation loss calculation formula, specifically includes:
[0167] The acoustic signature time-frequency spectrum corresponding to the selected multiple operating parameters is inversely transformed into time-domain sound pressure. Where g represents the g-th operating condition parameter, and t represents the time coordinate. Represents spatial coordinates;
[0168] Calculate the second-order partial derivatives of the time-domain sound pressure with respect to the time coordinate. and the second-order partial derivative of time-domain sound pressure with respect to spatial coordinates ;
[0169] The physical constraint loss value is calculated using the volatility loss calculation formula, whereby the volatility loss calculation formula is:
[0170] ;
[0171] In the formula, This represents the physical constraint loss value. G represents the total number of selected operating parameters, and c represents the speed of sound, which is the speed at which sound travels through the air.
[0172] In some possible implementations of the first aspect, the calculation of the generator's total loss value by combining the adversarial loss value and the physical constraint loss value specifically involves: ,in The value range is 0.01-0.1. In Embodiment 1 of this application... , ,in Reduce the physical constraint loss value to At the same time, it does not affect the preservation of fault features in the generated sample data.
[0173] In some possible implementations of the first aspect, the step of optimizing the generator parameters using an optimizer (Adam optimizer, exponential decay rate β1 = 0.5) based on the generator's total loss value to obtain an optimized generator specifically includes:
[0174] The optimized formula is: ;
[0175] In the formula, This represents the optimized generator parameters. This represents the old generator parameters. Indicates the learning rate. , This represents the gradient operator.
[0176] Based on the above scheme, by adding the fluctuation loss calculation formula as a regularization term to the generator's total loss value calculation, it is ensured that the generated sample data conforms to the laws of acoustic physics.
[0177] In some possible implementations of the first aspect, the classification of multiple pseudo-temporal spectrograms using a trained classifier to obtain classification results specifically includes:
[0178] Multiple pseudo-time spectrograms are input into the trained classifier to obtain the classification confidence of each pseudo-time spectrogram;
[0179] When the classification confidence of the pseudo-time spectrum is lower than the preset threshold (which can be set to 0.5-0.7, and here we choose 0.6), the working condition parameters corresponding to the pseudo-time spectrum are classified and marked as difficult cases and stored in the difficult case sample pool.
[0180] Otherwise, the operating parameters corresponding to the pseudo-time spectrum are divided and marked as regular examples, and stored in the regular sample pool.
[0181] Correspondingly, updating the input weights of the operating condition parameters in the input generator based on the classification results specifically includes:
[0182] Increase the input weight of operating condition parameters in the difficult example sample pool (in this first embodiment, the weight is increased by 2 times).
[0183] Then, the input weights of all operating condition parameters in the difficult sample pool and the regular sample pool are uniformly normalized globally.
[0184] Based on the above scheme, the specific method for classifying pseudo-time spectrograms in this application is as follows: The pseudo-time spectrograms are input into a pre-trained classifier to obtain the classification confidence of each pseudo-time spectrogram. When the confidence is below a preset threshold, it is classified as a difficult example; otherwise, it is classified as a normal example. Then, the input weight of difficult examples is increased, and then normalization is performed uniformly. Thus, in subsequent training rounds, the generator more frequently uses difficult examples as input, forcing the generator to focus on generating "difficult" samples that are easily confused by the classifier, improving the diversity and boundary quality of the overall generated data. Through this method, the fault identification accuracy was increased from 72% to 94%.
[0185] In some possible implementations of the first aspect, prior to the step of selecting multiple fault acoustic signature data of the server and their corresponding operating parameters from the sample library, the method further includes:
[0186] Obtain the number of fault voiceprint data of the server to be enhanced in the sample library;
[0187] When the number of fault acoustic fingerprint data of the server to be enhanced is ≥20, the step of selecting multiple fault acoustic fingerprint data of the server and their corresponding operating parameters from the sample library is executed.
[0188] Transfer learning is performed when the number of fault voiceprint data of the server to be enhanced is less than 20.
[0189] Specifically, the steps of transfer learning are as follows:
[0190] 1. Obtain publicly available fault voiceprint data of the same fault type and its corresponding operating parameters, or obtain fault voiceprint data of the same model server and its corresponding operating parameters. Then, based on the aforementioned method for constructing an enhancement model for server voiceprint data (trained 200 times), an initial enhancement model is pre-constructed.
[0191] This stage enables the generator to learn general voiceprint feature extraction capabilities and basic time-spectrum generation capabilities, providing good initial parameters for subsequent fine-tuning.
[0192] 2. Perform traditional data augmentation on the collected fault voiceprint data of the server to be augmented (e.g., 5 samples). Three methods can be used: time stretching (±5%), pitch shift (±1 semitone), and noise addition (SNR=25dB) to increase the number of samples from 5 to 20, forming a fine-tuned training set to prevent overfitting.
[0193] 3. Training the generator, specifically including:
[0194] The parameters of the bottom layers (encoder part, the first 8 convolutional layers) of the generator U-Net in the initial augmentation model are completely frozen (requires_grad=False, indicating that this tensor does not participate in backpropagation calculation); these bottom layers are responsible for extracting general acoustic features such as spectral edges and textures, and are applicable to different server models without needing to be relearned. The parameters of the higher layers of the generator (decoder part, the last 6 layers) remain trainable and are used to learn the fault acoustic texture details specific to the target domain.
[0195] 4. Discriminator and Learning Rate Settings
[0196] All layer parameters of the discriminator remain trainable (not frozen), but the learning rate is reduced to [a lower value]. (1 / 20th of the pre-training phase). The generator's higher layers use the same reduced learning rate. The remaining hyperparameters (optimizer Adam, exponential decay rate β1=0.5, batch size 16) remain consistent with those in the pre-build stage.
[0197] 5. Low-learning-rate fine-tuning training
[0198] The model undergoes 50 fine-tuning iterations using a significantly reduced learning rate. In each training epoch, the discriminator is updated twice and the generator once, with optimizations alternating between the two. This low learning rate strategy ensures that the model does not deviate drastically from the pre-trained general feature representations during fine-tuning.
[0199] 6. Difficult Case Mining and Fine-tuning
[0200] During the fine-tuning training process, hard case mining is still performed once every 10 iterations, but the threshold T is increased to 0.65 (originally 0.6) to compensate for the bias in the pre-classifier confidence estimation caused by the extremely small sample size.
[0201] 7. Model Validation and Storage
[0202] Every 10 iterations, the Frachtert initial distance (FID) and downstream classification F1 score of the generated samples are evaluated on the target domain validation set, and the model with the best performance is selected as the final augmented model and saved.
[0203] 8. Final Result
[0204] In an extreme scenario with only 5 fault voiceprint data points, 200 valid samples were generated, improving the fault detection F1-score from 0.35 to 0.81, thus solving the cold start problem of fault prediction in the early stage of the new model server's launch.
[0205] Based on the above scheme, generator crashes can be avoided, and valid fault samples can be generated stably. This solves the problems of scarce fault samples, poor adaptability with small samples, and the inability to train a stable model and resulting in decreased generation quality when the amount of fault voiceprint data is extremely small.
[0206] It should be understood that although the steps in the flowchart are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order constraint on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the diagram may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0207] In a second aspect of this application, a system for constructing an enhanced model of server voiceprint data is provided, including modules for implementing the method for constructing an enhanced model of server voiceprint data as described above.
[0208] like Figure 7 As shown, in some possible implementations of the second aspect, the system construction includes:
[0209] The selection module 10 is used to select multiple fault acoustic fingerprint data of the server and their corresponding operating parameters from the sample library.
[0210] Transformation module 20 is used to perform Fourier transform on each fault acoustic data to obtain the initial time spectrum.
[0211] Normalization module 30 is used to normalize the initial time spectrum to obtain the voiceprint time spectrum;
[0212] The classifier pre-training module 40 is used to pre-train the classifier in the deep learning model using multiple time-spectral maps of voiceprints and corresponding working conditions, so as to obtain the pre-trained classifier and the pre-trained deep learning model.
[0213] The enhancement training module 50 is used to enhance the training of the pre-trained deep learning model using multiple audioprint time-spectrum maps and corresponding operating parameters to obtain an enhanced model.
[0214] like Figure 8 As shown, in some possible implementations of the second aspect, the deep learning model is a conditional generative adversarial network model, including a generator and a judge; the augmentation training module 50 includes:
[0215] The alternating optimization unit 501 is used to alternately optimize the discriminator and the generator using multiple time-spectrum diagrams of acoustic signatures and corresponding operating parameters.
[0216] The pseudo-time spectrum generation unit 502 is used to select multiple operating parameters and input them into the generator to generate multiple pseudo-time spectrums every k alternating optimizations.
[0217] Classification unit 503 is used to classify multiple pseudo-time spectrograms using a trained classifier to obtain classification results;
[0218] Update weight unit 504, which is used to update the input weights of the working condition parameters in the input generator according to the classification results;
[0219] The judgment unit 505 is used to determine whether the total number of optimizations of the generator has reached the preset training threshold or whether the total loss of the generator has converged.
[0220] Training termination unit 506 is used to terminate training if the total number of optimizations of the generator reaches a preset training threshold or the total loss of the generator converges, thus obtaining the trained augmented model.
[0221] The loop execution unit 507 is used to repeat the above steps if the total number of optimizations of the generator has not reached the preset training threshold or if the total loss of the generator has not converged.
[0222] like Figure 9 As shown, in some possible implementations of the second aspect, the alternating optimization unit 501 includes:
[0223] Discriminator optimization unit 5011, used to optimize the discriminator, specifically includes:
[0224] The first construction unit 50111 is used to select multiple operating condition parameters and construct a corresponding first condition vector by combining them with the real-time acquired random noise vector.
[0225] The first generation unit 50112 is used to input multiple first condition vectors into the generator for processing to obtain multiple first pseudo-time spectrum maps.
[0226] The image definition unit 50113 is used to take the acoustic time spectrum map and the corresponding first pseudo time spectrum map corresponding to the selected multiple working condition parameters as the original image, and the image obtained by downsampling the original image as the preprocessed image.
[0227] The discriminator loss calculation unit 50114 is used to input each original image and the corresponding preprocessed image into the discriminator to calculate the loss value corresponding to each original image, and to calculate the total loss value of the discriminator through the loss value of each original image;
[0228] The discriminator optimization unit 50115 is used to optimize the discriminator parameters through an optimizer based on the total loss value of the discriminator, so as to obtain an optimized discriminator;
[0229] The discriminator freezing unit 50116 is used to freeze the discriminator parameters after optimizing the discriminator M times.
[0230] Generator optimization unit 5012 is used to optimize the generator:
[0231] The second construction unit 50121 is used to select multiple operating condition parameters and construct a corresponding second condition vector by combining them with the real-time acquired random noise vector.
[0232] The second generation unit 50122 is used to input multiple second condition vectors into the generator for processing to obtain multiple second pseudo-time spectrum maps.
[0233] The adversarial loss calculation unit 50123 is used to input the second pseudo-time spectrum corresponding to multiple selected working condition parameters into the optimized discriminator to calculate the adversarial loss value.
[0234] The physical constraint loss calculation unit 50124 is used to calculate the physical constraint loss value based on the acoustic time spectrum corresponding to multiple selected operating parameters and the fluctuation loss calculation formula.
[0235] The generator loss calculation unit 50125 is used to calculate the total loss value of the generator by combining the adversarial loss value and the physical constraint loss value.
[0236] The generator optimization unit 50126 is used to optimize the generator parameters by an optimizer based on the total loss value of the generator, so as to obtain an optimized generator.
[0237] like Figure 10 As shown, in some possible implementations of the second aspect, the classifier pre-training module 40 includes:
[0238] The partitioning unit 401 is used to divide multiple acoustic time-spectrum diagrams and corresponding operating parameters into training sets and validation sets;
[0239] The pre-training unit 402 is used to pre-train the classifier in the deep learning model using the training set;
[0240] The verification unit 403 is used to verify the classification results of the classifier using a verification set;
[0241] Return unit 404 is used to return the classifier in the deep learning model for the next round of pre-training using the training set when the accuracy of the verification result is less than 90%.
[0242] The classifier freezing unit 405 is used to freeze the classifier parameters when the accuracy of the verification result is not less than 90%, so as to obtain a pre-trained classifier.
[0243] In a third aspect of this application, a method for enhancing server voiceprint data is provided (which can be performed on a regular workstation (8GB video memory)), comprising:
[0244] Acquire the fault acoustic signature data to be enhanced, along with its corresponding operating parameters and the expected number of samples to be generated;
[0245] Perform Fourier transform on the enhanced fault acoustic data to obtain the initial fault time spectrum; normalize the initial fault time spectrum to obtain the fault acoustic data time spectrum.
[0246] The operating parameters and expected number of fault voiceprint data to be enhanced are input into the enhancement model, which is an enhancement model constructed using the method described above for constructing an enhancement model for server voiceprint data.
[0247] The model is enhanced to output the corresponding high-fidelity pseudo-time spectrum.
[0248] The faulty voiceprint time spectrum map and the corresponding high-fidelity pseudo-time spectrum map (the number of high-fidelity pseudo-time spectrum maps is consistent with the expected number of generated) are used to construct the enhanced voiceprint dataset.
[0249] Based on the above scheme, by inputting the operating parameters and expected number of fault voiceprint data to be enhanced into the enhancement model, a corresponding number of high-fidelity pseudo-time spectrum maps can be obtained. The fault voiceprint time spectrum map and the corresponding high-fidelity pseudo-time spectrum map are used to construct the enhanced voiceprint dataset, which solves the problem that in the existing technology, fault voiceprint data is scarce in actual operation and maintenance, accounting for less than 5%, resulting in poor generalization ability of the server fault early warning system.
[0250] To verify the accuracy of the enhanced model constructed in Embodiment 1 of this application, 50 fault voiceprint time-frequency spectrograms were extracted, and 50 high-fidelity pseudo-time-frequency spectrograms were extracted from the generated 500 high-fidelity pseudo-time-frequency spectrograms. These 50 fault voiceprint time-frequency spectrograms and 50 high-fidelity pseudo-time-frequency spectrograms were then mixed into a hybrid sample. Three operation and maintenance experts independently judged whether each sample in the hybrid sample was real or generated. The results showed that the average correct recognition rate of the three operation and maintenance experts was only 52.7%, and the subjective realism score of the generated sample (4.3 / 5) was very close to that of the real sample (4.5 / 5). The results indicate that the enhanced model constructed using the server voiceprint data enhancement model method provided in this application generates samples that are highly consistent with real samples in terms of auditory and time-frequency characteristics during use, possessing high fidelity, high diversity, and physical consistency. This model can effectively replace real samples for training downstream fault prediction models.
[0251] In a fourth aspect of the embodiments of this application, an electronic device is provided, comprising:
[0252] Memory;
[0253] Processor; and
[0254] Computer programs;
[0255] The computer program is stored in the memory and is configured to be executed by the processor to implement the method for constructing an enhanced model of server voiceprint data as described above, or to be configured to be executed by the processor to implement the method for enhancing server voiceprint data as described above.
[0256] In a fifth aspect of the present application, a computer-readable storage medium is provided having a computer program stored thereon; the computer program is executed by a processor to implement the method for constructing an enhanced model of server voiceprint data as described above, or to implement the method for enhancing server voiceprint data as described above.
[0257] Example 2
[0258] In Embodiment 2 of this application, the fault types are expanded from 8 dimensions to 12 dimensions, covering not only the common fault types in Embodiment 1, but also more fault types that may occur in the server.
[0259] Table 4. Fault Type Description Table 2
[0260]
[0261] Example 3
[0262] The difference between Example 3 and Example 2 is that the fault type in Example 3 is a mixed fault, including "fan bearing wear, blade cracks, abnormal noise from the magnetic head, and power supply whistling".
[0263] Each fault type includes 30-50 fault voiceprint data points, for a total of 180 fault voiceprint data points.
[0264] The server models are 2288H / 1288H (two types).
[0265] The difference between Comparative Example 1 and Example 2 is that there are no physical constraints, and only the adversarial loss value is used as the total loss value of the generator.
[0266] Table 5 Comparison of Results
[0267]
[0268] As shown in Table 5, by calculating the above indicators, it can be found that the wave equation residual (which measures the degree of deviation between the generated sample and the real acoustic physical law) decreased by 78.6%, the proportion of high-frequency noise decreased by 88.5%, and the signal-to-noise ratio increased by 8.7%.
[0269] A comparison between Example 3 and Comparative Example 1 revealed the following: 1. In the blade crack fault type, the pseudo-time spectrum generated in Example 3 exhibits a clear sideband modulation characteristic in the 1500-2000Hz range (the physical fault characteristic is a ±50Hz sideband of the blade's passing frequency), while Comparative Example 1 loses this characteristic. 2. In the power supply whistling fault type, the pseudo-time spectrum generated in Example 3 shows a narrow-band spike at 5kHz (fault characteristic), and the fluctuation amplitude over time conforms to the actual power supply load variation pattern, while Comparative Example 1 generates a non-physical continuous wide spectrum. 3. Using the generated samples as training data to train the classifier, tests were conducted on real data without known operating parameters (fan speed 7200RPM, CPU load 60%). The accuracy of Comparative Example 1 was 67.3%, while the accuracy of Example 3 was 85.7%, representing an improvement of 18.4%.
[0270] Based on the above scheme, it is shown that physical constraints help the model learn the essential fault characteristics that remain unchanged under the operating conditions, rather than overfitting to the training conditions. The participation of the physical constraint loss calculation unit 50124 can effectively suppress non-physical noise, ensure that the generated samples conform to the laws of sound wave propagation, and enable the model to maintain high generalization ability under unknown operating conditions.
[0271] Comparative Example 2
[0272] The difference from Example 2 is that the deep learning model is used as the basis for the generative adversarial network.
[0273] Table 6 Improvement Effect Table
[0274]
[0275] As shown in Table 6, the technical solution provided in this application can improve the diversity of generated samples, the accuracy of fault classification, and the stability of training; the reduction in physical consistency indicates that the generated samples conform to the physical propagation law of sound waves, and the physical fidelity is significantly improved, solving the problem of lack of physical consistency in the prior art.
[0276] Based on the above solution, multiple fault types are taken into account, which solves the problem that the existing technology does not take into account the multi-source noise aliasing characteristics of the server, and the use of GAN to generate bearing fault signals is only suitable for ideal scenarios with a single fault, and cannot meet the needs of real-world scenarios.
[0277] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of this application can be implemented in various computer languages, such as C, VHDL, Verilog, the object-oriented programming language Java, and the interpreted scripting language JavaScript.
[0278] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0279] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0280] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0281] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0282] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.
[0283] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A method for constructing an enhancement model for server voiceprint data, characterized in that, include: Select multiple fault acoustic fingerprint data of the server and their corresponding operating parameters from the sample library; Operating parameters include fan speed, CPU load, ambient temperature, fault type, and server model; Perform a Fourier transform on each fault voiceprint data to obtain the initial time spectrum; normalize the initial time spectrum to obtain the voiceprint time spectrum. The classifier in the deep learning model is pre-trained by multiple time-spectral maps of voiceprints and corresponding operating parameters to obtain a pre-trained classifier and a pre-trained deep learning model. The pre-trained deep learning model is enhanced by using multiple audioprint time-spectrum maps and corresponding operating parameters to obtain an enhanced model.
2. The method for constructing an enhanced model of server voiceprint data according to claim 1, characterized in that, The deep learning model is a conditional generative adversarial network model, which includes a generator and a judge. The process of enhancing the pre-trained deep learning model by using multiple audioprint time-spectrum maps and corresponding operating parameters to obtain an enhanced model specifically includes: The discriminator and generator are alternately optimized by using multiple time-spectrum diagrams of voiceprints and corresponding operating parameters. After each k iterations of alternating optimization, multiple operating parameters are selected and input into the generator to generate multiple pseudo-time spectrum diagrams. The trained classifier is used to classify multiple pseudo-time spectrograms to obtain the classification results; Update the input weights of the operating condition parameters in the input generator based on the classification results; Determine whether the total number of optimization attempts of the generator has reached the preset training threshold or whether the total loss of the generator has converged. If so, then training ends and a trained augmented model is obtained; Otherwise, repeat the above steps.
3. The method for constructing an enhanced model of server voiceprint data according to claim 2, characterized in that, The process of alternately optimizing the discriminator and generator using multiple time-spectrum diagrams of voiceprints and corresponding operating parameters specifically includes: Optimize the discriminator, specifically including: Multiple operating condition parameters are selected and constructed into a corresponding first condition vector by combining them with the real-time acquired random noise vector; Multiple first condition vectors are input into the generator for processing to obtain multiple first pseudo-time spectrum maps; The acoustic signature time spectrum and the corresponding first pseudo-time spectrum of the selected multiple operating parameters are used as the original image, and the image obtained by downsampling the original image is used as the preprocessed image. Each original image and its corresponding preprocessed image are input into the discriminator to calculate the loss value for each original image, and the total loss value of the discriminator is calculated using the loss value of each original image. Based on the total loss value of the discriminator, the discriminator parameters are optimized by the optimizer to obtain the optimized discriminator; After optimizing the discriminator M times, the discriminator parameters are frozen; and the generator is optimized according to the following steps: Multiple operating condition parameters are selected and constructed into a corresponding second condition vector by combining them with the real-time acquired random noise vector; Multiple second conditional vectors are input into the generator for processing to obtain multiple second pseudo-time spectrum maps; The second pseudo-time spectrum corresponding to the selected multiple operating parameters is input into the optimized discriminator to calculate the adversarial loss value; Based on the acoustic waveform time spectrum corresponding to the selected multiple operating parameters, the physical constraint loss value is calculated according to the fluctuation loss calculation formula. The total loss value of the generator is calculated by combining the adversarial loss value and the physical constraint loss value. Based on the generator's total loss value, the generator parameters are optimized by the optimizer to obtain the optimized generator.
4. The method for constructing an enhanced model of server voiceprint data according to claim 3, characterized in that, The physical constraint loss value is calculated based on the acoustic waveform time-spectrum diagrams corresponding to the selected multiple operating parameters, according to the fluctuation loss calculation formula, specifically including: The acoustic signature time-frequency spectrum corresponding to the selected multiple operating parameters is inversely transformed into time-domain sound pressure. Where g represents the g-th operating condition parameter, and t represents the time coordinate. Represents spatial coordinates; Calculate the second-order partial derivatives of the time-domain sound pressure with respect to the time coordinate. and the second-order partial derivative of time-domain sound pressure with respect to spatial coordinates ; The physical constraint loss value is calculated using the volatility loss calculation formula, whereby the volatility loss calculation formula is: ; In the formula, This represents the physical constraint loss value. G represents the total number of selected operating parameters, and c represents the speed of sound, which is the speed at which sound travels through the air.
5. The method for constructing an enhanced model of server voiceprint data according to claim 3, characterized in that, The step of inputting multiple first condition vectors into a generator for processing to obtain multiple first pseudo-time spectrum diagrams specifically includes: The first conditional vector is reshaped into an initial feature map through a fully connected mapping layer; The initial feature map is processed sequentially by D downsampling blocks in the encoder module to obtain a high-dimensional feature map. The high-dimensional feature map is processed sequentially by D upsampling blocks in the decoder module to obtain the low-dimensional feature map. The low-dimensional feature map is processed by the upsampling module to obtain the first pseudo-time spectrum.
6. The method for constructing an enhanced model of server voiceprint data according to claim 1, characterized in that, The process of pre-training the classifier in the deep learning model using multiple time-spectrum audiograms and corresponding operating parameters to obtain a pre-trained classifier specifically includes: Multiple audioprint time-spectrum maps and corresponding operating parameters are divided into training set and validation set; Pre-train the classifier in the deep learning model using the training set; The classification results of the classifier are validated using a validation set; If the accuracy of the validation result is less than 90%, return to the training set to pre-train the classifier in the deep learning model for the next round. When the accuracy of the verification results is not less than 90%, the classifier parameters are frozen to obtain a pre-trained classifier.
7. A system for constructing an enhanced model of server voiceprint data, characterized in that, The module includes a method for constructing an enhanced model of server voiceprint data as described in any one of claims 1 to 6.
8. A method for enhancing server voiceprint data, characterized in that, include: Acquire the fault acoustic signature data to be enhanced, along with its corresponding operating parameters and the expected number of samples to be generated; Perform a Fourier transform on the enhanced fault acoustic data to obtain the initial fault spectrum. The spectrum diagram at the initial fault is normalized to obtain the spectrum diagram at the fault sound signature. The operating parameters and the expected number of fault voiceprint data to be enhanced are input into the enhancement model, wherein the enhancement model is the enhancement model constructed by the method of constructing the enhancement model of server voiceprint data as described in any one of claims 1 to 6; The model is enhanced to output the corresponding high-fidelity pseudo-time spectrum. The faulty voiceprint time spectrum map and the corresponding high-fidelity pseudo-time spectrum map are used to construct the enhanced voiceprint dataset.
9. An electronic device, characterized in that, include: Memory; processor; as well as Computer programs; The computer program is stored in the memory and configured to be executed by the processor to implement the method for constructing an enhanced model of server voiceprint data as described in any one of claims 1 to 6, or configured to be executed by the processor to implement the method for enhancing server voiceprint data as described in claim 8.
10. A computer-readable storage medium, characterized in that, It stores a computer program; the computer program is executed by a processor to implement the method for constructing an enhanced model of server voiceprint data as described in any one of claims 1 to 6, or to implement the method for enhancing server voiceprint data as described in claim 8.