Motor sound quality intelligent evaluation method based on semi-supervised learning and variance guided screening

CN122818073APending Publication Date: 2026-09-25SHANGHAI UNIV OF ENG SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611010758.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-08
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0003]现有特征提取方法多采用固定时间窗内的统计特征或静态频谱特征,虽然能够表征一定的频带能量分布,但对于阶次轨迹迁移、幅值起伏以及局部高频成分突变等时间演化信息的描述仍不充分

Benefits of technology

本发明面向声品质标签获取成本高且有标签样本规模受限的实际约束,构建半监督 VG-FPS-ViT-KAN 声品质评价框架。特征表征层面,在静态 GFCC 倒谱系数基础上引入一阶差分与二阶差分通道,形成三通道动态特征图,使输入同时包含听觉尺度谱包络信息及其随时间演化的变化率与变化趋势,从而增强对非稳态工况下阶次迁移、幅值起伏与局部高频纹理突变等时频细节的表征能力。输入建模层面,提出VG-FPS策略,以分块方差作为区域信息量与结构复杂度的统计度量,对高方差关键区域采用细粒度 token 划分并在块内执行局部自注意力聚合以强化局部相关性建模,同时对低方差区域进行池化压缩以降低冗余 token 与计算负担,从而在控制序列长度的前提下提高有效判别线索的利用效率。训练机制层面,引入 FreeMatch 半监督学习,将未标注工况样本纳入训练过程,在弱增强视图上生成伪标签并通过自适应置信度阈值进行筛选,以抑制低可信伪监督信号对优化方向的干扰,同时在强增强视图下施加一致性约束以稳定输出并持续校正决策边界,使无标签数据的分布信息能够对模型参数更新形成长期约束,进而缓解监督信息集中于单一工况导致的分布偏置,提升多工况条件下声品质等级判别的泛化性与稳定性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122818073A_ABST
    Figure CN122818073A_ABST
Patent Text Reader

Abstract

The application relates to a motor sound quality intelligent evaluation method based on semi-supervised learning and variance guided screening, which comprises the following steps: collecting noise signals of a permanent magnet synchronous motor under multiple operating conditions, and constructing a noise data set containing labeled samples and unlabeled samples; extracting static GFCC cepstrum coefficients for each noise sample in the noise data set, introducing a time derivative feature, and constructing a dynamic feature map; based on a variance guided fine-grained feature block screening VG-FPS strategy, different regions of the dynamic feature map are differentially processed to generate a structured Token sequence; a ViT-KAN sound quality evaluation model is constructed; a semi-supervised learning strategy based on FreeMatch is adopted, and the noise data set is used to train the ViT-KAN sound quality evaluation model; and a motor noise signal to be evaluated is input into the trained ViT-KAN sound quality evaluation model to obtain a corresponding sound quality grade.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of drive motor noise evaluation technology, and in particular to an intelligent evaluation method for motor sound quality based on semi-supervised learning and variance-guided screening. Background Technology

[0002] Noise quality assessment of drive motors is a crucial aspect of noise, vibration, and acoustic roughness control in new energy vehicles. Under constant speed, acceleration, deceleration, and varying load conditions, the order components, amplitude distribution, and high-frequency texture of the noise in permanent magnet synchronous motors continuously change with rotational speed and load, resulting in significant non-steady-state characteristics in the noise signal. Current sound quality assessment methods typically obtain sound quality levels through subjective listening evaluation and establish a mapping relationship between noise characteristics and subjective evaluation results using psychoacoustic parameters, time-frequency characteristics, or cepstral features.

[0003] Existing feature extraction methods mostly employ statistical features or static spectral features within a fixed time window. While these can characterize a certain frequency band energy distribution, they are insufficient in describing temporal evolution information such as order trajectory shifts, amplitude fluctuations, and abrupt changes in local high-frequency components. In deep learning methods based on time-frequency feature maps, the same size and granularity of block processing or convolution is typically used for each region of the feature map. Different representation granularities are not allocated according to the information density and structural complexity of different regions, which easily leads to insufficient representation of high-information regions and redundant computation of low-information regions.

[0004] Furthermore, motor noise quality labels typically require manual listening tests, which are costly to organize and result in a limited number of labeled samples. In contrast, unlabeled noise data under different speeds and load conditions is relatively easy to collect. Supervised training using only a small number of labeled samples can easily lead to a bias towards the operating conditions covered by the labeled samples, leaving room for improvement in discriminative stability across speed and load conditions. Therefore, a motor noise quality evaluation method is needed that can simultaneously describe the dynamic evolution characteristics of non-steady-state noise, implement differentiated modeling based on the information content of feature map regions, and utilize unlabeled operating condition data to improve the model's generalization ability. Summary of the Invention

[0005] The purpose of this invention is to propose an intelligent evaluation method for motor sound quality based on semi-supervised learning and variance-guided screening, in order to solve the problems existing in the prior art. To achieve the above objectives, the present invention provides the following solution: A method for intelligent evaluation of motor sound quality based on semi-supervised learning and variance-guided screening includes: Noise signals of permanent magnet synchronous motors under various operating conditions were collected, and a noise dataset containing labeled and unlabeled samples was constructed. For each noise sample in the noise dataset, static GFCC cepstral coefficients are extracted, and time derivative features are introduced to construct a dynamic feature map; Based on the variance-guided fine-grained feature block screening VG-FPS strategy, different regions of the dynamic feature map are differentiated to generate a structured token sequence. A visual transformer-Kolmogorov-Arnold network sound quality evaluation model is constructed, denoted as the ViT-KAN sound quality evaluation model. The ViT-KAN sound quality evaluation model takes the structured token sequence as input, uses the ViT network to extract global time-frequency features, and uses KAN to replace the traditional multilayer perceptron (MLP) classification head to map the global time-frequency features into sound quality level prediction results. A semi-supervised learning strategy based on FreeMatch was adopted to train the ViT-KAN sound quality evaluation model using the noise dataset. The motor noise signal to be evaluated is input into the trained ViT-KAN sound quality evaluation model to obtain the corresponding sound quality level.

[0006] Optionally, constructing the dynamic feature map includes: The noise sample signal is sequentially processed by auditory filtering, subband energy calculation, logarithmic compression and cepstral transform to obtain the GFCC cepstral coefficients corresponding to each frame of signal. The first few GFCC cepstral coefficients are selected as static GFCC cepstral coefficients. The time derivative feature is introduced, including first-order difference and second-order difference; wherein, the first-order difference: by approximating the derivative of the change in the cepstral sequence of adjacent frames, it characterizes the rate of change of the spectral envelope with time; the second-order difference: describes the trend of the rate of change. The calculation formula is as follows: Where, Δ ci It is a first-order difference, Δ2 ci It is a second-order difference. ci Indicates the first i The static GFCC feature vector corresponding to the frame noise signal ci-1 and ci+ 1 represents the first i The static GFCC feature vectors corresponding to the previous and next frames, Δ ci Indicates the first i First-order temporal difference of frame static GFCC eigenvectors, Δ² ci This represents the second-order temporal difference of the static GFCC feature vector of the i-th frame; The first-order difference, second-order difference, and static GFCC feature vectors are concatenated to form a three-channel image feature representation.

[0007] Optionally, a variance-guided fine-grained feature block screening VG-FPS strategy is used to differentiate different regions of the dynamic feature map, generating a structured token sequence including: The dynamic feature map is spatially divided into multiple non-overlapping coarse-grained image blocks, the variance of the coarse block pixels is calculated, and the variance contribution rate is calculated. The dynamic feature map is divided into high-variance regions and low-variance regions based on the variance contribution rate. A dual-path token generation mechanism, based on fine-grained self-attention aggregation in high-variance regions and pooling compression in low-variance regions, generates a structured token sequence.

[0008] Optionally, calculating the variance contribution rate includes: The total variance of the entire image is obtained by summing the variances of all pixel blocks. Sort the variance contributions in descending order to obtain the sorted sequence; Based on the sorted sequence and the total variance of the entire graph, the variance contribution rate of coarse blocks is obtained.

[0009] Optionally, the dual-path token generation mechanism includes: Each high-variance coarse-grained image block is divided into multiple fine-grained sub-blocks, and each fine-grained sub-block is mapped to a corresponding sub-block Token; Layer normalization and linear mapping are performed on sub-block tokens belonging to the same high-variance coarse-grained image block to obtain query vector, key vector and value vector, and scaling dot product attention output is calculated based on the query vector, key vector and value vector; The scaled dot product attention output is added to the corresponding original sub-block token by residual addition to obtain multiple updated fine-grained sub-block tokens, and the multiple fine-grained sub-block tokens are used as fine-grained representations of the high-variance coarse-grained image block. Each low-variance coarse-grained image block is divided into multiple fine-grained sub-blocks, and the sub-block tokens corresponding to the multiple fine-grained sub-blocks are pooled to compress each low-variance coarse-grained image block into a coarse-grained token. The fine-grained sub-block tokens corresponding to all high-variance coarse-grained image blocks and the coarse-grained tokens corresponding to all low-variance coarse-grained image blocks are combined, and the two-dimensional position information of each token in the original dynamic feature map is retained to generate the structured token sequence.

[0010] Optionally, the ViT-KAN sound quality evaluation model includes: The Token embedding module is used to map the structured Token sequence to an embedding space of a preset dimension, and to add a classification marker CLS Token to the structured Token sequence. The two-dimensional rotation position encoding module is used to generate a two-dimensional rotation position encoding based on the horizontal and vertical coordinates of each token in the dynamic feature map, and inject the two-dimensional rotation position encoding into the query vector and key vector in the Transformer self-attention calculation respectively; The ViT feature extraction module includes a multi-layer Transformer encoder. Each Transformer encoder includes a multi-head self-attention sub-layer, a feedforward neural network sub-layer, residual connections, and layer normalization, which are used to model the global correlation between different tokens and obtain global time-frequency features through the CLS token of the final layer. The KAN classification head is used to convert the global time-frequency features into prediction results corresponding to different sound quality levels through nonlinear mapping based on spline basis functions.

[0011] Optionally, extracting global time-frequency features using a ViT network includes: The ViT network divides the input features into several sub-blocks and maps them into sequence form. It then uses a multi-head self-attention mechanism to model the correlation between the sub-blocks, thereby representing the overall structural information of the input and obtaining global time-frequency features.

[0012] Optionally, the KAN maps global time-frequency features to sound quality level prediction results, including: The global time-frequency features are normalized. A cubic B-spline basis function system is constructed over a preset interval to achieve localized approximation of the nonlinear mapping; For each input scalar dimension, its response on multiple cubic spline bases is calculated using the constructed basis functions, thereby achieving nonlinear expansion of the single-dimensional feature. Stack the expanded results of all dimensions to form a basis matrix, then vectorize the basis matrix in a fixed order and flatten it into column vectors; On a fixed spline basis, only linear weights are learned, and the class logits are obtained through linear combination; The class logits are converted into a probability distribution using Softmax to obtain the model's predicted output; For labeled samples, the cross-entropy loss function is used as the training objective.

[0013] Optionally, FreeMatch-based semi-supervised learning strategies include: Forward inference is performed on the original image of the unlabeled sample to obtain the class probability distribution of the network output. This distribution is used as a soft pseudo-label to characterize the model's prediction tendency for each class. Furthermore, the maximum value in the probability distribution and its corresponding class index are extracted and used as the confidence score and hard pseudo-label of the sample, respectively. Based on the aforementioned confidence level and hard pseudo-labels, an adaptive confidence threshold strategy is used to screen unlabeled samples. Perform multi-view analysis on the selected unlabeled samples. Figure 1 Dedication training; Calculate the total loss and update the model parameters.

[0014] Optionally, the adaptive confidence threshold strategy includes: Calculate the average of the maximum confidence scores for all current unlabeled samples; The global threshold is updated based on the average value using the EMA method. A class-based thresholding strategy is introduced to calculate the average confidence score for each class and scale the global threshold accordingly. The global threshold is scaled according to the relative magnitude of the average confidence of each category to obtain the threshold for each category.

[0015] The beneficial effects of this invention are as follows: This invention addresses the practical constraints of high cost and limited sample size in acquiring sound quality labels by constructing a semi-supervised VG-FPS-ViT-KAN sound quality evaluation framework. At the feature representation level, first-order and second-order difference channels are introduced on top of static GFCC cepstral coefficients to form a three-channel dynamic feature map. This allows the input to simultaneously include auditory scale spectral envelope information and its rate and trend of change over time, thereby enhancing the representation of time-frequency details such as order shifts, amplitude fluctuations, and local high-frequency texture abrupt changes under non-steady-state conditions. At the input modeling level, a VG-FPS strategy is proposed, using block variance as a statistical measure of regional information content and structural complexity. Fine-grained token partitioning is applied to high-variance key regions, and local self-attention aggregation is performed within blocks to strengthen local correlation modeling. Simultaneously, pooling compression is used for low-variance regions to reduce redundant tokens and computational burden, thereby improving the utilization efficiency of effective discrimination cues while controlling sequence length. At the training mechanism level, FreeMatch semi-supervised learning is introduced, which incorporates unlabeled working condition samples into the training process. Pseudo-labels are generated on the weakly enhanced view and filtered by an adaptive confidence threshold to suppress the interference of low-confidence pseudo-supervisory signals on the optimization direction. At the same time, consistency constraints are applied under the strongly enhanced view to stabilize the output and continuously correct the decision boundary. This allows the distribution information of unlabeled data to form a long-term constraint on the update of model parameters, thereby alleviating the distribution bias caused by the concentration of supervision information in a single working condition and improving the generalization and stability of sound quality level discrimination under multiple working conditions. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a schematic diagram illustrating the GFCC acoustic feature construction process with the introduction of dynamic features in an embodiment of the present invention. Figure 2 This is a schematic diagram of the feature block selection and token construction process based on variance-guided fine-grained screening (VG-FPS) according to an embodiment of the present invention; Figure 3 This is a schematic diagram of a ViT network according to an embodiment of the present invention; Figure 4 This is a schematic diagram of the multi-head self-attention calculation process for introducing RoPE position-aware rotational coding in an embodiment of the present invention; Figure 5 This is a schematic diagram of the sound quality evaluation model structure based on the ViT feature extraction layer and the KAN classification layer in an embodiment of the present invention; Figure 6 This is a schematic diagram of the semi-supervised learning training process based on the VG-FPS-ViT-KAN model in an embodiment of the present invention. Figure 7 This is a schematic diagram illustrating the impact of different numbers of attention heads on model training loss in an embodiment of the present invention; Figure 8 This is a schematic diagram showing the change in accuracy of the GFCC-ViT model with training rounds in an embodiment of the present invention. Detailed Implementation

[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0020] This embodiment proposes an intelligent evaluation method for motor sound quality based on semi-supervised learning and variance-guided screening, including: Noise signals of permanent magnet synchronous motors under various operating conditions were collected, and a noise dataset containing labeled and unlabeled samples was constructed. For each noise sample in the noise dataset, static GFCC cepstral coefficients are extracted, and time derivative features are introduced to construct a dynamic feature map; Based on the variance-guided fine-grained feature block screening VG-FPS strategy, different regions of the dynamic feature map are differentiated to generate a structured token sequence. A visual transformer-Kolmogorov-Arnold network sound quality evaluation model is constructed, denoted as the ViT-KAN sound quality evaluation model. The ViT-KAN sound quality evaluation model takes the structured token sequence as input, uses the ViT network to extract global time-frequency features, and uses KAN to replace the traditional multilayer perceptron (MLP) classification head to map the global time-frequency features into sound quality level prediction results. A semi-supervised learning strategy based on FreeMatch was adopted to train the ViT-KAN sound quality evaluation model using the noise dataset. The motor noise signal to be evaluated is input into the trained ViT-KAN sound quality evaluation model to obtain the corresponding sound quality level.

[0021] Specifically, in this embodiment, noise from a permanent magnet synchronous motor under various operating conditions, including constant speed, acceleration, and deceleration, is first collected to construct a noise dataset for sound quality evaluation and modeling analysis.

[0022] This embodiment conducts an acoustic acquisition experiment on a drive motor, and the basic parameters of the motor are shown in Table 1. The measurement methods used in the experiment strictly follow the relevant requirements of the national standard GB / T 10069.1-2006 "Methods for Determination and Limits of Noise in Rotating Electrical Machines - Part 1", and the number of test microphones is reasonably reduced according to the experimental conditions and experimental objectives.

[0023] Table 1 Basic Electrode Parameters In this embodiment, noise was collected on a dynamometer stand in a semi-anechoic chamber. The cutoff frequency of the semi-anechoic chamber was 50Hz, and the background noise was less than 20dB(A). The LMS Testlab data acquisition platform was used, with a PCB 378B02 microphone, a frequency response range of 3.15Hz-20kHz, a sensitivity of 50mV / Pa, and a sampling frequency of 51200Hz. The microphones were placed 1m in front of, to the left of, to the right of, and above the motor housing. All microphones were calibrated before the test.

[0024] The motor speed range was set to 500 RPM to 5000 RPM, and the torque load was set to 0 N·m, 50 N·m, 100 N·m, and 150 N·m, respectively. Under each torque load, noise signals were collected for constant speed, acceleration, and deceleration conditions. For the constant speed condition, 10 speed points were collected in 500 RPM increments: 500 RPM, 1000 RPM, ..., 5000 RPM, with each speed point collected continuously for 12 seconds. For the acceleration condition, the speed increased from 500 RPM to 5000 RPM, and for the deceleration condition, the speed decreased from 5000 RPM to 500 RPM, with an acceleration / deceleration rate of 100 RPM / s and a single-segment acquisition time of 45 seconds. Each condition was collected twice, and the one with the smaller fluctuation was selected as the analysis data.

[0025] All collected noise signals were replayed and verified, and the data collected by the front microphone was selected uniformly. In order to match the judgment time of subjective evaluation, the 12s noise signal of the constant speed condition was truncated into multiple 5s segments, and the 45s noise signal of the acceleration and deceleration conditions was divided into multiple 5s segments at 5s intervals. Thus, a total of 112 noise samples were obtained under 4 torque loads, each load containing 10 samples of the constant speed condition, 9 samples of the acceleration condition, and 9 samples of the deceleration condition.

[0026] All noise samples were selected as the subjects of subjective evaluation. However, due to limitations in the organization cost and time of the subjective listening evaluation, subjective sound quality scoring tests were only conducted on 28 noise samples under a 150 N·m torque load. Details are as follows: Subjective evaluations were used to obtain labeled samples: Twenty-eight noise samples under a 150 N·m torque load (including 10 constant speed conditions, 9 acceleration conditions, and 9 deceleration conditions) were selected, and subjective listening tests were conducted using a paired comparison method. In a semi-anechoic chamber, 20 evaluators used Sennheiser HD600 high-fidelity headphones and a Behringer UMC202HD external sound card to reproduce the signals and rate the annoyance level of the samples. After data validation (triangular cyclic misjudgment analysis and consistency weighting factor screening) and cross-group scale inversion unification, a subjective annoyance score was obtained for each sample, and this score was used as the sound quality level label for that sample. Thus, the 28 noise samples under a 150 N·m load were labeled as tagged samples.

[0027] Construction of unlabeled sample set: All noise samples under the remaining three torque loads (0 N·m, 50 N·m, and 100 N·m) were used as unlabeled samples, totaling 84 samples (28 for each load). These samples were not subjected to subjective listening tests and did not have corresponding sound quality level labels, serving as unlabeled data in semi-supervised training.

[0028] Isolation of unlabeled training pool: To ensure that the validation set is not leaked by unlabeled data during semi-supervised training, a rotational speed isolation strategy is adopted: the rotational speed values ​​contained in the validation set (a portion of the validation set divided from 28 labeled samples) are recorded; any unlabeled sample with the same rotational speed as any sample in the validation set under loads of 0 N·m, 50 N·m, or 100 N·m are removed from the unlabeled training pool. After this isolation process, the remaining unlabeled samples constitute the final unlabeled training set, which is used for consistency constraint training in semi-supervised learning.

[0029] At this point, a noisy dataset containing 28 labeled samples (150 N·m) and unlabeled samples (0, 50, 100 N·m) after isolation and filtering has been completed.

[0030] Furthermore, constructing the dynamic feature map includes: The noise sample signal is sequentially processed by auditory filtering, subband energy calculation, logarithmic compression and cepstral transform to obtain the GFCC cepstral coefficients corresponding to each frame of signal. The first few GFCC cepstral coefficients are selected as static GFCC cepstral coefficients. The time derivative feature is introduced, including first-order difference and second-order difference; wherein, the first-order difference: by approximating the derivative of the change in the cepstral sequence of adjacent frames, it characterizes the rate of change of the spectral envelope with time; the second-order difference: describes the trend of the rate of change. The first-order difference, the second-order difference, and the static GFCC cepstral coefficients are concatenated to form a three-channel image feature representation.

[0031] Furthermore, the extraction of global time-frequency features using ViT networks includes: The ViT network divides the input features into several sub-blocks and maps them into sequence form. It then uses a multi-head self-attention mechanism to model the correlation between the sub-blocks, thereby representing the overall structural information of the input and obtaining global time-frequency features.

[0032] Furthermore, the KAN maps global time-frequency features to sound quality level prediction results, including: The global time-frequency features are normalized. A cubic B-spline basis function system is constructed over a preset interval to achieve localized approximation of the nonlinear mapping; For each input scalar dimension, its response on multiple cubic spline bases is calculated using the constructed basis functions, thereby achieving nonlinear expansion of the single-dimensional feature. Stack the expanded results of all dimensions to form a basis matrix, then vectorize the basis matrix in a fixed order and flatten it into column vectors; On a fixed spline basis, only linear weights are learned, and the class logits are obtained through linear combination; The class logits are converted into a probability distribution using Softmax to obtain the model's predicted output; For labeled samples, the cross-entropy loss function is used as the training objective.

[0033] Specifically, in this embodiment, to compensate for the inadequacy of static spectra in fully reflecting the continuous temporal evolution of non-steady-state noise and to improve the model's ability to represent the time-varying laws of non-steady-state noise, this embodiment further introduces time derivative features, including first-order difference (Δ) and second-order difference (ΔΔ), based on the static GFCC cepstral coefficients. Static cepstral coefficients mainly reflect the overall shape of the short-time spectral envelope and can effectively describe the instantaneous characteristics of energy distribution in different frequency bands, but they are insufficient in characterizing phenomena such as transient amplitude fluctuations, order shifts, and local howling that evolve continuously over time. The Δ feature, by approximating the derivative of the changes in the cepstral sequence of adjacent frames, characterizes the rate of change of the spectral envelope over time and can be used to capture the dynamic fluctuations of noise signals under acceleration, deceleration, and load disturbance conditions. The ΔΔ feature further describes the trend of the rate of change, reflecting the acceleration information of spectral evolution, and is more sensitive to fine-grained temporal structures such as abrupt changes, inflection points, and dynamic modulation in non-steady-state processes. By introducing Δ and ΔΔ, the input features supplement multi-scale temporal dynamic cues while maintaining the auditory scale spectrum envelope information. This helps the model learn more discriminative evolution patterns in the joint time-frequency space, thereby improving the ability and robustness to identify the sound quality levels of motor noise under multiple operating conditions.

[0034] The calculation formula is as follows: Wherein, Δci is the first-order difference, representing the time rate of change of the cepstral coefficients, reflecting the trend of energy change. Δ2ci is the second-order difference, representing sudden shocks or discontinuous modes.

[0035] The two dynamic features are concatenated with the original cepstral coefficients to form a 105-dimensional three-channel image feature representation for each frame: The pseudo-image input consists of three channels: static GFCC, Δ, and ΔΔ, and its structure is shown in the diagram below. Figure 1 As shown.

[0036] Furthermore, based on the variance-guided fine-grained feature block screening VG-FPS strategy, different regions of the dynamic feature map are differentiated to generate a structured token sequence, including: The dynamic feature map is spatially divided into multiple non-overlapping coarse-grained image blocks, the variance of the coarse block pixels is calculated, and the variance contribution rate is calculated. The dynamic feature map is divided into high-variance regions and low-variance regions based on the variance contribution rate. A dual-path token generation mechanism, based on fine-grained self-attention aggregation in high-variance regions and pooling compression in low-variance regions, generates a structured token sequence.

[0037] Specifically, in this embodiment, to adapt to the input method of ViT based on block modeling and to solve the representation redundancy problem caused by the uneven information distribution of GFCC acoustic feature maps in the time-frequency domain, this embodiment introduces the VG-FPS strategy to implement differentiated processing for different regions. The rationale is that traditional uniform block modeling assumes that each region has similar information value, which easily leads to insufficient representation of high-information regions and repeated modeling of low-information regions, hindering the effective use of limited computing resources. Based on this, VG-FPS uses local statistics to evaluate the information content of regions, allocating more representational power to regions with more complex structural features and greater sensitivity to category discrimination, while using simplified representations for relatively sparse regions. This improves feature extraction efficiency and input representation quality without excessively increasing computational overhead.

[0038] Specifically, the input GFCC acoustic feature map is first spatially divided into 100 non-overlapping coarse-grained image blocks, each with a size of 64×64 pixels. For the pixel set Ωi corresponding to the i-th coarse block, the mean and variance within the block are calculated to obtain the information content index vi for that region. The pixel mean within the block is defined as: Based on this, the variance of coarse block pixels is defined as: in, I ( u,v ) indicates that the GFCC acoustic feature map is located at pixel coordinates ( u,v The value at position ) , |Ω i | for the first i The number of pixels contained in a coarse block.

[0039] Furthermore, summing the variances of all coarse blocks yields the total variance of the entire map: Then contribute the variance {vi} Perform a descending sort, and denote the sorted sequence as: To characterize the relative contribution of different coarse blocks to the overall information, the first... k The variance contribution rate of each coarse block is: Accordingly, ranked top k The cumulative variance contribution rate of each high-variance coarse block is defined as: The above indicators can be used to measure the spatial concentration of high information density regions. This embodiment statistically analyzes all GFCC acoustic feature maps and calculates different... k Below η ( k And a cumulative contribution curve was plotted. The results show that the information content of the GFCC acoustic feature map has obvious concentration characteristics in spatial distribution: when k When =37, the cumulative variance contribution rate η ( k The top 37 coarse-grained image patches, exceeding 90%, cover the main information components of the entire feature map. This result indicates that the information distribution of the GFCC dynamic feature map is not uniform but rather concentrated in a few key time-frequency regions. This means that assigning the same modeling granularity to all regions is not optimal under limited computational resources. Priority should be given to improving the representation accuracy of high-information-density regions, while appropriately compressing low-information-density regions. Therefore, the subsequent adoption of a variance-guided differential tokenization strategy has clear data support, rather than being based on empirical settings.

[0040] Based on the aforementioned variance contribution rate analysis results, this embodiment employs a differentiated tokenization strategy for different regions of the GFCC acoustic feature map within the VG-FPS framework. High-variance regions are considered key time-frequency subdomains with high information density and stronger discriminative cues, and are thus refined in characterization. Low-variance regions are considered background subdomains with relatively stable information and high redundancy, and are represented in a compressed manner. This design aims to prioritize the preservation of local time-frequency information that is more sensitive to differences in the structure of non-steady-state noise without significantly increasing sequence length and computational overhead, thereby improving the effectiveness of feature modeling and the efficiency of computational resource utilization. The overall process is as follows: Figure 2 As shown.

[0041] Specifically, the input feature map is first divided into 100 coarse-grained image blocks of 64×64 pixels, and these blocks are sorted in descending order based on their intra-block pixel variance. For the top 37 high-variance coarse blocks, a fine-grained path is adopted: each coarse block is further evenly divided into 16×16 fine-grained sub-blocks, resulting in 16 fine-grained tokens within that block. To explicitly model the correlation between fine-grained tokens within the local area of ​​the coarse block, this embodiment performs a round of self-attention aggregation within that block. Specifically, the sequence of 16 tokens is first input into the attention module, normalized, and linearly mapped to obtain query, key, and value vectors. The scaled dot product attention weights are calculated, and the value vectors are weighted and summed to achieve intra-block information interaction and feature integration. Subsequently, the attention output is added to the original tokens via a residual connection to obtain the fine-grained representation of the coarse block after a local self-attention update. This process enhances the aggregation expression of structural information within the coarse block without introducing cross-block computation, allowing for more comprehensive modeling of local patterns in high-information regions.

[0042] For the remaining 63 low-variance coarse blocks, a compressed path is used to reduce redundant tokens: each coarse block is still divided into 16 16×16 sub-blocks, but instead of performing intra-block self-attention calculations, mean pooling is directly performed on the 16 sub-block tokens, aggregating them into a single token representing the coarse block. In this way, low-information regions are compressed from 16 tokens to 1 token, significantly reducing the sequence length and computational burden on the subsequent Transformer backbone while maintaining coarse-grained background contour information.

[0043] The dual-path token generation mechanism includes: Each high-variance coarse-grained image block is divided into multiple fine-grained sub-blocks, and each fine-grained sub-block is mapped to a corresponding sub-block Token; Layer normalization and linear mapping are performed on sub-block tokens belonging to the same high-variance coarse-grained image block to obtain query vector, key vector and value vector, and scaling dot product attention output is calculated based on the query vector, key vector and value vector; The scaled dot product attention output is added to the corresponding original sub-block token by residual addition to obtain multiple updated fine-grained sub-block tokens, and the multiple fine-grained sub-block tokens are used as fine-grained representations of the high-variance coarse-grained image block. Each low-variance coarse-grained image block is divided into multiple fine-grained sub-blocks, and the sub-block tokens corresponding to the multiple fine-grained sub-blocks are pooled to compress each low-variance coarse-grained image block into a coarse-grained token. The fine-grained sub-block tokens corresponding to all high-variance coarse-grained image blocks and the coarse-grained tokens corresponding to all low-variance coarse-grained image blocks are combined, and the two-dimensional position information of each token in the original dynamic feature map is retained to generate the structured token sequence.

[0044] In summary, VG-FPS achieves structured utilization of the information density differences in GFCC acoustic feature maps through a dual-path token generation mechanism that combines fine-grained self-attention aggregation in high-variance regions and pooling compression in low-variance regions. The specific steps are as follows: The first step is to divide the high-variance coarse block into 16 sub-blocks (Tokens). The input sequence is first subjected to LayerNormalization, and then a query is generated through linear mapping. Q ),key( K ),value( V )vector: The second step is to calculate the weights of the Scaled Dot-Product Attention: The third step is to calculate the weighted summation vector: The fourth step is to add the attention output to the original token using the residual, which serves as the fine-grained representation of the coarse block after one self-attention aggregation and residual update. .

[0045] Furthermore, the ViT-KAN sound quality evaluation model includes: The Token embedding module is used to map the structured Token sequence to an embedding space of a preset dimension, and to add a classification marker CLS Token to the structured Token sequence. The two-dimensional rotation position encoding module is used to generate a two-dimensional rotation position encoding based on the horizontal and vertical coordinates of each token in the dynamic feature map, and inject the two-dimensional rotation position encoding into the query vector and key vector in the Transformer self-attention calculation respectively; The ViT feature extraction module includes a multi-layer Transformer encoder. Each Transformer encoder includes a multi-head self-attention sub-layer, a feedforward neural network sub-layer, residual connections, and layer normalization, which are used to model the global correlation between different tokens and obtain global time-frequency features through the CLS token of the final layer. The KAN classification head is used to convert the global time-frequency features into prediction results corresponding to different sound quality levels through nonlinear mapping based on spline basis functions.

[0046] Specifically, in this embodiment, the sound quality evaluation network structure based on ViT-KAN is designed as follows: ViT backbone network configuration: Vision Transformer (ViT) is a feedforward deep learning model based on a self-attention mechanism. Its core structure consists of a feature embedding module and a multi-layer Transformer encoder. This embodiment uses a Transformer network as the backbone for visual feature extraction. Each encoding block consists of modules such as Multi-Head Scaled Dot-Product Attention (MHA), Feed-Forward Network (FFN), Residual Connection, and Layer Normalization (LN). The specific steps are as follows: Figure 3 As shown, this model divides the input features into several sub-blocks and maps them into a sequence form. It then utilizes a multi-head self-attention mechanism to model the correlations between these sub-blocks, thereby representing the overall structural information of the input. ViT was chosen primarily because the time-frequency feature map of motor noise not only contains local texture differences but also exhibits cross-regional overall correlations. Relying solely on local convolutional receptive fields makes it difficult to fully characterize these global structural relationships. ViT, on the other hand, can explicitly model long-range dependencies between different regions during feature extraction, thus possessing stronger global feature modeling capabilities in complex time-frequency pattern recognition.

[0047] To enhance the stability and generalization performance of the model during training, ViT extensively employs layer normalization and residual connections in its network structure to mitigate gradient instability during deep network training. It also utilizes Dropout to regularize attention weights and feature maps, thereby reducing the risk of overfitting the model to training samples. The synergistic effect of these structural designs and optimization strategies enables ViT to maintain modeling flexibility while possessing good training stability and feature representation capabilities.

[0048] Ablation analysis was conducted on the depth of the Transformer encoder and the number of self-attention heads. The results show that, while maintaining consistent input construction and training settings, the 12-layer encoder exhibits faster loss descent and a more stable convergence process. While the 24-layer encoder still has some optimization potential in the mid-to-late stages, its overall convergence is slower. The 32-layer encoder is more prone to optimization difficulties and loss fluctuations under small sample conditions. Based on these comparisons, this embodiment sets the encoder depth to 12 layers by default and further examines the impact of the number of attention heads under a fixed depth. With the embedding dimension unchanged, the training curve of the 8-head configuration is smoother and the convergence efficiency is higher. The 12-head configuration often requires longer iterations to reach a stable region, while the 16-head configuration exhibits more pronounced non-stationarity and a higher loss plateau. Therefore, this embodiment uses 12 layers and 8 attention heads as the default configuration without changing the backbone structure.

[0049] 2D-RoPE location information embedding: After introducing VG-FPS, different regions, even after differentiated tokenization, still need to participate in global attention modeling within a unified sequence space. Without explicit positional constraints, while the model can identify token content similarity, it struggles to accurately distinguish the geometric differences when the same local pattern appears at different temporal frequencies. Therefore, this embodiment employs 2D Rotational Relative Position Encoding (2D-RoPE), extending RoFormer's rotational position encoding from one-dimensional sequences to a two-dimensional grid. By applying a rotational transformation determined by spatial coordinates to the query and key vectors, positional information is written into the dot product attention calculation in phase form, enabling the attention weights to respond to the relative displacement between tokens. Compared to directly superimposing absolute position embeddings onto content embeddings, this approach does not rely on additional learnable position vectors or require interpolation of the position table when the input scale changes, making positional modeling closer to the computational path of the attention operator. Furthermore, by adopting a two-dimensional form, horizontal and vertical displacement information can enter different subspaces, helping to maintain a consistent expression of two-dimensional geometric relationships. This approach has been used in research on visual Transformers to improve positional awareness and cross-scale adaptation performance. The process for constructing and injecting positional information is described in [link to relevant documentation]. Figure 4 As shown, the specific implementation is given below. First, the center point coordinates of each patch in the image grid are represented as (x, y), and normalized to the interval [0,1]: in, W Indicates the image width. H Indicates the image height.

[0050] In Transformer, for query vectors Q and key vector VFor each two-dimensional sub-vector, the corresponding rotation angle needs to be calculated based on the patch coordinates. θx and θy Let the total rotational encoding dimension be... d The frequency reference is defined as follows: Then the first m The horizontal and vertical rotation angles corresponding to each frequency are as follows: Then, the Q / K vector is divided into several consecutive 2D sub-vectors, and a 2×2 rotation transformation matrix is ​​constructed. R ( θ Inject the position-dependent phase into the subvector; in, θ Determined by the patch coordinates. In practice, the embedding dimension is divided into two: the first half of the sub-vector is injected with the horizontal angle. θ x The second half of the sub-vector is injected with the vertical angle. θ y .

[0051] For any frequency index m , for the 2nd m Rotate the sub-vectors according to the horizontal coordinate θx , 2nd m +1-dimensional subvector rotated according to vertical coordinates θy The resulting vector after positional encoding injection is obtained: in, θmaxis It depends on the axis to which the subvector belongs. (Key vector) K The processing method is the same. In this way, the increment of the two-dimensional position is injected in the form of a rotational phase. Q / K Within the sub-vectors.

[0052] KAN classification header design: Specifically, after the input token sequence is updated layer by layer through 12 Transformer encoding blocks, the CLS token continuously fuses the contextual information of each patch token between layers. Its final output can be regarded as the global representation vector of the sample. In this embodiment, the CLS token is used as the input feature for subsequent classification, with an embedding dimension of 512. In terms of classification head design, conventional MLPs usually rely on linear transformations and fixed-form nonlinear activations to achieve the mapping from the representation space to the category space. KAN, on the other hand, introduces learnable spline basis functions to approximate the mapping relationship, which has certain advantages in terms of parameter utilization efficiency and fitting flexibility. At the same time, related studies have shown that using the KAN structure to replace the MLP sublayer in the Transformer helps to enhance the expressive power of the channel mixing stage. The reason for introducing KAN into the classification head is that sound quality level discrimination is essentially a complex nonlinear mapping problem. The global representation contained in the CLS token and the final category are not a simple linear correspondence. Therefore, the classification end needs to have a stronger function approximation capability to improve the ability to characterize subtle category differences. Based on the above, this embodiment replaces the original MLP classification head with KAN, and inputs the normalized CLS token into KAN to complete the sound quality level discrimination. The corresponding calculation process and module connection relationship are as follows: Figure 5 As shown. Specific implementation details are as follows: The first step is to normalize the CLS Token to make the distribution of each dimension more numerically stable.

[0053] The second step is to construct a cubic [-3,3] interval. B A spline basis function system is used to achieve localized approximation of nonlinear mappings. This embodiment uses a cubic B-spline of order (…). p =3), and set on each dimension K =8 basis functions are used to strike a trade-off between representational power and parameter size. The node vectors are uniformly arranged to ensure that the basis functions have approximately uniform resolution within the domain. Additionally, the basis functions are repeated at both ends of the interval. p+1 order nodes are added to form clamping boundary conditions, ensuring the spline maintains complete basis function support and stable numerical behavior at the boundaries, avoiding fitting degradation caused by insufficient support in the endpoint regions. The basis functions are constructed according to the Cox-de Boor recursion. The zeroth-order basis function can be considered as an indicator function for the node interval, used to characterize the activation range of the local interval. Higher-order basis functions are obtained by weighted recursion of adjacent lower-order basis functions according to the node spacing, naturally forming a piecewise polynomial expression and ensuring the smooth continuity of the function at the nodes. Due to the tight support characteristic of B-splines, each basis function is non-zero only within a finite node interval, making the model's response to input perturbations local. This helps suppress global overfitting and improves the stability and interpretability of parameter learning. Simultaneously, its piecewise polynomial structure maintains fitting flexibility while possessing good numerical stability, making it suitable for nonlinear discriminant mapping modeling under small sample conditions.

[0054] The zeroth-order basis function is given by the following formula: The recursive relation is shown by the following formula: The third step is to input a scalar for each dimension. z i Using the aforementioned basis functions, its value in... K =Responses on 8 cubic spline bases, thus achieving nonlinear expansion of single-dimensional features: The fourth step is to stack the expanded results of all 512 dimensions to form a base matrix. B Its size is 512×8. The matrix's _th i The line represents the input of the first line. i The response of dimension on each basis function: To facilitate the subsequent use of a unified linear transformation to complete the discriminative mapping, this embodiment expands the feature matrix obtained from the spline basis. B The inputs are vectorized in a fixed order and flattened into a column vector of length 4096. This step concatenates the response coefficients of each input dimension on multiple B-spline basis functions, integrating the nonlinear expansion terms, which were originally scattered across different dimensions and basis function indices, into a single high-dimensional feature representation. Through this vectorized representation, subsequent logits calculations can weightedly combine all expansion terms in a consistent parameter form, thereby achieving a unified mapping from the high-dimensional nonlinear feature space to the class space and facilitating stable gradient backpropagation and parameter updates during training.

[0055] The fifth step is to learn only linear weights on a fixed spline basis.α j,i,k The category logits are obtained through linear combination: Step 6: Convert logits into a probability distribution using Softmax to obtain the model's predicted output. Step 7: On labeled samples, use the cross-entropy loss function as the training objective. Furthermore, FreeMatch-based semi-supervised learning strategies include: Forward inference is performed on the original image of the unlabeled sample to obtain the class probability distribution of the network output. This distribution is used as a soft pseudo-label to characterize the model's prediction tendency for each class. Furthermore, the maximum value in the probability distribution and its corresponding class index are extracted and used as the confidence score and hard pseudo-label of the sample, respectively. Based on the aforementioned confidence level and hard pseudo-labels, an adaptive confidence threshold strategy is used to screen unlabeled samples. Perform multi-view analysis on the selected unlabeled samples. Figure 1 Dedication training; Calculate the total loss and update the model parameters.

[0056] Specifically, this embodiment uses the FreeMatch semi-supervised framework. Through the adaptive update mechanism of global threshold and class-based threshold, the utilization intensity of unlabeled samples can be dynamically adjusted with the training stage and class learning state, thereby improving the reliability of pseudo-label screening and the stability of unlabeled learning, and enhancing the generalization performance of the model under cross-conditions.

[0057] Based on the 112 samples collected from 28 rotational speeds and 4 load combinations, GFCC acoustic feature maps were extracted. Of these, 28 samples under the 150 N·m load condition had subjective rating labels, while the remaining 84 samples under the 0 N·m, 50 N·m, and 100 N·m load conditions served as an unlabeled set for semi-supervised training. To avoid indirect use of validation set information during the semi-supervised stage, this embodiment employs a rotational speed isolation strategy when constructing the unlabeled training pool. For any rotational speed value appearing in the supervised validation set, the corresponding unlabeled samples under the 0 N·m, 50 N·m, and 100 N·m load conditions are removed from the training pool. This process ensures that the unlabeled data and the supervised validation set remain independent of key load variables, reducing the risk of information leakage and allowing subsequent evaluations to better reflect the model's true generalization performance across load conditions.

[0058] FreeMatch semi-supervised learning task design: The key to FreeMatch semi-supervised learning lies in using an exponential moving average mechanism to adaptively estimate the class confidence threshold, and combining it with multi-view... Figure 1 Consistency regularization continuously constrains model output on unlabeled data, thereby reducing reliance on manually set warm-up stages and fixed threshold strategies. Its basic process is as follows: Figure 6 As shown in the diagram. Specifically, the model first generates a predicted distribution for unlabeled samples on a weakly augmented view, and then generates pseudo-labels and corresponding confidence scores based on this distribution. Subsequently, the threshold is dynamically updated using EMA, allowing the selection criteria to adaptively adjust with the training process and the difficulty of class learning. This avoids early thresholds that are too strict, leading to insufficient effective samples, or too lenient, introducing noisy pseudo-labels. Based on this, a consistency constraint is introduced only for unlabeled samples that meet the current threshold conditions, requiring that the prediction and pseudo-label of the same sample remain consistent under a strongly augmented view. This forms a stable learning signal on the unlabeled data and gradually corrects the decision boundary.

[0059] Pseudo-labels and confidence generation: Specifically, the original image of the unlabeled sample is first forward-inferred to obtain the class probability distribution of the network output, and this distribution is used as a soft pseudo-label to characterize the model's prediction tendency for each class. Based on this, the maximum value in the probability distribution and its corresponding class index are further extracted, which are used as the confidence score and hard pseudo-label of the sample, respectively, for subsequent threshold screening and consistency training stages to determine and utilize the reliability of the unlabeled sample.

[0060] The first step is to extract the original GFCC acoustic feature map. ub ( w As input, a soft probability distribution is obtained through model prediction and used as pseudo-labels: in, fθ Indicates by parameters θ The prediction function of the control. qb It is a sample b Probability vectors belonging to each category.

[0061] The second step, according to q b Define the predicted category of the sample. y b and its confidence level conf b for q b The highest probability and its corresponding category: Weak enhancement prediction generates soft pseudo-labels for unlabeled samples, and hard pseudo-labels and confidence scores are extracted to provide a basis for subsequent threshold screening and consistency training.

[0062] Adaptive confidence threshold update: FreeMatch's adaptive confidence threshold strategy uses a global threshold driven by the intra-batch prediction confidence statistic and updated smoothly using an exponential moving average (EMA) to mitigate the impact of small-batch fluctuations on the screening criteria. Simultaneously, it scales the global threshold using the EMA statistics of class-specific confidence to obtain a class-related threshold, thereby achieving more stable unlabeled sample screening under different class difficulties and sample imbalance conditions.

[0063] The first step is to calculate the average of the maximum confidence scores of all current unlabeled samples. s t : in, M This represents the number of unlabeled samples in this batch.

[0064] The second step is to update the global threshold using EMA. τt : in, λ Let EMA momentum be 0.999. τ 0 initialized to 1 / C , C The number of categories is 9 in this embodiment.

[0065] The third step involves introducing a class-based thresholding strategy to account for imbalances and difficulty levels across different categories. This strategy involves calculating the average confidence level for each category and scaling the global threshold accordingly.

[0066] Maintain the EMA mean of the confidence level for each category. pt ( c ): in, q b,c It is a sample b Category c The probability of.

[0067] The fourth step is to scale the global threshold according to the relative magnitude of the average confidence scores for each category to obtain the categories. c threshold τ ( c ): When a certain category exhibits a high overall prediction confidence during the current training phase, its class-specific confidence statistic... pt ( c The relative size is larger, thus the category threshold obtained by scaling is larger. τt (c The confidence level will be correspondingly increased, thereby imposing stricter screening criteria on unlabeled samples to suppress the introduction of low-quality pseudo-labels and improve the reliability of pseudo-supervision signals. Conversely, for categories with generally low prediction confidence, it usually reflects that the discrimination boundary of that category has not yet been sufficiently stabilized or that the samples are more confusing. pt ( c Smaller τt ( c The threshold decreases accordingly, which can appropriately expand the coverage of available unlabeled samples of this category while controlling noise risk, thus alleviating the training bias caused by the imbalance in sample utilization between categories. Furthermore, to avoid insufficient effective samples due to an excessively high threshold in the early stages of training, or optimization oscillations caused by an excessively low threshold leading to an increased proportion of noise and false labels, this embodiment... τt ( c Apply upper and lower bound constraints to limit it to the interval [0.4, 0.95] to achieve threshold clamping operation and enhance the numerical stability of the threshold update process.

[0068] This class-adaptive strategy helps ensure fairness among different classes in unsupervised training, improves sample coverage, and enhances training stability.

[0069] Sample selection and consistency training: Based on the aforementioned pseudo-label generation, confidence-based filtering and multi-view techniques are employed. Figure 1 Consistency constraints are used to construct a semi-supervised training loss, enabling joint optimization of unlabeled samples.

[0070] The first step is to examine unlabeled samples. b Check if the prediction confidence score exceeds the corresponding class threshold. If the sample's confidence score is greater than the threshold, its pseudo-label is considered highly confident and reliable, and it is selected into the unsupervised training set. Otherwise, it is discarded. A filtering mask is recorded using an indicator function: If the above conditions are met, the value is 1; otherwise, it is 0. Let Nsel be the number of unlabeled samples retained after filtering. Only unlabeled samples that meet the confidence threshold are included in the subsequent calculation of unsupervised loss.

[0071] The second step involves FreeMatch using multiple views to process the selected unlabeled samples. Figure 1 Consistency loss is used to fully utilize unlabeled data. Specifically, for each unlabeled sample, 48 strongly enhanced views are generated through strong enhancement. Strong enhancement is obtained by adjusting the image Gamma {0.84, 0.92, 1.08, 1.16}, rotating {90°, 180°, 270°}, and mirroring, and the probability distribution of the strongly enhanced views is obtained by model prediction.

[0072] The third step, within the consistency training framework, assumes that the predicted distribution of the strongly augmented view should be consistent with the pseudo-labels of the original samples, and defines the average cross-entropy between the predicted distribution of the filtered unlabeled samples and their soft pseudo-labels as the unsupervised consistency loss: in, CE ( q b ,p s,b ) indicates with q b For the goal, p s,b This is the predicted cross-entropy.

[0073] The fourth step is to calculate the total training loss of the model, which consists of supervised and unsupervised components: in, L l represents the cross-entropy loss during supervised training. λu This is the unsupervised loss weight, with a value of 1. The model loss is a weighted sum of supervised and unsupervised terms, and the model parameters are updated via backpropagation.

[0074] Through the aforementioned semi-supervised training process, unlabeled samples are incorporated into parameter updates through pseudo-label supervision and consistency constraints. This ensures that the data seen during the training phase is no longer limited to a single labeled working condition, thus achieving more comprehensive sample coverage across the working condition dimension. The feature distribution information carried by unlabeled data continuously constrains the decision boundary during iteration, helping to mitigate the working condition bias that can easily arise when relying solely on a small number of labeled samples. Simultaneously, confidence threshold screening limits the interference of low-confidence pseudo-labels on the optimization direction, enabling multi-view... Figure 1 Consistency loss enables the model to maintain output stability under enhanced perturbations, improving the robustness of feature representation and the controllability of the training process. Overall, this strategy provides a more reliable training basis for sound quality level discrimination under different load and speed conditions.

[0075] Training of VG-FPS-ViT-KAN sound quality assessment model based on semi-supervised learning: During training, cross-entropy loss and classification accuracy are used as process evaluation metrics to characterize the model's convergence behavior and changes in its discriminative ability. Figure 7 and Figure 8 The curves showing the changes in loss and accuracy with the number of iterations on the training and validation sets are presented respectively.

[0076] Regarding the classification accuracy curve, the model exhibits a significant upward trend in the early stages of training, indicating that the network can capture the main discriminative information related to sound quality levels within a limited number of iterations. As training progresses, the rate of accuracy improvement gradually decreases and plateaus, with the training and validation set curves maintaining similar evolutionary trajectories overall, without a continuously widening gap with increasing iterations. This phenomenon suggests that the model possesses relatively stable generalization performance under the current data scale, enhancement strategy, and regularization configuration. At the end of training, the validation set accuracy stabilizes at 97.22%, reflecting a high level of consistency in the model's classification of validation samples.

[0077] Regarding the cross-entropy loss curve, both training and validation losses decrease overall with each iteration, with a larger decrease in the early stages, indicating that the deviation between the model's output distribution and the true labels is rapidly compressed. The mid-stage loss curve exhibits intermittent fluctuations and peaks, mainly due to the variance of mini-batch stochastic gradient estimation and the perturbations of the input distribution and feature pathways caused by mechanisms such as random augmentation and random deactivation during training. These fluctuations do not alter the overall decreasing trend of the loss, and in the later stages, the two curves gradually converge and stabilize at a low level. The final validation loss is approximately 0.04, indicating that the parameter updates have entered a stable range and the model has achieved effective convergence.

[0078] Based on the two types of process indicators, it can be concluded that the constructed sound quality evaluation model can achieve stable optimization under the given training configuration. The training process does not show a significant overfitting expansion trend, and the performance on the validation set is consistent with that on the training set.

[0079] To address the challenges of varying noise time-frequency structure and limited subjectively labeled samples under multiple speed and load conditions in permanent magnet synchronous motors, this embodiment constructs an acoustic feature map using GFCC and time-series difference features, employs ViT for feature extraction, and proposes a VG-FPS strategy-guided differential tokenization to improve the representation efficiency of key regions. KAN is introduced at the discrimination end to enhance nonlinear mapping capabilities. Furthermore, FreeMatch semi-supervised learning is combined with unlabeled operating condition samples for consistency constraints and adaptive threshold pseudo-label training. Finally, a motor noise sound quality evaluation model suitable for small sample sizes and multiple operating conditions is established and validated.

[0080] Based on noise acquisition experiments during the actual operation of a permanent magnet synchronous motor, this embodiment acquires noise samples under multiple operating conditions and establishes a sample library for typical speed ranges where motor noise is dominant. Given that the calculation of psychoacoustic objective parameters involves integration and time accumulation, directly applying the traditional process based on the steady-state assumption can easily introduce biases. Therefore, this embodiment characterizes signal stability from both the time and frequency domains. The rate of change of the coefficient of variation is used to characterize the time-varying amplitude of the relative dispersion of the amplitude, and the rate of change of spectral flatness is used to measure the time-varying degree of energy distribution uniformity. Automatic identification and segmentation of quasi-steady-state intervals are achieved by setting dual thresholds. Based on the quasi-steady-state segmentation results, and combined with the duration of each quasi-steady-state segment and traditional psychoacoustic calculation models, this embodiment constructs a method for calculating psychoacoustic objective parameters under non-steady-state conditions and establishes a five-dimensional psychoacoustic objective parameter database for noise. Based on this, in order to obtain a subjective scale corresponding to the objective parameters, this embodiment uses five-dimensional psychoacoustic features as the basis for sample characterization. First, nonlinear dimensionality reduction and clustering are used to achieve structured grouping of samples. Then, within each group, pairwise comparative listening and evaluation tests are organized to generate subjective annoyance scores. Finally, the subjective evaluation results are obtained by combining consistency tests and cross-group standard unification.

[0081] To address the practical constraints of high cost in acquiring sound quality labels and limited sample size, a semi-supervised VG-FPS-ViT-KAN sound quality evaluation framework is constructed. At the feature representation level, first-order and second-order difference channels are introduced on top of static GFCC cepstral coefficients to form a three-channel dynamic feature map. This allows the input to simultaneously include auditory scale spectral envelope information and its rate and trend of change over time, thereby enhancing the representation of time-frequency details such as order shifts, amplitude fluctuations, and local high-frequency texture abrupt changes under non-steady-state conditions. At the input modeling level, a VG-FPS strategy is proposed, using block variance as a statistical measure of regional information content and structural complexity. Fine-grained token partitioning is applied to high-variance key regions, and local self-attention aggregation is performed within blocks to strengthen local correlation modeling. Simultaneously, pooling compression is used for low-variance regions to reduce redundant tokens and computational burden, thereby improving the utilization efficiency of effective discrimination cues while controlling sequence length. At the training mechanism level, FreeMatch semi-supervised learning is introduced, which incorporates unlabeled working condition samples into the training process. Pseudo-labels are generated on the weakly enhanced view and filtered by an adaptive confidence threshold to suppress the interference of low-confidence pseudo-supervisory signals on the optimization direction. At the same time, consistency constraints are applied under the strongly enhanced view to stabilize the output and continuously correct the decision boundary. This allows the distribution information of unlabeled data to form a long-term constraint on the update of model parameters, thereby alleviating the distribution bias caused by the concentration of supervision information in a single working condition and improving the generalization and stability of sound quality level discrimination under multiple working conditions.

[0082] This embodiment focuses on the sound quality evaluation of permanent magnet synchronous motor noise under multiple operating conditions. It designs and implements a subjective evaluation experiment, constructs a sound quality evaluation framework that considers both auditory perception mechanisms and deep representation capabilities, and expands the scope of training data utilization through semi-supervised learning to improve the model's generalization stability under small sample sizes and cross-operating conditions. These methods not only provide a technical path for predicting motor noise sound quality but also offer a valuable modeling approach for research on the sound quality evaluation of other rotating machinery noise with significant time-varying characteristics.

[0083] This embodiment significantly improves evaluation accuracy: The constructed VG-FPS-ViT-KAN sound quality evaluation model, combined with semi-supervised learning, achieves a classification accuracy of 97.22% in the sound quality evaluation task of permanent magnet synchronous motor noise. This is an improvement over the basic ViT (88.61%), ViT-KAN (93.48%), and the version without semi-supervised learning (97.08%). Furthermore, the mean absolute error is lower, and the prediction results are closer to the subjective scores.

[0084] By introducing first-order and second-order differences of GFCC to construct a three-channel dynamic feature map, the time-varying evolution law of noise under unsteady conditions (such as order shift, amplitude fluctuation, and local high-frequency texture abrupt change) is effectively characterized, overcoming the shortcomings of traditional static features in representing dynamic information.

[0085] The proposed VG-FPS strategy uses block variance to measure information density, performs fine-grained tokenization and local self-attention aggregation on high-information regions, and pooling compression on low-information regions. This strategy enhances the expression of key discriminative clues while controlling computational overhead. Ablation experiments confirm that performance significantly decreases after removing this strategy.

[0086] A FreeMatch semi-supervised learning framework was introduced, utilizing a large number of unlabeled samples (0, 50, and 100 N·m loads) for adaptive threshold pseudo-label selection and consistency constraint training, effectively alleviating the operational bias caused by insufficient labeled samples (only 150 N·m load). Seven-fold cross-validation results show that the model exhibits low fluctuation and high stability under different data partitions.

[0087] This approach enables semi-supervised learning by making full use of readily available unlabeled operating data, even with a small number of labeled samples. It significantly reduces reliance on expensive and time-consuming subjective listening and evaluation experiments, and has good engineering application value.

[0088] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made to the technical solutions of the present invention by those skilled in the art without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.

Claims

1. A method for intelligent evaluation of motor sound quality based on semi-supervised learning and variance-guided screening, characterized in that, include: Noise signals of permanent magnet synchronous motors under various operating conditions were collected, and a noise dataset containing labeled and unlabeled samples was constructed. For each noise sample in the noise dataset, static GFCC cepstral coefficients are extracted, and time derivative features are introduced to construct a dynamic feature map; Based on the variance-guided fine-grained feature block screening VG-FPS strategy, different regions of the dynamic feature map are differentiated to generate a structured token sequence. A visual transformer-Kolmogorov-Arnold network sound quality evaluation model is constructed, denoted as the ViT-KAN sound quality evaluation model. The ViT-KAN sound quality evaluation model takes the structured token sequence as input, uses the ViT network to extract global time-frequency features, and uses KAN to replace the traditional multilayer perceptron (MLP) classification head to map the global time-frequency features into sound quality level prediction results. A semi-supervised learning strategy based on FreeMatch was adopted to train the ViT-KAN sound quality evaluation model using the noise dataset. The motor noise signal to be evaluated is input into the trained ViT-KAN sound quality evaluation model to obtain the corresponding sound quality level.

2. The intelligent evaluation method for motor sound quality based on semi-supervised learning and variance-guided screening as described in claim 1, characterized in that, Constructing the dynamic feature map includes: The noise sample signal is sequentially processed by auditory filtering, subband energy calculation, logarithmic compression and cepstral transform to obtain the GFCC cepstral coefficients corresponding to each frame of signal. The first few GFCC cepstral coefficients are selected as static GFCC cepstral coefficients. The time derivative feature is introduced, including first-order difference and second-order difference; wherein, the first-order difference: by approximating the derivative of the change in the cepstral sequence of adjacent frames, it characterizes the rate of change of the spectral envelope with time; the second-order difference: describes the trend of the rate of change. The calculation formula is as follows: ; ; Where, Δ ci It is a first-order difference, Δ2 ci It is a second-order difference. ci Indicates the first i The static GFCC feature vector corresponding to the frame noise signal ci-1 and ci+ 1 represents the first i The static GFCC feature vectors corresponding to the previous and next frames, Δ ci Indicates the first i First-order temporal difference of frame static GFCC eigenvectors, Δ² ci This represents the second-order temporal difference of the static GFCC feature vector of the i-th frame; The first-order difference, second-order difference, and static GFCC feature vectors are concatenated to form a three-channel image feature representation.

3. The intelligent evaluation method for motor sound quality based on semi-supervised learning and variance-guided screening as described in claim 1, characterized in that, Based on the variance-guided fine-grained feature block screening VG-FPS strategy, different regions of the dynamic feature map are differentiated to generate a structured token sequence, including: The dynamic feature map is spatially divided into multiple non-overlapping coarse-grained image blocks, the variance of the coarse block pixels is calculated, and the variance contribution rate is calculated. The dynamic feature map is divided into high-variance regions and low-variance regions based on the variance contribution rate. A dual-path token generation mechanism, based on fine-grained self-attention aggregation in high-variance regions and pooling compression in low-variance regions, generates a structured token sequence.

4. The intelligent evaluation method for motor sound quality based on semi-supervised learning and variance-guided screening as described in claim 3, characterized in that, Calculating the variance contribution rate includes: The total variance of the entire image is obtained by summing the variances of all block pixels. Sort the variance contributions in descending order to obtain the sorted sequence; Based on the sorted sequence and the total variance of the entire graph, the variance contribution rate of coarse blocks is obtained.

5. The intelligent evaluation method for motor sound quality based on semi-supervised learning and variance-guided screening according to claim 3, characterized in that, The dual-path token generation mechanism includes: Each high-variance coarse-grained image block is divided into multiple fine-grained sub-blocks, and each fine-grained sub-block is mapped to a corresponding sub-block Token; Layer normalization and linear mapping are performed on sub-block tokens belonging to the same high-variance coarse-grained image block to obtain query vector, key vector and value vector, and scaling dot product attention output is calculated based on the query vector, key vector and value vector; The scaled dot product attention output is added to the corresponding original sub-block token by residual addition to obtain multiple updated fine-grained sub-block tokens, and the multiple fine-grained sub-block tokens are used as fine-grained representations of the high-variance coarse-grained image block. Each low-variance coarse-grained image block is divided into multiple fine-grained sub-blocks, and the sub-block tokens corresponding to the multiple fine-grained sub-blocks are pooled to compress each low-variance coarse-grained image block into a coarse-grained token. The fine-grained sub-block tokens corresponding to all high-variance coarse-grained image blocks and the coarse-grained tokens corresponding to all low-variance coarse-grained image blocks are combined, and the two-dimensional position information of each token in the original dynamic feature map is retained to generate the structured token sequence.

6. The intelligent evaluation method for motor sound quality based on semi-supervised learning and variance-guided screening according to claim 1, characterized in that, The ViT-KAN sound quality evaluation model includes: The Token embedding module is used to map the structured Token sequence to an embedding space of a preset dimension, and to add a classification marker CLS Token to the structured Token sequence. The two-dimensional rotation position encoding module is used to generate a two-dimensional rotation position encoding based on the horizontal and vertical coordinates of each token in the dynamic feature map, and inject the two-dimensional rotation position encoding into the query vector and key vector in the Transformer self-attention calculation respectively; The ViT feature extraction module includes a multi-layer Transformer encoder. Each Transformer encoder includes a multi-head self-attention sub-layer, a feedforward neural network sub-layer, residual connections, and layer normalization, which are used to model the global correlation between different tokens and obtain global time-frequency features through the CLS token of the final layer. The KAN classification head is used to convert the global time-frequency features into prediction results corresponding to different sound quality levels through nonlinear mapping based on spline basis functions.

7. The intelligent evaluation method for motor sound quality based on semi-supervised learning and variance-guided screening according to claim 1, characterized in that, Extracting global time-frequency features using ViT networks includes: The ViT network divides the input features into several sub-blocks and maps them into sequence form. It then uses a multi-head self-attention mechanism to model the correlation between the sub-blocks, thereby representing the overall structural information of the input and obtaining global time-frequency features.

8. The intelligent evaluation method for motor sound quality based on semi-supervised learning and variance-guided screening according to claim 1, characterized in that, The KAN mapping of global time-frequency features to sound quality level prediction results includes: The global time-frequency features are normalized. A cubic B-spline basis function system is constructed over a preset interval to achieve localized approximation of the nonlinear mapping; For each input scalar dimension, its response on multiple cubic spline bases is calculated using the constructed basis functions, thereby achieving nonlinear expansion of the single-dimensional feature. Stack the expanded results of all dimensions to form a basis matrix, then vectorize the basis matrix in a fixed order and flatten it into column vectors; On a fixed spline basis, only linear weights are learned, and the class logits are obtained through linear combination; The class logits are converted into a probability distribution using Softmax to obtain the model's predicted output; For labeled samples, the cross-entropy loss function is used as the training objective.

9. The intelligent evaluation method for motor sound quality based on semi-supervised learning and variance-guided screening according to claim 1, characterized in that, Semi-supervised learning strategies based on FreeMatch include: Forward inference is performed on the original image of the unlabeled sample to obtain the class probability distribution of the network output. This distribution is used as a soft pseudo-label to characterize the model's prediction tendency for each class. Furthermore, the maximum value in the probability distribution and its corresponding class index are extracted and used as the confidence score and hard pseudo-label of the sample, respectively. Based on the aforementioned confidence level and hard pseudo-labels, an adaptive confidence threshold strategy is used to screen unlabeled samples. Perform multi-view consistency training on the selected unlabeled samples; Calculate the total loss and update the model parameters.

10. The intelligent evaluation method for motor sound quality based on semi-supervised learning and variance-guided screening according to claim 9, characterized in that, The adaptive confidence threshold strategy includes: Calculate the average of the maximum confidence scores for all current unlabeled samples; The global threshold is updated based on the average value using the EMA method. A class-based thresholding strategy is introduced to calculate the average confidence score for each class and scale the global threshold accordingly. The global threshold is scaled according to the relative magnitude of the average confidence of each category to obtain the threshold for each category.