Non-contact feature recognition method for on-state of main circuit of inverter power supply
Patent Information
- Application Number
- CN202611281566.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-24
- Publication Date
- 2026-09-25
AI Technical Summary
该类高维特征空间虽然能够保留较丰富的状态表征信息,但同时也包含大量与接通状态弱相关或无关的噪声特征,且不同特征之间往往存在较强重叠和冗余
[0019]本发明通过传感器阵列非接触式采集多通道电磁辐射信号,避免了接触式检测带来的潜在安全隐患;并结合小波包分解、排列熵和能量占比系数级联,刻画信号在不同频段上的复杂度变化与能量分布特征,增强了高维初始特征对主回路状态的表征能力。进一步地,利用最大信息系数筛除低相关特征,降低模型输入维度;采用近邻传播聚类挖掘特征之间的相似结构,并在序列前向浮动选择过程中引入由交叉熵损失和冗余度惩罚项构成的组合评价函数,对同簇待选特征施加约束,减少冗余特征重复入选,从而兼顾最优特征子集的多样性和判别能力。最后,基于筛选得到的最优特征子集训练梯度提升决策树辨识模型,在降低运算开销的同时提高分类精度,实现了逆变电源主回路接通状态的高效非接触辨识。
Smart Images

Figure CN122818098A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of inverter power supply, and in particular relates to a non-contact feature identification method for the on-state of the main circuit of an inverter power supply. Background Technology
[0002] Inverters, as core equipment in power conversion systems, are widely used in critical fields such as new energy power generation, industrial drives, aerospace, and rail transportation. Their operational reliability directly affects the safety and stability of the entire system. Monitoring and identifying the on-state of the inverter's main circuit is a crucial prerequisite for equipment condition assessment, fault early warning, and safety control. Traditional inverter condition monitoring typically relies on contact-based measurement methods such as voltage and current sensors. This not only requires altering the original circuit topology, increasing hardware design complexity and subsequent maintenance costs, but also presents challenges such as difficulty in electrical isolation and high potential safety risks under high-voltage and high-current environments. Non-contact identification methods, by detecting the spatial electromagnetic radiation signals generated by key components in the main circuit during switching or operation, infer the electrical state of key components and the main circuit. Without direct connection to the circuit under test, these methods offer advantages such as high security, non-invasiveness, and ease of online deployment, and have become an important development direction in the field of intelligent condition monitoring of power electronic equipment.
[0003] In real industrial environments, the electromagnetic signals radiated from the main circuit of an inverter power supply typically exhibit non-stationary, nonlinear, and strongly time-varying characteristics, making them susceptible to interference from complex background noise and compounded by electromagnetic coupling effects between multiple devices. To accurately identify the state information contained within these signals, it is usually necessary to rely on sensor arrays to acquire multi-channel time-series signals and extract spatiotemporal and time-frequency domain features such as multi-band energy and entropy values, thereby forming a high-dimensional feature vector set. While this type of high-dimensional feature space can retain relatively rich state representation information, it also contains a large number of noise features that are weakly correlated with or unrelated to the on / off state, and there is often strong overlap and redundancy between different features. If a high-dimensional mixed feature is directly used for classification model training without an effective screening mechanism, it will lead to increased computational overhead, decreased recognition response speed, and a high risk of the machine learning model falling into the "curse of dimensionality," resulting in overfitting and ultimately reducing the accuracy of the recognition results and the model's generalization ability. Therefore, how to extract and evaluate effective features from complex and ever-changing electromagnetic radiation signals, construct an improved feature selection strategy that can reduce redundant information and alleviate the problem of strong feature correlation, and realize efficient non-contact identification of the main circuit connection status of inverter power supply are technical problems that urgently need to be solved in this field. Summary of the Invention
[0004] According to one aspect of the present invention, a non-contact feature identification method for the on-state of the main circuit of an inverter power supply is proposed to solve the above-mentioned technical problem, comprising the following steps:
[0005] The system acquires multiple sets of time-series electromagnetic radiation signals from key components of the main circuit collected by the sensor array and constructs a multi-channel spatiotemporal signal matrix. Wavelet packet decomposition is performed on each channel signal of the multi-channel spatiotemporal signal matrix to obtain the reconstruction coefficient of each frequency band. The entropy of each frequency band is calculated and concatenated with the energy proportion coefficient to form a high-dimensional hybrid feature vector. The high-dimensional hybrid feature vectors of multiple sets of signals form a training feature set. The maximum information coefficient between each dimension of the training feature set and the connection status label is calculated as the relevance. Features with relevance values lower than the mean are filtered out to obtain a candidate feature set. The similarity between each feature in the candidate feature set is calculated to form a similarity matrix. The median of the elements in the similarity matrix is used as the nearest neighbor propagation clustering bias parameter to cluster the candidate feature set to obtain feature clusters and cluster affiliation information. The sequential forward floating selection algorithm is used to search the candidate feature set. A combined evaluation function is set, which is weighted by the gradient boosting decision tree cross-entropy loss and the redundancy penalty term. The redundancy penalty term is used to penalize the candidate features in the same cluster as the selected features to determine the optimal feature subset. The gradient boosting decision tree is trained based on the optimal feature subset and the on / off status label to obtain the identification model. The optimal feature subset of the real-time acquired signal is extracted and input into the identification model, and the identification result is output.
[0006] Optionally, the step of acquiring multiple sets of timing electromagnetic radiation signals from key components of the main circuit collected by the sensor array and constructing a multi-channel spatiotemporal signal matrix includes: For each group of signals, the rising edge of the insulated gate bipolar transistor turn-on control signal sent by the main controller of the inverter power supply system to the drive board is used as the synchronization trigger point. In response to the synchronous trigger point, the non-contact electromagnetic sensor array deployed above the key components of the main circuit of the inverter power supply is controlled to perform synchronous high-frequency sampling to obtain the timing electromagnetic radiation signal within a fixed time window. Each sensor in the sensor array collects a one-dimensional time-series signal as an independent channel. The time-series data of each channel are arranged according to the spatial positional relationship of the sensor array to construct a multi-channel spatiotemporal signal matrix containing spatial topology dimension and time sequence dimension.
[0007] Optionally, the step of performing wavelet packet decomposition on each channel signal of the multi-channel spatiotemporal signal matrix to obtain the reconstruction coefficients of each frequency band includes: Select wavelet basis functions and perform N-level wavelet packet decomposition on the time-series signals of each channel in the multi-channel spatiotemporal signal matrix to obtain the node coefficients of 2^N frequency bands in the Nth level. The data after retaining the node coefficients of a specific frequency band and setting the node coefficients of the remaining frequency bands to zero is subjected to inverse wavelet packet transform to obtain a reconstructed signal with the same length as the original time series signal. The reconstructed signals of each frequency band are extracted and used as the reconstruction coefficients of the corresponding frequency band.
[0008] Optionally, the calculation of the entropy of each frequency band arrangement and its concatenation with the energy proportion coefficient to form a high-dimensional hybrid feature vector includes: Set the embedding dimension and delay time for phase space reconstruction, and perform phase space reconstruction on the reconstruction coefficients of each frequency band; Arrange the elements in each reconstructed state vector in ascending order, obtain the index of the sorting pattern, count the probability of each sorting pattern and calculate the Shannon entropy of the probability distribution to obtain the permutation entropy. Calculate the sum of squares of the reconstruction coefficients for each frequency band, divide the sum of squares of the current frequency band by the sum of the sums of squares of the reconstruction coefficients of all frequency bands under this channel, and obtain the energy proportion coefficient of the corresponding frequency band in this channel; The permutation entropy and energy proportion coefficient of each frequency band in each channel are concatenated in frequency band order to construct a high-dimensional hybrid feature vector.
[0009] Optionally, calculating the maximum information coefficient between each dimension of the training feature set and the connection status label as the relevance includes: Obtain the numerical distribution of features in each dimension of the training feature set and the category distribution of the preset connection status labels; The number of grids in the label dimension is fixed according to the number of categories of the connected status label. The grid division scale is traversed along the feature dimension. At each grid division scale, the joint probability between each dimension feature and the connected status label and their respective marginal probabilities are calculated. The mutual information value is calculated based on the joint probability and marginal probability under the current grid partitioning scale and then normalized. Find the maximum value of the normalized mutual information value under all grid division scales, and use it as the maximum information coefficient between the corresponding dimensional feature and the on / off state label.
[0010] Optionally, the step of filtering out features below the relevance mean to obtain a candidate feature set includes: Sum the maximum information coefficients of all dimensional features and the connection status labels, and divide by the total number of feature dimensions to obtain the mean of the maximum information coefficients; Set the mean value as the filtering threshold; Compare the maximum information coefficient of each dimension feature with the filtering threshold one by one, and delete the dimension features whose maximum information coefficient is less than the filtering threshold; The dimensional features with a maximum information coefficient not less than the filtering threshold are retained to form the candidate feature set.
[0011] Optionally, the step of extracting the optimal feature subset of the real-time acquired signal and inputting it into the identification model to output the identification result includes: Acquire the timing electromagnetic radiation signal of the main circuit when the inverter is operating in real time, and calculate the high-dimensional hybrid feature vector of the real-time acquired signal; Based on the feature index corresponding to the optimal feature subset, extract the corresponding real-time feature subset from the high-dimensional hybrid feature vector; The real-time feature subset is input into the trained gradient boosting decision tree model for tree-by-tree traversal and ensemble prediction calculation, and the corresponding classification result label is output as the on / off state identification result of the main loop.
[0012] According to another aspect of the present invention, a non-contact feature identification system for the on-state of the main circuit of an inverter power supply is proposed, comprising the following modules: The acquisition module is used to acquire multiple sets of time-series electromagnetic radiation signals of key components in the main circuit collected by the sensor array and construct a multi-channel spatiotemporal signal matrix. Wavelet packet decomposition is performed on each channel signal of the multi-channel spatiotemporal signal matrix to obtain the reconstruction coefficient of each frequency band. The entropy of each frequency band is calculated and concatenated with the energy proportion coefficient to form a high-dimensional hybrid feature vector. The high-dimensional hybrid feature vectors of multiple sets of signals form a training feature set. The calculation module is used to calculate the maximum information coefficient between each dimension of the training feature set and the connection status label as the relevance, filter out features with a relevance lower than the mean to obtain a candidate feature set, calculate the similarity between each feature in the candidate feature set to form a similarity matrix, and use the median of the elements in the similarity matrix as the nearest neighbor propagation clustering bias parameter to cluster the candidate feature set to obtain feature clusters and cluster affiliation information. The output module is used to search the candidate feature set using a sequential forward floating selection algorithm, and sets a combined evaluation function weighted by gradient boosting decision tree cross-entropy loss and redundancy penalty term. The redundancy penalty term is used to penalize candidate features in the same cluster as the selected features to determine the optimal feature subset. Based on the optimal feature subset and the on / off status label, a gradient boosting decision tree is trained to obtain the identification model. The optimal feature subset of the real-time acquired signal is extracted and input into the identification model, and the identification result is output.
[0013] Preferably, the step of acquiring multiple sets of timing electromagnetic radiation signals from key components of the main circuit collected by the sensor array and constructing a multi-channel spatiotemporal signal matrix includes: For each group of signals, the rising edge of the insulated gate bipolar transistor turn-on control signal sent by the main controller of the inverter power supply system to the drive board is used as the synchronization trigger point. In response to the synchronous trigger point, the non-contact electromagnetic sensor array deployed above the key components of the main circuit of the inverter power supply is controlled to perform synchronous high-frequency sampling to obtain the timing electromagnetic radiation signal within a fixed time window. Each sensor in the sensor array collects a one-dimensional time-series signal as an independent channel. The time-series data of each channel are arranged according to the spatial positional relationship of the sensor array to construct a multi-channel spatiotemporal signal matrix containing spatial topology dimension and time sequence dimension.
[0014] Preferably, the step of performing wavelet packet decomposition on each channel signal of the multi-channel spatiotemporal signal matrix to obtain the reconstruction coefficients of each frequency band includes: Select wavelet basis functions and perform N-level wavelet packet decomposition on the time-series signals of each channel in the multi-channel spatiotemporal signal matrix to obtain the node coefficients of 2^N frequency bands in the Nth level. The data after retaining the node coefficients of a specific frequency band and setting the node coefficients of the remaining frequency bands to zero is subjected to inverse wavelet packet transform to obtain a reconstructed signal with the same length as the original time series signal. The reconstructed signals of each frequency band are extracted and used as the reconstruction coefficients of the corresponding frequency band.
[0015] Preferably, the step of calculating the entropy of each frequency band and concatenating it with the energy proportion coefficient to form a high-dimensional hybrid feature vector includes: Set the embedding dimension and delay time for phase space reconstruction, and perform phase space reconstruction on the reconstruction coefficients of each frequency band; Arrange the elements in each reconstructed state vector in ascending order, obtain the index of the sorting pattern, count the probability of each sorting pattern and calculate the Shannon entropy of the probability distribution to obtain the permutation entropy. Calculate the sum of squares of the reconstruction coefficients for each frequency band, divide the sum of squares of the current frequency band by the sum of the sums of squares of the reconstruction coefficients of all frequency bands under this channel, and obtain the energy proportion coefficient of the corresponding frequency band in this channel; The permutation entropy and energy proportion coefficient of each frequency band in each channel are concatenated in frequency band order to construct a high-dimensional hybrid feature vector.
[0016] Preferably, calculating the maximum information coefficient between each dimension of the training feature set and the connection status label as the relevance includes: Obtain the numerical distribution of features in each dimension of the training feature set and the category distribution of the preset connection status labels; The number of grids in the label dimension is fixed according to the number of categories of the connected status label. The grid division scale is traversed along the feature dimension. At each grid division scale, the joint probability between each dimension feature and the connected status label and their respective marginal probabilities are calculated. The mutual information value is calculated based on the joint probability and marginal probability under the current grid partitioning scale and then normalized. Find the maximum value of the normalized mutual information value under all grid division scales, and use it as the maximum information coefficient between the corresponding dimensional feature and the on / off state label.
[0017] Preferably, the step of filtering out features below the relevance mean to obtain a candidate feature set includes: Sum the maximum information coefficients of all dimensional features and the connection status labels, and divide by the total number of feature dimensions to obtain the mean of the maximum information coefficients; Set the mean value as the filtering threshold; Compare the maximum information coefficient of each dimension feature with the filtering threshold one by one, and delete the dimension features whose maximum information coefficient is less than the filtering threshold; The dimensional features with a maximum information coefficient not less than the filtering threshold are retained to form the candidate feature set.
[0018] Preferably, the step of extracting the optimal feature subset of the real-time acquired signal and inputting it into the identification model to output the identification result includes: Acquire the timing electromagnetic radiation signal of the main circuit when the inverter is operating in real time, and calculate the high-dimensional hybrid feature vector of the real-time acquired signal; Based on the feature index corresponding to the optimal feature subset, extract the corresponding real-time feature subset from the high-dimensional hybrid feature vector; The real-time feature subset is input into the trained gradient boosting decision tree model for tree-by-tree traversal and ensemble prediction calculation, and the corresponding classification result label is output as the on / off state identification result of the main loop.
[0019] This invention utilizes a sensor array to non-contactly acquire multi-channel electromagnetic radiation signals, avoiding the potential safety hazards associated with contact detection. Furthermore, it combines wavelet packet decomposition, permutation entropy, and energy proportion coefficient cascades to characterize the complexity variations and energy distribution features of the signal across different frequency bands, enhancing the representational ability of high-dimensional initial features for the main circuit state. Further, it employs the maximum information coefficient to filter out low-relevance features, reducing the model input dimensionality; it uses nearest-neighbor propagation clustering to mine similar structures between features; and it introduces a combined evaluation function consisting of cross-entropy loss and redundancy penalty terms during the sequential forward floating selection process to constrain candidate features within the same cluster, reducing redundant feature duplication and thus balancing the diversity and discriminative power of the optimal feature subset. Finally, based on the selected optimal feature subset, a gradient boosting decision tree identification model is trained, improving classification accuracy while reducing computational overhead, achieving efficient non-contact identification of the inverter power supply main circuit connection state. Attached Figure Description
[0020] Figure 1 A flowchart of the first embodiment; Figure 2 This is a schematic diagram of a transient electromagnetic radiation signal waveform; Figure 3 A schematic diagram showing the distribution of energy proportion coefficients in each frequency band; Figure 4A diagram showing the comparison of recognition results and time taken for each group. Detailed Implementation
[0021] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0022] In the first embodiment, the present invention proposes a non-contact feature identification method for the on-state of the main circuit of an inverter power supply, such as... Figure 1 It includes the following steps: S1. Acquire the time-series electromagnetic radiation signals of multiple main circuit key components collected by the sensor array and construct a multi-channel spatiotemporal signal matrix. Perform wavelet packet decomposition on the signals of each channel of the multi-channel spatiotemporal signal matrix to obtain the reconstruction coefficients of each frequency band. Calculate the entropy of each frequency band and concatenate it with the energy proportion coefficient to form a high-dimensional hybrid feature vector. The high-dimensional hybrid feature vectors of multiple signals form a training feature set.
[0023] A giant magnetoresistive sensor array, arranged around the insulated gate bipolar transistor (IGBT) of the inverter power supply, synchronously acquires the time-series electromagnetic radiation signals of the main circuit in both on and off states at a sampling rate of MHz. After analog-to-digital conversion, a multi-channel spatiotemporal signal matrix containing spatial channel dimensions and temporal sampling dimensions is constructed. The on-state label refers to the state corresponding to when the key switching device of the inverter power supply main circuit is driven effectively and the main power circuit forms a conducting current path; the off-state label refers to the state corresponding to when the key switching device is turned off or the main power circuit does not form a conducting current path. For each row of channel data in the multi-channel spatiotemporal signal matrix, a four-level wavelet packet decomposition is performed using the db4 wavelet basis function to obtain node data for sixteen frequency bands. The node reconstruction method is called to calculate the reconstruction coefficient time series for each frequency band. For the reconstruction coefficient series of each frequency band, the permutation entropy calculation function is used, with the time series embedding dimension parameter set to 4 and the time delay parameter set to 1, to calculate the scalar permutation entropy. The sum of the squares of the values of each element in the reconstruction coefficient sequence of this frequency band is calculated as the absolute frequency band energy, and divided by the sum of the absolute energies of the sixteen frequency bands to obtain the scalar energy proportion coefficient. The entropy of the sixteen frequency bands and the energy proportion coefficients of the sixteen frequency bands of all spatial channels are concatenated in a one-dimensional unfolded arrangement to form a single one-dimensional high-dimensional hybrid feature vector. When the sensor array contains nine spatial channels, a single channel forms a 32-dimensional feature, and the nine channels, cascaded, form a 288-dimensional high-dimensional hybrid feature vector. The one-dimensional high-dimensional hybrid feature vectors corresponding to the signal samples acquired under multiple different operating conditions are stacked row-wise to form a two-dimensional training feature set data matrix.
[0024] In an optional embodiment, the step of acquiring multiple sets of timing electromagnetic radiation signals from key components of the main circuit collected by the sensor array and constructing a multi-channel spatiotemporal signal matrix includes: For each group of signals, the rising edge of the insulated gate bipolar transistor turn-on control signal sent by the main controller of the inverter power supply system to the drive board is used as the synchronization trigger point. In response to the synchronous trigger point, the non-contact electromagnetic sensor array deployed above the key components of the main circuit of the inverter power supply is controlled to perform synchronous high-frequency sampling to obtain the timing electromagnetic radiation signal within a fixed time window. Each sensor in the sensor array collects a one-dimensional time-series signal as an independent channel. The time-series data of each channel are arranged according to the spatial positional relationship of the sensor array to construct a multi-channel spatiotemporal signal matrix containing spatial topology dimension and time sequence dimension.
[0025] The pulse width modulation control signal issued by the digital signal processor of the inverter power supply system is used as a hardware trigger source. For on-state samples, the rising edge of the hardware trigger source is detected as the synchronous trigger point. For off-state samples, the falling edge of the corresponding off control signal is detected, and sampling is started after a preset dead time delay after the falling edge, or sampling is started by a timed trigger signal generated by the data acquisition card while the main circuit remains in the off state. In this way, on-state samples and off-state samples are obtained respectively, ensuring the closed-loop data source of the binary classification training set. The aforementioned method of using the rising edge of the turn-on control signal as the synchronous trigger point is the preferred synchronous acquisition method for on-state samples or turn-on transient signals. The off-falling edge delay trigger or the timed trigger method during the off-hold phase, which are added in the specification for off-state samples, are used to cooperate in constructing binary classification training sets for on-state and off-state, and do not change the overall process of acquiring the timing electromagnetic radiation signals of multiple sets of key components in the main circuit. When the synchronous trigger point is detected, the high-speed data acquisition card connected via an external interrupt is triggered to control the 3×3 near-field magnetic field probe array deployed about 2 to 5 cm above the insulated gate bipolar transistor module to perform synchronous sampling. It should be noted that the above method of using the rising edge of the insulated gate bipolar transistor (IGBT) turn-on control signal as the synchronous trigger point is used for acquiring samples of the on-state or turn-on transient states. For samples of the off-state, the timing point can be the time after the falling edge of the turn-off control signal delayed by a preset dead time, or the timing trigger signal during the period when the main circuit remains off-state, as the synchronous trigger point. The sampling frequency is set to 50MHz to detect the transient characteristics of high-frequency electromagnetic radiation, and the sampling time window is fixed at 10μs, corresponding to acquiring 500 discrete sampling points per trigger. Among them, the sampling window of the on-state samples covers the electromagnetic transient process during the turn-on phase of the switching device, and the sampling window of the off-state samples covers the electromagnetic background response during the stable off-state phase or the off-state holding phase after turn-off. The one-dimensional time-series signals acquired by the nine non-contact electromagnetic sensors in the probe array are sequentially labeled as channels 1 to 9. According to the relative coordinate order of these nine probes in two-dimensional space, these nine one-dimensional time-series data with a length of 500 are vertically stacked and spliced. A multi-channel spatiotemporal signal matrix with dimensions of 9×500 is generated. The row dimension of 9 represents the distribution differences of the electromagnetic radiation field in the main loop spatial topology, and the column dimension of 500 maps the nanosecond-level time-series transient response characteristics of each monitoring point over time. A schematic diagram of the transient electromagnetic radiation signal waveform is shown below. Figure 2 As shown.
[0026] In an optional embodiment, the step of performing wavelet packet decomposition on each channel signal of the multi-channel spatiotemporal signal matrix to obtain the reconstruction coefficients of each frequency band includes: Select wavelet basis functions and perform N-level wavelet packet decomposition on the time-series signals of each channel in the multi-channel spatiotemporal signal matrix to obtain the node coefficients of 2^N frequency bands in the Nth level. The data after retaining the node coefficients of a specific frequency band and setting the node coefficients of the remaining frequency bands to zero is subjected to inverse wavelet packet transform to obtain a reconstructed signal with the same length as the original time series signal. The reconstructed signals of each frequency band are extracted and used as the reconstruction coefficients of the corresponding frequency band.
[0027] The Daubechies db4 wavelet, known for its compact support and smoothness, was selected as the basis function, and the decomposition level N=4. For each row of the multi-channel spatiotemporal signal matrix (i.e., 500 sampling points per channel), a discrete wavelet packet transform algorithm was used to perform a 4-level full-tree structure decomposition. After decomposition, 16 frequency bands of node coefficients were generated at the 4th level, each covering 1 / 16 of the original signal bandwidth. These node coefficients represent the energy distribution of the signal within 16 equal-width frequency band subsets from low to high frequencies. To extract frequency band features without changing the length of the time-domain sequence, a single-branch reconstruction algorithm was used to process the 16 frequency bands. Taking the i-th frequency band as an example, only the node coefficients of that band were retained in the coefficient matrix, while the node coefficient arrays of the other 15 frequency bands were all set to zero, and an inverse wavelet packet transform was performed. Thus, a reconstructed signal sequence of length 500 was generated for each frequency band. The 16 independent one-dimensional arrays of length 500 reconstructed from each channel were extracted and cached as the reconstruction coefficients of the corresponding channel in different frequency bands.
[0028] In an optional embodiment, the calculation of the entropy of each frequency band and its concatenation with the energy proportion coefficient to form a high-dimensional hybrid feature vector includes: Set the embedding dimension and delay time for phase space reconstruction, and perform phase space reconstruction on the reconstruction coefficients of each frequency band; Arrange the elements in each reconstructed state vector in ascending order, obtain the index of the sorting pattern, count the probability of each sorting pattern and calculate the Shannon entropy of the probability distribution to obtain the permutation entropy. Calculate the sum of squares of the reconstruction coefficients for each frequency band, divide the sum of squares of the current frequency band by the sum of the sums of squares of the reconstruction coefficients of all frequency bands under this channel, and obtain the energy proportion coefficient of the corresponding frequency band in this channel; The permutation entropy and energy proportion coefficient of each frequency band in each channel are concatenated in frequency band order to construct a high-dimensional hybrid feature vector.
[0029] For any frequency band reconstruction coefficient sequence of length 500, the embedding dimension m=4 and the delay time τ=1 are set for phase space reconstruction. The coefficient sequence is transformed into 497 state vectors of length 4 through phase space reconstruction. For these 497 state vectors, the four elements inside each vector are sorted in ascending order. Since m=4, there are theoretically 24 permutation patterns. The frequency of these 24 patterns in the 497 vectors is counted and divided by 497 to calculate the probability of each pattern. The permutation entropy is obtained and normalized by dividing by ln(24) so that the permutation entropy value of the frequency band is mapped to the interval [0,1]. The 500 data points of the frequency band reconstruction coefficient sequence are squared one by one and accumulated to obtain the absolute energy value of the frequency band. The absolute energy values of the 16 frequency bands under a single channel are accumulated to obtain the total channel energy. Then, the absolute energy value of the current frequency band is divided by the total channel energy to calculate the energy ratio coefficient accurate to four decimal places. Following the frequency band order from 0 to 15, one permutation entropy and one energy proportion coefficient for each frequency band are sequentially stored in a one-dimensional array, generating 32-dimensional features per channel. The features from the nine channels are concatenated and stitched together to output a high-dimensional mixed feature vector containing 288 floating-point feature elements. For example... Figure 3 The figure shows the distribution of the energy proportion of the 16 sub-bands after the single-channel signal is decomposed into 4 layers of db4 wavelet.
[0030] S2, calculate the maximum information coefficient between each dimension of the training feature set and the connection status label as the relevance, filter out features with a relevance lower than the mean to obtain a candidate feature set, calculate the similarity between each feature in the candidate feature set to form a similarity matrix, and use the median of the elements in the similarity matrix as the nearest neighbor propagation clustering bias parameter to cluster the candidate feature set to obtain feature clusters and cluster affiliation information.
[0031] The feature vectors of each column in the training feature set data matrix and the corresponding binary classification real label data of whether the inverter power supply is on or off are extracted. The MINE algorithm is used to calculate the maximum information coefficient between each feature column and the label data column, which is taken as the correlation between the two. The arithmetic mean of the maximum information coefficients of all dimensions is calculated. Using the Boolean index filtering mechanism of the NumPy library, feature columns with maximum information coefficients lower than the average are deleted, and the remaining highly correlated feature columns are recombined column-wise to obtain a two-dimensional candidate feature set. For any pair of feature columns in the candidate feature set, the Pearsonr function of the statistical analysis module in the SciPy library is called to calculate the Pearson correlation coefficient and take its absolute value as the numerical similarity between the two features. The similarity of all feature pairs is filled into the matrix according to the original feature column sorting index to obtain a symmetric two-dimensional feature similarity matrix. The statistical median of the off-diagonal elements in the similarity matrix is calculated, preferably the median of the upper triangular off-diagonal elements is taken as the preference parameter for nearest neighbor propagation clustering, and the similarity input mode of the clustering algorithm is set to the pre-calculated similarity matrix mode. The `fit` method of the clustering algorithm object is called, passing in a two-dimensional feature similarity matrix, to perform unsupervised clustering on the candidate feature set. Based on the `labels` attribute array output by the algorithm, the numerical index of the cluster center assigned to each feature is obtained, resulting in multiple feature clusters automatically divided for all features and the explicit cluster affiliation information for each feature.
[0032] In an optional embodiment, calculating the maximum information coefficient between each dimension of the training feature set and the connection status label as the relevance includes: Obtain the numerical distribution of features in each dimension of the training feature set and the category distribution of the preset connection status labels; The number of grids in the label dimension is fixed according to the number of categories of the connected status label. The grid division scale is traversed along the feature dimension. At each grid division scale, the joint probability between each dimension feature and the connected status label and their respective marginal probabilities are calculated. The mutual information value is calculated based on the joint probability and marginal probability under the current grid partitioning scale and then normalized. Find the maximum value of the normalized mutual information value under all grid division scales, and use it as the maximum information coefficient between the corresponding dimensional feature and the on / off state label.
[0033] For a training set containing 1000 samples, the input data is a 1000×288 dimensional feature matrix, and the output labels are preset to two types of on / off state labels, representing the main circuit on and off states respectively. Based on the number of label categories, the number of grids in the label dimension of the maximum information coefficient calculation algorithm is fixed at 2. According to the maximum grid division limit (the total number of grids is less than 0.6 times the number of samples), the upper limit of the feature dimension grid division scale is calculated to be 31, and the number of feature dimension grid divisions is set to iterate from 2 to 31. Under the specific grid division scale where the feature dimension is divided into 5 intervals, a 5×2 two-dimensional probability grid is constructed. The number of samples whose feature values fall into a specific interval and whose labels correspond to a specific category is counted, divided by 1000 to calculate the joint probability, and the total frequency of rows and columns is counted to obtain the marginal probability. The mutual information value is calculated according to the formula and divided by... Perform normalization. After the traversal loop is complete, compare the normalized mutual information results obtained under each grid division scale, and extract the highest value as the maximum information coefficient between the feature of that dimension and the on / off state label, representing the nonlinear mapping correlation between the two.
[0034] In an optional embodiment, the step of filtering out features below the relevance mean to obtain a candidate feature set includes: Sum the maximum information coefficients of all dimensional features and the connection status labels, and divide by the total number of feature dimensions to obtain the mean of the maximum information coefficients; Set the mean value as the filtering threshold; Compare the maximum information coefficient of each dimension feature with the filtering threshold one by one, and delete the dimension features whose maximum information coefficient is less than the filtering threshold; The dimensional features with a maximum information coefficient not less than the filtering threshold are retained to form the candidate feature set.
[0035] Obtain the maximum information coefficient value corresponding to each of the 288-dimensional mixed features calculated previously. Sum these 288 values in the [0,1] interval, divide the sum by the total number of dimensions (288), and calculate the global relevance mean of the features in the training set. For example, if the mean is 0.325, set this mean as the threshold for feature selection. During the selection phase, compare based on this threshold: extract the value of the first-dimensional feature; if the value is 0.18 (lower than 0.325), the feature is considered weakly relevant and is removed; if the value of the second-dimensional feature is 0.45 (greater than or equal to 0.325), the feature is added to the retention list. Following this logic, iterate through all 288 dimensions of data, reorganizing the retained features (such as the 125-dimensional feature) according to their original dimensional order to construct a candidate feature set.
[0036] S3. The sequential forward floating selection algorithm is used to search the candidate feature set. A combined evaluation function is set, which is weighted by the gradient boosting decision tree cross-entropy loss and the redundancy penalty term. The redundancy penalty term is used to penalize the candidate features in the same cluster as the selected features to determine the optimal feature subset. The gradient boosting decision tree is trained based on the optimal feature subset and the on / off status label to obtain the identification model. The optimal feature subset of the real-time acquired signal is extracted and input into the identification model, and the identification result is output.
[0037] An empty optimal feature subset is initialized in memory. Following the logical rules of the sequential forward floating selection algorithm, each feature in the candidate feature set is added forward and then eliminated backward for verification. When evaluating the currently constructed candidate feature subset, a combined evaluation function is explicitly defined in the code. The initial instance of the gradient boosting decision tree is constructed using the GradientBoostingClassifier class. When calling the cross-validation score calculation function, the scoring method is set to negative log loss (scoring='neg_log_loss'), and the average of the inverse scores is used as the cross-entropy loss term of the gradient boosting decision tree on the current candidate feature subset. All features contained in the current candidate feature subset are traversed. Using the cluster affiliation information obtained in the previous step, the number of features belonging to the same feature cluster in the current subset is counted. If multiple features belong to the same feature cluster, a redundancy penalty term is calculated. The value of this penalty term is equal to the sum of the total number of features within the same cluster minus 1. The cross-entropy loss term is added to the product of the set penalty weight coefficient and the redundancy penalty term to obtain the single-output score returned by the combined evaluation function.
[0038] The sequential forward floating selection algorithm uses minimizing the output score of the evaluation function as its global optimization objective. It continuously performs heuristic iterative searches until the evaluation score no longer decreases after five consecutive iterations or reaches the preset upper limit for the number of retained features. At this point, the algorithm terminates, retaining the feature combination that results in the lowest historical score for the combined evaluation function as the optimal feature subset. The specific column data corresponding to the optimal feature subset in the original dataset is extracted, along with the on / off state binary classification label. The `fit` method of the gradient boosting decision tree algorithm object is then called to complete the full-data training process of the non-contact identification model. In the practical application inference stage, after the sensor array acquires real-time unknown spatial electromagnetic signals, the same wavelet packet decomposition and multi-dimensional hybrid feature vector extraction processing flow as in the training stage is executed through the underlying code. Using the feature column index stored within the optimal feature subset, the corresponding real-time optimal feature subset data is selected and input as an independent variable into the solidified identification model. The `predict` function of the identification model is called to calculate and output the digital prediction identification result: whether the current inverter power supply main circuit is in the on or off state.
[0039] In an optional embodiment, the step of extracting the optimal feature subset of the real-time acquired signal and inputting it into the identification model to output the identification result includes: Acquire the timing electromagnetic radiation signal of the main circuit when the inverter is operating in real time, and calculate the high-dimensional hybrid feature vector of the real-time acquired signal; Based on the feature index corresponding to the optimal feature subset, extract the corresponding real-time feature subset from the high-dimensional hybrid feature vector; The real-time feature subset is input into the trained gradient boosting decision tree model for tree-by-tree traversal and ensemble prediction calculation, and the corresponding classification result label is output as the on / off state identification result of the main loop.
[0040] Under the load-bearing operation of the inverter power supply, data detection is triggered to acquire real-time multi-channel transient electromagnetic waveforms. A method involving four layers of wavelet packet reconstruction and entropy energy calculation is used to generate a high-dimensional hybrid feature vector of dimension 288 within 50 milliseconds. Based on the optimal feature subset index mapping table derived from the forward floating selection algorithm, which contains 18 optimal feature dimension indices (e.g., 12th, 45th, 88th, etc.), these 18 features are extracted from the real-time 288-dimensional feature vector to construct a 1×18 real-time feature subset vector. This subset vector is then input into a gradient boosting decision tree identification model. The model, based on the node splitting threshold rules of its internally integrated 100 decision trees, performs Boolean logic judgments on the input vector tree by tree. By traversing each tree, the residual prediction values of each leaf node are accumulated and fused, outputting a binary classification probability of the main circuit belonging to the on or off state. The class with the higher probability is taken as the current identification result of the main circuit.
[0041] The experimental setup consisted of a dataset of 1500 samples, with 1000 used as the training set and 500 as the test set. The samples covered both on-state and off-state states of the inverter main circuit. The hardware sampling environment was configured with a 50MHz sampling rate and a fixed 10μs time window, and the platform utilized a high-performance computing node with 16GB of memory. The model parameters were uniformly configured as a gradient boosting decision tree classifier integrating 100 weak learners, with a learning rate of 0.1. The experiment was divided into four groups: a basic group with 288-dimensional mixed features as input; a control group (using only the mean of the maximum information coefficient for initial screening); a control group (removing permutation entropy and utilizing only energy features with a complete screening algorithm); and a complete scheme group employing all multi-stage feature dimensionality reduction and mixed feature construction logic.
[0042] The specific data and results of the experiment are as follows: The basic group retained all 288-dimensional features for training, with a training time of 15.4 seconds, a test set state identification accuracy of 87.6%, and a single model inference time of 45 milliseconds. Control group 1, after filtering with the mean of the maximum information coefficient, retained 125-dimensional features, reducing training time to 8.2 seconds, increasing test set accuracy to 90.2%, and reducing single inference time to 28 milliseconds. Control group 2 reduced the input features to 12 dimensions, with a training time of only 2.5 seconds and an inference time of 15 milliseconds, but the identification accuracy dropped to 85.4%. The complete solution group, after initial screening and forward floating selection, extracted an optimal 18-dimensional feature subset, with a model training time of 3.1 seconds, a test set identification accuracy of 96.8%, and a stable single inference time of 18 milliseconds. The comparison of identification performance and time for each group is illustrated below. Figure 4 As shown.
[0043] The improvement analysis shows that the step-by-step dimensionality reduction strategy of the complete solution optimizes the computational efficiency of the model. Compared with the unselected basic group, the feature dimension is reduced by 93%, and the inference time is reduced by 60%, ensuring the real-time performance of electromagnetic radiation signal identification. Comparison with control group two shows that the addition of permutation entropy, a nonlinear temporal feature, compensates for the lack of pure energy features in expressing transient high-frequency microstructures, playing a core supporting role in improving accuracy. The overall solution, through the fusion of spatiotemporal topological matrix construction, wavelet packet decomposition hybrid feature extraction, and optimal subset optimization, improves the accuracy of state identification by 9.2 percentage points without increasing hardware overhead, achieving improvements in model generalization ability and response speed.
[0044] In the second embodiment, the present invention also proposes a non-contact feature identification system for the main circuit connection state of an inverter power supply, comprising the following modules: The acquisition module is used to acquire multiple sets of time-series electromagnetic radiation signals of key components in the main circuit collected by the sensor array and construct a multi-channel spatiotemporal signal matrix. Wavelet packet decomposition is performed on each channel signal of the multi-channel spatiotemporal signal matrix to obtain the reconstruction coefficient of each frequency band. The entropy of each frequency band is calculated and concatenated with the energy proportion coefficient to form a high-dimensional hybrid feature vector. The high-dimensional hybrid feature vectors of multiple sets of signals form a training feature set. The calculation module is used to calculate the maximum information coefficient between each dimension of the training feature set and the connection status label as the relevance, filter out features with a relevance lower than the mean to obtain a candidate feature set, calculate the similarity between each feature in the candidate feature set to form a similarity matrix, and use the median of the elements in the similarity matrix as the nearest neighbor propagation clustering bias parameter to cluster the candidate feature set to obtain feature clusters and cluster affiliation information. The output module is used to search the candidate feature set using a sequential forward floating selection algorithm, and sets a combined evaluation function weighted by gradient boosting decision tree cross-entropy loss and redundancy penalty term. The redundancy penalty term is used to penalize candidate features in the same cluster as the selected features to determine the optimal feature subset. Based on the optimal feature subset and the on / off status label, a gradient boosting decision tree is trained to obtain the identification model. The optimal feature subset of the real-time acquired signal is extracted and input into the identification model, and the identification result is output.
[0045] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0046] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A non-contact feature identification method for the on-state of the main circuit of an inverter power supply, characterized in that, include: The system acquires multiple sets of time-series electromagnetic radiation signals from key components of the main circuit collected by the sensor array and constructs a multi-channel spatiotemporal signal matrix. Wavelet packet decomposition is performed on each channel signal of the multi-channel spatiotemporal signal matrix to obtain the reconstruction coefficient of each frequency band. The entropy of each frequency band is calculated and concatenated with the energy proportion coefficient to form a high-dimensional hybrid feature vector. The high-dimensional hybrid feature vectors of multiple sets of signals form a training feature set. The maximum information coefficient between each dimension of the training feature set and the connection status label is calculated as the relevance. Features with relevance values lower than the mean are filtered out to obtain a candidate feature set. The similarity between each feature in the candidate feature set is calculated to form a similarity matrix. The median of the elements in the similarity matrix is used as the nearest neighbor propagation clustering bias parameter to cluster the candidate feature set to obtain feature clusters and cluster affiliation information. The sequential forward floating selection algorithm is used to search the candidate feature set. A combined evaluation function is set, which is weighted by the gradient boosting decision tree cross-entropy loss and the redundancy penalty term. The redundancy penalty term is used to penalize the candidate features in the same cluster as the selected features to determine the optimal feature subset. The gradient boosting decision tree is trained based on the optimal feature subset and the on / off status label to obtain the identification model. The optimal feature subset of the real-time acquired signal is extracted and input into the identification model, and the identification result is output.
2. The method according to claim 1, characterized in that, The process of acquiring multiple sets of timing electromagnetic radiation signals from key components of the main circuit collected by the sensor array and constructing a multi-channel spatiotemporal signal matrix includes: For each group of signals, the rising edge of the insulated gate bipolar transistor turn-on control signal sent by the main controller of the inverter power supply system to the drive board is used as the synchronization trigger point. In response to the synchronous trigger point, the non-contact electromagnetic sensor array deployed above the key components of the main circuit of the inverter power supply is controlled to perform synchronous high-frequency sampling to obtain the timing electromagnetic radiation signal within a fixed time window. Each sensor in the sensor array collects a one-dimensional time-series signal as an independent channel. The time-series data of each channel are arranged according to the spatial positional relationship of the sensor array to construct a multi-channel spatiotemporal signal matrix containing spatial topology dimension and time sequence dimension.
3. The method according to claim 1, characterized in that, The step of performing wavelet packet decomposition on each channel signal of the multi-channel spatiotemporal signal matrix to obtain the reconstruction coefficients of each frequency band includes: Select wavelet basis functions and perform N-level wavelet packet decomposition on the time-series signals of each channel in the multi-channel spatiotemporal signal matrix to obtain the node coefficients of 2^N frequency bands in the Nth level. The data after retaining the node coefficients of a specific frequency band and setting the node coefficients of the remaining frequency bands to zero is subjected to inverse wavelet packet transform to obtain a reconstructed signal with the same length as the original time series signal. The reconstructed signals of each frequency band are extracted and used as the reconstruction coefficients of the corresponding frequency band.
4. The method according to claim 3, characterized in that, The calculation of the entropy of each frequency band and its concatenation with the energy proportion coefficient to form a high-dimensional hybrid feature vector includes: Set the embedding dimension and delay time for phase space reconstruction, and perform phase space reconstruction on the reconstruction coefficients of each frequency band; Arrange the elements in each reconstructed state vector in ascending order, obtain the index of the sorting pattern, count the probability of each sorting pattern and calculate the Shannon entropy of the probability distribution to obtain the permutation entropy. Calculate the sum of squares of the reconstruction coefficients for each frequency band, divide the sum of squares of the current frequency band by the sum of the sums of squares of the reconstruction coefficients of all frequency bands under this channel, and obtain the energy proportion coefficient of the corresponding frequency band in this channel; The permutation entropy and energy proportion coefficient of each frequency band in each channel are concatenated in frequency band order to construct a high-dimensional hybrid feature vector.
5. The method according to claim 1 or 2, characterized in that, The calculation of the maximum information coefficient between each dimension of the training feature set and the connection status label as the relevance includes: Obtain the numerical distribution of features in each dimension of the training feature set and the category distribution of the preset connection status labels; The number of grids in the label dimension is fixed according to the number of categories of the connected status label. The grid division scale is traversed along the feature dimension. At each grid division scale, the joint probability between each dimension feature and the connected status label and their respective marginal probabilities are calculated. The mutual information value is calculated based on the joint probability and marginal probability under the current grid partitioning scale and then normalized. Find the maximum value of the normalized mutual information value under all grid division scales, and use it as the maximum information coefficient between the corresponding dimensional feature and the on / off state label.
6. The method according to claim 1, characterized in that, The process of filtering out features below the relevance mean yields a candidate feature set, including: Sum the maximum information coefficients of all dimensional features and the connection status labels, and divide by the total number of feature dimensions to obtain the mean of the maximum information coefficients; Set the mean value as the filtering threshold; Compare the maximum information coefficient of each dimension feature with the filtering threshold one by one, and delete the dimension features whose maximum information coefficient is less than the filtering threshold; The dimensional features with a maximum information coefficient not less than the filtering threshold are retained to form the candidate feature set.
7. The method according to claim 1, characterized in that, The optimal feature subset extracted from the real-time acquired signal is input into the identification model, and the identification result is output, including: Acquire the timing electromagnetic radiation signal of the main circuit when the inverter is operating in real time, and calculate the high-dimensional hybrid feature vector of the real-time acquired signal; Based on the feature index corresponding to the optimal feature subset, extract the corresponding real-time feature subset from the high-dimensional hybrid feature vector; The real-time feature subset is input into the trained gradient boosting decision tree model for tree-by-tree traversal and ensemble prediction calculation, and the corresponding classification result label is output as the on / off state identification result of the main loop.
8. A non-contact feature identification system for the on-state of an inverter power supply main circuit, characterized in that, include: The acquisition module is used to acquire multiple sets of time-series electromagnetic radiation signals of key components in the main circuit collected by the sensor array and construct a multi-channel spatiotemporal signal matrix. Wavelet packet decomposition is performed on each channel signal of the multi-channel spatiotemporal signal matrix to obtain the reconstruction coefficient of each frequency band. The entropy of each frequency band is calculated and concatenated with the energy proportion coefficient to form a high-dimensional hybrid feature vector. The high-dimensional hybrid feature vectors of multiple sets of signals form a training feature set. The calculation module is used to calculate the maximum information coefficient between each dimension of the training feature set and the connection status label as the relevance, filter out features with a relevance lower than the mean to obtain a candidate feature set, calculate the similarity between each feature in the candidate feature set to form a similarity matrix, and use the median of the elements in the similarity matrix as the nearest neighbor propagation clustering bias parameter to cluster the candidate feature set to obtain feature clusters and cluster affiliation information. The output module is used to search the candidate feature set using a sequential forward floating selection algorithm, and sets a combined evaluation function weighted by gradient boosting decision tree cross-entropy loss and redundancy penalty term. The redundancy penalty term is used to penalize candidate features in the same cluster as the selected features to determine the optimal feature subset. Based on the optimal feature subset and the on / off status label, a gradient boosting decision tree is trained to obtain the identification model. The optimal feature subset of the real-time acquired signal is extracted and input into the identification model, and the identification result is output.
9. The system according to claim 8, characterized in that, The process of acquiring multiple sets of timing electromagnetic radiation signals from key components of the main circuit collected by the sensor array and constructing a multi-channel spatiotemporal signal matrix includes: For each group of signals, the rising edge of the insulated gate bipolar transistor turn-on control signal sent by the main controller of the inverter power supply system to the drive board is used as the synchronization trigger point. In response to the synchronous trigger point, the non-contact electromagnetic sensor array deployed above the key components of the main circuit of the inverter power supply is controlled to perform synchronous high-frequency sampling to obtain the timing electromagnetic radiation signal within a fixed time window. Each sensor in the sensor array collects a one-dimensional time-series signal as an independent channel. The time-series data of each channel are arranged according to the spatial positional relationship of the sensor array to construct a multi-channel spatiotemporal signal matrix containing spatial topology dimension and time sequence dimension.
10. The system according to claim 8, characterized in that, The step of performing wavelet packet decomposition on each channel signal of the multi-channel spatiotemporal signal matrix to obtain the reconstruction coefficients of each frequency band includes: Select wavelet basis functions and perform N-level wavelet packet decomposition on the time-series signals of each channel in the multi-channel spatiotemporal signal matrix to obtain the node coefficients of 2^N frequency bands in the Nth level. The data after retaining the node coefficients of a specific frequency band and setting the node coefficients of the remaining frequency bands to zero is subjected to inverse wavelet packet transform to obtain a reconstructed signal with the same length as the original time series signal. The reconstructed signals of each frequency band are extracted and used as the reconstruction coefficients of the corresponding frequency band.