Photovoltaic equipment fault detection method and system based on voiceprint recognition
By collecting environmental sounds in photovoltaic equipment fault detection, establishing the peak characteristics of the spectrum of scene-related equipment, performing multi-subband division and short-time energy envelope calculation, combining Mel frequency cepspectral coefficient extraction and acoustic scene recognition, the problem of environmental noise impact is solved, and high-precision fault detection and adaptive monitoring are achieved.
Patent Information
- Application Number
- CN202510825467.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-08-15
AI Technical Summary
In the prior art, voiceprint recognition is susceptible to environmental noise in the monitoring of photovoltaic equipment failures, making it difficult to distinguish the sound source contribution of different types of components, and fixed feature templates are difficult to adapt to the dynamic operating environment, resulting in a decrease in recognition accuracy.
By collecting the surrounding environmental sounds of the photovoltaic power station, establishing the peak characteristics of the scene-related equipment, performing multi-subband division and short-term energy envelope calculations, combining the Mel frequency cepspectral coefficient extraction and acoustic scene recognition, multi-source feature modeling of the operating sound of the photovoltaic equipment, and improving the recognition accuracy through mutual correlation analysis.
It improves the sensitivity and stability of photovoltaic equipment fault detection, enhances system adaptability, effectively supports the perception and trend warning of potential faults, and avoids interference from misjudgment of a single indicator.
Smart Images

Figure CN120496575A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of voiceprint recognition, and in particular to a photovoltaic equipment fault detection method and system based on voiceprint recognition. Background Art
[0002] The photovoltaic equipment fault detection method based on voiceprint recognition introduces voiceprint recognition technology into the photovoltaic equipment operation status monitoring and fault diagnosis scenarios.
[0003] While existing technologies have applied voiceprint recognition to photovoltaic equipment fault monitoring, this approach, which relies solely on overall sound signals, is susceptible to environmental noise and struggles to distinguish the contributions of different types of components, resulting in a lack of discriminative feature extraction. Furthermore, this approach fails to consider the changing characteristics of voiceprints in different acoustic scenarios, making fixed feature templates difficult to adapt to dynamic operating environments. This results in reduced recognition accuracy during long-term monitoring. Therefore, improvements are needed. Summary of the Invention
[0004] The purpose of the present invention is to solve the shortcomings of the prior art and to propose a photovoltaic equipment fault detection method and system based on voiceprint recognition.
[0005] In order to achieve the above objectives, the present invention adopts the following technical solution: a photovoltaic equipment fault detection method based on voiceprint recognition, comprising the following steps:
[0006] Collect ambient sound around the current photovoltaic power station and establish the spectrum peak characteristics of scene-related equipment;
[0007] Based on the spectral peak characteristics of the scene-associated device, the current photovoltaic device operating sound is divided into multiple sub-bands, and the short-time energy envelope of each sub-band signal is calculated to obtain a multi-sub-band energy envelope curve set. Based on the multi-sub-band energy envelope curve set, the cross-correlation coefficient between the target envelope pairs is calculated to establish a reference envelope cross-correlation coefficient;
[0008] Receive newly collected real-time photovoltaic equipment operating sounds, extract several spectral peak frequency points and corresponding amplitudes from the real-time photovoltaic equipment operating sounds, calculate them with the spectral peak characteristics of the scene-associated equipment, and obtain spectral peak characteristic offsets; perform multi-subband energy envelope extraction on the newly collected real-time photovoltaic equipment operating sounds and calculate the cross-correlation coefficients between the new envelope lines; perform difference calculation on the cross-correlation coefficients with the reference envelope to obtain real-time synchronization differences, and generate a comprehensive equipment sound change index;
[0009] Based on the comprehensive change index of the device sound, the spectral peak feature offset is compared with the spectral peak offset threshold set for the current acoustic scene category, and the real-time synchronization difference is compared with the corresponding synchronization difference threshold to generate a threshold compliance state. Based on the threshold compliance state, a conclusion on the operation status of the photovoltaic device is output.
[0010] Preferably, the step of acquiring the spectrum peak characteristics of the scene-related device is:
[0011] The ambient sound around the current photovoltaic power station is collected. Based on the ambient sound, the spectrum distribution is obtained using Fourier transform. The target frequency band in the spectrum distribution is extracted. The frequency bands are merged using a Mel filter bank. The logarithmic power spectrum is calculated and then a discrete cosine transform is performed to obtain the Mel-frequency cepstrum coefficients of the ambient sound.
[0012] Based on the Mel-frequency cepstral coefficients of the ambient sound, the Mel-frequency cepstral coefficients are input into a pre-trained scene classification model, the characteristic distance between the Mel-frequency cepstral coefficients and the preset scene category is calculated, the scene category corresponding to the ambient sound is determined, and a preset acoustic scene category is generated;
[0013] Based on the preset acoustic scene category, the real-time operating sounds of the inverter and cooling fan components in the photovoltaic equipment are collected, and the spectrum information of the real-time operating sounds is extracted. By searching for the maximum points of the spectrum, multiple spectral peak frequency points in the operating sound spectrum of the inverter and cooling fan components are located respectively, and the corresponding amplitudes of the spectral peak frequency points are extracted. The spectral peak frequency points and the corresponding amplitudes are associated with the preset acoustic scene category to form the spectral peak characteristics of the scene-associated device.
[0014] Preferably, the steps of obtaining the multi-subband energy envelope curve set are:
[0015] Based on the spectral peak characteristics of the scenario-associated device, extract the spectral peak frequency points of the operating sound of the inverter and cooling fan components from the spectral peak characteristics of the scenario-associated device, use the spectral peak frequency points as the center reference frequency, expand the frequency spectrum of the current photovoltaic device operating sound to both sides of the frequency axis with a fixed width, determine the boundary range of each frequency sub-band, and obtain multiple frequency sub-bands of the photovoltaic device operating sound;
[0016] Based on the multiple frequency sub-bands of the photovoltaic device operation sound, each frequency sub-band is band-pass filtered to separate the time domain sound signal within the corresponding sub-band frequency range, and the filtered sub-band time domain sound signal is subjected to a sliding rectangular time window, and the integral of the square of the sound signal amplitude in each time window is gradually calculated to obtain the short-time energy envelope value of each sub-band frequency signal;
[0017] Based on the short-time energy envelope value of each sub-band frequency signal, a multi-sub-band energy envelope curve set corresponding to the current photovoltaic device operation sound is formed in sequence according to the sequence number of the sub-band frequency.
[0018] Preferably, the steps for obtaining the reference envelope correlation coefficient are:
[0019] Based on the multi-subband energy envelope curve set, two adjacent energy envelope curves are combined into a pair of curves in order of subband sequence numbers, the average value of the amplitude difference of each pair of energy envelope curves at the same time point within a fixed time period is calculated, and curve pairs whose average difference does not exceed the amplitude threshold and whose direction change signs are consistent are selected to obtain a set of target envelope curve pairs for cross-correlation calculation;
[0020] Calculating a target envelope pair set based on the cross-correlation, extracting the amplitude sequence of each pair of envelopes, calculating the time-averaged amplitude of each curve, and calculating the normalized cross-correlation coefficient;
[0021] Based on the normalized cross-correlation coefficient of each pair of energy envelope curves, all curve pairs are traversed, all normalized cross-correlation coefficients are extracted, and the largest cross-correlation coefficient is selected as the reference envelope cross-correlation coefficient in the current state.
[0022] Preferably, the step of obtaining the spectrum peak characteristic offset is:
[0023] Receive newly collected real-time photovoltaic equipment operation sound signals, divide the sound signals in a continuous time period into overlapping time windows of fixed length, perform Fourier transform on the signal in each window to extract the spectral power density, select the top three spectral peak frequency points with the highest amplitude, record the frequency values and corresponding amplitudes, and sort them from low to high according to the spectral peak frequency to form a set of spectral peak frequency points of the real-time photovoltaic equipment operation sound;
[0024] According to the spectrum peak frequency point set of the real-time photovoltaic device operating sound, the corresponding frequency points in the spectrum peak characteristics of the scene-associated device are paired one by one according to the frequency sorting, and the spectrum peak characteristic offset of each group of frequency points is calculated;
[0025] Based on the spectral peak feature offset, a spectral peak feature offset sequence is constructed. If there are more than three consecutive groups of spectral peak feature offsets exceeding the threshold, the segment is marked as an abnormal spectrum state, and an offset abnormality warning message is output. At the same time, the mean of all spectral peak feature offsets is calculated and output as the spectral peak feature offset.
[0026] Preferably, the steps for obtaining the device sound comprehensive change index are:
[0027] For newly collected real-time photovoltaic equipment operating sounds, the real-time sound signal is divided into multiple sub-bands in the spectrum according to the peak frequency points in the spectral peak characteristics of the scene-related equipment. The signals of each sub-band are separated through bandpass filtering. The short-time energy of the filtered signal of each sub-band is calculated and the corresponding short-time energy envelope is extracted to form a multi-sub-band energy envelope curve set of the real-time sound;
[0028] Based on the multi-subband energy envelope curve set of the real-time sound, extract the amplitude sequence of each subband in the energy envelope curve, pair them one by one with the energy envelope curve subband sequence corresponding to the reference envelope cross-correlation coefficient according to the subband frequency order, calculate the normalized cross-correlation coefficient of the energy envelope curve pair of the real-time sound pair by pair, calculate the absolute value of the difference with the reference envelope cross-correlation coefficient one by one, and then take the average to obtain the real-time synchronization difference;
[0029] Based on the real-time synchronization difference and in combination with the spectrum peak characteristic offset, the real-time synchronization difference and the spectrum peak characteristic offset are integrated to form a comprehensive device sound change index.
[0030] Preferably, the steps for obtaining the threshold compliance status are:
[0031] Extracting a spectral peak feature offset from the device sound comprehensive change index based on the device sound comprehensive change index, comparing the spectral peak feature offsets with a spectral peak offset threshold value pre-set according to the current acoustic scene category, determining whether the spectral peak feature offsets exceed the spectral peak offset threshold value one by one, and generating a spectral peak offset threshold determination result;
[0032] Based on the device sound comprehensive change index, extract the real-time synchronization difference in the device sound comprehensive change index, compare the real-time synchronization difference with the synchronization difference threshold according to a preset synchronization difference threshold for the current acoustic scene category, and mark it as abnormal if the value of the real-time synchronization difference exceeds the synchronization difference threshold; otherwise, mark it as normal, and generate a synchronization difference threshold determination result;
[0033] Based on the peak shift threshold determination result and the synchronization difference threshold determination result, logical judgment is performed simultaneously. When both are normal, the threshold compliance state is marked as normal. When either one is abnormal, the threshold compliance state is marked as abnormal, and a threshold compliance state is generated.
[0034] Preferably, the steps for obtaining the conclusion of the photovoltaic equipment operating status are:
[0035] Based on the threshold compliance state, determining whether the threshold compliance state is normal; if normal, extracting the spectrum peak feature of the current real-time sound, reading each spectrum peak frequency point and corresponding amplitude one by one, and replacing the corresponding spectrum peak frequency points and amplitudes of the spectrum peak feature of the scene-associated device one by one according to the spectrum peak frequency order of the spectrum peak feature of the scene-associated device, to generate an updated spectrum peak feature of the scene-associated device;
[0036] Based on the updated scene-associated device spectral peak characteristics, the normalized mutual correlation coefficient of the real-time sound corresponding to the current real-time sound is extracted, and the normalized mutual correlation coefficients of the real-time sound are replaced one by one with the values of the corresponding positions in the reference envelope mutual correlation coefficient to form an updated reference envelope mutual correlation coefficient, and a conclusion on the operation status of the photovoltaic equipment is output.
[0037] The present invention also provides a photovoltaic equipment fault detection system, comprising:
[0038] The voiceprint collection module collects the ambient sound around the current photovoltaic power station and establishes the spectrum peak characteristics of scene-related equipment;
[0039] An energy envelope calculation module divides the current photovoltaic equipment operating sound into multiple sub-bands based on the spectral peak characteristics of the scene-associated device, calculates the short-time energy envelope of each sub-band signal, and obtains a multi-sub-band energy envelope curve set; based on the multi-sub-band energy envelope curve set, calculates the mutual correlation coefficient between the target envelope pairs and establishes a reference envelope mutual correlation coefficient;
[0040] The real-time sound analysis module receives newly collected real-time photovoltaic equipment operating sounds, extracts several spectral peak frequency points and corresponding amplitudes from the real-time photovoltaic equipment operating sounds, calculates them with the spectral peak characteristics of the scene-associated equipment, and obtains the spectral peak characteristic offset; extracts multi-subband energy envelopes from the newly collected real-time photovoltaic equipment operating sounds and calculates the cross-correlation coefficients between the new envelope pairs; calculates the difference between the cross-correlation coefficients and the baseline envelope to obtain the real-time synchronization difference, and generates a comprehensive equipment sound change index;
[0041] The fault judgment module compares the spectral peak feature offset with the spectral peak offset threshold set for the current acoustic scene category based on the comprehensive change index of the device sound; compares the real-time synchronization difference with the corresponding synchronization difference threshold to generate a threshold compliance status; and outputs a conclusion on the operation status of the photovoltaic device based on the threshold compliance status.
[0042] Compared with the prior art, the advantages and positive effects of the present invention are:
[0043] This invention incorporates environmental sound classification during the sound acquisition phase, combining Mel-frequency cepstral coefficient extraction with acoustic scene recognition to contextually correlate photovoltaic system operating sounds with surrounding background sounds, providing a more targeted reference framework for subsequent analysis. By synchronously acquiring the operating sounds of the inverter and cooling fan within the device and extracting the spectral peak frequencies and amplitudes, multi-source sound feature modeling is achieved. During signal processing, multi-subband division is performed based on the scene-correlated spectral peak frequencies, aligning each subband with the device's acoustic characteristics. This improves the structural accuracy of short-term energy envelope extraction and ensures more reliable cross-correlation analysis results. After multidimensional parameter extraction, the new sounds collected in real time are differentially integrated using frequency offset and envelope synchronization differences to form a comprehensive sound change index, enhancing the sensitivity and stability of anomaly detection. A separate judgment method using a comparison threshold can separately address frequency offset and energy structure fluctuations, avoiding interference caused by misjudgment of a single indicator. Reference templates are actively updated when the device status stabilizes, enabling adaptive dynamic maintenance of the system, enhancing long-term monitoring accuracy and robustness, and effectively supporting the detection and trend warning of potential photovoltaic system failures. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 It is a schematic diagram of the steps of the present invention. DETAILED DESCRIPTION
[0045] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0046] See also Figure 1 The present invention provides a technical solution, a photovoltaic equipment fault detection method based on voiceprint recognition, comprising the following steps:
[0047] Collect ambient sound around the current photovoltaic power station and establish the spectrum peak characteristics of scene-related equipment;
[0048] Based on the spectral peak characteristics of the scene-related equipment, the current photovoltaic equipment operating sound is divided into multiple sub-bands, and the short-time energy envelope of each sub-band signal is calculated to obtain a multi-sub-band energy envelope curve set. Based on the multi-sub-band energy envelope curve set, the cross-correlation coefficient between the target envelope pairs is calculated to establish the baseline envelope cross-correlation coefficient;
[0049] Receive newly collected real-time photovoltaic equipment operating sounds, extract several spectral peak frequency points and corresponding amplitudes from the real-time photovoltaic equipment operating sounds, calculate the spectral peak characteristics of the scene-related equipment, and obtain the spectral peak characteristic offset. Perform multi-subband energy envelope extraction on the newly collected real-time photovoltaic equipment operating sounds and calculate the cross-correlation coefficient between the new envelope line pairs. Difference calculation is performed on the cross-correlation coefficient with the baseline envelope to obtain the real-time synchronization difference, and generate a comprehensive equipment sound change index;
[0050] Based on the comprehensive change index of equipment sound, the spectral peak feature offset is compared with the spectral peak offset threshold set for the current acoustic scene category, and the real-time synchronization difference is compared with the corresponding synchronization difference threshold to generate a threshold compliance state. Based on the threshold compliance state, the photovoltaic equipment operating status judgment conclusion is output.
[0051] The steps for obtaining the spectral peak characteristics of scene-related devices are as follows:
[0052] The ambient sound around the current photovoltaic power station is collected. Based on the ambient sound, the spectrum distribution is obtained using Fourier transform. The target frequency band in the spectrum distribution is extracted. The frequency bands are merged using a Mel filter bank. The logarithmic power spectrum is calculated and then a discrete cosine transform is performed to obtain the Mel-frequency cepstrum coefficients of the ambient sound.
[0053] Based on the Mel-frequency cepstral coefficients of the ambient sound, the model is input into the pre-trained scene classification model, the characteristic distance between the Mel-frequency cepstral coefficients and the preset scene category is calculated, the scene category corresponding to the ambient sound is determined, and the preset acoustic scene category is generated;
[0054] Based on the preset acoustic scene category, the real-time operating sounds of the inverter and cooling fan components in the photovoltaic equipment are collected, and the spectrum information of the real-time operating sounds is extracted. By searching for the maximum points of the spectrum, multiple spectral peak frequency points in the operating sound spectrum of the inverter and cooling fan components are located respectively, and the corresponding amplitudes of the spectral peak frequency points are extracted. The spectral peak frequency points and the corresponding amplitudes are associated with the preset acoustic scene category to form the spectral peak characteristics of the scene-associated equipment.
[0055] Specifically, the ambient sound around the current photovoltaic power station is collected by deploying at least 5 omnidirectional microphones around the photovoltaic array and at the center, 1.5 meters above the ground, with a spacing of no less than 50 meters between the microphones. The audio sampling rate is set to 44100 Hz and the bit depth is 16 bits. The sound signal is continuously collected and the collected continuous sound signal s(t) is framed. The Hamming window is used as the window function, the window length is set to 2048 sampling points, the frame shift is 512 sampling points, and the short-time Fourier transform is applied to each windowed signal frame to obtain the complex spectrum of each frame and calculate its power spectrum, thereby obtaining the time-varying sound signal. The sound spectrum distribution is quantified. Then, the target frequency band is extracted from the spectrum distribution. Specifically, the spectrum data in the range of 20Hz to 10000Hz is intercepted. This range covers the main environmental sound sources such as wind, rain, birdsong and long-distance traffic. Subsequently, the extracted target frequency band is merged through a Mel filter bank containing 40 triangular filters. The center frequency and bandwidth of each triangular filter are equidistantly distributed on the Mel scale. The energy value output by each filter is logarithmically operated to obtain the logarithmic power spectrum. Finally, the discrete cosine transform is performed on the 40 logarithmic power spectrum values of each audio frame. The calculation formula is: π is a constant, the ratio of pi, which is approximately 3.14159, and i is the index of the Mel frequency cepstral coefficient (i.e., the DCT coefficient to be calculated), where C i is the ith Mel frequency cepstral coefficient, N is the number of Mel filters, here 40, E j is the energy value output by the jth Mel filter. The 2nd to 13th coefficients in the transformation result, a total of 12 coefficients, are selected as the feature vector of the current frame. Finally, the feature vectors of all frames are combined in chronological order to obtain the Mel-frequency cepstrum coefficients representing the ambient sound of this segment.
[0056] Based on the Mel-frequency cepstral coefficients of environmental sounds, a Mel-frequency cepstral coefficient matrix consisting of 40 consecutive audio frames with a dimension of 12x40 is input as a sample to a pre-trained convolutional neural network scene classification model. The model has been trained on an acoustic scene dataset containing five preset scene categories: "sunny breeze", "rainy day", "strong wind", "nearby construction", and "quiet at night" through supervised learning. The model structure includes: an input layer that receives a 12x40 matrix, two sequentially connected convolutional layers, the first layer contains 32 3x3 convolution kernels, and the second layer contains 64 3x3 convolution kernels. Each convolution layer is connected to a rectified linear unit (ReLU) activation function and a 2x2 maximum pooling layer. The pooling layer is followed by a flattening layer for converting two-dimensional features into a single matrix. The feature map is converted into a one-dimensional vector, and then connected to a fully connected layer containing 256 neurons (also using ReLU activation), and finally an output layer containing 5 neurons, using the Softmax activation function. In the inference stage, the input Mel-frequency cepstral coefficient matrix is forward propagated through the model, and the output layer generates a vector containing 5 probability values. Each probability value corresponds to the confidence of a preset scene category. The model selects the category with the highest confidence as the classification result. For example, if the output vector is (0.85, 0.05, 0.05, 0.03, 0.02), it means that the probability that the current sound is judged to be the first category of scene "sunny breeze" is 85%, which is much higher than other categories. Therefore, the scene category corresponding to the current environmental sound is determined to be "sunny breeze", and the preset acoustic scene category is generated.
[0057] Based on the preset acoustic scene category, such as "sunny breeze", the system will trigger a targeted device sound collection, using a directional microphone to aim at the heat dissipation grille of the photovoltaic inverter shell and the air outlet of the cooling fan, respectively, to collect real-time operating sounds for 30 seconds. The collected sound signals of the inverter and cooling fan are processed by short-time Fourier transform respectively, with a window length of 4096 points and a frame shift of 1024 points to obtain high-resolution spectrum information. Subsequently, a peak search algorithm is used to analyze the power spectrum of each frame. The algorithm identifies the points with amplitude higher than the two adjacent frequency points on the left and right and the peak prominence exceeding a specific threshold as spectral peaks, where the peak prominence threshold is set to 1.5 times the average energy of the local background noise plus one standard deviation. The background noise energy is the energy at a distance from the current peak. The algorithm is calculated within a frequency window of 50 Hz on each side of the value point. To distinguish different components, the algorithm will search within the preset characteristic frequency band. For example, the high-frequency switching noise of the inverter is mainly concentrated in the range of 8000 Hz to 16000 Hz, while the rotation and aerodynamic noise of the cooling fan are mainly in the range of 200 Hz to 1500 Hz. The algorithm locates the top three spectral peak frequency points in the corresponding frequency band, and records their frequency values and linear amplitudes. The spectral peak frequencies and amplitudes obtained from all frames within 30 seconds are time-averaged to obtain stable values. Finally, the multiple spectral peak frequency points and corresponding amplitudes of the inverter and cooling fan obtained by averaging are bound to the current preset acoustic scene category of "sunny breeze" in the data structure to form the spectral peak characteristics of the scene-associated device.
[0058] The steps for obtaining the multi-subband energy envelope curve set are:
[0059] Based on the spectral peak characteristics of the scene-related equipment, the spectral peak frequency points of the operating sound of the inverter and cooling fan components are extracted from the spectral peak characteristics of the scene-related equipment. The spectral peak frequency points are used as the central reference frequency. In the spectrum of the current photovoltaic equipment operating sound, the boundary range of each frequency sub-band is determined by expanding it to both sides of the frequency axis with a fixed width, thereby obtaining multiple frequency sub-bands of the photovoltaic equipment operating sound;
[0060] Based on the multiple frequency sub-bands of the photovoltaic equipment operating sound, each frequency sub-band is band-pass filtered to separate the time domain sound signal within the corresponding sub-band frequency range. The filtered sub-band time domain sound signal is then subjected to a sliding rectangular time window, and the integral of the squared amplitude of the sound signal within each time window is gradually calculated to obtain the short-time energy envelope value of each sub-band frequency signal.
[0061] Based on the short-time energy envelope value of each sub-band frequency signal, a multi-sub-band energy envelope curve set corresponding to the current photovoltaic equipment operation sound is formed in sequence according to the sequence number of the sub-band frequency.
[0062] Specifically, based on the spectrum peak characteristics of the scene-associated device, the recorded spectrum peak frequency points of the inverter and cooling fan components are first read from the feature data structure. For example, for the cooling fan, the extracted spectrum peak frequency points may be 350Hz, 850Hz and 1200Hz, and for the inverter, they may be 9500Hz, 12500Hz and 15500Hz. Then, each extracted spectrum peak frequency point is used as the center reference frequency, and on the power spectrum of the photovoltaic device operation sound currently collected and completed Fourier transform, it is expanded to both sides around the center reference frequency to delineate the sub-band boundary. The fixed width here is set according to experience and is related to the center frequency. For 1 For low-frequency components below 500Hz, the frequency drift is small, and the fixed width is set to 10% of the center frequency. For example, for a spectral peak of 350Hz, the expansion width is 35Hz, and the sub-band boundary range is 315Hz to 385Hz. For high-frequency components above 1500Hz, the switching frequency is greatly affected by temperature and load, and the fixed width is set to 5% of the center frequency. For example, for a spectral peak of 9500Hz, the expansion width is 475Hz, and the sub-band boundary range is 9025Hz to 9975Hz. By performing this operation for all spectral peak frequency points, a list containing multiple non-overlapping frequency intervals is finally generated, and multiple frequency sub-bands of the operating sound of the photovoltaic equipment are obtained.
[0063] Based on the multiple frequency sub-bands of the photovoltaic equipment operation sound, the system uses a fourth-order Butterworth bandpass filter to filter the original, unprocessed time-domain signal of the current photovoltaic equipment operation sound for each sub-band. The passband cutoff frequency of the filter is strictly set to the upper and lower boundary frequencies of the corresponding sub-band. For example, for the sub-band from 315Hz to 385Hz, a corresponding bandpass filter is designed. This operation filters out the components of the original mixed sound signal that are not within this frequency range, thereby separating the time-domain sound signal containing only the information of this specific sub-band. Next, each sub-band time-domain sound signal obtained after filtering is analyzed using a sliding rectangular time window with a length of 20 milliseconds. The step length of the time window is set to 10 milliseconds, that is, there is a 50% overlap. Within each time window, the amplitude of all sound signal sampling points within the window is squared, and then all square values are added and summed. The calculation formula is: Among them, E i (q) represents the short-term energy of the i-th subband in the q-th time window, x i (n) is the amplitude of the time domain signal after filtering of the i-th subband at the n-th sampling point, L is the number of sampling points contained in each time window, and by gradually moving the time window along the duration of the entire subband signal and repeating this calculation, a series of energy values arranged in time order are obtained, that is, the short-time energy envelope value of each subband frequency signal.
[0064] Based on the short-term energy envelope values of each subband frequency signal, the system first organizes these value sets. Each set represents a trajectory of the energy change over time for a corresponding subband. Then, to facilitate subsequent cross-correlation calculations, each energy envelope trajectory is independently normalized. The specific normalization method is to find the maximum energy value in the energy envelope trajectory and then divide all energy values on the trajectory by this maximum value, so that the value range of each energy envelope trajectory is scaled to between 0 and 1. Next, these normalized energy envelope trajectories are sorted in order from low to high subband center frequency. For example, if the subbands correspond to 350Hz, 850Hz, 9500Hz, and 12500Hz, respectively, then their corresponding normalized energy envelope trajectories are also arranged in this order. Finally, these sorted, normalized energy envelope trajectories are combined into a data matrix. Each row of the matrix represents the energy envelope curve of a subband, and the column represents the time point, forming a multi-subband energy envelope curve set corresponding to the current photovoltaic equipment operating sound.
[0065] The steps to obtain the reference envelope correlation coefficient are:
[0066] Based on the multi-subband energy envelope curve set, two adjacent energy envelope curves are combined into a pair according to the subband sequence number. The average amplitude difference of each pair of energy envelope curves at the same time point within a fixed time period is calculated. The curve pairs whose average difference does not exceed the amplitude threshold and whose direction change signs are consistent are selected to obtain the target envelope curve pair set for cross-correlation calculation;
[0067] Calculate the target envelope pair set based on the cross-correlation, extract the amplitude sequence of each pair of envelopes, calculate the time average amplitude of each curve, and calculate the normalized cross-correlation coefficient at the same time. The calculation formula is:
[0068]
[0069] Among them, R d is the normalized cross-correlation coefficient of the d-th pair of envelopes, E d,j is the amplitude of the first energy envelope curve in the dth pair at the jth time point, F d,j is the amplitude of the second energy envelope curve at the jth time point, is the average amplitude of the first curve of the pair, is the average amplitude of the second curve of the pair, and m is the total number of time points used for calculation;
[0070] Based on the normalized cross-correlation coefficient of each pair of energy envelope curves, all curve pairs are traversed, all normalized cross-correlation coefficients are extracted, and the largest cross-correlation coefficient is selected as the benchmark envelope cross-correlation coefficient in the current state.
[0071] Specifically, based on the multi-subband energy envelope curve set, the system first sorts the energy envelope curves in descending order according to the sub-band center frequency corresponding to each energy envelope curve, and sequentially combines two adjacent energy envelope curves to form a curve pair. For example, if there are 6 energy envelope curves, 5 curve pairs will be formed. Subsequently, the system intercepts the data of the last 10 seconds for each curve pair for analysis, and calculates the average value of the absolute value of the amplitude difference of the two energy envelope curves at each same time point during this time period, and then compares the average difference with a preset amplitude threshold. The amplitude threshold is dynamically calculated based on the data of the equipment running continuously for 30 days in a healthy state. The specific method is: the average difference of all adjacent curve pairs is calculated every day to form a historical data set, and the amplitude threshold is set to the average value of the data set plus 1.5 times the standard deviation. For example For example, if the historical average difference is 0.05 (normalized amplitude) and the standard deviation is 0.01, the amplitude threshold is 0.05+1.50.01=0.065. Curve pairs with an average difference exceeding 0.065 will be eliminated. For curve pairs that pass the amplitude screening, the system further determines the consistency of their directional changes by calculating the amplitude change direction (increase, decrease, or unchanged) of the two curves at adjacent time points (for example, every 10 milliseconds) and counting the proportion of time points with consistent directions. Only when this proportion exceeds a fixed directional change consistency threshold is the curve pair considered valid. The directional change consistency threshold is set to 85% based on experience, that is, the two curves are required to fluctuate in the same direction for more than 85% of the time. Finally, all curve pairs that meet both the amplitude difference and directional change consistency conditions are summarized to obtain the target envelope pair set for cross-correlation calculation.
[0072] formula: The benefit of the formula is that it quantifies the degree of linear correlation between the two energy envelope curves by calculating the Pearson correlation coefficient. This coefficient is a linear measure of the amplitude and is not affected by the absolute energy of the two curves. It only focuses on the synchronization of their waveform changes over time. In equipment fault detection, the physical coupling relationship between components (such as vibration transmission) will produce highly synchronized sound energy fluctuations, forming a stable high-correlation feature. When the equipment has an early fault, this coupling relationship often changes, resulting in a decrease in synchronization. Therefore, by monitoring the changes in the normalized mutual correlation coefficient, changes in the dynamic characteristics of the equipment caused by potential faults can be sensitively captured.
[0073] d is the serial number of the energy envelope curve pair currently being processed. It is a positive integer that uniquely identifies a specific curve pair selected from the cross-correlation calculation target envelope curve pair set. This parameter is obtained by traversing the cross-correlation calculation target envelope curve pair set. For example, if there are three qualified curve pairs in the set, d will take the values of 1, 2, and 3 respectively. In this calculation, the first curve pair is processed, so the value of d is 1.
[0074] m is the total number of time points used for the calculation. This value is determined by the length of the analysis time window and the calculation frame rate of the short-term energy envelope. The system is configured with a 10-second analysis time window and a calculation frame rate of 100 points per second for the energy envelope (i.e., one energy value every 10 milliseconds). Therefore, the calculated value of m is 10 seconds × 100 points / second = 1000 points. This setting is based on long-term observations of the sound signals of photovoltaic equipment operation. The 10-second data length is sufficient to cover typical dynamic processes such as fan startup and shutdown and inverter power adjustment.
[0075] E d,j is the normalized amplitude of the first energy envelope curve in the dth pair of curves at the jth time point. This data comes directly from the multi-subband energy envelope curve set generated in the previous step. The amplitude of the curve has been normalized to the interval [0, 1]. Here, E 1,j Represents the amplitude sequence of the first curve extracted from the first qualified curve pair. For example, the first four points of the sequence are taken for example calculation, and their values are [0.5, 0.6, 0.4, 0.7].
[0076] F d,j is the normalized amplitude of the second energy envelope curve in the dth pair of curves at the jth time point, and is related to E d,j Similarly, this data is directly extracted from the multi-subband energy envelope curve set and represents the other curve in the same curve pair, where F 1,j Represents the amplitude sequence of the second curve extracted from the first qualified curve pair. For example, the first four points corresponding to it are taken for example calculation, and their values are [0.6, 0.7, 0.5, 0.8].
[0077] is the average amplitude of the first curve in the dth pair at all m time points, which is obtained by d,j The average of all data points is obtained, which reflects the average energy level of the subband during the analysis period. For the above example sequence E 1,j =[0.5,0.6,0.4,0.7], its average amplitude The calculation process is (0.5+0.6+0.4+0.7) / 4=0.55.
[0078] is the average amplitude of the second curve in the dth pair at all m time points, which is calculated in the same way as Exactly the same, it is sequence F d,j The average value reflects the average energy level of the sub-band corresponding to the second curve. For the above example sequence F 1,j =
[0079] [0.6, 0.7, 0.5, 0.8], its average amplitude The calculation process is (0.6+0.7+0.5+0.8) / 4=0.65.
[0080] Calculation process:
[0081] Calculate the target envelope pair set based on the cross-correlation, extract the amplitude sequence of the first pair of curves (d=1) for calculation, and take the data of m=4 time points:
[0082] The amplitude sequence is: E 1,j =[0.5,0.6,0.4,0.7], F 1,j =[0.6,0.7,0.5,0.8].
[0083] First calculate the average amplitude of the two curves:
[0084]
[0085] Next, calculate the numerator of the formula, which is the covariance of the two curves:
[0086]
[0087]
[0088] Then, calculate the denominator of the formula, which is the product of the standard deviations of the two curves:
[0089]
[0090] Denominator = 0.2236 × 0.2236 = 0.05.
[0091] Finally, calculate the normalized cross-correlation coefficient R1:
[0092]
[0093] The results show that at the four time points in the example, the two energy envelope curves show a completely positive correlation with a value of 1.0. A high positive value (close to 1.0) indicates that the two energy envelope curves have very strong synchronous fluctuations, which is an expected feature in normally operating equipment, while a low value (close to 0) or a negative value may indicate an abnormality in the physical coupling relationship within the device.
[0094] Based on the normalized cross-correlation coefficient of each pair of energy envelope curves, the system processes all normalized cross-correlation values calculated in the previous step (for example, if there are three qualified curve pairs, a list of three coefficient values is obtained, such as [0.89, 0.95, 0.91]). The system traverses this coefficient list, comparing them one by one to find and determine the maximum value. In this example, 0.89 and 0.95 are first compared to determine 0.95 as the current maximum value. Then 0.95 is compared with the next value in the list, 0.91, to finally confirm that 0.95 is the maximum value in the entire list. The maximum cross-correlation coefficient is selected because it represents the strongest and most stable coupling relationship between the acoustic characteristics of the physical components under the current operating state of the device. This strongest coupling relationship generally best reflects the operating state of the core components of the device and is most sensitive to minor abnormal disturbances. Therefore, using it as a baseline can provide the most reliable and representative health status indicator. This selected maximum value (0.95 in this example) is finally determined and output as the baseline envelope cross-correlation coefficient for the current state.
[0095] The steps to obtain the peak characteristic offset are:
[0096] Receive newly collected real-time photovoltaic equipment operation sound signals, divide the sound signals in a continuous time period into overlapping time windows of fixed length, perform Fourier transform on the signal in each window to extract the spectral power density, select the top three spectral peak frequency points with the highest amplitude, record the frequency values and corresponding amplitudes, and sort them from low to high according to the spectral peak frequency to form a set of spectral peak frequency points of the real-time photovoltaic equipment operation sound;
[0097] Based on the spectrum peak frequency point set of the real-time photovoltaic equipment operating sound, the corresponding frequency points in the spectrum peak characteristics of the scene-related equipment are sorted by frequency and paired one by one. The spectrum peak characteristic offset of each group of frequency points is calculated. The calculation formula is:
[0098]
[0099] Among them, D c is the characteristic offset of the c-th group of spectral peaks (unit: Hz), f c,k is the kth spectrum peak frequency point (Hz) of the cth group in the real-time photovoltaic equipment operation sound, g c,k is the kth spectral peak frequency point (Hz) of the cth group in the spectral peak characteristics of the scene-associated device, a c,k with b c,k are the amplitudes (linear amplitudes) of the spectrum peak frequency point in real time and in scene equipment, respectively. norm is the amplitude normalization constant, which is the average of all valid spectrum peak amplitudes of the same type of equipment in the past 30 days;
[0100] Based on the peak feature offset, a peak feature offset sequence is constructed. If there are more than three consecutive groups of peak feature offsets exceeding the threshold, the segment is marked as an abnormal spectrum state, and an offset abnormality warning message is output. At the same time, the mean of all peak feature offsets is calculated and output as the peak feature offset.
[0101] Specifically, the system receives newly collected real-time photovoltaic equipment operation sound signals, processes the continuously collected sound signals with every 5 seconds as an analysis unit, and further divides the 5-second sound signal into multiple time windows with a fixed duration of 50 milliseconds and an overlap rate of 50%, that is, each time window contains 25 milliseconds of new data. For the signal in each time window, the Hamming window function is applied and then the fast Fourier transform is performed to calculate the spectral power density of the window. Then, the system searches for the spectral peak in a preset frequency range (for example, 200Hz to 16000Hz), confirms it by finding the local maximum point and applying a dynamic threshold For effective spectrum peaks, the dynamic threshold is set to 2.5 times the overall mean of the current spectrum power density. The system will sort all the screened spectrum peaks from high to low according to the amplitude, and select the top three spectrum peaks with the highest amplitude as the feature points of the current time window, record their precise frequency values and corresponding linear amplitudes, and average all the spectrum peak frequency points and amplitudes extracted in 100 consecutive time windows (corresponding to 5 seconds) to obtain a set of stable and reliable spectrum peak data. Finally, the three averaged spectrum peak frequency points are sorted from low to high according to the frequency value to form a spectrum peak frequency point set of the real-time photovoltaic equipment operation sound in the current 5-second time period.
[0102] formula: The benefit of this formula is that it constructs a comprehensive spectrum peak characteristic offset indicator. This indicator not only takes into account the drift of the real-time spectrum peak frequency relative to the reference spectrum peak frequency, but also introduces amplitude as a weight, so that the frequency offset of spectrum peaks with higher amplitudes and more significant amplitudes contributes more to the final result. This is because changes in the dominant spectrum peak with concentrated energy often better reflect substantial changes in the operating status of the equipment. By multiplying the square of the frequency offset by the amplitude weight and then taking the square root, the two physical quantities of different dimensions, frequency (unit: Hz) and amplitude (linear unit) are effectively integrated into a unified, frequency-dimensional offset indicator.
[0103] c is the serial number of the device component group, which is used to distinguish different sound sources. For example, the cooling fan and inverter in a photovoltaic device. The system defines the cooling fan-related spectral peaks as group 1 (c=1) and the inverter-related spectral peaks as group 2 (c=2). This parameter is determined by consulting a pre-established scenario-related device spectral peak feature library, which has grouped and identified the spectral peak features of different components. In this calculation, the spectral peak of the cooling fan component is processed, so the value of c is 1.
[0104] k is the peak number within each group, with values of 1, 2, and 3, representing the three main peaks sorted from low to high frequency. This parameter is determined in the previous step by selecting and sorting the top three peaks with the highest amplitude. For example, for a cooling fan, k = 1 may correspond to its fundamental frequency, and k = 2 or 3 to its harmonics or the main peak of aerodynamic noise. In this calculation, these three peaks are processed sequentially, that is, k is 1, 2, and 3, respectively.
[0105] f c,k is the frequency value of the kth spectral peak of the cth group in the real-time photovoltaic equipment operation sound, in Hz. This value is obtained by performing spectrum analysis on the real-time sound signal and extracting the set of spectral peak frequency points. For example, for the cooling fan (c = 1), the three spectral peak frequencies measured in this real-time measurement are f 1,1 =355Hz, f 1,2 =852Hz, f 1,3 =1210Hz.
[0106] g c,k The reference frequency value of the kth spectral peak in the cth group of the scene-associated device spectral peak feature, in Hz, is read from the scene-associated device spectral peak feature established in the previous step and matching the current acoustic scene. It represents the standard spectral peak position of the device in a healthy state. For example, for a cooling fan (c = 1), its reference spectral peak frequency is g 1,1 =350Hz, g 1,2 =850Hz,g 1,3 =1200Hz.
[0107] a c,k is the linear amplitude of the kth spectrum peak of the cth group in the real-time sound, which is the same as f c,k The synchronous spectrum is obtained from the real-time spectrum analysis, which reflects the energy intensity of the spectrum peak. For example, the linear amplitudes of the three spectrum peaks measured in this measurement (after uniform scale transformation) are a 1,1 =120,a 1,2 =95,a 1,3 =70.
[0108] b c,k is the baseline linear amplitude of the kth spectral peak in the cth group of the scene-associated device spectral peak characteristics. This value is read from the scene-associated device spectral peak characteristics and represents the standard spectral peak amplitude in a healthy state. For example, the baseline spectral peak amplitude of a cooling fan is b 1,1 =110, b 1,2 =100, b 1,3 =80.
[0109] A normis the amplitude normalization constant, which is obtained by counting the linear amplitudes of all effective spectrum peaks recorded by all photovoltaic devices of the same type under normal operation in the past 30 days and calculating their average value. The purpose of this is to eliminate the amplitude baseline differences caused by gain settings or environmental influences of different devices and at different times, and to provide a stable reference scale. For example, by calculating the data of the past 30 days in the database, A norm =100.
[0110] Calculation process:
[0111] Based on the spectrum peak frequency point set of the real-time photovoltaic equipment operating sound, paired with the spectrum peak characteristics of the scene-related equipment, the spectrum peak characteristic offset D1 of the cooling fan component (c=1) is calculated:
[0112] Substitute parameter values:
[0113] f 1,1 =355Hz,g 1,1 =350Hz,a 1,1 =120,b 1,1 =110;
[0114] f 1,2 =852Hz,g 1,2 =850Hz,a 1,2 =95,b 1,2 =100;
[0115] f 1,3 =1210Hz,g 1,3 =1200Hz,a 1,3 =70,b 1,3 =80;
[0116] A norm =100;
[0117] The calculation process is as follows:
[0118] First, calculate the weighted frequency offset square term of each spectral peak (k=1,2,3):
[0119] For k=1:
[0120] For k=2:
[0121] For k=3:
[0122] Then add the terms together and take the square root:
[0123]
[0124] The results show that the comprehensive spectrum peak characteristic offset of the cooling fan component is 14.67Hz, which comprehensively reflects the comprehensive changes in the frequency and amplitude of its three main spectrum peaks relative to the healthy baseline.
[0125] Based on the spectral peak feature offset, the system stores the spectral peak feature offset values (such as 14.67Hz of the cooling fan and a certain value of the inverter) calculated by each analysis unit (every 5 seconds) in a queue in chronological order to form a spectral peak feature offset sequence. The system will continuously monitor this sequence and check whether there are three or more consecutive groups of spectral peak feature offsets that exceed the preset offset threshold of their corresponding components. The offset threshold is set independently for different components (such as fans and inverters) and different acoustic scenarios (such as "sunny breeze" and "strong wind"). The setting method is: collect a large amount of spectral peak feature offset data when the component is running healthily in the scenario to form a data set, and then calculate the data set. The average value and standard deviation are used to calculate the threshold value, which is the average value plus three times the standard deviation. For example, if the cooling fan has a historical offset average of 5 Hz and a standard deviation of 2 Hz in the "sunny day breeze" scenario, then its offset threshold is 5 + 32 = 11 Hz. If three or more consecutive values greater than 11 Hz appear in the sequence, the system will immediately mark the time period as an abnormal spectrum state and generate an offset abnormality warning message containing a timestamp, component name, and details of the threshold being exceeded. At the same time, the system will also calculate the average value of all spectral peak feature offsets (including fans and inverters) in the current analysis period (for example, the past 1 minute) and output this average value as the final spectral peak feature offset value for subsequent comprehensive indicator calculations.
[0126] The steps to obtain the comprehensive change index of device sound are as follows:
[0127] For newly collected real-time photovoltaic equipment operating sounds, the real-time sound signal is divided into multiple sub-bands in the spectrum according to the peak frequency points in the spectral peak characteristics of the scene-related equipment. The signals of each sub-band are separated through bandpass filtering. The short-time energy of the filtered signal of each sub-band is calculated and the corresponding short-time energy envelope is extracted to form a multi-sub-band energy envelope curve set of the real-time sound;
[0128] Based on the multi-subband energy envelope curve set of real-time sound, the amplitude sequence of each sub-band in the energy envelope curve is extracted, and the energy envelope curve sub-band sequence corresponding to the reference envelope cross-correlation coefficient is paired one by one according to the sub-band frequency order. The normalized cross-correlation coefficient of the energy envelope curve pair of real-time sound is calculated pair by pair, and the absolute value of the difference with the reference envelope cross-correlation coefficient is calculated one by one and then averaged to obtain the real-time synchronization difference;
[0129] Based on the real-time synchronization difference and the spectrum peak characteristic offset, the real-time synchronization difference and the spectrum peak characteristic offset are integrated to form a comprehensive device sound change index.
[0130] Specifically, for the newly collected real-time photovoltaic equipment operation sound, the system first calls the scene-associated device spectral peak characteristics that match the current acoustic scene, and obtains the central frequency points of all spectral peaks from them. Then, the system completely reproduces the sub-band division process when establishing the benchmark, that is, based on these spectral peak frequency points, the same fixed-width expansion rule is used to divide the spectrum of the real-time sound signal into a series of corresponding frequency sub-bands. For example, if the benchmark spectral peak frequency points are 350Hz and 9500Hz, the real-time sound spectrum will also be divided into frequency bands centered on these two frequencies and with widths of 35Hz and 475Hz respectively. Divide, then, for each demarcated sub-band, apply a fourth-order Butterworth bandpass filter that fully matches its boundary frequency to separate the signal components of the sub-band from the original real-time sound time domain signal. Subsequently, a 20-ms sliding rectangular time window and a 10-ms step are also used to calculate the short-time energy of the filtered signal of each sub-band to generate a series of energy values arranged in time, and these energy values are normalized to the maximum value. Finally, the normalized energy envelopes of all sub-bands are combined in order from low to high according to their center frequency to form a multi-sub-band energy envelope curve set of real-time sound.
[0131] Based on the multi-subband energy envelope curve set of real-time sound, the system first extracts the normalized amplitude sequence of each sub-band from the curve set. At the same time, the system calls the benchmark envelope correlation coefficient determined in the previous step, and finds the pair of energy envelope curves used to calculate the benchmark coefficient, and records the sub-band serial numbers corresponding to them. For example, if the benchmark coefficient is calculated based on the energy envelope curves of the 2nd and 3rd sub-bands, the system will pair the 2nd and 3rd energy envelope curves in the multi-subband energy envelope curve set of the current real-time sound. Then, the system uses the Pearson correlation coefficient formula that is exactly the same as that for calculating the benchmark coefficient to calculate the pair of real-time energy envelope curves. The normalized cross-correlation coefficient within the last 10 seconds is used to obtain a real-time coefficient value, such as 0.85. Next, the system compares this real-time calculated normalized cross-correlation coefficient with the stored baseline envelope cross-correlation coefficient (e.g., 0.95) and calculates the absolute value of the difference between the two, i.e., |0.85-0.95|=0.10. This difference represents the degree of change in the device operating status in terms of energy fluctuation synchronization. The system will repeat the above real-time calculation and difference calculation process for all curve pairs in the set of target envelope lines for cross-correlation calculation when establishing the baseline. Finally, all calculated absolute values of the difference are averaged to obtain the final real-time synchronization difference.
[0132] Based on the real-time synchronization difference and combined with the spectral peak feature offset calculated in the previous step, the system integrates these two core indicators to form a two-dimensional device sound comprehensive change index. Specifically, this indicator is constructed as a vector containing two elements, with the format of [spectral peak feature offset, real-time synchronization difference]. For example, if the calculated spectral peak feature offset is 14.67Hz and the real-time synchronization difference is 0.10, the generated device sound comprehensive change index is [14.67, 0.10]. This integration method retains the independent physical meaning of the two indicators. The spectral peak feature offset mainly reflects the static or slow drift of the equipment operating conditions (such as speed and switching frequency), while the real-time synchronization difference mainly captures the changes in the dynamic coupling relationship between the components within the equipment. Combining them together can comprehensively evaluate the changes in the equipment sound characteristics from both static and dynamic dimensions to form a device sound comprehensive change index.
[0133] The steps to obtain the threshold compliance status are:
[0134] Extracting the peak feature offset from the device sound comprehensive change index based on the device sound comprehensive change index, comparing the peak feature offsets with the peak offset threshold pre-set according to the current acoustic scene category, and determining whether the peak feature offsets exceed the peak offset threshold one by one to generate a peak offset threshold determination result.
[0135] Based on the comprehensive device sound change index, the real-time synchronization difference in the comprehensive device sound change index is extracted. According to the synchronization difference threshold pre-set according to the current acoustic scene category, the real-time synchronization difference is compared with the synchronization difference threshold. If the value of the real-time synchronization difference exceeds the synchronization difference threshold, it is marked as abnormal; if it does not exceed the synchronization difference threshold, it is marked as normal, and the synchronization difference threshold judgment result is generated;
[0136] Based on the peak offset threshold judgment result and the synchronization difference threshold judgment result, logical judgment is performed simultaneously. When both are normal, the threshold compliance state is marked as normal. When either one is abnormal, the threshold compliance state is marked as abnormal, and the threshold compliance state is generated.
[0137] Specifically, based on the comprehensive change index of equipment sound, the system first extracts the first element from the two-dimensional vector index, that is, the spectral peak feature offset, such as 14.67Hz. Then, according to the currently identified acoustic scene category (such as "sunny breeze"), the system retrieves the corresponding spectral peak offset threshold from a pre-built threshold library. This threshold library is established offline, which stores independent thresholds for different equipment components (such as cooling fans, inverters) in all preset acoustic scenes. The determination of each threshold is based on the continuous recording of at least 1000 hours of the component in this specific scene. The statistical analysis of healthy operation data specifically calculates the 99.7% quantile of all spectral peak feature offset values during this period (that is, the position of the mean plus three standard deviations) and uses this as the boundary for judging abnormalities. For example, the system queries that the spectral peak offset threshold of the cooling fan in the "sunny day and breeze" scenario is 11Hz. The system then compares the real-time calculated spectral peak feature offset of 14.67Hz with the threshold of 11Hz. Since 14.67Hz is greater than 11Hz, the system determines that the indicator exceeds the threshold and finally generates a Boolean type spectral peak offset threshold judgment result, marked as "abnormal".
[0138] Based on the comprehensive device sound variation index, the system then extracts the second element from this two-dimensional vector index, namely the real-time synchronization difference, for example, 0.10. The system then retrieves the synchronization difference threshold corresponding to the current acoustic scene category (e.g., "sunny day with a gentle breeze") from the same pre-built threshold library. This threshold is set similarly to the peak shift threshold and is derived through statistical analysis of a large amount of historical health data. The system collects all real-time synchronization difference values calculated when the device is operating normally in a specific acoustic scene, forming a distribution, and uses the 99.7% quantile of this distribution as the synchronization difference threshold. This method ensures that the threshold can effectively distinguish between normal fluctuations and significant changes caused by faults. For example, the system finds that the synchronization difference threshold for the "sunny day with a gentle breeze" scenario is 0.15. The system then compares the real-time synchronization difference of 0.10 with the threshold of 0.15. Since 0.10 does not exceed 0.15, the system determines that the indicator is within the normal range and marks the result as "normal," generating the synchronization difference threshold determination result.
[0139] Based on the peak shift threshold judgment result and the synchronization difference threshold judgment result, the system performs a logical "and" operation to integrate these two independent judgment conclusions. The system checks whether the peak shift threshold judgment result is "normal" and whether the synchronization difference threshold judgment result is also "normal". Only when these two conditions are met at the same time, the final comprehensive judgment is "normal". In the current example, the peak shift threshold judgment result is "abnormal", and the synchronization difference threshold judgment result is "normal". Since the requirement that both conditions are "normal" is not met at the same time, the result of the logical judgment is false. Therefore, the system will uniformly mark any one or two judgment results as "abnormal" as a comprehensive status abnormality. Finally, the system marks the threshold compliance status as "abnormal" and generates the conclusion.
[0140] The steps for obtaining the conclusion of the photovoltaic equipment operating status are as follows:
[0141] Based on the threshold compliance status, determine whether the threshold compliance status is normal. If it is normal, extract the spectrum peak features of the current real-time sound, read each spectrum peak frequency point and corresponding amplitude one by one, and replace the corresponding spectrum peak frequency points and amplitudes of the spectrum peak features of the scene-associated device one by one according to the spectrum peak frequency order of the spectrum peak features of the scene-associated device, to generate an updated spectrum peak feature of the scene-associated device;
[0142] Based on the updated scene-associated device spectral peak characteristics, the normalized mutual correlation coefficient of the real-time sound corresponding to the current real-time sound is extracted, and the normalized mutual correlation coefficient of the real-time sound is replaced one by one with the value of the corresponding position in the benchmark envelope mutual correlation coefficient to form an updated benchmark envelope mutual correlation coefficient, and the photovoltaic equipment operation status judgment conclusion is output.
[0143] Specifically, based on the threshold compliance state, the system first judges the state. If its value is "normal", it indicates that the current device sound characteristics have not deviated significantly in both the static spectrum peak and dynamic synchronization dimensions. The system then starts the adaptive benchmark update program and retrieves from the memory the spectrum peak frequency point set of the real-time photovoltaic equipment operating sound that has been generated when calculating the spectrum peak feature offset. The set contains the average frequency and corresponding average amplitude of the first three spectrum peaks of various components such as cooling fans and inverters sorted by amplitude. The system reads the data in this set one by one. For example, the spectrum peak data of the cooling fan is read as [351Hz, 118], [851Hz, 98], [1205Hz, 75], then the system loads the corresponding scene-associated device spectral peak feature data structure under the current acoustic scene, and matches the real-time spectral peak data with the benchmark data one by one according to the component name (such as "cooling fan") and the frequency order of the spectral peak from low to high. Then, the system writes the new real-time spectral peak frequency and amplitude data into the corresponding position in the data structure, directly overwriting the original benchmark value. This replacement operation will be performed on all spectral peaks of all identified device components in sequence to generate updated scene-associated device spectral peak features.
[0144] After the threshold meets the status of "normal" and the update of the spectral peak characteristics of the scene-related equipment is completed, the system continues to adaptively update the synchronization benchmark. The system first extracts the normalized mutual correlation values of the energy envelope curves of all real-time sounds calculated in the process of calculating the real-time synchronization difference to form a real-time coefficient list, such as [0.94, 0.92]. At the same time, the system loads the currently stored benchmark envelope mutual correlation coefficient, which is also a list containing historical benchmark values that correspond one to one with the real-time coefficient list, such as [0.95, 0.91]. The system replaces the corresponding positions in the benchmark envelope mutual correlation coefficient list with the values in the real-time coefficient list one by one according to the order in which they are paired with the benchmark coefficients during calculation. The old value set is replaced, for example, 0.95 is replaced by 0.94, and 0.91 is replaced by 0.92. After completing the update of all benchmark values, the updated benchmark envelope mutual correlation coefficient is formed. Finally, the system outputs the final conclusion on the operation status of the photovoltaic equipment based on the threshold compliance status judged this time. Since the threshold compliance status this time is "normal", the conclusion output by the system is "the equipment operation status is normal". If the threshold compliance status judged in the previous step is "abnormal", the system does not perform any benchmark update operation, and directly outputs the conclusion as "the equipment operation status is abnormal" and attaches specific abnormal indicators, such as "spectral peak characteristic offset exceeds the limit: 14.67Hz" or "real-time synchronization difference exceeds the limit: 0.21".
[0145] The above are merely preferred embodiments of the present invention and do not limit the present invention in any other form. Any technician familiar with the profession may use the technical content disclosed above to change or modify it into an equivalent embodiment with equivalent changes and apply it to other fields. However, any simple modification, equivalent change and modification made to the above embodiment based on the technical essence of the present invention without departing from the content of the technical solution of the present invention shall still fall within the scope of protection of the technical solution of the present invention.
Claims
1. A photovoltaic equipment fault detection method based on voiceprint recognition, characterized in that: The following steps are involved: Collect ambient sound around the current photovoltaic power station and establish the spectrum peak characteristics of scene-related equipment; Based on the spectral peak characteristics of the scene-associated device, the current photovoltaic device operating sound is divided into multiple sub-bands, and the short-time energy envelope of each sub-band signal is calculated to obtain a multi-sub-band energy envelope curve set. Based on the multi-sub-band energy envelope curve set, the cross-correlation coefficient between the target envelope pairs is calculated to establish a reference envelope cross-correlation coefficient; Receive newly collected real-time photovoltaic equipment operating sounds, extract several spectral peak frequency points and corresponding amplitudes from the real-time photovoltaic equipment operating sounds, calculate them with the spectral peak characteristics of the scene-associated equipment, and obtain spectral peak characteristic offsets; perform multi-subband energy envelope extraction on the newly collected real-time photovoltaic equipment operating sounds and calculate the cross-correlation coefficients between the new envelope lines; perform difference calculation on the cross-correlation coefficients with the reference envelope to obtain real-time synchronization differences, and generate a comprehensive equipment sound change index; Based on the comprehensive change index of the device sound, the spectral peak feature offset is compared with the spectral peak offset threshold set for the current acoustic scene category, and the real-time synchronization difference is compared with the corresponding synchronization difference threshold to generate a threshold compliance state. Based on the threshold compliance state, a conclusion on the operation status of the photovoltaic device is output.
2. The photovoltaic equipment fault detection method based on voiceprint recognition according to claim 1 is characterized in that: The steps for obtaining the spectrum peak characteristics of the scene-related device are as follows: The ambient sound around the current photovoltaic power station is collected. Based on the ambient sound, the spectrum distribution is obtained using Fourier transform. The target frequency band in the spectrum distribution is extracted. The frequency bands are merged using a Mel filter bank. The logarithmic power spectrum is calculated and then a discrete cosine transform is performed to obtain the Mel-frequency cepstrum coefficients of the ambient sound. Based on the Mel-frequency cepstral coefficients of the ambient sound, the Mel-frequency cepstral coefficients are input into a pre-trained scene classification model, the characteristic distance between the Mel-frequency cepstral coefficients and the preset scene category is calculated, the scene category corresponding to the ambient sound is determined, and a preset acoustic scene category is generated; Based on the preset acoustic scene category, the real-time operating sounds of the inverter and cooling fan components in the photovoltaic equipment are collected, and the spectrum information of the real-time operating sounds is extracted. By searching for the maximum points of the spectrum, multiple spectral peak frequency points in the operating sound spectrum of the inverter and cooling fan components are located respectively, and the corresponding amplitudes of the spectral peak frequency points are extracted. The spectral peak frequency points and the corresponding amplitudes are associated with the preset acoustic scene category to form the spectral peak characteristics of the scene-associated device.
3. The photovoltaic equipment fault detection method based on voiceprint recognition according to claim 1 is characterized in that: The steps of obtaining the multi-subband energy envelope curve set are: Based on the spectral peak characteristics of the scenario-associated device, extract the spectral peak frequency points of the operating sound of the inverter and cooling fan components from the spectral peak characteristics of the scenario-associated device, use the spectral peak frequency points as the center reference frequency, expand the frequency spectrum of the current photovoltaic device operating sound to both sides of the frequency axis with a fixed width, determine the boundary range of each frequency sub-band, and obtain multiple frequency sub-bands of the photovoltaic device operating sound; Based on the multiple frequency sub-bands of the photovoltaic device operation sound, each frequency sub-band is band-pass filtered to separate the time domain sound signal within the corresponding sub-band frequency range, and the filtered sub-band time domain sound signal is subjected to a sliding rectangular time window, and the integral of the square of the sound signal amplitude in each time window is gradually calculated to obtain the short-time energy envelope value of each sub-band frequency signal; Based on the short-time energy envelope value of each sub-band frequency signal, a multi-sub-band energy envelope curve set corresponding to the current photovoltaic device operation sound is formed in sequence according to the sequence number of the sub-band frequency.
4. The photovoltaic equipment fault detection method based on voiceprint recognition according to claim 1, characterized in that: The steps for obtaining the reference envelope correlation coefficient are: Based on the multi-subband energy envelope curve set, two adjacent energy envelope curves are combined into a pair of curves in order of subband sequence numbers, the average value of the amplitude difference of each pair of energy envelope curves at the same time point within a fixed time period is calculated, and curve pairs whose average difference does not exceed the amplitude threshold and whose direction change signs are consistent are selected to obtain a set of target envelope curve pairs for cross-correlation calculation; Calculating a target envelope pair set based on the cross-correlation, extracting the amplitude sequence of each pair of envelopes, calculating the time-averaged amplitude of each curve, and calculating the normalized cross-correlation coefficient; Based on the normalized cross-correlation coefficient of each pair of energy envelope curves, all curve pairs are traversed, all normalized cross-correlation coefficients are extracted, and the largest cross-correlation coefficient is selected as the reference envelope cross-correlation coefficient in the current state.
5. The photovoltaic equipment fault detection method based on voiceprint recognition according to claim 1, characterized in that: The steps for obtaining the peak characteristic offset are: Receive newly collected real-time photovoltaic equipment operation sound signals, divide the sound signals in a continuous time period into overlapping time windows of fixed length, perform Fourier transform on the signal in each window to extract the spectral power density, select the top three spectral peak frequency points with the highest amplitude, record the frequency values and corresponding amplitudes, and sort them from low to high according to the spectral peak frequency to form a set of spectral peak frequency points of the real-time photovoltaic equipment operation sound; According to the spectrum peak frequency point set of the real-time photovoltaic device operating sound, the corresponding frequency points in the spectrum peak characteristics of the scene-associated device are paired one by one according to the frequency sorting, and the spectrum peak characteristic offset of each group of frequency points is calculated; Based on the spectral peak feature offset, a spectral peak feature offset sequence is constructed. If there are more than three consecutive groups of spectral peak feature offsets exceeding the threshold, the segment is marked as an abnormal spectrum state, and an offset abnormality warning message is output. At the same time, the mean of all spectral peak feature offsets is calculated and output as the spectral peak feature offset.
6. The photovoltaic equipment fault detection method based on voiceprint recognition according to claim 1, characterized in that: The steps for obtaining the device sound comprehensive change index are as follows: For newly collected real-time photovoltaic equipment operating sounds, the real-time sound signal is divided into multiple sub-bands in the spectrum according to the peak frequency points in the spectral peak characteristics of the scene-related equipment. The signals of each sub-band are separated through bandpass filtering. The short-time energy of the filtered signal of each sub-band is calculated and the corresponding short-time energy envelope is extracted to form a multi-sub-band energy envelope curve set of the real-time sound; Based on the multi-subband energy envelope curve set of the real-time sound, extract the amplitude sequence of each subband in the energy envelope curve, pair them one by one with the energy envelope curve subband sequence corresponding to the reference envelope cross-correlation coefficient according to the subband frequency order, calculate the normalized cross-correlation coefficient of the energy envelope curve pair of the real-time sound pair by pair, calculate the absolute value of the difference with the reference envelope cross-correlation coefficient one by one, and then take the average to obtain the real-time synchronization difference; Based on the real-time synchronization difference and in combination with the spectrum peak characteristic offset, the real-time synchronization difference and the spectrum peak characteristic offset are integrated to form a comprehensive device sound change index.
7. The photovoltaic equipment fault detection method based on voiceprint recognition according to claim 1, characterized in that: The steps for obtaining the threshold compliance status are: Extracting a spectral peak feature offset from the device sound comprehensive change index based on the device sound comprehensive change index, comparing the spectral peak feature offsets with a spectral peak offset threshold value pre-set according to the current acoustic scene category, determining whether the spectral peak feature offsets exceed the spectral peak offset threshold value one by one, and generating a spectral peak offset threshold determination result; Based on the device sound comprehensive change index, extract the real-time synchronization difference in the device sound comprehensive change index, compare the real-time synchronization difference with the synchronization difference threshold according to a preset synchronization difference threshold for the current acoustic scene category, and mark it as abnormal if the value of the real-time synchronization difference exceeds the synchronization difference threshold; otherwise, mark it as normal, and generate a synchronization difference threshold determination result; Based on the peak shift threshold determination result and the synchronization difference threshold determination result, logical judgment is performed simultaneously. When both are normal, the threshold compliance state is marked as normal. When either one is abnormal, the threshold compliance state is marked as abnormal, and a threshold compliance state is generated.
8. The photovoltaic equipment fault detection method based on voiceprint recognition according to claim 1, characterized in that: The steps for obtaining the conclusion of the photovoltaic equipment operating status are as follows: Based on the threshold compliance state, determining whether the threshold compliance state is normal; if normal, extracting the spectrum peak feature of the current real-time sound, reading each spectrum peak frequency point and corresponding amplitude one by one, and replacing the corresponding spectrum peak frequency points and amplitudes of the spectrum peak feature of the scene-associated device one by one according to the spectrum peak frequency order of the spectrum peak feature of the scene-associated device, to generate an updated spectrum peak feature of the scene-associated device; Based on the updated scene-associated device spectral peak characteristics, the normalized mutual correlation coefficient of the real-time sound corresponding to the current real-time sound is extracted, and the normalized mutual correlation coefficients of the real-time sound are replaced one by one with the values of the corresponding positions in the reference envelope mutual correlation coefficient to form an updated reference envelope mutual correlation coefficient, and a conclusion on the operation status of the photovoltaic equipment is output.
9. The photovoltaic equipment fault detection system according to any one of claims 1 to 8, characterized in that: include: The voiceprint collection module collects the ambient sound around the current photovoltaic power station and establishes the spectrum peak characteristics of scene-related equipment; An energy envelope calculation module divides the current photovoltaic equipment operating sound into multiple sub-bands based on the spectral peak characteristics of the scene-associated device, calculates the short-time energy envelope of each sub-band signal, and obtains a multi-sub-band energy envelope curve set; based on the multi-sub-band energy envelope curve set, calculates the mutual correlation coefficient between the target envelope pairs and establishes a reference envelope mutual correlation coefficient; The real-time sound analysis module receives newly collected real-time photovoltaic equipment operating sounds, extracts several spectral peak frequency points and corresponding amplitudes from the real-time photovoltaic equipment operating sounds, calculates the spectral peak characteristics of the equipment associated with the scene, and obtains the spectral peak characteristic offset; extracts multi-subband energy envelopes from the newly collected real-time photovoltaic equipment operating sounds and calculates the mutual correlation coefficients between the new envelope line pairs; Calculate the difference between the reference envelope and the cross-correlation coefficient to obtain the real-time synchronization difference and generate the comprehensive change index of the device sound; a fault judgment module that compares the spectral peak feature offset with a spectral peak offset threshold set for the current acoustic scene category based on the device sound comprehensive change index; The real-time synchronization difference is compared with the corresponding synchronization difference threshold to generate a threshold compliance state; based on the threshold compliance state, a photovoltaic device operation state determination conclusion is output.
Citation Information
Cited By
Intelligent spot inspection management platform and method
CN120764860A
Device working state recognition method based on voiceprint recognition model
CN120853617A
A device working state recognition method based on a voiceprint recognition model
CN120853617B
Photovoltaic inverter fault processing method and device based on acoustic characteristics
CN121171259A