Method and device for constructing training data set of arc detection model, and model
Patent Information
- Application Number
- CN202610976740.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-01
- Publication Date
- 2026-09-25
AI Technical Summary
上述的模型训练数据集构建方法,需人工框选标注电弧发生区域,标注耗时且质量不稳定,不利于数据集批量制作
[0016]本申请,首先采集光伏直流回路原始电流信号进行预处理,分帧后经频域转换得到初始频谱序列;其次,基于电弧特征频段从初始频谱序列中截取特征频点构建特征频谱序列,滤除无关干扰;然后,根据预设的时间帧数堆叠特征频谱序列,生成时频特征矩阵,无需强制时间帧数与特征频点数对齐,可有效减少数据存储与计算消耗;最后,将时频特征矩阵转换为标准张量作为模型数据样本,模型数据样本仅需标注电弧故障标签与正常标签,无需人工框选故障区域。由此,可以提高样本制作效率,得到数据精简、有效,专用于拉弧检测模型的训练数据集。
Smart Images

Figure CN122817869A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of photovoltaic power generation technology, specifically to a method, apparatus, and model for constructing a training dataset for an arc detection model. Background Technology
[0002] Photovoltaic power generation systems are prone to DC arcing faults during operation, which are a major cause of electrical fires. Therefore, accurate and reliable detection of DC arcing is crucial for ensuring the safe operation of photovoltaic power generation systems. Deep learning models, as the primary method for DC arcing detection, have detection performance closely related to the quality of the training dataset.
[0003] Currently, model training datasets are often constructed using current waveforms or single-frame spectra to build training samples. Some schemes use complete time-frequency diagrams or full-spectrum signals and construct square two-dimensional time-frequency diagrams of fixed dimensions, then construct training samples by labeling the arc-generating regions. These methods require manual selection and labeling of the arc-generating regions, which is time-consuming and of inconsistent quality, making them unsuitable for batch dataset production.
[0004] Therefore, improving the efficiency of training dataset production is an urgent problem to be solved. Summary of the Invention
[0005] This application provides a method, apparatus, and model for constructing a training dataset for an arc detection model, aiming to solve the aforementioned technical problems.
[0006] Firstly, this application provides a method for constructing a training dataset for an arc detection model, comprising: acquiring the original current signal in a photovoltaic DC circuit as an initial current signal, and preprocessing the initial current signal to obtain a standard current signal; performing frame-by-frame processing on the standard current signal to obtain multiple single-frame signals, performing frequency domain transformation on each single-frame signal to obtain a corresponding initial spectrum sequence; determining the frequency point index range according to a preset arc characteristic frequency band, sampling frequency, and frequency domain transformation length, and truncating the initial spectrum sequence by characteristic frequency points according to the frequency point index range to obtain a characteristic spectrum sequence; aggregating and stacking the characteristic spectrum sequences according to a preset number of time frames to obtain a time-frequency feature matrix with the number of time frames as the first dimension and the number of characteristic frequency points as the second dimension; wherein the number of time frames and the number of characteristic frequency points are independent of each other; converting the time-frequency feature matrix into a standard tensor according to a preset transformation rule to obtain model data samples, which are used to construct a training dataset for training the arc detection model.
[0007] In some possible implementations, the standard current signal is segmented into multiple single-frame signals, and each single-frame signal is subjected to frequency domain transformation to obtain a corresponding initial spectrum sequence. This includes: segmenting the standard current signal into multiple single-frame signals according to a preset frame length and a preset frame shift; weighting each single-frame signal using a window function to obtain a weighted single-frame signal; and performing a fast Fourier transform on the weighted single-frame signal to obtain the initial spectrum sequence.
[0008] In some possible implementations, the method further includes: setting a sliding window according to a preset size, sliding the sliding window along the standard current signal in a time sequence, and sequentially capturing local sampling segments of the standard current signal; obtaining the instantaneous rate of change of the standard current signal based on the amplitude difference between adjacent sampling points within the local sampling segment; comparing the instantaneous rate of change with a preset rate of change threshold, and determining a preset frame length based on the comparison result.
[0009] In some possible implementations, performing a Fast Fourier Transform (FFT) on the weighted single-frame signal to obtain an initial spectral sequence includes: performing a FFT on the weighted single-frame signal to obtain a complex spectral sequence; splitting the complex spectral sequence to obtain a phase spectral sequence and an amplitude spectral sequence; expanding the amplitude spectral sequence to obtain an expanded data sequence, which is either a power spectral sequence or a logarithmic amplitude spectral sequence; and using any one of the complex spectral sequence, the amplitude spectral sequence, and the expanded data sequence as the initial spectral sequence.
[0010] In some possible implementations, a frequency point index range is determined based on a preset arc characteristic frequency band, sampling frequency, and frequency domain transformation length. The initial spectrum sequence is then truncated based on the frequency point index range to obtain a characteristic spectrum sequence. This includes: determining a first frequency point index and a second frequency point index corresponding to the arc characteristic frequency band based on the preset upper and lower limits of the arc characteristic frequency band, sampling frequency, and frequency domain transformation length; determining the frequency point index range based on the first and second frequency point indices; and truncating characteristic frequency points from the initial spectrum sequence based on the frequency point index range to obtain the characteristic spectrum sequence.
[0011] In some possible implementations, the method further includes: determining the observation frequency band of the electric arc and using the observation frequency band as the characteristic frequency band of the electric arc; or, acquiring the current signal of the photovoltaic DC circuit under normal operating conditions, performing a fast Fourier transform on the current signal to obtain full-band spectrum data, extracting the background noise amplitude corresponding to each frequency band based on the full-band spectrum data, and using the frequency band with the background noise amplitude lower than the noise threshold as the characteristic frequency band of the electric arc; or, setting multiple non-overlapping target frequency bands and using the target frequency bands as the characteristic frequency bands of the electric arc.
[0012] In some possible implementations, feature spectrum sequences are aggregated and stacked according to a preset number of time frames to obtain a time-frequency feature matrix with the number of time frames as the first dimension and the number of feature frequency points as the second dimension. This includes: collecting a corresponding number of feature spectrum sequences according to a preset number of time frames; sorting the feature spectrum sequences according to their chronological order; and splicing the sorted feature spectrum sequences to obtain the time-frequency feature matrix.
[0013] In some possible implementations, the time-frequency feature matrix is converted into a standard tensor according to a preset conversion rule to obtain model data samples, including: converting the time-frequency feature matrix into a standard tensor of the same size according to the conversion rule; mapping the standard tensor into a color image, wherein the color image is a grayscale image, an RGB color image, or an HSV color gamut image; and using the color image as a model data sample.
[0014] Secondly, this application provides a training dataset construction device for an arc detection model. The device includes: a signal preprocessing module for acquiring the original current signal in the photovoltaic DC circuit as the initial current signal and preprocessing the initial current signal to obtain a standard current signal; a frequency domain transformation module for performing frame-by-frame processing on the standard current signal to obtain multiple single-frame signals, and performing frequency domain transformation on each single-frame signal to obtain a corresponding initial spectrum sequence; a feature extraction module for determining the frequency point index range according to the preset arc characteristic frequency band, sampling frequency, and frequency domain transformation length, and extracting feature frequency points from the initial spectrum sequence according to the frequency point index range to obtain a feature spectrum sequence; a feature reconstruction module for aggregating and stacking the feature spectrum sequences according to the preset number of time frames to obtain a time-frequency feature matrix with the number of time frames as the first dimension and the number of feature frequency points as the second dimension; and a sample construction module for converting the time-frequency feature matrix into a standard tensor according to a preset conversion rule to obtain model data samples, which are used to construct a training dataset for training the arc detection model.
[0015] Thirdly, this application provides an arc detection model, which is trained based on a training dataset, which is obtained through the training dataset construction method described above.
[0016] This application first collects and preprocesses the raw current signal of the photovoltaic DC circuit, then performs frequency domain transformation after framing to obtain an initial spectrum sequence. Second, based on the characteristic frequency band of the electric arc, characteristic frequency points are extracted from the initial spectrum sequence to construct a characteristic spectrum sequence, filtering out irrelevant interference. Then, the characteristic spectrum sequences are stacked according to a preset number of time frames to generate a time-frequency feature matrix. This eliminates the need for mandatory alignment of the time frame number with the number of characteristic frequency points, effectively reducing data storage and computational costs. Finally, the time-frequency feature matrix is converted into a standard tensor as model data samples. These samples only need to be labeled with arc fault and normal labels, eliminating the need for manual selection of fault areas. This improves sample production efficiency, resulting in a concise and effective training dataset specifically designed for arc detection models. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a flowchart illustrating the training dataset construction method provided in the embodiments of this application; Figure 2 This is a schematic diagram of a photovoltaic DC circuit provided in an embodiment of this application; Figure 3 This is the Normal multi-frame time-frequency diagram provided in the application embodiment; Figure 4 This is the Acr multi-frame time-frequency diagram provided in the application embodiment; Figure 5 It is a tensor graph of the time-frequency feature matrix corresponding to the Normal single-frame signal provided in the application embodiment; Figure 6 It is a tensor graph of the time-frequency feature matrix corresponding to the Acr single-frame signal provided in the application embodiment; Figure 7 This is a schematic diagram of the structure of the training dataset construction device provided in the application embodiment; Figure 8 This is a schematic diagram of the deployment of the arc detection model provided in the application embodiment. Detailed Implementation
[0019] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0020] In the description of this application, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, features defined as "first" or "second" may explicitly or implicitly include one or more features. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified. "A and / or B" includes the following three combinations: only A, only B, and a combination of A and B.
[0021] The use of "applies to" or "configured to" in this application implies open and inclusive language, which does not preclude applicability to or configuration to devices performing additional tasks or steps. Furthermore, the use of "based on" implies openness and inclusivity, because processes, steps, calculations, or other actions "based on" one or more conditions or values may in practice be based on additional conditions or values beyond those conditions.
[0022] In this application, the term "exemplary" is used to mean "used as an example, illustration, or description." Any embodiment described as "exemplary" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use this application. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that this application can be made without using these specific details. In other instances, well-known structures and processes are not described in detail to avoid obscuring the description of this application with unnecessary detail. Therefore, this application is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed in this application.
[0023] Before introducing the training dataset construction method, apparatus and model for arc detection model provided in the embodiments of this application, the relevant technical content will be explained first so that those skilled in the art can fully understand the content of this application.
[0024] The relevant technology employs a deep learning model to detect DC arcing in photovoltaic power generation systems. However, the detection performance of the deep learning module is highly dependent on the quality of the training dataset. The relevant technology suffers from the following shortcomings in constructing the training dataset: Limited feature dimension and insufficient information utilization: Related technologies often directly input one-dimensional current waveforms or single-frame spectra into the model, failing to effectively construct and utilize the correlation between fault features in the two-dimensional space of time and frequency, thus limiting the upper limit of model performance.
[0025] Data redundancy and noise interference are serious: related technologies often directly use full-time, full-frequency signal input models, introducing a large amount of irrelevant information, which dilutes the effective features, makes model training difficult and prone to overfitting.
[0026] The annotation process is complex and costly: related technologies are based on two-dimensional time-frequency maps for arc detection, which requires manual selection and annotation of the arc regions in the two-dimensional time-frequency maps. The annotation efficiency is low and the annotation quality is inconsistent, which restricts the large-scale production of datasets and the efficiency of model iteration.
[0027] Rigid input shape and low computational efficiency: Related technologies for arc detection are based on two-dimensional convolutional neural network models, which typically construct a square matrix with a fixed aspect ratio as the model input, such as a 224×224 square matrix. Since the number of frequency points representing an arc can often reach hundreds, a large number of frames need to be stacked to construct the square matrix, making the model unsuitable for deployment on microcontroller units with limited resources in photovoltaic power generation systems, thus restricting the practical application of the model.
[0028] Based on this, this application provides a method for constructing a training dataset for an arc detection model. By extracting feature frequency points representing arc characteristics from the original current signal of a photovoltaic DC circuit, a two-dimensional time-frequency feature matrix is obtained by stacking these feature frequency points over time. A training dataset is then constructed based on this time-frequency feature matrix. Thus, the time-frequency feature matrix integrates the correlation features of the time and frequency domains, and all extracted feature frequency points are related to the arc. This simplifies the data without requiring additional annotation of the arc region. Furthermore, the time-frequency feature matrix can adopt an asymmetric structure, eliminating the need to stack a large number of frames, thereby obtaining a concise and effective training dataset specifically for the arc detection model.
[0029] Firstly, this embodiment provides a method for constructing a training dataset for an arc detection model, such as... Figure 1 As shown. The method for constructing the training dataset includes steps S101 to S105, which will be described in detail below.
[0030] Step S101: Collect the original current signal in the photovoltaic DC circuit as the initial current signal, and preprocess the initial current signal to obtain the standard current signal.
[0031] As an example, the raw current signal of the photovoltaic DC circuit can be acquired at sampling rates of 50kHz, 250kHz, or 500kHz, with the sampling rate selected to fully highlight the high-discrimination frequency band of the electric arc. Subsequently, wavelet threshold filtering or median filtering is used to remove impulse noise from the signal, while gain enhancement is applied to the 2kHz-8kHz arc characteristic frequency band, and power frequency interference and high-frequency switching noise are suppressed, completing the preprocessing of the initial current signal. The photovoltaic DC circuit is as follows: Figure 2As shown, the original current signal under normal operating conditions and the original current signal under arc fault conditions can be collected from the photovoltaic DC circuit.
[0032] In step S101, the original current signal is acquired by setting a sampling rate, and the original current signal is subjected to noise reduction and arc feature enhancement processing by combining filtering and enhancement processing, thereby obtaining a standard current signal with good signal-to-noise ratio.
[0033] Step S102: Perform frame segmentation on the standard current signal to obtain multiple single-frame signals, and perform frequency domain transformation on each single-frame signal to obtain the corresponding initial spectrum sequence.
[0034] As an example, firstly, a preset frame length and frame shift can be used to gradually slide along a standard current signal to obtain multiple single-frame signals. Secondly, a discrete Fourier transform is performed on each single-frame signal to obtain the initial spectrum sequence corresponding to each single-frame signal.
[0035] In some possible implementations, step S102 may include steps S1021 to S1023, which will be described in detail below.
[0036] Step S1021: Perform frame segmentation on the standard current signal according to the preset frame length and preset frame shift to obtain multiple single-frame signals.
[0037] As an example, the preset frame length can be a fixed frame length. For example, setting the fixed frame length to L and the frame shift R to a value of Therefore, the standard current signal is divided into sliding frames according to the preset frame length and preset frame shift. In a specific example, the frame length can be set to 20ms and the frame shift can be set to 10ms to achieve sliding frame division.
[0038] As another example, the preset frame length can also be a variable frame length to retain temporal resolution during steady-state periods and sufficient detail during non-steady-state periods. For example, a sliding window is set according to a preset size, and the sliding window is time-sequentially slid along a standard current signal to sequentially capture local sampling segments of the standard current signal; the instantaneous rate of change of the standard current signal is obtained based on the amplitude difference between adjacent sampling points within the local sampling segment; the instantaneous rate of change is compared with a preset rate of change threshold, and the preset frame length is determined based on the comparison result.
[0039] In a specific example, assuming a local sampling segment contains N sampling points, the amplitude difference between adjacent sampling points is calculated sequentially, and the average value is taken as the instantaneous rate of change of the local sampling segment. This instantaneous rate of change is compared with a preset rate of change threshold. When the instantaneous rate of change is less than the preset rate of change threshold, the signal is determined to be in a steady-state region, and a first frame length is selected as the preset frame length, such as 80 sampling points. A single frame signal is then extracted from the local sampling segment based on the preset frame length and a preset frame shift. When the instantaneous rate of change is greater than or equal to the preset rate of change threshold, the signal is determined to be in a non-steady-state region. To retain as much detail as possible, a second frame length is selected as the preset frame length, such as 20 sampling points. The second frame length is less than the first frame length, and a single frame signal is then extracted from the local sampling segment based on the preset frame length and a preset frame shift to capture as many instantaneous signal features as possible. Thus, by continuously traversing the entire standard current signal along the time sequence through a sliding window and dynamically switching the preset frame length according to the instantaneous rate of change of each local sampling segment, adaptive framing is completed.
[0040] Step S1022: Weight each single-frame signal using a window function to obtain the weighted single-frame signal.
[0041] As an example, the window function can be any of the Hanning Window, Hamming Window, Blackman Window, and Kaiser Window to weight each single frame signal, thereby suppressing spectral leakage during subsequent frequency domain transformation.
[0042] Step S1023: Perform a Fast Fourier Transform on the weighted single-frame signal to obtain the initial spectrum sequence.
[0043] As an example, the initial frequency sequence can be the complex spectrum sequence obtained by the fast Fourier transform, or it can be the amplitude spectrum sequence, logarithmic amplitude spectrum sequence, power spectrum sequence, etc., calculated from the complex spectrum sequence. This application does not impose any specific restrictions on this.
[0044] In a specific example, a 2048-point Fast Fourier Transform can be performed on a single frame signal. Since the single frame signal is a real-time signal, the output of its 2048-point Fast Fourier Transform has conjugate symmetry characteristics. The redundant frequency points in the second half are discarded, and only the amplitudes corresponding to the first 1024 frequency points are retained. A logarithmic transformation is then performed on these 1024 amplitude values, and the final output is a logarithmic amplitude spectrum sequence containing 1024 values as the initial spectrum sequence.
[0045] As a possible implementation of step S1023, step S1023 may include: performing a fast Fourier transform on the weighted single-frame signal to obtain a complex spectrum sequence; splitting the complex spectrum sequence to obtain a phase spectrum sequence and an amplitude spectrum sequence; expanding the amplitude spectrum sequence to obtain an expanded data sequence, wherein the expanded data sequence is a power spectrum sequence or a logarithmic amplitude spectrum sequence; and using any one of the complex spectrum sequence, the amplitude spectrum sequence, and the expanded data sequence as the initial spectrum sequence.
[0046] In a specific example, the amplitude spectrum sequence can be converted into a logarithmic spectrum sequence using the logarithmic transformation formula. Logarithmic spectrum sequences have a wider dynamic range and are easier for the human eye and neural networks to analyze. The logarithmic transformation formula is: ; in, This represents the amplitude corresponding to the k-th frequency point in the amplitude spectrum sequence; This is a preset setting to avoid extremely small positive numbers with log0. This is the logarithmic magnitude corresponding to the k-th frequency point, used to construct the logarithmic spectrum sequence.
[0047] In step S102, the standard current signal is first split into multiple single-frame signals by time sequence; then, frequency domain conversion and spectral leakage are achieved by using window function and fast Fourier transform; thus, the initial spectral sequence is obtained.
[0048] Step S103: Determine the frequency index range based on the preset arc characteristic frequency band, sampling frequency and frequency domain transformation length, and extract the characteristic frequency points from the initial spectrum sequence according to the frequency index range to obtain the characteristic spectrum sequence.
[0049] As an example, firstly, the corresponding frequency index range is determined based on the upper and lower limits of the arc's characteristic frequency band, the sampling frequency, and the frequency domain transform length. Then, based on this frequency index range, multiple characteristic frequency points are extracted from the initial spectrum sequence corresponding to each single frame signal to form a characteristic spectrum sequence, discarding other interfering frequency points. Thus, the constructed characteristic spectrum sequence retains only the characteristic frequency points related to the arc, simplifying the data without requiring additional labeling of the arc occurrence area.
[0050] In some possible implementations, step S103 may include steps S1031 to S1033, which will be described in detail below.
[0051] Step S1031: Determine the first frequency index and the second frequency index corresponding to the arc characteristic frequency band based on the preset upper and lower limits of the arc characteristic frequency band, the sampling frequency, and the frequency domain transformation length.
[0052] As an example, the upper and lower limits of the characteristic frequency band of the electric arc are taken as the target frequency ft. Substituting the target frequency ft, sampling frequency fs, and frequency domain transformation length L into the frequency index formula, the frequency index range formed by the first and second frequency indexes is obtained. The frequency index formula is: ; Where k is the frequency index in the initial spectrum sequence, and is an integer.
[0053] In the specific example, the characteristic frequency band of the electric arc is 2kHz-8kHz, the sampling frequency is 500kHz, and a single frame signal is transformed by Fast Fourier Transform to obtain 2048 frequency points, that is, the frequency domain transformation length is 2048. When 2kHz is used as the target frequency, the corresponding first frequency point index is the 8th frequency point index; when 8kHz is used as the target frequency, the corresponding second frequency point index is the 32nd frequency point index.
[0054] It should be noted that the characteristic frequency band of the electric arc can be obtained in several ways. For example, the observed frequency band of the electric arc can be determined and used as the characteristic frequency band of the arc; alternatively, the current signal of the photovoltaic DC circuit under normal operating conditions can be acquired, and a fast Fourier transform can be performed on the current signal to obtain full-band spectrum data. Based on the full-band spectrum data, the background noise amplitude corresponding to each frequency band can be extracted, and the frequency bands with background noise amplitudes below the noise threshold can be used as the characteristic frequency bands of the arc; alternatively, multiple non-overlapping target frequency bands can be set, and these target frequency bands can be used as the characteristic frequency bands of the arc.
[0055] In specific examples, the observation frequency band of DC arcing can be determined experimentally and used as the arc characteristic frequency band. This observation frequency band refers to the frequency band where DC arcing is most easily observed, for example, the observation frequency band is 2kHz-8kHz, where 8kHz and 2kHz are the upper and lower limits of the arc characteristic frequency band, respectively. Alternatively, a current signal under normal operating conditions can be analyzed using Fast Fourier Transform to identify the background noise amplitude corresponding to each frequency band, and the frequency band with the lower background noise amplitude can be extracted as the arc characteristic frequency band to avoid noise interference. Alternatively, multiple non-overlapping target frequency bands can be set as arc characteristic frequency bands, such as 1-5kHz and 25-35kHz, to detect different types of arcs and arcs with wideband characteristic distributions. Subsequently, the characteristic frequency points extracted from different target frequency bands can be spliced or stacked by channel to construct a characteristic spectrum sequence.
[0056] Step S1032: Determine the frequency index range based on the first frequency index and the second frequency index.
[0057] As an example, when the characteristic frequency band of the electric arc is 2kHz-8kHz, the sampling frequency is 500kHz, and the frequency domain transformation length is 2048, the first frequency point index is 8, and the second frequency point index is 32. Therefore, the frequency point index range is 8-32.
[0058] Step S1033: Based on the frequency index range, extract the characteristic frequency points from the initial spectrum sequence to obtain the characteristic spectrum sequence.
[0059] As an example, when the frequency index range is 8-32, the 8th to 32nd frequency points are extracted from the initial spectrum sequence corresponding to each single frame signal as feature frequency points, that is, 25 feature frequency points are extracted to construct a feature spectrum sequence.
[0060] Thus, in step S103, the frequency index range is determined by the characteristic frequency band of the electric arc, and characteristic frequency points are extracted from the frequency index range to construct a characteristic spectrum sequence. This characteristic spectrum sequence retains only the frequency point data related to the electric arc, which significantly compresses the data volume of the characteristic spectrum sequence and eliminates the need for additional labeling of the electric arc occurrence area.
[0061] Step S104: Aggregate and stack the feature spectrum sequences according to the preset number of time frames to obtain a time-frequency feature matrix with the number of time frames as the first dimension and the number of feature frequency points as the second dimension; wherein, the number of time frames and the number of feature frequency points are independent of each other.
[0062] As an example, assuming the number of feature frequency points in the feature spectrum sequence is N and the number of time frames is M, where M and N are positive integers, then M frames of feature spectrum sequences are stacked vertically in chronological order to generate a time-frequency feature matrix of size M×N. M and N are configured independently, and M can be preset according to the typical duration of the electric arc and the required model input size.
[0063] In some possible implementations, step S104 may include: collecting a corresponding number of feature spectrum sequences according to a preset number of time frames; sorting the feature spectrum sequences according to their chronological order; and concatenating the sorted feature spectrum sequences to obtain a time-frequency feature matrix.
[0064] As an example, when M=1, only single-frame feature spectrum sequences are stacked to form a special time-frequency feature matrix. When M is less than N, a two-dimensional time-frequency feature matrix of M rows and N columns is constructed. When M=N, the number of time frames M is equal to the number of feature frequency points N, constructing a symmetric two-dimensional time-frequency feature matrix. When M is greater than N, the two dimensions of the matrix are interchanged to construct an N-row, M-column two-dimensional time-frequency feature matrix to fit the fixed-length and width convolutional kernels of the model input layer, facilitating stable extraction of information from the time-frequency feature matrix by the convolutional kernels.
[0065] It should be noted that the number of time frames M and the number of feature frequency points N are independent of each other. The number of time frames M does not need to be consistent with the number of feature frequency points N. This allows for the construction of an M×N asymmetric time-frequency feature matrix to improve the computational efficiency of the model.
[0066] In other possible implementations, multiple time frames can be set, for example, 16 or 32 time frames, to generate time-frequency feature matrices at different time scales, which can identify arcs of different durations.
[0067] It should be noted that when the initial spectrum sequence is a complex spectrum sequence, the spectral amplitude and phase amplitude are retained in the extracted feature frequency points. This allows the construction of a three-dimensional time-frequency feature matrix of size M×N×2, which utilizes the additional information carried by the phase to improve the model's ability to distinguish electric arcs.
[0068] Thus, in step S104, the original current signal is reconstructed into a time-frequency feature matrix from two independent physical dimensions: the number of time frames M and the number of feature frequency points N. The number of time frames M and the number of feature frequency points N can be interchanged to match the requirements of the model's input layer. Simultaneously, by adjusting the number of time frames M, arc detection with different durations, both short and long, can be covered, thereby improving the model's generalization ability in arc recognition.
[0069] Step S105: Convert the time-frequency feature matrix into a standard tensor according to a preset conversion rule to obtain model data samples, which are used to construct a training dataset for training the arc detection model.
[0070] As an example, the time-frequency feature matrix of size M×N can first be normalized and standardized into a standard tensor of the same size; then, the standard tensor can be mapped to a grayscale image, an RGB (red, green, blue) color image, or an HSV (hue, saturation, brightness) color gamut image, depending on the requirements; finally, the mapped image tensor is used as a single model data sample.
[0071] Therefore, by repeating steps S101 to S105 on a large number of raw current signals collected under normal and arc fault conditions, arc data samples and normal data samples of uniform size are obtained, which constitute the model data samples. All arc data samples are assigned an arc fault label (Acr), and all normal data samples are assigned a normal label (Normal). Without separately labeling the arc occurrence area, a complete training dataset can be formed.
[0072] In some possible implementations, step S105 may include: converting the time-frequency feature matrix into a standard tensor of the same size according to the conversion rules; mapping the standard tensor into a color image, wherein the color image is a grayscale image, an RGB color image, or an HSV color gamut image; and using the color image as a model data sample.
[0073] As an example, if the mapping is selected as a single-channel grayscale image, the time-frequency feature matrix is normalized to values in the [0,1] interval and directly assigned, generating a single-channel tensor with dimensions [1,M,N]. The values within the tensor correspond to the grayscale levels of the image. If the mapping is selected as an RGB color image, the normalized values are simultaneously filled into the R, G, and B channels, generating a three-channel tensor with dimensions [3,M,N]. Color mapping can also be completed by segmenting colors according to the numerical range of the time-frequency feature matrix. If the mapping is selected as an HSV color gamut image, the hue (H) parameter is bound to the normalized values, and fixed thresholds are configured for saturation (S) and brightness (V). Then, the HSV image is converted to RGB format, outputting a three-channel tensor with dimensions [3,M,N].
[0074] Thus, in step S105, the time-frequency feature matrix is mapped to an image format tensor, transforming the time-frequency features into an image input format suitable for the model. Multiple image mapping schemes, including grayscale, RGB, and HSV, are provided to adapt to the training needs of different models. Furthermore, all model data samples only require classification labels, eliminating the need for pixel-level region annotation, reducing the manual annotation cost of dataset creation, and improving the efficiency and accuracy of training dataset creation.
[0075] The training dataset construction method provided in this application first collects the original current signal of the photovoltaic DC circuit for preprocessing, and then obtains an initial spectrum sequence after frame division and frequency domain transformation. Second, based on the characteristic frequency band of the electric arc, characteristic frequency points are extracted from the initial spectrum sequence to construct a characteristic spectrum sequence, and irrelevant interference is filtered out. Then, the characteristic spectrum sequences are stacked according to the preset number of time frames to generate an M×N time-frequency feature matrix. There is no need to force M and N alignment, which reduces data storage and computation consumption and is suitable for edge low-computing-power hardware. Finally, the time-frequency feature matrix is converted into a standard tensor as model data samples. The model data samples only need to be labeled with electric arc fault labels and normal labels, without the need for manual selection of fault areas, which reduces labeling costs and improves sample production efficiency.
[0076] The above provides a detailed explanation of the method for constructing the training dataset for the arc detection model. The following section uses photovoltaic inverter testing as an example to further explain the above method for constructing the training dataset.
[0077] First, a large number of raw current signals under normal and arc fault conditions are collected. The raw current signals under normal conditions are labeled "Normal," and the raw current signals under arc fault conditions are labeled "Acr." Next, the raw current signals are used as initial current signals and preprocessed to remove abnormal noise. Then, the standard current signals are segmented into frames to obtain multiple single-frame signals. Each single-frame signal is then subjected to frequency domain transformation to obtain the corresponding initial spectrum sequence.
[0078] In this process, multiple initial spectra can be stacked in time to obtain multiple time-frequency maps. For example... Figure 3 and Figure 4 As shown in the figure, the horizontal axis represents the signal acquisition time, the vertical axis represents the signal frequency, and the color bar on the right represents the signal amplitude at each time and frequency point in decibels. Figure 3 The Normal multi-frame time-frequency diagram presents the energy distribution pattern of the original current signal under normal operating conditions across all time periods and frequency bands. Figure 4 The Acr multi-frame time-frequency diagram shows the energy distribution pattern of the original current signal under the arc fault condition across the entire time period and frequency band.
[0079] Next, the frequency index range is determined according to the preset arc characteristic frequency band, sampling frequency and frequency domain transformation length. N characteristic frequency points are extracted from the initial spectrum sequence according to the frequency index range to construct a characteristic spectrum sequence. Finally, the characteristic spectrum sequences are aggregated and stacked according to the preset number of time frames M to obtain an M×N time-frequency feature matrix.
[0080] The time-frequency feature matrix corresponding to a single frame signal is converted into a standard tensor in image format according to a preset conversion rule, and used as model data samples. For example... Figure 5 and Figure 6 As shown in the figure, the horizontal axis represents the number of time frames M, the vertical axis represents the characteristic frequency points N, and the color bar on the right represents the signal amplitude corresponding to each frequency point with normalized values. Figure 5 This is a tensor plot of the time-frequency feature matrix corresponding to a normal single-frame signal, presenting a visualization of the M×N time-frequency feature matrix corresponding to a single-frame signal under normal operating conditions. Figure 6 This is a tensor plot of the time-frequency feature matrix corresponding to a single frame of the Acr signal, presenting a visualization of the M×N time-frequency feature matrix corresponding to a single frame of the signal under arc fault conditions.
[0081] Secondly, this embodiment provides a training dataset construction device for an arc detection model, such as... Figure 7 As shown, the training dataset construction device includes a signal preprocessing module, a frequency domain transformation module, a feature extraction module, a feature reconstruction module, and a sample construction module.
[0082] The signal preprocessing module is used to collect the original current signal in the photovoltaic DC circuit as the initial current signal, and to preprocess the initial current signal to obtain the standard current signal.
[0083] The frequency domain transformation module is used to perform frame-by-frame processing on the standard current signal to obtain multiple single-frame signals. The frequency domain transformation is then performed on each single-frame signal to obtain the corresponding initial spectrum sequence.
[0084] The feature extraction module is used to determine the frequency index range based on the preset arc characteristic frequency band, sampling frequency and frequency domain transformation length, and to extract the characteristic frequency points from the initial spectrum sequence according to the frequency index range to obtain the characteristic spectrum sequence.
[0085] The feature reconstruction module is used to aggregate and stack feature spectrum sequences according to a preset number of time frames to obtain a time-frequency feature matrix with the number of time frames as the first dimension and the number of feature frequency points as the second dimension.
[0086] The sample construction module is used to convert the time-frequency feature matrix into a standard tensor according to a preset transformation rule to obtain model data samples, which are used to construct a training dataset for training the arc detection model.
[0087] It is understood that the training dataset construction apparatus provided in the above embodiments is used to implement the training dataset construction method described above. The apparatus and method correspond one-to-one, and each apparatus possesses at least all the beneficial effects of the above method embodiments, which will not be elaborated upon further here.
[0088] Thirdly, embodiments of this application provide an arc detection model, which is trained based on a training dataset, which is obtained through the training dataset construction method described above.
[0089] As an example, the arc detection model can employ a two-dimensional CNN (Convolutional Neural Network) detection model.
[0090] For an M×N input tensor obtained from an asymmetric time-frequency feature matrix, the convolutional kernel size of the model's input layer can be set to kt×kf, where kt≠kf, and kt represents the width of the convolutional kernel and kf represents its height. Simultaneously, the temporal stride can be set to st, and the frequency stride can be set to sf, with the two being independent. Thus, the convolution operation on the input tensor is completed through an asymmetric convolutional kernel and independent time-frequency strides, decoupling the learning of temporal and frequency domain features.
[0091] As another example, the arc detection model can also employ a multi-input CNN detection model. In addition to an M×N input tensor, the model input can include the corresponding single-frame signal and initial spectral sequence. These three elements enter the model through three independent input channels, allowing the model to autonomously learn a multi-dimensional feature fusion strategy. For the M×N input tensor, a non-square convolution kernel and an asymmetric time-frequency stride are used to perform convolution operations, adapting to the transient change characteristics of the arc in the time domain and the broadband energy distribution characteristics of the arc in the frequency domain, thereby improving the accuracy of arc detection.
[0092] Understandably, the input tensor constructed in this application possesses lightweight and asymmetric characteristics, which can effectively compress model input and reduce model computational overhead. Therefore, the arc detection model can be deployed on edge computing devices with limited computing resources, without relying on high-performance servers.
[0093] For example, such as Figure 8 As shown, this arcing detection model can be deployed in the microcontrollers of photovoltaic inverters and energy storage inverters. The microcontroller can directly acquire the raw current signal in the photovoltaic DC circuit, and perform preprocessing, framing, frequency domain transformation, feature frequency point extraction, time-frequency feature matrix construction, and tensor conversion locally. The input tensor is then input into the arcing detection model to detect the presence of arcing faults. This enables real-time on-site arcing monitoring of photovoltaic and energy storage branches.
[0094] It should be noted that the existing arc detection model structure can be used for other parts of the model other than the input layer, and this application does not impose specific restrictions.
[0095] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0096] The above provides a detailed description of the training dataset construction method, apparatus, and model for arc detection model provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A method for constructing a training dataset for an arc detection model, characterized in that, The method includes: The original current signal in the photovoltaic DC circuit is collected as the initial current signal, and the initial current signal is preprocessed to obtain the standard current signal; The standard current signal is divided into frames to obtain multiple single-frame signals. Each single-frame signal is then transformed in the frequency domain to obtain a corresponding initial spectrum sequence. The frequency point index range is determined based on the preset arc characteristic frequency band, sampling frequency, and frequency domain transformation length. The initial spectrum sequence is then truncated based on the frequency point index range to obtain the characteristic spectrum sequence. The feature spectrum sequences are aggregated and stacked according to a preset number of time frames to obtain a time-frequency feature matrix with the number of time frames as the first dimension and the number of feature frequency points as the second dimension; wherein, the number of time frames and the number of feature frequency points are independent of each other; The time-frequency feature matrix is converted into a standard tensor according to a preset conversion rule to obtain model data samples, which are used to construct a training dataset for training the arc detection model.
2. The training dataset construction method according to claim 1, characterized in that, The standard current signal is segmented into multiple single-frame signals, and each single-frame signal is subjected to frequency domain transformation to obtain a corresponding initial spectrum sequence, including: The standard current signal is divided into frames according to a preset frame length and a preset frame shift to obtain multiple single-frame signals. Each single-frame signal is weighted using a window function to obtain the weighted single-frame signal. The initial spectral sequence is obtained by performing a Fast Fourier Transform on the weighted single-frame signal.
3. The training dataset construction method according to claim 2, characterized in that, The method further includes: A sliding window is set according to a preset size, and the sliding window is slid along the standard current signal in a time sequence to sequentially capture local sampling segments of the standard current signal. The instantaneous rate of change of the standard current signal is obtained based on the amplitude difference between adjacent sampling points within the local sampling segment. The instantaneous rate of change is compared with a preset rate of change threshold, and the preset frame length is determined based on the comparison result.
4. The training dataset construction method according to claim 2, characterized in that, The step of performing a Fast Fourier Transform on the weighted single-frame signal to obtain the initial spectrum sequence includes: The weighted single-frame signal is subjected to a Fast Fourier Transform to obtain a complex spectrum sequence; The complex spectrum sequence is split to obtain a phase spectrum sequence and an amplitude spectrum sequence; The amplitude spectrum sequence is expanded to obtain an expanded data sequence, which is a power spectrum sequence or a logarithmic amplitude spectrum sequence. The initial spectrum sequence is any one of the complex spectrum sequence, the amplitude spectrum sequence, and the extended data sequence.
5. The training dataset construction method according to claim 1, characterized in that, The process involves determining a frequency index range based on a preset arc characteristic frequency band, sampling frequency, and frequency domain transform length, and then truncating the initial spectrum sequence based on the frequency index range to obtain a characteristic spectrum sequence, including: Based on the preset upper and lower limits of the arc characteristic frequency band, the sampling frequency, and the frequency domain transformation length, determine the first frequency point index and the second frequency point index corresponding to the arc characteristic frequency band; The frequency point index range is determined based on the first frequency point index and the second frequency point index; Based on the frequency index range, the characteristic frequency points are extracted from the initial spectrum sequence to obtain the characteristic spectrum sequence.
6. The training dataset construction method according to claim 5, characterized in that, The method further includes: The observation frequency band of the electric arc is determined, and the observation frequency band is used as the characteristic frequency band of the electric arc; or, Acquire the current signal of the photovoltaic DC circuit under normal operating conditions, perform a Fast Fourier Transform on the current signal to obtain full-band spectrum data, extract the background noise amplitude corresponding to each frequency band based on the full-band spectrum data, and use the frequency bands with background noise amplitudes below the noise threshold as the arc characteristic frequency bands; or... Multiple non-overlapping target frequency bands are set, and the target frequency bands are used as the characteristic frequency bands of the electric arc.
7. The training dataset construction method according to claim 1, characterized in that, The step of aggregating and stacking the feature spectrum sequences according to a preset number of time frames to obtain a time-frequency feature matrix with the number of time frames as the first dimension and the number of feature frequency points as the second dimension includes: Collect a corresponding number of feature spectrum sequences according to the preset number of time frames; The characteristic spectrum sequence is sorted according to the chronological order; The sorted feature spectrum sequences are concatenated to obtain the time-frequency feature matrix.
8. The training dataset construction method according to claim 1, characterized in that, The step of converting the time-frequency feature matrix into a standard tensor according to a preset transformation rule to obtain model data samples includes: The time-frequency feature matrix is converted into a standard tensor of the same size according to the conversion rule. The standard tensor is mapped to a color image, wherein the color image is a grayscale image, an RGB color image, or an HSV color gamut image; The color image is used as a sample of the model data.
9. A device for constructing a training dataset for an arc detection model, characterized in that, The device includes: The signal preprocessing module is used to acquire the original current signal in the photovoltaic DC circuit as the initial current signal, and to preprocess the initial current signal to obtain the standard current signal. The frequency domain transformation module is used to perform frame-segmentation processing on the standard current signal to obtain multiple single-frame signals, and to perform frequency domain transformation on each single-frame signal to obtain the corresponding initial spectrum sequence. The feature extraction module is used to determine the frequency point index range based on the preset arc characteristic frequency band, sampling frequency and frequency domain transformation length, and to extract the feature frequency points from the initial spectrum sequence according to the frequency point index range to obtain the feature spectrum sequence. The feature reconstruction module is used to aggregate and stack the feature spectrum sequence according to a preset number of time frames to obtain a time-frequency feature matrix with the number of time frames as the first dimension and the number of feature frequency points as the second dimension. The sample construction module is used to convert the time-frequency feature matrix into a standard tensor according to a preset conversion rule to obtain model data samples, which are used to construct a training dataset for training the arc detection model.
10. An arc detection model, characterized in that, The arc detection model is trained based on a training dataset, which is obtained by the training dataset construction method as described in any one of claims 1 to 8.