Model training method and device applied to direct current arc discharge detection
By performing multiple data processing and iterative optimization on the time series data of DC arc stretch detection, a high-quality sample set is formed, and the optimal neural network model is selected, which solves the model accuracy problem caused by abnormal data in DC arc stretch detection, and achieves the safe and stable operation of the photovoltaic system.
Patent Information
- Application Number
- CN202510491588.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-18
AI Technical Summary
In the DC arc drawing detection, the abnormal data during data acquisition and labeling process is not effectively screened, resulting in a decrease in the quality of model training, which affects the detection accuracy.
By performing at least two data processing on the time series data related to DC arc-pull and non-pull-pull, an initial sample set is formed, and through the iterative optimization process, the optimal neural network model is selected to improve the accuracy and accuracy of the detection model.
Effectively remove abnormal data and noise interference, improve the accuracy and accuracy of model training, ensure the safe and stable operation of the photovoltaic system, and reduce the false alarm rate and missed alarm rate.
Smart Images

Figure CN120336855A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the technical field of fault arc detection, and in particular, to a model training method and device for DC arc detection. Background Art
[0002] DC arc detection based on a neural network model is crucial for ensuring the safety of photovoltaic systems. When training the model, it often relies on a single model training process, ignoring data quality and the quality of the generated model, resulting in a poor model effect.
[0003] The original data set is prone to anomalies in two stages: Acquisition stage: When collecting data, arc data and normal operation data are collected manually. Different staff have different judgment methods, and abnormal data may be collected during the process of collecting arc data due to the instability of the arc. Annotation stage: During manual annotation, it is inevitable that a large amount of abnormal data is mixed into the data set due to human error or software program loopholes in the annotation software. When using traditional model training schemes, these abnormal data will be introduced into the model training process, resulting in a decrease in the accuracy of the model quality during model training. Moreover, in the prior art, most schemes do not focus on data screening at the data source. Low-quality data will inevitably produce a low-quality model, resulting in poor model accuracy. Summary of the Invention
[0004] In view of this, the present invention provides a model training method and device for DC arc detection, which can improve the accuracy and precision of model training through a feedback optimization process to achieve accurate detection of DC arcs in photovoltaic systems and ensure the safe and stable operation of photovoltaic power generation systems.
[0005] According to one aspect of the present invention, an embodiment of the present invention provides a model training method for DC arc detection, the method comprising:
[0006] Obtaining first time series data related to DC arcing and second time series data related to non-DC arcing; wherein, both the first time series data and the second time series data are composed of a plurality of data frames;
[0007] Performing at least two data processes on the first time series data and the second time series data respectively to obtain processed target first time series data and target second time series data;
[0008] Forming an initial sample set with the target first time series data and the target second time series data;
[0009] Select the initial neural network model as the current detection model, use the initial sample set as the current test set, and test the current detection model using the current test set to obtain the current test output result; wherein, the initial neural network model is a neural network model obtained by training using the initial sample set;
[0010] Determine the next sample set based on the current test output result, train the next neural network model using the next sample set, use the next neural network model as the current detection model, use the next sample set as the current training set, and return to the step of testing the current detection model using the current test set to obtain the current test output result until the current test output result meets the preset requirements to obtain the optimal detection model for arcing detection.
[0011] According to another aspect of the present invention, an embodiment of the present invention further provides a model training device for DC arcing detection, and the device includes:
[0012] A data acquisition module, configured to acquire first time series data related to DC arcing and second time series data related to non-DC arcing; wherein, both the first time series data and the second time series data are composed of multiple data frames;
[0013] A data processing module, configured to perform at least two data processes on the first time series data and the second time series data respectively to obtain processed target first time series data and target second time series data;
[0014] A sample set formation module, configured to form an initial sample set from the target first time series data and the target second time series data;
[0015] A training and testing module, configured to select an initial neural network model as the current detection model, use the initial sample set as the current test set, and test the current detection model using the current test set to obtain the current test output result; wherein, the initial neural network model is a neural network model obtained by training using the initial sample set;
[0016] An iterative optimization module, configured to determine the next sample set based on the current test output result, train the next neural network model using the next sample set, use the next neural network model as the current detection model, use the next sample set as the current test set, and return to the step of testing the current detection model using the current test set to obtain the current test output result until the current test output result meets the preset requirements to obtain the trained optimal detection model, and deploy the optimal detection model to an electronic device for arcing detection.
[0017] According to another aspect of the present invention, an embodiment of the present invention further provides an electronic device, which includes:
[0018] at least one processor; and
[0019] a memory communicatively connected to the at least one processor; wherein,
[0020] the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the model training method for DC arc detection according to any embodiment of the present invention.
[0021] According to another aspect of the present invention, an embodiment of the present invention further provides a computer-readable storage medium, which stores computer instructions for causing a processor to implement the model training method for DC arc detection according to any embodiment of the present invention when executed.
[0022] The above technical solution of the embodiment of the present invention can effectively remove abnormal data and noise interference and improve the quality of the data set by performing at least two data processes on the first time series data and the second time series data respectively and forming an initial sample set; then select the initial neural network model as the current detection model, use the initial sample set as the current test set, test the current detection model with the current test set to obtain the current test output result. On this basis, determine the next sample set based on the current test output result, train the next neural network model with the next sample set, use the next neural network model as the current detection model, use the next sample set as the current test set, and return to the step of testing the current detection model with the current test set to obtain the current test output result until the current test output result meets the preset requirements to obtain the optimal detection model. It can optimize the process through data feedback, improve the accuracy and precision of model training, obtain the optimal detection model, thereby realizing the accurate detection of DC arcs in the photovoltaic system, ensuring the safe and stable operation of the photovoltaic power generation system, and this method does not require additional hardware costs and is economical and practical.
[0023] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0025] Figure 1 Flowchart of a model training method for DC arc detection provided by an embodiment of the present invention;
[0026] Figure 2 Schematic diagram of the test accuracy of a model trained multiple times provided by an embodiment of the present invention;
[0027] Figure 3 Schematic diagram of the false alarm and missed alarm rates of a model trained multiple times provided by an embodiment of the present invention;
[0028] Figure 4 Flowchart of another model training method for DC arc detection provided by an embodiment of the present invention;
[0029] Figure 5 Schematic diagram of the screening process of abnormal files provided by an embodiment of the present invention;
[0030] Figure 6 Flow schematic diagram of yet another model training method for DC arc detection provided by an embodiment of the present invention;
[0031] Figure 7 Flow schematic diagram of still another model training method for DC arc detection;
[0032] Figure 8 Structure block diagram of a model training device for DC arc detection provided by an embodiment of the present invention;
[0033] Figure 9 Structure schematic diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0034] To enable those skilled in the art to better understand the solutions of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention.
[0035] It should be noted that the terms "first", "second", etc. in the description, claims and above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0036] In one embodiment, Figure 1 is a flowchart of a model training method for DC arc detection provided by an embodiment of the present invention. This embodiment is applicable to the situation when training and testing a model for DC arc detection. This method can be executed by a model training device for DC arc detection, and the model training device for DC arc detection can be implemented in the form of hardware and / or software. The model training device for DC arc detection can be configured in an electronic device. As Figure 1 shown, the method includes:
[0037] S110. Obtain first time series data related to DC arcing and second time series data related to non-DC arcing.
[0038] Among them, the first time series data is voltage or current time series data related to DC arcing; the second time series data is voltage or current time series data related to non-DC arcing. In this embodiment, both the first time series data and the second time series data are composed of multiple data frames. Each data frame may include multiple sampling points.
[0039] In this embodiment, the first time series data related to DC arcing may include arcing data artificially marked and manufactured by users, as well as arcing data collected from photovoltaic systems, energy storage systems, etc.; the second time series data related to non-DC arcing includes non-arcing data collected from photovoltaic systems, energy storage systems, etc. In this embodiment, the arcing data is fault data, and the non-arcing data is normal data.
[0040] S120. Perform at least two data processes on the first time series data and the second time series data respectively to obtain processed target first time series data and target second time series data.
[0041] Among them, the target first time series data refers to the time series data obtained by performing at least two data processes on the first time series data; the target second time series data refers to the time series data obtained by performing at least two data processes on the second time series data. In this embodiment, the data process at least includes: multiple abnormal data screening and Fast Fourier Transformation (FFT).
[0042] In this embodiment, the data process for obtaining the target first time series data by performing at least two data processes on the first time series data is the same as that for obtaining the target second time series data by performing at least two data processes on the second time series data; when performing abnormal data screening, the first time series data and the second time series data can be roughly screened according to the data law first, and then the second data screening is performed by calculating the standard deviation and variance. On this basis, it can also be checked whether some sampling points in each data frame exceed the maximum sampling value of the ADC (Analog-to-Digital Converter). If some data points in the data frame exceed the maximum sampling value of the ADC (Analog-to-Digital Converter), it usually causes clipping, that is, the data is truncated to the maximum representable value of the ADC, which will affect the FFT result and introduce harmonic distortion or spectral leakage. Therefore, some sampling points in the data frame that exceed the maximum sampling value of the ADC (Analog-to-Digital Converter) will also be screened out; in some embodiments, since the second-order difference can highlight the trend and mutation of data changes, the second-order difference method can also be used to perform abnormal screening again. In this embodiment, for the time series data after the screening process, FFT or wavelet transform is performed to obtain the corresponding spectral feature data.
[0043] S130. Form an initial sample set from the target first time series data and the target second time series data.
[0044] Among them, the initial sample set may include each data frame in the target first time series data and the corresponding first label, each data frame in the target second time series data and the corresponding second label; among them, the first label represents that the data frame is an abnormal frame, and the second label represents that the data frame is a normal frame.
[0045] In this embodiment, the target first time series data and the target second time series data after data processing form an initial sample set. In some embodiments, the sample set can be divided into a test set and a training set according to a preset ratio; among them, the training set is used for model training and optimization of model quality and data quality, and the test set is used for final model evaluation. The preset ratio in this embodiment can be set according to user requirements. Exemplarily, training set:test set = 7:3, or training set:test set = 8:2, and this embodiment does not limit this here. In other embodiments, the initial sample set can also not be divided, and the initial sample set can be directly used as the initial training set to train the model until the parameters of the model reach the optimal to obtain a trained model, and this embodiment does not specifically limit this here.
[0046] S140. Select the initial neural network model as the current detection model, use the initial sample set as the current test set, and use the current test set to test the current detection model to obtain the current test output result; among them, the initial neural network model is a neural network model obtained by training with the initial sample set.
[0047] Among them, the current detection model at least includes any one of a binary classification model, a convolutional neural network model, and a recurrent neural network model. The convolutional neural network can also be a one-dimensional convolutional neural network or a two-dimensional convolutional neural network. The current training set is the training set at the current iteration, and each round of iteration corresponds to a corresponding current training set. Similarly, due to the existence of the iteration process, each round of iteration corresponds to a corresponding current detection model.
[0048] In this embodiment, a trained neural network model can be selected from a preset model library as the initial neural network model, and this initial neural network model is used as the current detection model, and the initial sample set is used as the current test set. Then, the current test set is used to test the current detection model to obtain the current test output result. In some embodiments, the initial neural network model is the initial neural network model obtained by training with the initial sample set. It should be noted that when using the initial sample set to train the initial neural network model, the initial sample set can be first divided into a test set and a training set. After the division of the test set and the training set, the training set is used to train the initial neural network model. During the training process, the test set can be used to test the initial neural network model. If the score of the data frame in the test result exceeds the set threshold, it is considered that the initial neural network model is trained well, and the training can be stopped to obtain a trained initial neural network model. This trained model can be added to the preset model library for subsequent selection of the initial neural network model from the model library. In other embodiments, it is also possible not to divide the training set and the test set, but directly use the initial sample set as the initial training set to train the initial neural network model until the model parameters of the initial neural network model reach the optimal, so as to obtain a trained initial neural network model, and then add this trained initial neural network model to the preset model library. This embodiment does not make specific restrictions here.
[0049] In this embodiment, when using the initial sample set as the current test set and then using the initial sample set to test the current detection model, the initial sample set can be preprocessed, including low-pass filtering, wavelet denoising, normalization, feature extraction, etc. Then, the preprocessed initial sample set is used as the current test set to test the current detection model. Before inputting to the current test model, each data frame and its corresponding first label in the target first time series data in the input current test set, and each data frame and its corresponding second label in the target second time series data can be divided into multiple data files according to certain requirements, and then input into the current detection model to be tested. The detection model will score each data frame in the input file, determine the data classification result corresponding to each data frame based on the scoring result and the preset binary classification threshold, so as to determine the classification result of the data file to which the data frame belongs, and count the number of abnormal files according to the classification result of the file to which it belongs. The number of abnormal files is used as the current test output result, and then the optimal detection model is determined based on the comparison between the number of abnormal files in the current test output result and the preset expected number reference value of abnormal files.
[0050] S150. Determine the next sample set based on the current test output result, train the next neural network model using the next sample set, use the next neural network model as the current detection model, use the next sample set as the current test set, and return to the step of using the current test set to test the current detection model in S140 to obtain the current test output result, until the current test output result meets the preset requirements, and obtain the optimal detection model for arcing detection.
[0051] Among them, the next sample set is the next sample set used in each next-round iteration process, which can be the sample set in the second-round iteration process, the sample set in the third-round iteration process,..., until the current test output result meets the preset requirements, and the optimal detection model, that is, the detection model that can be deployed, is obtained. It can be understood that each round of iteration corresponds to a corresponding next sample set, and this next sample set will be used as the current test set for subsequent operations. The next neural network model can be understood as the next neural network model selected from the preset model library, which is a neural network model obtained by training with the next sample set. Similarly, in each iteration process, a neural network model obtained by training with the next sample set can be selected as the current detection model. It can be understood that in the second-round iteration process, a neural network model obtained by training with the next sample set can be selected, and in the third-round iteration process, a neural network model obtained by training with the next sample set can be selected, until the current test output result output by the selected neural network model meets the preset requirements.
[0052] In this embodiment, the current test output result can be understood as the classification result of the current detection model for the input sample set, and this classification result can characterize whether the data frame is classified abnormally. The preset requirements include: the number of abnormal files is less than or equal to the preset expected number reference value of abnormal files. Among them, the preset expected number reference value of abnormal files is a reference value set by the user according to requirements, and the smaller this reference value is, the more accurate the finally obtained model is. Exemplarily, the preset expected number reference value of abnormal files is set to 0, 1, or 2.
[0053] In this embodiment, a next sample set is determined based on the current test output result, and a next neural network model is selected. The next neural network model is used as the current detection model, and the next sample set is used as the current test set. Then, return to the step in S140 of using the current test set to test the current detection model to obtain the current test output result, until the current test output result meets the preset requirements, so as to obtain the optimal detection model, and deploy the optimal detection model to the electronic device for arcing detection. In some embodiments, if the number of abnormal files characterized by the current test output result is greater than the preset expected number reference value of abnormal files, the abnormal data frames in the abnormal files are determined and deleted, and the normal data frames in the abnormal files are retained to obtain the target files. The target files and other normal files are re-formed into a new sample set, which is used as the next sample set. Then, a next neural network model obtained by training with the next sample set is selected, the next neural network model is used as the current detection model, and the next sample set is used as the current test set. Return to the step in S140 of using the current test set to test the current detection model to obtain the current test output result, until the current test output result meets the preset requirements; if the number of abnormal files is less than or equal to the preset expected number reference value of abnormal files, the current detection model is directly output as the optimal detection model.
[0054] In this embodiment, the next sample set obtained in each iteration process is used to train the next neural network model selected in this iteration. It should be noted that when using the next sample set to train the next neural network model, the training method used is similar to the training process of the above initial neural network model. Similarly, the next sample set can be first divided into a test set and a training set. After the division of the test set and the training set, the training set is used to train the next neural network model. During the training process, the test set can be used to test the next neural network model. If the score of the data frame in the test result exceeds the set threshold, it is considered that the next neural network model is trained well, and the training can be stopped to obtain the next neural network model, which can be added to the preset model library. In some other embodiments, it is also possible not to divide the training set and the test set, and directly use the next sample set as the next training set to train the next neural network model until the model parameters of the next neural network model reach the optimal, so as to obtain the next neural network model, and then add the next neural network model to the preset model library. This embodiment does not make specific limitations here.
[0055] Exemplarily, without partitioning, it can be as follows: Assume that the number of sample data in the sample set is N. Use this sample set to train a corresponding training model. Then, use this sample set as the test set and input it into the trained model to obtain the corresponding test results. Then, perform anomaly rejection on the test results and reorganize them into a new sample set to start the next round of loop, that is, use this new sample set to train the next model. After training, use the new sample set as the new test set to test the model trained in this iteration to obtain the corresponding test results, and repeatedly iterate until the requirements are met. In the case of partitioning, it can be as follows: Partition the sample set into a training set and a test set. Use the training set in this sample set to train the initial model. During the training process, use the test set to test the model. If the test results are above the preset threshold, stop training to obtain a trained model. Then, use the entire sample set (including the training set and the test set) as the current test set to obtain the corresponding test results. If the test results meet the requirements, retain the tested model. If not, perform anomaly rejection on the test results, reorganize them into a new sample set, then partition the new sample set, select the next model for training. After training, use the new sample set as the new test set to test the model trained in this iteration to obtain the corresponding test results, and repeatedly iterate until the requirements are met.
[0056] Through the above technical solutions of the embodiments of the present invention, by performing at least two data processes on the first time series data and the second time series data respectively and forming an initial sample set, it is possible to effectively remove abnormal data and noise interference, improve the quality of the data set, provide a more accurate and reliable data basis for model training, and help improve the performance of the model; by selecting the initial neural network model as the current detection model, using the initial sample set as the current test set, testing the current detection model with the current test set to obtain the current test output result, on this basis, determining the next sample set based on the current test output result, training the next neural network model with the next sample set, using the next neural network model as the current detection model, using the next sample set as the current test set, and returning to the step of testing the current detection model with the current test set to obtain the current test output result until the current test output result meets the preset requirements to obtain the optimal detection model after training, it is possible to optimize the process through data feedback, effectively reduce the false alarm rate and the missed alarm rate, improve the accuracy and precision of model training, thereby achieving precise detection of DC arcing in the photovoltaic system, ensuring the safe and stable operation of the photovoltaic power generation system, and helping to reduce potential safety hazards and economic losses.
[0057] In one embodiment, the method further includes: during each iteration of training and testing of the current detection model, recording the number of abnormal frames with abnormal classification of data frames; wherein, the abnormal classification of data frames includes: false alarms and missed detections of data frames.
[0058] Determine the accuracy rate of the current detection model under the current training set based on the number of abnormal frames, and generate a test accuracy rate graph.
[0059] In this embodiment, during each iteration of training and testing of the current detection model, record the number of abnormal frames of false alarms and missed detections during the iterative process. Calculate the accuracy rate, false alarm rate, and missed detection rate of the current detection model under the current training set in the current iteration through the number of abnormal frames, and generate corresponding graphs of accuracy rate, false alarm rate, and missed detection rate.
[0060] Exemplarily, for better understanding the generation process of the test accuracy rate graph, Figure 2 FIG. is a schematic diagram of the test accuracy rate of a multi-trained model provided by an embodiment of the present invention. Figure 3 FIG. is a schematic diagram of the test false alarm and missed detection rates of a multi-trained model provided by an embodiment of the present invention.
[0061] In this embodiment, first, data of arcing (Arc) and non-arcing (Normal) within the full current range is collected on-site at the power station and a data set is constructed. Before starting model training, set the reference value of the expected number of abnormal files to 2, and exit the model training process when the number of abnormal files detected by the model is less than or equal to this value. Start model training. After the model training outputs the best model, use the training set to test the model, detect abnormal data frames, record the number of abnormal frames of false alarms (Error_0) and missed detections (Error_1), calculate the accuracy rate (Accrracy) of the model under this training set, and plot a line graph of the test results of the model generated by each training. From Figure 2 , Figure 3 It can be seen that during the first training, due to the existence of cross-abnormalities (data that should be Arc data but appears as Normal data) after data screening, the accuracy rate is low and the false alarm / missed detection rate is high after model testing. After the data screening stage, these cross-abnormal data are screened out, then model training is carried out and tested. The test accuracy rate of the model shows an upward trend and finally stabilizes after multiple cycles. After the 8th training, the number of abnormal files is less than or equal to 2, and the model training terminates. The final result of this embodiment will output a high-quality model and a high-quality data set, use the training-test cycle, and perform feedback optimization based on the number of abnormal files detected by the model test.
[0062] In one embodiment, Figure 4The flowchart of another model training method for DC arc detection provided by an embodiment of the present invention. Based on the above embodiments, this embodiment performs at least two data processes on the first time series data and the second time series data respectively to obtain the processed target first time series data and target second time series data; uses the current test set to test the current detection model to obtain the current test output result; and further refines the determination of the next sample set based on the current test output result.
[0063] As Figure 4 shown, the model training method for DC arc detection in this embodiment may specifically include the following steps:
[0064] S410. Obtain the first time series data related to DC arcing and the second time series data related to non-DC arcing.
[0065] S420. Determine the mean and standard deviation of the data frames included in the first time series data and the second time series data respectively to preliminarily screen out the data frames with abnormal distributions.
[0066] In this embodiment, calculate the mean and standard deviation of the data frames included in the first time series data and the second time series data respectively, and preliminarily screen out the data frames with abnormal distributions through the calculation results of the mean and standard deviation, and retain other data frames with normal distributions.
[0067] S430. For the first time series data and the second time series data after preliminary screening, perform secondary screening on the data frames whose signal amplitudes of the respective corresponding data frames are greater than the maximum ADC sampling value.
[0068] In this embodiment, for the first time series data and the second time series data after preliminary screening, determine whether the signal amplitudes of the respective corresponding data frames of the first time series data and the second time series data are greater than the maximum ADC sampling value. If so, perform secondary screening on the data frames whose signal amplitudes of the respective corresponding data frames are greater than the maximum ADC sampling value, and retain the data frames with signal amplitudes less than the maximum ADC sampling value. Among them, the maximum ADC sampling value is the maximum sampling value in the analog-to-digital converter (ADC), that is, the maximum value of the digital output. It can be determined by the resolution and reference voltage of the ADC.
[0069] S440. For the first time series data and the second time series data after secondary screening, respectively use the second-order difference method to screen out abnormal spikes to obtain the screened first time series data and second time series data.
[0070] Among them, the second-order difference method is the result obtained by performing a difference operation on the first-order difference again, that is, performing a difference on the difference values of two consecutive intervals. In a discrete function, the second-order difference can reveal the acceleration of data change, that is, the change of the change rate.
[0071] In this embodiment, for the first time series data and the second time series data after the secondary screening, the second-order difference method is respectively used to screen out abnormal spikes to obtain the screened first time series data and the second time series data. Specifically, for time series data, the second-order difference can highlight the trend and mutation of data change. The specific calculation process is as follows: Let the data sequence be [x1, x2, x3, …, xn]. First, calculate the first-order difference sequence [y1, y2, …, yn-1], where yi = xi+1 - xi; then calculate the second-order difference sequence [z1, z2, …, zn-2], where zi = yi+1 - yi. By analyzing the second-order difference statistical characteristics of the data during normal operation, a suitable threshold range (for example, 100) is set. If the second-order difference result corresponding to a certain data point exceeds this threshold range, it is determined that this data point is an abnormal spike and is screened out from the data set. After these two steps of screening, a more accurate and high-quality data set for model training is constructed.
[0072] S450. Perform FFT transformation on the screened first time series data and the second time series data respectively to obtain the corresponding target first time series data and target second time series data.
[0073] In this embodiment, the screened first time series data and the second time series data are respectively filtered. For the filtered time series data, FFT transformation is respectively performed to obtain the corresponding target first time series data and target second time series data. The filtering in this embodiment adopts a threshold filtering method based on histogram statistics, and it can also be other filtering methods, which are not limited in this embodiment.
[0074] S460. Form an initial sample set from the target first time series data and the target second time series data.
[0075] S470. Select the initial neural network model as the current detection model and use the initial sample set as the current test set; among them, the initial neural network model is a neural network model obtained by training with the initial sample set.
[0076] S480. Divide each data frame in the target first time series data and the corresponding first label, and each data frame in the target second time series data and the corresponding second label into multiple data files.
[0077] In this embodiment, each data frame and its corresponding first label in the target first time series data in the current test set, and each data frame and its corresponding second label in the target second time series data are respectively divided into multiple data files. In this embodiment, the division of each data frame in the target first time series data and the file division of each data frame in the target second time series data can be divided according to requirements, which can be divided in the same way or in different ways. This embodiment does not limit this here.
[0078] S490. Input each data file into the current detection model to calculate the binary classification scores corresponding to each data frame in each data file, and map each binary classification value through the softmax function to obtain the score results of each data frame.
[0079] In this embodiment, each data file is input into the current detection model to calculate the binary classification scores corresponding to each data frame in each data file, and each binary classification value is mapped through the softmax function to obtain the score results of each data frame. Specifically, for each frame of data Dj (where 0 < j < total number of data frames) in file Fi (where 0 < i < total number of files), the binary classification score output by the model is calculated, and this score is mapped through the softmax function to obtain the data score Sij of this data frame. Set the binary classification threshold as T. If Sij < T, the classification result of this frame of data is 1, otherwise the classification result is 0. For file Fi, assuming the label of this file Fi is classification 0, then the expected result should be that the model test results of all data frames in this file are 0. In this case, this file is normal. However, in the actual test process, some data score results in this file may be determined as classification 1. In this case, this file is determined as an abnormal file.
[0080] S4100. Determine the data classification results corresponding to each data frame according to the score results of each data frame and the preset binary classification threshold.
[0081] Among them, the preset binary classification threshold is the threshold customized by the user. Exemplarily, the binary classification threshold is set as T.
[0082] In this embodiment, the data classification results corresponding to each of the data frames are determined according to the score results of each data frame and the preset binary classification threshold. Among them, the data classification results include: normal data frame classification and abnormal data frame classification. In this embodiment, abnormal data frame classification can include missed reports and false alarms.
[0083] S4110. Determine the classification result of the data file to which the data frame belongs based on the classification results of each data, and count the number of abnormal files according to the classification results of the files to which they belong. Use the number of abnormal files as the current test output result.
[0084] In this embodiment, determining the classification result of the data file to which the data frame belongs based on the classification results of each data can be understood as follows: as long as the classification result of one data frame in the file is an abnormal classification, it can be determined that the file to which the data frame belongs is an abnormal file. Then, count the number of abnormal files and use the number of abnormal files as the current test output result.
[0085] S4120. When the number of abnormal files is greater than the preset reference value of the expected number of abnormal files, determine the abnormal data frames in the abnormal files and delete the abnormal data frames. Retain the normal data frames in the abnormal files to obtain the target file. Re-form a new sample set with the target file and other normal files, and use the new sample set as the next sample set.
[0086] S4130. Train the next neural network model using the next sample set, use the next neural network model as the current detection model, use the next sample set as the current test set, and return to 480 until the number of abnormal files in the current test output result is less than or equal to the preset reference value of the expected number of abnormal files and meets the preset requirements, obtaining the trained optimal detection model for arcing detection.
[0087] In this embodiment, when the number of abnormal files is greater than the preset reference value of the expected number of abnormal files, determine the abnormal data frames in the abnormal files and delete the abnormal data frames. Retain the normal data frames in the abnormal files to obtain the target file. Re-form a new sample set with the target file and other normal files, and use the new sample set as the next sample set. Select the next neural network model, use the next neural network model as the current detection model, use the next sample set as the current test set, and return to 480 until the number of abnormal files in the current test output result is less than or equal to the preset reference value of the expected number of abnormal files and meets the preset requirements, obtaining the optimal detection model, and deploy the optimal detection model to the electronic device for arcing detection.
[0088] Exemplarily, for better understanding the screening process of abnormal files, Figure 5 FIG. is a schematic diagram of a screening process of abnormal files provided by an embodiment of the present invention. As Figure 5 shown, after the data set is screened, a binary classification (Normal - 0 (normal) / Arc - 1 (abnormal)) data set is constructed, and the training set and the test set are divided. Determine whether there is a model in the model library. If there is a model, input the current training set data into the model for testing. The testing process is as Figure 5As shown, the file format in the training set is in csv format. In each file, the data frame is divided by N points, and the total number of frames in each file is n, as in process ①. After reading the data in the file, set the binary classification threshold to T and input it into the model (Model), as in process ②. The model discriminates each frame of the input data and outputs the score of each frame of data (Score: if the label is Arc, then Score < T; if the label is Normal, then Score > T). According to the score results, obtain the abnormal data frames existing in the file (for the Arc label, it should be Score < T, but the result is Score > T; for the Normal label, it should be Score > T, but the result is Score < T), as in process ③. Traverse each frame of the abnormal data frames in the abnormal file, filter out the abnormal frames, and save the new anomaly-free data file in the new csv format, as in processes ④ and ⑤. Add the new csv file to the temporary folder to build a new training set for the next test, as in process ⑥.
[0089] S4140. When the number of abnormal files is less than or equal to the preset reference value of the expected number of abnormal files, directly output the current detection model as the optimal detection model to deploy the optimal detection model to the electronic device for arc detection.
[0090] In this embodiment, from the perspective of model training, the consistent single training method is changed. According to the comparison between the number of abnormal files output by the model during the test of the training set and the expected acceptable number of abnormal files, the number of abnormal files output by the model during the test is gradually made to approach the preset target value through multiple trainings. The new training scheme enables the trained model to exhibit excellent performance in the DC arc detection task of the photovoltaic system. Compared with the model trained by the traditional training method, the model trained by the present invention can effectively reduce the false alarm rate and the missed alarm rate, improve the detection accuracy, and accurately identify the DC arc phenomenon under complex working conditions, providing a reliable guarantee for the safe and stable operation of the photovoltaic system.
[0091] In the above technical solution of the embodiment of the present invention, by initially screening and secondarily screening the data frames with abnormal distributions, and respectively using the second-order difference method to screen out abnormal cusps for the first time series data and the second time series data after the second screening, it is possible to specifically remove the abnormal values and abnormal cusps in the operation data of the photovoltaic system, significantly improve the quality of the data set for model training, and thus provide a solid data foundation for training a high-precision arc detection model. Compared with the traditional data processing method, the data screening method of the present invention can more effectively retain the key information related to arcing, reduce noise interference, and improve the usability of the data; by inputting each data file into the current detection model to calculate the binary classification scores respectively corresponding to each data frame in each data file, and mapping each binary classification value through the softmax function to obtain the score results of each data frame, determining the data classification results respectively corresponding to each data frame according to the score results of each data frame and a preset binary classification threshold, and counting the number of abnormal files, in the case where the number of abnormal files is greater than the preset reference value of the expected number of abnormal files, determining the abnormal data frames in the abnormal files and deleting the abnormal data frames, retaining the normal data frames in the abnormal files to obtain the target file, reforming the target file and other normal files into a new sample set, and using the new sample set as the next sample set; in the case where the number of abnormal files is less than or equal to the preset reference value of the expected number of abnormal files, directly outputting the current detection model as the optimal detection model, which can further iteratively optimize the model performance step by step through multiple iterations, so that the obtained model shows excellent performance in the arc detection task, can adapt to the detection requirements under complex working conditions, significantly improve the accuracy and reliability of arc detection, and thus help to timely discover and handle the DC arcing phenomenon in the photovoltaic system, prevent it from causing more serious consequences, provide a reliable guarantee for the safe and stable operation of the photovoltaic system, and help to reduce potential safety hazards and economic losses.
[0092] In one embodiment, for better understanding of the model training method applied to DC arc detection, Figure 6 FIG. is a schematic flowchart of another model training method applied to DC arc detection provided by an embodiment of the present invention. Figure 7 FIG. is a schematic flowchart of yet another model training method applied to DC arc detection. The embodiment of the present invention proposes a new model training scheme - data feedback control for DC arc detection in a photovoltaic system. Through multiple training and testing processes, the number of abnormal files output by the model test is gradually approximated to a preset target value, and the model output through this process is finally deployed for testing. This technology realizes the precise detection of DC arcing in the photovoltaic system through data processing and model training optimization, ensures the safe and stable operation of the photovoltaic power generation system, and this method does not require additional hardware costs, and is economical and practical.
[0093] AsFigure 6 As shown below, the specific steps are as follows:
[0094] a1. Obtain the first time series data related to DC arc striking and the second time series data related to non-DC arc striking, calculate the mean standard deviation of the data, filter out data frames with abnormal distributions, and preliminarily remove abnormal data points that deviate significantly from the normal operating data distribution, laying a foundation for more accurate subsequent data processing.
[0095] a2. On the basis of the data after the first rough screening, use the second-order difference method for further processing. For time series data, the second-order difference can highlight the trend and mutation of data changes.
[0096] a3. Calculate the N-point FFT transform of the data frame to obtain the corresponding frequency spectrum data.
[0097] a4. Construct a training set and a test set from the screened data according to a preset ratio (e.g., training set: test set = 7:3). The training set is used for model training, model quality and data quality optimization, and the test set is used for final model evaluation.
[0098] Model training: The model used is a binary classification model. During the training process, the current model checkpoint will be saved in each iteration cycle to ensure that the training is terminated due to power outages or program exceptions and other emergencies. And the trained model will be saved during training. After the training is completed, the model will be added to a pre-set model library. Thus, during each iteration training and testing of the current detection model, the accuracy rate under the current training set is determined according to the number of abnormal frames classified by the data frame, and a test accuracy rate graph is generated.
[0099] a5. Select a model from the model library and use the sample set (training set and test set) to test the model: Model testing is used for data quality optimization. In this process, the sample set data is input into the model for testing. For each frame of data Dj (where 0 < j < total number of data frames) in file Fi (where 0 < i < total number of files), the binary classification score output by the model will be calculated, and this score will be mapped to 0 - 65535 through the softmax function to obtain the score Sij of this frame of data. Set the binary classification threshold as T. If Sij < T, the classification result of this frame of data is 1, otherwise the classification result is 0. For file Fi, assuming the label of this file is classification 0, then the expected result should be that the model test results of all data frames in this file are 0. In this case, this file is normal. However, in the actual test process, some data score results in this file may be determined as classification 1. In this case, this file is determined as an abnormal file.
[0100] a6. Abnormal file screening: Count the number of abnormal files and compare it with the reference value of the expected number of abnormal files (assumed to be 0). If the number of abnormal files is greater than the reference value, discard the existing model.
[0101] a7. Input the abnormal files into the abnormal file screening program, directly locate the abnormal data frames in the files, and directly delete the abnormal data, retaining the normal data in the abnormal files.
[0102] a8. Rebuild the next sample set, return to step a5 and continue to loop until a model that meets the requirements is obtained.
[0103] As Figure 7 shown, the model training method applied to DC arcing detection specifically includes the following steps:
[0104] b1. Build a data set.
[0105] b2. Use the data set as the current training set.
[0106] In this embodiment, the current training set is the initial sample set in the above embodiment.
[0107] b3. Determine whether there is a trained model in the preset model library. If not, execute b4; if so, execute b6.
[0108] b4. Train the initial model using the current training set.
[0109] b5. Add the trained model to the model library.
[0110] b6. Select a trained model from the preset model library as the initial neural network model, and use this initial neural network model as the current detection model.
[0111] b7. Use the current training set as the current test set, and use the current test set to test the current detection model to obtain the current test result.
[0112] b8. Count the number of abnormal files.
[0113] b9. Determine whether the number of abnormal files is less than or equal to the preset reference value of the expected number of abnormal files. If so, execute b10; if not, execute b12.
[0114] b10. Store the result of the current detection model in the database so as to draw a line chart of the test results of the models generated by each training.
[0115] b11. Determine the abnormal data frames in the abnormal file, delete the abnormal data frames, retain the normal data frames in the abnormal file to obtain the target file, reform the target file and other normal files into a new sample set, use the new sample set as the next sample set, first use the next sample set as the current training set respectively, return to b2 for loop, and then use the current training set as the current test set.
[0116] In this embodiment, each time a new round of loop starts, the selected model is the next neural network model, and the next neural network model is used as the current detection model for subsequent operations; the next neural network model is trained by the corresponding next sample set.
[0117] b12. Directly output the current detection model as the optimal detection model, so as to deploy the optimal detection model to the electronic device for arcing detection.
[0118] In one embodiment, Figure 8 is a structural block diagram of a model training device for DC arcing detection provided by an embodiment of the present invention. The device is applicable to the situation when training a model for DC arcing detection, and the device can be implemented by hardware / software. It can be configured in an electronic device to implement a model training method for DC arcing detection in an embodiment of the present invention. As Figure 8 shown, the device includes: a data acquisition module 810, a data processing module 820, a sample set formation module 830, a training and testing module 840, and an iterative optimization module 850.
[0119] Among them, the data acquisition module 810 is used to acquire the first time series data related to DC arcing and the second time series data related to non-DC arcing; wherein, both the first time series data and the second time series data are composed of multiple data frames;
[0120] The data processing module 820 is used to perform at least two data processes on the first time series data and the second time series data respectively to obtain the processed target first time series data and target second time series data;
[0121] The sample set formation module 830 is used to form an initial sample set from the target first time series data and the target second time series data;
[0122] The training and testing module 840 is used to select an initial neural network model as the current detection model, and use the initial sample set as the current test set, and test the current detection model with the current test set to obtain the current test output result; wherein, the initial neural network model is a neural network model obtained by training with the initial sample set;
[0123] An iterative optimization module 850 is used to determine the next sample set based on the current test output result, train the next neural network model using the next sample set, use the next neural network model as the current detection model, use the next sample set as the current test set, and return the step of testing the current detection model using the current test set to obtain the current test output result until the current test output result meets the preset requirements, so as to obtain an optimal detection model for arc striking detection.
[0124] In the embodiment of the present invention, by performing at least two data processes on the first time series data and the second time series data respectively and forming an initial sample set, abnormal data and noise interference can be effectively removed, and the quality of the data set can be improved; then an initial neural network model is selected as the current detection model, the initial sample set is used as the current test set, and the current test set is used to test the current detection model to obtain the current test output result. On this basis, the next sample set is determined based on the current test output result, the next neural network model is trained using the next sample set, the next neural network model is used as the current detection model, the next sample set is used as the current test set, and the step of testing the current detection model using the current test set to obtain the current test output result is returned until the current test output result meets the preset requirements, so as to obtain an optimal detection model, which can optimize the process through data feedback, improve the accuracy and precision of model training, obtain an optimal detection model, thereby realizing accurate detection of DC arc striking in a photovoltaic system, ensuring the safe and stable operation of the photovoltaic power generation system, and this method does not require additional hardware costs, and is economical and practical.
[0125] In one embodiment, the data process at least includes: abnormal data screening and FFT Fourier transform; correspondingly, the data process module 820 includes:
[0126] A preliminary screening unit is used to determine the mean and standard deviation of the data frames included in the first time series data and the second time series data respectively, so as to preliminarily screen the data frames with abnormal distributions.
[0127] A secondary screening unit is used to perform secondary screening on the data frames with signal amplitudes greater than the maximum ADC sampling value of the corresponding data frames for the first time series data and the second time series data after preliminary screening.
[0128] A cusp screening unit is used to respectively adopt a second-order difference method to screen out abnormal cusps for the first time series data and the second time series data after secondary screening to obtain the screened first time series data and second time series data.
[0129] A feature transformation module for performing FFT transformation on the screened first time series data and second time series data respectively to obtain corresponding target first time series data and target second time series data.
[0130] In one embodiment, the current detection model at least includes any one of a binary classification model, a convolutional neural network model, and a recurrent neural network model.
[0131] In one embodiment, the current test set includes each data frame in the target first time series data and the corresponding first label, and each data frame in the target second time series data and the corresponding second label; correspondingly, the training and testing module 840 includes:
[0132] A file partitioning unit for partitioning each data frame in the target first time series data and the corresponding first label, and each data frame in the target second time series data and the corresponding second label into multiple data files respectively;
[0133] A result determination unit for inputting each of the data files into the current detection model to calculate the binary classification scores respectively corresponding to each data frame in each of the data files, and mapping each of the binary classification values through a softmax function to obtain the score results of each of the data frames;
[0134] A classification unit for determining the data classification results respectively corresponding to each of the data frames according to the score results of each of the data frames and a preset binary classification threshold; wherein, the data classification results include: data frame classification normal and data frame classification abnormal;
[0135] A test output unit for determining the classification results of the data files to which the data frames belong according to each of the data classification results, and counting the number of abnormal files according to the classification results of the files to which they belong, and taking the number of abnormal files as the current test output result.
[0136] In one embodiment, the preset requirement includes: the number of abnormal files is less than or equal to a preset expected number reference value of abnormal files; the iterative optimization module 850 includes:
[0137] An abnormal frame deletion unit, in the case where the number of abnormal files is greater than the preset expected number reference value of abnormal files, determining the abnormal data frames in the abnormal files and deleting the abnormal data frames, retaining the normal data frames in the abnormal files to obtain target files, re-forming the target files and other normal files into a new sample set, and taking the new sample set as the next sample set;
[0138] A model output unit, configured to directly output the current detection model as the optimal detection model when the number of the abnormal files is less than or equal to a preset reference value of the expected number of abnormal files.
[0139] In one embodiment, the apparatus further includes:
[0140] A recording module, configured to record the number of abnormal frames with data frame classification anomalies during each iteration training and testing of the current detection model; wherein, the data frame classification anomalies include: data frame false alarms and missed reports;
[0141] A generating module, configured to determine the accuracy rate of the current detection model under the current training set according to the number of abnormal frames, and generate a test accuracy rate graph.
[0142] The model training apparatus for DC arcing detection provided by the embodiments of the present invention can execute the model training method for DC arcing detection provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method.
[0143] In one embodiment, Figure 9 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. The electronic device 10 is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device (such as a helmet, glasses, a watch, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0144] As Figure 9 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., wherein, the memory stores a computer program executable by at least one processor, and the processor 11 can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.
[0145] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0146] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the model training method applied to DC arcing detection.
[0147] In some embodiments, the model training method applied to DC arcing detection can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the model training method applied to DC arcing detection described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the model training method applied to DC arcing detection by any other suitable means (e.g., by means of firmware).
[0148] The various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0149] A computer program for implementing the method of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable model training devices for DC arc detection, such that when the computer programs are executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer programs can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.
[0150] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0151] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0152] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.
[0153] A computing system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs that run on respective computers and have a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0154] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.
[0155] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A model training method applied to DC arc detection, characterized in that, The described model training method includes: Obtaining first time series data related to DC arcing and second time series data related to non-DC arcing; wherein, both the first time series data and the second time series data are composed of multiple data frames; Performing at least two data processes on the first time series data and the second time series data respectively to obtain processed target first time series data and target second time series data; Forming an initial sample set from the target first time series data and the target second time series data; Selecting an initial neural network model as the current detection model, using the initial sample set as the current test set, and testing the current detection model with the current test set to obtain a current test output result; wherein, the initial neural network model is a neural network model trained using the initial sample set; Determining the next sample set based on the current test output result, training a next neural network model using the next sample set, using the next neural network model as the current detection model, using the next sample set as the current test set, and returning to the step of testing the current detection model with the current test set to obtain a current test output result until the current test output result meets the preset requirements to obtain an optimal detection model for arcing detection.
2. The method according to claim 1, wherein The data process at least includes: abnormal data screening and FFT Fourier transform; Performing at least two data processes on the first time series data and the second time series data respectively to obtain processed target first time series data and target second time series data, including: Determining the mean and standard deviation of the data frames included in the first time series data and the second time series data respectively to preliminarily screen out data frames with abnormal distributions; For the first time series data and the second time series data after preliminary screening, performing secondary screening on the data frames in which the signal amplitude of each corresponding data frame is greater than the maximum ADC sampling value; For the first time series data and the second time series data after secondary screening, respectively using the second-order difference method to screen out abnormal sharp points to obtain the screened first time series data and second time series data; Performing FFT transforms on the screened first time series data and second time series data respectively to obtain the corresponding target first time series data and target second time series data.
3. The method according to claim 1, wherein The current detection model at least includes any one of: a binary classification model, a convolutional neural network model, and a recurrent neural network model.
4. The method according to claim 1, wherein The current test set includes: each data frame in the target first time series data and the corresponding first label, each data frame in the target second time series data and the corresponding second label; Correspondingly, testing the current detection model with the current test set to obtain a test output result includes: Divide each data frame in the target first time series data and its corresponding first label, and each data frame in the target second time series data and its corresponding second label into multiple data files respectively; Input each of the data files into the current detection model to calculate the binary classification scores corresponding to each data frame in each data file, and map each of the binary classification values through the softmax function to obtain the score results of each data frame; Determine the data classification results corresponding to each data frame according to the score results of each data frame and a preset binary classification threshold; wherein, the data classification results include: data frame classification normal and data frame classification abnormal; Determine the classification results of the data files to which the data frames belong according to each of the data classification results, and count the number of abnormal files according to the classification results of the files to which they belong, and use the number of abnormal files as the current test output result.
5. The method according to claim 4, wherein The preset requirement includes: the number of abnormal files is less than or equal to a preset expected number reference value of abnormal files; Correspondingly, the determining the next sample set based on the current test output result includes: In the case that the number of abnormal files is greater than the preset expected number reference value of abnormal files, determine the abnormal data frames in the abnormal files and delete the abnormal data frames, retain the normal data frames in the abnormal files to obtain target files, and re-form a new sample set from the target files and other normal files, and use the new sample set as the next sample set; In the case that the number of abnormal files is less than or equal to the preset expected number reference value of abnormal files, directly output the current detection model as the optimal detection model.
6. The method according to claim 1, wherein The method further includes: during each iterative training and testing of the current detection model, record the number of abnormal frames with abnormal data frame classification; wherein, the abnormal data frame classification includes: false alarms and missed detections of data frames; Determine the accuracy rate of the current detection model under the current training set according to the number of abnormal frames and generate a test accuracy rate graph.
7. A model training device applied to DC arc striking detection, characterized in that, The device includes: A data acquisition module, configured to acquire first time series data related to DC arcing and second time series data related to non-DC arcing; wherein, both the first time series data and the second time series data are composed of multiple data frames; A data processing module, configured to perform at least two data processings on the first time series data and the second time series data respectively to obtain processed target first time series data and target second time series data; A sample set formation module, configured to form an initial sample set from the target first time series data and the target second time series data; A training and testing module, configured to select an initial neural network model as the current detection model, and use the initial sample set as the current test set, and test the current detection model with the current test set to obtain a current test output result; wherein, the initial neural network model is a neural network model obtained by training with the initial sample set; An iterative optimization module, configured to determine a next sample set based on the current test output result, train a next neural network model using the next sample set, use the next neural network model as the current detection model, use the next sample set as the current test set, and return the step of testing the current detection model using the current test set to obtain a current test output result, until the current test output result meets a preset requirement, to obtain an optimal detection model for arcing detection.
8. An electronic device, characterized in that, The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the model training method for DC arcing detection according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a processor to execute the model training method for DC arcing detection according to any one of claims 1-6 when executed.
10. A computer program product, characterized in that, The computer program product includes a computer program, and the computer program implements the model training method for DC arcing detection according to any one of claims 1-6 when executed by a processor.
Citation Information
Cited By
Arcing detection method and device of inverter, electronic equipment and storage medium
CN120928137A
Arc detection method and device of inverter, electronic equipment and storage medium
CN120928137B
Arcing detection method and device and electronic equipment
CN122241459A