Photovoltaic DC arc fault fire risk assessment method
By using iterative training and weighted fusion methods, an integrated classifier is constructed, sample weights are dynamically adjusted, and parameter combinations are optimized. This solves the problems of discrimination accuracy and stability in photovoltaic DC arc fault fire risk assessment, and realizes refined risk assessment and quantification of safety standards for photovoltaic systems.
Patent Information
- Application Number
- CN202511375549.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-25
- Publication Date
- 2026-02-06
AI Technical Summary
Existing photovoltaic DC arc fault fire risk assessment methods have low accuracy in the critical energy range, the models tend to favor the majority class, lack special processing for boundary samples, and lack dynamic adjustment of sample weights and reasonable control of the number of iterations during iterative training, resulting in unstable risk level classification.
Multiple basic classifiers are generated through iterative training, and an ensemble classifier is constructed by weighted fusion. The sample weights are dynamically adjusted, a decision boundary function is constructed, the risk level is classified using the ignition probability distribution, and the parameter combination is optimized to reduce misjudgment.
It improves the identification accuracy in high-risk and sensitive areas, reduces the probability of misjudgment due to blurred boundaries, and achieves fine and stable fire risk assessment for different arc power and arcing time conditions. It is adaptable to a variety of actual working conditions and provides a quantitative basis for the performance evaluation and safety standards of arc protection devices for photovoltaic systems.
Smart Images

Figure CN121481206A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of photovoltaic system safety, and in particular to a photovoltaic direct current arc fault risk assessment method. BACKGROUND
[0002] At present, the fire risk assessment of arc fault mostly relies on single threshold judgment method. For example, only according to whether the arc time or arc power exceeds the preset value to judge the risk. However, experiments show that when the arc energy is close to a certain critical value, the ignition and non-ignition phenomenon appears alternately, which leads to the significant decline of the discrimination accuracy of the traditional method in this interval.
[0003] In order to improve the discrimination ability, some methods have tried to introduce machine learning model for classification. However, these existing technologies generally have the following shortcomings: first, the ignition and non-ignition samples are not balanced, and the model is easy to be biased to the majority class; second, the boundary samples of the critical energy interval are not specially processed, which leads to poor performance of the model in the key risk interval; third, there is a lack of effective sample weight dynamic adjustment and reasonable iteration number control in the iterative training process; fourth, the risk level division still relies on fixed threshold, ignoring the continuity of probability distribution, and the results have the problems of large jump and poor stability.
[0004] Therefore, how to improve the discrimination accuracy of boundary samples, ensure the model accuracy on the basis of limited sample quantity and realize the rationality of risk level division in the photovoltaic direct current arc fault scene is a technical problem to be solved. SUMMARY
[0005] The present application provides a photovoltaic direct current arc fault fire risk assessment method.
[0006] In the first aspect, the present application provides a photovoltaic direct current arc fault fire risk assessment method, comprising:
[0007] Obtaining an arc fault basic sample set, each sample in the basic sample set comprising an arc time feature, an arc average power feature and label information corresponding to whether the combustible material is ignited;
[0008] For the basic sample set, an iterative training method is used to generate a plurality of basic classifiers;
[0009] The plurality of basic classifiers are fused according to their classification accuracy to construct an integrated classifier;
[0010] Using the integrated classifier to discriminate the combination input of arc power and arc time, and mapping the output result to an ignition probability distribution;
[0011] Based on the ignition probability distribution, a decision boundary function is constructed, and the arc fault fire risk is divided into different levels.
[0012] Optionally, the obtaining the arc fault basic sample set further comprises:
[0013] Optionally, the obtaining the arc fault basic sample set further comprises:
[0014] Optionally, the obtaining the arc fault basic sample set further comprises:
[0015] Optionally, the obtaining the arc fault basic sample set further comprises:
[0016] Optionally, the obtaining the arc fault basic sample set further comprises:
[0017] Optionally, the generating the plurality of basic classifiers by using the iterative training method further comprises:
[0018] Optionally, the generating the plurality of basic classifiers by using the iterative training method further comprises:
[0019] Optionally, the generating the plurality of basic classifiers by using the iterative training method further comprises:
[0020] Optionally, the generating the plurality of basic classifiers by using the iterative training method further comprises:
[0021] Optionally, the adjusting the sample weight according to the prediction result further comprises:
[0022] Optionally, the adjusting the sample weight according to the prediction result further comprises:
[0023] Optionally, the using the integrated classifier to discriminate the combination of the arc power and the arc time, and mapping the output result as the ignition probability distribution further comprises:
[0024] Optionally, the using the integrated classifier to discriminate the combination of the arc power and the arc time, and mapping the output result as the ignition probability distribution further comprises:
[0025] Optionally, the using the integrated classifier to discriminate the combination of the arc power and the arc time, and mapping the output result as the ignition probability distribution further comprises:
[0026] Optionally, the using the integrated classifier to discriminate the combination of the arc power and the arc time, and mapping the output result as the ignition probability distribution further comprises:
[0027] Optionally, the constructing a decision boundary function based on the ignition probability distribution and dividing the arc fault fire risk into different levels further comprises:
[0028] determining a candidate function form as the decision boundary function according to the combination relationship between different arc burning times and arc powers in the ignition probability distribution;
[0029] adjusting parameters of the candidate function, and obtaining the decision boundary function according to the optimized parameters;
[0030] dividing the region corresponding to the ignition probability into multiple risk level intervals according to the decision boundary function.
[0031] Optionally, the adjusting parameters of the candidate function and obtaining the decision boundary function according to the optimized parameters further comprises:
[0032] setting multiple parameter combinations according to discrete steps in a preset parameter space, and traversing the parameter combinations one by one to cover all possible cases of the parameter space;
[0033] calculating a discrimination effect of each parameter combination in different risk intervals based on the ignition probability distribution, the discrimination effect being used to measure the rationality of the candidate function in dividing high probability regions and low probability regions;
[0034] when calculating the discrimination effect, a penalty value is given to the case that a high probability sample is divided into a low risk region or a low probability sample is divided into a high risk region;
[0035] selecting a parameter combination with the best comprehensive score of the discrimination effect from the parameter space as a final parameter, and constructing the decision boundary function using the final parameter.
[0036] Optionally, the generating multiple base classifiers using the iterative training method for the base sample set further comprises:
[0037] merging the existing base classifiers to obtain a temporary integrated classifier in each iteration;
[0038] continuously monitoring the classification error rate of each temporary base classifier;
[0039] when there is no error rate decrease in a preset continuous iteration round, stopping iteration, and taking the iteration round corresponding to the historical lowest error rate as a final training termination point.
[0040] In a second aspect, the present disclosure provides an electronic device, comprising a processor and a memory connected to the processor in communication;
[0041] The memory stores computer-executable instructions;
[0042] The processor executes the computer-executable instructions stored in the memory to implement the method in the present disclosure.
[0043] In a third aspect, the present disclosure provides a computer-readable storage medium, the computer-readable storage medium storing computer-executable instructions, the computer-executable instructions being executed by a processor to implement the method of the present disclosure.
[0044] The present disclosure has the following advantages compared with the prior art:
[0045] 1) In the sample pretreatment stage, the present disclosure enhances the weight of misjudgment-prone samples near the energy threshold, so that the model can pay more attention to the features in the critical interval during the training process, thereby effectively improving the recognition accuracy in the high-risk sensitive area and reducing the misjudgment probability caused by fuzzy boundaries.
[0046] 2) The present disclosure adopts an iterative training + weighted fusion modeling mechanism: in each round of training, the sample weight is dynamically adjusted according to the classification result, so that the model continuously focuses on the critical interval and misjudgment-prone samples, and gradually forms a group of basic classifiers that are good at different feature regions; on this basis, the basic classifiers are weighted and fused according to the classification accuracy to build a robust integrated discriminant model. This mechanism, on the one hand, continuously corrects the deviation and suppresses overfitting through iteration, thereby improving the recognition accuracy of key boundary samples and the overall generalization ability; on the other hand, it significantly enhances the output stability by offsetting the random fluctuations of a single classifier, and provides a high-confidence input for subsequent mapping of the discriminant score to a continuous ignition probability and construction of a probability-based decision boundary, thereby achieving fine and stable fire risk assessment for different arc power-arc time working conditions.
[0047] 3) Unlike the traditional method of binary classification based on fixed threshold, the present disclosure constructs a risk assessment system by using the mapping probability distribution output by the integrated classifier, and uses a candidate function and parameter optimization method to construct a decision boundary function, thereby achieving dynamic division of risk levels under different combinations of arc power and arc time. This method can adapt to various actual working conditions, providing a quantitative basis for performance evaluation and safety standard formulation of arc protection devices for photovoltaic systems, and has strong engineering practicality. BRIEF DESCRIPTION OF DRAWINGS
[0048] The accompanying drawings, which are incorporated into and form part of the specification, illustrate embodiments consistent with the present disclosure and, together with the specification, serve to explain the principles of the present disclosure.
[0049] Figure 1 A schematic diagram of a photovoltaic direct-current arc fault fire risk assessment method provided by an embodiment of the present disclosure;
[0050] Figure 2 A flowchart of a method for obtaining an arc fault basic sample set is provided for the embodiments of the present disclosure.
[0051] Figure 3 A flowchart of a method for iteratively training to generate multiple basic classification methods is provided for the embodiments of the present disclosure.
[0052] Figure 4 A flowchart of a method for integrating the output results of the classifier to map the ignition probability distribution is provided for the embodiments of the present disclosure.
[0053] Figure 5 A flowchart of a method for constructing a decision boundary function and dividing arc fault fire risk levels is provided for the embodiments of the present disclosure. Now, specific embodiments of the present application will be further described. Figure 5
[0054] The specific embodiments of the present disclosure have been shown by the above-described drawings, and will be described in more detail hereinafter. These drawings and textual descriptions are not intended to limit the scope of the concept of the present disclosure by any means, but to illustrate the concept of the present disclosure to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION
[0055] The present disclosure will be further described below in conjunction with the drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present disclosure, and cannot be used to limit the protection scope of the present disclosure.
[0056] Figure 1 A schematic diagram of a method for evaluating photovoltaic DC arc fault fire risk is provided for the embodiments of the present disclosure. Referring to Figure 1 The specific steps will now be described in detail in conjunction with the present embodiments.
[0057] S100, obtaining an arc fault basic sample set, each sample in the basic sample set including an arc burning time feature, an arc average power feature, and label information corresponding to whether a combustible is ignited.
[0058] In the present embodiment, the construction of the arc fault basic sample set is the premise and basis of the entire method. Each sample in the basic sample set includes three aspects of information: first, an arc burning time feature, used to represent the duration of the arc from generation to extinction; second, an arc average power feature, used to reflect the energy intensity level of the arc during the arc burning process; and third, label information, indicating whether the arc ignites a combustible under the corresponding test conditions, serving as a supervision signal for subsequent classification model training.
[0059] In the specific acquisition process, a plurality of arc current conditions (for example, 3A, 5A, 8A, 12A, 16A, etc.) can be set in an experimental environment, and arc voltage and current signals can be collected; wherein when the arc voltage exceeds a set threshold (for example, 10V), the arc starting time is determined; when the arc current decreases to a preset threshold (for example, 0.25A), the arc ending time is determined. The arc time is determined by the difference between the above-mentioned starting time and ending time, and the arc average power is calculated by the arc voltage and current during the arc time. Based on the above calculation results, combined with the observation of whether the combustible is ignited in the test, the label information of the sample is generated.
[0060] S200, for the basic sample set, an iterative training method is used to generate a plurality of basic classifiers.
[0061] In this embodiment, in order to improve the generalization ability and stability of the model in the arc fault risk discrimination, an iterative training mechanism is introduced for the basic sample set, and a plurality of basic classifiers are gradually generated. Specifically, at the beginning of training, an initial weight is assigned to each sample; then the samples are trained with the current weight distribution to obtain a new basic classifier.
[0062] After each round of training is completed, the basic sample set is predicted by using the basic classifier to identify correctly classified and misclassified samples; wherein the weight of the misclassified sample is improved in the next round, and the weight of the correctly classified sample is correspondingly reduced. Through the above dynamic weight updating mechanism, the subsequent generated basic classifier can pay more attention to the difficult to distinguish samples, thereby making up for the shortcomings of the previous round of classifier.
[0063] Through continuous multiple rounds of iterative training, a plurality of basic classifiers complementary in feature intervals are finally obtained, which provide support for subsequent weighted fusion of the classifiers and stable output of the ignition probability distribution.
[0064] S300, a plurality of the basic classifiers are weighted and fused according to their classification accuracy to construct an integrated classifier;
[0065] In this embodiment, after the iterative training generates a plurality of basic classifiers, in order to avoid the deviation of a single classifier in some conditions, the present application further constructs an integrated classifier by weighted fusion. Specifically, the prediction performance of each basic classifier is first evaluated to obtain its classification accuracy information on the basic sample set. The classification accuracy reflects the overall performance of the classifier in processing the relationship between the arc burning time, the arc average power and the ignition label.
[0066] Subsequently, weights are assigned to each base classifier according to the classification accuracy, and the higher the classification accuracy, the greater the contribution weight of the classifier in the integrated model. The output results of multiple base classifiers are fused into a whole discriminant result through weighting, thereby forming an integrated classifier. The integrated classifier inherits the advantages of multiple base classifiers as a whole, and can realize complementary advantages in different sample regions.
[0067] Through the above fusion step, the integrated classifier has higher stability and robustness compared to any single base classifier. It not only reduces errors caused by individual base classifier bias, but also maintains strong discriminant ability in the critical energy interval, thereby providing high-confidence output results for subsequent ignition probability distribution calculation.
[0068] S400, using the integrated classifier to discriminate the combined input of arc power and arc time, and mapping the output result as an ignition probability distribution;
[0069] In this embodiment, the integrated classifier constructed is used to discriminate new arc fault data. The input data is composed of a feature pair of arc time and arc average power, which is used as the input variable of the integrated classifier. When the integrated classifier discriminates the input, it will output a comprehensive score, which comprehensively reflects the discrimination results of multiple base classifiers on the input and their respective weights.
[0070] In order to convert the comprehensive score into a more intuitive and continuous risk measure, this embodiment further introduces a probability mapping mechanism. Specifically, the discrimination score is converted into an ignition probability distribution after a nonlinear mapping function. The probability distribution has a value range of 0%-100%, reflecting the possibility of igniting combustible materials under different combinations of arc time and arc power.
[0071] S500, based on the ignition probability distribution, constructing a decision boundary function, and dividing arc fault fire risk into different levels.
[0072] In this embodiment, the ignition probability distribution output by the integrated classifier is used as the basis for constructing risk division. First, for the combined input of arc time and arc power, multiple candidate functions are determined on the probability distribution as the form of the decision boundary function, such as the improved energy threshold function. Subsequently, the parameter combinations of the candidate functions are evaluated one by one in the preset parameter space, and the discrimination effect in different regions is calculated based on the ignition probability distribution, and a penalty mechanism is used to avoid high-probability samples being classified into low-risk regions or low-probability samples being classified into high-risk regions. Through this optimization process, the decision boundary function with optimal parameters is finally determined.
[0073] Based on the decision boundary function, the arcing time and the arc power plane can be divided into multiple risk level intervals. For example, the region with an ignition probability less than 20% is determined as a no-fire risk region, the region with 20%-50% is determined as a low risk region, the region with 50%-70% is determined as a medium risk region, and the region with more than 70% is determined as a high risk region. The division results of different risk levels can provide a basis for performance evaluation and safety strategy optimization of the photovoltaic direct current arc fault protection device.
[0074] Figure 2 A method flow diagram for obtaining an arc fault basic sample set is provided for the embodiments of the present disclosure. The specific embodiments of the present application are further described below in combination with Figure 2 , further describe the specific embodiments of the present application.
[0075] S110, sample balancing processing is performed on the basic sample set, so that the number of samples igniting the combustible and the number of samples not igniting the combustible are balanced.
[0076] In the present embodiment, the arc fault basic sample set is obtained from multiple groups of arc ignition tests. Each group of tests is performed under different current conditions, for example, 3 A, 5 A, 8 A, 12 A, and 16 A, the arc voltage and current signals are recorded, and the starting point is determined based on the arc voltage exceeding 10 V, and the ending point is determined based on the arc current falling below 0.25 A, so as to calculate the arcing time and the average arc power. Subsequently, combined with field observation, whether the sample ignites the combustible is labeled.
[0077] However, there are significant deviations in the original sample set directly collected: the proportion of samples not igniting in most conditions is much higher than that of samples igniting. For example, in a typical test, a total of 1000 groups of arc samples are collected, of which about 720 groups are not ignited, accounting for 72%, and only 280 groups are ignited, accounting for 28%. If directly used for model training, it is easy to cause the model to be biased towards predicting "not ignited", thereby misjudging in high-risk conditions.
[0078] Therefore, the present embodiment randomly down-samples the non-ignited samples to reduce the number, so as to realize sample balancing processing and provide stable data input for subsequent iterative training.
[0079] S120, taking the arcing energy of the arcing energy greater than the preset energy threshold but not igniting the combustible in the basic sample set as the maximum energy value.
[0080] In the present embodiment, the arcing energy is calculated by the product of the arcing time and the average arc power. The formula is expressed as: arcing energy = arcing time x average arc power.
[0081] The arc burning time is determined by the duration from the arc voltage rising to the arc ignition threshold (e.g. 10 V) to the arc current falling to the cutoff threshold (e.g. 0.25 A); the average power is calculated by the integral mean of the arc voltage and current in the time interval. Through the arc ignition experiment of combustible materials, it is found that the boundary between ignition and non-ignition is not obvious when the arc energy is about 200 J, and there is a risk of ignition when the arc energy is less than 200 J, and there is a possibility of non-ignition when the arc energy is higher than 200 J, which is very unfavorable for model identification, so the initial weight of these difficult-to-distinguish samples needs to be redesigned. Therefore, in this embodiment, the samples whose arc burning energy exceeds the preset threshold (for example, 200 J) but do not ignite the combustible material are taken as a reference, and the maximum energy value E_max is extracted therefrom. For example, in a typical test, the arc burning energy of some samples reaches 250 J or even exceeds 280 J, but the adjacent cable sheath material is still not ignited. By traversing the non-ignition samples, the maximum arc burning energy is finally determined to be 285 J, so E_max = 285 J.
[0082] In S130, the arc burning energy of the basic sample set that is less than the preset energy threshold but ignites the combustible material is taken as the minimum energy value.
[0083] In this embodiment, the arc burning energy is still calculated by the product of the arc burning time and the average arc power. Unlike S120, this time we are concerned about special samples that have low energy levels but can still cause ignition. In the actual collected basic sample set, some samples have arc burning energy lower than the preset threshold (e.g. 200 J), but still successfully ignited the adjacent combustible material. Such samples reveal the uncertainty of arc fire risk, i.e. even in low energy conditions, a fire may occur due to factors such as concentrated sparks, local high temperature, or flammable material quality. Therefore, this embodiment takes these samples with insufficient energy but ignition as the key selection object, and extracts the minimum arc burning energy value therefrom, defined as the minimum energy value E_min. For example, in a typical test, the arc burning energy of some arc samples is only 120 J or 135 J, but still ignites the PVC sheath material. By traversing these "low energy ignition" samples, the minimum arc burning energy is finally determined to be 118 J, so E_min = 118 J. Through this step, a lower limit boundary can be formed in the energy dimension, i.e. below this value it is almost impossible to ignite, and above this value there is a potential fire risk. Together with the maximum energy value E_max in S120, it forms the energy interval [E_min, E_max] for arc fire risk discrimination, providing key parameter reference for subsequent iterative training and decision boundary construction.
[0084] S140, samples with the arcing energy in the interval between the maximum energy value and the minimum energy value are regarded as misjudgment-prone samples, and the sample weight of the misjudgment-prone samples is increased.
[0085] In the embodiment, the maximum energy value E_max and the minimum energy value E_min have been determined through S120 and S130 respectively. The energy interval [E_min, E_max] between the two values corresponds to the area where the ignition and non-ignition phenomena appear alternately. The samples in this area may exhibit different ignition results at the same or similar energy level, which belongs to the typical boundary fuzzy area.
[0086] For example, in one test, E_min = 118 J and E_max = 285 J. In this interval, some samples ignite the combustible at about 150 J, while others do not ignite even at 250 J. If these samples are treated the same as ordinary samples during model training, the classifier is prone to error learning, resulting in a significant decrease in discrimination accuracy in the critical energy interval.
[0087] Therefore, the embodiment defines the samples in the interval [E_min, E_max] as misjudgment-prone samples, and increases their weight during training, so that the model can pay more attention to these samples during iterative training. The specific strategy is to multiply the weight of the samples in this interval based on the initial weight distribution of the samples, so that the classifier is forced to learn its features more fully in each iteration. In this way, the classifier performs better in distinguishing between ignition and non-ignition samples in the critical energy interval, effectively reducing the misjudgment rate.
[0088] In the embodiment, a specific weight enhancement method is given as follows:
[0089]
[0090]
[0091] wherein E_i is the arc energy value of the i th sample, represents the weight of the i th sample in the first training process.
[0092] Figure 3 The embodiment of the present disclosure provides an iterative training to generate a plurality of basic classification method flowcharts, which are combined with Figure 3 to further illustrate the specific embodiments of the present application.
[0093] S210, training the basic sample set with the current sample weight to obtain a new basic classifier.
[0094] In this embodiment, a classifier template with undetermined parameters is prepared first. The template structure is relatively simple, for example, it can be a single threshold discriminator based on arc burning time or average power, or a two-dimensional discriminator formed by linear combination of time and power. The characteristics of such templates are lightweight structure, which can quickly learn the differences between samples, but the individual accuracy is usually not high.
[0095] The data used for training is the basic sample set after sample balancing and weight adjustment. Initially, all sample weights are the same, for example, each sample weight is about one ten-thousandth of the total. After the preliminary preprocessing step, samples in the critical energy interval are given higher weights, so that these "easily misjudged" samples will be paid more attention during training. For example, the boundary samples, which account for less than 30% of the total samples, may account for more than half of the overall training after weight amplification.
[0096] During the training process, the classifier will try different parameter combinations in turn. For example, if it is a threshold classifier based on power, it will be divided at different power thresholds; if it is a threshold classifier based on time, it will test the division effect at different burning time points. Each parameter combination will bring a certain classification error rate, and since the samples have weights, the calculation of the error rate emphasizes the performance of the boundary samples with higher weights. Finally, a set of parameters with the lowest error rate under the current weight distribution will be selected as the final parameters of the current classifier.
[0097] For example, if it is found during training that when the power threshold is set at about 430 watts, the classification result can correctly distinguish most of the high-weight critical samples, then this threshold will be determined as the optimal parameter for this round of training. The new classifier obtained in this way, although the overall accuracy may be only about 70%, but its performance on samples in the critical interval is obviously better than other parameter combinations.
[0098] The output of this step is a new basic classifier, which has a determined discrimination rule and can predict the combination of input arc burning time and arc power. At the same time, the classification performance of this classifier under the current weight distribution is also recorded, so as to reference its reliability for subsequent fusion.
[0099] S220, using the new basic classifier to predict the basic sample set, and adjusting the sample weight according to the prediction result.
[0100] In this embodiment, the main goal is to identify the samples predicted to be wrong and the samples predicted to be correct, and to increase the weight of the samples predicted to be wrong and to reduce the weight of the samples predicted to be correct.
[0101] Specifically, after a new base classifier is generated from one round of training, it is immediately used to make predictions on the entire base sample set, or a test set, a validation set, or a training set from the base sample set. That is, the samples containing arc time and arc power are input into the classifier again, and the output results are compared with the true labels one by one.
[0102] Through this process, it can be identified which samples are correctly classified and which samples are still misclassified. In particular, attention should be paid to the critical interval samples marked as "easy to misjudge" in the sample preprocessing stage, which often still have a high error rate under the current classifier. At this time, the weights of these misclassified samples are increased, so that they are given higher importance in the next round of training; correspondingly, the weights of the samples that have been correctly classified are decreased, so that their influence is weakened in the next round of training.
[0103] This weight redistribution mechanism causes the subsequent training process to gradually focus on the difficult-to-distinguish samples. In other words, the model will continuously shift its attention from easy-to-classify samples to boundary samples and difficult-to-classify samples in the iteration process, thereby improving the overall generalization ability and robustness.
[0104] After this adjustment, a new sample weight distribution is generated. For example, if some samples with a power of about 450 watts and a time of about one second are repeatedly misclassified in the first round, the weights of these samples may account for more than one-tenth of the total weight in the second round of training, so that the new classifier is more inclined to correct this type of error.
[0105] The final output of this step is the updated sample weight distribution, which will be used as input for the next round of training to provide a basis for the generation of a new classifier.
[0106] S230, use the updated sample weights for the next round of training until the iteration is complete.
[0107] In this embodiment, when the sample weights are updated, the new weight distribution is directly used for the next round of training. Specifically, the new training round still starts with the same base classifier template, but due to the change in sample weights, the classifier will pay more attention to the samples that were misclassified or difficult to distinguish in the previous round during the learning process.
[0108] As the iteration progresses, a new base classifier is generated in each round, and its classification effect on the weighted samples is recorded. In the initial few rounds of training, the classifier may tend to use simple partition rules, such as a low power threshold to correctly identify most samples; but as the weights of difficult samples continue to increase, subsequent classifiers will gradually adjust the discrimination boundary to better distinguish complex situations in the arc energy critical interval.
[0109] This process will continue for multiple iterations, and the result of each iteration includes a new base classifier and an updated sample weight distribution. Eventually, the iteration will terminate under a preset stopping condition, and after the iteration is completed, the system will obtain a set of base classifiers arranged in the order of training, each with a corresponding performance record. These classifiers collectively form the basis for subsequent weighted fusion and provide the necessary components for building an integrated classifier.
[0110] Next, the embodiment also discloses a method for judging the completion of iteration. The idea is that each iteration combines the existing base classifiers to obtain a temporary integrated classifier, and the classification error rate of each temporary base classifier is continuously monitored. When there is no error rate decrease in a preset number of consecutive iterations, the iteration is stopped, and the iteration round corresponding to the historical minimum error rate is taken as the final training termination point.
[0111] Specifically, after the training of each new base classifier is completed, the performance of the classifier is not evaluated separately, but is weighted and fused with all the base classifiers generated before to form a temporary integrated classifier. At this time, the weights during fusion can be allocated according to the classification accuracy of each base classifier. In this way, the overall discrimination ability in the current state can be obtained at any time during training.
[0112] The temporary integrated classifier discriminates the validation set or the reserved samples to calculate the overall error rate of the current round. With the increase of the number of iterations, the error rate will theoretically gradually decrease, but in actual situations, when the iteration exceeds a certain number of times, the model may oscillate or tend to be stable, and at this time, further iteration has limited benefits. When it is found that the error rate in consecutive iterations is not better than the previous minimum value, it is considered that the model has reached the best state, and further iteration will not bring substantial improvement, at which time the early stopping mechanism is triggered. After stopping the iteration, the system will backtrack to the iteration round corresponding to the historical minimum error rate, and the integrated classifier generated in this round is taken as the final result. In this way, the generalization performance of the model can be guaranteed while avoiding the overfitting problem caused by overtraining.
[0113] Figure 4 An integrated classifier output result mapping point ignition probability distribution process schematic diagram is provided for the embodiments of the present disclosure. The specific embodiments of the present application are further described in combination with Figure 4
[0114] S410, the discrimination results of the plurality of base classifiers are weighted and summed according to the fusion weights to obtain a score value for representing the discrimination strength of the arc fault.
[0115] In this embodiment, each base classifier will obtain a fusion weight after training, which depends on the classification accuracy of the classifier. The classification accuracy refers to the proportion of correct samples to the total number of samples in a given sample set, which is used to measure the reliability of the classification result of the classifier. For example, when a certain classifier correctly identifies 85 out of 100 samples, its classification accuracy is 85%. The higher the accuracy, the more reliable the classifier in identifying the risk and non-risk state of the arc ignition, and the greater its weight in the overall fusion.
[0116] In the fusion process, the classification results given by each base classifier for the same input sample are multiplied by their corresponding weights and then summed to obtain a continuous score value. This score value is no longer a single "ignition / non-ignition" label, but a comprehensive reflection of the opinions of all base classifiers and their reliability. The size of the score value can intuitively represent the risk intensity of the arc fault: the larger the value, the more the overall classification result tends to be the risk of ignition; the smaller the value, the more it tends to be the non-risk state.
[0117] Specifically, in the base sample set or other data set, input data containing two input features (arc ignition time, arc power) can be found, and after inputting into the integrated classifier, a score for representing the ignition expectation will be obtained. After selecting a sufficient number of inputs, the ignition expectation score corresponding to various different conditions (arc ignition time, arc power) will be obtained.
[0118] S420, performing nonlinear function mapping on the score value, so that the score value is converted into a probability value with a value range of zero to one.
[0119] The score value obtained by the foregoing weighted sum is essentially a continuous discrimination strength indicator, but the range of this value is not fixed, and the comparability between different samples is limited. In order to make the discrimination result have the ability of probabilistic interpretation, the score value needs to be mapped by a nonlinear function.
[0120] Specifically, the score value is input into a monotonically increasing nonlinear mapping function, so that it is compressed and normalized between 0 and 1. The mapped result can be used as the ignition probability value, and the value closer to 1 indicates a higher possibility of igniting the combustible material; the value closer to 0 indicates a lower risk. Through this conversion, the output of the model not only has an intuitive probability meaning, but also provides a quantitative basis for subsequent risk level classification based on probability distribution.
[0121] Specifically, the present application discloses a calculation method as follows:
[0122]
[0123] wherein h m(x) represents the output of the mth base classifier under this input (e.g., ignite output 1, unignite output -1), a m represents the weight of the mth base classifier. represents the probability of ignition under the input of the arc time and the arc power.
[0124] S430, the probability value is taken as an indicator of the possibility of arc fault igniting combustible, for representing the ignition probability distribution of arc power and arc time under different combination conditions.
[0125] In this embodiment, the probability value obtained by mapping the score value output by the integrated classifier through a nonlinear function is directly taken as an indicator of the ignition possibility, for representing the risk level of arc power and arc time under different combination conditions. The probability value ranges from 0 to 1, and the closer the value is to 1, the higher the possibility of igniting combustible is; the closer the value is to 0, the lower the risk is.
[0126] For example, when the input parameters are arc time 0.05 s and power 80 W, the probability given by the model is about 0.01, indicating that there is almost no risk of ignition; while in the case of arc time 0.20 s and power 1000 W, although the total energy is about 200 J, due to the extremely high instantaneous power, the ignition probability is evaluated as 0.60, which belongs to a high risk. For example, when the arc time is 1.00 s and the power is 200 W, the total energy is also about 200 J, and the output probability of the model is 0.50, which is in a critical state. In contrast, when the arc time is 2.00 s and the power is 100 W, the total energy is still about 200 J, but due to the low power, the local temperature rise is insufficient, and the output probability is only 0.45.
[0127] Further, even in the case of energy less than 200 J, there can be a high probability. For example, when the arc time is 0.30 s and the power is 600 W, the total energy is about 180 J, but the model still gives a probability of 0.65, indicating that high-power short-time arc also has a risk of ignition. In contrast, when the arc time is 1.20 s and the power is 150 W, the total energy is also about 180 J, but the output probability decreases to 0.25, indicating that the risk is significantly reduced.
[0128] In the case of high energy, the probability given by the model is closer to 1. For example, when the arc time is 1.50 s and the power is 500 W, the total energy is about 750 J, and the corresponding probability is 0.90; when the arc time is 2.00 s and the power is 500 W, the total energy is about 1000 J, and the output probability reaches 0.98, which can be basically determined as certain ignition. At the same time, the model also retains the characterization of uncertainty, for example, when the arc time is 0.60 s and the power is 350 W, the total energy is about 210 J, and the output probability is about 0.55, which is only slightly higher than the median, reflecting the uncertainty of the actual ignition process.
[0129] Through the above probability distribution result, a three-dimensional mapping relationship of "arcing time-power-ignition probability" can be constructed, and a decision boundary function can be further fitted on the basis to divide the risk level and give a safety warning.
[0130] Figure 5 A flowchart of a process for constructing a decision boundary function and dividing an arc fault fire risk level is provided for the embodiments of the present disclosure. The specific embodiments of the present application will be further described in combination with Figure 5
[0131] S510, according to the combination relationship of different arcing times and arc powers in the ignition probability distribution, a candidate function form is determined as the decision boundary function.
[0132] In the present embodiment, first, based on the overall characteristics of the ignition probability distribution, the probability change trend under the combination of different arcing times and arc powers is observed. Through the analysis of the distribution result, it can be found that there is a significant difference in ignition risk between high-power short-time arcing and low-power long-time arcing, and the critical region presents a nonlinear characteristic. In order to adapt to this characteristic, the present embodiment proposes several candidate function forms as the decision boundary function, for example, a function with the inverse proportional relationship of arcing power and arcing time as the core, or a segmented function with adjustable parameters, to approximately describe the actual risk distribution situation.
[0133] The setting of the candidate function not only needs to cover the main trend in the probability distribution, but also must retain a certain flexibility, so as to minimize the number of boundary misclassified samples in the subsequent parameter adjustment process. Specifically, the candidate function considers the constraint idea of "high probability samples should not fall into the low risk area, and low probability samples should not fall into the high risk area" when constructed, thereby laying the foundation for the subsequent introduction of the penalty mechanism.
[0134] In an embodiment, the traditional experience 200J is taken as the dangerous dividing line, that is:
[0135]
[0136] Based on this form, a decision boundary function with a similar structure can be improved, such as:
[0137]
[0138] S520, adjusting the parameters of the candidate function, and obtaining the decision boundary function according to the optimized parameters.
[0139] Firstly, a plurality of parameter combinations are set according to discrete steps within a preset parameter space, and the parameter combinations are traversed one by one to cover all possible cases of the parameter space. For example, upper and lower limits and step sizes are set for coefficient parameters, offset parameters, etc. in the function, thereby generating a complete set of candidate parameters. This method can ensure comprehensive traversal of a low-dimensional parameter space and avoid the occurrence of local optimal solutions.
[0140] Then, based on the ignition probability distribution, the discrimination effect of each parameter combination in different risk intervals is calculated, and the discrimination effect is used to measure the rationality of the division of high-probability and low-probability regions by the candidate function. In this process, a penalty mechanism is introduced, and a penalty value is applied to the case where a high-probability sample is divided into a low-risk region or a low-probability sample is divided into a high-risk region. The size of the penalty value is proportional to the degree of deviation, and when a high-probability point is seriously misclassified into a risk-free region, the penalty value increases exponentially. On the contrary, when a low-probability point is divided into a high-risk region, it will also be given a certain penalty. In this way, the optimization process not only focuses on the overall discrimination, but also focuses on key error points. When calculating the discrimination effect, a penalty value is given to the case where a high-probability sample is divided into a low-risk region or a low-probability sample is divided into a high-risk region, so that the discrimination effect comprehensively reflects the accuracy and rationality of the division.
[0141] Finally, the parameter combination with the optimal comprehensive score of the discrimination effect is selected from the parameter space as the final parameter, and the final parameter is used to construct the decision boundary function. Through this method, not only the fitting accuracy of the boundary function as a whole is guaranteed, but also the occurrence of misclassified samples is effectively suppressed, thereby improving the reliability and engineering applicability of the arc fault fire risk grade division.
[0142] In this embodiment, the fire risk of arc fault is divided into four categories, namely: no fire risk, low fire risk, medium fire risk and high fire risk. Among them, no fire risk means that arc fault in this area basically will not cause fire except for the special case of extremely close flammable material distance, and the ignition probability in the area is required to be less than 20% (decision boundary 1). If the arc fault detection algorithm or device can cut off the arc fault in this area, it is considered that it has excellent protection effect on electrical fire and excellent performance. Low fire risk means that arc fault in this area has a certain probability to cause fire, and the ignition probability in the area is required to be between 20% and 50%. If the arc fault detection algorithm or device can cut off the arc fault in this area, it is considered that it has good protection effect on electrical fire and good performance. Medium fire risk means that arc fault in this area has a high probability to cause fire, and the ignition probability in the area is required to be between 50% and 70%. If the arc fault detection algorithm or device can cut off the arc fault in this area, it is considered that it has a certain protection effect on electrical fire and qualified performance. High fire risk means that arc fault in this area will basically cause fire in the presence of flammable materials around, and the ignition probability in the area is required to be higher than 70%. If the arc fault detection algorithm or device can cut off the arc fault in this area, it is considered that it has very limited protection effect on electrical fire and unqualified performance.
[0143] The improved decision boundary function similar to the traditional 200J dangerous line is constructed, the parameter space is shown in Table 1, and the performance of each group of parameters is evaluated one by one, and finally the parameter configuration that can optimize the objective function is selected.
[0144] Table 1 Parameter space
[0145] Parameter Range Parameter Range a [0,500] c [0,5] b [0,5] d [-100,100]
[0146] Taking the parameter determination of decision boundary 1 as an example, when performing grid search, the objective function is to minimize the points whose ignition probability exceeds 20% in the lower left of the decision boundary, and to minimize the points whose ignition probability is less than 20% in the upper right of the decision boundary. The points whose ignition probability exceeds 20% in the lower left and the points whose ignition probability is less than 20% in the upper right are punished, and the punishment degree increases exponentially with the degree that the ignition probability exceeds the threshold, and the calculation formula is as follows:
[0147]
[0148] In the formula, P is the total penalty value; M is the total number of misjudgment data under a certain decision boundary; a is the penalty coefficient, which is 10; p_i is the ignition probability of the misjudgment data point.
[0149] Finally, the parameter with the minimum total penalty is selected as the decision boundary parameter, so as to avoid classifying the points with relatively high ignition probability into the no fire risk area.
[0150] S530, according to the decision boundary function, the region corresponding to the ignition probability is divided into multiple risk level intervals.
[0151] According to the embodiment, the less than 20% interval, the 20% to 50% interval, the 50% to 70% interval, and the greater than 70% interval are respectively positioned as a no-fire risk area, a low fire risk area, a medium fire risk area, and a high fire risk area. To this end, it is necessary to combine the method in S520 to fit three curves to segment the probability distribution graph.
[0152] In a specific embodiment, the parameters in Table 2 are constructed as follows:
[0153] Table 2 Parameter set of decision boundary function
[0154] Number of groups a b c d 1 97.12 0.56 0.15 -30.17 2 143.01 0.58 0.18 -39.53 3 167.86 0.56 0.17 -35.05
[0155] By constructing three groups of decision boundaries with the three groups of parameters shown in Table 2, the fire risk of arc fault is divided into four regions according to the arc burning time and arc power, which correspond to the no-fire risk area, the low fire risk area, the medium fire risk area, and the high fire risk area, respectively.
[0156] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, which can include a processor, a communications interface, a memory, and a communications bus, wherein the processor, the communications interface, and the memory complete mutual communication through the communications bus. The processor can invoke the logic instructions in the memory to execute the soft authorization implementation method based on the configuration software.
[0157] In addition, the logic instructions in the memory described above can be implemented in the form of a software function unit and sold or used as an independent product, which can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present disclosure essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the method described in various embodiments of the present disclosure. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.
[0158] In another aspect, the disclosure also provides a non-transitory computer readable storage medium having stored thereon a computer program, which, when executed by a processor, implements the method for implementing the soft authorization based on configuration software.
[0159] The device embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separated, and the components displayed as units can or can not be physical units, i.e., can be located in one place or distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the embodiment scheme according to actual needs. Those skilled in the art can understand and implement without creative labor.
[0160] From the above description of the embodiments, those skilled in the art can clearly understand that the embodiments can be realized by means of software plus necessary universal hardware platforms, and of course can also be realized by hardware. Based on such understanding, the above technical solutions, essentially or in other words, the part of the prior art that makes a contribution, can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0161] It should be understood that the above embodiments are only used to illustrate the technical solutions of the disclosure, rather than limit them; although the disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements for some technical features therein; and these modifications or replacements do not make the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the disclosure.
Claims
1. A method for assessing the fire risk of photovoltaic DC arc faults, characterized in that, include: Obtain a basic sample set of electric arc faults. Each sample in the basic sample set includes arc time characteristics, average arc power characteristics, and label information indicating whether or not a combustible material has been ignited. For the aforementioned basic sample set, an iterative training method is used to generate multiple basic classifiers; The multiple base classifiers are weighted and fused according to their classification accuracy to construct an ensemble classifier; The integrated classifier is used to discriminate the combined input of arc power and arcing time, and the output is mapped to an ignition probability distribution. Based on the ignition probability distribution, a decision boundary function is constructed, and the risk of arc fault fires is classified into different levels.
2. The method for assessing the fire risk of photovoltaic DC arc faults according to claim 1, characterized in that, The acquisition of the basic sample set of arc faults also includes: Perform sample balancing on the basic sample set to ensure that the number of samples of ignited combustibles is equal to the number of samples of unignited combustibles. The maximum energy value is defined as the arc energy of the basic sample set that is greater than the preset energy threshold but fails to ignite the combustible material. The minimum energy value is defined as the arc energy of a combustible material that is less than a preset energy threshold but still ignites the basic sample set. Samples whose arc energy falls between the maximum and minimum energy values are considered as samples prone to misjudgment, and their sample weights are increased.
3. The method for assessing the fire risk of photovoltaic DC arc faults according to claim 1, characterized in that, The method of generating multiple basic classifiers using iterative training on the basic sample set also includes: A new base classifier is obtained by training on the base sample set with the current sample weights. The base sample set is predicted using a new base classifier, and the sample weights are adjusted based on the prediction results. The updated sample weights are used in the next round of training until the iteration is complete.
4. The method for assessing the fire risk of photovoltaic DC arc faults according to claim 3, characterized in that, The step of adjusting sample weights based on prediction results also includes: Identify samples that are incorrectly predicted and those that are correctly predicted, increase the weight of samples that are incorrectly predicted, and decrease the weight of samples that are correctly predicted.
5. The photovoltaic DC arc fault fire risk assessment method according to claim 3, characterized in that, The step of using the ensemble classifier to discriminate the combined input of arc power and arcing time, and mapping the output to an ignition probability distribution, further includes: The discrimination results of multiple basic classifiers are weighted and summed according to their fusion weights to obtain a score value used to characterize the discrimination strength of arc faults; A nonlinear function mapping is applied to the score value, so that the score value is transformed into a probability value with a range between zero and one. The probability value is used as an index of the likelihood of an arc fault igniting combustibles, and is used to characterize the ignition probability distribution under different combinations of arc power and arc time.
6. The method for assessing the fire risk of photovoltaic DC arc faults according to claim 1, characterized in that, The step of constructing a decision boundary function based on the ignition probability distribution and classifying arc fault fire risks into different levels also includes: Based on the combination relationship between different arcing times and arc power in the ignition probability distribution, the candidate functional forms are determined as decision boundary functions; The parameters of the candidate function are adjusted, and the decision boundary function is obtained based on the optimized parameters; Based on the decision boundary function, the region corresponding to the ignition probability is divided into multiple risk level intervals.
7. The method for assessing the fire risk of photovoltaic DC arc faults according to claim 6, characterized in that, The step of adjusting the parameters of the candidate function and obtaining the decision boundary function based on the optimized parameters further includes: Within a preset parameter space, multiple parameter combinations are set according to the step size, and each parameter combination is traversed one by one to cover all possible cases in the parameter space. Based on the ignition probability distribution, the differentiation effect of each parameter combination in different risk intervals is calculated. The differentiation effect is used to measure the rationality of the candidate function in dividing the high probability region and the low probability region. When calculating the differentiation effect, a penalty value is assigned to cases where a high-probability sample is classified into a low-risk area or a low-probability sample is classified into a high-risk area. The parameter combination with the best overall score for discrimination effect is selected from the parameter space as the final parameter, and the decision boundary function is constructed using the final parameter.
8. The method for assessing the fire risk of photovoltaic DC arc faults according to claim 3, characterized in that, The method of generating multiple basic classifiers using iterative training on the basic sample set also includes: Each iteration merges the existing base classifiers to obtain a temporary ensemble classifier; The classification error rate of each of the temporary base classifiers is continuously monitored. If the error rate does not decrease within a preset number of consecutive iterations, the iteration is stopped, and the iteration corresponding to the lowest historical error rate is taken as the final training termination point.
9. An electronic device, comprising: A processor, and a memory communicatively connected to the processor; characterized in that: The memory stores the instructions that the computer executes; The processor executes computer execution instructions stored in memory, in accordance with the steps of the method according to any one of claims 1-8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the steps of the method according to any one of claims 1-8.