Pipe network burr data detection method based on Focal Loss improved LightGBM and memory
By introducing FocalLoss loss function and model fusion technology in the LightGBM model, the problem of low accuracy of glitch data detection in the existing technology is solved, and a higher recall and accuracy rate is achieved, which improves the overall effect of glitch data detection in the pipeline network.
Patent Information
- Application Number
- CN202411741510.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-29
- Publication Date
- 2025-05-30
AI Technical Summary
In the prior art, glitch data detection methods based on statistical features are difficult to effectively identify multiple fault factors, and detection based on artificial intelligence models can easily lead to overfitting of the model and reduce accuracy.
The pipeline network glitch data detection method based on FocalLoss is adopted to improve LightGBM. The glitch data set is constructed by extracting neighborhood features, multiple model verifications are carried out and the loss function parameters are adjusted, and finally the model fusion is carried out to improve the detection effect.
By adjusting the loss function parameters and model fusion, the recall and accuracy can be better taken into account in the detection of pipeline network glitch data, improving the overall detection effect, and reducing the risk of model overfitting.
Smart Images

Figure CN120067667A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a pipeline network burr data detection method and a memory based on FocalLoss improved LightGBM. Background Art
[0002] In the current era of big data, data has become an important basis for decision-making in enterprises and organizations. However, in the process of data collection, burr data, that is, abnormal values or error values, are often generated. Behind these burr data, there may be key information hidden, which is of great significance for in-depth understanding and analysis of data. At the same time, the generation of these abnormal data may also lead to serious consequences, such as pipeline leakage, pipeline burst and other problems. These problems have brought huge economic losses to enterprises, and also led to the waste of water resources, affecting the normal water supply of society. Therefore, the detection of burr data has become an indispensable key step in the process of data cleaning and preprocessing.
[0003] In the prior art, there are usually two types of methods for detecting burr data: a method based on statistical features and a detection method based on artificial intelligence technology.
[0004] For example, Chinese patent CN202310984887.1 discloses a method, system, equipment and medium for detecting glitch data of water flow, the method comprising obtaining water flow data, processing the water flow data according to a time series to obtain change value data of the water flow data, and obtaining an average fluctuation coefficient according to the change value data; processing the water flow data and the average fluctuation coefficient according to a time series to obtain the fluctuation multiple of the water flow data; and determining the flow glitch data according to the fluctuation multiple. By performing time series processing and fluctuation multiple calculation on the water flow data, the glitch data in the time series data of the pipe network flow, pressure, and water quality equipment can be detected with high accuracy. Find the glitch data and repair it in a targeted manner to significantly reduce the false alarm rate of the monitoring point.
[0005] For another example, in Chinese Patent CN202110338988.2, a method and device for detecting abnormal electricity consumption in power big data based on BRB and LSTM models. The method specifically includes the following steps: extracting the user electricity consumption fluctuation characteristics and the abnormal characteristics of the user electricity consumption curve from the big data of electricity consumption; establishing a belief rule-based reasoning BRB system to perform confidence conversion on the sum of the electricity fluctuation coefficient and the burr width; according to the belief rule base in the belief rule-based reasoning BRB system, using the evidence reasoning ER algorithm to compare the converted confidence degrees to obtain the trust degree of each reference value in the output result of the non-technical loss NTL abnormality of the user; calibrating the abnormal electricity consumption of the non-technical loss NTL of the user; establishing a long short-term memory LSTM model based on the calibrated data, and using the LSTM model to effectively extract and detect the abnormal electricity consumption characteristics, and finally accurately diagnosing the NTL abnormal situation. It can effectively identify abnormal electricity consumption situations.
[0006] However, in the actual implementation process, the inventor found that the detection methods designed based on statistical features usually can only effectively discriminate a certain type of burr data with significant features. In the actual online scenario, more fault factors need to be configured with corresponding statistical discrimination strategies respectively. For the same reason, during the training process of the artificial intelligence model, due to the large number of fault factors, and the burr data itself is an accidental data mutation situation in a large amount of normal data, the number of data samples available for training the model is also small. If random oversampling is used, it is easy to cause the problem of model overfitting, further reducing the accuracy. Summary of the Invention
[0007] In view of the above problems existing in the prior art, a method for detecting burr data in a pipeline network based on Focal Loss-improved LightGBM is provided.
[0008] The specific technical solution is as follows:
[0009] A method for detecting burr data in a pipeline network based on Focal Loss-improved LightGBM includes:
[0010] Step S1: Extracting neighborhood features from the original pipeline network data and constructing a burr data set;
[0011] Step S2: Based on the burr data set, verifying the detection model multiple times, and adjusting the loss function parameters of the loss function of the detection model during each verification process;
[0012] Step S3: Selecting multiple detection models according to the verification results to form a fusion model, and using the fusion model to identify the pipeline network data.
[0013] On the other hand, the step S1 includes:
[0014] Step S11: Manually annotate the original pipeline network data to obtain data labels;
[0015] The data labels include burr data and normal data;
[0016] Step S12: Extract neighborhood features from the burr data and the normal data respectively;
[0017] Step S13: Generate the burr data set according to the neighborhood features and the data corresponding to the data labels.
[0018] On the other hand, in the step S12, the neighborhood features include neighborhood values, unilateral neighborhood differences, and bilateral neighborhood differences.
[0019] On the other hand, after executing the step S12 and before executing the step S13, it further includes:
[0020] Step A13: Classify the burr data and the normal data according to time partitions to form time annotations;
[0021] In the step S13, the time annotations are further added to the burr data set.
[0022] On the other hand, the step S2 includes:
[0023] Step S21: Set the parameter adjustment range of the loss function parameters, and generate multiple groups of loss function parameter combinations within the parameter adjustment range;
[0024] Step S22: Adjust the detection model according to the loss function parameter combinations to obtain a model to be verified;
[0025] Step S23: Verify the model to be verified by using the burr data set to obtain the verification result.
[0026] On the other hand, the loss function is:
[0027] FL(p t )=-α t (1-p t ) γ log(p t );
[0028] In the formula, FL(p t ) is the loss function, p t is the probability of the occurrence of the burr data, γ is the adjustment factor parameter, and α is the weight factor.
[0029] On the other hand, in the step S21, the adjustment factor parameter and the weight factor are used as the loss function parameters.
[0030] On the other hand, the detection model is a LightGBM model.
[0031] On the other hand, the step S3 includes:
[0032] Step S31: Determine evaluation indicators from the verification results, and select the detection model under the corresponding loss function parameters according to the evaluation indicators;
[0033] Step S32: Fuse the detection models to obtain the fused model.
[0034] A memory, the memory includes computer instructions, when a computer device runs the computer instructions, the above pipeline burr data detection method is executed.
[0035] The above technical solution has the following advantages or beneficial effects:
[0036] Aiming at the problems of unbalanced burr data samples and poor model training effect in the prior art, in this solution, a verification process of pre-verifying the loss function and adjusting the loss function parameters is introduced. By adjusting the parameters of the loss function and performing multiple rounds of verification, it is possible to determine a parameter combination that better balances the recall rate and the precision rate during the training of a specific burr data set, and finally perform fusion, thereby improving the overall detection effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] With reference to the accompanying drawings, the embodiments of the present invention are described more fully. However, the accompanying drawings are only for illustration and explanation, and do not constitute a limitation to the scope of the present invention.
[0038] Figure 1 It is a schematic diagram of the whole of the embodiment of the present invention;
[0039] Figure 2 It is a schematic diagram of step S1 in the embodiment of the present invention;
[0040] Figure 3 It is a schematic diagram of step A13 in the embodiment of the present invention;
[0041] Figure 4 It is a schematic diagram of step S2 in the embodiment of the present invention;
[0042] Figure 5 It is the Focal Loss curve under different adjustment factors in the embodiment of the present invention;
[0043] Figure 6 It is a schematic diagram of step S2 in the embodiment of the present invention;
[0044] Figure 7Schematic diagram of the improvement effect of the model precision rate relative to the baseline under different parameters in the embodiments of the present invention;
[0045] Figure 8 Schematic diagram of the improvement effect of the model recall rate relative to the baseline under different parameters in the embodiments of the present invention;
[0046] Figure 9 Schematic diagram of the improvement effect of the harmonic mean of Focal Loss under different parameters in the embodiments of the present invention;
[0047] Figure 10 Schematic diagram for comparing the improvement effects of the preferred parameters in the embodiments of the present invention;
[0048] Figure 11 Schematic diagram for comparing the effects of the fusion model and the base model in the embodiments of the present invention. Detailed implementation manners
[0049] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0050] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.
[0051] Next, the present invention will be further described in conjunction with the accompanying drawings and specific embodiments, but it is not limited to the present invention.
[0052] The present invention includes:
[0053] A pipeline network burr data detection method based on improved LightGBM with Focal Loss, as Figure 1 shown, includes:
[0054] Step S1: Extract neighborhood features from the original pipeline network data and construct a burr data set;
[0055] Step S2: Verify the detection model multiple times based on the burr data set, and adjust the loss function parameters of the loss function of the detection model during each verification process;
[0056] Step S3: Select multiple detection models according to the verification results to form a fusion model, and use the fusion model to identify the pipeline network data.
[0057] Specifically, to address the problem of unbalanced burr data samples and poor model training effect in the prior art, in this solution, a verification process for pre-verifying the loss function and adjusting the loss function parameters is introduced to improve the recall rate and precision rate during actual prediction.
[0058] Specifically, for the pipe network to be predicted, a large amount of original pipe network data is collected according to different pipe network types, such as secondary water supply pipe networks, domestic sewage pipe networks, etc. This type of original data usually contains time series data in multiple different dimensions, such as dimensions of pipe network flow, pipe network pressure, pollutant concentration, etc., and is arranged in chronological order to form time series data. Usually, this type of time series data randomly varies within the fluctuation range of a typical value, or varies according to a relatively long period, etc. This situation is called a normal sample. In some cases, abnormal fluctuation values will appear in the time series data, such as a sudden change in the back pressure of the pipe network caused by a pump failure. At this time, the peak value that appears is regarded as the burr data in this time series data. By identifying and classifying the burr data, it can provide an effective basis for the fault monitoring and identification of the pipe network.
[0059] For this part of the burr data, in this embodiment, first, the burr data is marked from a large amount of original pipe network data, and then the corresponding burr features are extracted from this part of the burr data, so as to define the burr data and construct a burr data set for training the artificial intelligence model, so that the artificial intelligence model can identify the burr part in this type of time series data.
[0060] In the prior art, usually, the burr data set is directly used to train the model. However, due to the imbalance of the samples themselves, for example, in a certain set of pressure measurement point data, the total number of data is 26,042, and only 116 burr data are marked, accounting for 0.44% of the total samples. Therefore, training for a small number of samples is extremely likely to cause overfitting of the model.
[0061] To solve the above problems, in this embodiment, a step of adjusting the loss function is introduced. By adjusting the parameters of the loss function and performing multiple rounds of verification, it is possible to determine a parameter combination that better balances the recall rate and precision rate during the training process of a specific burr data set, and finally fuse them to improve the overall detection effect.
[0062] In one embodiment, as Figure 2 shown, step S1 includes:
[0063] Step S11: Manually label the original pipe network data to obtain data labels;
[0064] The data labels include burr data and normal data;
[0065] Step S12: Extract neighborhood features from the burr data and normal data respectively;
[0066] Step S13: Generate a burr data set according to the data corresponding to the neighborhood features and data labels.
[0067] Specifically, to achieve a better model training effect, in this embodiment, after collecting the original pipeline network data, the original pipeline network data is first manually labeled to obtain data labels.
[0068] Specifically, data within a continuous time period is extracted for analysis and manually labeled, and at the same time, some missing data is excluded. The labeling personnel are preferably experienced duty personnel. Before labeling, the duty personnel perform trial labeling on the data in other time periods, and formal labeling is only carried out when the consistency of the labeling results meets the requirements. In the formal labeling, for data with inconsistent opinions, a voting method is used for decision-making to ensure the quality of manual labeling.
[0069] Through this step, burr data and normal data are identified. In particular, the number of burr data is significantly less than that of normal data. By pre-labeling, unnecessary normal data can be excluded in advance, and the proportion of data that is more effective for training can be determined, and then the extraction of neighborhood data is carried out.
[0070] Generally speaking, burr data detection belongs to the category of time series outlier detection. In the field of pipeline network monitoring, it specifically refers to three consecutive points in the time series, where the value of the first point mutates to the value of the second point, and the value of the third point returns to a value similar to that of the first point.
[0071] For example, at a pressure measurement point of a pipeline network, the value is 0.25 Mpa at 00:00, 1.5 Mpa at 00:01, and 0.26 Mpa at 00:02. It mutates from 0.25 Mpa to 1.5 Mpa and then returns to 0.26 Mpa. In this case, the pressure value at 00:01 is identified as burr data and needs to be corrected or excluded.
[0072] It can be seen that the above burr data has certain differences from other data outliers and pays more attention to local outliers. Therefore, from the perspective of feature design, it is mainly based on the neighborhood. The value of the previous moment at the current moment is the left neighborhood, and the value of the next moment is the right neighborhood. By extracting neighborhood features, a better representation of burr data can be achieved.
[0073] After determining the data ratio, the burr data and normal data are represented by extracting neighborhood features respectively, and finally a burr data set is constructed. According to the training requirements of the model, the burr data set includes a training set and a test set.
[0074] In one embodiment, in step S12, the neighborhood features include neighborhood values, unilateral neighborhood differences, and bilateral neighborhood differences.
[0075] Specifically, to better represent the neighborhood features of the data, in this embodiment, the above-mentioned indicators are selected as the neighborhood features.
[0076] Among them, the neighborhood values include the current value of the data, the left neighborhood value of the data, and the right neighborhood value of the data. The neighborhood values are the data at the nearest sampling points at the previous or next moment of the data.
[0077] The unilateral neighborhood differences include the left neighborhood ratio, the right neighborhood ratio, the left neighborhood difference, and the right neighborhood difference. The left neighborhood ratio is the quotient of the current value and the left neighborhood, the left neighborhood difference is the difference between the current value and the left neighborhood, the right neighborhood ratio is the quotient of the current value and the right neighborhood, and the right neighborhood difference is the difference between the current value and the right neighborhood.
[0078] The bilateral neighborhood differences include the bilateral neighborhood difference and the bilateral neighborhood difference ratio. The bilateral neighborhood difference represents the absolute value of the sum of the left and right neighborhood differences, and the bilateral neighborhood difference ratio is the absolute value of the quotient of the bilateral neighborhood difference and the average value of the left and right neighborhoods.
[0079] In one embodiment, as Figure 3 shown, after executing step S12 and before executing step S13, it further includes:
[0080] Step A13: Classify the glitch data and normal data according to time partitions to form time annotations;
[0081] In step S13, time annotations are also added to the glitch dataset.
[0082] Specifically, considering that in the process of generating glitch data, some glitch data may be related to the non-randomness of time. In this embodiment, the glitch data and normal data are also classified according to time partitions to form time annotations, that is, 24 hours a day are divided into 8 segments, with each segment being 3 hours. The time of data acquisition is encoded in one-hot form. The time period corresponding to the acquisition moment is 1, and the remaining time periods are 0. This part of the encoding is added to the glitch dataset as time annotations to capture this part of the features.
[0083] In one embodiment, as Figure 4 shown, step S2 includes:
[0084] Step S21: Set the parameter adjustment range of the loss function parameters, and generate multiple groups of loss function parameter combinations within the parameter adjustment range;
[0085] Step S22: Adjust the detection model according to the loss function parameter combinations to obtain the model to be verified;
[0086] Step S23: Use the burr data set to verify the model to be verified to obtain a verification result.
[0087] Specifically, to achieve a better model training effect, in this embodiment, first, a parameter adjustment range is generated for the adjustable loss function parameters in the loss function, and then the numerical values of the corresponding loss function parameters are randomly output within the parameter range, and multiple loss function parameters are randomly combined to form a loss function parameter combination. Subsequently, the loss function parameter combination is combined with the detection model to form a model to be verified.
[0088] To achieve a better recognition effect on the time series data of the pipe network, the LightGBM model is selected as the detection model in this solution. Subsequently, the model to be verified is trained and verified using the burr data set respectively to determine the changes in the accuracy and recall rate of the model under different loss functions. Finally, the optimal combination is determined as the actual loss function parameters, and the corresponding detection model is selected for fusion to form a detection fusion model for training and deployment.
[0089] In one embodiment, the loss function is:
[0090] FL(p t )=-α t (1-p t ) γ log(p t );
[0091] In the formula, FL(p t ) is the loss function, p t is the probability of the occurrence of burr data, γ is the adjustment factor parameter, and α is the weight factor.
[0092] Generally speaking, since the detection problem targeted by this solution is to determine burr data in long-term time series data, which is an obvious binary classification problem, and the recognition results are burr data and normal data. For this scenario, the cross-entropy loss function is usually used to train the model in the prior art, which is defined as follows:
[0093]
[0094] In the formula, y∈{0,1} represents the true label of the data, where 1 is burr data and 0 is non-burr data. p∈[0,1] represents the probability that the data to be classified is label = 1. For the convenience of marking, let CE(p,y)=CE(p t )=-log(p t ), where the definition of p t is as follows:
[0095]
[0096] That is, glitch data and non-glitch data are mutually exclusive events.
[0097] For all samples, the loss function L is defined as follows:
[0098]
[0099] Where m is the number of samples, n is the number of negative samples, L is the total number of samples, m+n=N. When the samples are unbalanced, the distribution of the loss function L becomes skewed, especially when m<<n. In this case, negative samples will dominate the loss function. During training, the model will tend to favor the category with more samples, resulting in poor performance of the model on minority samples.
[0100] Corresponding to the situation of glitch data detection in the pipeline network, there are too few glitch sample data, which causes the model to be more inclined to the large number of normal data.
[0101] To solve the above problems, this solution introduces a weight factor α∈[0,1] for positive samples, i.e., burr data, and the weight of negative samples is 1-α. In practice, α is usually set to the frequency of the opposite category, or tuned as a hyperparameter through cross-validation. The cross-entropy loss function with α introduced is defined as a balanced cross-entropy loss function, which can be regarded as a simplified version of the Focal Loss introduced in the present invention. However, the balanced cross-entropy loss only solves the imbalance problem from the perspective of sample distribution, and cannot distinguish the difficulty of samples.
[0102] Therefore, the present invention proposes to introduce the FocalLoss function into the burr data detection model to solve the problem of class imbalance. Its essence is to reconstruct the loss function in the detection model, that is, to add an adjustment factor (1-p t ) γ , (γ≥0), so that the model focuses more on samples that are difficult to classify during training. FocalLoss is defined as:
[0103] FL(p t )=-(1-p t ) γ log(p t ) (2)
[0104] When an example is misclassified and p t When p is small, the value of the modulation factor is close to 1, and the loss function has little effect. t When it approaches 1, the value of the modulation factor is close to 0, and the weight of the samples that are easier to distinguish will be reduced. t It reflects the difficulty of classification. When p tThe larger it is, the higher the confidence of classification, indicating that the sample is easier to classify; p t The smaller it is, the lower the confidence of classification, indicating that the sample is more difficult to classify. Therefore, FocalLoss is equivalent to increasing the weight of difficult-to-classify samples in the loss function, making the loss function tend to difficult-to-classify samples. When γ = 0, this is the traditional cross-entropy loss. The parameter γ can adjust the proportion of the weight reduction of easy-to-classify samples. When γ ∈ [0, 5], the loss function is as Figure 5 shown Figure 5 is the Focal Loss curve for different γ.
[0105] Based on the above results, it can be seen that combining the adjustment factor and the α weight coefficient can achieve better performance. Therefore, the loss function adopted in this scheme is:
[0106] FL(p t ) = -α t (1 - p t ) γ log(p t ).
[0107] Balanced cross-entropy solves the problem of sample imbalance from the perspective of sample distribution. Focal Loss starts from the difficulty level of sample classification, making the loss more focused on difficult-to-classify samples.
[0108] In practice, difficulty itself is a dynamic concept, that is, it changes with the training process. The originally difficult-to-classify samples may change into easy-to-classify samples with the training process. On the contrary, there are also cases where the originally easy-to-classify samples become difficult-to-classify samples. At this time, in order to prevent the frequent change of difficult and easy samples, a smaller learning rate should be selected.
[0109] Correspondingly, in step S21, the adjustment factor parameter and the weight factor are used as the loss function parameters.
[0110] Specifically, in step S2, mainly under the same other conditions, the parameters of the adjustment factor parameter γ and the weight factor α are jointly analyzed, where α represents the weight of the minority class samples, and α ∈ {0.5, 0.6, 0.7, 0.8, 0.9, 0.99} is set. When α = 0.5, it means that the weight factors of samples of different classes are the same. When α = 0.99, it represents an extremely imbalanced situation. γ ∈ {0, 0.1, 0.2, 0.5, 1, 2, 5} is set. When the parameters are α = 0.5 and γ = 0, the FocalLoss loss function will degenerate into the cross-entropy loss function, which is used as the baseline for the comparative experiment of the present invention. Studying the model effects under different parameters is also the basis for model fusion, which is beneficial to selecting models with different effects under different evaluation indicators during model fusion, so as to improve the model effect.
[0111] In one embodiment, as Figure 6 shown, step S3 includes:
[0112] Step S31: Determine evaluation metrics from the verification results, and select a detection model under the corresponding loss function parameters according to the evaluation metrics;
[0113] Step S32: Fuse the detection models to obtain a fused model.
[0114] Specifically, after verifying the models to be verified under different combinations of loss parameters, the verification results usually include multiple evaluation metrics, including precision, recall, and F1 value. After obtaining different combinations of loss parameters, the corresponding detection models can be selected according to different evaluation metrics, and the parameters of the detection models are fused to obtain a fused model. For example, if it is desired to detect more burr data and be able to eliminate the impact of false detection by combining the filling of abnormal data, a relatively high α value can be selected as much as possible to achieve a higher recall rate. If it is relatively more conservative, the balance between precision and recall is considered comprehensively, and the F1 value is considered.
[0115] In a parameter experiment, the above process was verified using the pressure measurement point data of a certain site in the pipe network of a certain water service group.
[0116] In this embodiment, the sensor device collects data at multiple time points per minute, and the data collected at the last time point within each minute is selected as the data of this site for this minute. That is, one piece of data per minute, and a total of 1440 pieces of data per day. A continuous time period with more burr data is selected for research, and the data within the continuous time period from May 1 to 15:00 on May 19, 2023 is extracted for analysis and manually labeled. The annotators are three experienced duty officers at the pipe network end of the water service group, and the consistency of the trial annotation results is as high as 99.95%. The distribution of the data is shown in Table 1:
[0117] Table 1 Dataset Distribution
[0118] Quantity Non-burr 25956 Burr 116 Total 26042
[0119] It can be seen from this that there are 116 pieces of burr data, accounting for only 0.44% of the total data. The extremely large data imbalance poses a challenge to the detection of burr data. 80% of this is selected as the training set, and 20% as the test set. Subsequent comparative experiments are all based on this dataset.
[0120] To address the problem of imbalanced dataset samples, the present invention adopts the above-mentioned method for detecting pipeline burr data by improving LightGBM based on Focal Loss. When the effect of LightGBM is good, next, the effects of LightGBM using FocalLoss and cross-entropy loss are compared respectively, and at the same time, the balance between the recall rate and the precision rate of FocalLoss under different weight factors α and influencing parameters γ is further studied.
[0121] Therefore, in the present invention, under the condition that other parameters are the same, the parameters of α and γ are jointly analyzed, with α ∈ {0.5, 0.6, 0.7, 0.8, 0.9, 0.99} and γ ∈ {0, 0.1, 0.2, 0.5, 1, 2, 5}. The precision rate and recall rate of the experimental results under different parameters are sorted into Figure 7 、 8 as shown, where Figure 7 represents the improvement effect of the model precision rate relative to the baseline under different parameters, and Figure 8 represents the improvement effect of the model recall rate relative to the baseline under different parameters.
[0122] Different colors represent different values of α, and the abscissa corresponds to different values of γ.
[0123] Figure 8 In , when α = 0.99, the recall rate improvement is the highest, exceeding 30%, and the actual recall rate values all remain above 93%. When γ is 0.5, the recall rate even reaches 100%. The recall rate improvement of α = 0.9 is overall lower than that of α = 0.99 and ranks second in the figure. The difference in the recall rate improvement between α = 0.8 and 0.7 is not significant, and it generally tends to be in the middle position in the figure. When α is 0.6 and 0.5, the overall recall rate improvement is the lowest, and in some cases, the recall rate even decreases, and the recall rate values do not exceed 70%. According to the relative improvement effects of different α, it can be seen that the weight coefficient α has a greater impact on the recall rate.
[0124] In addition, the larger α is, the higher the recall rate of positive samples. Most of the improvement effects in the figure are above the baseline, indicating that the weight coefficient α in the Focal loss function can effectively improve the recall rate of positive samples. However, from the perspective of the precision rate improvement, when α is 0.99, the precision rate of positive samples decreases instead, and the decrease amplitude is relatively large.
[0125] This means that in order to ensure that all positive samples are recalled, the model misclassifies some negative samples. This is mainly because when the weight coefficient exceeds a certain threshold, the model pays too much attention to positive samples and insufficiently learns negative samples, resulting in misclassification.
[0126] On the other hand, when α is 0.7, the improvement effect of the precision rate of positive samples is the most obvious, and all exceed the baseline. This indicates that through reasonable selection of α, Focal Loss can also improve the precision rate of the model to a certain extent.
[0127] Precision rate and recall rate are two contradictory indicators. Usually, when we try to improve the precision rate, the recall rate will decrease, and vice versa. The F1 value combines the two through the harmonic mean, synthesizes their performances, and can more comprehensively evaluate the performance of the model. Especially for imbalanced datasets, the precision rate and recall rate may produce misleading results. Therefore, the present invention further analyzes the harmonic mean (F1) value of the positive samples in the above experiments.
[0128] The analysis results are as Figure 9 shown. When α is 0.9, the improvement effect of the overall F1 value is relatively stable, only lower than the point where α = 0.7 when γ = 0.2, and the optimal F1 value is obtained at other points. And when α is 0.7, 0.8, and 0.9, they all comprehensively exceed the baseline. It can be seen that under the same α, different γ values also have a greater impact on the results. When α = 0.7, the recall rate when γ = 0.2 is more than 8% higher than the recall rate when γ = 0.1, and the improvement effect is obvious. The experimental results show that reasonable selection of α and γ values is the key factor for Focalloss to achieve performance improvement. The present invention preferably selects two groups of parameter combinations, namely α = 0.7, γ = 0.2 and α = 0.9, γ = 0.5 to further compare the improvement effect of the model relative to the baseline, as Figure 10 shown.
[0129] After adopting Focal Loss, the precision rate, recall rate, and F1 value have all been significantly improved. Especially the improvement amplitude of the recall rate is the most obvious.
[0130] It is worth noting that when α = 0.9 and γ = 0.5 are selected, the improvement effect of the recall rate is the most prominent, reaching a significant increase of 30%. Further observing the F1 value, it is found that under the two groups of preferred parameters, the overall F1 value has an improvement effect of 16.2%. This is mainly due to the significant increase in the recall rate, which effectively improves the performance of the model in positive sample classification.
[0131] In practical applications, appropriate parameter combinations can be selected according to different requirements. If it is desired to detect more glitch data and be able to eliminate the impact of false detections by combining the filling of abnormal data, a relatively high α value can be selected as much as possible to achieve a higher recall rate. If it is relatively more conservative, the balance relationship between precision and recall can be comprehensively considered and considered from the perspective of the F1 value. The research on the model effects under different parameters is also the basis for model fusion, which is conducive to selecting models with different effects under different evaluation indicators when performing model fusion, so as to improve the model effect.
[0132] Based on the research on precision and recall under different parameters of FocalLoss, the present invention will further study how to select models for multi-model fusion. The comparison between the fusion model and the base models is shown in Table 2.
[0133] Research shows that when performing multi-model fusion, models with differences in evaluation indicators should be selected as much as possible, because the results obtained by homogeneous models may have strong correlations and cannot bring about an improvement in effect. Since the more models are fused, the greater the demand for computing resources and computing time, the present invention only selects three models for fusion attempts. Select α = 0.7, γ = 0.2, α = 0.9, γ = 0.5, α = 0.99, γ = 0.5 as the base models, and the fused model as Experimental Group 1. Among them, when α = 0.99, γ = 0.5, the recall rate of the model is the highest, while when α = 0.7, γ = 0.2 and α = 0.9, γ = 0.5, the F1 value of the model is the highest, which can ensure the effect after model fusion and there are certain differences between the models. Select α = 0.7, γ = 0.2, α = 0.9, γ = 0.5, α = 0.9, γ = 0.1, whose F1 values rank among the top three in Section 5.2, as the base models of Experimental Group 2. Select α = 0.6, γ = 2, α = 0.99, γ = 0.5, α = 0.9, γ = 0.5, which are the highest in precision, recall, and F1 value respectively, as the base models of Experimental Group 3. Each experimental group is compared with each other. The present invention uses the voting method for model fusion, that is, voting on the results of multiple base models, and the minority obeys the majority. The effects of the base models and the fusion model are shown in Table 3.
[0134] Table 2 Fusion Model and Base Models
[0135]
[0136]
[0137] Such as Figure 11As shown in the figure, both Experimental Group 1 and Experimental Group 2 achieved improvements in the F1 value compared to the base model. Among them, Experimental Group 1 made full use of the advantages between different base models, and exceeded the values of its base model with α = 0.9 and γ = 0.5 in terms of precision, recall, and F1 value. Moreover, its precision was second only to the base model with α = 0.7 and γ = 0.2, and its recall was second only to the base model with α = 0.99 and γ = 0.5. However, the F1 value exceeded all its base models, and the F1 value increased by 1.8% compared to the highest F1 value of the base model.
[0138] The experimental results show that Experimental Group 1 achieved a new balance between precision and recall through model fusion. Compared with Experimental Group 1, in Experimental Group 2, α = 0.99 and γ = 0.5 in the base model were replaced with α = 0.9 and γ = 0.1 with a higher F1 value. However, the F1 value of the fused model only increased by 1.2% compared to the highest F1 value of the base model, which was lower than that of Experimental Group 1. By comparing and analyzing the prediction results of the fused model and the base model in Experimental Group 2, it can be seen that the results of its base model with α = 0.9 and γ = 0.1 were similar to those with α = 0.9 and γ = 0.5, and could not provide more differential information for the fused model. Therefore, the improvement effect was not as obvious as that of Experimental Group 1. Compared with Experimental Group 1, in Experimental Group 3, α = 0.7 and γ = 0.2 were replaced with α = 0.6 and γ = 2. However, the results of the fused model did not break through its base model. Although the precision of α = 0.6 and γ = 2 was higher, its overall F1 value differed greatly from other base models, making it difficult to improve the effect after fusion.
[0139] Table 3 Comparison of the results of optimized parameters
[0140]
[0141]
[0142] The results of the control experiment of the fused model show that the effect of the base model is the premise to ensure the effect of the fused model, and the differential performance between models can provide more information for the fused model to further improve the effect. The fused model proposed in the present invention can effectively improve the effect of the base model. Compared with the model with α = 0.9 and γ = 0.5, Experimental Group 1 improved by 0.7% in precision, 3.3% in recall, and 1.8 in the overall F1 value. Compared with the model with α = 0.5 and γ = 0, it improved by 5.1% in precision, 33.3% in recall, and 18% in the overall F1 value.
[0143] The present invention also provides a memory, which includes computer instructions. When a computer device runs the computer instructions, the above-mentioned pipeline burr data detection method is executed.
[0144] Those of ordinary skill in the art will understand that various aspects of the present invention, or possible implementations of various aspects, can be embodied as a system, method, or computer program product. Therefore, various aspects of the present invention, or possible implementations of various aspects, can take the form of a complete hardware embodiment, a complete software embodiment (including firmware, resident software, etc.), or an embodiment combining software and hardware aspects, which are collectively referred to herein as "circuits", "modules", or "systems". In addition, various aspects of the present invention, or possible implementations of various aspects, can take the form of a computer program product, which refers to computer instructions stored in a memory.
[0145] The memory can be a computer-readable signal medium or a computer-readable storage medium. Computer-readable storage media include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or apparatuses, or any suitable combination of the foregoing, such as random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable read-only memory (CD-ROM).
[0146] The processor in the computer reads the computer instructions stored in the memory, enabling the processor to perform the functional actions specified in each step, or combination of steps, in the flowchart; generating means for performing the functional actions specified in each block, or combination of blocks, in the block diagram.
[0147] It should be understood that the processor in the computer can be understood as being implemented by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers (MCUs), microprocessors, or other electronic components for executing the aforementioned computer instructions.
[0148] The computer instructions can be executed entirely on the user's local computer, partially on the user's local computer, as a separate software package, partially on the user's local computer and partially on a remote computer, or entirely on a remote computer or server. It should also be noted that in certain alternative embodiments, the functions noted in each step in the flowchart, or each block in the block diagram, may not occur in the order noted in the figure. For example, depending on the functions involved, two consecutive steps, or two blocks shown, may actually be executed substantially simultaneously, or these blocks may sometimes be executed in the reverse order.
[0149] The above are only the preferred embodiments of the present invention, and do not limit the implementation manners and protection scope of the present invention accordingly. For those skilled in the art, it should be realized that all the solutions obtained by equivalent substitution and obvious changes made by using the description and illustrations of the present invention should be included in the protection scope of the present invention.
Claims
1. A pipeline network burr data detection method based on Focal Loss improved LightGBM, characterized in that: include: Step S1: extract neighborhood features from the original data of the pipe network and construct a burr dataset; Step S2: verifying the detection model multiple times based on the burr data set, and adjusting the loss function parameters of the loss function of the detection model during each verification process; Step S3: Select multiple detection models according to the verification results to fuse them into a fusion model, and use the fusion model to identify the pipe network data.
2. The method for detecting burr data in a pipe network according to claim 1, characterized in that: The step S1 comprises: Step S11: manually labeling the original data of the pipe network to obtain data labels; The data labels include burr data and normal data; Step S12: extracting the neighborhood features from the burr data and the normal data respectively; Step S13: Generate the burr data set according to the neighborhood features and the data corresponding to the data labels.
3. The method for detecting burr data in a pipe network according to claim 2, characterized in that: In step S12, the neighborhood features include neighborhood values, unilateral neighborhood differences, and bilateral neighborhood differences.
4. The method for detecting burr data in a pipe network according to claim 2, characterized in that: After executing step S12 and before executing step S13, the method further includes: Step A13: classifying the burr data and the normal data according to time partitions to form time annotations; In the step S13, the time mark is also added to the burr data set.
5. The method for detecting burr data in a pipe network according to claim 1, characterized in that: The step S2 comprises: Step S21: setting a parameter adjustment range of the loss function parameter, and generating a plurality of loss function parameter combinations within the parameter adjustment range; Step S22: adjusting the detection model according to the loss function parameter combination to obtain a model to be verified; Step S23: using the burr data set to verify the model to be verified to obtain the verification result.
6. The method for detecting burr data in a pipe network according to claim 5, characterized in that: The loss function is: FL(p t )=-a t (1-p t ) γ log(p t ); In the formula, FL(p t ) is the loss function, p t is the probability of the burr data appearing, γ is the adjustment factor parameter, and α is the weight factor.
7. The method for detecting burr data in a pipe network according to claim 6, characterized in that: In the step S21, the adjustment factor parameter and the weight factor are used as the loss function parameters.
8. The method for detecting burr data in a pipe network according to claim 1, characterized in that: The detection model is the LightGBM model.
9. The method for detecting burr data in a pipe network according to claim 1, characterized in that: The step S3 comprises: Step S31: determining an evaluation index from the verification result, and selecting the detection model under the corresponding loss function parameter according to the evaluation index; Step S32: Fusing the detection models to obtain the fusion model.
10. A memory comprising computer instructions, characterized in that: When the computer device runs the computer instructions, the pipeline network burr data detection method according to any one of claims 1 to 9 is executed.
Citation Information
Patent Citations
A Method and Device for Detecting Electricity Consumption Anomalies Based on BRB and LSTM Models using Big Data in the Power Industry
CN112926686B
Water affair flow burr data detection method, system and equipment and readable medium
CN117194752A