An insulator pollution flash identification and contamination degree prediction method based on multi-feature fusion
By using consistent acquisition session configuration and multi-feature fusion processing, the inconsistency problem in waveform acquisition and processing throughout the insulator discharge process was solved. This achieved a synchronous closed loop for discharge state identification and pollution level prediction, improving the accuracy and reliability of power equipment monitoring. It also ensured the effectiveness of the technical means and achieved a synchronous closed loop for discharge state identification and pollution level prediction, thus improving the accuracy and reliability of power equipment monitoring.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NANJING INST OF TECH
- Filing Date
- 2026-03-26
- Publication Date
- 2026-05-29
AI Technical Summary
Existing technologies for acquiring and processing spatial electric field waveforms throughout the entire insulator discharge process suffer from problems such as difficulty in unifying acquisition session configurations, inconsistent feature fusion links, and unstable model training and prediction values, resulting in insufficient accuracy and reliability of insulator flashover identification and pollution degree prediction.
By binding probe numbers, registering geometric relationships, calibrating channels, registering sampling rates, implementing synchronous triggering strategies, and generating a unified time axis, a consistent acquisition session configuration is formed. Combining time domain, frequency domain, and time-frequency energy feature extraction, channel health and reliability scoring are performed, high-dimensional feature vectors are generated, and redundancy is eliminated. Finally, parallel training and consistency verification are performed to generate alarm escalation strategies.
This achieves a synchronous closed loop for discharge state identification and pollution level prediction, improving the accuracy and reliability of power equipment monitoring and ensuring the consistency and traceability of online output results.
Smart Images

Figure CN121903096B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power equipment condition assessment and intelligent early warning, and in particular to a method for insulator flashover identification and pollution degree prediction based on multi-feature fusion. Background Technology
[0002] In the field of power equipment condition assessment and intelligent early warning, existing solutions for obtaining spatial electric field waveforms throughout the insulator discharge process typically rely on lists of spatial electric field sensors and environmental sensors for data acquisition. These data undergo denoising, normalization, filtering, and segmentation, which presents limitations such as inconsistent acquisition session configurations, unstable extraction criteria for stage anchor points, and inconsistent feature fusion links. Existing methods often rely on a single feature dimension or fixed feature combinations after segmentation for judgment and modeling. Under online input of real-time feature vectors, inconsistencies arise in the synchronous output of development sections and pollution prediction values, and consistency checks are difficult to close, making it challenging to achieve stable development sections and pollution prediction values. For the joint processing of acquisition session configuration and spatial electric field waveforms throughout the discharge process, existing technologies generally suffer from insufficient integration in areas such as unified time axis generation and timestamp alignment, channel health and credibility scoring, feature credibility gating and environmental parameter fusion. This makes it difficult to form a consistent process for acquisition session configuration generation, sliding time window segmentation, segment label alignment, discrimination index calculation, redundancy resolution records and version number association in application scenarios where insulator structure parameters, spatial electric field sensor lists and environmental sensor lists are collaboratively constrained. As a result, it is difficult to stably build the optimal feature ledger for segments and support the continuous operation of parallel training and development of segment identification models and pollution prediction models. Summary of the Invention
[0003] Purpose of the invention: To provide a method for identifying flashover and predicting pollution levels in insulators based on multi-feature fusion, achieving synchronous closed-loop identification of discharge state and prediction of pollution levels under non-contact monitoring conditions, aiming to improve the accuracy and reliability of power equipment monitoring.
[0004] To achieve the above objectives, this invention proposes a method for insulator flashover identification and pollution degree prediction based on multi-feature fusion, which is implemented through the following process:
[0005] The system acquires insulator structure parameters, a list of space electric field sensors, and a list of environmental sensors. It then binds probe numbers, registers geometric relationships, and calibrates channels for these sensors. Finally, it performs sampling rate registration, synchronous triggering strategy registration, unified time axis generation, and timestamp alignment to generate the acquisition session configuration.
[0006] Based on the acquisition session configuration, the spatial electric field waveform of the entire discharge process is accessed and time-aligned, and noise reduction and normalization, sliding time window segmentation and stage anchor point candidate extraction are performed to generate a segmented slice set and a stage anchor point candidate set.
[0007] Extract the time-domain statistical features, frequency-domain spectral features, and time-frequency energy features of the segmented slice set and the stage anchor point candidate set, and perform channel health and credibility scoring calculation, feature credibility gating, and environmental parameter fusion processing to generate high-dimensional feature vector records;
[0008] Calculate the feature discrimination of predetermined fields in the high-dimensional feature vector record according to the segment and perform redundancy elimination processing to generate the segment-optimal feature ledger;
[0009] Based on the optimal feature ledger of the aforementioned segment, parallel training and version management, online inference and consistency verification of the segment identification model and the pollution degree prediction model are performed to generate alarm escalation strategy instructions.
[0010] As a preferred embodiment, the probe number binding, deployment geometry registration, and channel calibration processing specifically include:
[0011] Generate a channel number table containing channel number, sensor identifier, and installation point identifier based on the list of space electric field sensors;
[0012] Based on the insulator structural parameters and installation point identification, calculate and register the distance and orientation parameters of each channel relative to the insulator reference point to form a channel geometric binding record;
[0013] Zero-point bias measurement, gain coefficient calibration, and saturation threshold registration are performed on the space electric field sensor channel.
[0014] As a preferred embodiment, the execution sampling rate registration, synchronization triggering strategy registration, unified timeline generation, and timestamp alignment processing specifically include:
[0015] Based on the monitoring task type and sensor range, write the sampling rate field and the segmented slice window length field to generate a synchronous trigger strategy field containing trigger source identifier, trigger source type, and trigger delay compensation value;
[0016] A unified time axis field is generated based on the clock status of each channel, and clock offset estimation, offset compensation and alignment verification are performed on each channel. A collection session configuration is generated, which includes session number, channel number table, distance parameter, azimuth parameter, calibration parameter, sampling rate field, synchronization trigger strategy field, unified time axis field, timestamp alignment status field, segmented slice window length field, sliding step size field and collection session configuration version number.
[0017] As a preferred embodiment, the denoising and normalization, sliding time window segmentation, and stage anchor point candidate extraction processes specifically include:
[0018] The time-aligned waveform is sequentially subjected to filtering, baseline drift correction, outlier repair, and amplitude normalization to obtain a preprocessed sequence.
[0019] The preprocessed sequence is truncated based on the segmented slice window length field and the sliding step field to generate a segmented slice set containing the segment number, window start timestamp, window end timestamp and segmented slice data;
[0020] For each segment slice, calculate the rising edge mutation index, local high-frequency energy surge index, and duration index, and generate a candidate set of stage anchor points containing candidate time window numbers.
[0021] As a preferred embodiment, the channel health and credibility score calculation includes: calculating the noise ratio, saturation ratio and drift index as channel health fields based on the waveform sequence, and calculating the credibility score by combining the fluctuation amplitude of the channel health fields and features in adjacent windows;
[0022] The feature credibility gating process includes: performing gating discrimination on feature records based on the credibility score, candidate credibility markers of stage anchor candidates, and feature missing markers, and outputting a gating pass feature set;
[0023] The environmental parameter fusion processing includes: obtaining temperature and humidity sequences based on the environmental sensor list, and aligning and splicing them with the gating control according to the feature set and time window.
[0024] As a preferred embodiment, the high-dimensional feature vector record includes:
[0025] The fields include session number, channel number, segment number, window start timestamp, mean, variance, skewness, kurtosis, spectral centroid, bandwidth, spectral entropy, wavelet packet energy distribution, noise percentage, saturation ratio, drift index, credibility score, temperature sequence statistics, humidity sequence statistics, gating result, and acquisition session configuration version number.
[0026] As a preferred embodiment, the step of calculating the feature discriminative power of predetermined fields in the high-dimensional feature vector record by segment and performing redundancy resolution specifically includes:
[0027] Historical samples with segment labels are grouped according to segment labels. Inter-class distance and intra-class divergence ratio or mutual information are calculated for the fields of mean, variance, skewness, kurtosis, spectral centroid, bandwidth, spectral entropy and wavelet packet energy distribution as discrimination indicators. The credibility score is introduced as a sample weight in the calculation process. Valid samples are selected based on noise ratio, saturation ratio and drift index.
[0028] A candidate feature pool is established based on the discriminant index. Redundancy is determined by correlation threshold or mutual information overlap. The feature pool is sorted according to the discriminant index and the mean, variance, spectral centroid and wavelet packet energy distribution are retained as minimum set fields to form the optimal feature subset of the segment. A redundancy resolution record containing redundancy determination path markers, threshold version markers, field removal reference markers and field retention reference markers is generated.
[0029] As a preferred embodiment, the segment-optimal feature ledger includes:
[0030] Segment labels, discrimination indexes, redundancy resolution records, optimal feature subsets of segments, ledger version numbers, and collection session configuration version numbers.
[0031] As a preferred embodiment, the parallel training and version management includes: parsing feature entry mapping records containing feature field reference markers from the optimal feature ledger of the segment; assembling a training feature matrix from historical samples with segment labels based on these records to train the development segment identification model and the contamination prediction model in parallel; and generating a model parameter version number that is associated with the ledger version number and the collection session configuration version number.
[0032] The online inference includes: performing a matching verification between the acquisition session configuration version number and the ledger version number on the high-dimensional feature vector record acquired in real time; assembling a real-time feature vector based on the feature entry mapping record; and inputting it into the development segment identification model and the pollution degree prediction model respectively to obtain the development segment field and the pollution degree prediction value field.
[0033] The consistency verification process includes: calling the mapping table between the development segment and the pollution threshold interval, comparing the pollution prediction value field with the pollution threshold interval mapped by the development segment field, and generating an alarm escalation strategy instruction based on the consistency verification result.
[0034] As a preferred embodiment, the alarm escalation policy instruction includes:
[0035] Session ID, Development Segment Field, Pollution Prediction Value Field, Alarm Level Field, Policy Version Number Field, and Action Reference Marker.
[0036] Compared with the prior art, the present invention has the following beneficial effects:
[0037] (1) To address the problem of difficulty in unifying the acquisition session configuration, this invention binds the probe number, registers the layout geometry, calibrates the channel, registers the sampling rate, registers the synchronous triggering strategy, and solidifies the generation of a unified time axis and timestamp alignment into the acquisition session configuration. This ensures that the spatial electric field waveform and the output of the environmental sensor list during the entire discharge process form a consistent timing and channel aperture on the acquisition side, reducing the sample alignment deviation caused by differences in channel calibration, asynchronous triggering, and inconsistent timestamps in the acquisition link. This also ensures that subsequent noise reduction and normalization, filtering, baseline drift correction, and sliding time window segmentation have verifiable input boundaries.
[0038] (2) By simultaneously organizing time-domain statistical features, frequency-domain spectral features and time-frequency energy features at the segmented slice level, and introducing channel health and credibility scores to form feature credibility gating constraints, and then performing environmental parameter fusion under the gating constraints to obtain high-dimensional feature vectors, feature extraction, quality assessment and environmental parameter fusion form a consistent data organization relationship within the same link, reducing feature mismatch caused by residual anomaly repair, drift index fluctuations or noise ratio changes, and improving the input stability of subsequent segment-based feature discrimination calculation.
[0039] (3) This invention takes the optimal feature ledger of the segment as the core, and manages the segment label, discrimination index, redundancy elimination record, optimal feature subset of the segment, ledger version number and collection session configuration version number in a unified manner, so that the parallel training development segment identification model and the pollution degree prediction model can obtain consistent feature entry and version constraints; and through consistency verification and matching the mapping table of development segment and pollution degree threshold interval, output resampling instructions or alarm upgrade strategy instructions, so that the online output results and collection session configuration form a traceable closed loop link, reducing the judgment drift caused by changes in sample distribution or configuration evolution during online operation, and maintaining the coordination and consistency of development segment and pollution degree prediction value within the same operating framework. Attached Figure Description
[0040] Figure 1 This is a schematic diagram of the overall process of an insulator flashover identification and pollution degree prediction method based on multi-feature fusion, provided for an embodiment of this application.
[0041] Figure 2 The execution flowchart of step S400 provided in the embodiments of this application is shown. Detailed Implementation
[0042] In the following description, numerous specific details are set forth in order to provide a more thorough understanding of the invention. However, it will be apparent to those skilled in the art that the invention can be practiced without one or more of these details. In other instances, certain technical features well-known in the art have not been described in order to avoid obscuring the invention.
[0043] Reference Figure 1 This is a flowchart illustrating a method for insulator flashover identification and pollution degree prediction based on multi-feature fusion provided in an embodiment of the present invention. The process may include at least steps S100-S500:
[0044] S100, based on insulator structure parameters, space electric field sensor list and environmental sensor list, performs probe number binding, deployment geometry registration and channel calibration processing, and performs sampling rate registration, synchronous trigger strategy registration, unified time axis generation and timestamp alignment processing to generate acquisition session configuration;
[0045] In S100, the insulator structure parameters, the list of spatial electric field sensors and the list of environmental sensors are first obtained. Then, probe number binding, deployment geometry registration, channel calibration, sampling rate registration, synchronous triggering strategy registration, unified time axis generation and timestamp alignment are performed to generate the acquisition session configuration.
[0046] Specifically, this step is triggered by the acquisition session management module in the online insulator monitoring scenario. Triggering conditions include the issuance of a monitoring task, successful equipment power-on self-test, arrival of scheduled polling, or alarm-triggered resampling to enter session reconstruction mode. The acquisition session management module obtains insulator structural parameters from the equipment asset archive and field installation records. These parameters include insulator string type, skirt dimensions, string length, relative hardware positions, installation height, and adjacent phase spacing. Simultaneously, it obtains a list of spatial electric field sensors and an environmental sensor list from the sensor registry. The spatial electric field sensor list includes sensor identifier, mounting bracket identifier, measurement range, number of channels, communication address, and power supply status fields. The environmental sensor list includes temperature sensor identifier, humidity sensor identifier, installation point identifier, sampling interval, and communication address fields. The acquisition session management module loads the above inputs into a session initialization input package and writes it into the draft acquisition session configuration to be generated, for subsequent probe number binding and deployment geometry registration.
[0047] Specifically, probe number binding is performed by the channel mapping unit. The channel mapping unit assigns a channel number to each space electric field sensor based on the space electric field sensor list and generates a channel number table. The channel number table fields include the channel number, sensor identifier, installation point identifier, communication address, measurement range, and power supply status. Deployment geometry registration is performed by the geometry registration unit. The geometry registration unit reads the insulator structural parameters and installation point identifier, calculates and registers the distance and orientation parameters of each channel relative to the insulator reference point. The distance parameter represents the straight-line distance from the sensor's sensitive surface to the insulator reference point, and the orientation parameter represents the angular direction encoding between the sensor's sensitive surface direction and the insulator's axis direction. The geometry registration unit writes the distance and orientation parameters into the channel number table to form a channel geometry binding record and registers the deployment geometry version number in the acquisition session configuration draft. Understandably, when there are redundant deployments of multiple probes on site, the geometric registration unit performs conflict checks on the channel geometric binding records. Conflict checks include three categories: duplicate installation points, identical orientation parameters and distance parameter differences below the threshold, and duplicate communication addresses. When a conflict occurs, a conflict flag is generated and written to the abnormal record field of the acquisition session configuration draft, while keeping the channel number table from overwriting the original record, and then proceeding to manual confirmation or automatic reassignment process.
[0048] Specifically, channel calibration is performed by the calibration unit. The calibration unit performs zero-point offset measurement, gain coefficient calibration, and saturation threshold registration on the space electric field sensor channel. Zero-point offset measurement acquires baseline segments and calculates the baseline mean during periods of no discharge or low disturbance. Gain coefficient calibration is written into the gain field based on preset calibration sources or factory calibration parameters. Saturation threshold registration is written into the upper and lower limit threshold fields based on the range setting. Simultaneously, the calibration unit performs range verification and drift checks on the environmental sensor and writes the verification results into the environmental channel status field. Sampling rate registration is performed by the sampling configuration unit. The sampling configuration unit writes the sampling rate field based on the monitoring task type and the space electric field sensor range setting, and also writes the segmented slice window length and sliding step size configuration fields. The segmented slice window length and sliding step size serve as inputs for subsequent sliding time window segmented slicing steps. The registration of the synchronous triggering strategy is performed by the synchronous triggering unit. The synchronous triggering unit generates the synchronous triggering strategy field, which includes the trigger source identifier, trigger source type, trigger delay compensation value, and re-trigger suppression interval. The trigger source identifier includes two types: hard trigger source and software trigger source of the space electric field sensor. The hard trigger source is provided by the coaxial trigger line or isolated trigger interface, and the software trigger source is provided by the acquisition session management module issuing trigger commands and recording command timestamps.
[0049] Specifically, the unified timeline generation and timestamp alignment are performed by the time alignment unit. The time alignment unit reads the sampling rate field, the synchronization trigger strategy field, and the clock status of each channel, generates a unified timeline field, and writes it into the acquisition session configuration draft. The unified timeline field includes the start timestamp, time dimension, time window number, and cumulative offset fields. The time alignment unit performs timestamp alignment on the spatial electric field sensor channel and the environmental sensor channel. The timestamp alignment process includes three stages: clock offset estimation, offset compensation writing, and alignment verification writing. Clock offset estimation is calculated based on the difference between the trigger time recorded in the synchronization trigger strategy field and the arrival time of each channel. Offset compensation writing writes the compensation value into the offset field of each channel. Alignment verification writing generates a verification flag by comparing the stability of the difference between consecutive trigger times and writes it into the alignment status field. In the engineering embodiment, the space electric field sensor is installed on the lateral support of the insulator string below the crossarm of the tower, and the environmental sensor is fixed on the leeward side of the tower body; the acquisition session management module triggers the calibration unit to complete the zero-point offset measurement during the low-disturbance window at night, and then the synchronous triggering unit sends out the software trigger source and generates the trigger delay compensation value. After the time alignment unit writes the unified time axis field, an acquisition session configuration that can be directly called for the acquisition of space electric field waveforms throughout the discharge process is formed.
[0050] The output of this step is the acquisition session configuration. The acquisition session configuration fields include session number, channel number table, distance parameter, orientation parameter, calibration parameter, sampling rate field, synchronization trigger strategy field, unified time axis field, timestamp alignment status field, segmented slice window length field, and sliding step size field. The acquisition session configuration version number is written into the version number subfield of the acquisition session configuration field. Among them, the session number and channel number table serve as the index entry for subsequent steps to obtain the spatial electric field waveform of the entire discharge process. The unified time axis field and timestamp alignment status field serve as the alignment basis for subsequent steps of denoising and normalization, filtering, baseline drift correction, outlier repair, and sliding time window segmented slicing. The calibration parameter and sampling rate field serve as the input constraints for subsequent steps of segmented slicing. The environmental sensor list related fields serve as the input source for subsequent steps of performing environmental parameter fusion based on feature credibility gating.
[0051] This step completes the collaborative registration of channel number, deployment geometry, channel calibration, sampling rate, synchronization triggering strategy, and unified time axis through the acquisition session configuration, forming a consistent acquisition and alignment basic data structure across channels, and providing traceable configuration input for subsequent steps.
[0052] S200: Based on the acquisition session configuration, the spatial electric field waveform of the entire discharge process is accessed and time-aligned, and noise reduction and normalization, sliding time window segmentation and stage anchor point candidate extraction processing are performed to generate a segmented slice set and a stage anchor point candidate set.
[0053] Specifically, this step is triggered after the acquisition session management module completes the generation of the acquisition session configuration. Triggering conditions include a change in the acquisition session configuration version number, a mismatch marker appearing in the timestamp alignment status field, the online monitoring task entering the discharge process spatial electric field waveform acquisition window, or the previous round of consistency verification outputting a resampling instruction entering the resampling window. The input sources for this step include the acquisition session configuration and the discharge process spatial electric field waveform. The acquisition session configuration, as the constraint carrier for waveform access and alignment processing, includes at least the session number, channel number table, calibration parameters, sampling rate field, synchronization trigger strategy field, unified time axis field, timestamp alignment status field, segmented slice window length field, and sliding step size field. The discharge process spatial electric field waveform is the time-series data continuously acquired by each channel within the time period covered by the unified time axis field. The data fields include the session number, channel number, sampling timestamp, original sampled value, and channel status marker.
[0054] Specifically, waveform access is performed by the waveform access unit of the edge acquisition terminal. The waveform access unit polls the spatial electric field sensor channels according to the channel number table, drives sampling according to the sampling rate field, and writes the original sampled values. Simultaneously, it reads the channel status flags and writes saturation, disconnection, and reconnection flags. When the synchronization trigger strategy field points to a hard trigger source, the waveform access unit writes a trigger time flag and records the trigger delay compensation value after the trigger edge arrives. When the synchronization trigger strategy field points to a software trigger source, the waveform access unit receives the trigger command issued by the acquisition session management module and records the command timestamp. Understandably, the spatial electric field waveform during the entire discharge process may experience channel sampling loss, sampling jitter, or short-term saturation. When the waveform access unit detects consecutive disconnection flags or saturation flags exceeding a threshold, it writes an anomaly record field and archives it with the session number for subsequent anomaly point repair and segmented boundary constraint calls.
[0055] Specifically, time alignment is performed by the time alignment unit. This unit reads a unified time axis field and performs alignment mapping on the sampling timestamps of the spatial electric field waveform throughout the discharge process. The mapping process includes three stages: timestamp verification, offset compensation writing, and alignment index generation. Timestamp verification determines whether a channel has entered the alignment processing path based on the timestamp alignment status field. Offset compensation writing reads the offset field from the acquisition session configuration to correct the sampling timestamps. Alignment index generation maps the corrected sampling timestamps to a unified time window sequence number and segment sequence number, forming the alignment index field. The alignment index field, along with the session number and channel number, constitutes the index input for subsequent sliding time window segmentation. When the time alignment unit detects a mismatch in the timestamp alignment status field, it writes an alignment failure flag and calls the acquisition session management module to trigger session reconstruction. During reconstruction, the original alignment index field is retained, and a rollback flag is written, forming a traceable operation record.
[0056] Specifically, the denoising and normalization operation is performed by the preprocessing unit. Denoising and normalization is defined as performing noise suppression and amplitude uniformity processing on the aligned spatial electric field waveform of the entire discharge process, outputting a preprocessed sequence that can be used for subsequent feature extraction. The preprocessing unit first performs filtering, which involves applying bandpass or low-pass constraints to the original sampled values using digital filtering operations. The filtering constraint parameters are derived from the calibration parameters and sampling rate field configured in the acquisition session. In scenarios with power frequency interference or switching operation interference, the preprocessing unit switches the filtering configuration to a multi-segment filtering configuration. The switching condition comes from the interference flag in the channel status flag or the interference type field in the anomaly record field. Then, baseline drift correction is performed. Baseline drift correction is defined as subtracting the low-frequency offset trend of the filtered output sequence. The correction process reads the zero-point offset field from the calibration parameters, combines it with the unified time axis field to generate a baseline segment index, calculates the baseline offset from the baseline segment and writes it into the drift compensation field, and then performs drift subtraction on the sequence to obtain the drift-corrected sequence. Then, anomaly repair is performed. Anomaly repair is defined as replacing the sampling points corresponding to the disconnection marker, lost sample marker, and saturation marker. The replacement process includes three types of paths: neighborhood interpolation, segment endpoint truncation, and saturation value backoff. Path selection is driven by the channel status marker and the anomaly record field. When an anomaly is near the trigger time marker, anomaly repair prioritizes endpoint truncation and writes a truncation marker. The truncation marker is used by subsequent anchor point candidate extraction to bypass this interval. After anomaly repair is completed, the preprocessing unit performs amplitude normalization. Amplitude normalization maps the drift correction sequence to a uniform amplitude scale based on the gain coefficient calibration field in the calibration parameters and writes it to the normalization parameter field. The normalization parameter field is associated with the session number and saved for subsequent reliability scoring traceability.
[0057] Specifically, the sliding time window segmentation slicing is performed by the slicing unit. Sliding time window segmentation slicing is defined as, under a unified time axis field constraint, performing continuous window truncation on the normalized drift correction sequence according to the segmented slice window length field and the sliding step size field to generate segmented slices. The slicing unit determines the starting time window number based on the alignment index field, and proceeds from the starting time window number according to the sliding step size field, truncating the set of sampling points covered by the window length field window by window to generate segmented slices. For each segmented slice, a segment number, window start timestamp, and window end timestamp fields are generated. When an abnormal record field has a truncation mark or a disconnection mark covering the window interval, the slicing unit records the window as an invalid window and writes an invalid window mark. The invalid window mark is used as input for subsequent feature confidence gating. The slicing unit aggregates the segmented slices by channel number to form a segmented slice set. The segmented slice set fields at least include the session number, channel number, segment number, window start timestamp, window end timestamp, segmented slice data, and invalid window mark, where the segmented slice data is the sequence of sampled values within the window.
[0058] Specifically, the stage anchor point candidate extraction is performed by the anchor point extraction unit. The stage anchor point candidate is defined as the set of candidate positions in the segmented slice set used to characterize the transition between stages of the discharge process. It is characterized by the rising edge abrupt change index, the local high-frequency energy surge index, and the duration index. The anchor point extraction unit calculates the rising edge abrupt change index for each segmented slice. The rising edge abrupt change index is a composite of the slope abrupt change intensity and the stability of the abrupt change position in the sampled value sequence within the window. The slope abrupt change intensity is obtained statistically from the differential sequences of adjacent sampling points, and the stability of the abrupt change position is obtained from the concentration of the differential peak within the window. The anchor point extraction unit calculates the local high-frequency energy surge index for each segmented slice. The local high-frequency energy surge index is a composite of the energy density of the filtered output sequence in the high-frequency subband and the energy difference between the preceding and following windows. The high-frequency subband boundary is derived from the sampling rate field and the filtering parameters. The anchor point extraction unit calculates the duration index for each segmented slice. The duration index is a composite of the number of windows where the rising edge abrupt change index or the local high-frequency energy surge index continuously exceeds a threshold and the continuous coverage time. Subsequently, the anchor point extraction unit performs combined discrimination on the rising edge mutation index, the local high-frequency energy surge index, and the duration index. The combined discrimination includes three stages of processing: threshold comparison, adjacent window consistency check, and inter-channel commonality check. The threshold comparison is based on the threshold configuration field registered in the acquisition session configuration or the threshold version field generated by historical operation statistics. The inter-channel commonality check is based on the channel number table to align and merge multiple channel candidates within the same time window to form stage anchor point candidates. When the channel status flag has a saturation flag or an invalid window flag, the anchor point extraction unit reduces the weight of the channel candidate and writes a candidate trust flag. The candidate trust flag is used as an input reference field for subsequent channel health and trustworthiness score calculation.
[0059] The output of this step includes a segmented slice set and a stage anchor candidate set. The segmented slice set fields include session ID, channel ID, segment number, window start timestamp, window end timestamp, segmented slice data, and invalid window marker. The stage anchor candidate set fields include session ID, channel ID, candidate number, candidate time window number, rising edge mutation index, local high-frequency energy surge index, duration index, and candidate trustworthy marker. The segmented slice set is used as input for segmented slices in subsequent steps when performing time-domain statistical features, frequency-domain spectral features, and time-frequency energy features extraction. The candidate trustworthy marker and invalid window marker are used as scoring input in subsequent steps when calculating channel health and trustworthiness scores. The candidate time window number is used as a time alignment index reference field in subsequent steps when performing feature trustworthiness gating and environmental parameter fusion. At the same time, the session ID is used throughout subsequent steps to associate the ledger version number and the acquisition session configuration version number in the segment optimal feature ledger.
[0060] This step involves aligning, denoising, normalizing, filtering, baseline drift correction, outlier repair, and sliding time window segmentation of the spatial electric field waveform throughout the discharge process, forming a traceable set of segmented slices. The output of the stage anchor point candidate extraction is a candidate set that can be directly connected to the subsequent feature processing link.
[0061] S300: Based on the segmented slice set and the stage anchor point candidate set, perform time-domain statistical features, frequency-domain spectral features and time-frequency energy features extraction processing, and perform channel health and credibility score calculation, feature credibility gating and environmental parameter fusion processing to generate high-dimensional feature vector records;
[0062] Specifically, this step is triggered by the feature engineering module after the generation of stage anchor point candidates. The triggering conditions include the arrival of the segment slice set writing completion marker, the stage anchor point candidate set candidate credibility markers meeting the minimum number threshold, or the acquisition session management module receiving a resampling instruction and entering the fast feature recalculation path. The input sources for this step include the segment slice set, the stage anchor point candidate set, and the acquisition session configuration. The segment slice set provides segment slice data, window start timestamp, and window end timestamp fields. The stage anchor point candidate set provides candidate time window sequence number, rising edge mutation index, local high-frequency energy surge index, duration index, and candidate credibility marker. The acquisition session configuration provides a channel number table, calibration parameters, sampling rate field, unified time axis field, and timestamp alignment status field.
[0063] Specifically, the extraction of time-domain statistical features, frequency-domain spectral features, and time-frequency energy features is performed by the feature extraction unit. The feature extraction unit uses segmented data as the basic input object and the segment number as an index to generate three types of feature sub-vectors for each segmented data and writes them into the feature record. Time-domain statistical features are defined as representations obtained by performing statistical operations on the segmented data in the time domain, and include at least four fields: mean, variance, skewness, and kurtosis. The mean field is calculated from the sample average of the segmented data; the variance field is calculated from the sample dispersion; the skewness field is calculated from the asymmetry of the sample distribution; and the kurtosis field is calculated from the kurtosis of the sample distribution. The feature extraction unit associates these fields with the session number, channel number, segment number, and window start timestamp fields and writes them into the time-domain statistical feature record. The frequency domain spectral shape feature is defined as a characteristic quantity obtained by statistically analyzing the spectral shape after frequency domain transformation of the segmented slice data. It includes at least three fields: spectral centroid, bandwidth, and spectral entropy. The frequency domain transformation is performed by the spectrum calculation unit under the constraint of the sampling rate field. The spectral centroid field is calculated based on the weighted center of the spectral energy distribution. The bandwidth field is calculated based on the bandwidth of the main energy concentration interval. The spectral entropy field is calculated based on the energy distribution dispersion. The spectrum calculation unit also writes a spectrum validity marker field to mark the validity of the frequency domain transformation. The time-frequency energy feature is defined as a characteristic quantity obtained by quantizing the energy distribution of the decomposed sub-bands after time-frequency decomposition of the segmented slice data. It includes at least a wavelet packet energy distribution field. The time-frequency decomposition is performed by the wavelet packet decomposition unit. Under the constraint of the sampling rate field, the wavelet packet decomposition unit performs multi-scale decomposition of the segmented slice data, calculates the energy proportion of each sub-band at each scale, and summarizes it in the wavelet packet energy distribution field. Understandably, for segments with invalid window markings or truncation markings, the feature extraction unit writes the segment into the feature missing marking field and uses the missing marking field in subsequent confidence score calculations to avoid missing segments directly entering environmental parameter fusion.
[0064] Furthermore, to form a feature organization method consistent with the stage anchor point candidates, the feature extraction unit performs anchor point association registration on the segment numbers corresponding to the candidate time window numbers. Anchor point association registration is defined as establishing a correspondence between the candidate time window numbers in the stage anchor point candidate set and the segment numbers in the feature records and writing them into the anchor point association field. The anchor point association field and the candidate credibility marker together form the index input for subsequent feature credibility gating. In the engineering embodiment, the spatial electric field sensor channel experiences an increase in noise ratio before and after rainfall. When the effective spectral marker field is abnormal, the feature extraction unit writes the corresponding field of the frequency domain spectral shape feature into the missing measurement marker, and retains the time domain statistical features and wavelet packet energy distribution fields to form a record of cross-domain feature availability differences, providing a traceable basis for subsequent gating.
[0065] Specifically, channel health and reliability scoring are performed by the health scoring unit. The health scoring unit gathers inputs from the segmented slice set, acquisition session configuration, and feature records to generate channel health and reliability scores. Channel health is defined as a set of channel-level indicators quantified from the acquisition quality status of the space electric field sensor channel within the current session window. It includes at least three fields: noise percentage, saturation ratio, and drift index. The noise percentage field is obtained by statistically analyzing the noise component percentage in the output sequence of the preprocessing unit; the saturation ratio field is obtained by statistically analyzing the cumulative ratio of saturation markers in the channel status flags; and the drift index field is obtained by statistically analyzing the changing trend of the drift compensation field formed during baseline drift correction. The health scoring unit writes the channel health data into a channel health record and archives the channel health record along with the session number, channel number, and unified time axis field. The credibility score is defined as a score field that quantifies the credibility of feature records for subsequent fusion. Its calculation process includes three stages: health constraint mapping, stability calculation, and score synthesis. Health constraint mapping maps noise proportion, saturation proportion, and drift index to channel health penalty factors. Stability calculation calculates the feature stability field from the fluctuation amplitude of the mean, variance, spectral centroid, and wavelet packet energy distribution fields in adjacent windows from the feature records. Score synthesis combines the channel health penalty factor and the feature stability field to obtain the credibility score. Understandably, when the timestamp alignment status field is marked as mismatched or failed, the health scoring unit writes an alignment penalty flag into the credibility score and uses this flag, along with candidate credibility flags, to participate in gating discrimination, forming an automatic weighting processing link for time-inconsistent scenarios.
[0066] Specifically, feature credibility gating is executed by a gating unit. The gating unit selectively outputs values to feature records based on credibility scores, candidate credibility markers, and missing data markers. The gating output is defined as dividing the feature records into a gating pass feature set and a gating reject feature set, and writing these into the gating result field. During gating, the gating unit follows a minimum set constraint, which refers to the necessary set of fields supporting the core logic of this invention. This set includes mean, variance, skewness, kurtosis, spectral centroid, bandwidth, spectral entropy, and wavelet packet energy distribution fields, as well as their corresponding session number, channel number, segment number, and window start timestamp fields. When the gating process detects a missing data marker for any segment corresponding to any field in the minimum set constraint and the credibility score is below the threshold, the gating unit writes the segment into the gating reject feature set and registers the rejection reason field. The rejection reason field includes the missing data reason, alignment penalty marker, and a summary of the channel health penalty factor. For the preferred extended function fields, the gating unit allows appending output when the confidence score meets the threshold and the channel health has no abnormal records. These preferred extended function fields include enhanced index information for the spectrum valid label field and the anchor point association field, used for sample selection and backtracking during subsequent model training. The gating unit associates the gating result field with the acquisition session configuration version number and writes it to the gating log record. The gating log record is used for redundancy resolution records and ledger version number tracing in the subsequent segment optimal feature ledger.
[0067] Specifically, environmental parameter fusion is performed by the fusion unit. After the gating is generated through the feature set, the fusion unit obtains the environmental parameter sequence corresponding to the environmental sensor list and performs time alignment. The environmental parameter sequence includes temperature and humidity sequences, and their timestamps are mapped to the window start timestamp field consistent with the segmented slice set according to a unified time axis field. The fusion unit extracts the temperature and humidity sequence segments of the corresponding window based on the segment number and window start timestamp field output by the feature confidence gating, generates environmental parameter segments, and writes them into the environmental parameter record. The environmental parameter record fields include session number, window start timestamp field, temperature sequence statistical value field, and humidity sequence statistical value field. The temperature and humidity sequence statistical value fields are obtained by statistically analyzing the mean and fluctuation amplitude of the environmental parameters within the window. The fusion unit performs field-level concatenation of the gating control through the feature set and environmental parameter records, based on the session number, channel number, segment sequence number, and window start timestamp fields, to generate a high-dimensional feature vector. The high-dimensional feature vector fields include time-domain statistical feature fields, frequency-domain spectral feature fields, time-frequency energy feature fields, channel health fields, credibility score fields, and environmental parameter fields, and are written into the high-dimensional feature vector record. Understandably, when the environmental sensor list contains offline or missing data markers, the fusion unit writes the environmental parameter record into the environmental missing data marker field and simultaneously reduces the environmental weight marker of the credibility score. The environmental weight marker is used for threshold interval mapping table matching and correction paths during subsequent consistency verification, while retaining the core field groups of the high-dimensional feature vector to maintain the input continuity for subsequent steps calculating feature discrimination by segment.
[0068] The output of this step is a high-dimensional feature vector record, whose fields include session number, channel number, segment number, window start timestamp, window end timestamp, mean, variance, skewness, kurtosis, spectral centroid, bandwidth, spectral entropy, wavelet packet energy distribution, noise ratio, saturation ratio, drift index, credibility score, temperature sequence statistics, humidity sequence statistics, gating result, and acquisition session configuration version number. The high-dimensional feature vector record serves as the direct input for subsequent steps to calculate feature discrimination based on the high-dimensional feature vector and historical samples with segment labels, and to perform redundancy resolution. The gating log record association information and acquisition session configuration version number field in the high-dimensional feature vector record serve as the redundancy resolution record and ledger version number traceability input for the subsequent step of constructing the optimal feature ledger for each segment. Simultaneously, the credibility score field and channel health field group serve as reference input fields in the subsequent step's consistency verification output of resampling instructions or alarm escalation strategy instructions, participating in the resampling trigger judgment.
[0069] This step extracts multi-domain features from segmented slices and combines channel health and credibility scores to complete feature credibility gating. Then, it fuses these features with environmental parameters to generate high-dimensional feature vector records, forming a unified input data structure for constructing the optimal feature ledger for subsequent segments and training dual models.
[0070] S400: Based on high-dimensional feature vector records, feature discrimination is calculated by segment and redundancy resolution is performed to generate the segment-optimal feature ledger.
[0071] Figure 2 The execution flowchart of step S400 provided in the embodiments of this application mainly includes the following six key steps:
[0072] Triggered execution: The ledger construction module starts the process after certain conditions are met, including incrementing the collection session configuration version number, starting the model training task, or receiving a resampling instruction to trigger sample supplementation.
[0073] Input source preparation: Prepare two core input data sets, namely, high-dimensional feature vector records containing multi-domain features and environmental parameters from previous steps, and paired historical samples with segment labels that indicate the discharge development stage.
[0074] Sample alignment check: Before calculation, the input data is strictly validated to ensure that the acquisition session configuration version number is consistent, the gating result is valid and there are no key data missing. The check conclusion will be written into the sample validation flag field.
[0075] The discrimination index is calculated by segment: The discrimination index calculation unit groups historical samples by segment label and calculates the discrimination index (such as inter-class distance and intra-class divergence ratio) for each feature field (such as mean, variance, spectral centroid, etc.). This process introduces the credibility score as the sample weight and selects effective samples based on channel health constraints.
[0076] Redundancy Reduction: Based on the calculated discrimination index, the redundancy reduction unit identifies and removes feature fields that are repetitive or highly correlated. While ensuring the minimum set of fields (such as mean, variance, spectral centroid, wavelet packet energy distribution), it forms the optimal feature subset of the segment and generates a complete redundancy reduction record.
[0077] Constructing the segment-optimal feature ledger: The ledger generation unit integrates segment labels, discrimination index records, redundancy resolution records, and the segment-optimal feature subset to generate a structured segment-optimal feature ledger, assigns it a ledger version number, and associates it with the version number configured in the collection session to form a traceable and versionable feature entry standard.
[0078] The above processes are interconnected, ultimately producing a standardized feature ledger that supports subsequent model training and online inference.
[0079] Specifically, this step is triggered by the ledger construction module after the high-dimensional feature vector records are continuously added to the database and the gating result field reaches a usable state. The triggering conditions include the increment of the collection session configuration version number, the arrival of the model training task start marker, or the consistency check output resampling instruction triggering sample supplementation to enter the ledger incremental update path. The input source for this step includes high-dimensional feature vector records and historical samples with segment labels. The high-dimensional feature vector records are the output obtained from the previous step by fusing environmental parameters based on feature credibility gating. The fields include at least session number, channel number, segment number, window start timestamp, mean, variance, skewness, kurtosis, spectral centroid, bandwidth, spectral entropy, wavelet packet energy distribution, noise ratio, saturation ratio, drift index, credibility score, temperature sequence statistics, humidity sequence statistics, gating result, and acquisition session configuration version number. The historical samples with segment labels are a set of paired samples of high-dimensional feature vector records and segment labels corresponding to the historical acquisition session configuration. The segment labels are the label fields obtained by manually or offline rule-based annotation of the stage anchor point candidates and their adjacent segments of the spatial electric field waveform of the entire discharge process. Understandably, historical samples with segment labels and high-dimensional feature vector records are aligned in the window start timestamp field, segment number, and session number dimensions. The ledger construction module performs a sample alignment check before entering the discrimination calculation. The alignment check includes a consistency check of the collection session configuration version number, a validity check of the gating result field, and a check to avoid missing test markers. The check results are written into the sample verification marker field, which serves as one of the source fields for subsequent redundant resolution records.
[0080] Specifically, the feature discrimination index is calculated by segment, which is performed by the discrimination index calculation unit. The discrimination index calculation unit groups the historical samples with segment labels according to the segment labels, generates segment group sample sets, and calculates the discrimination index for each feature field based on the segment group sample sets. The discrimination index is defined as a quantitative result that measures the ability of a feature field to separate different segment labels. In this embodiment, the discrimination index includes at least two paths: inter-class distance and intra-class divergence ratio or mutual information. The inter-class distance and intra-class divergence ratio paths are generated by the discrimination index calculation unit by calculating the difference in segment mean values for each segment group sample set and combining it with the degree of dispersion within the segment to generate a ratio field. The mutual information path is generated by the discrimination index calculation unit by the degree of correlation between the distribution of feature field values and the distribution of segment labels to generate a mutual information field. The discrimination calculation unit calculates discrimination indices for each of the fields: mean, variance, skewness, kurtosis, spectral centroid, bandwidth, spectral entropy, and wavelet packet energy distribution. During the calculation, a credibility score is introduced as a sample weight field, mapped from the credibility score, with the mapping rules written into the weight mapping record field. Simultaneously, the discrimination calculation unit uses noise percentage, saturation ratio, and drift index as channel health constraint fields to screen valid samples for discrimination calculation. The screening criteria are that samples with channel health constraint fields below a threshold and a passing gating result field enter the main calculation path. Samples that fail the screening are written into a rejection flag field, and the rejection reason field is recorded. Furthermore, the discrimination calculation unit calculates discrimination indices for the temperature sequence statistics field and the humidity sequence statistics field and writes them into the environmental discrimination field. This is used for the subsequent version evolution of the optimal feature subset of the segment under different environmental windows. The environmental discrimination field and the acquisition session configuration version number field are jointly written into the discrimination index record.
[0081] Specifically, redundancy elimination is performed by the redundancy elimination unit. After the discrimination index record is generated, the redundancy elimination unit performs redundancy identification and screening on the feature fields of the high-dimensional feature vector record. Redundancy identification is defined as determining that multiple feature fields have information duplication or high correlation under the same acquisition session configuration version number. Screening is defined as removing redundant fields and forming the optimal feature subset of the segment while retaining the fields with higher discrimination index. The redundancy resolution unit first establishes a candidate feature pool, which consists of feature fields whose discrimination index reaches a threshold in the discrimination index record. The candidate feature pool at least covers the fields of mean, variance, skewness, kurtosis, spectral centroid, bandwidth, spectral entropy, and wavelet packet energy distribution, and writes the candidate feature pool into the candidate pool record field. Then, redundancy determination is performed, which adopts two paths: correlation threshold determination or mutual information overlap determination. The correlation threshold determination path is where the redundancy resolution unit calculates the correlation of the joint change trend of the feature fields in the candidate feature pool on historical samples with segment labels and compares it with the threshold. The mutual information overlap determination path is where the redundancy resolution unit determines the overlap between the feature fields in the candidate feature pool and the mutual information fields of the segment labels. When the redundancy determination is successful, the redundancy resolution unit retains the top-ranked feature fields according to the discrimination index ranking in the discrimination index record and writes the eliminated fields into the redundancy resolution record. Understandably, to maintain the minimum set constraint of the logical link of this invention, the redundancy elimination unit retains the mean, variance, spectral centroid, and wavelet packet energy distribution fields as minimum set fields when performing screening, and writes the minimum set fields into the minimum set marker field; the skewness, kurtosis, bandwidth, and spectral entropy fields are included as preferred extended fields in the screening candidates. Whether to retain the preferred extended fields is determined by the discrimination index threshold and the redundancy judgment result, and the decision result is written into the preferred extended marker field and the version evolution reason field is recorded. After each round of redundancy elimination, the redundancy elimination unit generates a redundancy elimination summary field. The redundancy elimination summary field includes reference markers for the retained field list, reference markers for the removed fields, redundancy judgment path markers, and threshold version markers. The threshold version markers and the weight mapping record field together form an auditable ledger construction basis.
[0082] Specifically, the segment-optimal feature ledger is constructed and generated by the ledger generation unit. This unit combines segment labels, discrimination index records, redundancy resolution records, and the segment-optimal feature subset to form the segment-optimal feature ledger, and generates the association between the ledger version number and the acquisition session configuration version number. The segment-optimal feature ledger is defined as a feature subset management data structure oriented towards segment labels. Its fields include segment labels, discrimination indexes, redundancy resolution records, the segment-optimal feature subset, the ledger version number, and the acquisition session configuration version number. The segment-optimal feature subset is a set of feature fields obtained under the constraint of the minimum set label field and supplemented by optimized extended label fields. The discrimination index corresponds to the inter-class distance and intra-class divergence ratio or mutual information field of each feature field in each segment group sample set. The redundancy resolution records correspond to the redundancy decision path marker, threshold version marker, removal field reference marker, and retention field reference marker. The ledger generation unit introduces an incremental update strategy when generating the ledger version number. This strategy is defined as follows: when the acquisition session configuration version number remains unchanged and the discrimination index record does not cross a threshold, no new ledger version number is generated; only the redundant resolution summary field of the runtime log is appended. When the acquisition session configuration version number increments or the environmental discrimination field crosses a threshold, a new ledger version number is generated and written to the version increment reason field. The version increment reason field includes a summary of the acquisition session configuration version number change flag, the environmental discrimination field change flag, and the sample verification flag field. Understandably, when generating the optimal feature ledger for a segment, the ledger generation unit simultaneously writes a reference index field. This reference index field is used to quickly locate the optimal feature subset of the corresponding segment label when subsequently training and developing the segment identification model and the contamination prediction model in parallel. The reference index field is also associated with the gating log record for archiving, allowing for backtracking of resampling instructions during consistency verification.
[0083] In an engineering example, an outdoor support insulator at a substation operates under alternating foggy and light rain conditions. The acquisition session management module triggers the acquisition session configuration generation and performs spatial electric field waveform acquisition throughout the discharge process daily. After the feature engineering module forms a high-dimensional feature vector record, the ledger construction module triggers an incremental update path weekly. Historical samples with segment labels are grouped by segment labels, and the discriminative indexes of fields such as mean, variance, spectral centroid, spectral entropy, and wavelet packet energy distribution are calculated and recorded. When the humidity sequence statistical value field changes across thresholds and the environmental discriminative field changes significantly, the ledger generation unit generates a new ledger version number and associates it with a new acquisition session configuration version number field. At the same time, in the redundant resolution record, some frequency domain spectral feature fields that are highly redundant with the wavelet packet energy distribution are removed. The mean, variance, spectral centroid, and wavelet packet energy distribution are retained as minimum set fields, and skewness and kurtosis are selected as preferred extended fields according to the threshold version label, forming a segment-optimal feature ledger that can be directly used for subsequent model training.
[0084] The output of this step is the segment optimal feature ledger. The segment optimal feature ledger fields include segment label, discrimination index, redundancy resolution record, segment optimal feature subset, ledger version number, and acquisition session configuration version number. Among them, the segment optimal feature subset serves as the feature entry reference field for the subsequent steps of parallel training and development of segment identification model and contamination prediction model. The discrimination index and redundancy resolution record serve as the basis fields for the subsequent steps of model training sample selection and consistency verification threshold mapping table maintenance. The ledger version number and acquisition session configuration version number serve as the version matching condition when inputting real-time feature vectors online in the subsequent steps and participate in the tracing of resampling instructions or alarm escalation strategy instructions.
[0085] This step calculates segment-based discriminant resolution and eliminates redundancy by performing segment-based resolution on high-dimensional feature vector records in the dimension of historical samples with segment labels, forming a segment-optimal feature ledger with version numbers and traceable records, and providing a stable feature entry point and version association for subsequent model training and online consistency verification.
[0086] In one embodiment, in S400, based on the high-dimensional feature vector and historical samples with segment labels, the feature discrimination degree is calculated by segment and redundancy resolution is performed to form the segment-optimal feature ledger.
[0087] This step is executed by the ledger construction module, with input sources being high-dimensional feature vector records and historical samples with segment labels. The high-dimensional feature vector records are derived from the environmental parameter fusion output of the previous step, and their fields include session ID, channel ID, segment sequence number, window start timestamp, mean, variance, skewness, kurtosis, spectral centroid, bandwidth, spectral entropy, wavelet packet energy distribution, noise percentage, saturation ratio, drift index, confidence score, temperature sequence statistics, humidity sequence statistics, gating result, and acquisition session configuration version number. The historical samples with segment labels are a set of paired high-dimensional feature vector records and segment labels, with the segment labels corresponding to the annotation results of anchor point candidates for the discharge process stage. The ledger construction module first performs sample alignment checks on the input data, including acquisition session configuration version number consistency checks, gating result field validity checks, and missing measurement marker avoidance checks, generating sample verification marker fields. Subsequently, the discrimination calculation unit groups the historical samples with segment labels according to the segment labels, forming segment-grouped sample sets. For each feature field (such as mean, variance, etc.), its discrimination index needs to be calculated to quantify the ability to separate segment labels.
[0088] Formula (1) is used to calculate the intra-class scatter matrix, which serves as the basis for discrimination evaluation: for segment label grouping ( , (Total number of segments), feature fields ( Fields taken from high-dimensional feature vector records, such as mean and variance, have their within-class scatter matrix. Defined as the covariance matrix of the feature values of the samples within this segment.
[0089]
[0090] in, Indicates a section The number of samples; It is a section The Middle Each sample in the feature field The value of ; It is a section Feature fields The sample mean; The scatter matrix is the intra-class scatter matrix. For sample index; Section The number of samples; Feature field index; : Segment index; : Transpose operator, which represents the transpose of a matrix or vector.
[0091] Data source mapping: Extracting feature fields from high-dimensional feature vector records. The value of is denoted as Segment labels are extracted from historical samples with segment labels and then grouped to obtain... and Formula (1) This serves as input for subsequent discrimination calculations. Simple numerical example: Assuming a segment... have One sample, feature fields The mean is taken as the value. If the range is [0.1, 0.2, ..., 0.5], then... Calculated It is a 2×2 matrix (if the features are multidimensional, we assume a single feature here, and the divergence is a scalar variance of about 0.02).
[0092] Formula (2) calculates the inter-class distance and the intra-class divergence ratio based on the intra-class divergence in Formula (1), which serves as the core metric for discrimination. Formula (2) uses the output of Formula (1). In conjunction with inter-class divergence, the ratio of inter-class distance to intra-class divergence is defined. :
[0093]
[0094] in, The inter-class scatter matrix; Represents the trace of a matrix; As a distinguishing indicator; : Intra-class scatter matrix.
[0095] Data source mapping: derived from formula (1) Aggregate And extracted from high-dimensional feature vector records. and Formula (2) As a component of the discrimination index record. Simple numerical example: assuming feature fields... exist In each segment, the sum of scatter traces within a class is 0.1, and the sum of scatter traces between classes is 0.5. Then... .
[0096] This section outputs the discrimination index records, including These fields are input to the subsequent redundancy resolution unit for redundancy determination.
[0097] Following the aforementioned discrimination index, the redundancy elimination unit performs redundancy identification and screening. Formula (3) is used for redundancy determination and to calculate the correlation threshold between feature fields.
[0098] Formula (3) directly uses the discrimination index of Formula (2). As the basis for weighting, feature fields are defined. and correlation coefficient :
[0099]
[0100] in, This represents the total number of samples; It is the first Each sample in the feature field The value of ; For feature fields The global mean; The correlation coefficient; For sample index; Feature field index.
[0101] Data source mapping: Extracting feature fields from high-dimensional feature vector records. and The sequence of values is denoted as and Formula (3) Used for redundancy determination, when ( Redundancy is marked when the threshold is 0.8. Formula (3) depends on formula (2). As a priority reference for feature selection. Simple numerical example: Feature fields and The value sequences are [0.1, 0.2, 0.3] and [0.15, 0.25, 0.35], respectively. , Calculated , indicating high redundancy.
[0102] Formula (4) optimizes feature selection based on the discriminant index of Formula (2) and the correlation of Formula (3), constructing the optimal feature subset for each segment. Formula (4) employs a linear programming model to maximize overall discriminant strength while minimizing redundancy. The objective function is:
[0103]
[0104] Constraints: ,
[0105] in, The total number of feature fields. Select a binary variable (1 indicates selection). Minimum set constraint (e.g.) (corresponding to mean, variance, spectral centroid, and wavelet packet energy distribution). This is the redundancy penalty coefficient; From formula (2), From formula (3).
[0106] Data source mapping: From formula (2) From formula (3), we get Substitute this into the optimization model. The output of formula (4) Define the optimal feature subset of a segment. Simple numerical example: Let... , , The matrix shows high redundancy in features 1 and 3; after optimization... Select features 1, 2, and 4.
[0107] This section outputs redundant resolution records and the optimal feature subset of the segment, which are input into the ledger generation unit for ledger construction.
[0108] Following the aforementioned discriminant index records and redundancy resolution records, the ledger generation unit constructs the segment-optimal feature ledger. Formula (5) is used for ledger version number generation, based on the version number and environment discriminant field configured in the acquisition session. Formula (5) defines the version increment condition:
[0109]
[0110] in, and For the old and new ledger version numbers; Configure the version number change amount for the data collection session; This represents the change in the environmental discrimination field. For the threshold; and The data is sourced from the version number field and environment differentiation field in the data collection session configuration. : Logical OR operator.
[0111] Data source mapping: extracted from the version number field configured in the collection session. Extracted from the environmental discrimination field in the discrimination index record. The output of formula (5) As a ledger version number. Simple numerical example: Old version number , , , ,but .
[0112] Formula (6) is used for ledger construction, integrating segment labels, distinguishability indicators, redundant resolution records, optimal feature subsets of segments, and version numbers. Formula (6) defines the ledger data structure as tuples:
[0113]
[0114] in, The segment-optimal feature ledger data structure represents a tuple containing multiple components; For a set of segment labels, For the discrimination index record (including formula (2)) ), For redundant elimination records (including formula (3)) And formula (4) ), The optimal feature subset of the segment. From formula (5), Configure a version number for the data collection session.
[0115] Data source mapping: Directly integrate the aforementioned output. Formula (6) This is the optimal feature ledger for the segment. A simple numerical example: , , , , , .
[0116] This step forms a versioned feature ledger by segmented discrimination calculation and redundancy resolution, providing a stable feature entry point; the formula chain ensures that the discrimination index drives redundancy judgment and optimizes feature subsets; the ledger construction integrates version management to support the consistency of subsequent model training.
[0117] S500, based on the optimal feature ledger of the segment, performs parallel training and version management of the segment identification model and the pollution degree prediction model, online inference and consistency verification, and generates alarm escalation strategy instructions.
[0118] Specifically, this step is triggered by the integrated model training and online inference module. Its input sources include the segment-optimal feature ledger, high-dimensional feature vector records, and historical samples with segment labels. The segment-optimal feature ledger includes at least the segment label, discrimination index, redundancy resolution record, segment-optimal feature subset, ledger version number, and acquisition session configuration version number. The high-dimensional feature vector records include at least the session number, channel number, segment sequence number, window start timestamp field, mean, variance, skewness, kurtosis, spectral centroid, bandwidth, spectral entropy, wavelet packet energy distribution, noise ratio, saturation ratio, drift index, credibility score, temperature sequence statistics field, humidity sequence statistics field, gating result field, and acquisition session configuration version number field. The historical samples with segment labels contain the pairing relationship between high-dimensional feature vector records and segment labels. The integrated model training and online inference module starts the training path when it detects an increase in the ledger version number, the addition of new historical samples with segment labels reaching the batch threshold, or a consistency check triggering a resampling instruction that causes an update in the sample distribution. At the same time, it only starts the online inference path when the collection session configuration version number remains unchanged and the ledger version number does not change. The triggering reason is registered as a training trigger flag field or an inference trigger flag field and written to the session running log for subsequent traceability verification.
[0119] Specifically, parallel training is implemented by the training orchestration unit. The training orchestration unit parses the optimal feature subset of the segment in the segment optimal feature ledger into feature entry mapping records. The feature entry mapping records contain feature field reference markers, field value position indices, and ledger version numbers, which serve as the source of field constraints for training sample assembly. The training orchestration unit then reads the sample set that matches the ledger version number from the historical samples with segment labels. Based on the feature entry mapping records, it extracts fields such as mean, variance, skewness, kurtosis, spectral centroid, bandwidth, spectral entropy, wavelet packet energy distribution, temperature sequence statistics, and humidity sequence statistics from the high-dimensional feature vector records to generate a training feature matrix, and writes the segment labels into the classification label vector field. The development segment identification model is defined as a classification model that takes a training feature matrix as input and outputs development segments. Internally, it consists of a feature normalization processing submodule, a classifier backbone submodule, and an output decoding submodule. The feature normalization processing submodule reads the training feature matrix and performs scale normalization registration. The classifier backbone submodule reads the normalization result and generates a segment probability vector field. The output decoding submodule maps the segment probability vector field to a development segment field and registers a confidence label field. The contamination prediction model is defined as a regression model that takes a training feature matrix as input and outputs contamination prediction values. Internally, it consists of a feature normalization processing submodule, a regressor backbone submodule, and an output calibration submodule. The regressor backbone submodule outputs a continuous value vector field. The output calibration submodule maps the continuous value vector field to a contamination prediction value field and registers a value caliber label field. The training orchestration unit performs parallel training scheduling for the two types of models. The training process is recorded as a model training record, which includes a training data summary field, ledger version number, acquisition session configuration version number, feature entry mapping record reference marker, parameter version number, and training completion marker field. The parameter version number is generated by the version management submodule. The version management submodule writes the ledger version number and acquisition session configuration version number together into the parameter version number association field, and writes the success or failure status into the training completion marker field. The failure status also registers the failure reason field and rolls back to the previous parameter version number, forming an automatic rollback path.
[0120] Furthermore, the online input of real-time feature vectors is implemented by the online inference unit. The online inference unit receives high-dimensional feature vector records from the output of the preceding steps, performs version matching and feature assembly. Version matching includes the association verification of the acquisition session configuration version number field and the ledger version number field, and the verification result is written to the version matching flag field. Feature assembly includes extracting the corresponding fields of the optimal feature subset of the segment based on the feature entry mapping record to form a real-time feature vector, and writing the confidence score and gating result fields into the inference context field. The online inference unit calls the development segment identification model to output the development segment field and the confidence flag field, and calls the contamination prediction model to output the contamination prediction value field and the value caliber flag field. Simultaneously, the two types of outputs are associated with the session number, channel number, and window start timestamp fields and written into the inference output record. The inference output record serves as the direct input for subsequent consistency verification. Understandably, when the version matching flag field is in a mismatch state or the gating result field is in a rejection state, the online inference unit writes the real-time feature vector into the inference pending review queue field and triggers the consistency check to enter the conservative discrimination path. The conservative discrimination path writes the low confidence flag field into the inference output record and retains the inference context field for the resampling instruction generation call.
[0121] Specifically, consistency verification is performed by a consistency verification unit. The consistency verification unit reads the inference output record and calls the development segment to contamination threshold interval mapping table to perform mapping table matching. The development segment to contamination threshold interval mapping table is defined as a rule data structure that maps the development segment field to the contamination threshold interval field. Its fields include the development segment field, the contamination threshold interval field, the mapping table version number field, and the ledger version number association field. The consistency verification unit retrieves the contamination threshold interval field based on the development segment field, then performs an interval comparison between the contamination prediction value field and the contamination threshold interval field, outputs the consistency verification result field, and registers the inconsistency reason field. The inconsistency reason field includes the version mismatch flag, the low confidence flag, the interval out-of-bounds flag, and the channel health summary field. Based on the consistency verification result fields, the consistency verification unit generates a resampling instruction or an alarm escalation strategy instruction. The resampling instruction is defined as a data structure for the acquisition adjustment instruction to the acquisition session management module. Its fields include session number, channel number, target sampling rate, synchronous trigger strategy registration reference flag, resampling duration window, and resampling reason field. The resampling instruction is then sent to the acquisition session configuration update channel, triggering the reconfiguration process of sampling rate registration and synchronous trigger strategy registration. The alarm escalation strategy instruction is defined as a data structure for the strategy instruction to the alarm processing module. The alarm escalation strategy instruction includes session number, development segment field, pollution level prediction value field, alarm level field, strategy version number field, and action reference flag. The alarm escalation strategy instruction is written to the alarm event queue for subsequent system to orchestrate action execution according to the strategy version number field.
[0122] In the engineering implementation, the list of spatial electric field sensors and the list of environmental sensors are deployed around the outdoor support insulators of the substation. The acquisition session management module outputs high-dimensional feature vector records aligned with a unified time axis and writes them into the optimal feature ledger matching index of the section. The training orchestration unit assembles historical samples with section labels to form a training feature matrix as the ledger version number increments, and completes the parameter version number registration of the development section identification model and the pollution degree prediction model. The online inference unit reads the real-time feature vector and outputs the inference output record when each sliding time window arrives. The consistency verification unit completes the interval comparison according to the mapping table between the development section and the pollution degree threshold interval. When a version mismatch mark or interval out-of-bounds mark appears, a resampling instruction is generated and the acquisition session configuration is triggered to update. At the same time, when the number of consecutive inconsistencies reaches the policy threshold, an alarm escalation policy instruction is generated and written into the alarm event queue. The relevant logs are synchronously written into the session operation log and the model training record, forming a closed-loop traceability link from the ledger version number to the parameter version number and then to the inference output record.
[0123] This step achieves synchronous output of development segments and pollution prediction values through parallel training and online inference driven by the segment-optimal feature ledger. It also generates resampling instructions or alarm escalation strategy instructions through consistency verification, completing the closed-loop connection with the previous feature ledger construction and subsequent collection session configuration updates.
[0124] In one embodiment, S500, a segment identification model and a contamination prediction model are trained in parallel based on the segment optimal feature ledger. Real-time feature vectors are input online, and the predicted values of the developed segments and contamination levels are output synchronously. Consistency checks are performed, and a resampling instruction or an alarm escalation strategy instruction is output based on the consistency check results.
[0125] This step is triggered by the integrated model training and online inference module. The input source includes the segment-optimal feature ledger, high-dimensional feature vector records, and historical samples with segment labels. The segment-optimal feature ledger provides segment labels, discrimination index, redundancy resolution records, segment-optimal feature subsets, ledger version number, and acquisition session configuration version number. The high-dimensional feature vector records provide session number, channel number, segment sequence number, window start timestamp field, mean, variance, skewness, kurtosis, spectral centroid, bandwidth, spectral entropy, wavelet packet energy distribution, noise ratio, saturation ratio, drift index, confidence score, temperature sequence statistics field, humidity sequence statistics field, gating result field, and acquisition session configuration version number field. The historical samples with segment labels provide the pairing relationship between high-dimensional feature vector records and segment labels. The integrated model training and online inference module initiates the training path when it detects an incrementing ledger version number, the addition of new historical samples with segment tags reaching the batch threshold, or a consistency check triggering a resampling instruction. Otherwise, it initiates the online inference path and registers the triggering reason as either the training trigger flag field or the inference trigger flag field, writing it to the session runtime log. The online inference unit first performs version matching to verify the consistency between the collection session configuration version number and the ledger version number.
[0126] Formula (7) is used in the version matching function to calculate the matching flag:
[0127]
[0128] in, The matching flag has a value range of {0, 1}, where 1 indicates a successful match. Configure a version number for the data collection session; For ledger version number; This is an indicator function that outputs 1 when the condition is true and 0 otherwise.
[0129] Data source mapping: The acquisition session configuration version number is extracted from the high-dimensional feature vector record and denoted as... The ledger version number is extracted from the optimal feature ledger of the segment and recorded as . Together, they form the input to formula (7). A simple numerical example: , ,but If a match fails, the online inference unit writes the real-time feature vector into the inference pending review queue field; if a match succeeds, feature assembly continues. Feature assembly extracts the corresponding fields of the optimal feature subset of the segment from the high-dimensional feature vector record based on the feature entry mapping record, forming the real-time feature vector.
[0130] Formula (8) is used for the feature assembly function to define the construction of real-time feature vectors:
[0131]
[0132] in, For real-time feature vectors; The feature matrix is a record of high-dimensional feature vectors; The binary mask vector for the feature entry mapping record takes the value range {0,1}, where 1 indicates that the feature is selected and comes from the segment optimal feature subset of the segment optimal feature ledger; dot product represents element-wise multiplication.
[0133] Data source mapping: The feature matrix is extracted from the high-dimensional feature vector record and denoted as... The mask vector is extracted from the feature entry mapping record and denoted as... Together, they form the input to formula (8). A simple numerical example: (Corresponding to mean, variance, and skewness) ,but Matching markers of formula (7) Used to control the execution conditions of formula (8), when Assembly is performed in real time. This section outputs real-time feature vectors. It is input into the prediction function of the subsequent development segment identification model and the pollution degree prediction model.
[0134] Following the aforementioned real-time feature vectors, the online inference unit invokes the development segment identification model and the pollution level prediction model to perform parallel predictions. The development segment identification model is a classification model, outputting the probability of the development segment; the pollution level prediction model is a regression model, outputting the predicted pollution level value.
[0135] Formula (9) is used to develop the prediction function of the segment identification model and calculate the segment probability vector:
[0136]
[0137] in, This represents the segment probability vector; This is the weight matrix for the classification model; This is the bias vector for the classification model; The softmax function normalizes the input vector into a probability distribution; The output from formula (8).
[0138] Data source mapping: Real-time feature vectors derived from formula (8) As input, combined with pre-trained parameters and Calculate the probability. Simple numerical example: , , The linear output is [0.11, 0.14], after softmax. .
[0139] Formula (10) is used as the prediction function of the filth degree prediction model to calculate the filth degree value:
[0140]
[0141] in, This is a predicted value for the level of filth. This represents the weight vector of the regression model; This is the bias scalar of the regression model; The output from formula (8).
[0142] Data source mapping: Real-time feature vectors derived from formula (8) As input, combined with pre-trained parameters and Calculate the predicted value. Simple numerical example: , , ,but The segment probability vector of formula (9) Decoding the development segment field used for subsequent consistency verification, the predicted contamination value of formula (10) Used for verification and comparison. This section outputs the development zone field and the pollution level prediction value field, which are input by the consistency verification unit.
[0143] Following the aforementioned development segment field and pollution level prediction value field, the consistency verification unit performs mapping table matching and interval comparison. The consistency verification unit calls the development segment and pollution level threshold interval mapping table, which contains the development segment field, pollution level threshold interval field, mapping table version number field, and ledger version number association field.
[0144] Formula (11) is used in the consistency check function to calculate the consistency check result:
[0145]
[0146] in, For the verification result, the value range is {0,1}, where 1 indicates that the predicted value is within the threshold range; The predicted filth level is derived from the output of formula (10); This represents the dirtiness threshold range; This is an indicator function.
[0147] Data source mapping: the predicted foulness value from formula (10) Compare with the threshold range retrieved from the mapping table. Simple numerical example: If the threshold range is [0.3, 0.5], then... If the verification fails, register the inconsistency reason field, such as version mismatch flag or range out-of-bounds flag.
[0148] Formula (12) is used in the complex sampling instruction generation function to define the instruction triggering condition:
[0149]
[0150] in, This is a flag for triggering a complex sampling instruction, with a value range of {0,1}, where 1 indicates a trigger instruction. Output from formula (11); This represents the number of consecutive verification failures. The policy threshold takes the value of a positive integer. For logical AND operator.
[0151] Data source mapping: Verification results from formula (11) And the calculation of historical failure counts. Simple numerical example: , , ,but .when If the condition is met, a resampling instruction is generated, which includes the target sampling rate field and the synchronous trigger strategy registration reference flag; otherwise, an alarm escalation strategy instruction is generated.
[0152] Verification results of formula (11) It is directly used for decision-making in formula (12). This part outputs a resampling instruction or an alarm escalation strategy instruction, which is input by the acquired session management module or the alarm processing module. In the engineering embodiment, the substation outdoor post insulator monitoring system performs online reasoning when the sliding time window arrives: after the real-time feature vector is assembled by formula (8), the development section and pollution degree value are output by formulas (9) and (10); when formula (11) checks and finds that the pollution degree value exceeds the limit, formula (12) triggers a resampling instruction to adjust the acquisition configuration.
[0153] This step achieves a closed-loop operation from version matching to instruction generation through formula chain, ensuring the coordination of online reasoning and consistency verification; the feature assembly of formula (8) ensures the consistency of feature entry, and the verification and instruction generation of formulas (11) and (12) enhance the system's adaptive capability.
[0154] As described above, although the invention has been shown and described with reference to specific preferred embodiments, it should not be construed as limiting the invention itself. Various changes in form and detail may be made without departing from the spirit and scope of the invention as defined in the appended claims.
Claims
1. A method for identifying insulator flashover and predicting pollution level based on multi-feature fusion, characterized in that, Includes the following steps: The system acquires insulator structure parameters, a list of space electric field sensors, and a list of environmental sensors. It then binds probe numbers, registers geometric relationships, and calibrates channels for these sensors. Finally, it performs sampling rate registration, synchronous triggering strategy registration, unified time axis generation, and timestamp alignment to generate the acquisition session configuration. Based on the acquisition session configuration, the spatial electric field waveform of the entire discharge process is accessed and time-aligned, and noise reduction and normalization, sliding time window segmentation and stage anchor point candidate extraction are performed to generate a segmented slice set and a stage anchor point candidate set. The time-domain statistical features, frequency-domain spectral features, and time-frequency energy features of the segmented slice set and the stage anchor point candidate set are extracted, and channel health and credibility score calculation, feature credibility gating, and environmental parameter fusion processing are performed to generate high-dimensional feature vector records. The channel health and credibility score calculation includes: calculating the noise ratio, saturation ratio, and drift index as channel health fields based on the waveform sequence, and calculating the credibility score by combining the channel health fields and the fluctuation amplitude of features in adjacent windows. The feature credibility gating process includes: performing gating discrimination on feature records based on the credibility score, candidate credibility markers of stage anchor candidates, and feature missing markers, and outputting a gating pass feature set; The environmental parameter fusion processing includes: obtaining temperature and humidity sequences based on the environmental sensor list, and aligning and splicing them with the gating control according to the feature set and time window; Calculate the feature discrimination of predetermined fields in the high-dimensional feature vector record according to the segment and perform redundancy elimination processing to generate the segment-optimal feature ledger; Based on the optimal feature ledger of the aforementioned segment, parallel training and version management, online inference and consistency verification of the segment identification model and the pollution degree prediction model are performed to generate alarm escalation strategy instructions.
2. The method for insulator flashover identification and pollution degree prediction based on multi-feature fusion according to claim 1, characterized in that, The probe number binding, deployment geometry registration, and channel calibration processing specifically include: Generate a channel number table containing channel number, sensor identifier, and installation point identifier based on the list of space electric field sensors; Based on the insulator structural parameters and installation point identification, calculate and register the distance and orientation parameters of each channel relative to the insulator reference point to form a channel geometric binding record; Zero-point bias measurement, gain coefficient calibration, and saturation threshold registration are performed on the space electric field sensor channel.
3. The insulator flashover identification and pollution degree prediction method based on multi-feature fusion according to claim 1, characterized in that, The execution sampling rate registration, synchronization triggering strategy registration, unified timeline generation, and timestamp alignment processing specifically include: Based on the monitoring task type and sensor range, write the sampling rate field and the segmented slice window length field to generate a synchronous trigger strategy field containing trigger source identifier, trigger source type, and trigger delay compensation value; A unified time axis field is generated based on the clock status of each channel, and clock offset estimation, offset compensation and alignment verification are performed on each channel. A collection session configuration is generated, which includes session number, channel number table, distance parameter, azimuth parameter, calibration parameter, sampling rate field, synchronization trigger strategy field, unified time axis field, timestamp alignment status field, segmented slice window length field, sliding step size field and collection session configuration version number.
4. The method for insulator flashover identification and pollution degree prediction based on multi-feature fusion according to claim 1, characterized in that, The specific processing steps of denoising and normalization, sliding time window segmentation, and stage anchor point candidate extraction include: The time-aligned waveform is sequentially subjected to filtering, baseline drift correction, outlier repair, and amplitude normalization to obtain a preprocessed sequence. The preprocessed sequence is truncated based on the segmented slice window length field and the sliding step field to generate a segmented slice set containing the segment number, window start timestamp, window end timestamp and segmented slice data; For each segment slice, calculate the rising edge mutation index, local high-frequency energy surge index, and duration index, and generate a candidate set of stage anchor points containing candidate time window numbers.
5. The insulator flashover identification and pollution degree prediction method based on multi-feature fusion according to claim 1, characterized in that, The high-dimensional feature vector record includes: The fields include session number, channel number, segment number, window start timestamp, mean, variance, skewness, kurtosis, spectral centroid, bandwidth, spectral entropy, wavelet packet energy distribution, noise percentage, saturation ratio, drift index, credibility score, temperature sequence statistics, humidity sequence statistics, gating result, and acquisition session configuration version number.
6. The method for insulator flashover identification and pollution degree prediction based on multi-feature fusion according to claim 1, characterized in that, The step of calculating the feature discrimination of predetermined fields in the high-dimensional feature vector record by segment and performing redundancy elimination specifically includes: Historical samples with segment labels are grouped according to segment labels. Inter-class distance and intra-class divergence ratio or mutual information are calculated for the fields of mean, variance, skewness, kurtosis, spectral centroid, bandwidth, spectral entropy and wavelet packet energy distribution as discrimination indicators. The credibility score is introduced as a sample weight in the calculation process. Valid samples are selected based on noise ratio, saturation ratio and drift index. A candidate feature pool is established based on the discriminant index. Redundancy is determined by correlation threshold or mutual information overlap. The feature pool is sorted according to the discriminant index and the mean, variance, spectral centroid and wavelet packet energy distribution are retained as minimum set fields to form the optimal feature subset of the segment. A redundancy resolution record containing redundancy determination path markers, threshold version markers, field removal reference markers and field retention reference markers is generated.
7. The insulator flashover identification and pollution degree prediction method based on multi-feature fusion according to claim 1, characterized in that, The optimal feature ledger for the segment includes: Segment labels, discrimination indexes, redundancy resolution records, optimal feature subsets of segments, ledger version numbers, and collection session configuration version numbers.
8. The insulator flashover identification and pollution degree prediction method based on multi-feature fusion according to claim 7, characterized in that, The parallel training and version management includes: parsing the feature entry mapping record containing feature field reference markers from the optimal feature ledger of the segment; assembling the training feature matrix from historical samples with segment labels based on this record to train the development segment identification model and the contamination prediction model in parallel; and generating a model parameter version number that is associated with the ledger version number and the collection session configuration version number. The online inference includes: performing a matching verification between the acquisition session configuration version number and the ledger version number on the high-dimensional feature vector record acquired in real time; assembling a real-time feature vector based on the feature entry mapping record; and inputting it into the development segment identification model and the pollution degree prediction model respectively to obtain the development segment field and the pollution degree prediction value field. The consistency verification process includes: calling the mapping table between the development segment and the pollution threshold interval, comparing the pollution prediction value field with the pollution threshold interval mapped by the development segment field, and generating an alarm escalation strategy instruction based on the consistency verification result.
9. The insulator flashover identification and pollution degree prediction method based on multi-feature fusion according to claim 1, characterized in that, The alarm escalation strategy instructions include: Session ID, Development Segment Field, Pollution Prediction Value Field, Alarm Level Field, Policy Version Number Field, and Action Reference Marker.