An abnormal working condition data self-healing transmission method and system for a power collection terminal
Patent Information
- Application Number
- CN202610994318.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-06
- Publication Date
- 2026-09-04
- Estimated Expiration
- 2046-07-06
AI Technical Summary
[0009]针对电力采集终端在设备损伤与信道劣化并发的极端工况下,现有技术无法将工况感知、数据质量评估与传输策略决策形成闭环联动,导致终端在并发异常工况下无法实现降级但不中断的自愈数据传输的技术问题,本发明提供一种用于电力采集终端的异常工况数据自愈传输方法及系统
[0082] The core benefit of the three-dimensional per-sampling-point confidence scoring mechanism lies in refining the quality assessment of power sampling data from the granularity of the acquisition cycle to the granularity of a single sampling point. Furthermore, the assessment dimensions cover the data content layer (physical rationality), the temporal structure layer (temporal continuity), and the hardware awareness layer (channel health). The abnormal patterns captured by these three dimensions are physically independent and complementary. Specific effects are reflected in the following aspects.
Smart Images

Figure CN122513256B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of artificial intelligence and Internet of Things power acquisition technology, and in particular to a method and system for self-healing transmission of abnormal operating condition data for power acquisition terminals. Background Technology
[0002] Power data acquisition terminals are the infrastructure of the power system's information sensing layer. Widely deployed in outdoor pole-mounted transformer substations, distribution rooms, metering boxes, and other field environments, they are responsible for the real-time acquisition and periodic uploading of key power parameters such as voltage, current, power, and energy consumption. The master station system relies on the continuous time-series data provided by the acquisition terminals to support core services such as power grid operation status awareness, electricity billing and settlement, load forecasting, and distribution network dispatching. Under normal operating conditions, the acquisition terminals complete data sampling and channel transmission according to a fixed acquisition cycle, allowing the master station to stably obtain a complete and continuous time-series data stream, meeting the data quality requirements of the aforementioned services.
[0003] However, the actual deployment environment of power data acquisition terminals is extremely harsh. Extreme weather or electrical events such as lightning strikes, floods, sustained high temperatures, and strong electromagnetic interference can cause physical impact or electrical damage to on-site terminal equipment. Under these extreme conditions, the terminal hardware's sensors, sampling circuits, and main control units often experience varying degrees of performance degradation simultaneously, manifesting as increased noise levels in the sampling channel, decreased sampling integrity, and numerical bias or abrupt changes. Simultaneously, the same extreme physical triggers also lead to the degradation of external communication channels, resulting in decreased signal-to-noise ratio, increased transmission delay jitter, narrowed available bandwidth, and even periodic interruptions. This concurrent operating mode of terminal equipment damage and communication channel degradation driven by the same physical event represents a unique scenario faced by power data acquisition systems in actual operation. In this scenario, the data collected by the terminal itself suffers from decreased reliability, while the channel transmitting data to the main station is simultaneously constrained, creating a dual constraint of compromised data quality and scarce transmission resources.
[0004] Existing technologies have yielded certain research results regarding the handling of communication channel anomalies, including link fault-tolerant methods based on channel switching and data retransmission mechanisms based on breakpoint resumption. However, these methods only focus on the recovery at the communication link level, treating the terminal as a signal source with unchanged data quality. They fail to perceive changes in the operating conditions of the terminal equipment itself due to physical damage, nor do they assess the degree of quality degradation of the collected data under extreme conditions. In scenarios where equipment damage and channel degradation occur simultaneously, simply restoring the channel connection cannot guarantee the reliability of the transmitted content. The data received by the master station may contain a large number of distorted or missing sampling points, thereby affecting subsequent power grid analysis and decision-making.
[0005] In terminal equipment fault diagnosis, existing technologies focus on classifying and identifying fault types, categorizing faults into several predefined categories by analyzing equipment status characteristics. However, the output of these methods is usually discrete fault category labels, lacking the ability to continuously quantify the degree of equipment degradation, and also unable to dynamically assess the reliability of data at each sampling point during the acquisition process. Furthermore, a systematic solution for transforming fault diagnosis results into a linkage mechanism for adjusting transmission strategies has not yet been developed in existing technologies.
[0006] In terms of data quality restoration, existing time-series interpolation and anomaly correction techniques are typically run offline in batch processing on the main site or in the backend system, relying on global analysis of complete historical data. These methods are difficult to run locally on the acquisition terminal in real time, cannot proactively intervene in data quality assessment and restoration after data generation and before transmission, and do not have the ability to combine restoration results with transmission strategy decisions.
[0007] Furthermore, in the fields of IoT edge computing and industrial IoT, existing research has explored the combination of edge-side fault detection and adaptive communication strategies, which can trigger alarms or perform binary communication mode switching after a fault is detected. However, the fault response mechanism of such methods is relatively coarse-grained, only realizing the switching between "normal transmission" and "alarm transmission". It cannot perform fine-grained evaluation of data quality at the sampling point level, nor does it support the joint orchestration capability of dynamically adjusting the data filtering range, compression strategy and transmission priority sorting according to the severity of the fault.
[0008] The existing technologies at each of these levels are independent and fragmented: communication self-healing technology is unaware of data quality, fault diagnosis technology does not output quantitative scores that can be used for transmission decisions, and data repair technology cannot operate in real time at the terminal and coordinate with the transmission strategy. Currently, no existing technology solution forms a closed-loop linkage between terminal equipment condition awareness, sampling point-level data confidence assessment, and transmission strategy decision-making to cope with extreme operating conditions where equipment damage and channel degradation occur simultaneously. Therefore, under the dual constraints of limited local computing power at the terminal and scarce channel resources, how to achieve real-time integrated decision-making for condition awareness, data quality assessment, and transmission orchestration, and ensure that collected data is transmitted to the master station without interruption even under degraded conditions in extreme situations, remains a technical problem to be solved in the field of power data acquisition. Summary of the Invention
[0009] To address the technical problem that existing technologies cannot establish a closed-loop linkage between operating condition perception, data quality assessment, and transmission strategy decision-making in power data acquisition terminals under extreme operating conditions of concurrent equipment damage and channel degradation, resulting in the inability of terminals to achieve degraded but uninterrupted self-healing data transmission under concurrent abnormal operating conditions, this invention provides a method and system for self-healing data transmission under abnormal operating conditions in power data acquisition terminals. This method constructs a three-dimensional joint perception and adaptive transmission decision-making mechanism for operating conditions, data, and channel on the terminal side. It achieves real-time detection and severity measurement of abnormal operating conditions through a lightweight concurrent anomaly identification model, thereby driving end-to-end collaborative processing of confidence scoring per sampling point, online quality pre-repair, and adaptive transmission priority orchestration. This ensures the continuity and reliability of data transmission under the dual constraints of limited terminal computing power and scarce channel resources.
[0010] In one aspect, the present invention provides a method for self-healing data transmission under abnormal operating conditions for power acquisition terminals, comprising the following steps.
[0011] Step S1: The acquisition terminal synchronously acquires the terminal device layer status signal and the communication channel layer status signal in each acquisition cycle, concatenates them into a joint state perception vector, and performs online normalization based on historical statistics using a sliding window. Specifically, the terminal device layer status signal includes the noise level, sampling integrity rate, main control unit temperature, and power supply voltage fluctuation amplitude of each sampling channel; the communication channel layer status signal includes the signal-to-noise ratio, round-trip delay, packet loss rate, and available bandwidth estimate of the current channel; the online normalization standardizes the component using the mean and standard deviation over the most recent preset number of acquisition cycles. The terminal concatenates the observations of the above indicators in the current acquisition cycle in a preset order to form a fixed-dimensional joint state perception vector to eliminate the influence of differences in the dimensions of each indicator on subsequent inference. Preferably, the length of the sliding window is set to the most recent thirty acquisition cycles. This length can cover sufficient historical statistical samples to obtain stable mean and standard deviation estimates, without causing an over-smoothing effect on the gradual change trend of the operating condition due to an excessively long window.
[0012] Step S2 involves inputting the normalized joint state-aware vector into the lightweight concurrent anomaly identification model on the terminal side, outputting a condition type label and a condition severity score. The lightweight concurrent anomaly identification model employs a dual-branch shared embedding structure. After being mapped to a low-dimensional embedding representation through the shared embedding layer, it outputs four condition probability distributions via a classification branch and a severity scalar from zero to one via a regression branch. Steps S3 to S6 are triggered when the condition type label indicates concurrent anomaly of both device and channel. Specifically, the shared embedding layer consists of two fully connected layers with a ReLU activation function in between, mapping the joint state-aware vector from its original dimension to a low-dimensional embedding space. The classification branch consists of a fully connected layer connected to a Softmax activation function, outputting the probability distributions for four conditions: normal operation, device degradation only, channel degradation only, and concurrent anomaly of both device and channel. The category with the highest probability is taken as the condition type label. The regression branch consists of a fully connected layer followed by a Sigmoid activation function, outputting a continuous scalar value ranging from zero to one as the severity score of the operating condition, where zero represents completely normal and one represents extremely severe. Furthermore, the training of this lightweight concurrent anomaly recognition model employs a multi-task joint loss function: cross-entropy loss for the classification branch and mean squared error loss for the regression branch. The two losses are weighted and summed with preset weights to form the total loss function for end-to-end training. The training data comes from a dataset of terminal and channel status annotations accumulated from historical operation records. The annotation method involves experts retrospectively annotating the operating condition type and severity for each collection cycle based on equipment maintenance records and communication logs. Preferably, the hidden dimension of the shared embedding layer is set to thirty-two, the training uses the Adam optimizer, the initial learning rate is set to 8.5 × 10^-4, a cosine annealing learning rate decay strategy is adopted, the batch size is set to sixty-four, the maximum number of training epochs is set to one hundred and twenty, and an early stopping mechanism is used to terminate training when the validation set loss does not decrease for ten consecutive epochs to prevent overfitting. When the operating condition type label is normal operating condition, equipment degradation condition only, or channel degradation condition only, the terminal performs data reporting according to the normal transmission logic.
[0013] Step S3: Calculate a confidence score for each sampling point in the raw data of the current acquisition cycle. The confidence score is obtained by weighting and summing the physical rationality score, temporal continuity score, and channel health score with preset weights. Specifically, the physical rationality score is determined by judging whether the sampling point value falls within the preset physical rationality range of the corresponding power parameter. If it falls within the range, the score is one; if it exceeds the range, the score decreases linearly to zero according to the degree of exceedance. The temporal continuity score is determined by calculating the absolute value of the difference between the sampling point and the sampling point before and after it, and comparing it with a preset normal change threshold. The smaller the absolute value of the difference, the higher the score. The channel health score is the inverse value of the normalized noise level index of the sampling channel in step S1. Further, the preset weights of the physical rationality score, temporal continuity score, and channel health score are 0.4, 0.35, and 0.25, respectively. The above three dimensions comprehensively evaluate the credibility of each sampling point from three independent perspectives: the rationality of the data itself, the consistency of the temporal neighborhood, and the reliability of the acquisition channel. The confidence score ranges from zero to one, and the higher the score, the stronger the credibility of the sampling point.
[0014] Step S4: Perform online pre-repair for sampling points with confidence scores lower than a preset confidence threshold. Specifically, for sampling points with completely missing values, an imputation method based on locally weighted linear regression is used. A preset number of valid sampling points before and after the missing point are used as anchor points, and the confidence scores of each anchor point are used as regression weights to fit a locally linear model and predict the missing value. For sampling points with existing values but confidence scores lower than the preset confidence threshold, the locally weighted linear regression method is used to calculate the estimated value at that location. The original value and the estimated value of the sampling point are weighted and fused using the original confidence score of the sampling point and the repair reliability determined based on the number of anchor points and the mean confidence score of the anchor points, respectively, to obtain a corrected value. The confidence score of the pre-repaired sampling points is recalculated. The repaired confidence score is the smaller of the product of the original confidence score and the repair gain factor and the preset repair upper limit. Furthermore, the repair gain factor is determined based on the number of anchor points involved in the repair and the average confidence level of the anchor points. A larger number of anchor points and a higher average confidence level result in a larger repair gain factor, indicating a higher reliability of the repair result. The preset repair upper limit is used to constrain the confidence score after repair from exceeding this upper limit, preventing the repair operation from excessively boosting the confidence score. Preferably, the preset confidence threshold is 0.5, and the number of anchor points in the locally weighted linear regression is three valid sampling points before and after the regression.
[0015] Step S5: Generate a transmission orchestration scheme based on the operational severity score and channel layer status signal, including the data filtering range, data compression granularity, and transmission priority queue. Specifically, in the data filtering range dimension, a dynamic confidence transmission threshold is set based on the operational severity score. The dynamic confidence transmission threshold is determined by the sum of a base threshold value and the product of the operational severity score and a threshold adjustment coefficient. Only sampling points with confidence scores higher than the dynamic confidence transmission threshold are included in the dataset to be transmitted. Preferably, the base threshold value is 0.3, and the threshold adjustment coefficient is 0.4. That is, when the operational severity score is 0.5, the dynamic confidence transmission threshold is 0.3 plus 0.5 multiplied by 0.4 equals 0.5. In this case, only sampling points with confidence scores higher than 0.5 are included in the transmission. Furthermore, in terms of data compression granularity, the compression level is determined based on the ratio of the estimated available bandwidth of the current channel to the amount of data to be transmitted. When the ratio is less than one, a lossy compression strategy based on confidence-weighted sampling point aggregation is used. This lossy compression combines multiple adjacent sampling points into a single representative value using a confidence-weighted average. The width of the aggregation window increases with the compression level, and high-confidence sampling points receive greater weight during aggregation to retain more reliable information. Lossless compression is used when available bandwidth is sufficient. In terms of transmission priority queues, the transmission order is determined by the service criticality of power parameter types as the first ranking criterion and the confidence score as the second ranking criterion. The priority of power parameter types is ordered according to a preset service criticality, with voltage and current parameters having higher priority than power parameters, and power parameters having higher priority than electricity parameters. Within the same parameter type, data segments with higher confidence scores are transmitted first. The terminal sends data packets sequentially according to the priority queue. When the channel condition deteriorates further and transmission is interrupted, the transmitted data packets are the subset of data with the highest confidence and the highest business criticality under the current operating conditions.
[0016] Step S6: Generate a condition summary metadata data packet and transmit it to the master station as the highest priority data packet. Specifically, the condition summary metadata data packet includes the timestamp of the current collection period, the condition type label and condition severity score, the total number of sampling points in the original data, the distribution of the number of sampling points in each confidence interval, the number of sampling points to be repaired and the average change in confidence before and after repair, the dynamic confidence transmission threshold and compression level, and the proportion of the number of sampling points actually included in the transmission to the total number of sampling points. After receiving the condition summary metadata data packet, the master station can use it to determine the integrity and reliability status of the current batch of data, providing a basis for subsequent data fusion processing and post-processing of missing data. Further, when the concurrent abnormal condition continues to exceed a preset duration threshold, the cumulative abnormal duration and the trend information of the change in the condition severity score of each collection period are appended to the condition summary metadata data packet for the master station to assess the damage evolution of the terminal equipment. Preferably, the duration threshold is set to five consecutive collection periods.
[0017] In another aspect, the present invention provides a self-healing data transmission system for abnormal operating conditions of a power acquisition terminal, comprising: a joint state perception module, used to synchronously acquire terminal device layer state signals and communication channel layer state signals in each acquisition cycle, concatenate the two into a joint state perception vector and perform online normalization based on sliding window historical statistics; and a concurrent anomaly identification module, used to input the normalized joint state perception vector into a lightweight concurrent anomaly identification model on the terminal side, and output an operating condition type label and an operating condition severity score; the lightweight concurrent anomaly identification model adopts a dual-branch shared embedding structure, which is mapped to a low-dimensional embedding representation through a shared embedding layer, and then outputs four types of operating condition probability distributions through a classification branch and a severity scalar from zero to one through a regression branch; when the operating condition type label is a device and channel concurrent anomaly, the system can be configured to perform the corresponding function. The system triggers the sequential operation of the sampling point confidence scoring module, online pre-repair module, adaptive transmission orchestration module, and operating condition summary generation module. The sampling point confidence scoring module calculates a confidence score for each sampling point in the raw data of the current acquisition cycle. This confidence score is obtained by weighting and summing the physical rationality score, temporal continuity score, and channel health score using preset weights. The online pre-repair module performs online pre-repair on sampling points with confidence scores below a preset confidence threshold. The adaptive transmission orchestration module generates a transmission orchestration scheme based on the operating condition severity score and channel layer status signal, including the data filtering range, data compression granularity, and transmission priority queue. The operating condition summary generation module generates an operating condition summary metadata data packet and transmits it to the master station as the highest priority data packet. All these modules are deployed on the power acquisition terminal side and operate collaboratively locally to achieve self-healing data transmission under abnormal operating conditions.
[0018] This invention constructs a three-dimensional joint sensing mechanism of operating condition, data, and channel on the terminal side, fusing terminal device layer status signals and communication channel layer status signals into a unified joint state sensing vector. Based on a lightweight dual-branch shared embedding model, it simultaneously completes operating condition type identification and severity quantification, achieving real-time sensing and graded response to concurrent abnormal operating conditions. This invention calculates a confidence score for each sampling point from three dimensions: physical rationality, temporal continuity, and channel health. Based on the confidence score, it drives online pre-repair using locally weighted linear regression, achieving point-by-point quantitative evaluation and proactive repair of data quality collected under abnormal operating conditions on the terminal side, improving the overall credibility of the data to be transmitted. This invention achieves adaptive matching of transmission strategy to operating condition status and channel resources by dynamically adjusting the confidence transmission threshold based on the operating condition severity score, adaptively selecting compression granularity based on available channel bandwidth, and jointly determining the transmission priority queue based on service criticality and confidence score. Under conditions of limited channel resources, it prioritizes the transmission of the data subset with the highest confidence and greatest service value. This invention generates a working condition summary metadata data package containing working condition status and data quality statistics, and transmits it to the master station as the highest priority data package. This provides the master station with a transparent perception of the integrity and reliability of the current batch of data, facilitating subsequent data fusion processing and post-processing of missing data. Attached Figure Description
[0019] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0020] Figure 1 This is an overall flowchart of the abnormal operating condition data self-healing transmission method for power acquisition terminals provided in this embodiment of the invention.
[0021] Figure 2 This is a schematic diagram of the dual-branch shared embedding structure of the lightweight concurrent anomaly identification model provided in this embodiment of the invention.
[0022] Figure 3 This is a data flow diagram of confidence scoring and online pre-repair provided in an embodiment of the present invention.
[0023] Figure 4 This is a structural block diagram of an abnormal operating condition data self-healing transmission system for power acquisition terminals provided in an embodiment of the present invention.
[0024] Figure 5 This is a schematic diagram of the three-dimensional joint sensing space and decision boundary of the working condition-data-channel of the present invention.
[0025] Figure 6 This is a schematic diagram of the confidence-driven dynamic threshold filtering and priority transmission orchestration of the present invention.
[0026] Figure 7This is a comparison chart of the effective sampling rate and data reliability of different schemes in this invention.
[0027] Figure 8 This is a graph showing the impact of the severity of the operating conditions on the effective sampling rate and transmission threshold of this invention. Detailed Implementation
[0028] To make the objectives and technical solutions of this invention clearer, the embodiments of this invention will be further described in detail below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of this invention and should not be construed as limiting the scope of protection of this invention.
[0029] Example 1
[0030] This embodiment provides a self-healing data transmission method for abnormal operating conditions in power acquisition terminals, such as... Figure 1 As shown, this method is deployed on the power acquisition terminal side. Under the dual constraints of limited terminal computing power and scarce channel resources, it achieves the continuity and reliability of data transmission under abnormal operating conditions through a three-dimensional joint perception and adaptive transmission decision mechanism of operating conditions, data, and channel. The method includes the following steps.
[0031] Step S1: Construct the terminal operating condition-channel joint state awareness vector.
[0032] During each data acquisition cycle, the acquisition terminal synchronously acquires the terminal device layer status signal and the communication channel layer status signal, and merges them into a unified joint status perception vector. The acquisition of the terminal device layer status signal covers several key dimensions of the terminal hardware: the noise level of each sampling channel reflects the signal quality status of the sensor and sampling circuit; the sampling completeness rate reflects the proportion of successful sampling completed by each channel in the current acquisition cycle; the main control unit temperature reflects the internal thermal state of the terminal; and the power supply voltage fluctuation amplitude reflects the stability of the terminal power supply system. In actual deployment, the noise level can be obtained by performing a fast short-time Fourier transform on the background noise of each channel during the sampling interval and taking the average spectral energy; the sampling completeness rate is the ratio of the actual number of sampling points completed by the channel in the current cycle to the planned number of sampling points; the main control unit temperature is directly read by the temperature sensor on the terminal motherboard; and the power supply voltage fluctuation amplitude is the ratio of the peak-to-peak value of the power supply voltage sampling sequence in the current cycle to the rated voltage.
[0033] The acquisition of communication channel layer status signals covers key transmission quality indicators of currently active communication channels. Signal-to-noise ratio (SNR) is obtained by the ratio of signal power to noise power in the terminal's modulation and demodulation module. Round-trip time (RTD) is obtained by sending probe messages from the terminal to the master station and timing the process. Packet loss rate is calculated based on the proportion of data packets sent within the most recent statistical window that have not received acknowledgments. Available bandwidth is estimated using an online bandwidth probing algorithm on the terminal side based on the relationship between data packet transmission rate and acknowledgment rate. In typical power data acquisition terminal deployment scenarios, communication channels may be of various types, such as GPRS, 4G, Ethernet, or power line carrier. The aforementioned indicators differ in magnitude across different channel types, thus requiring subsequent normalization processing to eliminate these differences.
[0034] The terminal concatenates the observations of the above indicators in the current acquisition cycle in a preset order to form a fixed-dimensional joint state perception vector. Assuming the terminal has Nc sampling channels, the device-level state signal contains Nc noise level values, Nc sampling integrity rate values, one main control unit temperature value, and one power supply voltage fluctuation amplitude value, for a total of 2Nc+2 components; the channel-level state signal contains four components: signal-to-noise ratio, round-trip delay, packet loss rate, and available bandwidth estimate. The total dimension of the joint state perception vector is 2Nc+6. In actual power acquisition terminals, the number of sampling channels Nc is typically four to eight, corresponding to parameter channels such as voltage, current, active power, and reactive power. Therefore, the dimension of the joint state perception vector is typically fourteen to twenty-two dimensions.
[0035] To eliminate the impact of differences in the dimensions of various indicators on subsequent inference, the terminal performs online normalization processing based on historical statistics using a sliding window for each component in the joint state perception vector. For the i-th component xi in the joint state perception vector, the terminal maintains the observation sequence of this component over the most recent W acquisition cycles, calculates the mean mui and standard deviation sigmai of this sequence, and the normalized value zi = (xi - mui) / (sigmai + epsilon), where epsilon is a very small positive number (1.0 × 10^-8) to avoid division by zero error when the standard deviation is zero. The length W of the sliding window is set to the most recent thirty acquisition cycles, corresponding to 7.5 hours of historical data at a typical acquisition frequency of once every fifteen minutes. This length can cover sufficient historical statistical samples to obtain stable mean and standard deviation estimates without causing excessive smoothing of the gradual trend of the operating condition due to an excessively long window. The terminal only needs to maintain the sliding window cumulative sum and cumulative sum of squares for each component, without storing all historical original values, to achieve incremental updates of mean and standard deviation, thus meeting the deployment requirements of terminal with limited storage.
[0036] Step S2: Perform concurrent abnormal condition identification and severity measurement based on the joint state-aware vector.
[0037] The normalized joint state-aware vector obtained in step S1 is input into a lightweight concurrent anomaly detection model deployed on the terminal side. This model simultaneously outputs the condition type label and condition severity score for the current data collection period. Figure 2 As shown, this lightweight concurrent anomaly recognition model adopts a dual-branch shared embedding structure. Its design idea is to reduce the number of model parameters by sharing the underlying feature extraction, while maintaining the output accuracy of the classification and regression tasks through independent task branches.
[0038] The model's front end is a shared embedding layer, consisting of two cascaded fully connected networks. The first fully connected network maps the normalized joint state-aware vector from its original dimension (2Nc+6 dimensional) to a 64-dimensional intermediate representation. After introducing a nonlinear transformation using the ReLU activation function, the second fully connected network further compresses the 64-dimensional intermediate representation into a 32-dimensional low-dimensional embedding space. The choice of 32 dimensionality as the embedding dimension is based on a balance between terminal-side computing power constraints and feature representation capabilities. Under this dimension, the total number of model parameters can be controlled within the tens of thousands, meeting the inference deployment requirements of terminal-side microcontrollers or low-power processors.
[0039] The 32-dimensional low-dimensional embedding representation of the shared embedding layer output is then input into two parallel branches. The first branch is the condition type classification branch, which maps the 32-dimensional embedding to a four-dimensional output using a fully connected network. The output is then converted into a probability distribution for four conditions: normal condition, equipment degradation condition, channel degradation condition, and concurrent equipment and channel abnormal condition using a Softmax activation function. The terminal selects the category with the highest probability as the condition type label for the current acquisition period. The second branch is the severity regression branch, which maps the 32-dimensional embedding to a one-dimensional output using a fully connected network. The output is compressed to a continuous scalar from zero to one using a Sigmoid activation function, serving as the condition severity score, where zero represents completely normal and one represents extremely severe. The two branches share the parameters of the embedding layer but each has independent output layer parameters. During the inference phase, the model can simultaneously complete condition type identification and severity measurement with a single forward propagation, avoiding the additional computational overhead of separate inference by two independent models.
[0040] When the operating condition type label is "Dual abnormal operating condition of equipment and channel," it indicates that the terminal is currently facing an extreme operating condition with concurrent equipment damage and channel degradation, triggering the self-healing transmission process in subsequent steps S3 to S6. When the operating condition type label is "Normal operating condition," "Equipment-only deterioration operating condition," or "Channel-only deterioration operating condition," the terminal performs data reporting according to the conventional transmission logic. Among these, the "Equipment-only deterioration operating condition" and "Channel-only deterioration operating condition" can be handled by the existing single-dimensional fault handling mechanisms, respectively.
[0041] like Figure 5As shown, this invention constructs a three-dimensional joint sensing space with device layer health, channel layer quality, and data quality index as coordinate axes. In this space, four operating conditions exhibit distinct clustering characteristics: scatter points for normal operating conditions are concentrated in the high-value region of the three-dimensional space, representing operating conditions where the device is in good condition, the channel quality is stable, and the data reliability is high; scatter points for conditions with only device degradation shift along the low-value region of the device layer health axis, while the channel layer quality remains at a high level; scatter points for conditions with only channel degradation shift along the low-value region of the channel layer quality axis; and scatter points for concurrent abnormal operating conditions cluster in the low-value corners of the three-dimensional space, where both device layer health and channel layer quality are at a low level.
[0042] The lightweight concurrent anomaly detection model learns a decision boundary surface in this three-dimensional space to automatically classify four types of operating conditions. The decision boundary, based on a threshold of 0.5, divides the perception space into four operating condition regions. The scatter plot size encodes the severity score of the operating condition; the higher the severity, the larger the scatter plot, intuitively reflecting the severity of the abnormal operating condition. When the detection result falls into a concurrent anomaly region, a subsequent self-healing transmission process is triggered.
[0043] This lightweight concurrent anomaly detection model needs to be trained offline before deployment to the terminal. The training data comes from a dataset of terminal and channel status annotations accumulated from historical operation records. The annotation method involves domain experts retrospectively labeling the operating condition types and severity for each collection cycle based on equipment maintenance records and communication logs: the operating condition type label is determined by experts based on a comprehensive assessment of the fault categories in the maintenance records and the abnormal events in the channel logs, categorizing them into one of four types; the severity label is assigned a continuous value between zero and one by experts based on a comprehensive assessment of the degree of equipment damage and channel degradation. The training dataset consists of tens of thousands to hundreds of thousands of labeled samples, divided into training, validation, and test sets in a 7:1.5:1.5 ratio. During the data preprocessing stage, the training data undergoes the same normalization process as online inference, and the SMOTE oversampling method is used to augment the data for concurrent anomaly operating condition categories with smaller sample sizes to alleviate class imbalance.
[0044] The model training employs a multi-task joint loss function. The classification branch uses cross-entropy loss (Lce), and the regression branch uses mean squared error loss (Lmse). The total loss function is L = alpha × Lce + beta × Lmse, where alpha and beta are the weighting coefficients for the classification and regression losses, respectively. Alpha is set to 0.6 and beta to 0.4; this weighting ratio ensures that the classification task receives a slightly higher optimization priority during joint training, guaranteeing the accuracy of the work condition type labels. Training uses the Adam optimizer with an initial learning rate of 8.5 × 10⁻⁴. A cosine annealing learning rate decay strategy is employed, smoothly decreasing the learning rate from the initial value to a minimum of 1.0 × 10⁻⁶ according to a cosine function during training. The batch size is set to 64, and the maximum number of training epochs is set to 120. An early stopping mechanism is used to terminate training if the total loss of the validation set does not decrease for ten consecutive epochs to prevent overfitting. Training can be completed on a workstation equipped with a single GPU, with typical training time within tens of minutes. After training, the model's evaluation metrics on the test set include: the macro-average F1 score and recall rate for concurrent outlier categories in the classification branch, and the root mean square error and mean absolute error in the regression branch. The model is then quantized and compressed before being deployed to the inference engine on the terminal side.
[0045] Step S3: Perform a confidence score for each sampling point on the raw data of the current acquisition period.
[0046] After step S2 identifies a concurrent abnormal condition, the terminal calculates a confidence score for each sampling point in the original data sequence obtained within the current acquisition cycle. For example... Figure 3 As shown, this confidence scoring method comprehensively evaluates the credibility of each sampling point from three independent dimensions.
[0047] The first dimension is the physical rationality score. The terminal pre-stores the physically reasonable ranges corresponding to each power parameter type, and the range is determined according to the operation specifications of the power system and the measuring range of the terminal. For voltage parameters, the physically reasonable range is usually set to positive and negative 30 percent of the rated voltage; for current parameters, the physically reasonable range is set from zero to three times the rated current; for power parameters, the upper limit of the physically reasonable range is determined according to the product of the rated voltage and the rated current. When the value of a sampling point falls within the physically reasonable range of the corresponding parameter, the score of this dimension is 1. When the value of the sampling point is out of the range, the score linearly decays to zero according to the degree of exceeding. The specific calculation method is as follows: let the upper and lower bounds of the physically reasonable range be Vmax and Vmin respectively, and the sampling value be x. When x > Vmax, the physical rationality score is max(0, 1-(x-Vmax) / (Vmax-Vmin); when x < Vmin, the physical rationality score is max(0, 1-(Vmin-x) / (Vmax-Vmin)). This calculation method enables sampling points that exceed the range farther obtain lower physical rationality scores, and the score decays to zero when the degree of exceeding reaches one time the width of the range.
[0048] The second dimension is the time-series continuity score. For the k-th sampling point in the original collected data sequence, the terminal calculates the absolute difference delta_k_prev=|x_k-x_{k-1}| between it and the previous sampling point, and the absolute difference delta_k_next=|x_k-x_{k+1}| between it and the next sampling point, and takes the larger value of the two delta_k_max=max(delta_k_prev,delta_k_next) as the time-series change measurement of this sampling point. The terminal presets a normal change threshold delta_th for each power parameter type. This threshold is determined according to the statistical quantile of the absolute difference of adjacent sampling points in historical normal operation data, and the 95th percentile is usually taken. The time-series continuity score is calculated as follows: when delta_k_max is less than or equal to delta_th, the score is 1; when delta_k_max is greater than delta_th, the score is max(0, 1-(delta_k_max-delta_th) / delta_th). For sampling points at the head and end of the sequence, only one-sided difference is calculated. This design enables the time-series continuity score of sampling points whose mutation amplitude exceeds twice the normal change threshold to decay to zero.
[0049] The third dimension is the channel health score. This dimension directly utilizes the noise level index of the sampled channel obtained in step S1. The terminal normalizes the current noise level value through a sliding window and takes the inverse value as the channel health score, i.e., the channel health score is 1 / (1+noise_norm), where noise_norm is the normalized noise level value. When the channel noise level is within the historical normal range, the normalized noise level value is close to zero, and the channel health score is close to one; when the channel noise level rises abnormally, the normalized value increases, and the channel health score decreases accordingly.
[0050] The terminal sums the scores from the three dimensions using preset weights to obtain the confidence score for each sampling point: C = 0.40 × S_phy + 0.35 × S_temp + 0.25 × S_ch, where S_phy is the physical plausibility score, S_temp is the temporal continuity score, and S_ch is the channel health score. The physical plausibility score has the highest weight, reflecting the design principle that the physical plausibility of the data itself is the primary basis for confidence assessment under abnormal operating conditions. The temporal continuity score has the second highest weight, while the channel health score has the lowest weight but is still not negligible. This is because even if the value of a sampling point appears reasonable and is continuous with its neighbors, if the noise level of its channel has deteriorated significantly, the sampling point may still carry a large measurement error.
[0051] Step S4: Perform online pre-repair on low-confidence sampling points based on confidence scores.
[0052] For sampling points with confidence scores lower than the preset confidence threshold in step S3, the terminal performs an online pre-repair operation locally to improve the overall credibility of the data to be transmitted. The preset confidence threshold is set to 0.5, meaning that sampling points with confidence scores lower than 0.5 are identified as low-confidence sampling points and require pre-repair processing.
[0053] For sampling points with completely missing values, the terminal employs an imputation method based on locally weighted linear regression for repair. Specifically, three valid sampling points are searched forward and backward from the location of the missing point as anchor points. Valid sampling points are those with existing sampling values and a confidence score of not less than 0.5. If there are fewer than three valid anchor points within a preset search range in a certain direction, all available valid anchor points are used. A weighted least squares linear regression model y = a × t + b is constructed using the confidence score of each anchor point as regression weights, where t is the time index of the sampling point, y is the sampling value, and the model parameters a and b are solved by minimizing the weighted residual sum of squares sum(C_i × (y_i - a × t_i - b)^2). The predicted value of the missing point is determined by the output of the linear model at the time index of the missing point. This method utilizes local temporal trends and anchor point confidence information for imputation; high-confidence anchor points gain greater influence in the regression, making the imputation results more dependent on reliable neighborhood data.
[0054] For sampling points where a value exists but the confidence score is below 0.5, the terminal employs a confidence-weighted fusion correction method. First, the estimated value y_est for that location is calculated using the locally weighted linear regression method described above. Then, the original value y_raw and the estimated value y_est are weighted and fused: the corrected value y_fix = C_raw × y_raw + (1 - C_raw) × y_est, where C_raw is the original confidence score for that sampling point. The intention behind this fusion strategy is that sampling points with higher confidence retain more original information, while sampling points with extremely low confidence are more often replaced by neighborhood estimates. When the original confidence is close to 0.5, the corrected value is close to the equal-weighted average of the original and estimated values; when the original confidence is close to 0, the corrected value is almost entirely determined by the estimated value.
[0055] For the pre-repaired sampling points, the terminal recalculates their confidence scores. The repaired confidence score C_fix is determined as follows: C_fix = min(C_raw × G_repair, C_cap), where G_repair is the repair gain factor and C_cap is the preset repair upper limit. The repair gain factor G_repair is determined based on the number of anchors involved in the repair, n_anchor, and the mean confidence score of the anchors, C_anchor_mean: G_repair = 1 + 0.5 × (n_anchor / 6) × C_anchor_mean. The repair gain factor reaches its maximum value of 1.5 when all six anchors are available and the mean confidence score of the anchors is one. The preset repair upper limit C_cap is set to 0.75 to constrain the repaired confidence score from not exceeding this upper limit, thus avoiding an overly optimistic increase in confidence during the repair operation. This constraint reflects a conservative principle: the confidence level of online repaired sampling points, no matter how thorough the repair process, should not reach the level of unrepaired normal sampling points, because the repair introduces additional estimation errors.
[0056] Step S5: Generate an adaptive transmission orchestration scheme based on the severity of the operating conditions and the channel status.
[0057] Based on the operational severity score obtained in step S2 and the channel layer status signal extracted in step S1, the terminal dynamically generates a transmission orchestration scheme for the current data acquisition period. This transmission orchestration scheme covers three decision dimensions: the scope of data filtering, the data compression granularity, and the transmission priority queue. These three dimensions work together to achieve adaptive matching of the transmission strategy to the operational status and channel resources.
[0058] In terms of the data transmission filtering scope, the terminal sets a dynamic confidence transmission threshold Th_dyn based on the severity score of the operating condition. The threshold is calculated as Th_dyn = Th_base + S_severity × K_adj, where Th_base is the base threshold value, set to 0.3; S_severity is the severity score output in step S2; and K_adj is the threshold adjustment coefficient, set to 0.4. This design causes the transmission threshold to increase linearly with the severity of the operating condition: when the severity score is zero, the transmission threshold equals the base threshold value of 0.3, and almost all sampling points can be included in the transmission; when the severity score is one, the transmission threshold increases to 0.7, and only the subset of data with the highest confidence is retained for transmission. Taking a typical scenario as an example, when the severity score is 0.5, the dynamic confidence transmission threshold is 0.3 plus 0.5 multiplied by 0.4, which equals 0.5. At this point, sampling points with a confidence score lower than 0.5 are excluded from the transmission range, and the terminal only transmits data with medium to high confidence. This strategy ensures that the terminal rigorously filters data as the severity of the situation increases, avoiding the occupation of already strained channel resources by a large amount of low-confidence data.
[0059] At the data compression granularity level, the terminal determines the compression level based on the ratio of the estimated available bandwidth of the current channel, BW_avail, to the amount of data to be transmitted, D_trans, R_bw = BW_avail / D_trans. When R_bw is greater than or equal to 1, the available bandwidth is sufficient, and the terminal uses lossless compression to transmit all the data to be transmitted. When R_bw is less than 1, the terminal switches to lossy compression mode, and the compression level L_comp is determined according to R_bw: R_bw between 0.5 and 1 is level 1 compression, R_bw between 0.25 and 0.5 is level 2 compression, and R_bw less than 0.25 is level 3 compression. Lossy compression uses a confidence-weighted sampling point aggregation strategy. Level 1 compression aggregates two adjacent sampling points into a representative value, level 2 compression aggregates four adjacent sampling points, and level 3 compression aggregates eight adjacent sampling points. The aggregated representative value is calculated as a confidence-weighted average, meaning the representative value equals the weighted average of the sampled values and confidence scores of each sample point within the aggregation window: sum(C_i×x_i) / sum(C_i). High-confidence sample points receive greater weight during aggregation, allowing the representative value to reflect more information from reliable data and reducing the interference of low-confidence sample points on the aggregation result. The confidence score attached to the aggregated representative value is the weighted average of the confidence scores of each sample point within the aggregation window.
[0060] At the priority queue level, the terminal organizes the filtered and compressed data to be transmitted into data packets. Each data packet contains a set of sampling points of the same power parameter type within a continuous time segment. The terminal determines the transmission order of each data packet according to a joint ranking of service criticality and confidence score. The ranking rule for service criticality is: voltage and current parameters have the highest service criticality, followed by power parameters, and then electrical quantity parameters have the lowest. Within the same service criticality level, data packets are ranked from highest to lowest according to the average confidence score of the sampling points contained in the data packet. The terminal sends data packets sequentially according to this priority queue order. If the channel condition deteriorates further during transmission, causing a transmission interruption, the successfully transmitted data packets represent the subset of data with the highest confidence and service criticality under the current operating conditions, maximizing the reachability of critical service data. When the channel is restored, the terminal continues transmitting the remaining data packets from the point of interruption in the priority queue.
[0061] like Figure 6 As shown, the transmission orchestration mechanism of the present invention includes two core components: confidence filtering and priority arrangement. Figure 6 The left half displays a time-series heatmap of the confidence scores for each channel's sampling points. The horizontal axis represents the sampling time index, and the vertical axis represents the four parameter channels: voltage, phase A current, phase B current, and phase C current. The color intensity indicates the confidence score for each sampling point. In this example, the phase A current and phase B current channels exhibit large areas of low confidence, marked as sampling points below the dynamic transmission threshold Th_dyn.
[0062] Figure 6 The right half shows the orchestration results of the transmission priority queue. The queue is sorted according to the service criticality level, with voltage data packets having the highest priority, followed by current data packets, power calculation packets, and power statistics packets. Each data packet is labeled with an average confidence score. Although the confidence scores of the pre-repaired A-phase and B-phase current data packets are relatively low, they still meet the transmission threshold requirements. The bottom shows the key parameters of this transmission orchestration: operating condition severity is 0.63, dynamic confidence transmission threshold Th_dyn is 0.552, corresponding to a compression level of level two.
[0063] Step S6: Generate a working condition summary metadata data package and transmit it synchronously with the data.
[0064] While performing data transmission in step S5, the terminal generates a condition summary metadata packet and transmits it to the master station as the highest priority packet. The data size of this condition summary metadata packet is much smaller than the collected data itself, and it still has a high probability of successfully reaching the master station even under extremely degraded channel conditions, providing the master station with transparent information about the current terminal status and data quality.
[0065] The operating condition summary metadata data package includes the following fields: the timestamp of the current collection period, used for time alignment on the master station side; the operating condition type label and operating condition severity score output in step S2, allowing the master station to understand the type and severity of abnormal operating conditions currently faced by the terminal; the total number of sampling points of the original data in the current collection period, serving as a benchmark reference for data integrity; the distribution of the number of sampling points in each confidence interval after scoring in step S3, typically statistically divided into four intervals: 0 to 0.25, 0.25 to 0.5, 0.5 to 0.75, and 0.75 to 1, enabling the master station to intuitively understand the quality distribution of the current batch of data; the number of sampling points pre-repaired in step S4 and the average change in confidence before and after repair, reflecting the processing scope and effect of the pre-repair operation on the terminal side; the dynamic confidence transmission threshold and compression level determined in step S5, allowing the master station to understand the screening and compression strategies for the current batch of data; and the proportion of the number of sampling points actually included in the transmission to the total number of sampling points, directly reflecting the data transmission integrity rate.
[0066] After receiving the operational condition summary metadata data packet, the master station can determine the integrity and reliability status of the current batch of data. When the transmission integrity rate is low or the confidence distribution is biased towards the low range, the master station can initiate a post-event refined completion process, combining associated data from adjacent terminals or historical data from the same period to more accurately reconstruct missing and low-confidence data. When the operational condition severity score remains high, the master station can trigger remote maintenance commands or dispatch on-site repair tasks.
[0067] When a terminal detects concurrent abnormal operating conditions that persist for an extended period exceeding a preset duration threshold, it indicates that the terminal may be experiencing continuous environmental stress or progressive equipment damage. The cumulative duration of the abnormality and the trend information of the severity score for each acquisition cycle are appended to the operating condition summary metadata. The trend information specifically includes the severity score sequence and its linear regression slope for the most recent consecutive abnormal cycles. When this slope is positive and exceeds a preset alarm threshold, it indicates that the operating condition is continuously deteriorating. The master station can use this information to assess the damage evolution of the terminal equipment and plan emergency response measures in advance. The duration threshold is set to five consecutive acquisition cycles, corresponding to approximately one hour and fifteen minutes of continuous abnormal duration at a typical acquisition frequency. This duration is sufficient to eliminate false alarms caused by transient disturbances.
[0068] Example 2
[0069] In a preferred embodiment of the present invention, the three-dimensional sampling point confidence scoring mechanism and its synergistic effect in the entire process of online pre-repair and transmission orchestration are described in detail. This embodiment takes a power acquisition terminal deployed in an outdoor distribution substation as the operating scenario. The terminal simultaneously acquires eight power parameters, including three-phase voltage, three-phase current, active power, and reactive power, with an acquisition cycle of 1 second. Each parameter generates sixty sampling points per cycle.
[0070] After an abnormal concurrent operation is identified, the terminal calculates a confidence score for each sampling point within the current acquisition cycle. The confidence score is obtained by weighting three components—physical rationality score, temporal continuity score, and channel health score—with preset weights. The weighting formula is as follows: C(i) = w1 P(i) + w2 T(i) + w3 H(i) Wherein, C(i) is the confidence score of the i-th sampling point, P(i) is the physical rationality score, T(i) is the temporal continuity score, H(i) is the channel health score, and w1, w2, and w3 are the corresponding preset weights, the sum of the three weights equals 1. Preferably, w1 is 0.4, w2 is 0.35, and w3 is 0.25. The basis for this weight configuration is as follows: physical rationality reflects whether the sampling point meets the basic physical constraints of the power parameters, and is the most direct basis for judging the validity of the data, so it has the highest weight; temporal continuity reflects the consistency of the sampling point in its temporal neighborhood, and for the smoothly changing power frequency signal in the actual power system, this dimension can effectively identify abrupt sampling anomalies; channel health introduces the equipment layer status information into the data layer evaluation, reflecting the linkage between operating condition perception and data quality evaluation, but since channel health itself is based on the indirect estimation of the channel noise level, it is not as direct as the first two dimensions, so its weight is relatively the smallest.
[0071] The physical reasonableness score P(i) is calculated based on a pre-configured physical reasonable range for power parameters. For three-phase voltage parameters, the physical reasonable range is configured to be 0.8 to 1.2 times the rated voltage, i.e., in a scenario with a rated voltage of 220 volts, the reasonable range is 176 volts to 264 volts. When the sampled value falls within the reasonable range, P(i) is set to 1; when the sampled value exceeds the reasonable range, P(i) decays linearly to zero according to the proportion of the excess to the width of the reasonable range, using the following decay formula: P(i) = max(0, 1 - |x(i) - x_mid| / (x_range / 2 + delta)) Where x(i) is the sampled value, x_mid is the median of the reasonable range, x_range is the width of the reasonable range, and delta is the tolerance coefficient, preferably 0.1 times the width of the reasonable range, used to introduce a smooth transition near the boundary of the reasonable range to avoid abrupt changes in the score. For sampled points with completely missing sampled values, P(i) is directly set to zero.
[0072] The temporal continuity score T(i) is calculated based on the difference relationship between a sampling point and its neighboring sampling points in its temporal neighborhood. Specifically, the mean of the absolute values of the differences between the i-th sampling point and its preceding and following sampling points is calculated, denoted as D(i): D(i) = (|x(i) - x(i-1)| + |x(i+1) - x(i)|) / 2 Furthermore, the 90ths of the absolute difference between all adjacent sampling points of the power parameter in the most recent complete acquisition cycle is used as the normal variation threshold theta. The formula for calculating T(i) is as follows: T(i) = max(0, 1 - D(i) / (k theta)) Where k is the adjustment coefficient, preferably 2.0, indicating that T(i) decays to zero when D(i) reaches twice the normal variation threshold. The reason for using the decimal place instead of the mean as the normal variation threshold is that a certain proportion of large fluctuation sampling points exist in the normally collected data. Using the decimal place can preserve this normal fluctuation space and avoid misjudging normal power system transient processes. For the first and last sampling points of the sequence, when there is no predecessor or successor, D(i) is calculated only by the difference on one side.
[0073] The channel health score H(i) is directly taken as the inverse value of the noise level index of the channel to which the sampling point belongs after online normalization in step S1. Let the normalized noise level be n, then H(i) = 1 / (1+n). This score takes the same value for all sampling points of the same channel within the same acquisition period, reflecting the overall hardware health status of the acquisition channel. When the noise level index of a certain channel is at a historical high, the H(i) of all sampling points in that channel is low, reflecting the systematic impact of channel degradation on data reliability.
[0074] After calculating the three components separately, the comprehensive confidence score C(i) for each sampling point is obtained according to the weighting formula mentioned above, with a value ranging from 0 to 1. For a typical concurrent anomaly acquisition cycle, the eight parameters generate a total of 480 sampling points. Among them, sampling points with a physical rationality score below 0.5 are usually concentrated in the most severely disturbed channels, while sampling points with a temporal continuity score below 0.5 mainly appear at the transition positions of the sampling sequence. The superposition of the two makes the distribution of the comprehensive confidence score show obvious temporal locality—anomalies tend to occur in a concentrated manner within a specific time period and a specific channel, rather than being evenly distributed throughout the entire cycle. This characteristic is a prerequisite for the subsequent online pre-repair step to effectively utilize local neighborhood information for interpolation.
[0075] After the confidence score is completed, online pre-repair is performed on sampling points whose scores are lower than the preset confidence threshold of 0.5. The pre-repair is handled in two separate cases.
[0076] The first scenario involves sampling points with completely missing values. For such sampling points, we use the three valid sampling points before and after the missing point in time (i.e., sampling points with a score of at least 0.5 and existing values; if there are fewer than three valid points, we use the actual number available) as anchor points, and use the confidence scores of each anchor point as regression weights to construct a locally weighted linear regression model. Let the set of anchor points be {(t_j, x_j, C_j)}, where t_j is the time index, x_j is the sample value, and C_j is the confidence score. Then, the parameter estimation objective of the weighted linear regression is to minimize the weighted sum of squared residuals: min_{a, b} sum_j [C_j (x_j - a t_j - b)^2] After obtaining parameters a and b, the time indices of the missing points are substituted into the linear model to obtain the predicted imputation values. The technical rationale of this method lies in the fact that within a short time window (the time span corresponding to each of the three sampling points usually does not exceed 0.1 seconds), most power parameters approximately satisfy the local linearity assumption under normal operation and gradual degradation scenarios; and by using confidence as the regression weight, the influence of high-confidence anchor points on the fitting results is greater, thereby reducing the interference of adjacent outlier sampling points on the interpolation results.
[0077] The second scenario involves sampling points where the sampled value exists but the confidence score is below the threshold. For these sampling points, a locally weighted linear regression model is constructed using three valid anchor points before and after the original value to obtain the estimated value x_est. The corrected value x_cor is obtained by weighting and fusing the original value x_raw and the estimated value x_est according to their respective confidence scores. x_cor = (C_raw x_raw + C_est x_est) / (C_raw + C_est) Where C_raw is the original confidence score of the sampling point, and C_est is the repair reliability estimated based on the number of anchor points and the mean confidence score of the anchor points, and its calculation formula is: C_est = N_anchor / N_max C_anchor_mean Where N_anchor is the actual number of available anchors, N_max is the maximum number of anchors (N_max is six when there are three on each side), and C_anchor_mean is the mean confidence score of all anchors. The meaning of this weighted fusion method is: when the confidence of the original sampled value is relatively high, the corrected value retains more of the contribution of the original value; when the repair reliability is high (sufficient anchors and high mean confidence score of anchors), the corrected value adopts the estimated value more.
[0078] After the repair is completed, the confidence score is recalculated for the pre-repaired sampling points. The repaired confidence score C_rep is the smaller of the product of the original confidence score C_raw and the repair gain factor gamma, and the preset repair upper limit C_cap. C_rep = min(C_raw gamma, C_cap) Where gamma = 1 + alpha C_est and alpha are the repair gain coefficients, preferably 0.6, and C_cap is preferably 0.75. The design intent of the above repair upper limit is that: the pre-repair is based on local linear approximation and neighborhood anchor information, and its repair quality must have a certain upper bound. By constraining the upper limit of the repaired confidence to 0.75, it is ensured that the confidence of the pre-repaired sampling points is always lower than that of the original high-confidence sampling points, thus maintaining a reasonable priority order in subsequent transmission orchestration.
[0079] It is noteworthy that when multiple low-confidence sampling points appear consecutively within a certain time window under concurrent abnormal operating conditions, the number of effective anchor points that can be utilized by the locally weighted linear regression will decrease accordingly. When the total number of effective anchor points on both sides is less than two, the fitting accuracy of the locally linear regression will decrease significantly. In this case, it is necessary to revert to using the weighted average of the available anchor points directly as the filler value, and reduce C_est to 0.2 to reflect the decrease in repair reliability. This boundary condition handling ensures the robustness of the pre-repair module under extreme data sparsity scenarios.
[0080] Furthermore, a local consistency check step is introduced after the pre-repair is completed. For each repaired sampling point, the mean and standard deviation of its repaired value are calculated compared to the mean and standard deviation of all high-confidence sampling points (confidence level not lower than 0.7) within the time window (five sampling points before and after the point centered on that point). The standardized deviation of the repaired value is then calculated. z= |x_cor - mu_local| / (sigma_local + epsilon) Wherein, mu_local and sigma_local are the mean and standard deviation of the high-confidence sampling points within the time window, respectively, and epsilon is a small value to prevent division by zero, preferably 0.001. When z exceeds the preset deviation threshold (preferably 3.0, corresponding to three times the standard deviation), it is determined that the repaired value is inconsistent with the local distribution. The confidence score after repair is reduced to 0.7 times the original confidence score, and the verification anomaly mark of the sampling point is recorded in the working condition summary metadata data package for the main station to refer to in subsequent refined processing. This local consistency verification forms a closed-loop quality assurance process of "scoring-repairing-verification-re-scoring", effectively avoiding the risk of the repaired value deviating from the actual distribution that may occur in the one-way repair process.
[0081] Beneficial effects
[0082] The core benefit of the three-dimensional per-sampling-point confidence scoring mechanism lies in refining the quality assessment of power sampling data from the granularity of the acquisition cycle to the granularity of a single sampling point. Furthermore, the assessment dimensions cover the data content layer (physical rationality), the temporal structure layer (temporal continuity), and the hardware awareness layer (channel health). The abnormal patterns captured by these three dimensions are physically independent and complementary. Specific effects are reflected in the following aspects.
[0083] In terms of data quality assessment accuracy, single-dimensional quality assessment methods (such as judging solely based on physical range) cannot distinguish between systematic shifts caused by hardware degradation and isolated jumps caused by transient interference. However, the combination of three-dimensional weighted scoring results in the former receiving a low score in the channel health dimension and the latter receiving a low score in the temporal continuity dimension. Both types of anomalies are accurately reflected in the confidence score, and the scoring results have a more robust ability to distinguish between different anomaly types.
[0084] Regarding online pre-repair performance, the confidence score serves as the weighting criterion for locally weighted linear regression, enabling the repair process to automatically reduce the interference from adjacent low-quality sampling points and focus on fitting using truly reliable anchor points within the local neighborhood. Compared to equal-weighted interpolation methods, the confidence-weighted approach produces estimation results closer to the true values in scenarios with inconsistent anchor point quality.
[0085] In terms of transmission orchestration effectiveness, the confidence score serves as the basis for determining the transmission threshold. This ensures that when channel resources are strained, the lowest quality data is discarded first, while the highest quality and most valuable subset of data is retained. The dynamic confidence transmission threshold increases accordingly with the severity of the operating conditions, ensuring that the overall quality of the transmitted dataset remains at a reasonable level under extreme conditions, rather than experiencing proportional sparsity that leads to equal data loss across all quality levels.
[0086] Regarding the utilization of data on the main station side, the working condition summary metadata data package contains statistical information such as the distribution of the number of sampling points in each confidence interval, the change in average confidence before and after repair, and the count of abnormal points in local consistency verification. It provides the main station with a complete quality profile of the current batch of data. Based on this, the main station can configure reasonable confidence weights for subsequent data fusion algorithms or initiate supplementary sampling requests for data periods with insufficient confidence.
[0087] Principle Analysis The technical principle behind the three-dimensional confidence scoring mechanism achieving the above-mentioned effect can be analyzed from the following perspectives.
[0088] The effectiveness of physical rationality scores stems from the physical constraints of power parameters. Within normal operation and permissible fault ranges, voltage and current parameters in a power system are constrained by the impedance characteristics of the power supply and distribution network, and their numerical ranges have definite physical upper and lower bounds. When gain drift occurs in the acquisition channel, analog-to-digital converter malfunctions, or sensor damage occurs, the probability of sampled values exceeding this range increases significantly. Therefore, physical rationality scores can effectively capture such systemic anomalies at the hardware level.
[0089] The effectiveness of the timing continuity score stems from the assumption of smoothness in power frequency signals. Under normal operating conditions, the variation of the power frequency parameters (50 Hz) of the power system within adjacent sampling intervals (approximately 0.5 ms when 16.7 ms corresponds to 60 sampling points in one acquisition cycle) is constrained by system inertia and will not exhibit jumps exceeding the normal variation threshold by several times. When the sampling channel is subjected to impulsive interference, electromagnetic noise, or concentrated digital quantization errors, the absolute value of the difference between the sampling point and its neighboring sampling points will significantly deviate from the normal level, thus lowering the timing continuity score. This dimension has a strong ability to identify isolated pulse-type sampling anomalies, compensating for the insufficient sensitivity of the physical rationality dimension to such anomalies.
[0090] The effectiveness of channel health scores stems from the systematic correlation between device-level status signals and data quality. Increased noise levels in a acquisition channel typically indicate a deterioration in the channel's signal-to-noise ratio (SNR). Samples generated on channels with deteriorated SNR have lower overall reliability; this is a priori quality discount that can be determined before the specific sample value is calculated. Incorporating this prior information into the confidence score in the form of a channel health score allows the confidence score to simultaneously possess the dual capabilities of post-hoc evaluation of data content (physical plausibility and temporal continuity) and prior evaluation of the data source (channel health). This provides a valid quality reference even when data content-level information is insufficient (e.g., missing sampling points).
[0091] The local consistency verification step utilizes the signal distribution described by high-confidence sampling points within a local time window as a reference benchmark. In most practical power data acquisition scenarios, even under concurrent abnormal conditions, not all sampling points will simultaneously experience severe anomalies within a time window. The distribution characteristics exhibited by the local high-confidence sampling points can represent the true signal statistical characteristics of that period. If the repaired value deviates from this distribution by more than three standard deviations, it indicates that the repair result is likely to deviate from the true signal trajectory. Timely reduction of its confidence level can effectively prevent erroneous repaired values from misleading the main station's data fusion processing after entering the transmission dataset.
[0092] Example 3
[0093] In another embodiment of the present invention, the abnormal operating condition data self-healing transmission method replaces the dual-branch shared embedding structure with an operating condition perception model based on multi-scale temporal feature fusion, and replaces the per-sample-point confidence score with sliding window segmented statistical features, thus constituting an alternative solution with different technical paths in both the core aspects of operating condition perception modeling and data quality assessment. This embodiment is applicable to hardware configuration scenarios where the terminal processor supports lightweight one-dimensional convolution operations but is not suitable for deploying fully connected layer networks, such as the upgrading of old terminals using low-power digital signal processors.
[0094] In constructing the working condition perception model, this embodiment replaces the joint state perception vector with a joint state temporal feature matrix with a time-series input format. Specifically, the joint state perception vectors from the most recent sixteen acquisition cycles are stacked in chronological order into a two-dimensional feature matrix, where the row dimension corresponds to the time step and the column dimension corresponds to each state index component. The working condition perception model adopts a multi-scale one-dimensional convolutional network structure, with three sets of one-dimensional convolutional branches with different kernel widths set in parallel at the input layer. The three sets of convolutional kernel widths are 2, 4, and 8, respectively, to address the perception capabilities of short-range, medium-range, and long-range change patterns of temporal features. Each of the three sets of convolutional branches performs depthwise separable convolution on all feature channels, and the number of channels in the output feature map of each set is preferably set to sixteen. The outputs of the three sets of branches are concatenated along the channel dimension and then compressed into a fixed-length feature vector by global average pooling, which is then respectively connected to the working condition type classification head and the severity regression head.
[0095] The use of depthwise separable convolution can reduce the number of parameters by about eight to ten times compared to ordinary convolution, making it suitable for deployment on computing-constrained terminals. The design motivation for the multi-scale parallel convolution branch is that the abnormal operating conditions of power terminals are usually manifested in both short-range abrupt changes (such as voltage drops caused by instantaneous overload) and long-range trends (such as the slow increase in sampling noise levels caused by equipment aging). A single convolution kernel width is difficult to cover feature extraction at both time scales, while the multi-scale parallel structure can cover the perception capabilities of different time scales without increasing the network depth.
[0096] The training dataset for this multi-scale one-dimensional convolutional condition-aware model is the same as that for the dual-branch shared embedding model in Example 2, both derived from terminal and channel state annotation data in historical operation records. Since the model in this example uses a time-series sequence of length sixteen as input instead of a single vector, the annotation sequence needs to be cut into samples using a sliding window during the training dataset construction phase. The time step offset between adjacent samples is set to one, i.e., a dense sliding window with a stride of one is used to maximize the number of training samples. Training uses the same Adam optimizer as in Example 2, with a learning rate of 0.0005 and a batch size of thirty-two. Because the input format is a time-series matrix, the batch size is halved to control memory usage. The number of training epochs is set to eighty, and an early stopping mechanism is also used.
[0097] In terms of data quality assessment, this embodiment replaces the three-dimensional confidence score for each sampling point with a sliding window segmented statistical feature. Specifically, the sampling sequence of each power parameter in the current acquisition period is divided into several segments by a sliding window of length ten. The time step offset between adjacent segments is five (i.e., there is an overlap of five sampling points between adjacent segments). The following statistical features are calculated for each segment: the mean, standard deviation, mean absolute difference of the sampling points within the segment, and the proportion of sampling points within the segment that exceed the physically reasonable range. The above four statistical features together constitute the quality feature vector of the segment. The Euclidean distance between this quality feature vector and the quality feature vector of the corresponding segment in the historical normal period is used as the anomaly measure of the segment. The smaller the anomaly measure, the closer the statistical characteristics of the segment are to the normal state.
[0098] The anomaly measure of segmented quality features is normalized, and its complement is taken as the segmented confidence score. That is, the segmented confidence score equals one minus the normalized anomaly measure. The segmented confidence score also ranges from zero to one, but its granularity is segment-level rather than sampling point-level. All sampling points within each segment share the confidence score value of that segment. Compared with the sampling point-by-sampling confidence score in Example 2, the computational cost of segmented confidence score is significantly reduced, but at the same time, the ability to identify uneven anomaly distributions within segments is lost. When only a few sampling points in a segment are abnormal while most sampling points are normal, the segmented confidence score is dominated by the statistical characteristics of the majority of normal points, and its sensitivity to a few abnormal points is not as good as the sampling point-by-sampling score method. This applicable boundary means that when the terminal's operating conditions are mainly characterized by isolated sampling point abrupt anomalies, the sampling point-by-sampling confidence score scheme (Example 2) is better than the segmented statistical scheme; while when the data anomalies caused by the operating conditions are mainly manifested as a systematic deterioration of the entire data quality, the difference in performance between the two schemes is relatively small.
[0099] In the pre-repair phase, this embodiment employs a segmented replacement strategy based on segmented confidence levels rather than point-by-point interpolation. For segments with confidence levels below a preset segmentation threshold (preferably 0.4), the weighted average statistical characteristics of the three most recent high-confidence segments in the same path parameter are used as a reference. The sampling points within the low-confidence segments are then replaced entirely with a random simulation sequence based on the reference statistical characteristics. Specifically, a normally distributed random sequence of the same length as the original segment is generated, centered on the mean of the reference segment and measured by the standard deviation of the reference segment, as the replacement value sequence. The generation of this random simulation sequence does not rely on precise estimation of specific missing values, but rather maintains the overall statistical distribution of sampling points within the segment consistent with historical normal segments, thus meeting the requirements of the master station for the rationality of data distribution when performing statistical data analysis (such as power quality statistics and load characteristic analysis).
[0100] In terms of transmission orchestration, this embodiment uses segment confidence instead of sampling point confidence in determining the dynamic transmission threshold. The minimum granularity of transmission screening is increased from a single sampling point to a single segment (ten sampling points). The compression orchestration strategy remains consistent with step S5, still dynamically adjusting the transmission threshold based on the severity score and determining the compression level based on the available bandwidth estimate. The priority queue sorting rules are also the same as in step S5. Since the screening granularity is segmented, the transmitted data is composed of complete segments as the basic unit, avoiding the non-continuous sparse sampling sequences that may be generated by screening point by point. This allows the master station to directly perform statistical analysis on a segment-by-segment basis after receiving the data without having to process data gaps beforehand.
[0101] In terms of the content of the operating condition summary metadata, this embodiment adds segment-level quality statistics, including the number of high-confidence segments, the number of low-confidence segments, and the number of segments that have undergone segment replacement processing. Compared with the sampling point-level quality statistics in Embodiment 2, the segment-level statistics are more compact in terms of metadata volume, which is beneficial to reduce the transmission overhead of metadata in extreme channel resource-constrained scenarios.
[0102] Example 4
[0103] This embodiment provides a self-healing data transmission system for abnormal operating conditions in power acquisition terminals, such as... Figure 4 As shown, the system is deployed on the power acquisition terminal side and includes a joint state awareness module, a concurrent anomaly identification module, a sampling point confidence scoring module, an online pre-repair module, an adaptive transmission orchestration module, and a working condition summary generation module. Each module operates collaboratively locally on the terminal to achieve self-healing data transmission under abnormal working conditions. The system can be deployed and run on embedded power acquisition terminals equipped with low-power processors or microcontrollers. The hardware platform does not require dedicated acceleration chips; ordinary ARM Cortex-M series processors or processors with equivalent computing power can meet the real-time inference requirements of each module.
[0104] The joint state perception module synchronously acquires terminal device layer state signals and communication channel layer state signals in each data acquisition cycle. These signals are then concatenated in a preset order to form a fixed-dimensional joint state perception vector. Each component of the joint state perception vector undergoes online normalization processing based on sliding window historical statistics. The terminal device layer state signals include the noise level, sampling integrity rate, main control unit temperature, and power supply voltage fluctuation amplitude of each sampling channel. The communication channel layer state signals include the current channel's signal-to-noise ratio, round-trip delay, packet loss rate, and estimated available bandwidth. When the terminal has Nc sampling channels, the total dimension of the joint state perception vector is 2Nc+6. The online normalization processing maintains sliding window statistics for each component over the most recent thirty acquisition cycles, saving only the cumulative sum and cumulative squared sum for incremental updates, meeting the deployment requirements of terminal storage constraints. The output information of the joint state perception module flows to the concurrent anomaly identification module, while the noise level indicators of each sampling channel are retained for direct reuse by the downstream sampling point confidence scoring module, avoiding duplicate acquisition.
[0105] The concurrent anomaly identification module is used to input the normalized joint state-aware vector into a lightweight concurrent anomaly identification model deployed on the terminal side, and output the operating condition type label and the operating condition severity score. The lightweight concurrent anomaly identification model adopts a dual-branch shared embedding structure. Through a shared embedding layer, the joint state-aware vector is mapped from its original dimension to a 32-dimensional low-dimensional embedding representation. Then, a operating condition type classification branch outputs the probability distributions for four types of operating conditions, and a severity regression branch outputs a severity scalar from zero to one. The shared embedding layer consists of two cascaded fully connected networks with a ReLU activation function in between. The classification branch consists of a fully connected network connected to a Softmax activation function, outputting four probability distributions: normal operating condition, equipment degradation only, channel degradation only, and equipment and channel concurrent anomaly operating condition. The regression branch consists of a fully connected network connected to a Sigmoid activation function, outputting the operating condition severity score scalar. The two branches share the embedding layer parameters, completing both operating condition type identification and severity quantification simultaneously through a single forward propagation. When the operating condition type label is "device and channel concurrent abnormal operating condition", the concurrent abnormal identification module sends a trigger signal to the subsequent modules and transmits the operating condition severity score to the adaptive transmission orchestration module. When the operating condition type label is "normal operating condition", "device only deterioration condition", or "channel only deterioration condition", the terminal performs data reporting according to the conventional transmission logic.
[0106] The sampling point confidence scoring module calculates a confidence score for each sampling point in the raw data of the current acquisition cycle. The confidence score is obtained by weighting the physical rationality score, temporal continuity score, and channel health score with preset weights of 0.4, 0.35, and 0.25, respectively. The physical rationality score is determined by whether the sampling point value falls within a preset physical rationality range of the corresponding power parameter; if it falls within the range, the score is one; if it exceeds the range, the score linearly decays to zero according to the degree of exceedance. The temporal continuity score is determined by calculating the absolute value of the difference between the sampling point and the sampling point before and after it, and comparing it with a preset normal change threshold; the smaller the absolute value of the difference, the higher the score. The channel health score is directly taken as the normalized inverse value of the noise level index of the sampling channel already acquired by the joint state perception module, 1 / (1+noise_norm), without additional acquisition. The confidence score results are simultaneously transmitted to the online pre-repair module and the adaptive transmission orchestration module in the form of a point-by-point scoring vector.
[0107] The online pre-repair module is used to perform online pre-repair on sampling points with confidence scores lower than a preset confidence threshold, which is 0.5. For sampling points with completely missing values, an imputation method based on locally weighted linear regression is configured to be used. This method uses three valid sampling points before and after the missing point as anchor points, and the confidence score of each anchor point as the regression weight to fit a locally linear model y=a×t+b and predict the missing value. For sampling points with existing values but confidence scores lower than the preset confidence threshold, the locally weighted linear regression method is configured to calculate the estimated value at that location. The original value and the estimated value of the sampling point are then weighted and fused using the original confidence score of the sampling point and the repair reliability determined based on the number of anchor points and the mean confidence score of the anchor points, respectively, to obtain a corrected value. For sampling points that have undergone pre-repair processing, the confidence score is recalculated. The repaired confidence score is the smaller of the product of the original confidence score and the repair gain factor, and the preset repair upper limit of 0.75, to ensure that the repair operation does not produce an overly optimistic confidence increase. The output of the online pre-repair module is the repaired data sequence and the updated point-by-point confidence score vector, which are then passed to the adaptive transmission orchestration module.
[0108] The adaptive transmission orchestration module generates a transmission orchestration scheme based on the operational severity score output by the concurrent anomaly identification module and the channel layer state signal collected by the joint state awareness module. This scheme includes three decision dimensions: the scope of data filtering, the data compression granularity, and the transmission priority queue. In the data filtering scope dimension, a dynamic confidence transmission threshold Th_dyn = Th_base + S_severity × K_adj is calculated based on the operational severity score. Only sampling points with confidence scores higher than the dynamic transmission threshold are included in the dataset to be transmitted; the more severe the operational condition, the higher the threshold and the stricter the filtering. In the data compression granularity dimension, the compression level is determined based on the ratio R_bw of the estimated available bandwidth to the amount of data to be transmitted. When R_bw is less than one, a confidence-weighted sampling point aggregation strategy is used for lossy compression. The aggregation window width varies with the compression level from level one to level three, being two, four, and eight points respectively, with high-confidence sampling points receiving greater weight in the aggregation. In the transmission priority queue dimension, it is configured to be sorted according to the service criticality level of voltage, current, power, and energy. Within the same level, it is sorted from high to low according to the average confidence score to generate the final transmission priority queue. The adaptive transmission orchestration module passes the key parameters of the transmission orchestration scheme to the operating condition summary generation module, and at the same time drives the terminal communication module to send data packets in sequence.
[0109] The operating condition summary generation module generates operating condition summary metadata data packets and transmits them to the main station as the highest priority data packets. The operating condition summary metadata data packets include the timestamp of the current acquisition period, operating condition type label and operating condition severity score, the total number of sampling points in the raw data, the distribution of sampling point numbers in each confidence interval, the number of sampling points to be pre-repaired and the average confidence change before and after repair, the dynamic confidence transmission threshold and compression level, and the proportion of sampling points actually included in the transmission to the total number of sampling points. When concurrent abnormal operating conditions persist for more than five consecutive acquisition periods, the operating condition summary generation module is also configured to append the cumulative abnormal duration and the linear regression slope of the operating condition severity score sequence for each period to the operating condition summary metadata data packets, allowing the main station to assess the damage evolution of the terminal equipment and plan emergency response measures in advance. The information sources of the operating condition summary generation module are aggregated from the concurrent abnormality identification module, the sampling point confidence scoring module, the online pre-repair module, and the adaptive transmission orchestration module, forming a transparent record of the entire data processing process for this acquisition period.
[0110] Example 5
[0111] This embodiment describes the complete implementation process and technical effects of the method described in Embodiment 1 in a concurrent abnormal scenario where mechanical vibration of a power acquisition terminal causes damage to multiple channels of sensors, and at the same time, severe fading of the wireless channel occurs.
[0112] The data acquisition terminal is deployed outdoors and is responsible for periodically acquiring four parameter channels (Nc=4): voltage, A-phase current, B-phase current, and C-phase current of the power distribution line, at a frequency of once every fifteen minutes. During a period of strong winds, mechanical vibration caused the sampling circuits of the A-phase and B-phase current channels of the terminal to loosen. Simultaneously, the local wireless communication network experienced severe channel fading, resulting in an extreme situation of simultaneous equipment damage and channel degradation. Under these conditions, conventional methods struggle to simultaneously address the dual constraints of data distortion and channel limitations: directly transmitting the raw acquired data would cause distorted data to interfere with data analysis at the master station; waiting for the channel to recover before transmitting the full data could lead to prolonged data loss, affecting the continuity of power grid condition monitoring.
[0113] The execution process of step S1 is as follows. The terminal collects the noise level values of the four channels within this collection period. The normalized noise level of the A-phase current channel is 2.37, the B-phase current channel is 1.84, and the voltage channel and C-phase current channel are 0.12 and 0.19 respectively, which are within the historical normal range. The signal-to-noise ratio of the wireless channel drops to 8.3 dB (historical average is about 22 dB), the round-trip delay increases to 420 ms (historical average is about 80 ms), the packet loss rate reaches 18.5%, and the estimated available bandwidth drops to about 31% of the historical average. The terminal concatenates the above indicators with the sampling integrity rate (A-phase 0.73, B-phase 0.81, C-phase 0.97, voltage 0.99), the main control unit temperature (normalized value 0.08), and the power supply voltage fluctuation amplitude (normalized value 0.11) in a preset order to generate a normalized joint state perception vector with a dimension of 2×4+6=14.
[0114] The execution process of step S2 is as follows. The lightweight concurrent anomaly identification model performs a forward propagation on the above joint state-aware vector. The classification branch outputs the probability distribution of four operating conditions: normal operating condition 0.03, equipment degradation only 0.11, channel degradation only 0.07, and equipment and channel concurrent anomaly operating condition 0.79. The category with the highest probability is selected, and the operating condition type label is identified as equipment and channel concurrent anomaly operating condition. The regression branch outputs an operating condition severity score of 0.63, reflecting that the current operating condition is at a medium-to-high severity level. The above identification results trigger the self-healing transmission process from steps S3 to S6.
[0115] The execution process of step S3 is as follows. The terminal calculates the confidence score for each of the 960 sampling points (240 for each of the four channels) in this acquisition cycle. Taking a sampling point of the A-phase current channel as an example: the sampled value of this point is within the reasonable range according to the physical rationality judgment, and the physical rationality score is 1.00; the absolute value of the difference with the previous and subsequent sampling points exceeds the normal change threshold by about 1.8 times, and the timing continuity score is max(0, 1-0.8)=0.20; the channel health score is 1 / (1+2.37)=0.30. The comprehensive confidence score C=0.40×1.00+0.35×0.20+0.25×0.30=0.40+0.07+0.075=0.545. Although the sampled value of this sampling point is within the reasonable range, the confidence score is only 0.545 due to the sudden change in timing and the high channel noise. Of the 960 sampling points, 214 had a confidence score below 0.50. These were mainly concentrated in the A-phase current channel (103 points) and the B-phase current channel (87 points), with 12 points each in the voltage channel and the C-phase current channel.
[0116] The execution process of step S4 is as follows. For the aforementioned 214 low-confidence sampling points, the terminal performs online pre-repair. Among them, 61 sampling points are completely missing (37 in phase A current channels and 24 in phase B current channels), which are processed using an interpolation method based on local weighted linear regression. The number of anchor points for each missing point is between 4 and 6, and the corresponding repair gain factor is between 1.22 and 1.50. 153 sampling points have sampling values but a confidence level below 0.50, which are processed using a confidence-weighted fusion correction method. After repair, the confidence score for all repaired points is recalculated. Constrained by a preset repair upper limit of 0.75, the confidence score of all repaired sampling points does not exceed 0.75. The average confidence level of the 214 low-confidence sampling points before repair is 0.29, which increases to 0.61 after repair, representing an average confidence level increase of 0.32.
[0117] The execution process of step S5 is as follows. Based on the severity score of 0.63, the dynamic confidence transmission threshold Th_dyn = 0.30 + 0.63 × 0.40 = 0.552 is calculated. Sampling points with a confidence score below 0.552 are excluded from the transmission range. After screening, a total of 731 sampling points are actually included in the transmission, accounting for 76.1% of the total 960 sampling points. The estimated available bandwidth of the channel is approximately 0.38 times the bandwidth required for the amount of data to be transmitted, corresponding to R_bw = 0.38, falling within the secondary compression range (0.25 to 0.50). The terminal performs secondary lossy compression on the screened data, aggregating four adjacent sampling points into a representative value, and calculating the aggregated representative value using a confidence-weighted average. After compression, the actual number of data packets transmitted is approximately one-quarter of the number of sampling points after screening, further reducing channel occupancy. The transmission priority queue is sent sequentially in the order of voltage channel (highest service criticality), C-phase current channel (optimal confidence distribution), B-phase current channel, and A-phase current channel.
[0118] The execution process of step S6 is as follows. The terminal generates a working condition summary metadata data packet and sends it first as the highest priority data packet. The metadata data packet contains the above-mentioned statistical fields and has a data volume of approximately 380 bytes. Under the current packet loss rate of 18.5%, it is successfully delivered to the master station after retransmission. After receiving the working condition summary metadata data packet, the master station determines that the transmission integrity rate of this batch of data is 76.1%, the confidence level is concentrated in the range of 0.55 to 0.75, and learns that the data quality of the A-phase current channel and the B-phase current channel is low. Based on this, the post-event refined completion process is initiated, and the missing segments are reconstructed more accurately by combining the correlation data of adjacent terminals.
[0119] To verify the technical contributions of each innovative step in the method of this invention, comparative experiments were conducted under the same simulation scenario. The traditional direct transmission scheme was used as the baseline comparison method. This method does not perform condition identification and data quality assessment, but directly compresses and transmits the original data. The ablation scheme after removing the confidence scores of sampling points and the online pre-repair steps (i.e., removing steps S3 and S4) was used as another baseline comparison. The evaluation metrics selected were the effective sampling rate of the data received at the main station (the proportion of sampling points with a confidence score of not less than 0.50 out of all sampling points that should be collected) and the data credibility at the main station (the average confidence score of the received dataset).
[0120] Experimental results show that the traditional direct transmission scheme has an effective sampling rate of 41.3% under this concurrency anomaly scenario, with an average data confidence score of 0.38 on the master station side. A large amount of low-confidence data occupies channel resources, resulting in a low effective information delivery rate. The ablation scheme removing steps S3 and S4 improves the effective sampling rate to 58.7% and the average data confidence score on the master station side to 0.51. The dynamic confidence threshold and compression strategy in the transmission orchestration scheme play a filtering role, but because no pre-repair is performed on low-confidence data, some repairable data is directly excluded by the threshold, leaving room for further improvement in the effective sampling rate. The complete method of this invention achieves an effective sampling rate of 71.8% and an average data confidence score on the master station side of 0.63, representing improvements of 30.5 percentage points and 0.25 points respectively compared to the traditional direct transmission scheme.
[0121] like Figure 7 As shown, the horizontal axis represents the three comparison schemes, the left vertical axis (grayscale bar chart) represents the effective sampling rate (%), and the right vertical axis (black solid line chart with square markers) represents the average data reliability. The effective sampling rate of the traditional direct transmission scheme is 41.3%, and the average data reliability is 0.38; the effective sampling rate of the S3 / S4 ablation removal scheme is improved to 58.7% (an improvement of +17.4 percentage points compared to the traditional scheme), and the average data reliability is improved to 0.51 (+0.13); the effective sampling rate of the complete method of this invention reaches 71.8% (an improvement of +30.5 percentage points compared to the traditional scheme), and the average data reliability reaches 0.63 (+0.25). The above results show that the confidence score per sampling point (step S3) and online pre-repair (step S4) make significant contributions to improving data quality. Compared with the traditional scheme, the complete method of this invention achieves a systematic improvement in both effective sampling rate and reliability. Under different severity conditions (severity scores of 0.30, 0.50, 0.70, and 0.90), the effective sampling rates of the method of this invention were 87.2%, 76.4%, 68.3%, and 58.1%, respectively, showing a reasonable trend of moderate decrease with increasing severity. This reflects the mechanism characteristics of ensuring critical data transmission through screening and aggregation under extremely limited resource conditions. The above results demonstrate that the method of this invention can effectively improve the validity and reliability of transmitted data under concurrent abnormal operating conditions, and provides sufficient evidence for post-event completion by the master station through the operating condition summary metadata.
[0122] To analyze the adaptive performance of this invention under different severity levels, system tests were conducted on a continuous interval of severity scores from 0.0 to 1.0 (step size 0.1). The test method involved fixing channel parameters and injecting artificial disturbances at different amplitudes into the device and channel layers, corresponding to different severity levels. The actual values of the effective sampling rate and the dynamic confidence transmission threshold Th_dyn of the complete method of this invention were recorded under each condition. To reflect the performance differences between channels, the distribution range of the effective sampling rate for the four parameter channels was statistically analyzed as the performance fluctuation range.
[0123] like Figure 8 As shown, the horizontal axis represents the severity score (0.0~1.0), the left vertical axis (solid black line + circular marker, including light gray shading) represents the effective sampling rate (%), and the right vertical axis (dashed black line) represents the dynamic confidence transmission threshold Th_dyn. As severity increases, the effective sampling rate gradually decreases from approximately 87.2% in the mild degradation zone (0~0.3) to approximately 58.1% in the severe degradation zone (0.7~1.0), exhibiting controlled, gradual degradation rather than abrupt failure, demonstrating the adaptive transmission characteristic of this invention without interruption during degradation. The dynamic threshold Th_dyn increases linearly according to the formula Th_dyn = 0.3 + severity × 0.4, increasing linearly from 0.30 to 0.70, automatically raising the screening threshold as severity increases, ensuring that the main station prioritizes receiving high-quality data. The light gray shaded bands reflect the performance fluctuation range between different channels. Each channel has a difference of approximately ±4% to ±7% at the same severity level, indicating that the present invention achieves fine-grained confidence differentiation at the channel level.
[0124] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A self-healing data transmission method for abnormal operating conditions in power acquisition terminals, characterized in that, Includes the following steps: Step S1: The acquisition terminal synchronously acquires the terminal device layer status signal and the communication channel layer status signal in each acquisition cycle, concatenates the two into a joint status perception vector and performs online normalization based on sliding window historical statistics. Step S2: Input the normalized joint state awareness vector into the lightweight concurrent anomaly identification model on the terminal side, and output the working condition type label and the working condition severity score; the lightweight concurrent anomaly identification model adopts a dual-branch shared embedding structure. After being mapped to a low-dimensional embedding representation through the shared embedding layer, it outputs four types of working condition probability distributions through the classification branch and outputs a severity scalar from zero to one through the regression branch; when the working condition type label is a device and channel concurrent anomaly working condition, steps S3 to S6 are triggered; Step S3: Calculate the confidence score for each sampling point in the raw data of the current acquisition cycle. The confidence score is obtained by weighting the physical rationality score, the temporal continuity score, and the channel health score with preset weights. Step S4: Perform online pre-repair for sampling points with confidence scores lower than the preset confidence threshold; Step S5: Generate a transmission orchestration scheme based on the severity score of the operating condition and the channel layer status signal, including the filtering range of transmission data, the compression granularity of the data, and the transmission priority queue. Step S6: Generate a working condition summary metadata data packet and transmit it to the main station as the highest priority data packet.
2. The method according to claim 1, characterized in that, In step S1, the terminal device layer status signal includes the noise level, sampling integrity rate, main control unit temperature, and power supply voltage fluctuation amplitude of each sampling channel; the communication channel layer status signal includes the signal-to-noise ratio, round-trip delay, packet loss rate, and available bandwidth estimate of the current channel; the online normalization is standardized by the mean and standard deviation of the components within the most recent preset number of sampling periods.
3. The method according to claim 2, characterized in that, In step S3, the physical rationality score is determined by judging whether the value of the sampling point falls within the preset physical rationality range of the corresponding power parameter. If it falls within the range, the score is one; if it exceeds the range, the score is linearly reduced to zero according to the degree of exceedance. The temporal continuity score is determined by calculating the absolute value of the difference between the sampling point and the sampling point before and after it and comparing it with the preset normal change threshold. The smaller the absolute value of the difference, the higher the score. The channel health score is the inverse value of the noise level index of the sampling channel in step S1 after normalization.
4. The method according to claim 3, characterized in that, The preset weights for the physical rationality score, temporal continuity score, and channel health score are 0.4, 0.35, and 0.25, respectively.
5. The method according to claim 1, characterized in that, In step S4, for sampling points with completely missing values, an imputation method based on locally weighted linear regression is used. A preset number of valid sampling points before and after the missing point are used as anchor points, and the confidence score of each anchor point is used as the regression weight to fit a locally linear model and predict the missing value. For sampling points with existing values but a confidence score lower than the preset confidence threshold, the locally weighted linear regression method is used to calculate the estimated value of the sampling point at that location. The original value of the sampling point and the estimated value are then weighted and fused using the original confidence score of the sampling point and the repair reliability determined based on the number of anchor points and the mean confidence score of the anchor points, respectively, to obtain a corrected value. The confidence score is recalculated for the pre-repaired sampling points. The restored confidence score is the smaller of the product of the original confidence score and the repair gain factor and the preset repair upper limit.
6. The method according to claim 5, characterized in that, The preset confidence threshold is 0.5, and the number of anchor points for the locally weighted linear regression is three valid sampling points before and after.
7. The method according to claim 1, characterized in that, In step S5, regarding the data transmission range, a dynamic confidence transmission threshold is set based on the severity score of the operating condition. The dynamic confidence transmission threshold is determined by the sum of the product of the base threshold value, the severity score of the operating condition, and the threshold adjustment coefficient. Only sampling points with confidence scores higher than the dynamic confidence transmission threshold are included in the dataset to be transmitted. Regarding the data compression granularity, the compression level is determined based on the ratio of the estimated available bandwidth of the current channel to the amount of data to be transmitted. When the ratio of the estimated available bandwidth to the amount of data to be transmitted is less than one, a lossy compression strategy based on confidence-weighted sampling point aggregation is adopted. Regarding the transmission priority queue, the transmission order is determined based on the service criticality of the power parameter type as the first sorting criterion and the confidence score as the second sorting criterion.
8. The method according to claim 7, characterized in that, The basic threshold value is 0.3, and the threshold adjustment coefficient is 0.
4.
9. The method according to claim 7, characterized in that, In step S6, the operating condition summary metadata includes the timestamp of the current collection period, the operating condition type label and the operating condition severity score, the total number of sampling points of the original data, the distribution of the number of sampling points in each confidence interval, the number of sampling points to be repaired and the average change in confidence before and after repair, the dynamic confidence transmission threshold and compression level, and the proportion of the number of sampling points actually included in the transmission to the total number of sampling points. When the concurrent abnormal operating condition continues to exceed the preset duration threshold, the cumulative abnormal duration and the trend information of the change in the operating condition severity score of each collection period are added to the operating condition summary metadata.
10. A self-healing data transmission system for abnormal operating conditions in power acquisition terminals, characterized in that, include: The joint state awareness module is used to synchronously acquire the terminal device layer state signal and the communication channel layer state signal in each acquisition cycle, concatenate the two into a joint state awareness vector and perform online normalization based on sliding window historical statistics; The concurrency anomaly identification module is used to input the normalized joint state awareness vector into the lightweight concurrency anomaly identification model on the terminal side, and output the working condition type label and the working condition severity score. The lightweight concurrency anomaly identification model adopts a dual-branch shared embedding structure. After being mapped to a low-dimensional embedding representation through the shared embedding layer, it outputs four types of working condition probability distributions through the classification branch and outputs a severity scalar from zero to one through the regression branch. When the working condition type label is a device and channel concurrency anomaly, the sampling point confidence scoring module, the online pre-repair module, the adaptive transmission orchestration module, and the working condition summary generation module are triggered to run in sequence. The sampling point confidence scoring module is used to calculate a confidence score for each sampling point in the raw data of the current acquisition cycle. The confidence score is obtained by weighting the physical rationality score, the temporal continuity score, and the channel health score with preset weights. The online pre-repair module is used to perform online pre-repair on sampling points with confidence scores lower than a preset confidence threshold; The adaptive transmission orchestration module is used to generate a transmission orchestration scheme based on the severity score of the operating condition and the channel layer status signal, including the filtering range of transmission data, the compression granularity of data, and the transmission priority queue. The operating condition summary generation module is used to generate operating condition summary metadata data packets and transmit them to the main station as the highest priority data packets.
Citation Information
Patent Citations
Computer communication method and system based on Internet of Things
CN121356969A
Cooperative control method and system based on power distribution internet of things low-voltage intelligent switch
CN121485307A