Power utilization inspection hidden danger identification method and system based on light-weight model
Patent Information
- Application Number
- CN202610567637.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-27
- Publication Date
- 2026-09-04
- Estimated Expiration
- 2046-04-27
AI Technical Summary
[0004]针对上述中的相关技术,利用固化的轻量化识别模型直接对现场采集的全量图像数据进行持续的实时推理,产生大量的无效算力损耗,导致边缘计算设备长期处于高负荷与高功耗状态;同时,面对复杂多变的现场工况,轻量化模型对疑难边缘场景的隐患判定准确率较低,极易产生误报或漏报,识别可靠性较差,不利于用电检查业务的长期有效开展
1、通过建立无隐患基准图像,并与现场图像执行像素级形态学减影及连通域外扩,提取出发生物理形态改变的差异化裁剪图;
Smart Images

Figure CN122289800B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the intersection of computer vision and smart grid, and in particular to a method and system for identifying potential hazards in electricity inspection based on a lightweight model. Background Technology
[0002] In the construction of smart grids, safety inspections of distribution areas and user-side electrical facilities are crucial for ensuring the stable operation of the power system. In recent years, utilizing computer vision technology and edge computing devices to identify potential hazards from environmental images at power consumption sites has become a primary technical means for power grid inspections.
[0003] Among related technologies, Chinese invention patent application CN116846059A discloses an edge detection system for power grid inspection and monitoring, including a hardware module, a data processing module, an algorithm module, and a communication module. This system deploys lightweight intelligent terminals around power facility scenarios and utilizes a lightweight recognition model that has undergone model pruning and weight quantization to extract features and classify image data acquired by image acquisition devices. This allows it to identify power grid operation risks such as foreign objects on power lines and equipment damage, and transmits the identification results to the power grid cloud platform.
[0004] Regarding the aforementioned technologies, using a fixed, lightweight recognition model to directly perform continuous real-time inference on all image data collected on-site generates a large amount of ineffective computing power loss, causing edge computing devices to be in a state of high load and high power consumption for a long time. At the same time, in the face of complex and ever-changing on-site conditions, the lightweight model has a low accuracy rate in judging potential hazards in difficult edge scenarios, which is prone to false alarms or missed alarms, resulting in poor recognition reliability and hindering the long-term effective development of electricity inspection services. Summary of the Invention
[0005] To address the aforementioned issues, this invention provides a method and system for identifying potential hazards in electricity inspections based on a lightweight model. It employs morphological initial screening and an edge-cloud collaborative self-evolution mechanism, which can effectively filter out invalid background calculations and resolve misjudgments under difficult operating conditions, thereby improving the computing power utilization of edge computing devices and the accuracy of hazard identification.
[0006] The above objectives can be achieved through the following approach: A method for identifying potential hazards in electricity inspections based on a lightweight model includes: controlling a front-end acquisition device to acquire initial scene images of the power distribution area; performing static background modeling to generate a hazard-free baseline image; and using the front-end acquisition device to receive on-site environmental image data. The on-site environmental image data and the hazard-free baseline image are then subjected to pixel morphological subtraction to extract differential pixel clusters. Connectivity boundary expansion is then performed on the differential pixel clusters to generate a morphologically differentiated cropping map. This morphologically differentiated cropping map is input into a preset lightweight identification model, where feature vector mapping is performed to output initial hazard category labels and numerical confidence scores. The numerical confidence scores are continuously written to a local state buffer. The system constructs a time-sliding window confidence sequence; it extracts the sample mean and sample variance using the time-sliding window confidence sequence, calculates the difficult judgment interval, and when the numerical confidence score falls into the difficult judgment interval, it sends the morphological difference cropping image and the initial hazard category label to the cloud server; it performs high-dimensional feature extraction on the morphological difference cropping image through the cloud server to obtain the verification judgment label, and synchronously receives the incremental fine-tuning weight package through the cloud server; it loads the incremental fine-tuning weight package into the lightweight recognition model to perform local weight update, and performs association recording with the verification judgment label and the on-site environmental image data to generate hazard dispatch order data.
[0007] Optionally, receiving on-site environmental image data using the front-end acquisition device includes: driving the front-end acquisition device to continuously acquire initial scene images during the initialization phase, extracting temporal pixel values at the same coordinate positions according to spatial pixel index, and constructing a temporal feature vector sequence; performing median filtering on the temporal feature vector sequence to extract the background pixel median matrix, and performing spatial feature stitching on the background pixel median matrix to generate a baseline image free of hidden dangers.
[0008] Optionally, the generation of the morphologically differentiated cropping map includes: performing absolute pixel difference calculation on the on-site environmental image data and the no-hazard baseline image to generate a difference feature matrix, and performing binarization processing on the difference feature matrix to extract connected non-zero pixel sets and construct difference pixel clusters; obtaining the two-dimensional spatial extreme coordinates of the difference pixel clusters to construct a minimum bounding rectangle, extending the vertex coordinates of the minimum bounding rectangle outward to generate a cropping anchor frame, and using the cropping anchor frame to perform matrix slicing processing on the on-site environmental image data to generate a morphologically differentiated cropping map.
[0009] Optionally, constructing the time-sliding window confidence sequence includes: calling the lightweight recognition model to perform tensor convolution and pooling operations on the morphological difference cropping image to generate a high-dimensional feature vector; performing linear transformation and normalization operations on the high-dimensional feature vector; extracting the maximum probability value and mapping it to a numerical confidence score and an initial hazard category label; creating a local state buffer queue; appending the numerical confidence score to the tail of the local state buffer queue in time sequence; releasing the head node data when the node is full; and extracting the queue values in sequence to construct the time-sliding window confidence sequence.
[0010] Optionally, sending the morphological difference cropping image and the initial hazard category label to the cloud server includes: calculating the sample mean and sample variance using the time sliding window confidence sequence, and performing a combination operation on the sample mean and sample variance to generate a difficult judgment interval; when the numerical confidence score falls into the difficult judgment interval, calling the communication interface to encapsulate the morphological difference cropping image and the initial hazard category label into an anomaly reporting data packet and sending it to the cloud server.
[0011] Optionally, the method further includes: extracting the image encoding sequence of the morphologically differentiated cropped image, parsing the verification judgment label into a structured semantic field, performing key-value pair mapping and multimodal data encapsulation on the image encoding sequence and the structured semantic field, and constructing a multimodal hidden danger evidence archive.
[0012] Optionally, the step of synchronously receiving the incremental fine-tuning weight package through the cloud server includes: calling the cloud server to perform high-dimensional feature extraction on the morphological difference cropping image and performing a nearest neighbor index retrieval to generate a verification judgment label; performing backpropagation operation using the morphological difference cropping image and the verification judgment label to generate a parameter offset matrix; and performing structural alignment and packaging using the parameter offset matrix to generate the incremental fine-tuning weight package.
[0013] Optionally, generating the hazard dispatch order data includes: parsing the incremental fine-tuning weight package into a state dictionary, loading it into the tensor address of the lightweight recognition model to perform local weight updates; attaching the on-site environmental image data as a panoramic context to the multimodal hazard evidence archive, performing structured mapping through the multimodal hazard evidence archive, executing the association record between the verification judgment label and the on-site environmental image data, and performing business encapsulation to generate the hazard dispatch order data.
[0014] Based on the same inventive concept, this invention also provides a power inspection hazard identification system based on a lightweight model. The system includes: a scene baseline acquisition module, used to control a front-end acquisition device to acquire an initial scene image of the power distribution area, perform static background modeling, generate a hazard-free baseline image, and receive on-site environmental image data using the front-end acquisition device; a morphological difference extraction module, used to perform pixel morphological subtraction on the on-site environmental image data and the hazard-free baseline image, extract difference pixel clusters, and perform connected component boundary expansion on the difference pixel clusters to generate a morphological difference cropping map; and a lightweight inference buffer module, used to input the morphological difference cropping map into a preset lightweight identification model, perform feature vector mapping, output an initial hazard category label and a numerical confidence score, and convert the numerical confidence score into a data structure. A local state buffer queue is continuously written to construct a time-sliding window confidence sequence. A difficult-to-determine interval module is used to extract the sample mean and sample variance from the time-sliding window confidence sequence, calculate the difficult-to-determine interval, and when the numerical confidence score falls into the difficult-to-determine interval, send the morphological difference cropping image and the initial hazard category label to the cloud server. A cloud verification and synchronization module is used to perform high-dimensional feature extraction on the morphological difference cropping image through the cloud server to obtain verification and determination labels, and synchronously receive incremental fine-tuning weight packages through the cloud server. A weight update and distribution module is used to load the incremental fine-tuning weight package into the lightweight recognition model to perform local weight updates, and associate the verification and determination labels with the on-site environmental image data to generate hazard distribution order data.
[0015] Compared with the prior art, the present invention has the following advantages: 1. By establishing a baseline image free of hidden dangers and performing pixel-level morphological subtraction and connected component expansion on the on-site image, differential cropping images that have undergone physical morphological changes are extracted. 2. The interception criteria are automatically adjusted according to the "fluctuation degree" of the model under extreme lighting, occlusion and other interference, which eliminates the defects of manually preset rigid thresholds, effectively intercepts the misjudgment that the lightweight model is prone to in long-tail complex scenarios, and improves reliability and fault tolerance. 3. The data-driven feedback mechanism transforms the shortcomings of small models into growth drivers, enabling them to autonomously adapt to new types of potential problems and reducing long-term algorithm maintenance costs.
[0016] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures pointed out in the description, claims and drawings. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a flowchart illustrating the method for identifying potential electrical hazards based on a lightweight model, according to an embodiment of the present invention.
[0019] Figure 2 This is a visualization time-series distribution diagram of the fluctuation of the time-sliding window confidence score and the dynamic difficulty judgment interval in an embodiment of the present invention.
[0020] Figure 3 This is a violin plot showing the distribution of confidence scores before and after incremental fine-tuning of the lightweight recognition model in this embodiment of the invention.
[0021] Figure 4 This is a schematic diagram of the electrical inspection hazard identification system based on a lightweight model, according to an embodiment of the present invention. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0023] Reference Figure 1 One embodiment of the present invention proposes a method for identifying potential hazards in electricity inspection based on a lightweight model. It adopts a morphological initial screening and an edge-cloud collaborative self-evolution mechanism, which can effectively filter invalid background calculations and solve the problem of misjudgment in difficult working conditions, thereby improving the computing power utilization rate of edge computing devices and the accuracy of hazard identification.
[0024] The method described in this embodiment specifically includes: The system controls the front-end acquisition device to acquire initial scene images of the power distribution area, performs static background modeling, generates a baseline image free of hidden dangers, and uses the front-end acquisition device to receive on-site environmental image data. Optionally, receiving on-site environmental image data using the front-end acquisition device includes: The driver front-end acquisition device continuously acquires initial scene images during the initialization phase, extracts temporal pixel values at the same coordinate positions according to spatial pixel index, and constructs a temporal feature vector sequence. The hardware devices deployed in the power distribution area to capture optical images, i.e., front-end acquisition equipment, continuously acquire image data (initial scene image) during the initialization phase when no anomalies or changes occur. Based on the spatial pixel index (located by the intersection of horizontal and vertical coordinates) in the image, the brightness values (i.e., temporal pixel values) recorded at different time points for the same coordinate location are extracted. The temporal pixel values located at the same spatial pixel index are arranged in chronological order to construct a temporal feature vector sequence. This temporal feature vector sequence reflects the light and shadow fluctuation characteristics of a specific physical location over a period of time. This process can be described mathematically as follows: , Among them, letters Represents spatial pixel index and The sequence of temporal feature vectors at the location; letters Represents the time node Temporal pixel values of the initial scene image acquired; letters This represents the total number of initial scene images acquired. Total number The value is determined based on sensor noise stability test. Experiments show that when the number of samples is greater than or equal to 50, the probability distribution of background pixels tends to be stable. Therefore, it is set to 50 to ensure coverage of dynamic fluctuations in ambient light.
[0025] For example, the control front-end acquisition device continuously acquires 50 initial scene images during the initialization phase. For a specific location with spatial pixel indices of x-coordinate 50 and y-coordinate 100, the corresponding grayscale value is extracted from each of the 50 images. These 50 grayscale values are then arranged in chronological order of acquisition time into an array of 50 elements, generating a temporal feature vector sequence for that location. This sequence records all brightness characteristics of that specific spatial location within the initialization time period, providing raw data support for extracting stable background features.
[0026] Median filtering is performed on the temporal feature vector sequence to extract the background pixel median matrix. Spatial feature stitching is then performed on the background pixel median matrix to generate a base image free of hidden dangers.
[0027] The median filtering operation involves performing a mathematical process on the temporal feature vector sequence corresponding to each spatial pixel index, arranging the values in the sequence by magnitude and extracting the median value. Median filtering effectively removes extreme discrete noise from the sequence. This operation extracts a two-dimensional data array, the background pixel median matrix, composed of the median values obtained from all processed spatial pixel indices, arranged in their original spatial order. The calculation relationship satisfies the following formula: , Among them, letters Represents the median matrix of background pixels in spatial pixel index and The numerical value at the location; the letter This represents the median-taking function. This function calculates the median value from the time-series feature vector sequence. Sort the given numbers in ascending order and extract the value that is in the middle position after sorting. When When the number is odd, take the center value. When the median is even, the average of the two center values is taken. The median function can eliminate sudden high-value interference caused by flying insects and momentary reflections during the initialization phase. Subsequently, the values in the background pixel median matrix are recombined according to their corresponding spatial pixel indices to form a complete image, i.e., spatial feature stitching is performed, generating a reference image representing the normal static environment of the power distribution area, i.e., a no-hazard baseline image. The calculation relationship of the no-hazard baseline image satisfies the following formula: , Among them, letters Represents a baseline image with no potential hazards; letters Represents the total number of pixels representing the width of the initial scene image; letters Represents the total number of height pixels in the initial scene image; symbol This represents the spatial feature stitching operation. After completing the background modeling, the front-end acquisition equipment receives real-time images of the power distribution area captured during daily monitoring, i.e., the on-site environmental image data.
[0028] For example, the edge computing terminal reads the temporal feature vector sequence corresponding to the coordinate point (50, 100). Assume this sequence contains an abnormal brightness value of 255 caused by instantaneous reflection and a normal background brightness value of 120. The 50 values in the sequence are sorted in ascending order using median filtering, and the 25th and 26th values are extracted. If the extracted median value is 121, the abnormally high value 255 is discarded because it is at the end of the sequence. 121 is used as the median value of the background pixel matrix at that coordinate, and this process is repeated for all pixel points. Finally, the medians of all points are spatially stitched together at a resolution of 1920×1080 to generate a baseline image free of hazard signals, which serves as a benchmark for hazard identification.
[0029] The on-site environmental image data and the no-hazard baseline image are subjected to pixel morphological subtraction operation to extract the difference pixel clusters, and the connected component boundary expansion is performed on the difference pixel clusters to generate a morphological difference cropping image. Optionally, the generated morphologically differentiated cropping map includes: The absolute pixel difference is calculated between the on-site environmental image data and the baseline image without hidden dangers to generate a difference feature matrix. The difference feature matrix is then binarized to extract the connected non-zero pixel set and construct the difference pixel cluster. After acquiring the on-site environmental image data and the baseline image without potential hazards, an absolute pixel difference calculation is performed on the on-site environmental image data and the baseline image without potential hazards. This calculation process compares the brightness values of the on-site environmental image data and the baseline image without potential hazards at the same spatial pixel index, and calculates the absolute value of the difference. This calculation generates a two-dimensional numerical array, namely the difference feature matrix, that records the real-time changes in the image. The mathematical expression is: , in, Represents the difference feature matrix in spatial pixel index and The value at that location; Representing on-site environmental image data in spatial pixel index and The value at that location; The spatial pixel index represents the baseline image without hidden dangers. and The values at each point are then analyzed. Subsequently, the difference feature matrix is binarized. The binarization process converts values in the difference feature matrix that exceed a set threshold to a high logic level, and values that are lower than or equal to the threshold to a low logic level. Threshold for judgment. The numerical determination method is as follows: Based on statistical analysis of 100 sets of on-site measured noise data from power substations under different lighting conditions and electromagnetic interference environments, the maximum envelope value of the background fluctuation amplitude is extracted and set to 30 to effectively filter out interference from subtle light fluctuations and electronic noise. Through binarization processing, the sets of connected non-zero pixels that are spatially adjacent and in a logic high-level state are extracted. This set is used as the pixel combination representing the abnormal change area of the image, i.e., the differential pixel cluster. The mathematical relationship of this process is expressed as: , in, Represents the spatial pixel index after binarization. and The logical state value at the location; This represents the threshold for judgment.
[0030] For example, acquire on-site environmental image data and a baseline image without potential hazards. At the spatial pixel index coordinates of 150 (x-coordinate) and 200 (y-coordinate), the brightness value of the on-site environmental image data is 200, while the brightness value of the baseline image without potential hazards is 120. The absolute value of the difference is calculated to be 80, and 80 is used as the value of the difference feature matrix at that coordinate point. Since the value 80 is greater than the judgment threshold of 30, the logical state value of that coordinate point is set to 1. By traversing the entire image area, all pixels with a logical state value of 1 and physically connected are extracted to construct a difference pixel cluster representing the morphological changes of the substation equipment.
[0031] The two-dimensional spatial extreme coordinates of the differential pixel clusters are obtained to construct the minimum bounding rectangle. The vertex coordinates of the minimum bounding rectangle are extended outward to generate a cropping anchor frame. The cropping anchor frame is used to perform matrix slicing processing on the on-site environment image data to generate a morphologically differentiated cropping image.
[0032] Obtain the minimum and maximum position indices of the differential pixel clusters in the horizontal and vertical directions, i.e., the two-dimensional spatial extremum coordinates. Construct a bounding box (minimum bounding rectangle) that completely encloses the boundary of the differential pixel clusters using these two-dimensional spatial extremum coordinates. To preserve the abnormal variation region and its surrounding edge context features, thereby improving the robustness of the model's recognition, extend the vertex coordinates of the minimum bounding rectangle horizontally and vertically by a set pixel displacement to generate the truncated anchor box. (Extend pixel distance) The value was determined through 50 sets of measured data on the blurring degree of edge imaging of power equipment components. To ensure that the equipment outline can be completely included under various focusing conditions, the following parameters were selected: Set to 15 pixels. The mathematical expression for the stretching process is: , , , , in, and These represent the starting and ending coordinates of the anchor frame in the horizontal direction, respectively. and These represent the starting and ending coordinates of the anchor frame in the longitudinal direction, respectively. and The two-dimensional spatial extremum coordinates representing the horizontal direction of the differential pixel cluster; and The two-dimensional spatial extremum coordinates representing the vertical direction of the differential pixel cluster; Represents the distance between extended pixels; and These represent the total number of pixels in width and the total number of pixels in height of the on-site environmental image data, respectively. This represents the function that takes the maximum value. This represents the minimum value function. Matrix slicing is performed on the on-site environmental image data using a cropping anchor frame. The matrix slicing process extracts corresponding local pixel subarrays from the original full HD on-site image based on the coordinate range defined by the cropping anchor frame. After extraction, a local image containing only the abnormally changing target area is generated, i.e., a morphologically differentiated cropped image.
[0033] For example, the two-dimensional spatial extreme coordinates of the differential pixel clusters are obtained. The differential pixel clusters occupy a range from index 140 to 180 in the horizontal direction and from index 190 to 240 in the vertical direction. A minimum bounding rectangle is then constructed. The four vertices of the minimum bounding rectangle are extended outwards by 15 pixels. The calculated horizontal range of the cropped anchor frame is 125 to 195, and the vertical range is 175 to 255. Based on the coordinate range of this cropped anchor frame, the edge analysis terminal performs matrix slicing on the 1920×1080 resolution on-site environmental image data, extracting the local image matrix of this specific area and generating a morphological differential cropping map. This cropping map is then fed into the model for hazard category determination.
[0034] The morphologically differentiated cropped image is input into a preset lightweight recognition model, feature vector mapping is performed, and initial hazard category labels and numerical confidence scores are output. The numerical confidence scores are continuously written into a local state buffer queue to construct a time-sliding window confidence sequence. Optionally, the construction of the time sliding window confidence sequence includes: The lightweight recognition model is invoked to perform tensor convolution and pooling operations on the morphological difference cropping image to generate a high-dimensional feature vector. The high-dimensional feature vector is then subjected to linear transformation and normalization operations, and the maximum probability value is extracted and mapped to a numerical confidence score and an initial hazard category label. After obtaining the morphologically differentiated cropped image, a lightweight recognition model deployed in the memory of the edge computing terminal is invoked to process the image. The lightweight recognition model is a converged deep neural network model trained using a dataset of historical power consumption site hazard images. The lightweight recognition model performs tensor convolution and pooling operations on the morphologically differentiated cropped image. This operation extracts features from the pixel matrix through multiple sliding windows with learnable weights and performs spatial dimensionality reduction and compression using max pooling or average pooling, thereby generating a numerical sequence representing the deep semantic information of the local image, i.e., a high-dimensional feature vector. The mathematical expression is: , Among them, letters Represents a high-dimensional feature vector; letters Represents the generated morphological differentiation cropping diagram; symbol Tensor convolution operation represents the extraction of local pixel features; symbol Pooling operation represents dimensionality reduction and compression. After obtaining the high-dimensional feature vector, the high-dimensional feature vector is input into the fully connected layer of the model to perform linear transformation and normalization operations. The linear transformation maps the high-dimensional feature vector to the logistic regression score corresponding to each predetermined risk category through the weight matrix and bias vector. The normalization operation uses the Softmax function to transform the logistic regression score into a probability distribution vector with a sum of a constant 1. The mathematical relationship of this process satisfies the following formula: , , Among them, letters Represents the logistic regression score vector after linear transformation; letters and These represent the weight matrix and bias vector, respectively, obtained by the lightweight recognition model during the training phase through backpropagation; the letters... This represents the total number of hazard categories set, which is 6, including unauthorized wiring, equipment damage, obstruction by foreign objects, missing meter covers, missing safety signs, and normal status; the letter represents the total number of hazard categories set, which is 6. Represents the th after Softmax mapping Probability value of class; letter and These represent the first and second parts of the logistic regression score vector. Item and the The numerical value of the item; the symbol This represents an exponential function with the natural constant as its base. After normalization, a probability distribution vector containing six probability values is obtained, and the maximum probability value is extracted. This maximum probability value is directly mapped to a numerical confidence score, and the corresponding dimension index in the probability distribution vector is locked. This dimension index is then mapped to the initial hazard category label. The mathematical expression is: , , Among them, letters Represents numerical confidence score; letters The representative includes all The set of probability distributions; symbols This represents the operation of retrieving the largest element in a set; letters Represents the initial hazard category label; symbol This represents a function that retrieves the array index corresponding to the maximum value.
[0035] For example, the edge computing terminal calls a lightweight recognition model to perform tensor convolution and pooling operations on a morphologically differentiated cropped image of size 128×128. After extraction and spatial dimensionality reduction by a 5-layer convolutional network, a high-dimensional feature vector containing 512 dimensions is generated. This 512-dimensional feature vector undergoes linear transformation and normalization. The weight matrix of the fully connected layer is used to map the dimensions to 6 logistic regression scores, which are then converted into 6 probability values using the Softmax function: 0.02, 0.01, 0.85, 0.05, 0.04, and 0.03. The maximum probability value, 0.85, is extracted and mapped to the numerical confidence score for this recognition. Simultaneously, 0.85 is located at the 3rd index position of the probability distribution. Assuming that the preset 3rd index represents "foreign object occlusion," "foreign object occlusion" is mapped to the initial hazard category label.
[0036] Create a local state buffer queue, append the numerical confidence scores to the tail of the local state buffer queue in time sequence, release the head node data when the node is full, and extract the queue values in sequence to construct a time sliding window confidence sequence.
[0037] After generating a single numerical confidence score, a one-dimensional array data structure, namely a local state buffer queue, is created in the dynamic memory space of the edge device, following the "first-in, first-out" principle. The local state buffer queue is used to cache the recognition confidence status of consecutive video frames. The edge terminal appends newly generated numerical confidence scores to the tail of the local state buffer queue in sequence. Simultaneously, the maximum node capacity of the local state buffer queue is set. The value is based on model jitter period tests on 30 sets of on-site monitoring videos during sudden changes in lighting. To ensure that the sliding window can smooth out approximately 1 second of video jitter, combined with the camera's sampling rate of 10 frames per second, the value is... The value is set to 10. Each time new data is written, the queue length is monitored. When a node is full, the head node data is automatically released to maintain a constant queue length. Finally, all values in the current local state buffer queue are extracted sequentially to construct a time-sliding window confidence sequence for evaluating the stability of identification during the current period. The mathematical expression is: , Among them, letters Represents the current system time node Constructed time-sliding window confidence sequence; letters Represents the maximum node capacity of the local state buffer queue; letter Represents the current time point The latest numerical confidence score written; letter , etc. represent the confidence scores of historical moments cached.
[0038] For example, the edge terminal creates a local state buffer queue in memory and sets the maximum node capacity to 10. At the current moment... Previously, the queue had cached nine historical numerical confidence scores: 0.88, 0.87, 0.89, 0.85, 0.82, 0.84, 0.86, 0.85, and 0.83. At time [time value missing]... The newly generated numerical confidence score of 0.85 is appended to the tail of the local state buffer queue. At this point, the queue length reaches 10. At time [time value missing]... A new score of 0.75 is generated and appended to the tail. At this point, the node is full, and the oldest data at the head of the queue is automatically released (0.88). Subsequently, the remaining 10 most recent values in the queue are extracted in sequence to construct an updated time-sliding window confidence sequence: [0.87, 0.89, 0.85, 0.82, 0.84, 0.86, 0.85, 0.83, 0.85, 0.75], for use in the next step of time-series fluctuation characteristic analysis.
[0039] The sample mean and sample variance are extracted using the time sliding window confidence sequence, and the difficult judgment interval is calculated. When the numerical confidence score falls into the difficult judgment interval, the morphological difference cropping image and the initial hidden danger category label are sent to the cloud server. Optionally, sending the morphologically differentiated cropped image and the initial hazard category label to the cloud server includes: The sample mean and sample variance are calculated using the time sliding window confidence sequence, and the sample mean and sample variance are combined to generate a difficult judgment interval. After obtaining the time-sliding window confidence sequence, the sample mean, reflecting the central trend of the data, is calculated using the values in the sequence. This sample mean is obtained by summing all the values in the time-sliding window confidence sequence and dividing by the total sequence length. Then, the sample variance, reflecting the degree of data fluctuation, is calculated. This sample variance is obtained by summing the squares of the differences between each value in the time-sliding window confidence sequence and the sample mean, and dividing by the total sequence length. To define the reasonable fluctuation range of the lightweight recognition model under dynamic environmental interference, the sample mean and sample variance are combined to generate a numerical range with adaptive upper and lower limits, i.e., the difficult-to-determine interval. The mathematical relationship of this process satisfies the following formula: , , , Among them, letters Represents the sample mean; letter The total number of elements in the time-sliding window confidence sequence; this parameter directly corresponds to the maximum node capacity of the local state buffer queue; letters Represents the first in the time sliding window confidence sequence Item numerical value; letter Represents the sample standard deviation, which is the positive square root of the sample variance; the letter... Represents the range of difficult judgments; letters This represents the interval adjustment coefficient. Interval adjustment coefficient This is not a baseless assumption, but rather derived from statistical analysis of recognition deviation data extracted from edge terminals under 1000 real-world tests in low-light conditions such as rain and fog. To ensure that this range encompasses over 85% of anomalous fluctuation samples, [the following data was collected / analyzed]. The scalar value is set to 1.5.
[0040] For example, the edge computing terminal obtains a time-sliding window confidence sequence containing 10 values: 0.81, 0.80, 0.82, 0.79, 0.81, 0.83, 0.80, 0.81, 0.82, and 0.81. Summing these 10 values yields a total of 8.10, which, divided by 10, gives the sample mean of 0.81. Further calculation of the sum of squares of the differences between each value and 0.81, divided by 10, yields the sample variance of 0.00013. Taking its positive square root gives the sample standard deviation of 0.0114. Adding and subtracting the sample mean of 0.81 from 1.5 times the sample standard deviation of 0.0171, the lower limit is calculated to be 0.7929, and the upper limit is 0.8271. This generates a difficult-to-determine interval with floating boundaries, ranging from 0.7929 to 0.8271.
[0041] When the numerical confidence score falls into the difficult judgment range, the communication interface is called to encapsulate the morphological difference cropping image and the initial hidden danger category label into an abnormal reporting data packet and send it to the cloud server.
[0042] The numerical confidence score to be evaluated is extracted and compared with the calculated lower and upper limits of the difficult-to-determine interval. When the numerical confidence score falls within the closed range of the difficult-to-determine interval, it is determined that the current lightweight identification model on the edge side is in an ambiguous and uncertain state regarding the identification confidence of the hazard target. After this condition is triggered, the edge computing terminal calls the wireless communication interface in the underlying hardware module, i.e., the 4G or 5G radio frequency antenna module, to encapsulate the morphological difference cropping image and the initial hazard category label into an anomaly reporting data packet according to the standard network transmission protocol. The encapsulation action packages the image byte stream and string label into the same data payload and sends the anomaly reporting data packet to the cloud server through the communication link to request the cloud's more powerful high-parameter model to perform secondary feature verification.
[0043] For example, the newly generated numerical confidence score at the current time point is 0.80. The edge computing terminal compares 0.80 with the lower limit of the difficulty judgment interval (0.7929) and the upper limit (0.8271) generated in the previous step. Since 0.80 is greater than 0.7929 and less than 0.8271, it meets the trigger condition of falling within the interval. At this time, the edge computing terminal calls the 4G communication interface to encapsulate the 128×128 morphological difference cropped image and the initial hidden danger category label with the content "equipment damage" into a complete TCP network data packet, namely the anomaly reporting data packet, and pushes it to the cloud server for high-precision verification. Figure 2 As shown, an adaptive floating envelope constructed from the mean and variance of sliding window samples is presented. Compared with the traditional fixed rigid threshold, this dynamic range can capture transient ambiguous samples caused by sudden changes in illumination or electromagnetic interference, effectively triggering the cloud verification mechanism, thereby eliminating the phenomenon of batch false alarms and false alarms caused by sudden changes in the environment at the source.
[0044] Optionally, the image encoding sequence of the morphologically differentiated cropped image is extracted, and the verification judgment label is parsed into a structured semantic field. Key-value pair mapping and multimodal data encapsulation are performed on the image encoding sequence and the structured semantic field to construct a multimodal hidden danger evidence archive.
[0045] After the cloud server completes secondary verification and returns a high-precision verification judgment label, the edge computing terminal performs a structured combination of the returned data and local data. First, it extracts the underlying binary byte stream sequence of the morphologically differentiated cropped image, i.e., the image encoding sequence. Simultaneously, it parses the received verification judgment label into data fields containing explicit attribute names, i.e., structured semantic fields, such as parsing text key-value pairs containing hazard classification and severity alarm level. Then, it performs key-value pair mapping and multimodal data encapsulation on the image encoding sequence and structured semantic fields. The key-value pair mapping operation uses memory address pointers, using the structured semantic text as the index key and the underlying image byte sequence as its associated value. Multimodal data encapsulation merges the visual array data bound to memory pointers and the text semantic data into a continuous data block, thereby constructing a multimodal hazard evidence archive that simultaneously contains high-definition local images of the site and the final qualitative conclusion.
[0046] For example, the edge computing terminal receives a verification label from the cloud stating "severely damaged insulator". It extracts the Base64 format string (image encoding sequence) from the underlying layer of the current morphological differential cropping image. The verification label is parsed into two structured semantic fields: a key named "defect classification" with a value of "insulator damage", and a key named "hazard level" with a value of "high". Key-value pair mapping is performed, physically binding these two text fields to the extracted Base64 string at the database level. Through multimodal data encapsulation, the image-text bound structure is written into the local embedded database, ultimately generating a multimodal hazard evidence file that can be directly accessed and verified by management personnel.
[0047] The cloud server performs high-dimensional feature extraction on the morphologically differentiated cropping image to obtain verification and judgment labels, and synchronously receives incremental fine-tuning weight packets through the cloud server. Optionally, the step of synchronously receiving incremental fine-tuning weight packets through the cloud server includes: The cloud server is invoked to perform high-dimensional feature extraction on the morphologically differentiated cropped image and to perform a nearest neighbor index retrieval to generate a verification and judgment label. After receiving an anomaly report data packet, the cloud server invokes an internally deployed high-parameter visual network to perform high-dimensional feature extraction on the morphologically differentiated cropped image. Through multi-scale convolution operations, deep texture and semantic features of the local image are extracted, generating a floating-point one-dimensional array representing the image features—the cloud-based high-dimensional feature vector. Subsequently, a proximity index retrieval is performed in the historical hazard feature map database built in the cloud. This retrieval process measures the degree of feature matching in the vector space—that is, cosine similarity—by calculating the cosine of the spatial angle between the cloud-based high-dimensional feature vector and each historical hazard feature vector in the database. The mathematical expression for this process is: , Among them, letters Represents the high-dimensional feature vector in the cloud and the first Cosine similarity between the feature vectors of historical hidden dangers; letters Represents a high-dimensional feature vector generated in the cloud through feature extraction by a visual network with a large number of parameters; letters The first one represents the pre-stored data in the cloud-based historical hidden danger feature map database. Historical hidden danger feature vector; symbol and These represent the Euclidean norms of the corresponding high-dimensional eigenvectors; (symbols omitted) This represents the dot product operation of vectors. After calculating the cosine similarity of all vector comparisons, the classification label corresponding to the maximum cosine similarity is extracted. This classification label is used as the standard qualitative basis for correcting the front-end edge recognition results, and a verification judgment label is generated through mapping and matching.
[0048] For example, the cloud server invokes a high-parameter visual network to perform high-dimensional feature extraction on the received 128×128 resolution morphological difference cropped image. After deep convolution calculations, a cloud-based high-dimensional feature vector containing 1024 floating-point values is generated. A nearest neighbor index search is performed in a hidden danger feature map database containing 50,000 historical confirmed cases, and the cosine similarity between this 1024-dimensional feature vector and each historical hidden danger feature vector in the database is calculated. Through low-level matrix operations, the cosine similarity between this vector and the 852nd historical hidden danger feature vector in the database reaches a maximum value of 0.96. The cloud server extracts the classification business field "missing cover" associated with the 852nd data item and matches it to generate the verification judgment label for this high-precision review.
[0049] The backpropagation operation is performed using the morphological difference cropping image and the verification judgment label to generate a parameter offset matrix. The parameter offset matrix is then used to perform structural alignment and packing to generate an incremental fine-tuning weight package.
[0050] After obtaining the verification labels, the cloud server calls a cloud-based isomorphic shadow model with the exact same network architecture as the front-end lightweight recognition model. It then performs backpropagation using a morphologically differentiated cropped image and the verification labels. The backpropagation operation first calculates the cross-entropy error between the forward prediction distribution of the cloud-based isomorphic shadow model and the actual verification labels. Then, it calculates the gradient matrix of this cross-entropy error with respect to the weights of each layer of the model using the chain rule. Finally, it multiplies the gradient matrix by the learning rate scalar to generate the weight adjustment amount, i.e., the parameter offset matrix, used to specifically correct model errors. The mathematical expression for this update process is: , Among them, letters Represents the calculated parameter offset matrix; letters Represents the cross-entropy error between the forward prediction distribution and the validation label; the letter This represents the current network weight matrix of the cloud-based isomorphic shadow model; letters The learning rate represents the backpropagation rate; symbol This represents the multiplication operation between a scalar and a matrix. Learning rate. The value of directly determines the stability and convergence of model fine-tuning. Based on 100 fine-tuning stability experiments on long-tailed difficult samples, to avoid "catastrophic forgetting" caused by excessively large single parameter updates, The specific value is rigorously set to 0.002. After generating the parameter offset matrix, structural alignment and packaging are performed using the parameter offset matrix. The structural alignment process rearranges the tensor dimensions of the parameter offset matrix to ensure an absolute match with the memory structure of a specific convolutional or fully connected layer of the lightweight recognition model in the front end. The packaging process performs lossless data compression encoding on the rearranged parameter offset matrix to generate a small file package containing only parameter change information, namely the incremental fine-tuning weight package.
[0051] For example, the cloud server inputs the morphologically differentiated cropped image into the cloud-based isomorphic shadow model for forward inference. Combined with the matching-generated verification label "missing cover," the current cross-entropy error is calculated to be 0.55. Based on this error, a chain-like backward derivative is performed to calculate the gradient matrix of the weights of the last two network layers. Each value in the gradient matrix is multiplied by a set learning rate of 0.002 to calculate the parameter offset matrix used for minor correction of the classification boundary. Subsequently, structural alignment is performed on the parameter offset matrix, strictly reshaping its tensor shape to the 128×64 weight matrix format required by the corresponding edge front-end network. Finally, the ZIP algorithm is used to perform lossless packaging and compression on the aligned tensor array, generating an incremental fine-tuning weight packet of only 150KB. The front-end acquisition system synchronously receives this 150KB incremental fine-tuning weight packet via a 4G wireless network channel, thus eliminating the need to download and fully replace the tens of megabytes of the complete model file, achieving the evolution of edge-side recognition capabilities at extremely low communication bandwidth costs.
[0052] The incremental fine-tuning weight package is loaded into the lightweight recognition model to perform local weight updates, and the verification judgment label is associated with the on-site environmental image data to generate hazard dispatch order data.
[0053] Optionally, the data for generating the hazard dispatch order includes: The incremental fine-tuning weight package is parsed into a state dictionary and loaded into the tensor address of the lightweight recognition model to perform local weight updates; After receiving the incremental fine-tuning weight packet from the cloud, the edge computing terminal uses underlying deserialization and decompression algorithms to restore the packet into a set of key-value pairs (i.e., a state dictionary) containing the names of specific network layers and their corresponding weight values. Then, using a deep learning framework like PyTorch, the terminal retrieves and locates the physical memory block (tensor address) where the lightweight recognition model is currently running. The weight values from the parsed state dictionary are loaded into the corresponding tensor address. By performing matrix addition between the original weights in memory and the parameter offsets in the state dictionary, the self-evolution of the edge-side lightweight recognition model—i.e., local weight update—is completed. The mathematical relationship of this parameter overwrite process is as follows: , Among them, letters This represents the new weight matrix for a specific network layer of the lightweight recognition model after performing local weight updates; letters This represents the original weight matrix residing at the tensor address of the network layer; the letter... This represents the deserialization parameter offset matrix extracted from the state dictionary.
[0054] For example, the edge computing terminal receives a 150KB incremental weight adjustment packet and parses it into a state dictionary using a deserialization algorithm. The key name of this state dictionary is "Conv_Layer_5.weight", and the key value is a 128×64 parameter offset matrix. The address of the tensor corresponding to the lightweight recognition model "Conv_Layer_5" is located through memory pointer addressing. The original 128×64 weight matrix stored at this address is extracted and element-wise added to the parameter offset matrix in the state dictionary. The resulting matrix is written back to the tensor address. This completes the local weight update of the model at the edge without restarting the device, enabling the model to identify complex hidden dangers such as "missing watch covers".
[0055] The on-site environmental image data is mounted as a panoramic context to the multimodal hazard evidence archive. The multimodal hazard evidence archive is used for structured mapping. The association record between the verification and judgment tags and the on-site environmental image data is executed, and business encapsulation is performed to generate hazard dispatch order data.
[0056] After completing the model closed-loop update, the edge computing terminal extracts and records high-definition images of the entire original field of view of the power distribution area, i.e., on-site environmental image data. This on-site environmental image data serves as a reference for providing information on the overall physical environment and spatial orientation when a hazard occurs, i.e., a panoramic context, and is appended to the multimodal hazard evidence archive constructed in the previous steps, i.e., a mounting operation. The multimodal hazard evidence archive, as a data container, undergoes underlying relational database table structure processing, i.e., structured mapping. During the structured mapping process, the verification and judgment tags extracted from the multimodal hazard evidence archive are used as the primary key, and the mounted on-site environmental image data and the original morphological difference cropping image are used as related items, performing a primary and foreign key binding operation at the database level, i.e., as related records. The mathematical logic of this data association and merging process can be abstracted as follows: , Among them, letters Represents the associated records generated in the database; letters Represents the verification and judgment label extracted from the multimodal hazard evidence archive; letters Represents the mounted on-site environmental image data as a panoramic context; letters Represents the original data set in the multimodal hazard evidence archive; symbols This represents the structured mapping function for binding primary and foreign keys in the database. After binding the data relationships, the associated records and multimodal hazard evidence files are uniformly formatted, reassembled, and serialized according to the power grid standard work order communication protocol—that is, business encapsulation. After business encapsulation, a file payload containing complete structured data blocks, including equipment coordinates, hazard type, panoramic view, and local evidence, is generated—that is, hazard dispatch order data.
[0057] For example, the edge computing terminal extracts on-site environmental image data with a resolution of 1920×1080. This large-size image data is used as a panoramic context and appended to the previously generated multimodal hazard evidence archive via the file system. Subsequently, the verification and judgment tag "insulator damage" recorded within the multimodal hazard evidence archive is read. Structured mapping is executed using SQL commands in a relational database, setting "insulator damage" as the primary key of the hazard event and establishing a primary-foreign key binding with the physical storage path of the on-site environmental image data, generating a complete association record. Finally, business encapsulation is performed according to the JSON format protocol, encapsulating the association record, device ID, occurrence time, etc., into hazard dispatch order data containing fields such as "fault_type: insulator damage", "panoramic_image: / data / env_img.jpg", and "cropped_image: / data / crop_img.jpg". This dispatch order data is then directly pushed to the mobile terminal devices of power maintenance personnel, guiding them to the precise location for physical repair. Figure 3 As shown, the initial lightweight model exhibits variance diffusion and confidence decay under complex interference conditions. However, the evolved model after loading the incremental fine-tuning weight package not only shows a significant return of the mean confidence score to a high level, but also effectively converges the dispersion of the data distribution. This verifies the technical superiority of the edge-cloud collaborative closed-loop mechanism in overcoming the environmental generalization problem.
[0058] Based on the same inventive concept, this invention also provides a power inspection hazard identification system based on a lightweight model, such as... Figure 4 As shown, the system includes: The scene baseline acquisition module is used to control the front-end acquisition device to acquire the initial scene image of the power distribution area, perform static background modeling, generate a baseline image without hidden dangers, and use the front-end acquisition device to receive on-site environmental image data. The morphological difference extraction module is used to perform pixel morphological subtraction operation on the on-site environmental image data and the no-hazard benchmark image to extract the difference pixel clusters, and perform connected domain boundary expansion on the difference pixel clusters to generate a morphological difference cropping map. The lightweight inference buffer module is used to input the morphological difference cropping image into the preset lightweight recognition model, perform feature vector mapping, output the initial hazard category label and numerical confidence score, and continuously write the numerical confidence score into the local state buffer queue to construct a time sliding window confidence sequence. The difficult interval determination module is used to extract the sample mean and sample variance using the time sliding window confidence sequence, calculate the difficult determination interval, and when the numerical confidence score falls into the difficult determination interval, send the morphological difference cropping image and the initial hidden danger category label to the cloud server. The cloud-based verification and synchronization module is used to perform high-dimensional feature extraction on the morphologically differentiated cropping image through the cloud server to obtain verification and judgment labels, and to synchronously receive incremental fine-tuning weight packets through the cloud server. The weight update and assignment module is used to load the incremental fine-tuning weight package into the lightweight recognition model to perform local weight updates, and to associate the verification judgment label with the on-site environmental image data to generate hazard assignment order data.
[0059] It should be noted that the functional division and information interaction between the various modules described above are logical, but in terms of physical implementation, they can be integrated on the same software platform or deployed in a distributed manner. The connections between them represent data flow and control flow, aiming to collaboratively achieve the objectives of this invention. The above descriptions are merely exemplary embodiments of this invention and should not be construed as limiting the scope of protection of this invention.
Claims
1. A method for identifying potential electrical hazards based on a lightweight model, characterized in that, The method includes: The system controls the front-end acquisition device to acquire initial scene images of the power distribution area, performs static background modeling, generates a baseline image free of hidden dangers, and uses the front-end acquisition device to receive on-site environmental image data. The on-site environmental image data and the no-hazard baseline image are subjected to pixel morphological subtraction operation to extract the difference pixel clusters, and the connected component boundary expansion is performed on the difference pixel clusters to generate a morphological difference cropping image. The morphologically differentiated cropped image is input into a preset lightweight recognition model, feature vector mapping is performed, and initial hazard category labels and numerical confidence scores are output. The numerical confidence scores are continuously written into a local state buffer queue to construct a time-sliding window confidence sequence. The construction of the time-sliding window confidence sequence includes: calling the lightweight recognition model to perform tensor convolution and pooling operations on the morphologically differentiated cropped image to generate high-dimensional feature vectors; performing linear transformation and normalization operations on the high-dimensional feature vectors; extracting the maximum probability value and mapping it to numerical confidence scores and initial hazard category labels; creating a local state buffer queue; appending the numerical confidence scores to the tail of the local state buffer queue in time sequence; releasing the head node data when the node is full; and extracting queue values in sequence to construct a time-sliding window confidence sequence. The sample mean and sample variance are extracted using the time sliding window confidence sequence, and the difficult judgment interval is calculated. When the numerical confidence score falls into the difficult judgment interval, the morphological difference cropping image and the initial hidden danger category label are sent to the cloud server. The cloud server performs high-dimensional feature extraction on the morphologically differentiated cropping image to obtain verification and judgment labels, and synchronously receives incremental fine-tuning weight packets through the cloud server. The incremental fine-tuning weight package is loaded into the lightweight recognition model to perform local weight updates, and the verification judgment label is associated with the on-site environmental image data to generate hazard dispatch order data.
2. The method for identifying potential electrical hazards based on a lightweight model according to claim 1, characterized in that, The process of receiving on-site environmental image data using the front-end acquisition device includes: The driver front-end acquisition device continuously acquires initial scene images during the initialization phase, extracts temporal pixel values at the same coordinate positions according to spatial pixel index, and constructs a temporal feature vector sequence. Median filtering is performed on the temporal feature vector sequence to extract the background pixel median matrix. Spatial feature stitching is then performed on the background pixel median matrix to generate a base image free of hidden dangers.
3. The method for identifying potential electrical hazards based on a lightweight model according to claim 1, characterized in that, The generated morphologically differentiated cropping map includes: The absolute pixel difference is calculated between the on-site environmental image data and the baseline image without hidden dangers to generate a difference feature matrix. The difference feature matrix is then binarized to extract the connected non-zero pixel set and construct the difference pixel cluster. The two-dimensional spatial extreme coordinates of the differential pixel clusters are obtained to construct the minimum bounding rectangle. The vertex coordinates of the minimum bounding rectangle are extended outward to generate a cropping anchor frame. The cropping anchor frame is used to perform matrix slicing processing on the on-site environment image data to generate a morphologically differentiated cropping image.
4. The method for identifying potential electrical hazards based on a lightweight model according to claim 1, characterized in that, Sending the morphologically differentiated cropping image and the initial hazard category label to the cloud server includes: The sample mean and sample variance are calculated using the time sliding window confidence sequence, and the sample mean and sample variance are combined to generate a difficult judgment interval. When the numerical confidence score falls into the difficult judgment range, the communication interface is called to encapsulate the morphological difference cropping image and the initial hidden danger category label into an abnormal reporting data packet and send it to the cloud server.
5. The method for identifying potential electrical hazards based on a lightweight model according to claim 1, characterized in that, The method further includes: The image encoding sequence of the morphologically differentiated cropped image is extracted, and the verification judgment label is parsed into a structured semantic field. Key-value pair mapping and multimodal data encapsulation are performed on the image encoding sequence and the structured semantic field to construct a multimodal hidden danger evidence file.
6. The method for identifying potential electrical hazards based on a lightweight model according to claim 1, characterized in that, The step of synchronously receiving incremental fine-tuning weight packets through the cloud server includes: The cloud server is invoked to perform high-dimensional feature extraction on the morphologically differentiated cropped image and to perform a nearest neighbor index retrieval to generate a verification and judgment label. The backpropagation operation is performed using the morphological difference cropping image and the verification judgment label to generate a parameter offset matrix. The parameter offset matrix is then used to perform structural alignment and packing to generate an incremental fine-tuning weight package.
7. The method for identifying potential electrical hazards based on a lightweight model according to claim 5, characterized in that, The data for generating the hazard dispatch order includes: The incremental fine-tuning weight package is parsed into a state dictionary and loaded into the tensor address of the lightweight recognition model to perform local weight updates; The on-site environmental image data is mounted as a panoramic context to the multimodal hazard evidence archive. The multimodal hazard evidence archive is used for structured mapping. The association record between the verification and judgment tags and the on-site environmental image data is executed, and business encapsulation is performed to generate hazard dispatch order data.
8. A lightweight model-based electrical inspection hazard identification system, applied to the lightweight model-based electrical inspection hazard identification method as described in any one of claims 1-7, characterized in that, The system includes: The scene baseline acquisition module is used to control the front-end acquisition device to acquire the initial scene image of the power distribution area, perform static background modeling, generate a baseline image without hidden dangers, and use the front-end acquisition device to receive on-site environmental image data. The morphological difference extraction module is used to perform pixel morphological subtraction operation on the on-site environmental image data and the no-hazard benchmark image to extract the difference pixel clusters, and perform connected domain boundary expansion on the difference pixel clusters to generate a morphological difference cropping map. A lightweight inference buffer module is used to input the morphologically differentiated cropped image into a preset lightweight recognition model, perform feature vector mapping, output initial hazard category labels and numerical confidence scores, and continuously write the numerical confidence scores into a local state buffer queue to construct a time-sliding window confidence sequence. The construction of the time-sliding window confidence sequence includes: calling the lightweight recognition model to perform tensor convolution and pooling operations on the morphologically differentiated cropped image to generate high-dimensional feature vectors; performing linear transformation and normalization operations on the high-dimensional feature vectors; extracting the maximum probability value and mapping it to a numerical confidence score and an initial hazard category label; creating a local state buffer queue; appending the numerical confidence scores to the tail of the local state buffer queue in time sequence; releasing the head node data when the node is full; and extracting queue values in sequence to construct a time-sliding window confidence sequence. The difficult interval determination module is used to extract the sample mean and sample variance using the time sliding window confidence sequence, calculate the difficult determination interval, and when the numerical confidence score falls into the difficult determination interval, send the morphological difference cropping image and the initial hidden danger category label to the cloud server. The cloud-based verification and synchronization module is used to perform high-dimensional feature extraction on the morphologically differentiated cropping image through the cloud server to obtain verification and judgment labels, and to synchronously receive incremental fine-tuning weight packets through the cloud server. The weight update and assignment module is used to load the incremental fine-tuning weight package into the lightweight recognition model to perform local weight updates, and to associate the verification judgment label with the on-site environmental image data to generate hazard assignment order data.
Citation Information
Patent Citations
Cloud-edge co-learning power transmission inspection method and system
CN115272981A
Edge detection system for power grid inspection and monitoring
CN116846059A