Substation Intelligent Inspection Method and System Based on Cloud-Edge Collaboration of Sample Library
By extracting the commonality and personality characteristics of the substation inspection data set, conducting strategic sampling, building a balanced data set, and training the global inspection identification model, the problems of sample imbalance and poor model robustness in substation inspections are solved, and the generalization ability and accuracy of the inspection system are improved.
Patent Information
- Application Number
- CN202510278963.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-03-11
AI Technical Summary
In the prior art, substation inspections have problems with unbalanced sample data and low-quality samples. The cloud model is poorly robust and cannot meet the needs of multi-scenario applications, resulting in low generalization capabilities of inspection models and cannot support the construction of high-precision model libraries.
By acquiring multiple inspection data sets, extracting the comprehensive features of the data set, calculating similarity and clustering, obtaining common and personal characteristics, performing strategic sampling, building a balanced data set, and training the global inspection identification model.
It improves the generalization ability of the global model, reduces the overfitting of local models, improves the adaptability and accuracy of the inspection system in multiple scenarios, and reduces the working pressure of operation and maintenance personnel in the site.
Smart Images

Figure CN119784367B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power inspection, and particularly to a substation intelligent inspection method and system based on cloud-edge collaborative construction of a sample library. Background Art
[0002] Currently, substation inspection is in a severe situation where the scale of the station side is growing rapidly, the external environment is becoming increasingly complex, and the strength of inspection personnel is relatively insufficient. The large-scale application of inspection urgently needs to improve the application ability of digital technologies such as artificial intelligence.
[0003] Thanks to the promotion of the Internet of Things, 5G, and artificial intelligence, cloud-edge collaborative technology has developed rapidly in recent years. Combining the powerful data processing ability of cloud computing and the instant response ability of edge computing, it has opened up a new path for intelligent substation inspection.
[0004] However, there are still many difficulties in the global digital technology support of cloud-edge collaborative technology, facing challenges such as cloud-edge resource scheduling, upload and download management of sample libraries and model libraries. Currently, the large-scale sample and model transmission links are not unblocked. At the sample level, there are problems such as excessive data volume, sample imbalance, and low-quality samples in multiple sample sets from the edge sides of multiple different actual inspection environments, which is not conducive to the professional sample collection of substation inspection data and cannot support the construction of a high-quality sample library; at the model level, the global model in the cloud has problems of poor robustness and low generalization ability, cannot meet the actual inspection scenario application requirements of many station sides, is not conducive to the rapid iteration of substation inspection models, and cannot support the construction of a high-precision model library. Summary of the Invention
[0005] Aiming at the problems existing in the above-mentioned prior art, the present invention provides a substation intelligent inspection method and system based on cloud-edge collaborative construction of a sample library, aiming to perform sample data analysis on multiple sample sets from the edge sides of multiple different actual inspection environments, extract common features and individual features, streamline the numerous data from multiple sources, so that when the global model is trained on this sample data set, it can tend to learn the categories with a larger number of samples, and at the same time not completely ignore the individual categories with a smaller number, which can effectively improve the generalization ability of the global model and reduce the overfitting problem of the local model.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] In a first aspect, the present invention provides a substation intelligent inspection method based on cloud-edge collaborative construction of a sample library, including:
[0008] Obtain multiple inspection data sets, and extract the comprehensive features of each inspection data set;
[0009] Calculate the similarity of multiple inspection data sets based on the comprehensive features of the data set, and merge the multiple inspection data sets according to the similarity to form several similar data sets;
[0010] Perform clustering processing within each similar data set, and extract the common features and individual features of the inspection data set based on the clustering results;
[0011] Sample the similar data set based on the common features and individual features to obtain a balanced data set;
[0012] Train a global inspection recognition model based on the balanced data set.
[0013] Preferably, the comprehensive feature of the data set is the mean of the data comprehensive features of all sample data in the data set. The data comprehensive features include the basic features, frequency features, and association features of the sample data. The basic feature is the feature point of the sample data, the frequency feature represents the occurrence frequency information of the feature point of the sample data, and the association feature represents the association information of the feature point of the sample data in the neighborhood.
[0014] Preferably, when the sample data is image data, the frequency feature is the gray frequency of the image feature points, and the association feature is the neighborhood feature of the image feature points;
[0015] When the sample data is text data, the frequency feature is the word frequency of the keywords, and the association feature is the context feature of the keywords;
[0016] When the sample data is audio data, the frequency feature is the frequency of the audio key frames, and the association feature is the spectral difference amplitude of the audio key frames.
[0017] Preferably, the calculating the similarity of multiple inspection data sets based on the comprehensive features of the data set, and merging the multiple inspection data sets according to the similarity to form several similar data sets includes:
[0018] Calculate the mean vector of the data comprehensive feature vectors of all sample data in a single inspection data set as the comprehensive feature vector of the inspection data set;
[0019] Calculate the similarity between the comprehensive feature vectors of multiple inspection data sets based on cosine similarity, merge the sample data in several inspection data sets with similarity less than the first threshold, and add a source label to each sample data to obtain at least one similar data set; the source label represents the inspection data set from which the sample data comes.
[0020] Preferably, the performing clustering processing within each similar data set, and extracting the common features and individual features of the inspection data set based on the clustering results includes:
[0021] Performing K-means clustering on the sample data in each similar data set based on the data comprehensive feature vector of the sample data to obtain K clustering clusters;
[0022] For each clustering cluster, calculate the proportion of the number of each source label in this clustering cluster;
[0023] When the proportion of the number of any one of the source labels m in a clustering cluster is greater than the second threshold, mark this clustering cluster as the personalized clustering cluster of the source label m, and the sample data in multiple personalized clustering clusters of the source label m form the personalized data of the source label m;
[0024] Mark the clustering clusters in which the proportion of the number of each source label is less than the second threshold as common clustering clusters, and the data comprehensive feature of the common clustering cluster is used as the common feature of this similar data set;
[0025] The data comprehensive feature of the clustering cluster is the mean feature of the data comprehensive features of all sample data in this clustering cluster.
[0026] Preferably, the balanced data set includes the common data and personalized data after quantity sampling;
[0027] The steps of sampling to obtain the common data include:
[0028] Count the number of sample data in the common clustering cluster and calculate the variance of the sample data in this common clustering cluster;
[0029] Determine the first sampling number of the common clustering cluster based on the sample data variance and the number of sample data in the common clustering cluster, where the larger the sample data variance and the larger the number of sample data, the larger the first sampling number;
[0030] Calculate the distance between the data comprehensive feature of each sample data in the common clustering cluster and the data comprehensive feature of the common clustering cluster, denoted as the first distance;
[0031] In each common clustering cluster, sample from small to large according to the magnitude of the first distance to obtain the first sampling number of sample data, denoted as common data.
[0032] Preferably, the steps of sampling to obtain the personalized data include:
[0033] Statistically calculate the proportion cm of the total number of samples with the source label m in the total number of samples in all patrol data sets, and denote the product of the proportion cm and the total number of personalized data with the source label m as the second sampling number;
[0034] Randomly sample in the personalized data with the source label m to obtain the second sampling number of sample data, denoted as personalized data.
[0035] In a second aspect, the present invention provides a substation intelligent inspection system constructed based on cloud-edge collaboration of a sample library, including:
[0036] A data acquisition module for acquiring a plurality of inspection data sets;
[0037] A data set feature extraction module for extracting the comprehensive features of each inspection data set;
[0038] A similar data integration module for calculating the similarity of a plurality of inspection data sets based on the comprehensive features of the data sets, and merging the plurality of inspection data sets into several similar data sets according to the similarity;
[0039] A distribution feature extraction module for performing clustering processing within each similar data set, and extracting the common features and individual features of the inspection data sets based on the clustering results;
[0040] A sample sampling module for sampling the similar data sets based on the common features and individual features to obtain a balanced data set;
[0041] A model training module for training a global inspection recognition model based on the balanced data set.
[0042] In a third aspect, the present invention provides an electronic device, including:
[0043] A memory for storing executable instructions;
[0044] A processor, when running the executable instructions stored in the memory, implements a substation intelligent inspection method constructed based on cloud-edge collaboration of a sample library as described above.
[0045] In a fourth aspect, the present invention provides a computer-readable storage medium storing executable instructions, and when the executable instructions are executed by a processor, implements a substation intelligent inspection method constructed based on cloud-edge collaboration of a sample library as described above.
[0046] The substation intelligent inspection method and system constructed based on cloud-edge collaboration of a sample library of the present invention have the following beneficial effects:
[0047] 1. The present invention applies cloud-edge collaboration technology to the substation inspection field, strategically samples and uploads sample data, constructs a high-quality sample database, reduces the processing of repetitive data, reduces the computing burden on the cloud side and the edge side, can realize the training of a high-quality recognition model, and the intelligent model recognition technology can effectively reduce the daily work pressure and inspection burden of in-station operation and maintenance personnel.
[0048] 2. Considering the differences in the distribution of sample data uploaded from different edge sides, the present invention extracts the frequency characteristics and correlation characteristics of numerous sample data respectively to form the overall characteristics of the sample. Through clustering processing within the dataset with similar overall characteristics, the sample data is divided into common characteristics and individual characteristics. By oversampling within the common characteristic data, the global model's capture of universal laws is strengthened, enhancing its adaptability and generalization ability in various scenarios, thereby improving the overall performance of the global substation patrol recognition model. For individual characteristic data, weighted sampling is performed according to the number of samples, which can effectively balance the distribution of individual data and avoid ignoring a small number of samples in the diverse data collected at the corresponding station end when the local model is applied to that station end. Implementing common and individual sampling strategies for multi-source sample datasets can effectively address the problem of uneven data distribution, effectively improve the generalization ability of the global model and meet the special requirements of adapting to local scenarios, and enhance the wide application ability of the patrol system in the global model and the accuracy in specific environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 FIG. is an overall cloud-edge collaboration architecture diagram applying the method of the present invention provided by an embodiment of the present invention;
[0050] Figure 2 FIG. is a schematic diagram of a specific application scenario provided by an embodiment of the present invention;
[0051] Figure 3 FIG. is a schematic flowchart of a substation intelligent patrol method based on cloud-edge collaboration construction of a sample library according to the present invention;
[0052] Figure 4 FIG. is a structural block diagram of a substation intelligent patrol system based on cloud-edge collaboration construction of a sample library according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0053] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0054] Before describing the specific embodiments of this specification in detail, the application scenarios of this specification are introduced as follows:
[0055] The substation intelligent inspection method and system based on cloud-edge collaborative construction of a sample library provided in this specification can be applied to large-scale multi-terminal substation intelligent inspection, and can complete intelligent analysis of substation inspection in three major aspects, including but not limited to abnormal meter readings, abnormal oil seals, discoloration of breather silica gel, closing of pressure plates, opening of pressure plates, crossing the line and breaking in, not wearing a safety helmet, not wearing work clothes, smoking, blurred meter dials, damaged dials, damaged meter housings, cracked insulators, broken insulators, surface contamination, oil stains on the surfaces of oil leakage parts, oil stains on the ground due to oil leakage, metal corrosion, damaged silica gel barrels, abnormal closing of cabinet doors, foreign objects hanging in the air as suspended objects, bird nests, damage to doors, windows, walls and floors, damaged covers, unlocked structure ladders, damaged or fallen signs, detection of unknown abnormal events, calculation of high active power values, identification of bias magnetic conditions, event detection, etc. for the identification of the status of thirty types of substation equipment, personnel safety risks, and equipment defects. As Figure 1 shown, the overall architecture includes the cloud side, the edge side, and the station side. By distributing the substation inspection recognition model uniformly managed by the cloud to each station side, it supports the in-situ analysis of the station side, realizes the large-scale application of substation intelligent inspection, and comprehensively improves the intelligent level of substation inspection.
[0056] For the convenience of understanding, the structures at all levels in the overall architecture are first explained as follows:
[0057] Cloud side: The data aggregation layer and optimization layer of cloud-edge collaboration, used to execute more complex analysis and calculation tasks than the edge side, such as data aggregation, model training, and trend prediction. In this embodiment, the cloud side is specifically used to receive the preprocessed data set uploaded by the edge side, establish a unified high-quality sample library, use the sample library to train the global substation inspection recognition model, and distribute the trained model to the edge side in the form of an image through the cloud-edge collaboration component. The substation inspection recognition model is a general term for multiple recognition models, which can include artificial intelligence models for recognizing various substation states, and can include text recognition models, image recognition models, and audio recognition models. It is also used to update the inspection recognition model based on the historical sample data stored in the cloud and the real-time feedback from the station side, thereby establishing a model library for storing historical inspection recognition models, and visualizing the model library and the sample library to the inspection personnel through the operation platform for regular inspection application quality analysis and algorithm improvement.
[0058] Edge side: The data analysis layer for cloud-edge collaboration, which is used to integrate and process sample data under the condition of limited hardware resources. In this embodiment, it is specifically used to receive the data set from the station side for data preprocessing, identify common features and individual features, and upload the preprocessed data set to the cloud. The data preprocessing may include sample data classification, automatic annotation of sample data, and sample quality analysis. It is also used to receive the global patrol recognition model sent from the cloud, adaptively adjust the model parameters according to the historical application situation of the station side to form a local patrol recognition model, and send the model to the Internet of Things management platform through the HTTP Restful protocol. Through the Internet of Things management platform, the local model is sent down and deployed to the substation intelligent patrol system in each station side.
[0059] Terminal side: The data collection layer and model application layer for cloud-edge collaboration, including terminal devices and patrol systems. Among them, the terminal devices are used to collect data, which can be inspection robots, high-definition dome cameras, voiceprint collection devices, and collect sample data such as text, images, and audio in the application scenario. The patrol system is the actual application of the patrol recognition model, and based on the local patrol recognition model sent from the edge side, it identifies the patrol situation in the actual scenario and completes the patrol task.
[0060] Among them, in the above architecture, the flow of sample data is as follows: The sample data collected by the substation intelligent patrol system of the station side is transmitted to the edge side for feature analysis and collected by the cloud into the sample library; The high-quality samples in the cloud are sent down to the edge side, which can be used to improve the actual model of the station side and realize the co-construction and sharing of the high-quality sample library for the substation specialty.
[0061] The flow of model data is as follows: The cloud sends the substation patrol recognition model to the edge side, and then sends it to the substation intelligent patrol system of the station side through the Internet of Things management platform channel. For the individual recognition needs of each station side, the cloud conducts model customization and development to support the ability to send specific models to specific locations, realizing regional service support; The edge side pulls the model of the station side to obtain the model parameters and uploads the recognition results to the cloud for analysis, and constructs a high-quality model library for the substation specialty to achieve unified management.
[0062] The flow of management data is as follows: The substation intelligent patrol system uploads the model operation monitoring data to the cloud, and displays and analyzes the model application data on the operation platform of the cloud.
[0063] This embodiment provides a specific application scenario, such as Figure 2As shown in the figure, the above architecture is applied to the discoloration recognition of the respirator silica gel. That is, when receiving the task of discoloration recognition of the respirator silica gel, based on the high-definition dome camera or inspection robot installed in the substation, the data of the respirator silica gel is collected into the intelligent inspection system of the substation, and the sample data is intelligently analyzed through the discoloration recognition model of the respirator silica gel, and an abnormal warning is sent in time for the discoloration situation of the respirator silica gel, helping the operation and maintenance personnel to grasp the equipment status in real time. Specifically, data collection is carried out through the intelligent sensing terminal of the high-definition dome camera or robot. Among them, the high-definition dome camera collects video data of the status of the respirator silica gel at a fixed position, and the robot collects pictures and video data of the status of the respirator silica gel through dynamic inspection, and uploads the pictures and video data to the intelligent inspection system of the substation to realize the collection of the status data of the respirator silica gel at the substation end. The edge side aggregates the status data of the respirator silica gel collected at each station end, extracts the sample feature information therein, divides the common feature and individual feature data of the sample data from different sources, and uploads them to the cloud through different channels. The cloud trains the respirator silica gel status recognition model based on the common features and individual feature data and distributes it. The intelligent inspection system of the substation conducts intelligent analysis based on the received discoloration recognition model of the respirator silica gel to realize the local analysis and application at the substation end side.
[0064] To facilitate the understanding of this embodiment, first, a substation intelligent inspection method based on cloud-edge collaborative construction of a sample library disclosed in the embodiments of the present invention will be introduced in detail.
[0065] As Figure 3 shown, the method is applied in the above cloud-edge collaborative architecture, and includes the following steps:
[0066] S1, Obtain multiple inspection data sets collected from multiple station ends.
[0067] The edge side performs the following data processing operations:
[0068] S2, Extract the comprehensive features of each inspection data set;
[0069] S3, Calculate the similarity of multiple inspection data sets based on the comprehensive features of the data sets, and merge the multiple inspection data sets into several similar data sets according to the similarity;
[0070] S4, Perform clustering processing within each similar data set, and extract the common features and individual features of the inspection data set based on the clustering results;
[0071] S5, Based on the common features and individual features, sample the similar data sets to obtain a balanced data set and upload it to the cloud.
[0072] S6, The cloud trains the inspection recognition model based on the balanced data set and distributes it to the edge side and the station end.
[0073] The present invention uses acquisition terminals arranged at the station end to collect multiple sample data. All the sample data collected by each station end constitutes the inspection dataset of that station end. Each station end uploads its respective inspection dataset to the edge end. The edge end first integrates similar datasets from multiple station ends to reduce the amount of data processing. For the clustered similar datasets, it extracts the common features shared by the data from different station ends and the individual features unique to a certain station end. It extracts a part of the data in proportion from the common feature data corresponding to the common features to represent the original numerous data, and extracts a part of the data in proportion to the individual data ratio from the individual feature data corresponding to the individual features. The two parts of the data are merged and uploaded to the cloud for training to obtain a global inspection recognition model. The present invention only transmits key sample data to the cloud, reducing the data transmission pressure. The key sample data not only includes the general laws of numerous sample data but also includes a small number of individual samples, balancing the diverse distribution of sample data. The global model takes into account both the wide application ability and the accuracy in a specific environment.
[0074] The comprehensive feature of the dataset is the mean of the data comprehensive features of all sample data in the dataset. The data comprehensive features include the basic feature, frequency feature, and correlation feature of the sample data. The basic feature is the feature point of the sample data. The frequency feature represents the occurrence frequency information of the feature point of the sample data. The correlation feature represents the correlation information of the feature point of the sample data in the neighborhood.
[0075] It should be noted that an inspection dataset is a set of all sample data collected by a station end, including multiple sample data. Each sample data has a corresponding basic feature. For example, if a sample data is an image, the basic feature is the feature point of the image, such as a corner point. The basic feature, frequency feature, and correlation feature of the sample data constitute all the features of the sample data. The feature point, frequency feature, and correlation feature of each sample data are spliced into the comprehensive feature of the sample data.
[0076] It should be noted that the above comprehensive feature of the dataset is the feature that can represent the central sample data of the dataset. That is, after extracting the comprehensive features of each sample data in the dataset, for example, the comprehensive feature can be represented by a vector, and then the mean vector of the comprehensive feature vectors of all the data in the dataset is calculated as the comprehensive feature of the dataset.
[0077] In the substation inspection task, the sample data in the collected dataset includes one or more of pictures, videos, texts, and audios.
[0078] When the sample data is image data, the frequency feature is the gray frequency of the image feature points, and the correlation feature is the neighborhood feature of the image feature points.
[0079] That is, for an image dataset, feature points of each image need to be extracted. The specific extraction method is not limited in the present invention. For example, corner extraction algorithms, SURF algorithms, etc. in the prior art can be adopted. The frequency feature described in the present invention is the number of occurrences of the gray value of the feature point in the image, and the associated feature is the neighborhood feature of the image feature point. For example, local binary pattern LBP represents the connection between the pixel point and its surrounding pixel points.
[0080] When the sample data is text data, the frequency feature is the word frequency of the keyword, and the associated feature is the context feature of the keyword.
[0081] That is, for a text dataset, keywords of each text data are extracted as basic features. The frequency feature is the word frequency of the keyword. For example, keywords can be extracted by the TF-IDF algorithm, and the word frequency of the keyword can be obtained during the extraction process. The associated feature is the context feature of the keyword. For example, the semantic relationship between words is captured by FastText.
[0082] When the sample data is audio data, the frequency feature is the frequency of the audio key frame, and the associated feature is the spectral difference amplitude of the audio key frame.
[0083] That is, for an audio dataset, first, key frames of each audio segment are extracted as basic features. The frequency feature is the frequency of the audio key frame, such as zero-crossing rate, MFCC, etc. The associated feature is the spectral difference amplitude of the audio key frame, such as spectral difference amplitude, which represents the average change amount of the spectrum between two adjacent frames in an audio segment.
[0084] By extracting the feature manifestation amounts of the basic features of the sample data, namely the frequency feature and the associated feature, the present invention can represent the actual manifestation ability of the basic feature on the sample data. Compared with directly comparing the basic features, the feature manifestation amount can better reflect the inner features of the sample data, and can provide a more comprehensive clustering basis for clustering within subsequent similar datasets. In this way, even if the basic features, i.e., the feature points, are similar, but the performance differences of the feature points in the entire sample data are large, they can be distinguished, making the individual characteristics of single station data more obvious.
[0085] Preferably, step S3 of calculating the similarity of multiple inspection datasets based on the comprehensive feature of the dataset and merging the multiple inspection datasets into several similar datasets according to the similarity includes:
[0086] Calculating the mean vector of the data comprehensive feature vectors of all sample data within a single inspection dataset as the comprehensive feature vector of the inspection dataset;
[0087] Calculate the similarity between the dataset comprehensive feature vectors of multiple inspection datasets based on cosine similarity, merge the sample data in several inspection datasets with similarity less than the first threshold, and add source labels to each sample data to obtain at least one similar dataset; the source label represents the inspection dataset from which the sample data comes.
[0088] It should be noted that in a specific implementation process, there are N station terminals, namely station terminal A, station terminal B,..., station terminal N. Then the above source label indicates which station terminal the sample data comes from. For example, the source label of the sample data of station terminal A can be recorded as source label A.
[0089] The present invention combines and processes datasets with similar overall features. Similar data can be better clustered, reducing the number of clustering times, avoiding excessive single inspection datasets clustering alone, reducing the data processing volume, and reducing the computing burden at the edge. Also, on the premise of similar datasets, it is easier to obtain the individual data unique to a single inspection dataset.
[0090] Preferably, the clustering process performed in each similar dataset and the extraction of the common features and individual features of the inspection dataset based on the clustering results include:
[0091] Perform K-means clustering on the sample data in each similar dataset based on the data comprehensive feature vectors of the sample data to obtain K clustering clusters;
[0092] For each clustering cluster, calculate the proportion of the number of each source label in the clustering cluster;
[0093] When the proportion of the number of any one of the source labels A in a clustering cluster is greater than the second threshold, record this clustering cluster as the individual clustering cluster of the source label A, and the sample data in multiple individual clustering clusters of the source label A form the individual data of the source label A;
[0094] It should be understood that when the proportion of the number of any one of the source labels A in a clustering cluster is greater than the second threshold, it means that almost all the sample data in this clustering cluster are the sample data of the source label A, indicating that this part of the data is unique to station terminal A and is significantly different from the data of other station terminals, that is, the individual data of the source label A (the individual data of the source label A and the individual data of station terminal A refer to the same part of the data). And the individual data of station terminal A may not be similar to each other, that is, these individual data are also clustered into multiple clusters, and these clusters form all the individual data of station terminal A. To highlight the uniqueness of the individual data, the second threshold in this embodiment is selected as 95%.
[0095] Clusters in which the proportion of the number of each of the source tags is less than the second threshold are denoted as common clusters, and the comprehensive data features of the common clusters are used as the common features of the similar data set.
[0096] For example, in a certain cluster of a similar data set composed of sample data from station A, station B, and station C, the sample data of station A accounts for 80%, the sample data of station B accounts for 19%, and the sample data of station C accounts for 1%. Since they are all less than the second threshold of 95%, even though the sample data of station A in this cluster accounts for a relatively large proportion, it is not the individual data unique to station A.
[0097] The comprehensive data feature of the cluster is the mean feature of the comprehensive data features of all sample data in the cluster.
[0098] Preferably, the balanced data set includes common data and individual data after quantity sampling.
[0099] The steps of sampling to obtain common data include:
[0100] Count the number of sample data in the common cluster and calculate the variance of the sample data in the common cluster.
[0101] Determine the first sampling number of the common cluster based on the sample data variance and the number of sample data in the common cluster.
[0102] The larger the sample data variance and the more the number of sample data, it indicates that although the data in this cluster tend to be similar on a large scale and can be grouped into one category, there are still fluctuations and differences among them. Therefore, in order to better achieve data diversity, it is necessary to sample more data in this cluster to improve expressiveness.
[0103] In one implementation, the first sampling number is calculated as follows: calculate the product of the sample data variance of the common cluster and the logarithm of the number of sample data in the common cluster, and multiply the proportion of this product parameter in the overall parameter of the sample of the similar data set by the preset number of samples to be sampled to obtain the first sampling number.
[0104] Calculate the distance between the comprehensive data feature of each sample data in the common cluster and the comprehensive data feature of the common cluster, which is denoted as the first distance.
[0105] In each common cluster, sample from smallest to largest according to the magnitude of the first distance, and obtain the first sampling number of sample data, which is denoted as common data. One common data is obtained in each common cluster, and all the common data obtained in each common cluster in each similar data set are merged and uploaded to the cloud.
[0106] Among them, the first distance from small to large indicates that the data is from more similar to more different, that is, the data that can represent the common cluster as a whole is retained first.
[0107] The present invention multi-samples the common feature data and selects them in proportion according to the variance measurement of the data's ability to represent the whole in the cluster. With the support of the aforementioned basic features, frequency features and associated features, the common data of each part found can represent the local information of the original whole sample and can replace the original data to the greatest extent. It can not only reduce the amount of data uploaded to the cloud, but also strengthen the model's capture of universal laws while streamlining the data, enhance its adaptability and generalization ability in various scenarios, and thus improve the overall performance of the substation patrol recognition model.
[0108] In one embodiment, the step of sampling and acquiring personality data includes:
[0109] Count the ratio cm of the total number of samples with source label m to the total number of samples in the entire patrol data set, and the product of the ratio cm and the total number of personality data with source label m is recorded as the second sampling number;
[0110] Randomly sample the personality data with source label m to obtain the second sampling number of sample data, which are recorded as personality data.
[0111] In a specific implementation process, the total number of samples of all stations is 50,000, and the total number of samples of station A is 500. Among them, in the cluster T1 in the similar data set T composed of the sample data of stations A, station B, and station C, station A has 20 individual data, and in the cluster T2 in the similar data set T, station A has 180 individual data. Then the second sampling number of station A is 500 / 50,000*(20+180)=2.
[0112] It should be noted that personalized data is unique to the site, so random sampling is more in line with the data characteristics.
[0113] The data features collected by different terminals may vary greatly. Each terminal has its own unique features. Not every feature is suitable for uploading to the cloud for global training. If local features determined by the location of the terminal or its other attributes are included in the global training, the computational complexity and classification difficulty of the global classifier will be increased, resulting in an increase in time complexity and the cost of obtaining useful information, while reducing the classification accuracy of the cloud global recognition model. However, a small amount of personalized data can retain the uniqueness of some terminals, making the global model applicable to different terminals.
[0114] The present invention samples personality feature data according to the number of samples, which can effectively balance the data distribution, avoid neglecting minority samples, make full use of the diverse data collected at the edge, and improve the data utilization efficiency.
[0115] In some embodiments, when the recognition accuracy of the local patrol recognition model of the station end p is less than the third threshold, all the personality data in the patrol data set of the station end p are recorded as the correction data set. The edge end uploads the correction data set and the current local model parameters of the station end p to the cloud. The cloud corrects the current local model parameters of the station end p according to the correction data set, and the corrected parameters are sent to the station end p through the edge end.
[0116] It should be understood that although the steps in the flowcharts involved in the above embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps or stages in other steps.
[0117] Based on the same inventive concept, the embodiments of the present application also provide a system for implementing the substation intelligent patrol method based on cloud-edge collaboration for building a sample library. The implementation solutions provided by this system to solve problems are similar to the implementation solutions recorded in the above method. Therefore, the specific limitations in the embodiments of the substation intelligent patrol system based on cloud-edge collaboration for building a sample library provided below can refer to the limitations on this method in the above text, and will not be repeated here.
[0118] As Figure 4 shown, the present invention also provides a substation intelligent patrol system based on cloud-edge collaboration for building a sample library, including:
[0119] A data acquisition module, configured to acquire a plurality of patrol data sets;
[0120] A data set feature extraction module, configured to extract the comprehensive features of each patrol data set;
[0121] A similar data integration module, configured to calculate the similarity of a plurality of patrol data sets based on the comprehensive features of the data sets, and merge the plurality of patrol data sets into several similar data sets according to the similarity;
[0122] A distribution feature extraction module, configured to perform clustering processing within each similar dataset, and extract the common features and individual features of the inspection dataset based on the clustering results;
[0123] A sample sampling module, configured to sample the similar datasets based on the common features and individual features to obtain a balanced dataset;
[0124] A model training module, configured to train a global inspection recognition model based on the balanced dataset.
[0125] It should be noted that: when the substation intelligent inspection system constructed based on the cloud-edge collaboration of the sample library processes substation inspection tasks, only the division of the above functional modules is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the system is divided into different functional modules. Each of the functional modules can be composed of a single execution unit, or can be integrated by two or more execution units into a functional module to implement all the functions of the functional module.
[0126] Those skilled in the art can understand that all or part of the above-mentioned modules can be implemented through software, hardware, and their combinations. The above-mentioned modules can be embedded in the processor of the computer device in hardware form or independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above-mentioned modules.
[0127] An embodiment of the present invention also provides an electronic device, including:
[0128] A memory, configured to store executable instructions;
[0129] A processor, when running the executable instructions stored in the memory, implements a substation intelligent inspection method based on cloud-edge collaboration of the sample library as described above. The processor can be a central processing unit, other general-purpose processors, digital signal processors, application-specific integrated circuits, field programmable gate arrays, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. chips, or a combination of the above types of chips.
[0130] The specific details of the above computer device during substation intelligent inspection can be understood by referring to the relevant descriptions and effects corresponding to the previous method, and will not be elaborated here.
[0131] Another embodiment of the present invention also provides a computer-readable storage medium, storing executable instructions, and when the executable instructions are executed by a processor, implementing a substation intelligent inspection method based on cloud-edge collaboration of the sample library as described above.
[0132] The storage medium may be a variety of computer-readable storage media such as a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk, or an optical disc that can store program codes.
[0133] The present invention is not limited to the above specific embodiments. Those of ordinary skill in the art, starting from the above concepts and without creative labor, can make various transformations, all of which fall within the protection scope of the present invention.
Claims
1. A substation intelligent inspection method based on cloud-edge collaboration construction of a sample library, characterized in that, It includes the following steps: Obtain multiple inspection data sets, and extract the comprehensive data set features of each inspection data set; The comprehensive data set feature is the mean value of the comprehensive data features of all sample data in the data set. The comprehensive data feature includes the basic feature, frequency feature, and association feature of the sample data. The basic feature is the feature point of the sample data. The frequency feature represents the occurrence frequency information of the feature points of the sample data. The association feature represents the association information of the feature points of the sample data in the neighborhood; Calculate the similarity of multiple inspection data sets based on the comprehensive data set features, and merge multiple inspection data sets according to the similarity to form several similar data sets; Perform clustering processing within each similar data set, and extract the common features and individual features of the inspection data set based on the clustering results; Sample the similar data sets based on the common features and individual features to obtain a balanced data set; Train a global inspection recognition model based on the balanced data set.
2. The substation intelligent inspection method based on cloud-edge collaboration construction of a sample library according to claim 1, wherein, When the sample data is image data, the frequency feature is the gray scale frequency of the image feature points, and the association feature is the neighborhood feature of the image feature points; When the sample data is text data, the frequency feature is the word frequency of the keywords, and the association feature is the context feature of the keywords; When the sample data is audio data, the frequency feature is the frequency of the audio key frames, and the association feature is the spectral difference amplitude of the audio key frames.
3. The substation intelligent inspection method based on cloud-edge collaboration construction of a sample library according to claim 2, characterized in that The calculating the similarity of multiple inspection data sets based on the comprehensive data set features and merging multiple inspection data sets according to the similarity to form several similar data sets includes: Calculate the mean vector of the comprehensive data feature vectors of all sample data in a single inspection data set as the comprehensive data set feature vector of the inspection data set; Calculate the similarity between the comprehensive data set feature vectors of multiple inspection data sets based on cosine similarity, merge the sample data in several inspection data sets with a similarity less than the first threshold, and add a source label to each sample data to obtain at least one similar data set; The source label represents the inspection data set from which the sample data comes.
4. The substation intelligent inspection method based on cloud-edge collaboration construction of a sample library according to claim 3, characterized in that, The performing clustering processing within each similar data set and extracting the common features and individual features of the inspection data set based on the clustering results includes: Perform K-means clustering on the sample data in each similar data set based on the comprehensive data feature vectors of the sample data to obtain K clustering clusters; For each clustering cluster, calculate the proportion of the number of each source label in the clustering cluster; When the proportion of the number of any source label m in a clustering cluster is greater than the second threshold, record the clustering cluster as the individual clustering cluster of the source label m, and the sample data in multiple individual clustering clusters of the source label m form the individual data of the source label m; Record the clustering clusters in which the proportion of the number of each source label in the clustering cluster is less than the second threshold as common clustering clusters, and the comprehensive data feature of the common clustering clusters is used as the common feature of the similar data set; The comprehensive data feature of the clustering cluster is the mean feature of the comprehensive data features of all sample data in the clustering cluster.
5. The intelligent inspection method for a substation constructed based on cloud-edge collaboration of a sample library according to claim 4, wherein, The balanced data set includes the common data and individual data after quantity sampling; The steps of sampling to obtain common data include: Count the number of sample data in the common clustering cluster, and calculate the variance of the sample data in the common clustering cluster; Determine the first sampling number of the common clustering cluster based on the sample data variance and the number of sample data in the common clustering cluster, where the larger the sample data variance and the larger the number of sample data, the larger the first sampling number; Calculate the distance between the data comprehensive feature of each sample data in the common clustering cluster and the data comprehensive feature of the common clustering cluster, denoted as the first distance; In each common clustering cluster, sample from small to large according to the magnitude of the first distance, and obtain the first sampling number of sample data, denoted as common data.
6. The substation intelligent inspection method based on cloud-edge collaboration construction of a sample library according to claim 5, wherein The steps of sampling to obtain personalized data include: Statistically calculate the proportion cm of the total number of samples with source label m in the total number of samples in all inspection data sets. The product of the proportion cm and the total number of personalized data with source label m is denoted as the second sampling number; Randomly sample from the personalized data with source label m to obtain the second sampling number of sample data, denoted as personalized data.
7. A substation intelligent inspection system constructed based on cloud-edge collaboration of a sample library, characterized in that, Including: A data acquisition module for acquiring multiple inspection data sets; A data set feature extraction module for extracting the data set comprehensive feature of each inspection data set; Among them, the data set comprehensive feature is the mean value of the data comprehensive features of all sample data in the data set. The data comprehensive feature includes the basic feature, frequency feature, and association feature of the sample data. The basic feature is the feature point of the sample data. The frequency feature characterizes the occurrence frequency information of the feature point of the sample data. The association feature characterizes the association information of the feature point of the sample data in the neighborhood; A similar data integration module for calculating the similarity of multiple inspection data sets based on the data set comprehensive feature, and merging the multiple inspection data sets into several similar data sets according to the similarity; A distribution feature extraction module for performing clustering processing in each similar data set, and extracting the common features and personalized features of the inspection data set based on the clustering result; A sample sampling module for sampling the similar data set based on the common features and personalized features to obtain a balanced data set; A model training module for training a global inspection recognition model based on the balanced data set.
8. An electronic device, characterized in that, The electronic device includes: A memory for storing executable instructions; A processor for implementing a substation intelligent inspection method based on cloud-edge collaboration construction of a sample library according to any one of claims 1 to 6 when running the executable instructions stored in the memory.
9. A computer-readable storage medium storing executable instructions, characterized in that, The executable instructions, when executed by the processor, implement a substation intelligent inspection method based on cloud-edge collaboration construction of a sample library according to any one of claims 1 to 6.
Citation Information
Patent Citations
Activity recognition model and system considering universality and individuation
CN111985650A
Video processing method, system and device based on end-cloud collaboration and storage medium
CN117789075A