Power grid operation difficulty video mining and behavior recognition method, system and device and medium

By collecting power grid operation video data, generating pseudotext and extracting fine-grained features, calculating a comprehensive difficulty score, filtering difficult videos, and optimizing recognition model parameters, the problems of equipment changes and sample scarcity in power grid operation video analysis are solved, thereby improving the safety supervision and management efficiency of the power industry.

CN120808248APending Publication Date: 2025-10-17GUIZHOU POWER GRID CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510621806.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Traditional methods are difficult to meet the safety supervision requirements of the power industry when faced with dynamically changing equipment and standards, scarce high-risk low-frequency event samples, and accurate mining of difficult samples. In particular, in the analysis of power grid operation videos, there are problems such as difficulty in model updating, scarcity of samples, and inaccurate mining of difficult samples.

Method used

Video data of power grid operations is collected, pseudo-text is generated through a multimodal model, fine-grained features are extracted and a comprehensive difficulty score is calculated, difficult videos are screened, and the recognition model is tested. The model parameters are optimized using a weighted cross-entropy loss function to improve the model's adaptability and robustness to complex situations.

Benefits of technology

It enables accurate analysis of power grid operation videos, improves the reliability and efficiency of safety supervision and management in the power industry, and enhances the model's adaptability and recognition accuracy in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808248A_ABST
    Figure CN120808248A_ABST
Patent Text Reader

Abstract

The invention discloses a power grid operation difficulty video mining and behavior recognition method, system and device and a medium. The method comprises the steps that video data are collected, and behavior category labeling is conducted on the video data at different times; generating a pseudo text of an operation scene from the labeled video data through a multi-modal model; extracting fine-grained features of the pseudo text, obtaining a comprehensive difficulty score of the video, and comparing the comprehensive difficulty score with a preset threshold to obtain a difficult video; and inputting the difficult video into the recognition model to obtain recognition results of different time behavior categories in the power grid operation scene. According to the method, the pseudo text containing rich information is generated by using the multi-modal model, and the characteristics of the operation scene can be comprehensively reflected. According to the method, pseudo-text fine-grained features are extracted, a comprehensive difficulty score is calculated, and difficult videos are screened, so that the problem of a traditional method in power grid operation video analysis is solved, strict safety supervision requirements of the power industry are met, and power grid operation safety and management efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of power grid operation identification, and in particular to a power grid operation difficulty video mining and behavior identification method, system, device and medium. BACKGROUND

[0002] With the progress of artificial intelligence and computer vision technology, video analysis is widely used in industrial scenarios such as power grid outage operation, for monitoring the operation of personnel operation specification and process compliance, and evaluating the influence of environmental factors on operation safety. However, complex operation environment brings serious challenges, mainly in three aspects: equipment and specifications change constantly, requiring continuous model updates; high-risk low-frequency event samples are scarce; it is difficult to accurately mine difficult samples. These problems are interrelated, making it difficult for traditional methods to meet the strict safety supervision requirements of the power industry.

[0003] In terms of model updating, the particularity of power operation requires the learning system to not only quickly adapt to new knowledge, such as new insulation tools and safety procedures, but also to remember historical operation specifications. Traditional incremental learning techniques cannot do this. When the model adjusts to new operations, it is easy to forget traditional operation specifications. Even with strategies such as elastic weights solidification, in complex environments, the model still cannot accurately determine the cause of feature changes, leading to misjudgment of risk patterns. Small sample learning techniques also face difficulties. In the high-risk operation scenario of power, samples are scarce and conflict with operation specifications. Traditional meta-learning methods do not conform to the actual risk distribution, data augmentation may generate samples that violate physical properties, and generative adversarial networks lack understanding of power safety procedures, which may generate samples that violate regulations. These all affect the reliability of the model in real scenarios. Traditional methods of difficult sample mining also have limitations. Manual annotation relies on experts, is low in efficiency and highly subjective; the screening mechanism based on loss function cannot distinguish the root cause of sample difficulty; active learning methods have the problem of fragmentation of multi-modal information and lack of domain knowledge, making it difficult to deeply understand complex risk patterns and evaluate sample value from a fine-grained perspective. SUMMARY

[0004] In view of the above existing problems, the present application is proposed.

[0005] Therefore, the present application provides a power grid operation difficulty video mining and behavior identification method to solve the problems that traditional methods have limitations in dealing with dynamically changing equipment and specifications, high-risk low-frequency event sample scarcity, and accurate mining of difficult samples, making it difficult to meet the strict safety supervision requirements of the power industry.

[0006] To solve the above technical problems, the present application provides the following technical solutions:

[0007] In a first aspect, the present application provides a power grid operation difficulty video mining and behavior identification method, comprising:

[0008] collecting video data in a power grid operation scene, and labeling different time behavior categories of the video data;

[0009] performing inference learning on the labeled video data by using a multi-modal model to generate pseudo-text of the operation scene;

[0010] extracting fine-grained features of the pseudo-text, obtaining a comprehensive difficulty score of the video according to the fine-grained features, comparing the comprehensive difficulty score with a preset threshold, and obtaining a difficult video;

[0011] inputting the difficult video into a recognition model, testing by using the recognition model, and obtaining an identification result of different time behavior categories in the power grid operation scene.

[0012] As a preferred scheme of the power grid operation difficult video mining and behavior recognition method, wherein: performing inference learning on the labeled video data by using a multi-modal model to generate pseudo-text of the operation scene, comprising:

[0013] setting a power grid operation slogan based on the video data;

[0014] obtaining key frames of the video data, performing visual-text inference on the key frames by using a multi-modal large model based on the power grid operation slogan, and generating pseudo-text including scene information, target information, action information and time information.

[0015] The beneficial effects of the preferred technical scheme are: by setting the power grid operation slogan, the multi-modal large model can be focused on the power grid operation scene, so that the pseudo-text generated by the model is more targeted. By obtaining the key frames of the video data and performing visual-text inference based on the slogan, the effective information in the video can be efficiently utilized, and irrelevant or redundant content generated by the model can be avoided.

[0016] As a preferred scheme of the power grid operation difficult video mining and behavior recognition method, wherein: extracting fine-grained features of the pseudo-text, and obtaining a comprehensive difficulty score of the video according to the fine-grained features, comprising:

[0017] performing structured analysis on the pseudo-text, setting a feature category, and extracting fine-grained features;

[0018] performing numerical conversion on the fine-grained features by using a mapping function to obtain a difficulty score of each feature;

[0019] performing weighted summation on the difficulty scores of the features to obtain a comprehensive difficulty score of each video.

[0020] The beneficial effects of the preferred technical solution are: the pseudo-text is subjected to structured analysis and setting of characteristic categories to extract fine-grained features, which can deeply analyze the pseudo-text information from multiple dimensions, comprehensively capture factors affecting the difficulty level of the video sample, and convert the fine-grained feature values into difficulty scores by using a mapping function, thereby realizing quantitative expression of the difficulty level of different features and making the evaluation more objective and accurate.

[0021] As a preferred scheme of the power grid operation difficult video mining and behavior recognition method, the identification model is used for testing to obtain the identification results of different time behavior categories in the power grid operation scene, including:

[0022] The difficult video is divided into a training set and a test set, the identification model is trained using the training set, the test set is input into the trained identification model, and the identification results of different time behavior categories in the power grid operation scene are obtained.

[0023] The identification model is optimized by using a weighted cross-entropy loss function, and the identification model parameters are updated.

[0024] The beneficial effects of the preferred technical solution are: the difficult video is divided into a training set and a test set, which can effectively verify the performance and accuracy of the model. At the same time, the identification model is optimized by using a weighted cross-entropy loss function and the model parameters are updated, which can adjust the training weight of the model according to the difficulty level of the sample, make the model pay more attention to difficult samples, and enhance the adaptability and robustness of the model to complex situations.

[0025] As a preferred scheme of the power grid operation difficult video mining and behavior recognition method, the pseudo-text is subjected to structured analysis, and the characteristic categories are set to extract fine-grained features, including:

[0026] The set characteristic categories include target quantity, weather condition, time, and light condition.

[0027] The difficulty scores of the target quantity, weather condition, time, and light condition are calculated respectively, and the comprehensive difficulty score is obtained by combining the weighted parameters.

[0028] As a preferred scheme of the power grid operation difficult video mining and behavior recognition method, the calculation formula of the comprehensive difficulty score is:

[0029]

[0030] Wherein, alpha, beta, lambda is a weighted parameter, satisfying F1F2F3 represents the joint influence of the three, F1 is the target quantity difficulty score, F2 is the weather condition difficulty score, and F3 is the time and light condition difficulty score.

[0031] As a preferred scheme of the power grid operation difficulty video mining and behavior recognition method, wherein: the comprehensive difficulty score is compared with a preset threshold to obtain a difficult video, comprising:

[0032] According to the calculated comprehensive difficulty score, it is judged whether the video is a difficult sample, which is expressed as:

[0033]

[0034] Wherein, M is the number of video samples, is the comprehensive difficulty score of the video sample V j , and τ is the set comprehensive difficulty score threshold, is the jth difficult sample, S har is a difficult video set.

[0035] Secondly, the application provides a power grid operation difficulty video mining and behavior recognition system, comprising: a data acquisition module for collecting video data in a power grid operation scene, and labeling different time behavior categories of the video data;

[0036] A data processing module is configured to infer and learn the labeled video data through a multi-modal model to generate pseudo-text of the operation scene;

[0037] A feature acquisition module is configured to extract fine-grained features of the pseudo-text, acquire a comprehensive difficulty score of the video according to the fine-grained features, compare the comprehensive difficulty score with a preset threshold, and obtain a difficult video;

[0038] An identification module is configured to input the difficult video into an identification model, test the identification model, and obtain an identification result of different time behavior categories in the power grid operation scene.

[0039] Thirdly, the application provides an electronic device, comprising:

[0040] a memory and a processor;

[0041] The memory is configured to store computer executable instructions, and the processor is configured to execute the computer executable instructions, which implement the steps of the power grid operation difficulty video mining and behavior recognition method.

[0042] Fourthly, the application provides a computer readable storage medium storing computer executable instructions, which implement the steps of the power grid operation difficulty video mining and behavior recognition method when executed by a processor.

[0043] Compared with the prior art, the present application has the beneficial effects that: the present application collects and labels various power grid operation video data, generates pseudo-text containing rich information by using a multi-modal model, and can comprehensively reflect the characteristics of the operation scene. The pseudo-text fine-grained features are extracted and the comprehensive difficulty score is calculated to screen difficult videos, solving the problems faced by traditional methods in power grid operation video analysis, meeting the strict safety supervision requirements of the power industry, and improving the safety and management efficiency of power grid operation. BRIEF DESCRIPTION OF DRAWINGS

[0044] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed to be used in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0045] Figure 1 The overall flowchart of the power grid operation difficult video mining and behavior recognition method described in an embodiment of the present application.

[0046] Figure 2 The recognition model schematic diagram of the power grid operation difficult video mining and behavior recognition method described in an embodiment of the present application.

[0047] Figure 3 The test result schematic diagram of the power grid operation difficult video mining and behavior recognition method described in an embodiment of the present application. DETAILED DESCRIPTION

[0048] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the specific embodiments of the present application will be described in detail below with reference to the drawings in the specification. Obviously, the described embodiments are part of the embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should be within the scope of protection of the present application.

[0049] Embodiment 1, refer to Figure 1 An embodiment of the present application provides a power grid operation difficult video mining and behavior recognition method, comprising:

[0050] S100: collect video data in the power grid operation scene, and label the video data in different time behavior categories;

[0051] S102: infer and learn the labeled video data by a multi-modal model to generate pseudo-text of the operation scene;

[0052] S104: Extract the fine-grained features of the pseudo text, obtain the comprehensive difficulty score of the video according to the fine-grained features, compare the comprehensive difficulty score with a preset threshold, and obtain a difficult video;

[0053] S106: input the difficult video into the recognition model, test using the recognition model, and obtain the recognition result of different time behavior categories in the power grid operation scene.

[0054] It should be noted that the technical solutions of collecting power grid operation video data and labeling, generating pseudo text using a multi-modal model, extracting pseudo text features to obtain a comprehensive difficulty score to screen difficult videos and inputting the difficult videos into a recognition model, and testing using the model are designed because the traditional method faces problems such as dynamic changes in equipment and specifications, scarcity of high-risk low-frequency samples, and inaccurate difficult sample mining when dealing with power grid operation video analysis, and it is difficult to meet the safety supervision requirements. The present application can provide a basis for subsequent analysis by collecting and labeling video data; the multi-modal model can fully mine video information by generating pseudo text; the fine-grained features of the pseudo text can be extracted and the comprehensive difficulty score can be calculated to accurately screen difficult videos, providing more valuable data for model training and enhancing the adaptability of the model to complex situations; the recognition model can be used for testing to obtain accurate behavior category recognition results. The present application can improve the accuracy of power grid operation behavior recognition, improve the reliability and adaptability in complex environments, effectively meet the strict safety supervision requirements of the power industry, and effectively improve the safety and management efficiency of power grid operations.

[0055] Embodiment 2, refer to Figures 1-3 For an embodiment of the present application, based on the above embodiment, a power grid operation difficult video mining and behavior recognition method is provided.

[0056] In the embodiment of the present application, the video data in the power grid operation scene collected in step S100 is labeled for different time behavior categories, which specifically includes:

[0057] In the power grid outage operation scene, data collection is performed by combining field video recording with real-time monitoring systems. The video recording uses 1080P high-definition resolution, and the average duration of each video is about 10 minutes to ensure recording the entire operation process. The video data collection covers various complex environments, including changes in lighting, time, location, and weather; for example, data is recorded at different time periods within a day to capture scenes from bright sunlight to weak light environments; data is collected at different operation sites to ensure diversity of locations; and various weather conditions such as sunny, cloudy, and rainy days are included to simulate the variability of actual working conditions. All collected video data is selected by professional personnel for key frames and labeled for categories. The category labeling is for labeling the time sequence behavior of the operation personnel in the video, such as climbing and hanging ground wires.

[0058] In an optional embodiment, the video data is preprocessed, which includes denoising, data enhancement, and normalization. Denoising removes interference such as salt and pepper noise and Gaussian noise in the video using algorithms such as Gaussian filtering and median filtering, improves video clarity, and avoids noise affecting subsequent behavior recognition. Data enhancement expands the data volume by rotating, cropping, adjusting brightness, and other operations on the video, alleviates the problem of data scarcity, enhances the model generalization ability, and makes it adapt to more complex situations. Normalization processes the brightness, contrast, and other parameters of the video to a unified standard, maps the data to a specific interval, and reduces data deviation caused by device differences and different environments.

[0059] In an embodiment of the present application, the labeled video data in step S102 is learned by a multi-modal model to generate pseudo text of the work scene, which further includes sub-steps A1-A2:

[0060] A1: Based on the video data, set the power grid work slogan;

[0061] A2: Obtain the key frames of the video data, and based on the power grid work slogan, use a multi-modal large model to perform visual-text inference on the key frames to generate pseudo text including scene information, target information, action information, and time information.

[0062] In an embodiment of the present application, setting the power grid work slogan includes:

[0063] 1) Scene description prompt: guide the large model to describe the overall environmental information of the video;

[0064] 2) Behavior description prompt: guide the large model to identify the operation details of the work personnel;

[0065] 3) Target description prompt: guide the large model to focus on the work objects (such as power lines and tools) in the video;

[0066] 4) Time / weather prompt: analyze external environmental information (such as day / night and sunny / rainy).

[0067] In an embodiment of the present application, visual-text inference is performed on the video frames to generate video pseudo text T, which is represented by the formula:

[0068]

[0069] where V = {v1, v2, …, v n} is the sequence of video key frames, n is the number of frames, P is the power grid work slogan, represents a multi-modal large model.

[0070] In an alternative embodiment, the multi-modal large model can be GPT-4V, BLIP-2, Video-LLaMA, etc.

[0071] For example, if GPT-4V is selected as the multi-modal large model, the process of obtaining pseudo text includes: according to the characteristics of video data and analysis requirements, selecting appropriate prompt combinations from the above scene description prompt, behavior description prompt, target description prompt and time / weather prompt to form the power grid operation slogan P. For example, for a video showing a complex outdoor power grid maintenance operation, the scene description prompt "Please describe the overall environment in the video, including lighting, weather and work site conditions" can be combined with the behavior description prompt "Please describe the operation steps of the work personnel on various tools" to form the slogan P.

[0072] The key frame sequence V = {v_1, v_2, …, v_n} of the video data is obtained through video processing technology, which can represent important actions and scene changes in the video. Then, the power grid operation slogan P and the key frame sequence V are input into GPT-4V. GPT-4V, based on its powerful multi-modal fusion capability and pre-training knowledge, conducts in-depth visual-text reasoning on the key frames. It identifies the image content in the key frames, combines the guidance of the slogan, and extracts scene information such as "the video was shot in an outdoor open space with sufficient sunlight and surrounding power transmission towers"; target information such as "there are 3 work personnel in the video, carrying gloves, pliers and other tools, and the work object is a damaged power line"; action information such as "one work personnel is using pliers to disassemble the damaged power line, and another is assisting in transferring tools"; and time / weather information such as "the video was shot in the morning with clear weather". Finally, these information is integrated to generate pseudo text T containing the above types of information, providing a rich and detailed data basis for subsequent fine-grained feature extraction and analysis.

[0073] In an alternative embodiment, the scene information includes descriptions of weather, lighting, shooting environment, etc.; the target information includes work personnel, tools, power grid equipment, etc.; the action information includes the operation process of the work personnel; and the time information includes time-related elements such as time, weather, lighting, etc.

[0074] It should be explained that the present application can make the large model focus on the key information of the power grid operation scene by setting the power grid operation slogan and using different types of prompt language such as scene, behavior, target, time and weather to guide the large model, avoid generating irrelevant or redundant content, and enhance the pertinence of information acquisition. The key frames of the video are acquired, and visual-text reasoning is performed based on the slogan by using a multi-modal large model, which takes advantage of the multi-modal large model to fuse visual and text information, generates pseudo text covering scene, target, action and time information, and obtains the characteristics of the operation scene, which provides rich and structured data support for subsequent extraction of fine-grained features, evaluation of video difficulty and training of behavior recognition model.

[0075] In the embodiment of the present application, the fine-grained features of the pseudo text are extracted in step S104, the comprehensive difficulty score of the video is obtained according to the fine-grained features, the comprehensive difficulty score is compared with the preset threshold, and the difficult video is obtained, which further includes sub-steps B1-B3:

[0076] B1: performing structured analysis on the pseudo text, and setting feature categories to extract fine-grained features;

[0077] B2: converting the fine-grained features into numerical values by using a mapping function to obtain the difficulty score of each feature;

[0078] B3: weighted summing the difficulty scores of each feature to obtain the comprehensive difficulty score of each video.

[0079] In the embodiment of the present application, the set feature categories include target quantity, weather condition, time and illumination condition; the difficulty scores of the target quantity, weather condition, time and illumination condition are calculated respectively, and the comprehensive difficulty score is obtained in combination with the weighting parameter.

[0080] For example, let F={F1, F2, F3} be the difficulty scores of the fine-grained features, and the calculation method of the target quantity F1 is as follows:

[0081]

[0082] Wherein, C is the maximum impact score, N is the target quantity parsed by the pseudo text, k controls the growth rate, and N0 is the impact inflection point.

[0083] The calculation method of the weather condition F2 is as follows:

[0084] F2=C x μ1(T),

[0085] Wherein, μ1(T)=0 represents sunny, μ1(T)=0.3 represents overcast, μ1(T)=0.7 represents foggy, and μ1(T)=1 represents rain and snow.

[0086] Based on the Gamma transformation, the calculation of the time and light condition F3 is as follows:

[0087]

[0088] wherein L is the light level, L=100 in the daytime, L=50 in the evening, L=20 at night, and L max =100 is the maximum light, and γ is a control factor.

[0089] In the embodiment of the present application, the calculation formula of the comprehensive difficulty score is as follows:

[0090]

[0091] wherein α, β, λ is a weighted parameter, and satisfies F1F2F3 represents the joint influence of the three, F1 is the target quantity difficulty score, F2 is the weather condition difficulty score, and F3 is the time and light condition difficulty score.

[0092] In the embodiment of the present application, α=0.3, β=0.3, λ=0.1.

[0093] In the embodiment of the present application, according to the calculated comprehensive difficulty score, it is determined whether the video is a difficult sample, which is expressed as:

[0094]

[0095] wherein M is the number of video samples, is the comprehensive difficulty score of the video sample V j , τ is the set comprehensive difficulty score threshold, is the jth difficult sample, and S hard is the difficult video set.

[0096] In an optional embodiment, the acquisition of the difficult video can be a clustering analysis method. Specifically, all video samples are first subjected to data standardization processing based on the comprehensive difficulty score, so as to be mapped to the same dimension and eliminate the influence of different characteristic scale differences. Then, the K-Means clustering algorithm is used to cluster the video samples. Through multiple tests, the appropriate clustering number K is determined, so that the sample difference between different clusters is as large as possible, and the sample difference within the same cluster is as small as possible. In the clustering process, the algorithm will continuously adjust the cluster center until the convergence condition is reached. The video samples contained in the cluster with a higher comprehensive difficulty score are regarded as difficult videos.

[0097] In another optional embodiment, difficult videos can also be obtained by utilizing the anomaly detection method in deep learning. Specifically, an autoencoder model based on a convolutional neural network is constructed. The comprehensive difficulty score and related feature data of all video samples are used as input, and the autoencoder is trained so that it can learn the feature distribution of normal video samples. The autoencoder consists of an encoder and a decoder. The encoder compresses the input data into a low-dimensional representation, and the decoder reconstructs it back to the original data. During the training process, the model minimizes the reconstruction error. After the training is completed, for new video samples, the reconstruction error after passing through the autoencoder is calculated. If the reconstruction error is greater than the preset anomaly threshold, the video is determined to be a difficult video.

[0098] It should be noted that the present invention performs structured analysis on pseudo-text and sets specific feature categories to extract fine-grained features, deeply analyzing video information from multiple key dimensions such as target quantity, weather conditions, time and lighting conditions, ensuring that important factors affecting the difficulty of the video are not missed. Fine-grained features are digitized using a mapping function, and features that are difficult to measure directly are converted into quantitative difficulty scores, so that the difficulty levels of different features are comparable, achieving an objective and accurate evaluation. Finally, a comprehensive difficulty score is obtained through weighted summation, which comprehensively considers the relative importance of different features in judging the difficulty level of the video, and cross terms are added to the comprehensive difficulty score formula to reflect the joint influence of multiple factors, reflecting the complexity of the video more comprehensively. It can accurately select truly challenging and difficult videos from a large number of video samples, provide data for subsequent model training, and improve the model's adaptability and recognition accuracy to complex power grid operation scenarios.

[0099] In the embodiment of the present invention, step S106 inputs the difficult video into the recognition model, and uses the recognition model to perform testing to obtain recognition results of different time behavior categories in the power grid operation scene, which also includes sub-steps C1-C2:

[0100] C1: Divide the difficult videos into training and test sets, use the training set to train the recognition model, and input the test set into the trained recognition model to obtain recognition results for different temporal behavior categories in power grid operation scenarios;

[0101] C2: Use the weighted cross entropy loss function to optimize the recognition model and update the recognition model parameters.

[0102] In an embodiment of the present invention, the weighted cross entropy loss function is expressed as:

[0103]

[0104] Among them, by scoring the comprehensive difficulty Normalize to get ω jis the weight factor of the difficult sample, M is the number of samples, y j is the true value label of the jth sample, is the prediction result of the jth sample by the model.

[0105] In an embodiment of the present application, the update of the identification model parameter is represented as:

[0106]

[0107] wherein, is the optimized model parameter of the model, θ is the model parameter of the previous period of the model, and γ is the learning rate.

[0108] In an embodiment of the present application, the threshold τ = 60, and a large language model is used as the backbone network of the specific time sequence behavior identification model. As shown in Figure 2 , the large language model includes a visual encoder for processing video frames and a text encoder for processing text descriptions. In addition to the classification task, the visual-text alignment task is also performed by using the visual-text contrast learning and visual-text matching learning provided by the visual language model. The SGD optimizer is used as the training optimizer, the initial value of the learning rate γ is set to 1e-3, the batch size M is set to 128, and the learning rate is reduced to 0.1 times of the original value after the 40th and 70th training periods. The total training period is 80, and the test result is shown in Figure 3 .

[0109] In an alternative embodiment, the identification model can be a visual Transformer model based on a convolutional neural network. The identification model first uses the convolutional layer of the convolutional neural network to extract features from the key frames of the difficult video. The convolution kernel in the convolutional layer can effectively capture local features in the video image, such as the action posture of the worker and the shape of the tool. After multiple convolution and pooling operations, a feature map with a certain degree of abstraction is obtained. Then, these feature maps are input into the visual Transformer module. The visual Transformer can simultaneously focus on the features of different regions through the multi-head attention mechanism, better model the global information in the video, and thus learn the correlation and feature representation between different time behavior categories. In the training process, the difficult video is divided into a training set and a test set. The training set is used to train the model, so that it continuously adjusts the parameters to fit the data. The test set is used to evaluate the performance of the model. At the same time, a weighted cross-entropy loss function is used for optimization. The sample weight factor is obtained by normalizing the comprehensive difficulty score, so that the model pays more attention to difficult samples, and the learning rate γ controls the step size of the model parameter update, avoiding the problem of overfitting or slow convergence of the model in the training process.

[0110] In another alternative embodiment, the recognition model can also be an LSTM model based on a recurrent neural network. Specifically, the key frame sequence of the difficult video is sequentially input into the LSTM model. The LSTM can effectively handle long-term dependencies through its unique gating mechanism, including the forget gate, the input gate, and the output gate, and remember the behavior characteristics at different time points in the video. In the training phase, the training set and the test set are divided, and the LSTM model is trained using the training set. The model continuously adjusts its parameters during the training process to minimize the weighted cross-entropy loss function. For the input at each time step, the model makes a prediction based on the current input and the previously remembered information. The prediction result and the true label are used to calculate the error through the weighted cross-entropy loss function, and the model parameters are updated according to the formula. The learning rate γ determines the magnitude of each parameter update. By setting γ reasonably, the model can gradually converge to a better state during the training process, thereby accurately identifying the different time behavior categories of the videos in the test set.

[0111] It should be noted that the present application collects power grid operation video data in various complex environments and annotates it, providing basic data for subsequent analysis. Pseudo-text is generated using a multi-modal model, comprehensively covering various aspects of operation scene information. Fine-grained features are extracted from the pseudo-text, and a comprehensive difficulty score is calculated. The difficulty level of the video sample is scientifically quantified, and difficult videos are selected to provide efficient data for model training. The present application solves the problems faced by traditional methods in power grid operation video analysis, such as model updating due to changes in equipment and specifications, sample scarcity, and inaccurate difficult sample mining, thereby improving the reliability of power industry safety supervision and providing strong technical support for intelligent power grid operation, ensuring the safety and efficiency of power grid operation.

[0112] Embodiment 3, with reference to Tables 1-2, is an embodiment of the present application. Based on the above embodiments, a power grid operation difficult video mining and behavior recognition method is provided. To verify the beneficial effects of the present application, scientific demonstration is carried out through experiments.

[0113] In this embodiment, the collected data includes 870 videos and 2107 action examples, including 760 pole climbing examples, 313 ground wire hanging examples, and an action data retention rate of about 1 / 6. The training set and the validation set are divided at a ratio of about 3:1. The training set includes 642 videos, including 592 pole climbing actions and 272 ground wire hanging actions. The validation set includes 228 videos, including 168 pole climbing actions and 41 ground wire hanging actions. The collected data is used to train the recognition model.

[0114] The trained recognition model was deployed on a test system, and the model's recognition results in typical power grid operation scenarios were obtained using test data samples. The test system's software configuration included Python 3.8, PyTorch 1.9.0, CUDA 11.2, and PyCharm 2023.2, and the hardware configuration included an RTX-3090 graphics card (24GB of video memory).

[0115] The recognition model was used to conduct 10 tests using 150 newly collected sample data from the power grid outage operation scenario, and the average accuracy of the recorded sample classification was taken. The results are shown in Table 1.

[0116] Table 1 Comparison of model results on test data

[0117] Action category Baseline method (accuracy) This method (accuracy) Climbing 87.18% 89.64% Connecting the ground wire 81.23% 83.56%

[0118] A comparison between the baseline method and this method was conducted using a visual language model comprising a visual encoder and a text encoder. As shown in Table 1, the action recognition model of this embodiment was tested on newly collected data from a power outage scenario. The accuracy of this method for climbing and grounding wires was 89.64% and 83.56%, respectively, with an average accuracy of 86.6%. This represents a 2.39% improvement over the baseline method's average accuracy of 84.21%. While the baseline method does not select difficult samples, the results of this method demonstrate that by integrating the semantic understanding capabilities of a large model and drawing on higher-level semantic information, the selected difficult samples effectively improve the performance of the deep learning model in complex environments, thereby enhancing recognition accuracy.

[0119] Table 2 Comparison of model training efficiency

[0120]

[0121] A comparison of our method with a baseline method without a text encoder was conducted. As shown in Table 2, our method also achieves significant improvements in model training efficiency. By introducing pseudo-text generated based on a large model to mine fine-grained difficult samples, the training set includes more challenging data, accelerating model convergence and improving overall training efficiency. This demonstrates that our method can improve the automation and efficiency of training while maintaining recognition performance.

[0122] Embodiment 4, the above is a schematic scheme of a power grid operation difficulty video mining and behavior recognition method. It should be noted that the technical scheme of the power grid operation difficulty video mining and behavior recognition system belongs to the same concept as the technical scheme of the power grid operation difficulty video mining and behavior recognition method described above. The technical details of the power grid operation difficulty video mining and behavior recognition system in this embodiment are not described in detail. Please refer to the description of the technical scheme of the power grid operation difficulty video mining and behavior recognition method described above.

[0123] The embodiment also provides a power grid operation difficulty video mining and behavior recognition system, comprising:

[0124] A data acquisition module is configured to collect video data in a power grid operation scene, and label the video data with different time behavior categories.

[0125] A data processing module is configured to perform inference learning on the labeled video data by using a multi-modal model, and generate pseudo-text of the operation scene.

[0126] A feature acquisition module is configured to extract fine-grained features of the pseudo-text, acquire a comprehensive difficulty score of the video according to the fine-grained features, compare the comprehensive difficulty score with a preset threshold, and obtain a difficult video.

[0127] An identification module is configured to input the difficult video into an identification model, test the identification model, and obtain an identification result of different time behavior categories in the power grid operation scene.

[0128] The embodiment also provides an electronic device suitable for power grid operation difficulty video mining and behavior recognition, comprising a memory and a processor. The memory is configured to store computer executable instructions, and the processor is configured to execute the computer executable instructions to implement the power grid operation difficulty video mining and behavior recognition method proposed in the above embodiment.

[0129] The embodiment also provides a storage medium having a computer program stored thereon. When the program is executed by a processor, the power grid operation difficulty video mining and behavior recognition method proposed in the above embodiment is implemented.

[0130] The storage medium proposed in the embodiment and the power grid operation difficulty video mining and behavior recognition method proposed in the above embodiment belong to the same inventive concept. The technical details not described in detail in the embodiment can be referred to the above embodiment, and the embodiment has the same beneficial effects as the above embodiment.

[0131] Through the above description of the implementation methods, those skilled in the art can clearly understand that the present invention can be implemented with the help of software and necessary general hardware, and of course can also be implemented by hardware. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory (FLASH), hard disk or optical disk, etc., including a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods of various embodiments of the present invention.

[0132] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A method for mining and identifying difficult power grid operation videos, characterized in that: include: Collect video data from power grid operation scenes and label the video data with different time behavior categories; The labeled video data is used for inference learning through a multimodal model to generate pseudo-text of the work scenario; Extracting fine-grained features of the pseudo-text, obtaining a comprehensive difficulty score of the video based on the fine-grained features, and comparing the comprehensive difficulty score with a preset threshold to obtain a difficult video; The difficult video is input into a recognition model, and the recognition model is used for testing to obtain recognition results of different time behavior categories in the power grid operation scene.

2. The method for mining difficult power grid operation videos and identifying behaviors according to claim 1, characterized in that: The labeled video data is used for inference learning through a multimodal model to generate pseudo text for the task scenario, including: Set power grid operation slogans based on video data; Key frames of video data are obtained, and based on the power grid operation slogans, visual-text reasoning is performed on the key frames using a multimodal large model to generate pseudo text including scene information, target information, action information and time information.

3. The method for mining difficult power grid operation videos and identifying behaviors according to claim 2, characterized in that: Extracting fine-grained features of the pseudo-text and obtaining a comprehensive difficulty score of the video based on the fine-grained features, including: Performing structural analysis on the pseudo-text, setting feature categories and extracting fine-grained features; The fine-grained features are numerically converted using a mapping function to obtain a difficulty score for each feature; The difficulty scores of each feature are weighted and summed to obtain the comprehensive difficulty score of each video.

4. The method for mining difficult power grid operation videos and identifying behaviors according to claim 3, characterized in that: The recognition model was used to conduct tests and obtain recognition results for different time behavior categories in power grid operation scenarios, including: Dividing the difficult video into a training set and a test set, using the training set to train the recognition model, inputting the test set into the trained recognition model, and obtaining recognition results of different time behavior categories in the power grid operation scene; The weighted cross entropy loss function is used to optimize the recognition model and update the recognition model parameters.

5. The method for mining difficult power grid operation videos and identifying behaviors according to claim 4, characterized in that: Perform structural analysis on the pseudo text and set feature categories to extract fine-grained features, including: The set feature categories include number of targets, weather conditions, time of day and lighting conditions; The difficulty scores of the number of targets, weather conditions, time and lighting conditions are calculated separately, and combined with weighting parameters to obtain the comprehensive difficulty score.

6. The method for mining difficult power grid operation videos and identifying behaviors according to claim 2, characterized in that: The formula for calculating the comprehensive difficulty score is: Among them, α, β, λ is a weighting parameter that satisfies F1F2F3 represents the combined influence of the three factors, where F1 is the difficulty score of target quantity, F2 is the difficulty score of weather conditions, and F3 is the difficulty score of time and light conditions.

7. The method for mining difficult power grid operation videos and identifying behaviors according to claim 5, characterized in that: Comparing the comprehensive difficulty score with a preset threshold to obtain a difficult video includes: According to the calculated comprehensive difficulty score, determine whether the video is a difficult sample, which is expressed as: Where M is the number of video samples, Is the video sample V j The comprehensive difficulty score of τ is the set comprehensive difficulty score threshold. is the jth difficult sample, S har For difficult video collection.

8. A power grid operation difficulty video mining and behavior recognition system, applying the method according to any one of claims 1 to 7, characterized in that: include: A data acquisition module is used to collect video data in power grid operation scenes and label the video data with different time behavior categories; The data processing module is used to perform inference learning on the annotated video data through a multimodal model to generate pseudo-text of the operation scene; a feature acquisition module, configured to extract fine-grained features of the pseudo-text, obtain a comprehensive difficulty score of the video based on the fine-grained features, and compare the comprehensive difficulty score with a preset threshold to obtain a difficult video; The recognition module is used to input the difficult video into the recognition model, use the recognition model to perform testing, and obtain recognition results of different time behavior categories in the power grid operation scene.

9. An electronic device, characterized in that: include: memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the method for mining difficult power grid operation videos and identifying behaviors as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that It stores computer-executable instructions, which, when executed by a processor, implement the steps of the method for mining difficult power grid operation videos and identifying behaviors as described in any one of claims 1 to 7.