Power maintenance risk operation identification method, system and device and medium

By extracting key points of the skeleton during power maintenance and optimizing human motion sequences using a frequency domain encoder and a noise prediction network, the problem of inaccurate motion sequences in existing technologies is solved, enabling accurate identification and early warning of risky operations during power maintenance.

CN120823447APending Publication Date: 2025-10-21GUANGZHOU POWER SUPPLY BUREAU GUANGDONG POWER GRID CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511137200.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-14
Publication Date
2025-10-21

AI Technical Summary

Technical Problem

The existing human motion prediction method in power maintenance uses fixed value filling and DCT transformation, which leads to inaccurate future human motion sequences and inaccurate risk operation identification results.

Method used

By extracting the skeleton key point information from power maintenance video data, a real-time motion sequence is generated. A target predictor is used for prediction, and frequency domain encoder and noise prediction network are combined for frequency domain transformation and denoising. The overall trend and detailed changes are separated, and noise interference of different frequency components is optimized to finally generate an accurate future motion sequence.

Benefits of technology

It improves the accuracy of future human movement sequences, enhances the ability to identify risky operations in power maintenance, provides a timely early warning basis, and ensures the safety of power maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120823447A_ABST
    Figure CN120823447A_ABST
Patent Text Reader

Abstract

The invention discloses an electric power maintenance risk operation identification method, system and device and a medium, and belongs to the field of electric power systems.The method comprises the steps that a real-time sequence extracted from electric power maintenance video data and a predicted initial prediction sequence are spliced to obtain a spliced sequence, and frequency domain conversion is conducted on the spliced sequence to obtain frequency domain features; extracting a first frequency component and a second frequency component from the frequency domain feature as condition input, sampling from standard normal distribution to obtain first frequency sampling noise and second frequency sampling noise, and respectively inputting the first frequency sampling noise and the second frequency sampling noise into a first noise prediction network and a second noise prediction network for iterative denoising, the first de-noising feature and the second de-noising feature are fused to obtain a fused feature, and the fused feature is decoded to obtain a target prediction sequence; and carrying out risk operation identification according to the target prediction sequence, and if so, carrying out early warning. Therefore, by implementing the method and the device, accurate prediction of future motion can be realized, and the accuracy of risk operation identification is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of power systems, and in particular to a method, system, equipment and medium for identifying risky operations in power maintenance. Background Art

[0002] In the power industry, maintenance operations often involve high-risk scenarios such as working at height and operating with live power lines. Risky operations such as irregular movements or unexpected behaviors by operators can potentially lead to accidents. Therefore, using computer vision technology to monitor human motion in real time and predict and identify risky operations has become a crucial tool for ensuring power maintenance safety. Human motion prediction, a core technology for identifying risky operations in power maintenance, infers future motion trends based on historical motion sequences, providing a data foundation for risk warnings.

[0003] Currently, existing human motion prediction techniques first roughly fill the future motion sequence with fixed values ​​to form a complete motion sequence. This complete sequence is then mapped to the frequency domain using the Discrete Cosine Transform (DCT) and diffused in the frequency domain. Multiple noise samples are then sampled and denoised to generate multiple predicted future human motion sequences. However, using fixed values ​​to fill the future sequence disrupts the natural connection between historical and future motion, creating a significant "gap" in the time domain for the initial complete sequence, introducing invalid noise into subsequent frequency domain processing. Furthermore, after mapping the sequence to the frequency domain, the DCT transform fails to differentiate between the first frequency component (high-frequency features corresponding to overall motion trends, such as body movement direction) and the second frequency component (low-frequency features corresponding to detailed movements, such as finger movement amplitude). During the frequency domain diffusion process, the first frequency component is prone to aliasing or oversmoothing, resulting in either distortion of overall motion trends (e.g., sudden reverse movement of the body) or loss of local details (e.g., blurred tool grasping movements) when restored to the time domain.

[0004] In summary, some future human motion sequences generated by this method are not accurate enough and have distortion and floating phenomena, which leads to inaccurate identification results of power maintenance risk operations. Summary of the Invention

[0005] The present invention provides a method, system, device and medium for identifying electric power maintenance risk operations, which can solve the problem of inaccurate identification results of electric power maintenance risk operations caused by inaccurate human motion sequences in the future.

[0006] The present invention provides a method for identifying risky operations in power maintenance, comprising:

[0007] Extracting key point information of the human skeleton from the collected power maintenance video data to obtain a real-time motion sequence, and predicting the real-time motion sequence through a target predictor to obtain an initial predicted motion sequence;

[0008] splicing the real-time motion sequence with the initial predicted motion sequence to obtain a spliced ​​sequence, performing frequency domain conversion on the spliced ​​sequence using a target frequency domain encoder to obtain frequency domain features, and extracting a first frequency component and a second frequency component from the frequency domain features, wherein the frequency domain features are sorted in ascending order of frequency, the first frequency refers to a preset number of frequency components in the frequency domain features, and the second frequency refers to a preset number of frequency components uniformly selected from frequency components of the frequency domain features other than the first frequency component;

[0009] Sampling frequency domain noise from a standard normal distribution, dividing the frequency domain noise into first frequency sampling noise and second frequency sampling noise, inputting the first frequency component as a condition into a first frequency noise prediction network, and inputting the first frequency sampling noise and the current time step into the first frequency noise prediction network, so as to iteratively denoise the first frequency sampling noise through the first frequency noise prediction network; at the same time, inputting the second frequency component as a condition into a second frequency noise prediction network, and inputting the second frequency sampling noise and the current time step into the second frequency noise prediction network, so as to iteratively denoise the second frequency sampling noise through the second frequency noise prediction network, until the number of iterations equals a preset number, stopping the iteration, and obtaining a first frequency denoising feature output by the first frequency noise prediction network, and a second frequency denoising feature output by the second frequency noise prediction network, respectively;

[0010] The first frequency denoising feature and the second frequency denoising feature are spliced ​​and input into the target feature fusion device for feature fusion to obtain a fused feature, and the fused feature is input into the target frequency domain decoder for feature decoding to obtain a target predicted motion sequence corresponding to the real-time motion sequence; risk operations are identified according to the target predicted motion sequence, and when the risk operation occurs, a warning message is issued.

[0011] The embodiment of the present invention converts unstructured video data into structured motion sequence data by extracting skeleton key points, providing a modelable input basis for subsequent prediction and retaining the temporal and spatial characteristics of human motion; generates a preliminary estimate of future motion based on the historical motion sequence, provides an initial reference for subsequent frequency domain optimization, and reduces the prediction difficulty of the subsequent model; by splicing and fusing historical and preliminary future information, the frequency domain conversion maps the time domain motion features to the frequency domain, realizes the separate modeling of "overall trend (first frequency)" and "detail change (second frequency)", and improves the targeted feature expression; through separate denoising in different frequency bands, the noise interference of different frequency components is targetedly optimized, the distortion of the first frequency trend and the loss of the second frequency details are reduced, and the purity of the frequency domain features is improved; the fused and optimized first frequency component is decoded and restored to the time domain to generate a future motion sequence with both overall trend accuracy and detail continuity, thereby improving the integrity of the prediction results; based on accurate future motion prediction, early identification of potential dangerous actions is achieved, providing timely warning basis for power maintenance safety.

[0012] Furthermore, the iterative denoising of the first frequency sampling noise by the first frequency noise prediction network is specifically performed as follows:

[0013] The first frequency component in the frequency domain feature is used as a first frequency condition, and is input into the first frequency noise prediction network together with the first frequency sampling noise and the current time step;

[0014] In each round of iteration of the first frequency noise prediction network, for the current time step, based on the first frequency condition, the current first frequency sampling noise of the current time step is predicted by the first frequency noise prediction network to obtain the current first frequency noise residual corresponding to the current time step, and the current first frequency sampling noise is denoised according to the current first frequency noise residual to obtain the first frequency denoising feature of the previous time step, and the first frequency denoising feature of the previous time step is converted back to the time domain to obtain the first time domain denoising feature of the previous time step; at the same time, the first frequency component is added to the previous time step and converted back to the time domain to obtain the first frequency component of the previous time step. The first time domain noise feature of the previous time step is obtained, and the real-time motion sequence part in the first time domain noise feature of the previous time step is spliced ​​with the predicted motion sequence part in the first time domain denoising feature of the previous time step to obtain a spliced ​​sequence of the previous time step, and the spliced ​​sequence is converted into the frequency domain by the target frequency domain encoder to obtain the frequency domain feature of the previous time step, and the first frequency component of the frequency domain feature of the previous time step is extracted to obtain an updated first frequency sampling noise. The updated first frequency sampling noise is used for the next iteration of the first frequency noise prediction network until the number of iterations is equal to a preset number, and the first frequency denoising feature is obtained.

[0015] By using the first frequency of the initial predicted frequency domain features as a conditional input and constraining the first frequency component of the initial prediction, the noise prediction network is guided to focus on the difference between the initial prediction and the actual future, improving the targetedness of the conditional modeling. By sampling the first and second frequency noise inputs into the initial noise prediction network to predict noise, simulating the noise addition and denoising process, the noise prediction network learns to identify and remove the second frequency detail noise and the first frequency trend noise, improving denoising capabilities. By minimizing the difference between the predicted noise and the actual noise, the network's noise recognition accuracy is iteratively optimized, making the denoised first frequency component closer to the actual distribution.

[0016] Furthermore, the human skeleton key point information is extracted from the collected power maintenance video data to obtain a real-time motion sequence, specifically:

[0017] Extracting video frames from each power maintenance video data using a video processing tool, and adjusting the frame rate of each extracted video frame to a preset frequency to obtain multiple video frames;

[0018] Each video frame is processed using a 3D human pose estimation model to extract skeleton key point information. The corresponding 2D array for each video frame is then determined based on the number of joints and the 3D coordinate dimensions of each joint.

[0019] The two-dimensional arrays corresponding to all video frames are sorted in chronological order to obtain a key point sequence, and the key point sequence is converted into a time series data tensor according to a four-dimensional structure of batch, time step, joint point and coordinate dimensions, and the time series data tensor is standardized to obtain the real-time motion sequence.

[0020] In this way, by extracting and adjusting the video frames to the preset frame rate, the time scale of the video data is unified, the inconsistent timing features of different videos caused by frame rate differences are eliminated, and the temporal continuity of the motion sequence is ensured. Skeleton key points are extracted through three-dimensional human posture estimation to generate a single-frame two-dimensional array (joint point × three-dimensional coordinates). The three-dimensional posture estimation accurately captures the spatial position of the joints, converts the single-frame image into quantized joint coordinate data, and retains the spatial structure information of the human posture. The two-dimensional array is sorted by time to obtain a key point sequence, which is converted into a four-dimensional time series data tensor to structure the discrete frame data into a time series tensor (batch × time step × joint point × coordinate, which is adapted to the input format of the deep learning model and facilitates batch processing and time-dependent modeling. The real-time motion sequence is obtained by standardizing the time series data tensor. Standardization eliminates data noise and scale deviation, improves the consistency and stability of the motion sequence, and provides high-quality input for subsequent model training and prediction.

[0021] Furthermore, the time series data tensor is normalized to obtain the real-time motion sequence, specifically:

[0022] Taking a preset central key point as a spatial reference benchmark, calculating the offsets of all relevant nodes in the time series data tensor relative to the central key point to perform centralization processing and obtain a centralized tensor;

[0023] For the three-dimensional coordinates of all relevant nodes in the centralized tensor, calculate the scale distribution characteristics of all three-dimensional coordinates, determine a target scale range based on the scale distribution characteristics, and map all three-dimensional coordinates to the target scale range to obtain a first time series data tensor;

[0024] A sliding average algorithm or a filtering algorithm is used to perform time series smoothing processing on the key point trajectories in the first time series data tensor to obtain the real-time motion sequence.

[0025] In this way, by calculating the offset (centralization) based on the central key point, the global translation deviation of the human body in space (such as the difference in human coordinates shot at different positions) is eliminated, the coordinates of the joint points are unified relative to the center, and the spatial consistency of the motion characteristics is enhanced. By mapping the three-dimensional coordinates to the target scale range (scaling) based on the scale distribution characteristics, the spatial scales of different human bodies / scenes are unified (such as the difference in coordinate values ​​caused by differences in height and shooting distance), and the coordinates of the joint points are distributed within the same order of magnitude, improving the model's adaptability to scale changes. The key point trajectory is smoothed and denoised through a sliding average / filtering algorithm to remove noise points caused by occlusion and jitter during the posture estimation process, reduce the interference of the second frequency jitter, and enhance the temporal continuity and smoothness of the motion trajectory.

[0026] Furthermore, the target predictor is obtained by training the motion sequence prediction of historical power maintenance video data based on the initial predictor, wherein the initial predictor adopts a multi-layer perceptron structure, specifically including:

[0027] Acquire a historical motion sequence corresponding to the historical power maintenance video data, and divide the historical motion sequence according to time sequence to obtain a continuous first motion sequence and a second motion sequence;

[0028] The first motion sequence is used as the input of the initial predictor to obtain a first prediction sequence, and the second motion sequence is used as the expected output of the initial predictor, an error between the first prediction sequence and the second motion sequence is determined by a joint loss function, and the initial predictor is trained by minimizing the error to obtain the target predictor, wherein the joint loss function includes an average per-joint position error and an average per-joint velocity error.

[0029] In this way, by dividing the historical motion sequence into a continuous first motion sequence (input) and a second motion sequence (label) according to time, a supervised learning sample pair of "past→future" is constructed to simulate the temporal dependency of the real prediction scenario and provide the model with training data that conforms to actual logic; an initial predictor with a multi-layer perceptron structure is constructed, with the first sequence as input and the second sequence as the expected output. Through the nonlinear mapping capability of the multi-layer perceptron, the basic association law from historical motion to future motion is learned to provide model architecture support for initial prediction; through training with the joint loss function (MPJPE+MPJVE), MPJPE constrains the spatial deviation between the predicted joint position and the true value, thereby improving position accuracy; MPJVE constrains the continuity of joint velocity changes, reduces motion mutations, and makes the initial prediction sequence both accurate and smooth.

[0030] Furthermore, the target frequency domain encoder, target feature fusion device and target frequency domain decoder are obtained by training the historical motion sequence according to the initial frequency domain encoder, initial feature fusion device and initial frequency domain decoder, respectively, and specifically include:

[0031] Processing the historical motion sequence by the initial frequency domain encoder to obtain initial frequency domain features, and extracting an initial first frequency component and an initial second frequency component from the initial frequency domain features;

[0032] splicing the initial first frequency component and the initial second frequency component and inputting them into the initial feature fusion device to obtain an initial fusion feature; and inputting the initial fusion feature into the initial frequency domain decoder to obtain an initial motion sequence;

[0033] Determine the average per-joint position error between the historical motion sequence and the initial motion sequence, and determine the mean square error between the initial frequency domain feature and the encoding result obtained by discrete cosine transform. By minimizing the average per-joint position error and the mean square error, the parameters of the initial frequency domain encoder, the initial feature fuser and the initial frequency domain decoder are optimized to obtain the target frequency domain encoder, the target feature fuser and the target frequency domain decoder respectively.

[0034] In this way, the frequency domain features of the historical sequence are processed by the initial frequency domain encoder, and the initial historical first frequency (first N) and second frequency (evenly selected N) are extracted. Through frequency segment extraction, the frequency domain features are strengthened to express the hierarchical "trend-detail" of the historical motion, laying the characteristic foundation for subsequent reconstruction and optimization. By fusing the first frequency component and decoding it back to the time domain, the effectiveness of the frequency domain encoding and decoding process is verified to ensure that the frequency domain features can accurately restore the original motion sequence. MPJPE ensures the consistency of the reconstructed sequence with the original sequence, improving the encoding and decoding accuracy; DCT mean square error constrains the frequency domain features to conform to the actual frequency domain distribution law, enhancing the physical rationality of the frequency domain features.

[0035] Furthermore, the first frequency noise prediction network and the second frequency noise prediction network specifically include: an input projection layer, a feature decoder 1, a feature decoder 2, and an output projection layer connected in sequence;

[0036] The input projection layer is used to project the frequency domain features of the current time step from the initial dimensional feature space to the target dimensional feature space;

[0037] The feature decoder 1 includes a Mamba2 module, a first self-attention layer, and a first fully connected layer. The Mamba2 module implements selective feature processing based on a structured state-space dual framework. The feature decoder 1 injects information of the current time step into the feature processing process through adaptive layer normalization to model long-term dependency information.

[0038] The second feature decoder includes a second self-attention layer and a second fully connected layer. The second feature decoder injects the information of the current time step into the feature processing process through feature linear modulation to simultaneously model long-term dependency information and short-term dependency information;

[0039] The output projection layer is used to project the frequency domain features output by the second feature decoder from the target dimensional feature space to the initial dimensional feature space to obtain a predicted noise component.

[0040] In this way, the input projection layer projects frequency domain features from their initial dimensions to a high-dimensional space, expanding the feature expression dimension through high-dimensional space, enhancing the model's ability to capture complex frequency domain patterns and improving the richness of feature representation. The Mamba2 module efficiently models long-term temporal dependencies, while the self-attention layer captures feature associations. AdaLN dynamically adapts time-step information to strengthen targeted processing of features at different denoising stages, improving the accuracy of long-term dependency modeling. The self-attention layer deepens feature association modeling, while FiLM considers the impact of time steps on both long-term and short-term dependencies, achieving a joint optimization of "overall trends and local details." By restoring high-dimensional processed features to their original frequency domain dimensions, the noise prediction results are ensured to match the input feature dimensions, improving the practicality and accuracy of noise prediction.

[0041] Another embodiment of the present invention further provides a system for identifying risky operations of power maintenance, comprising: a prediction module, an encoding module, a denoising module, a decoding module, and an identification module;

[0042] The prediction module is used to extract human skeleton key point information from the collected power maintenance video data to obtain a real-time motion sequence, and predict the real-time motion sequence through a target predictor to obtain a predicted motion sequence;

[0043] The encoding module splices the real-time motion sequence and the predicted motion sequence to obtain a spliced ​​sequence, performs frequency domain conversion on the spliced ​​sequence through a target frequency domain encoder to obtain frequency domain features, and extracts a first frequency component and a second frequency component from the frequency domain features, wherein the frequency domain features are sorted in ascending order of frequency, the first frequency refers to a preset number of frequency components in the frequency domain features, and the second frequency refers to a preset number of frequency components uniformly selected from frequency components of the frequency domain features other than the first frequency component;

[0044] The denoising module is configured to generate frequency domain noise by sampling from a standard normal distribution, divide the frequency domain noise into first frequency sampling noise and second frequency sampling noise, input the first frequency component as a condition into a first frequency noise prediction network, and input the first frequency sampling noise and the current time step into the first frequency noise prediction network, so as to iteratively denoise the first frequency sampling noise through the first frequency noise prediction network; at the same time, input the second frequency component as a condition into a second frequency noise prediction network, and input the second frequency sampling noise and the current time step into the second frequency noise prediction network, so as to iteratively denoise the second frequency sampling noise through the second frequency noise prediction network, until the number of iterations equals a preset number, then stop the iteration, and obtain a first frequency denoising feature output by the first frequency noise prediction network, and a second frequency denoising feature output by the second frequency noise prediction network, respectively;

[0045] The decoding module is used to splice the first frequency denoising feature and the second frequency denoising feature and input them into the target feature fusion device for feature fusion to obtain a fused feature, and input the fused feature into the target frequency domain decoder for feature decoding to obtain a target predicted motion sequence corresponding to the real-time motion sequence;

[0046] The identification module is used to identify risky operations based on the target predicted motion sequence, and issue a warning message when the risky operation occurs.

[0047] The embodiment of the present invention converts unstructured video data into structured motion sequence data by extracting skeleton key points, providing a modelable input basis for subsequent prediction and retaining the temporal and spatial characteristics of human motion; generates a preliminary estimate of future motion based on the historical motion sequence, provides an initial reference for subsequent frequency domain optimization, and reduces the prediction difficulty of the subsequent model; by splicing and fusing historical and preliminary future information, the frequency domain conversion maps the time domain motion features to the frequency domain, realizes the separate modeling of "overall trend (first frequency)" and "detail change (second frequency)", and improves the targeted feature expression; through separate denoising in different frequency bands, the noise interference of different frequency components is targetedly optimized, the distortion of the first frequency trend and the loss of the second frequency details are reduced, and the purity of the frequency domain features is improved; the fused and optimized first frequency component is decoded and restored to the time domain to generate a future motion sequence with both overall trend accuracy and detail continuity, thereby improving the integrity of the prediction results; based on accurate future motion prediction, early identification of potential dangerous actions is achieved, providing timely warning basis for power maintenance safety.

[0048] Another embodiment of the present invention further provides a terminal device, comprising: a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the steps of the power maintenance risk operation identification method of the present invention are implemented.

[0049] Another embodiment of the present invention further provides a computer-readable storage medium item, comprising: a stored computer program, which controls the device where the computer-readable storage medium is located to execute the steps of the power maintenance risk operation identification method of the present invention when the computer program is running. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the technical solution of the present application, the following is a brief introduction to the drawings required for use in the implementation. Obviously, the drawings described below are only some implementation methods of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0051] Figure 1 This is a flow chart of a method for identifying risky operations in power maintenance provided by an embodiment of the present invention;

[0052] Figure 2 1 is a flow chart of the training steps of a noise prediction network provided by an embodiment of the present invention;

[0053] Figure 3 1 is a schematic diagram of an iterative update process of denoising a first frequency component and a second frequency component provided by an embodiment of the present invention;

[0054] Figure 41 is a schematic diagram of the structure of a noise prediction network provided by an embodiment of the present invention;

[0055] Figure 5 This is a flow chart of the training steps of a target frequency domain encoder, a target feature fuser, and a target frequency domain decoder provided by an embodiment of the present invention;

[0056] Figure 6 The present invention provides a schematic diagram of the structure of a power maintenance risk operation identification system. DETAILED DESCRIPTION

[0057] To make the objectives, technical solutions, and advantages of this application more clear, the technical solutions in this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this application.

[0058] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application belongs; the terms used herein are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above-mentioned figure descriptions are intended to cover non-exclusive inclusions.

[0059] In the description of the embodiments of this application, the technical terms "first" and "second" are used only to distinguish different objects and should not be understood to indicate or imply relative importance or implicitly specify the quantity, specific order, or primary and secondary relationship of the indicated technical features. In the description of the embodiments of this application, the meaning of "plurality" is more than two, unless otherwise clearly and specifically defined.

[0060] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0061] In the description of the embodiments of this application, the term "and / or" is simply a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent the following three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this document generally indicates that the associated objects are in an "or" relationship.

[0062] In the description of the embodiments of the present application, the term "multiple" refers to more than two (including two). Similarly, "multiple groups" refers to more than two groups (including two groups), and "multiple pieces" refers to more than two pieces (including two pieces).

[0063] In the description of the embodiments of the present application, unless otherwise expressly specified or limited, technical terms such as "installed," "connected," "connected," and "fixed" should be understood in a broad sense. For example, they can refer to fixed connections, detachable connections, or integration; mechanical connections or electrical connections; direct connections or indirect connections through an intermediate medium; internal connections between two components or interactions between two components. Those skilled in the art can understand the specific meanings of the above terms in the embodiments of the present application based on specific circumstances.

[0064] See also Figure 1 To solve the problem of inaccurate identification of power maintenance risk operations in the prior art, an embodiment of the present invention provides a method for identifying power maintenance risk operations, comprising:

[0065] Step S101: extracting key point information of a human skeleton from collected power maintenance video data to obtain a real-time motion sequence, and predicting the real-time motion sequence through a target predictor to obtain a predicted motion sequence.

[0066] In this embodiment, the collected power maintenance video data can be extracted using a video processing tool and adjusted to a uniform frame rate to produce a continuous video frame sequence. Three-dimensional human pose estimation technology is then used to extract the coordinate information of key points of the human skeleton from each video frame, forming time-series key point data reflecting human pose. This time-series key point data is then standardized, for example, by eliminating spatial position offsets, unifying the coordinate scale range, and smoothing and denoising, ultimately resulting in a real-time motion sequence. A target predictor is constructed based on the human motion sequence data in the power maintenance video data. Training samples are constructed by dividing past motion segments into corresponding future motion segments. A prediction model is constructed and trained using a multi-layer perceptron architecture to produce a target predictor capable of predicting future motion trends based on historical motion sequences. The real-time motion sequence is input into the target predictor, which then predicts the future motion trend of the real-time motion sequence and outputs an initial predicted motion sequence.

[0067] As an example of an embodiment of the present invention, the human skeleton key point information is extracted from the collected power maintenance video data to obtain a real-time motion sequence, specifically:

[0068] Video frames are extracted from each power maintenance video data through a video processing tool, and the frame rate of each extracted video frame is adjusted to a preset frequency to obtain multiple video frames; each video frame is processed through a three-dimensional human posture estimation model to extract skeleton key point information, and the two-dimensional array corresponding to each video frame is confirmed based on the number of joint points and the three-dimensional coordinate dimension of each joint point; the two-dimensional arrays corresponding to all video frames are sorted in chronological order to obtain a key point sequence, and the key point sequence is converted into a time series data tensor according to a four-dimensional structure of batch, time step, joint point and coordinate dimension, and the time series data tensor is standardized to obtain the real-time motion sequence.

[0069] In this embodiment, the collected power maintenance video data is first extracted using video processing tools such as OpenCV, and the frame rate of the extracted video frames is uniformly adjusted to 50fps to obtain a continuous video frame sequence. Then, a three-dimensional human posture estimation model such as MediaPipe Pose is used to process each frame of video to extract the key point information of the human skeleton. Each frame of video outputs a two-dimensional array, for example, the array size is 17×3 (17 joints, each joint contains three-dimensional coordinate information). Afterwards, the two-dimensional arrays corresponding to all video frames are arranged in chronological order and organized into a time series data tensor with a shape of (B, T, N, C), where B is the batch size, T is the time step, N is the number of joints (17), and C is the coordinate dimension (3 dimensions). In order to improve the training quality of the subsequent prediction model, the time series data tensor is standardized to eliminate noise and scale deviation in the data, and finally obtain a real-time motion sequence.

[0070] As an example of an embodiment of the present invention, the normalization processing of the time series data tensor to obtain the real-time motion sequence is specifically as follows:

[0071] Taking the preset central key point as the spatial reference benchmark, calculate the offset of all relevant nodes in the time series data tensor relative to the central key point to perform centralization processing and obtain a centralized tensor; for the three-dimensional coordinates of all relevant nodes in the centralized tensor, calculate the scale distribution characteristics of all three-dimensional coordinates, determine the target scale range according to the scale distribution characteristics, and map all three-dimensional coordinates to the target scale range to obtain a first time series data tensor; use a sliding average algorithm or a filtering algorithm to perform time series smoothing processing on the key point trajectories in the first time series data tensor to obtain the real-time motion sequence.

[0072] In this embodiment, a preset central key point such as the pelvis or spine is used as a reference point, and the offsets of all related nodes relative to the central key point are calculated to complete centralization and eliminate the global translation deviation of different human bodies in spatial positions; then the three-dimensional coordinates of all joint points are standardized and scaled, and the target scale range is determined based on the scale distribution characteristics of the joint point coordinates in the training set, and the coordinates are mapped to this range to unify the spatial scale and enhance the stability and generalization ability of the model training; in addition, a sliding average or filtering algorithm is used to denoise the key point trajectory to reduce the noise impact caused by posture estimation and improve the temporal continuity and stability of the data.

[0073] As an example of an embodiment of the present invention, the target predictor is obtained by training the historical power maintenance video data for motion sequence prediction based on the initial predictor, wherein the initial predictor adopts a multi-layer perceptron structure and specifically includes:

[0074] Obtain a historical motion sequence corresponding to the historical power maintenance video data, divide the historical motion sequence in chronological order, and obtain a continuous first motion sequence and a second motion sequence; use the first motion sequence as the input of the initial predictor to obtain a first prediction sequence, and use the second motion sequence as the expected output of the initial predictor, determine the error between the first prediction sequence and the second motion sequence through a joint loss function, train the initial predictor by minimizing the error, and obtain the target predictor, wherein the joint loss function includes an average per-joint position error and an average per-joint velocity error.

[0075] In this embodiment, after obtaining historical power maintenance video data, the key point information of the human skeleton is extracted from it to form a historical motion sequence. The historical motion sequence is divided into two continuous parts in chronological order, wherein the first motion sequence is used as the input part of the initial predictor during training (representing "past motion") to obtain the first prediction sequence output by the model. The second motion sequence is used as the expected output, and the error between the first prediction sequence and the second motion sequence is calculated by a joint loss function. The joint loss function is composed of the mean per-joint position error (MPJPE, which measures the joint spatial position deviation) and the mean per-joint velocity error (MPJVE, which constrains motion continuity and physical rationality). The specific form is expressed as follows:

[0076] L = MPJPE + 0.5 × MPJVE;

[0077] The error is minimized through the optimization algorithm, and the network parameters of the initial predictor are iteratively adjusted until the model performance converges, and finally the trained target predictor is obtained.

[0078] Step S102: splice the real-time motion sequence and the predicted motion sequence to obtain a spliced ​​sequence, perform frequency domain conversion on the spliced ​​sequence through a target frequency domain encoder to obtain frequency domain features, and extract a first frequency component and a second frequency component from the frequency domain features, wherein the frequency domain features are sorted in ascending order of frequency, the first frequency refers to a preset number of frequency components in the frequency domain features, and the second frequency refers to a preset number of frequency components uniformly selected from the frequency components of the frequency domain features other than the first frequency component.

[0079] In this embodiment, the target frequency domain encoder, target feature fusion device and target frequency domain decoder are obtained by pre-training based on the collected motion sequence data set. The real-time motion sequence extracted from the power maintenance video (i.e., the past human skeleton sequence) and the predicted motion sequence obtained by the target predictor (i.e., the preliminarily predicted future human skeleton sequence) are spliced ​​in the time dimension to form a complete time series sequence, i.e., a spliced ​​sequence, which contains the past and preliminarily predicted future motion information, and the spliced ​​sequence is input into the target frequency domain encoder. The target frequency domain encoder encodes the input spliced ​​sequence to obtain frequency domain features. When extracting the first frequency component and the second frequency component from the frequency domain features, the first n frequency components in the frequency domain features are taken as the first frequency component, representing the overall motion trend, and m frequency components are uniformly taken from the remaining ln frequency components as the second frequency components to retain fine-grained motion change information.

[0080] Step S103: Sampling frequency domain noise from a standard normal distribution, dividing the frequency domain noise into first frequency sampling noise and second frequency sampling noise, inputting the first frequency component as a condition into a first frequency noise prediction network, and inputting the first frequency sampling noise and the current time step into the first frequency noise prediction network, so as to iteratively denoise the first frequency sampling noise through the first frequency noise prediction network. Simultaneously, inputting the second frequency component as a condition into a second frequency noise prediction network, and inputting the second frequency sampling noise and the current time step into the second frequency noise prediction network, so as to iteratively denoise the second frequency sampling noise through the second frequency noise prediction network, until the number of iterations equals a preset number, stopping the iteration, and obtaining first frequency denoising features output by the first frequency noise prediction network and second frequency denoising features output by the second frequency noise prediction network, respectively.

[0081] In this embodiment, the first frequency noise prediction network and the second frequency noise prediction network It is a pre-trained model. Starting from the denoising time step T, the first frequency component and the second frequency component corresponding to the predicted motion sequence are used as the first frequency noise prediction network and the second frequency noise prediction network Conditional input, and the Classifier-Free Guidance strategy is introduced to discard the conditional input with a certain probability. In each round of iterative denoising, for a trained noise prediction network, such as the first frequency noise prediction network, at the current time step t, the first frequency sampling noise, the first frequency component and the current time step are input into the first noise prediction network to obtain the first frequency denoising feature of the t-1 step. However, the real-time motion sequence part of the first frequency denoising feature of the t-1 step is not the given real-time motion sequence. Therefore, the first frequency component is directly added to the t-1 step to obtain the first frequency domain denoising feature of the t-1 step. For the trained noise prediction network, the noise obtained The first frequency domain noise feature of step t-1 and the first frequency denoised feature of step t-1 obtained by denoising should be the same. Therefore, the noise result is converted back to the time domain, and the real-time motion sequence portion is taken. It is spliced ​​with the predicted motion sequence portion after the denoising result is restored to the time domain. The spliced ​​sequence is transformed into the frequency domain using a frequency domain encoder, and its first frequency component is extracted as the denoising input of time step t-1. The above denoising and feature update steps are repeated until the denoising time step changes from T to 1 and the number of iterations reaches the preset number (i.e., time step T). The iteration is stopped to obtain the first frequency denoised feature. The process of denoising the second frequency noise prediction network to obtain the second frequency denoised feature is similar.

[0082] It should be noted that if Figure 2 The figure shows a flow chart of the training steps of a noise prediction network.

[0083] It should be noted that the first frequency can be a low frequency, and the second frequency can be a high frequency. The low frequency is the first n frequency components in the frequency domain feature, and the high frequency is the n frequency components uniformly selected from the frequency components of the frequency domain feature other than the first n. During the training process of the first frequency noise prediction network and the second frequency noise prediction network, the mean square error between the first frequency noise and the current first frequency noise component is used as a loss function to update the parameters of the first frequency noise prediction network, and the mean square error between the second frequency noise and the current second frequency noise component is used as a loss function to update the parameters of the second frequency noise prediction network.

[0084] As an example of an embodiment of the present invention, the iterative denoising of the first frequency sampling noise by the first frequency noise prediction network is specifically performed as follows:

[0085] The first frequency component in the frequency domain feature is used as a first frequency condition, and is input into the first frequency noise prediction network together with the first frequency sampling noise and the current time step;

[0086] In each round of iteration of the first frequency noise prediction network, for the current time step, based on the first frequency condition, the current first frequency sampling noise of the current time step is predicted by the first frequency noise prediction network to obtain the current first frequency noise residual corresponding to the current time step, and the current first frequency sampling noise is denoised according to the current first frequency noise residual to obtain the first frequency denoising feature of the previous time step, and the first frequency denoising feature of the previous time step is converted back to the time domain to obtain the first time domain denoising feature of the previous time step; at the same time, the first frequency component is denoised to obtain the first frequency domain denoising feature of the previous time step, and the first frequency domain denoising feature is converted back to the time domain to obtain the first time domain denoising feature of the previous time step; at the same time, the first frequency component is denoised to obtain the first frequency domain denoising feature of the previous time step, and the first frequency domain denoising feature is converted back to the time domain to obtain the first time domain denoising feature of the previous time step. Convert back to the time domain to obtain the first time domain noisy feature of the previous time step; splice the real-time motion sequence part in the first time domain noisy feature of the previous time step with the predicted motion sequence part in the first time domain denoising feature of the previous time step to obtain a spliced ​​sequence of the previous time step, convert the spliced ​​sequence to the frequency domain through the target frequency domain encoder to obtain the frequency domain feature of the previous time step, and extract the first frequency component of the frequency domain feature of the previous time step to obtain an updated first frequency sampling noise, and the updated first frequency sampling noise is used for the next iteration of the first frequency noise prediction network until the number of iterations is equal to a preset number, and the first frequency denoising feature is obtained.

[0087] In this embodiment, if Figure 3 The figure shows a schematic diagram of an iterative update process of adding and denoising the first frequency component and the second frequency component. The noise is sampled from the standard normal distribution and the first frequency noise ∈ is obtained by dividing the noise. l and the second frequency noise ∈ h . The first frequency noise ∈ l , the corresponding first frequency condition input and the current time step t input first frequency noise prediction network The network predicts the first frequency noise ∈ l The noise in the current first frequency noise component ∈ l ′, and according to the current first frequency noise component ∈ l ′ for the first frequency noise ∈ l Perform denoising to obtain the first frequency denoising feature of step t-1 The second frequency noise ∈ h , the corresponding second frequency condition input and the current time step t input second frequency noise prediction network Also predict the second frequency noise ∈ h The noise in the current second frequency noise component ∈ h ′, and according to the current second frequency noise component ∈ h ′ for the second frequency noise ∈ hPerform denoising to obtain the current second frequency denoising feature of step t-1 At the same time, the extracted first frequency component is denoised to directly obtain the first frequency denoised feature of step t-1 The second frequency noise feature of the extracted second frequency component is obtained in step t-1 After each denoising, take the first frequency noise feature of step t-1 The real-time motion sequence part and the first frequency denoising feature of step t-1 The predicted motion sequence part updates the first frequency component to obtain the updated first frequency component for the next round of iterative denoising of the first frequency component; at the same time, the second frequency noise feature of step t-1 is taken The real-time motion sequence part and the second frequency denoising feature of step t-1 The second frequency component is updated by the predicted motion sequence part to obtain an updated second frequency component for the next round of iterative denoising of the second frequency component.

[0088] Step S104: splice the first frequency denoising feature and the second frequency denoising feature and input them into the target feature fusion device for feature fusion to obtain a fused feature, and input the fused feature into the target frequency domain decoder for feature decoding to obtain a target predicted motion sequence corresponding to the real-time motion sequence.

[0089] In this embodiment, after completing a preset number of iterative denoising, the denoised first frequency denoising features and the second frequency denoising features are obtained. These two types of denoising features are spliced ​​in the frequency domain dimension to form a target frequency domain feature that contains the overall motion trend and fine-grained motion changes. Subsequently, the target frequency domain feature is input into a pre-trained target feature fusion device, which integrates and optimizes the first frequency feature and the second frequency feature in the target frequency domain feature through a multi-layer perceptron structure, strengthens the correlation between the features, and finally outputs the fused feature. The fused feature is input into the target frequency domain decoder, and the decoder maps the fused feature of the frequency domain back to the time domain space through an inverse frequency domain conversion operation to generate a target predicted motion sequence containing a real-time motion sequence and a predicted motion sequence.

[0090] Step S105: identifying risky operations according to the target predicted motion sequence, and issuing a warning message when the risky operation occurs.

[0091] In this embodiment, based on the obtained target predicted motion sequence (including the time-series change data of the key points of the human skeleton of power maintenance personnel over a period of time in the future), the position, speed and movement continuity of each joint are analyzed to determine whether there are motion patterns that meet the characteristics of risky operations, such as abnormal distortion of joint trajectories and movements beyond the safe operating range. When the system recognizes the presence of such risky operation characteristics in the predicted sequence, it triggers the early warning mechanism and issues an early warning message through the monitoring terminal, sound and light alarm device, or the mobile terminal of the relevant personnel, prompting on-site management personnel or operation and maintenance personnel to pay timely attention and take intervention measures to avoid safety accidents.

[0092] As an example of an embodiment of the present invention, the first frequency noise prediction network specifically includes: an input projection layer, a feature decoder 1, a feature decoder 2 and an output projection layer connected in sequence; the input projection layer is used to project the frequency domain features of the current time step from the initial dimensional feature space to the target dimensional feature space; the feature decoder 1 includes a Mamba2 module, a first self-attention layer and a first fully connected layer, the Mamba2 module implements selective feature processing based on the structured state space dual framework, the feature decoder 1 injects the information of the current time step into the feature processing process through adaptive layer normalization, for modeling long-term dependency information; the feature decoder 2 includes a second self-attention layer and a second fully connected layer, the feature decoder 2 injects the information of the current time step into the feature processing process through feature linear modulation, for simultaneously modeling long-term dependency information and short-term dependency information; the output projection layer is used to project the frequency domain features output by the feature decoder 2 from the target dimensional feature space to the initial dimensional feature space to obtain a predicted noise component.

[0093] In this embodiment, if Figure 4Figure 1 shows the structure of a noise prediction network. The first frequency noise prediction network consists of an input projection layer, feature decoder 1, feature decoder 2, and an output projection layer, connected in sequence. The input projection layer receives the frequency domain features of the current time step and maps them from the initial dimensional feature space to the target dimensional feature space through a linear transformation, achieving dimensionality conversion to adapt to subsequent processing. Feature decoder 1 consists of a Mamba2 module, a first self-attention layer, and a first fully connected layer. The Mamba2 module, based on the structured state-space duality framework, achieves targeted processing by selectively filtering and enhancing input features. This decoder incorporates temporal information from the current time step into the feature processing pipeline through adaptive layer normalization. Combining the sequence modeling capabilities of the Mamba2 module with the global correlation capture of the first self-attention layer, and the feature integration of the first fully connected layer, it focuses on modeling long-term dependencies between features. Feature decoder 2 consists of a second self-attention layer and a second fully connected layer. It injects information from the current time step into the feature processing through feature linear modulation. The second self-attention layer is responsible for capturing long-term dependencies, while the second fully connected layer strengthens local feature interactions. Together, these two layers simultaneously model both long-term and short-term dependencies. The output projection layer performs a linear transformation on the target dimension feature output by feature decoder 2, projects it from the target dimension feature space back to the initial dimension feature space, and finally outputs the predicted noise component.

[0094] It should be noted that the first frequency noise prediction network and the second frequency noise prediction network have the same structure.

[0095] As an example of an embodiment of the present invention, the target frequency domain encoder, the target feature fuser, and the target frequency domain decoder are respectively obtained by training the historical motion sequence based on the initial frequency domain encoder, the initial feature fuser, and the initial frequency domain decoder, specifically including:

[0096] The historical motion sequence is processed by the initial frequency domain encoder to obtain an initial frequency domain feature, and an initial first frequency component and an initial second frequency component are extracted from the initial frequency domain feature; the initial first frequency component and the initial second frequency component are spliced ​​and input into the initial feature fusion device to obtain an initial fusion feature; and the initial fusion feature is input into the initial frequency domain decoder to obtain an initial motion sequence; the average per-joint position error between the historical motion sequence and the initial motion sequence is determined, and the mean square error between the initial frequency domain feature and the encoding result obtained by discrete cosine transform is determined, and the parameters of the initial frequency domain encoder, the initial feature fusion device and the initial frequency domain decoder are optimized by minimizing the average per-joint position error and the mean square error to obtain the target frequency domain encoder, the target feature fusion device and the target frequency domain decoder respectively.

[0097] In this embodiment, if Figure 5 The figure shows a flow chart of the training steps for a target frequency domain encoder, a target feature fusion unit, and a target frequency domain decoder. After obtaining a historical motion sequence, it is input into the initial frequency domain encoder for encoding, resulting in initial frequency domain features, which are sorted from low to high frequency. The first n frequency components of the initial frequency domain features are extracted as the initial first frequency components, and n frequency components are uniformly selected from the remaining frequency components as the initial second frequency components. These two types of initial historical frequency components are concatenated in the frequency domain and input into the initial feature fusion unit for feature integration, resulting in the initial fused features. The initial fused features are then input into the initial frequency domain decoder, and the initial motion sequence is generated through decoding. The mean per-joint position error (MPJPE) is calculated between the historical motion sequence and the initial motion sequence. The mean squared error (L2 loss) between the initial frequency domain features and the result obtained by encoding the historical motion sequence using the discrete cosine transform (DCT) is also calculated. These two errors are combined to form a loss function (L = average per-joint position error + mean square error). The loss function is minimized through an optimization algorithm, and the network parameters of the initial frequency domain encoder, initial feature fuser, and initial frequency domain decoder are iteratively adjusted until the model converges. Finally, the trained target frequency domain encoder, target feature fuser, and target frequency domain decoder are obtained.

[0098] like Figure 6 As shown, based on the above method embodiment, a corresponding system embodiment is provided;

[0099] An embodiment of the present invention provides a power maintenance risk operation identification system 600, comprising: a prediction module 601, an encoding module 602, a denoising module 603, a decoding module 604 and an identification module 605;

[0100] The prediction module 601 is used to extract human skeleton key point information from the collected power maintenance video data to obtain a real-time motion sequence, and predict the real-time motion sequence through a target predictor to obtain a predicted motion sequence;

[0101] The encoding module 602 splices the real-time motion sequence and the predicted motion sequence to obtain a spliced ​​sequence, performs frequency domain conversion on the spliced ​​sequence using a target frequency domain encoder to obtain frequency domain features, and extracts a first frequency component and a second frequency component from the frequency domain features, wherein the frequency domain features are sorted in ascending order of frequency, the first frequency refers to a preset number of frequency components in the frequency domain features, and the second frequency refers to a preset number of frequency components uniformly selected from frequency components of the frequency domain features other than the first frequency component;

[0102] The denoising module 603 samples frequency domain noise from a standard normal distribution, divides the frequency domain noise into first frequency sampling noise and second frequency sampling noise, inputs the first frequency component as a condition into a first frequency noise prediction network, and inputs the first frequency sampling noise and the current time step into the first frequency noise prediction network, so as to iteratively denoise the first frequency sampling noise through the first frequency noise prediction network. Simultaneously, the second frequency component is input as a condition into a second frequency noise prediction network, and the second frequency sampling noise and the current time step are input into the second frequency noise prediction network, so as to iteratively denoise the second frequency sampling noise through the second frequency noise prediction network, until the number of iterations equals a preset number, then stops the iteration, and obtains a first frequency denoising feature output by the first frequency noise prediction network, and a second frequency denoising feature output by the second frequency noise prediction network, respectively.

[0103] The decoding module 604 is configured to concatenate the first frequency denoising feature and the second frequency denoising feature and input the concatenated feature into a target feature fusion module for feature fusion to obtain a fused feature, and input the fused feature into a target frequency domain decoder for feature decoding to obtain a target predicted motion sequence corresponding to the real-time motion sequence;

[0104] The identification module 605 is used to identify risky operations according to the target predicted motion sequence, and issue a warning message when the risky operation occurs.

[0105] It can be understood that the above-mentioned system embodiment corresponds to the method embodiment of the present invention, which can implement the power maintenance risk operation identification method provided by any of the above-mentioned method embodiments of the present invention.

[0106] It should be noted that the system embodiments described above are merely illustrative, and some or all of the modules may be selected to achieve the objectives of the present embodiments as needed. Furthermore, in the drawings of the system embodiments provided herein, the connection relationships between modules indicate that they have communication connections, which may be implemented as one or more communication buses or signal lines. Persons of ordinary skill in the art can understand and implement the present invention without inventive effort.

[0107] Based on the above-mentioned embodiment of the method for identifying risky operations of electric maintenance, another embodiment of the present invention provides a terminal device, which includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the method for identifying risky operations of electric maintenance of any embodiment of the present invention is implemented.

[0108] For example, in this embodiment, the computer program may be divided into one or more modules, which are stored in the memory and executed by the processor to implement the present invention. The one or more module elements may be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program in the terminal device.

[0109] The terminal device may be a computing device such as a desktop computer, a notebook computer, a PDA, a cloud server, etc. The terminal device may include, but is not limited to, a processor and a memory.

[0110] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the terminal device, connecting various parts of the entire terminal device using various interfaces and lines.

[0111] Based on the above-mentioned method embodiments, another embodiment of the present invention provides a computer-readable storage medium, including a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the power maintenance risk operation identification method described in any one of the above-mentioned method embodiments of the present invention.

[0112] Wherein, the module / unit integrated in the device / terminal equipment, if implemented in the form of a software functional unit and sold or used as an independent product, can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the process in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and when the computer program is executed by the processor, it can implement the steps of the above-mentioned various method embodiments. Wherein, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc.

[0113] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.

Claims

1. A method for identifying risky operations in power maintenance, characterized in that: include: Extracting key point information of the human skeleton from the collected power maintenance video data to obtain a real-time motion sequence, and predicting the real-time motion sequence through a target predictor to obtain an initial predicted motion sequence; splicing the real-time motion sequence with the initial predicted motion sequence to obtain a spliced ​​sequence, performing frequency domain conversion on the spliced ​​sequence using a target frequency domain encoder to obtain frequency domain features, and extracting a first frequency component and a second frequency component from the frequency domain features, wherein the frequency domain features are sorted in ascending order of frequency, the first frequency refers to a preset number of frequency components in the frequency domain features, and the second frequency refers to a preset number of frequency components uniformly selected from frequency components of the frequency domain features other than the first frequency component; Sampling frequency domain noise from a standard normal distribution, dividing the frequency domain noise into first frequency sampling noise and second frequency sampling noise, inputting the first frequency component as a condition into a first frequency noise prediction network, and inputting the first frequency sampling noise and the current time step into the first frequency noise prediction network, so as to iteratively denoise the first frequency sampling noise through the first frequency noise prediction network; at the same time, inputting the second frequency component as a condition into a second frequency noise prediction network, and inputting the second frequency sampling noise and the current time step into the second frequency noise prediction network, so as to iteratively denoise the second frequency sampling noise through the second frequency noise prediction network, until the number of iterations equals a preset number, stopping the iteration, and obtaining a first frequency denoising feature output by the first frequency noise prediction network, and a second frequency denoising feature output by the second frequency noise prediction network, respectively; The first frequency denoising feature and the second frequency denoising feature are spliced ​​and input into a target feature fusion device for feature fusion to obtain a fused feature, and the fused feature is input into a target frequency domain decoder for feature decoding to obtain a target predicted motion sequence corresponding to the real-time motion sequence; Risky operations are identified based on the target predicted motion sequence, and when the risky operations occur, a warning message is issued.

2. The method for identifying risky power maintenance operations according to claim 1, wherein: The iterative denoising of the first frequency sampling noise by the first frequency noise prediction network is specifically performed as follows: The first frequency component in the frequency domain feature is used as a first frequency condition, and is input into the first frequency noise prediction network together with the first frequency sampling noise and the current time step; In each round of iteration of the first frequency noise prediction network, for the current time step, based on the first frequency condition, the current first frequency sampling noise of the current time step is predicted by the first frequency noise prediction network to obtain the current first frequency noise residual corresponding to the current time step, and the current first frequency sampling noise is denoised according to the current first frequency noise residual to obtain the first frequency denoising feature of the previous time step, and the first frequency denoising feature of the previous time step is converted back to the time domain to obtain the first time domain denoising feature of the previous time step; at the same time, the first frequency component is denoised to obtain the first frequency domain denoising feature of the previous time step, and the first frequency domain denoising feature is converted back to the time domain to obtain the first time domain denoising feature of the previous time step; at the same time, the first frequency component is denoised to obtain the first frequency domain denoising feature of the previous time step, and the first frequency domain denoising feature is converted back to the time domain to obtain the first time domain denoising feature of the previous time step. Convert back to the time domain to obtain the first time domain noisy feature of the previous time step; splice the real-time motion sequence part in the first time domain noisy feature of the previous time step with the predicted motion sequence part in the first time domain denoising feature of the previous time step to obtain a spliced ​​sequence of the previous time step, convert the spliced ​​sequence to the frequency domain through the target frequency domain encoder to obtain the frequency domain feature of the previous time step, and extract the first frequency component of the frequency domain feature of the previous time step to obtain an updated first frequency sampling noise, and the updated first frequency sampling noise is used for the next iteration of the first frequency noise prediction network until the number of iterations is equal to a preset number, and the first frequency denoising feature is obtained.

3. The method for identifying risky power maintenance operations according to claim 1, wherein: The human skeleton key point information is extracted from the collected power maintenance video data to obtain a real-time motion sequence, specifically: Extracting video frames from each power maintenance video data using a video processing tool, and adjusting the frame rate of each extracted video frame to a preset frequency to obtain multiple video frames; Each video frame is processed using a 3D human pose estimation model to extract skeleton key point information. The corresponding 2D array for each video frame is then determined based on the number of joints and the 3D coordinate dimensions of each joint. The two-dimensional arrays corresponding to all video frames are sorted in chronological order to obtain a key point sequence, and the key point sequence is converted into a time series data tensor according to a four-dimensional structure of batch, time step, joint point and coordinate dimensions, and the time series data tensor is standardized to obtain the real-time motion sequence.

4. The method for identifying risky power maintenance operations according to claim 3, wherein: The normalization process of the time series data tensor to obtain the real-time motion sequence is specifically as follows: Taking a preset central key point as a spatial reference benchmark, calculating the offsets of all relevant nodes in the time series data tensor relative to the central key point to perform centralization processing and obtain a centralized tensor; For the three-dimensional coordinates of all relevant nodes in the centralized tensor, calculate the scale distribution characteristics of all three-dimensional coordinates, determine a target scale range based on the scale distribution characteristics, and map all three-dimensional coordinates to the target scale range to obtain a first time series data tensor; A sliding average algorithm or a filtering algorithm is used to perform time series smoothing processing on the key point trajectories in the first time series data tensor to obtain the real-time motion sequence.

5. The method for identifying risky power maintenance operations according to claim 1, wherein: The target predictor is obtained by training the motion sequence prediction of historical power maintenance video data based on the initial predictor, wherein the initial predictor adopts a multi-layer perceptron structure, specifically including: Acquire a historical motion sequence corresponding to the historical power maintenance video data, and divide the historical motion sequence according to time sequence to obtain a continuous first motion sequence and a second motion sequence; The first motion sequence is used as the input of the initial predictor to obtain a first prediction sequence, and the second motion sequence is used as the expected output of the initial predictor, an error between the first prediction sequence and the second motion sequence is determined by a joint loss function, and the initial predictor is trained by minimizing the error to obtain the target predictor, wherein the joint loss function includes an average per-joint position error and an average per-joint velocity error.

6. The method for identifying risky power maintenance operations according to claim 5, wherein: The target frequency domain encoder, target feature fusion device and target frequency domain decoder are respectively obtained by training the historical motion sequence according to the initial frequency domain encoder, the initial feature fusion device and the initial frequency domain decoder, and specifically include: Processing the historical motion sequence by the initial frequency domain encoder to obtain initial frequency domain features, and extracting an initial first frequency component and an initial second frequency component from the initial frequency domain features; splicing the initial first frequency component and the initial second frequency component and inputting them into the initial feature fusion device to obtain an initial fusion feature; and inputting the initial fusion feature into the initial frequency domain decoder to obtain an initial motion sequence; Determine the average per-joint position error between the historical motion sequence and the initial motion sequence, and determine the mean square error between the initial frequency domain feature and the encoding result obtained by discrete cosine transform. By minimizing the average per-joint position error and the mean square error, the parameters of the initial frequency domain encoder, the initial feature fuser and the initial frequency domain decoder are optimized to obtain the target frequency domain encoder, the target feature fuser and the target frequency domain decoder respectively.

7. The method for identifying risky power maintenance operations according to claim 1, wherein: The first frequency noise prediction network specifically includes: an input projection layer, a feature decoder 1, a feature decoder 2, and an output projection layer connected in sequence; The input projection layer is used to project the frequency domain features of the current time step from the initial dimensional feature space to the target dimensional feature space; The feature decoder 1 includes a Mamba2 module, a first self-attention layer, and a first fully connected layer. The Mamba2 module implements selective feature processing based on a structured state-space dual framework. The feature decoder 1 injects information of the current time step into the feature processing process through adaptive layer normalization to model long-term dependency information. The second feature decoder includes a second self-attention layer and a second fully connected layer. The second feature decoder injects the information of the current time step into the feature processing process through feature linear modulation to simultaneously model long-term dependency information and short-term dependency information; The output projection layer is used to project the frequency domain features output by the second feature decoder from the target dimensional feature space to the initial dimensional feature space to obtain a predicted noise component.

8. A power maintenance risk operation identification system, characterized by: include: Prediction module, encoding module, denoising module, decoding module and recognition module; The prediction module is used to extract human skeleton key point information from the collected power maintenance video data to obtain a real-time motion sequence, and predict the real-time motion sequence through a target predictor to obtain an initial predicted motion sequence; The encoding module splices the real-time motion sequence with the initial predicted motion sequence to obtain a spliced ​​sequence, performs frequency domain conversion on the spliced ​​sequence through a target frequency domain encoder to obtain frequency domain features, and extracts a first frequency component and a second frequency component from the frequency domain features, wherein the frequency domain features are sorted in ascending order of frequency, the first frequency refers to a preset number of frequency components in the frequency domain features, and the second frequency refers to a preset number of frequency components uniformly selected from the frequency components of the frequency domain features other than the first frequency component; The denoising module is configured to generate frequency domain noise by sampling from a standard normal distribution, divide the frequency domain noise into first frequency sampling noise and second frequency sampling noise, input the first frequency component as a condition into a first frequency noise prediction network, and input the first frequency sampling noise and the current time step into the first frequency noise prediction network, so as to iteratively denoise the first frequency sampling noise through the first frequency noise prediction network; at the same time, input the second frequency component as a condition into a second frequency noise prediction network, and input the second frequency sampling noise and the current time step into the second frequency noise prediction network, so as to iteratively denoise the second frequency sampling noise through the second frequency noise prediction network, until the number of iterations equals a preset number, then stop the iteration, and obtain a first frequency denoising feature output by the first frequency noise prediction network, and a second frequency denoising feature output by the second frequency noise prediction network, respectively; The decoding module is used to splice the first frequency denoising feature and the second frequency denoising feature and input them into the target feature fusion device for feature fusion to obtain a fused feature, and input the fused feature into the target frequency domain decoder for feature decoding to obtain a target predicted motion sequence corresponding to the real-time motion sequence; The identification module is used to identify risky operations based on the target predicted motion sequence, and issue a warning message when the risky operation occurs.

9. A terminal device, characterized in that: The method comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the method for identifying electric power maintenance risk operations according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium, characterized in that include: A stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the power maintenance risk operation identification method according to any one of claims 1 to 7.