Distribution network operation monitoring data processing method, system and equipment and medium

Through multi-modal large model processing of distribution network operation monitoring data, the problems of inefficiency and high danger in the existing technology are solved, and efficient and safe live-operated operation monitoring is achieved.

CN120279462APending Publication Date: 2025-07-08SICHUAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510415703.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The live operations of the existing distribution network mainly rely on manual inspection and high-altitude operations, which are inefficient and have high risk factors, and cannot meet the needs of high-quality live operations.

Method used

The multimodal large model is used to process the distribution network operation monitoring data. By obtaining on-site video monitoring data, feature extraction and interactive processing are performed, the processing results of the data set are generated, including video question-and-answer and classification modes, and the multimodal large model is used for data analysis.

Benefits of technology

It improves the efficiency of distribution network operation monitoring, reduces the risk coefficient and quality uncontrollable phenomenon, and meets the needs of high-quality live operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279462A_ABST
    Figure CN120279462A_ABST
Patent Text Reader

Abstract

The invention relates to a distribution network operation monitoring data processing method, system and device and a medium, and the method comprises the steps: obtaining a plurality of video monitoring data of a first visual angle of a distribution network operation site and a model execution mode of a pre-trained multi-mode large model, and carrying out the processing of the distribution network operation monitoring data according to the plurality of video monitoring data; according to the method, the data set of the first view angle of the distribution network operation site is generated according to the model execution mode, the data set is processed through the multi-mode large model according to the model execution mode, the processing result corresponding to the data set can be effectively obtained, the efficiency of distribution network operation monitoring is improved, and the user experience is improved. And the danger coefficient and the quality uncontrollable phenomenon of the distribution network operation are effectively reduced, so that the requirement of the current high-quality hot-line operation of a large number of distribution networks is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of distribution network operation, and in particular to a method, system, equipment and medium for processing distribution network operation monitoring data. Background Art

[0002] In the power system, live working in the distribution network plays a vital role. The increase in the mileage of the distribution network transmission lines has led to frequent circuit damage and component shedding due to long-term exposure to the outdoor environment, which seriously threatens the safe operation of the distribution network transmission lines. At present, live working on transmission lines mainly adopts manual inspection and high-altitude working methods. This method is inefficient, the quality of work is uncontrollable, and the risk factor is high. It cannot meet the current needs of high-quality live working in a large number of distribution networks. Summary of the invention

[0003] In order to solve the problems existing in the prior art, the present invention provides a method, system, device and medium for processing distribution network operation monitoring data.

[0004] A first aspect of the present invention provides a method for processing distribution network operation monitoring data, the method comprising:

[0005] Obtain multiple video monitoring data from the first-person perspective of the distribution network operation site and the model execution mode of the pre-trained multi-modal large model;

[0006] Generate a data set of the first-person perspective of the distribution network operation site according to the multiple video monitoring data;

[0007] According to the model execution mode, the data set is processed by the multimodal large model to obtain a processing result corresponding to the data set.

[0008] Optionally, the data set includes labeling data of each to-be-used video monitoring data, and the data set is processed by the multimodal large model according to the model execution mode to obtain a processing result corresponding to the data set, including:

[0009] When the model execution mode is the video question-answering mode, the multimodal large model interactively processes multiple video monitoring data and multiple annotated data in the data set to obtain an answer result corresponding to each annotated data.

[0010] Optionally, the interactive processing of the multiple video monitoring data and the multiple annotated data in the data set by the multimodal large model to obtain an answer result corresponding to each annotated data includes:

[0011] Extracting features from the plurality of video monitoring data using the multimodal large model to obtain a first spatiotemporal feature of each video monitoring data;

[0012] Interactively process the multiple first spatiotemporal features and the multiple annotated data to obtain an answer result corresponding to each annotated data.

[0013] Optionally, the processing of the data set by the multimodal large model according to the model execution mode to obtain a processing result corresponding to the data set includes:

[0014] When the model execution mode is a video classification mode, feature extraction is performed on the plurality of video monitoring data by using the multimodal large model to obtain a second spatiotemporal feature of each video monitoring data;

[0015] According to the plurality of second spatiotemporal features, the plurality of video monitoring data in the data set are classified and processed to generate a classification result corresponding to each of the plurality of video monitoring data.

[0016] Optionally, the training process of the multimodal large model includes:

[0017] Acquire multiple standby video data of the first-person perspective of the distribution network operation site, and standby annotation data of each video data;

[0018] Generate a training set of the multimodal large model according to the plurality of unused video data and the unused annotation data of each video data;

[0019] The training set is input into an initial multimodal large model, and the initial multimodal large model is trained to obtain the multimodal large model.

[0020] Optionally, the initial modal large model includes multiple convolution kernels, multimodal rotation position embedding and a cross-modal attention mechanism, and the training set is input into the initial multimodal large model, the initial multimodal large model is trained, and the multimodal large model is obtained, including:

[0021] Extracting features of multiple unused video data in the training set by using the multiple convolution kernels to obtain unused spatiotemporal features corresponding to each unused video data;

[0022] Encoding a plurality of standby spatiotemporal features by means of the multimodal rotation position embedding to obtain a plurality of encoded standby spatiotemporal features;

[0023] The cross-modal attention mechanism is used to interact the encoded multiple unused spatiotemporal features with the multiple unused labeled data in the training set to obtain the multimodal large model.

[0024] Optionally, generating a data set of the first-person perspective of the distribution network operation site according to the plurality of video monitoring data includes:

[0025] Performing image enhancement, video parameter adjustment and denoising processing on the plurality of video monitoring data to obtain a plurality of stand-by video monitoring data;

[0026] Obtaining labeling data of each of the plurality of video monitoring data to be used;

[0027] A data set of the first-person perspective of the distribution network operation site is generated according to the multiple stand-by video monitoring data and the annotation data of each stand-by video monitoring data.

[0028] A second aspect of the present invention provides a system for processing distribution network operation monitoring data, the system comprising:

[0029] An acquisition module is configured to acquire a plurality of video monitoring data of a first-person perspective of a distribution network operation site and a model execution mode of a pre-trained multi-modal large model;

[0030] A generating module is configured to generate a data set of the first-person perspective of the distribution network operation site according to the plurality of video monitoring data;

[0031] The determination module is configured to process the data set through the multimodal large model according to the model execution mode to obtain a processing result corresponding to the data set.

[0032] Optionally, the data set includes labeling data of each to-be-used video monitoring data, and the determination module is configured to:

[0033] When the model execution mode is the video question-answering mode, the multimodal large model interactively processes multiple video monitoring data and multiple annotated data in the data set to obtain an answer result corresponding to each annotated data.

[0034] Optionally, the determining module is configured to:

[0035] Extracting features from the plurality of video monitoring data using the multimodal large model to obtain a first spatiotemporal feature of each video monitoring data;

[0036] Interactively process the multiple first spatiotemporal features and the multiple annotated data to obtain an answer result corresponding to each annotated data.

[0037] Optionally, the determining module is configured to:

[0038] When the model execution mode is a video classification mode, feature extraction is performed on the plurality of video monitoring data by using the multimodal large model to obtain a second spatiotemporal feature of each video monitoring data;

[0039] According to the plurality of second spatiotemporal features, the plurality of video monitoring data in the data set are classified and processed to generate a classification result corresponding to each of the plurality of video monitoring data.

[0040] Optionally, the training process of the multimodal large model includes:

[0041] Acquire multiple standby video data of the first-person perspective of the distribution network operation site, and standby annotation data of each video data;

[0042] Generate a training set of the multimodal large model according to the plurality of unused video data and the unused annotation data of each video data;

[0043] The training set is input into an initial multimodal large model, and the initial multimodal large model is trained to obtain the multimodal large model.

[0044] Optionally, the initial modal large model includes multiple convolution kernels, multimodal rotation position embedding and a cross-modal attention mechanism, and the training set is input into the initial multimodal large model, the initial multimodal large model is trained, and the multimodal large model is obtained, including:

[0045] Extracting features of multiple unused video data in the training set by using the multiple convolution kernels to obtain unused spatiotemporal features corresponding to each unused video data;

[0046] Encoding a plurality of standby spatiotemporal features by means of the multimodal rotation position embedding to obtain a plurality of encoded standby spatiotemporal features;

[0047] The cross-modal attention mechanism is used to interact the encoded multiple unused spatiotemporal features with the multiple unused labeled data in the training set to obtain the multimodal large model.

[0048] Optionally, the generating module is configured to:

[0049] Performing image enhancement, video parameter adjustment and denoising processing on the plurality of video monitoring data to obtain a plurality of stand-by video monitoring data;

[0050] Obtaining labeling data of each of the plurality of video monitoring data to be used;

[0051] A data set of the first-person perspective of the distribution network operation site is generated according to the multiple stand-by video monitoring data and the annotation data of each stand-by video monitoring data.

[0052] A third aspect of the present invention provides a computer device, comprising: one or more processors;

[0053] The processor is used to store one or more programs;

[0054] When the one or more programs are executed by the one or more processors, a method for processing distribution network operation supervision data described in the first aspect of the present invention above is implemented.

[0055] In the fourth aspect of the present invention, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed, a method for processing distribution network operation supervision data described in the first aspect of the present invention above is implemented.

[0056] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0057] The present invention provides a method, a system, a device and a medium for processing distribution network operation supervision data. The method for processing distribution network operation supervision data obtains a plurality of video supervision data from the first perspective of the distribution network operation site and the model execution mode of a pre-trained multi-modal large model. According to the plurality of video supervision data, a data set from the first perspective of the distribution network operation site is generated. According to the model execution mode, the multi-modal large model processes the data set, and can effectively obtain the processing result corresponding to the data set, improve the efficiency of distribution network operation supervision, effectively reduce the risk coefficient and the phenomenon of uncontrollable quality of distribution network operation, so as to meet the current needs of a large number of high-quality live distribution network operations. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 It is a flowchart of a method for processing distribution network operation supervision data provided by the present invention;

[0059] Figure 2 It is a flowchart of a method for processing distribution network operation supervision data provided by the present invention;

[0060] Figure 3 It is a flowchart of a training process of a multi-modal large model provided by the present disclosure;

[0061] Figure 4 It is a block diagram of a system for processing distribution network operation supervision data provided by the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0062] Example 1:

[0063] Figure 1 It is a flowchart of a method for processing distribution network operation supervision data provided by the present invention. As Figure 1 shown, the method for processing distribution network operation supervision data may include the following steps:

[0064] In step 101, a plurality of video monitoring data of the first-person perspective of the distribution network operation site and a model execution mode of a pre-trained multimodal large model are obtained.

[0065] Among them, the video data of the operation site is collected in real time at a preset resolution and frame rate through the first-person perspective acquisition device, and the data is sent to the back-end processing system through the data transmission module.

[0066] It should be noted that the original long video data set is segmented according to the video data collected by the first-person perspective acquisition device, and is segmented into multiple video monitoring data of preset duration. For example: use a video processing tool (such as FFmpeg) to segment the collected video data, set the segmentation time to 3 to 5 seconds, and pre-process the segmented video data, including video format conversion, resolution adjustment, etc., to ensure the consistency and availability of the video data.

[0067] Optionally, the segmented multiple video monitoring data are preprocessed, including video format conversion, resolution adjustment, etc.

[0068] In step 102, a data set of the first-person perspective of the distribution network operation site is generated based on the multiple video monitoring data.

[0069] A possible implementation of this step may include: performing image enhancement, video parameter adjustment and denoising processing on the multiple video monitoring data to obtain multiple stand-by video monitoring data; obtaining annotation data of each stand-by video monitoring data in the multiple stand-by video monitoring data; and generating a data set of the first-person perspective of the distribution network operation site based on the multiple stand-by video monitoring data and the annotation data of each stand-by video monitoring data.

[0070] The video parameters include brightness, contrast, color saturation, etc. The dataset is a JSON array, each element represents a piece of data, including multiple rounds of conversation messages and video data paths. Conversation messages include message content and message roles.

[0071] It should be noted that the image enhancement processing is performed on the collected multiple video monitoring data, and the visibility of the video under different lighting conditions is improved by adjusting parameters such as brightness, contrast and color saturation. The video data is denoised using a filtering algorithm to remove noise points in the video and further improve the clarity of the video. The frame rate of the video can also be adjusted according to actual needs to ensure the smoothness and stability of the video data. In order to improve the efficiency of data transmission and storage, the video data can also be compressed, and an efficient compression algorithm can be used to reduce the amount of data while trying to maintain the quality of the video. Through these preprocessing steps, the quality of the video data is significantly improved, providing a solid foundation for subsequent analysis and application.

[0072] One possible implementation of this step is to classify the multiple video monitoring data obtained by segmentation according to the corresponding distribution network operation standard, uniformly name the same type of video data, and place the same type of data sets in the same folder. For example: formulate video classification rules according to the distribution network operation standard, and classify multiple video monitoring data according to the video classification rules to ensure that each video data belongs to the correct category. Uniformly name the same type of video data, such as a_1, a_2, ... a_n, and so on. Put the same type of data sets in the same folder for easy management and use.

[0073] In step 103, the data set is processed by the multimodal large model according to the model execution mode to obtain a processing result corresponding to the data set.

[0074] The data set includes the labeling data of each video monitoring data to be used.

[0075] A possible implementation of this step may include: when the model execution mode is a video question-answering mode, the multimodal large model interactively processes multiple video monitoring data and multiple annotated data in the data set to obtain an answer result corresponding to each annotated data.

[0076] Among them, the specific implementation method of the above implementation method can be to extract features of the multiple video monitoring data through the multimodal large model to obtain the first spatiotemporal features of each video monitoring data; interactively process the multiple first spatiotemporal features and the multiple labeled data to obtain the answer result corresponding to each labeled data.

[0077] It should be noted that the fused multimodal features are further processed according to the specific application task (i.e., the model execution mode). In the video question-answering task (i.e., the video question-answering mode), the multimodal features can interact with the question text to generate answers. The multimodal features and the question text can be input into a multi-layer Transformer architecture, and the text representation of the answer can be generated through the self-attention mechanism and the cross-modal attention mechanism. Then, the text representation is converted into the probability distribution of the answer through a linear layer to obtain the final answer. In the video classification task, the multimodal features can be classified to obtain the category label of the video.

[0078] For example, in the video question answering task, the fused multimodal features are first represented as F = {f1,f2,...,f n}, where f i is the ith feature vector, and n is the total number of features. After preprocessing (such as word segmentation, word embedding, etc.), the question text is represented as T = {t1, t2, ..., t m}, t j is the embedding vector of the j-th word, and m is the number of words in the text. The multi-modal feature F and the question text T are input into a multi-layer Transformer architecture. The calculation formula of the self-attention mechanism in the Transformer architecture is as follows:

[0079]

[0080] In the formula, Q, K, and V are the Query, Key, and Value matrices respectively, and d k is the dimension of the key. During the calculation process, the expression ability of the model is further enhanced through the multi-head attention mechanism. The multi-head attention mechanism divides Q, K, and V into multiple heads, performs attention calculations separately, and then splices them together. Suppose there are h heads head h , and the dimension of each head is Then the output of the multi-head attention mechanism is:

[0081] MultiHead(Q, K, V) = Concat(head1,..., head h )W O

[0082] Among them, the calculation of each head is:

[0083]

[0084] In the formula, h is the number of attention heads, usually set to 8 or 16, W i Q , is the learnable weight matrix independent for each head, used to generate representations in different subspaces, and W O is the weight matrix of the output layer, which maps the spliced multi-head results back to the original dimension. The cross-modal attention mechanism is used to enable interaction between the multi-modal feature and the question text. Its calculation method is similar to the self-attention mechanism, except that the query matrix comes from one modality, while the key and value matrices come from another modality. Taking the question text as the query and the multi-modal feature as the key and value for cross-modal attention calculation, the feature representation after interaction is obtained, and then through multiple layers of such Transformer structures, the feature representation is continuously updated, and finally the text representation A of the answer is generated.

[0085] The above scheme obtains multiple video monitoring data of the first-person perspective of the distribution network operation site and the model execution mode of the pre-trained multimodal large model, and generates a data set of the first-person perspective of the distribution network operation site according to the multiple video monitoring data. According to the model execution mode, the data set is processed by the multimodal large model to effectively obtain the processing result corresponding to the data set, thereby improving the efficiency of distribution network operation monitoring, effectively reducing the risk factor and quality uncontrollable phenomenon of distribution network operations, thereby meeting the current demand for high-quality live operations in a large number of distribution networks.

[0086] Figure 2 A flowchart of a method for processing distribution network operation monitoring data provided by the present invention, such as Figure 2 As shown above Figure 2 Another possible implementation of step 103 may include the following steps:

[0087] In step 1031, when the model execution mode is the video classification mode, the multimodal large model is used to perform feature extraction on the multiple video monitoring data to obtain the second spatiotemporal features of each video monitoring data.

[0088] In step 1032, the multiple video monitoring data in the data set are classified according to the multiple second spatiotemporal features to generate a classification result corresponding to each video monitoring data in the multiple video monitoring data.

[0089] It should be noted that in the video classification task, the category label of the video is output. These results can be presented to the user in the form of text, image or voice, so as to achieve effective understanding of the video content.

[0090] Figure 3 A flowchart of the training process of a multimodal large model provided by the present disclosure, such as Figure 3 As shown, the training process of the multimodal large model may include the following steps:

[0091] In step S1, a plurality of standby video data of a first-view angle of a distribution network operation site and standby annotation data of each video data are obtained.

[0092] In step S2, a training set of the multimodal large model is generated according to the plurality of unused video data and the unused annotation data of each video data.

[0093] In step S3, the training set is input into an initial multimodal large model, and the initial multimodal large model is trained to obtain the multimodal large model.

[0094] Among them, the initial modal large model includes multiple convolution kernels, multimodal rotation position embedding and cross-modal attention mechanism.

[0095] A possible implementation method of the above step S3 may be: extracting features from the multiple stand-by video data in the training set through the multiple convolution kernels to obtain stand-by spatiotemporal features corresponding to each stand-by video data; encoding the multiple stand-by spatiotemporal features through the multimodal rotation position embedding to obtain the encoded multiple stand-by spatiotemporal features; interacting the encoded multiple stand-by spatiotemporal features with the multiple stand-by annotated data in the training set through the cross-modal attention mechanism to obtain the multimodal large model.

[0096] Example: Above Figure 3 Specific implementations of the illustrated scheme may include:

[0097] Video data preprocessing: First, the input first-person video data is frame sampled and the resolution is adjusted. Specifically, the video can be sampled at two frames per second, and then the resolution of the video frame can be dynamically adjusted according to the original resolution of the video frame and the processing capability of the model. For example, for a high-resolution video frame, it can be adjusted to a lower resolution to reduce the amount of calculation; for a low-resolution video frame, it can be kept at the original resolution or appropriately enlarged to retain more detail information.

[0098] ② Visual feature extraction: Use a visual encoder to extract features from the video frames after adjusting the resolution. In the first-person view video data, ViT is used as the visual encoder. The video frames are divided into multiple patches, each with a size of 14×14, and then each patch is embedded into a 768-dimensional vector space. Specifically, the input image (such as an RGB image of size H×W) is divided into multiple non-overlapping small blocks (patches), and the size of the small blocks is usually set to P×P. The size of each P×P block is flattened into a one-dimensional vector of size P2×C, where C is the number of channels of each image block (such as the three RGB channels). Then, these vectors are mapped to the dimension required by the model (usually the same as the dimension of the hidden state in the Transformer model, such as 768 dimensions) through a linear layer (projection layer). Then, these embedded vectors are processed through a multi-layer Transformer architecture to extract the visual features of the video frames. Use a standard Transformer encoder, which includes a multi-head self-attention (MSA) and a multi-layer perceptron (MLP) block. Each Transformer layer contains an MSA module and an MLP module, which are used to capture global dependencies and perform feature transformation respectively. An additional learnable "classification token" ([class]token) is added to the sequence output by the Transformer, and its output is used as the image representation, and then the final classification result is output through a classification head (usually a fully connected layer). The embedded vectors are input into a multi-layer Transformer architecture for processing to extract the visual features of the video frames. The Transformer architecture captures global dependencies through the self-attention mechanism and performs feature transformation through the MLP.

[0099] In addition, to better capture the spatio-temporal features in the video, 3D convolution can be used to process the video frames. Specifically, two-layer deep 3D convolution can be used, with a kernel size of 3×3×3 and a stride of 1, so that the video frames are processed as three-dimensional tubes, retaining more spatio-temporal information. Specifically, taking the video frames as the input, each video frame can be regarded as a three-dimensional tensor with dimensions T×H×W×C, where T is the time dimension, H and W are the spatial dimensions, and C is the number of channels. Use a 3D convolution kernel to perform a convolution operation on the input video frames. The kernel size is 3×3×3 and the stride is 1, indicating that the convolution is performed with a stride of 1 in the three dimensions of time, height, and width. The 3D convolution operation can be expressed as:

[0100]

[0101] where y t,h,w,c is the value of the output feature map at position (t, h, w) and channel c, and x t,h,w,cis the value of the input video frame at position (t, h, w) and channel c. is the weight of the convolutional kernel at position (k t , k h , k w ) and channel c, where k t , k h and k w are the sizes of the convolutional kernel in the time, height, and width dimensions. After a 3D convolutional operation, an activation function such as ReLU is usually applied to introduce non-linearity and enhance the model's expressive power. To reduce the size of the feature map and computational complexity, a pooling operation such as max pooling or average pooling can be performed after the 3D convolution. The pooling operation can be carried out in the three dimensions of time, height, and width to further extract features.

[0102] ③ Multimodal fusion: Multimodal Rotary Position Embedding (M-RoPE) is used to encode the spatio-temporal position information of video frames. Specifically, the traditional rotary position embedding is decomposed into three parts: time, height, and width, and the time ID, height ID, and width ID of the video frame are encoded respectively. For example, in the live working business scenario, the video data obtained by the first-person high-definition audio-visual synchronous camera needs to be processed efficiently. To fuse the time and space information of the video data, for a video frame, its time ID can be assigned according to the position of the frame in the video, and the height ID and width ID can be assigned according to the position of the patch in the video frame. Specifically, in the original Transformer self-attention mechanism, the position encoding is usually represented by sine and cosine functions. RoPE transforms the absolute position encoding into relative position encoding by introducing a rotation operation. Specifically, for the query vector q and the key vector k, the position encoding operation of RoPE can be expressed as:

[0103] f(q, m) = q · R m

[0104] f(k, n) = k · R n

[0105] where R m and R n are rotation matrices used to encode the position information into vectors. The rotation matrix R m can be expressed as:

[0106]

[0107] Where m and n are position indices, and θ is the frequency parameter. M-ROPE extends the rotary position encoding of RoPE to multimodal data by introducing multiple rotation matrices to process data in different dimensions. Specifically, for a one-dimensional text sequence, a two-dimensional visual image, and a three-dimensional video, M-ROPE introduces rotation matrices for time, height, and width respectively to capture and integrate the position information in multimodal data. For example, for three-dimensional video data, M-ROPE can be expressed as:

[0108] f(q, m, h, w) = q · R m · R h · R w

[0109] where R m 、R h and R w are the rotation matrices for time, height, and width respectively. In this way, the spatio-temporal position information of video frames can be explicitly modeled, so as to better fuse the spatio-temporal information and semantic information of the video. In the multimodal model, a cross-modal attention mechanism is used to realize the interaction between video features and text features. Specifically, the video features and text features can be input into a multi-layer Transformer architecture, and through the cross-modal attention mechanism, the model can simultaneously focus on the visual information in the video and the semantic information in the text.

[0110] All or part of the various modules in the first-person perspective real-time auxiliary equipment for live working on the distribution network provided in the above Embodiment 1 can be implemented through software, hardware, and their combinations. Each module of this technology can be embedded in the processor of a computer device in hardware form or be independent of it, or be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each module. Specifically, the module composition logic of this equipment is as follows: As shown in the figure, the above modules form a self-organizing network system, and realize efficient communication and data transmission between devices through self-organizing network technology. During the operation, the first-person perspective high-definition audio-visual synchronous camera device installed on the safety helmet acquires the first-person perspective audio-visual data, and transmits the data to the portable edge computing terminal through the signal relay transmission device. After receiving the data, the portable edge computing terminal uses the multimodal large model for real-time analysis and processing to provide real-time auxiliary decision-making support for the operators.

[0111] Specifically, the composition of each module is as follows:

[0112] ① Portable Edge Computing Terminal: Used to run multi-modal large models, it has powerful computing capabilities and portability, and can efficiently process the collected multi-modal data. The terminal includes a processor, a memory, an input / output interface, and a communication interface. The processor is a graphics processing unit (GPU), which is specifically used for AI model deployment and has powerful parallel computing capabilities to efficiently handle complex computing tasks in multi-modal large models. The specific component composition is as follows:

[0113] Processor: The processor is a graphics processing unit (GPU), which is specifically used for AI model deployment. The GPU has powerful parallel computing capabilities and can efficiently handle complex computing tasks in multi-modal large models. The GPU can be NVIDIA's Tesla series, A100, or other high-performance graphics processing units suitable for AI model deployment.

[0114] Memory: It includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores the operating system, computer programs, and databases. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium.

[0115] Input / Output Interface: Used to exchange information between the processor and external devices. For example, devices such as cameras, microphones, and displays can be connected through the input / output interface to obtain and output multi-modal data.

[0116] Communication Interface: Used to communicate with external terminals through a network connection. For example, data can be transmitted with external devices through wireless communication technologies such as Wi-Fi, Bluetooth, 4G / 5G, etc. The terminal is connected to the first-person view high-definition camera through a self-organizing network to obtain video data. Specifically, the communication interface of the terminal establishes a connection with the first-person view high-definition camera through self-organizing network technology to obtain video data in real time. The self-organizing network technology enables the terminal and the camera to automatically establish and maintain a network connection without relying on external network infrastructure, thus ensuring the stability and real-time nature of video data transmission.

[0117] ② First-person View High-definition Audio-Visual Synchronized Camera Device: Installed on a safety helmet, it is used to obtain the first-person view audio-visual data of the operator, has high-definition video and audio collection functions, and can collect video and audio data of the operation site in real time. The specific component composition is as follows:

[0118] Lens: It uses a micro gimbal mechanical anti-shake lens, which can effectively improve the stability of the picture, supports 5 million pixels and a 120-degree wide angle, making the picture clearer and wider. The micro gimbal mechanical anti-shake technology can effectively reduce the picture blur caused by device shaking during operation, ensuring that the collected video data is clear and stable. The 120-degree wide-angle lens can cover a wider field of view, capture more details of the operation site, and provide more comprehensive information for subsequent data analysis.

[0119] Microphone: Two high-sensitivity microphones are built-in, one main and one secondary. The main microphone is used to pick up human voices, and the secondary microphone is used to pick up ambient sounds. Combined with an intelligent noise reduction chip, it can effectively achieve environmental noise reduction, improve the clarity of human voices, and at the same time support echo cancellation. The high-sensitivity microphones can clearly capture the voice commands of the operators and the on-site environmental sounds, and the intelligent noise reduction chip can effectively filter out background noise to ensure the clarity and accuracy of audio data.

[0120] Video Encoding Module: The maximum supported resolution is 1920×1080, the frame rate is 30 frames, and H.264 encoding is used. H.264 encoding is an efficient video compression standard that can greatly reduce the storage space and transmission bandwidth requirements of video data while ensuring video quality. The resolution of 1920×1080 and the frame rate of 30 frames can ensure the clarity and smoothness of video data, meeting the high-quality requirements for video data in live distribution network energized operation sites.

[0121] Audio Encoding Module: Supports encoding formats such as 64Kbps (G.711), 16Kbps (G.726), 16 - 64Kbps (AAC), etc. These encoding formats can meet the audio data transmission requirements in different scenarios, ensuring the clarity and transmission efficiency of audio data. G.711 encoding is suitable for high-quality audio transmission, G.726 encoding is suitable for audio transmission in low-bandwidth environments, and AAC encoding has a high compression efficiency while ensuring audio quality.

[0122] Storage Module: Built-in 32G storage, enabling video recording for no less than 15 hours. The large-capacity storage can meet the needs of long-term operations, ensuring that the video data collected during the operation will not be lost due to insufficient storage space. The 15-hour video recording ability can cover most of the duration of live distribution network energized operations, providing sufficient materials for subsequent data analysis and playback.

[0123] Network Module: Supports WiFi network, Bluetooth 4.2, supports the full Netcom of Mobile / Unicom / Telecom 4G, and supports the wireless network of MESH self-organizing network. Among them, for the antenna part, the antenna factory will do special adaptation and optimization to ensure that the wireless signal of the device reaches the best. The support for multiple networks can ensure the network connection stability of the device in different environments. The MESH self-organizing network function can realize the self-organizing communication between devices, ensuring the real-time and reliable data transmission. The special adaptation and optimization of the antenna can further improve the strength and stability of the wireless signal, ensuring the efficiency of data transmission.

[0124] Positioning module: It supports GPS positioning and can display the real-time position in the form of a map on the software. The GPS positioning function can obtain the position information of the device in real time. Through the map display of the software, operators and managers can clearly understand the position distribution of the operation site, which is convenient for operation scheduling and safety management.

[0125] Battery: It is equipped with a polymer lithium battery with a capacity of 2000 mA / H, which can achieve a continuous video recording time of not less than 4 hours and a standby time of not less than 12 hours. The large-capacity battery can meet the needs of long-term operation and ensure that the device will not be interrupted due to insufficient power during operation. The 4-hour video recording endurance and 12-hour standby endurance can meet the duration requirements of most live distribution network operations and provide guarantee for the continuity of operations.

[0126] Indicator lights: It includes a power indicator light, a network indicator light, a video recording indicator light, and a laser light. The power indicator light is used to display the power status of the device, the network indicator light is used to display the network connection status of the device, the video recording indicator light is used to display the video recording status of the device, and the laser light can be used for auxiliary positioning and indication. These indicator lights can provide intuitive device status information for operators, which is convenient for operation and management.

[0127] ③ Signal relay transmission device: It is used to transmit the data collected by the first-person view high-definition audio and video synchronous camera device to the portable edge computing terminal, and has stable data transmission capabilities, which can ensure the real-time and integrity of data transmission. The specific component composition is as follows:

[0128] Wireless communication module: It supports WiFi network and is used to establish wireless connections with the first-person view high-definition audio and video synchronous camera device and the portable edge computing terminal to achieve wireless data transmission; it supports Bluetooth 4.2 and is used to transmit and communicate data with nearby devices to ensure the short-distance connection stability between devices; it supports mobile / unicom / telecom 4G full network access and can provide high-speed data transmission services in areas covered by cellular networks to ensure the real-time and reliability of data transmission; it adopts intelligent MESH networking technology, which can achieve rapid component networking and high-speed data interconnection within several kilometers without manual intervention. This module can automatically establish and maintain network connections and adapt to complex network environments and operation scenarios.

[0129] Network management module: It uses proprietary technology to ensure stable networking and supports seamless networking of up to 64 nodes. This module is responsible for managing the routing protocols in the network to ensure that data can be efficiently and stably transmitted between various nodes; it supports a point-to-point peak bandwidth of up to 62 Mbps and can dynamically adjust bandwidth allocation according to network conditions to ensure the efficiency and stability of data transmission. This module also has anti-interference capabilities and can maintain the integrity of data transmission in complex electromagnetic environments.

[0130] Data Encryption Module: Supports multiple encryption algorithms such as AES, RSA, etc., which are used to encrypt the transmitted data to prevent data leakage and ensure transmission security. This module can select the appropriate encryption algorithm according to different security requirements to ensure the confidentiality and integrity of the data during transmission.

[0131] Power Management Module: Provides a stable power supply for the device, supports multiple power input methods such as mains power, battery, etc., to ensure the normal operation of the device in different environments; monitors the power status of the device in real time, including information such as voltage, current, and power, to ensure power safety during device operation. This module also has a power failure alarm function, which can send an alarm in time when the power is abnormal to remind the user to handle it.

[0132] Antenna Module: Adopts high-gain antennas, which can effectively improve the transmission distance and coverage of wireless signals. The antenna part is specially adapted and optimized by the antenna factory to ensure the optimal wireless signal of the device; supports multi-antenna switching, and can automatically select the optimal antenna for data transmission according to the network conditions and signal strength to ensure the stability and reliability of data transmission.

[0133] In the scenario of live working on the distribution network, the specific implementation methods are as follows:

[0134] Device Deployment: At the live working site of the distribution network, the operator wears a safety helmet equipped with a first-person perspective high-definition audio-visual synchronous camera device, the portable edge computing terminal is placed near the operator, and the signal relay transmission device is deployed at a suitable position at the working site to ensure smooth communication between devices.

[0135] Data Acquisition: The first-person perspective high-definition audio-visual synchronous camera device collects video and audio data of the working site in real time and transmits the data to the signal relay transmission device through self-organizing network technology.

[0136] Data Transmission: The signal relay transmission device receives the data transmitted by the first-person perspective high-definition audio-visual synchronous camera device and transmits the data to the portable edge computing terminal through self-organizing network technology.

[0137] Data Processing: After receiving the data, the portable edge computing terminal uses a multi-modal large model to analyze and process the data in real time, extracts the key information of the working site, and provides real-time auxiliary decision-making support for the operator.

[0138] Result Feedback: The portable edge computing terminal feeds back the processing results to the operator, and the operator performs corresponding operations according to the feedback information to ensure the safe and efficient progress of the work.

[0139] The main steps for collecting first-person perspective video data of live working on the distribution network are as follows:

[0140] ① Start the acquisition device: Before the operator starts the operation, first start the first-person view acquisition device. After the device is started, it will automatically perform initialization and self-check procedures to ensure that all functions are running properly. Specifically, the device will load preset configuration parameters, including the resolution, frame rate, exposure time, etc. of the camera, to adapt to different operation environments and requirements. At the same time, the device will perform self-checks on its own hardware and software to ensure that key components such as cameras, sensors, and data transmission modules are working properly. If any abnormalities are found, the device will issue an alarm and record error information for timely repair or replacement. In addition, the device will inform the operator through the user interface (such as a display screen or voice prompt) that the device is ready to start the operation. This process ensures that the device is in the best state before the operation starts, providing a reliable guarantee for subsequent data acquisition.

[0141] ② Real-time acquisition of video data: During the operation, the first-person view acquisition device continuously acquires video data of the operation site and sends the data to the backend processing system through the data transmission module. Specifically, the device acquires video data of the operation site in real time at the preset resolution and frame rate to ensure the clarity and smoothness of the video. The camera's field of view covers the operator's field of vision, enabling it to comprehensively and accurately capture every detail of the operation site. The acquired video data is sent to the backend processing system in real time through the data transmission module. The data transmission module uses 5G wireless transmission technology to ensure the stability and real-time nature of data transmission. To ensure the security of data transmission, the data is encrypted during transmission to prevent data from being intercepted or tampered with. In addition, the device also has a certain buffering capacity to cope with network fluctuations or temporary interruptions to ensure the integrity and continuity of data.

[0142] ③ Data storage and management: The video data will be stored in the storage module, and at the same time, a data indexing and management mechanism will be established for subsequent data query and analysis. Specifically, the storage module has the characteristics of large capacity and high-speed reading and writing, which can meet the data storage requirements of long-term operations. The system will establish an index for each video data file, including metadata such as timestamps, operator information, and operation locations, to facilitate users to quickly locate and retrieve the required data. In addition, the system also provides data management software through which users can perform operations such as backing up, restoring, and deleting the stored video data to ensure the security and integrity of the data. The system also supports data query and analysis functions. Users can query video data according to conditions such as time, location, and operator, and perform operations such as playing, editing, and annotating to extract useful information to support decision-making in live distribution network operations. Through these storage and management measures, the value of video data is fully utilized, providing strong guarantee for the safety and efficiency of operations.

[0143] Example 2:

[0144] Figure 4 A block diagram of a distribution network operation monitoring data processing system provided by the present invention, the system comprising:

[0145] The acquisition module 401 is configured to acquire a plurality of video monitoring data of the first-person perspective of the distribution network operation site and a model execution mode of a pre-trained multi-modal large model;

[0146] A generating module 402 is configured to generate a data set of the first perspective of the distribution network operation site according to the plurality of video monitoring data;

[0147] The determination module 403 is configured to process the data set through the multimodal large model according to the model execution mode to obtain a processing result corresponding to the data set.

[0148] Optionally, the data set includes labeling data of each to-be-used video monitoring data, and the determination module 403 is configured to:

[0149] When the model execution mode is the video question-answering mode, the multimodal large model interactively processes multiple video monitoring data and multiple annotated data in the data set to obtain an answer result corresponding to each annotated data.

[0150] Optionally, the determining module 403 is configured to:

[0151] Extracting features from the plurality of video monitoring data using the multimodal large model to obtain a first spatiotemporal feature of each video monitoring data;

[0152] Interactively process the multiple first spatiotemporal features and the multiple annotated data to obtain an answer result corresponding to each annotated data.

[0153] Optionally, the determining module 403 is configured to:

[0154] When the model execution mode is a video classification mode, feature extraction is performed on the plurality of video monitoring data by using the multimodal large model to obtain a second spatiotemporal feature of each video monitoring data;

[0155] According to the plurality of second spatiotemporal features, the plurality of video monitoring data in the data set are classified and processed to generate a classification result corresponding to each of the plurality of video monitoring data.

[0156] Optionally, the training process of the multimodal large model includes:

[0157] Acquire multiple standby video data of the first-person perspective of the distribution network operation site, and standby annotation data of each video data;

[0158] Generate a training set of the multimodal large model according to the plurality of unused video data and the unused annotation data of each video data;

[0159] The training set is input into an initial multimodal large model, and the initial multimodal large model is trained to obtain the multimodal large model.

[0160] Optionally, the initial modal large model includes multiple convolution kernels, multimodal rotation position embedding and a cross-modal attention mechanism, and the training set is input into the initial multimodal large model, the initial multimodal large model is trained, and the multimodal large model is obtained, including:

[0161] Extracting features of a plurality of unused video data in the training set by using the plurality of convolution kernels to obtain unused spatiotemporal features corresponding to each unused video data;

[0162] Encoding a plurality of standby spatiotemporal features by means of the multimodal rotation position embedding to obtain a plurality of encoded standby spatiotemporal features;

[0163] The cross-modal attention mechanism is used to interact the encoded multiple unused spatiotemporal features with the multiple unused labeled data in the training set to obtain the multimodal large model.

[0164] Optionally, the generating module 402 is configured to:

[0165] Performing image enhancement, video parameter adjustment and denoising processing on the plurality of video monitoring data to obtain a plurality of stand-by video monitoring data;

[0166] Obtaining labeling data of each of the plurality of video monitoring data to be used;

[0167] A data set of the first-person perspective of the distribution network operation site is generated according to the multiple stand-by video monitoring data and the annotation data of each stand-by video monitoring data.

[0168] Embodiment 3:

[0169] Based on the same inventive concept, the present invention further provides a computer device, which includes a processor and a memory. The memory is used to store a computer program, and the computer program includes program instructions. The processor is used to execute the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions in the computer storage medium to implement the corresponding method flow or corresponding function, so as to implement the steps of a method for processing distribution network operation monitoring data in the above embodiments.

[0170] Embodiment 4:

[0171] Based on the same inventive concept, the present invention further provides a storage medium, specifically a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in a computer device and is used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device and, of course, the extended storage medium supported by the computer device. The computer-readable storage medium provides a storage space, and this storage space stores the operating system of the terminal. And, one or more instructions suitable for being loaded and executed by the processor are also stored in this storage space. These instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. The one or more instructions stored in the computer-readable storage medium can be loaded and executed by the processor to implement the steps of a method for processing distribution network operation monitoring data in the above embodiments.

[0172] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0173] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0174] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0175] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0176] The above are only embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are included in the scope of the claims of the present invention pending approval.

Claims

1. A method for processing distribution network operation monitoring data, characterized in that, The method comprises: Obtain multiple video monitoring data from the first-person perspective of the distribution network operation site and the model execution mode of the pre-trained multi-modal large model; Generate a data set of the first-person perspective of the distribution network operation site according to the multiple video monitoring data; According to the model execution mode, the data set is processed by the multimodal large model to obtain a processing result corresponding to the data set.

2. The method according to claim 1, characterized in that, The data set includes the labeled data of each to-be-used video monitoring data, and the data set is processed by the multimodal large model according to the model execution mode to obtain the processing result corresponding to the data set, including: When the model execution mode is the video question-answering mode, the multimodal large model interactively processes multiple video monitoring data and multiple annotated data in the data set to obtain an answer result corresponding to each annotated data.

3. The method according to claim 2, characterized in that The interactive processing of the multiple video monitoring data and the multiple annotated data in the data set by the multimodal large model to obtain the answer result corresponding to each annotated data includes: Extracting features from the plurality of video monitoring data using the multimodal large model to obtain a first spatiotemporal feature of each video monitoring data; Interactively process the multiple first spatiotemporal features and the multiple annotated data to obtain an answer result corresponding to each annotated data.

4. The method according to claim 1, characterized in that The step of processing the data set by using the multimodal large model according to the model execution mode to obtain a processing result corresponding to the data set includes: When the model execution mode is a video classification mode, feature extraction is performed on the plurality of video monitoring data by using the multimodal large model to obtain a second spatiotemporal feature of each video monitoring data; According to the plurality of second spatiotemporal features, the plurality of video monitoring data in the data set are classified and processed to generate a classification result corresponding to each of the plurality of video monitoring data.

5. The method according to claim 1, characterized in that, The training process of the multimodal large model includes: Acquire multiple standby video data of the first-person perspective of the distribution network operation site, and standby annotation data of each video data; Generate a training set of the multimodal large model according to the plurality of unused video data and the unused annotation data of each video data; The training set is input into an initial multimodal large model, and the initial multimodal large model is trained to obtain the multimodal large model.

6. The method according to claim 2, wherein The initial modality large model includes multiple convolution kernels, multimodal rotation position embedding and cross-modality attention mechanism, and the training set is input into the initial multimodal large model, and the initial multimodal large model is trained to obtain the multimodal large model, including: Extracting features of a plurality of unused video data in the training set by using the plurality of convolution kernels to obtain unused spatiotemporal features corresponding to each unused video data; Encoding a plurality of standby spatiotemporal features by means of the multimodal rotation position embedding to obtain a plurality of encoded standby spatiotemporal features; The cross-modal attention mechanism is used to interact the encoded multiple unused spatiotemporal features with the multiple unused labeled data in the training set to obtain the multimodal large model.

7. The method according to claim 1, characterized in that The step of generating a data set of the first-person perspective of the distribution network operation site according to the plurality of video monitoring data includes: Performing image enhancement, video parameter adjustment and denoising processing on the plurality of video monitoring data to obtain a plurality of stand-by video monitoring data; Obtaining labeling data of each of the plurality of video monitoring data to be used; A data set of the first-person perspective of the distribution network operation site is generated according to the multiple stand-by video monitoring data and the annotation data of each stand-by video monitoring data.

8. A processing system for distribution network operation monitoring data, characterized in that, The system comprises: An acquisition module is configured to acquire a plurality of video monitoring data of a first-person perspective of a distribution network operation site and a model execution mode of a pre-trained multi-modal large model; A generating module is configured to generate a data set of the first-person perspective of the distribution network operation site according to the plurality of video monitoring data; The determination module is configured to process the data set through the multimodal large model according to the model execution mode to obtain a processing result corresponding to the data set.

9. A computer device, characterized in that, include: one or more processors; The processor is used to store one or more programs; When the one or more programs are executed by the one or more processors, a method for processing distribution network operation monitoring data as described in any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and when the computer program is executed, a method for processing distribution network operation monitoring data as described in any one of claims 1 to 7 is implemented.