Power distribution network emergency operation ticket generation method and system based on multi-modal fusion
Through multimodal fusion technology, combined with the improved Jieba and Transformer models to process text data, and combined with the improved YOLOv8 model to process image data, the information asymmetry problem of traditional distribution network invoice system in burst scenarios is solved, and efficient and accurate generation of operation tickets is achieved, reducing safety hazards.
Patent Information
- Application Number
- CN202510184123.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-07-11
AI Technical Summary
Traditional distribution network invoice system lacks intelligence and automation methods, it is difficult to deal with on-site image information in burst scenarios, and it is impossible to automatically generate operation tickets that match the actual situation, resulting in insufficient invoice specifications and safety hazards.
The distribution network emergency operation ticket generation method based on multimodal fusion is adopted, and text data is processed through improved Jieba word segmentation technology and Transformer model, and image data is processed in combination with the improved YOLOv8 model to achieve deep semantic understanding of text and images, and operation tickets that meet safety regulations are automatically generated.
It improves the timeliness and accuracy of operation tickets, ensures that operation tickets that meet the actual situation are generated in emergencies, reduces the risk of missed tickets and missed tickets, and improves the accuracy and efficiency of on-site invoice issuance.
Smart Images

Figure SMS_1 
Figure SMS_2 
Figure SMS_3
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning, and more specifically, to a method and system for generating emergency operation tickets for distribution networks based on multimodal fusion. Background Art
[0002] The workload of issuing operation tickets for distribution network operations is large. At the same time, unplanned temporary repair tasks occur frequently, and the operating personnel are short on time and have heavy tasks. Most traditional ticket-issuing processes rely on manual filling and auditing, which are both time-consuming and prone to problems such as missing tickets, incorrect tickets, and wrong tickets. Especially in the case of emergency repairs, manual operations are extremely likely to lead to insufficient ticket-issuing specifications, and even unsafe phenomena such as "operating without a ticket". By implementing functions such as automatic ticket generation by the system, on-site electronic ticket issuance, and automatic verification of operation tickets, the accuracy of on-site ticket issuance can be guaranteed, the efficiency of on-site ticket issuance can be improved, the repair time can be reduced, and the burden on grass-roots units can also be relieved. However, traditional ticket-issuing systems lack intelligent and automated means, making it difficult to effectively support the accuracy and timeliness of on-site ticket issuance, increasing the time burden of distribution network repair and maintenance.
[0003] The prior art discloses an intelligent ticket-issuing system and method for distribution network dispatching operation tickets based on a single-line diagram. The ticket-issuing system includes an operation ticket generation module, a database server, and a single-line diagram review module that are connected to each other. The database server includes an operation rule database, an operation ticket database, and a single-line diagram database. The single-line diagram review module is connected to the single-line diagram management database; the single-line diagram review module converts the single-line diagram in the single-line diagram management database into single-line diagram data and stores it in the single-line diagram database. When the single-line diagram after a power outage repair project changes, the changed single-line diagram can also be used for intelligent ticket issuance through the ticket-issuing system; the scope of use of the ticket-issuing system is improved, and the popularization of the ticket-issuing system is promoted. However, this method is difficult to process on-site image information in emergency scenarios and cannot automatically generate operation tickets that match the actual situation. Summary of the Invention
[0004] The purpose of the present invention is to disclose a method, system, and storage medium for generating emergency operation tickets for distribution networks based on multimodal fusion with better dispatching effects.
[0005] To achieve the above purpose, the present invention provides a method for generating emergency operation tickets for distribution networks based on multimodal fusion, including:
[0006] S1: Obtain historical work ticket data as a source sequence, and clean the source sequence to obtain a cleaned source sequence;
[0007] S2: Construct a ticket-issuing model, extract keywords from the cleaned source sequence, and then train the ticket-issuing model to obtain a text ticket-issuing model;
[0008] S3: Obtain the emergencies of the distribution network, construct an image dataset and an image recognition model, and pre-train the image recognition model according to the image dataset to obtain a pre-trained image recognition model;
[0009] S4: Construct a specific task dataset and perform data augmentation processing to obtain an augmented specific task dataset;
[0010] S5: Train the pre-trained image recognition model and the text invoicing model according to the augmented specific task dataset to obtain a multi-modal distribution network emergency operation ticket generation model.
[0011] Furthermore, in step S1, cleaning the source sequence includes: correcting typos, removing irrelevant characters, and standardizing the expression by unifying the term format.
[0012] Furthermore, in step S2, extracting keywords from the cleaned source sequence includes: performing word segmentation through an improved version of Jieba.
[0013] Furthermore, it includes:
[0014] S2.1: In the improved version of Jieba, the domain adaptation dictionary and the dynamic update module construct an adaptive dictionary and an automatic update mechanism for the distribution network operation ticket according to the common words and specific scenario terms of historical work tickets. By analyzing a large number of historical work ticket texts, extract specific professional terms, equipment names, and operation action words, and automatically update them to the custom dictionary according to the frequency. And during the operation process, monitor the newly input unrecognized words, calculate the word frequency of each new word ω, and weight it according to the domain characteristics of the vocabulary to enhance the weights of professional terms and high-frequency equipment names. The calculation formula is as follows:
[0015]
[0016] In the formula, F(ω) is the weighted word frequency, f(ω) is the frequency of the word in all texts, N is the total number of words, f domain (ω) is the frequency of the word ω in the domain dictionary, N domain is the total number of words in the domain dictionary, α and β are adjustment parameters used to balance the influence of the domain dictionary and the ordinary word frequency, usually set α + β = 1,
[0017] Set a dynamic threshold so that new words with a frequency exceeding this threshold automatically enter the domain dictionary. The update formula for the threshold is:
[0018]
[0019] In the formula, θ t is the threshold for the t-th update, γ is the adjustment coefficient of the threshold, θ t-1 is the previous threshold, and η is the learning rate;
[0020] Let \(M\) be the number of newly emerged words in the current round, and \(F(\omega i )\) be the weighted word frequency of these new words;
[0021] Threshold update ensures the adaptability of the system, enabling the system to automatically adapt to different scenarios and perform dynamic updates as new devices are added;
[0022] S2.2: Custom word segmentation rules and named entity recognition modules in Jieba can accurately identify device names, action instructions, locations, and times. By using the custom word segmentation rule function of Jieba, key entities in the operation ticket are set as words to be preferentially recognized. At the same time, combined with a deep learning-based NER model, the recognition of uncommon nouns is improved, and word segmentation rules such as "circuit breaker number + serial number" and action word recognition rules such as "cut off + device name" are defined. For specific tasks, the following weighted dictionary optimization formula is defined:
[0023] weight(\(\omega i ) = \beta\cdot TF - IDF(\omega i )+\gamma\cdot X(\omega i )
[0024] Where \(\beta\) and \(\gamma\) are hyperparameters used to balance the influence of word frequency and context factors; \(TF - IDF(\omega i )\) is the traditional word weight based on word frequency and document frequency; \(X(\omega i )\) is the context-based weight, reflecting the importance of the word in a specific operation step;
[0025] Define the word set as \(W\), where each word \(\omega i \in W\), and its corresponding named entity category \(C(\omega i )\in\{device, action, location, time\}\). The system represents the category of each word through the NER model. According to the domain adaptive dictionary \(D\) and the dynamic update mechanism, the word frequency and entity category are weighted and adjusted:
[0026]
[0027] Where \(score(\omega i )\) represents the original score of the word in the source text, based on word frequency, context information, etc.; \(\alpha(C(\omega i ))\) is the coefficient weighted according to the entity category of the word, meeting the requirements of a specific domain; the weight adjustment of action instructions and device names is dynamically learned through frequency analysis of historical data;
[0028] S2.3: The context-based word segmentation optimization module in the improved Jieba dynamically optimizes the word segmentation strategy according to different contexts of the operation ticket. By analyzing different contexts in historical work tickets, the word segmentation strategy is adjusted to pay more attention to identifying device status words in the safety inspection steps and action words in the operation steps, and automatically increases the weights of the corresponding words.
[0029] Further, in step S2, the ticket-issuing model is a Transformer model.
[0030] Further, it includes:
[0031] S2.4: Perform word embedding in the Transformer model, that is, convert keyword information into a vector of a fixed size;
[0032] S2.5: Add positional encoding in the Transformer model to provide the position information of the words, and add the positional encoding to the word embedding vector;
[0033] S2.6: The self-attention mechanism in the Transformer model allows the model to consider all other words in the sequence while processing a certain word, so as to capture long-distance dependencies. This feature gives the Transformer great advantages in dealing with the sorting problem in the operation ticket. In addition, each layer in the Transformer contains a self-attention mechanism and a feed-forward neural network, which can gradually extract the deep features of the text, and then accurately understand the polysemous words and specific operation intentions in the work ticket;
[0034] S2.7: A series of vectors output by the Transformer model are sent into the decoder for decoding to generate the target text, that is, to generate a standardized and sequential electronic operation ticket.
[0035] Further, in step S3, it includes: The image recognition model is an improved YOLOv8 model, specifically:
[0036] S3.1: Replace the Bottleneck in C2f of YOLOv8 with GhostBottleneck;
[0037] The described GhostBottleneck consists of two stacked Ghost modules. Batch normalization and ReLu activation function are applied after the first Ghost module, and only batch normalization is performed after the second Ghost module;
[0038] S3.2: Add deformable convolution (DCNv2) after C2f in the Backbone. The features output in the standard convolution mode for the center point a0 are:
[0039]
[0040] In the formula, x is the input feature map, y is the output feature map, a0 is the true coordinate of the center point in the output feature map, and a n is the position of the convolution kernel, is the weight at the position of a n and N is the number of sampling points;
[0041] After the deformable convolution samples the input feature map, an offset Δa n is introduced;
[0042]
[0043] S3.3: A modulation mechanism is introduced, enabling the deformable convolutional neural network module to not only adjust the offset of the perceived input features but also adjust the amplitudes of the input features from different spatial positions;
[0044]
[0045] In the formula, Δl n is the weight modulation parameter at the sampling point a n position, and p(a n ) is the penalty term;
[0046] S3.4: Add a channel-spatial attention integration module with adaptive ability after SPPF in the Backbone;
[0047] The channel attention module first processes the input feature map with global average pooling and global max pooling to extract channel information, then generates channel weight coefficients through a convolutional layer, and finally passes through a sigmoid layer to apply the weights to the original feature map. The channel attention module can effectively enhance the key feature channels, improve the sensitivity of the model to important information, and thus improve the detection and recognition accuracy of on-site pictures in the emergency situation of the distribution network;
[0048] The spatial attention module first performs global average pooling and global max pooling on the input feature map to generate two-channel feature maps, then inputs them into a convolutional layer through concatenation to generate a spatial attention map, and finally applies the spatial attention map to the original feature map using the sigmoid function. The module can improve the model's evaluation of the importance of spatial regions, thereby improving the localization accuracy of the model in complex scenarios of the distribution network;
[0049] The CSA mentioned above can adaptively select and adjust the channel weights and spatial weights of the feature map, which can capture image features better while reducing the computational amount;
[0050] Add the CSA module before CBS in the Neck;
[0051] S3.5: The classification loss function is VFL, and the regression loss adopts the form of CIoU Loss + DFL;
[0052]
[0053] In the formula, S R and S p are the areas of the ground truth box and the predicted box respectively, IoU is the intersection over union of the predicted box and the ground truth box, m is the predicted class score, n is the predicted object score. If it is the true class, n = IoU; if it is other classes, n = 0;
[0054]
[0055] In the formula, c and c t represent the centers of the predicted target box and the ground truth target box respectively, l is the diagonal length of the smallest box that can enclose both the predicted box and the ground truth box, β is the penalty coefficient, w t and w are the widths of the ground truth box and the predicted box respectively, h t and h are the lengths of the ground truth box and the predicted box;
[0056]
[0057] In the formula, y i is the predicted value, and y is the true value;
[0058] S3.6: Use transfer learning technology to optimize the model. The early layers of the pre-trained model are used to extract general features such as edges, textures, and shapes in the image, and the weights of the early layers are frozen after preliminary pre-training.
[0059] Furthermore, in step S4, data augmentation processing includes horizontal flipping, scaling, cropping, color adjustment, and adding slight blur, brightness change, and color distortion to the pictures taken on site or by drones.
[0060] Furthermore, it includes:
[0061] In step S3, during the training of the image recognition model, freeze the weights of the early layers of the image recognition model;
[0062] In step S5, during the training of the image recognition model, adjust the weights of the later layers of the image recognition model.
[0063] In addition, the present invention also provides a distribution network emergency operation ticket generation system based on multimodal fusion, including:
[0064] Cleaning module: Obtain historical work ticket data as the source sequence, and clean the source sequence to obtain the cleaned source sequence;
[0065] Text training module: Build an invoicing model, extract keywords from the cleaned source sequence, and then train the invoicing model to obtain a text invoicing model;
[0066] Image pre-training module: Obtain power distribution network emergencies, build an image dataset and an image recognition model, and pre-train the image recognition model according to the image dataset to obtain a pre-trained image recognition model;
[0067] Specific module: Build a specific task dataset and perform data augmentation processing to obtain an enhanced specific task dataset;
[0068] Multi-modal module: Train the pre-trained image recognition model and the text invoicing model according to the enhanced specific task dataset to obtain a multi-modal power distribution network emergency operation ticket generation model.
[0069] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:
[0070] The multi-modal power distribution network emergency operation ticket generation model in the present invention not only receives text tokenization results, but also integrates image recognition results to achieve deep semantic understanding of text and images. The model can accurately identify the device operation relationship and automatically generate the content of the operation ticket that complies with the safety regulations. This model solves the problem of information asymmetry caused by data fragmentation in traditional methods, enabling the system to accurately adapt to the actual situation when generating operation tickets. In particular, the timeliness and accuracy of generating operation tickets in case of emergencies have been significantly improved. Brief Description of the Drawings
[0071] Figure 1 It is a flowchart of the power distribution network emergency operation ticket generation method based on multi-modal fusion described in Embodiment 1;
[0072] Figure 2 It is the improved C2f diagram described in Embodiment 2;
[0073] Figure 3 It is the improved Backbone diagram described in Embodiment 2;
[0074] Figure 4 It is the structure diagram of CSA described in Embodiment 2;
[0075] Figure 5 It is the improved Neck diagram described in Embodiment 2;
[0076] Figure 6 It is the improved YOLOv8 diagram described in Embodiment 2;
[0077] Figure 7 It is the block diagram of the power distribution network emergency operation ticket generation system based on multi-modal fusion described in Embodiment 3; Specific embodiments
[0078] The accompanying drawings are only for illustrative purposes and should not be construed as limiting the patent;
[0079] The technical solutions of the present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0080] Embodiment 1:
[0081] This embodiment provides a method for generating an emergency operation ticket for a distribution network based on multi-modal fusion as shown in Figure 1 and includes:
[0082] S1: Obtain historical work ticket data as the source sequence, and clean the source sequence to obtain the cleaned source sequence;
[0083] S2: Construct an operation ticket generation model, extract keywords from the cleaned source sequence, and then train the operation ticket generation model to obtain a text-based operation ticket generation model;
[0084] S3: Obtain the emergency situation of the distribution network, construct an image data set and an image recognition model, and pre-train the image recognition model according to the image data set to obtain a pre-trained image recognition model;
[0085] S4: Construct a specific task data set and perform data augmentation processing to obtain an augmented specific task data set;
[0086] S5: Train the pre-trained image recognition model and the text-based operation ticket generation model according to the augmented specific task data set to obtain a multi-modal emergency operation ticket generation model for the distribution network.
[0087] In this embodiment, the multi-modal emergency operation ticket generation model for the distribution network not only receives the text tokenization results, but also fuses the image recognition results to achieve deep semantic understanding of text and images. The model can accurately identify the equipment operation relationship and automatically generate the content of the operation ticket that complies with the safety regulations. This model solves the problem of information asymmetry caused by data fragmentation in traditional methods, enabling the system to accurately adapt to the actual situation when generating the operation ticket. In particular, the timeliness and accuracy of generating the operation ticket in case of emergencies have been significantly improved.
[0088] Embodiment 2:
[0089] This embodiment further discloses on the basis of Embodiment 1:
[0090] Further, in step S1, cleaning the source sequence includes: correcting typos, removing irrelevant characters, and standardizing the expression by unifying the term format.
[0091] Further, in step S2, extracting keywords from the cleaned source sequence includes: performing word segmentation through an improved version of Jieba.
[0092] Furthermore, it includes:
[0093] S2.1: In the domain adaptation dictionary and dynamic update module of the improved Jieba, an adaptive dictionary and an automatic update mechanism for distribution network operation tickets are constructed according to the common words and specific scenario terms in historical work tickets. By analyzing a large number of historical work ticket texts, specific professional terms, equipment names, and operation action words are extracted, automatically updated to the custom dictionary according to the frequency, and during the operation process, new unrecognized words input are monitored, the word frequency of each new word ω is calculated, and weighted according to the domain characteristics of the vocabulary to enhance the weights of professional terms and high-frequency equipment names. The calculation formula is as follows:
[0094]
[0095] In the formula, F(ω) is the weighted word frequency, f(ω) is the frequency of the word in all texts, N is the total number of words, f domain (ω) is the frequency of the word ω in the domain dictionary, N domain is the total number of words in the domain dictionary, α and β are adjustment parameters used to balance the influence of the domain dictionary and the ordinary word frequency, usually set as α + β = 1,
[0096] Set a dynamic threshold so that new words with a frequency exceeding this threshold automatically enter the domain dictionary. The update formula for the threshold is:
[0097]
[0098] In the formula, θ t is the threshold for the t-th update, γ is the adjustment coefficient of the threshold, θ t-1 is the previous threshold, and η is the learning rate;
[0099] M is the number of new words that appear in the current round, F(ω i ) is the weighted word frequency of these new words;
[0100] The threshold update ensures the self-adaptability of the system, enables the system to automatically adapt to different scenarios, and performs dynamic updates as new devices are added;
[0101] S2.2: The custom word segmentation rules and named entity recognition module in Jieba can accurately identify equipment names, action instructions, locations, and times. Using the custom word segmentation rules function of Jieba, the key entities in the operation ticket are set as the words to be preferentially recognized. At the same time, combined with the deep learning-based NER model, the recognition of uncommon nouns is improved, and the word segmentation rules such as "circuit breaker number + serial number" and the action word recognition rules such as "cut off + equipment name" are defined. For specific tasks, the following weighted dictionary optimization formula is defined:
[0102] weight(ω i ) = β·TF-IDF(ω i ) + γ·X(ω i )
[0103] where β and γ are hyperparameters used to balance the influence of term frequency and context factors; TF-IDF(ω i ) is the traditional word weight based on term frequency and document frequency; X(ω i ) is the context-based weight that reflects the importance of a word in a specific operation step;
[0104] Definition Let the word set be W, where each word ω i ∈W, and its corresponding named entity category C(ω i ) ∈ {device, action, location, time}. The system represents the category of each word through the NER model, and according to the domain adaptive dictionary D and the dynamic update mechanism, the term frequency and entity category are weighted and adjusted:
[0105]
[0106] where score(ω i ) represents the original score of the word in the source text, based on term frequency, context information, etc.; α(C(ω i )) is the coefficient weighted according to the entity category of the word, adapting to the needs of a specific domain; the weight adjustment of action instructions and device names is dynamically learned through the frequency analysis of historical data;
[0107] S2.3: The context-based word segmentation optimization module in the improved Jieba dynamically optimizes the word segmentation strategy according to different contexts of the operation ticket. By analyzing different contexts in historical work tickets, the word segmentation strategy is adjusted, paying more attention to identifying device status vocabulary in the safety inspection step and more attention to identifying action vocabulary in the operation step, and automatically increasing the weight of the corresponding vocabulary.
[0108] Furthermore, in step S2, the invoicing model is a Transformer model.
[0109] Furthermore, it includes:
[0110] S2.4: Perform word embedding in the Transformer model, that is, convert the keyword information into a fixed-size vector;
[0111] S2.5: Add position encoding in the Transformer model to provide the position information of the word, and add the position encoding to the word embedding vector;
[0112] S2.6: The self-attention mechanism in the Transformer model allows the model to consider all other words in the sequence while processing a certain word, thereby capturing long-range dependencies. This feature gives Transformer a great advantage in dealing with the sorting problem in operation tickets. In addition, each layer in the Transformer contains a self-attention mechanism and a feed-forward neural network, which can gradually extract the deep features of the text, and then accurately understand the polysemous words and specific operation intentions in the work ticket;
[0113] S2.7: A series of vectors output by the Transformer model are sent into the decoder for decoding to generate the target text, that is, to generate a standardized and sequential electronic operation ticket.
[0114] Furthermore, in step S3, it includes: The image recognition model is an improved YOLOv8 model, specifically:
[0115] S3.1: As Figure 2 shown, replace the Bottleneck in C2f with GhostBottleneck, thereby reducing the computational load and the model volume, making it easier for the system to be deployed lightly on different platforms including edge devices with limited hardware resources, and improving the computational efficiency of the system, which is very beneficial for the handling of emergency situations in the distribution network with high real-time requirements;
[0116] The described GhostBottleneck consists of two stacked Ghost modules. Batch normalization (BN) and ReLu activation function are applied after the first Ghost module, and only batch normalization is performed after the second Ghost module;
[0117] S3.2: As Figure 3 shown, add deformable convolution (DCNv2) after C2f in the Backbone to improve the adaptability of the network to target objects with irregular shapes and expand its receptive field. In the real scenarios of emergencies in the distribution network such as tree collapse, pole collapse, line disconnection, and insulator damage, the shapes of many target objects are irregular. Deformable convolution can adaptively adjust the shape and size of the receptive field according to the irregular shape of the target object, thereby improving the robustness of the system;
[0118] The feature of the center point a0 output in the standard convolution mode is:
[0119]
[0120] In the formula, x is the input feature map, y is the output feature map, a0 is the true coordinate of the center point in the output feature map, a n is the position of the convolution kernel, is a nThe weight at the position, where N is the number of sampling points;
[0121] After sampling the input feature map by deformable convolution, an offset Δa is introduced n ;
[0122]
[0123] S3.3: To further enhance the control ability of the deformable convolutional neural network over the spatial support region, a modulation mechanism is introduced, enabling the deformable convolutional neural network module to not only adjust the offset of the perceived input features but also adjust the amplitudes of the input features from different spatial positions;
[0124]
[0125] In the formula, Δl n is the weight modulation parameter at the sampling point a n position, and p(a n ) is the penalty term.
[0126] S3.4: As Figure 2 shown, a channel-spatial attention integration (CSA) module with adaptive ability is added after SPPF in the Backbone;
[0127] The channel attention module first processes the input feature map with global average pooling and global max pooling to extract channel information, then passes through a convolutional layer to generate channel weight coefficients, and finally, after a sigmoid layer, applies the weights to the original feature map. The described channel attention module can effectively enhance key feature channels, improve the model's sensitivity to important information, and thus improve the detection and recognition accuracy of on-site pictures in emergency situations of the distribution network;
[0128] The spatial attention module first performs global average pooling and global max pooling on the input feature map to generate two-channel feature maps, then inputs them into a convolutional layer through concatenation to generate a spatial attention map, and finally applies the spatial attention map to the original feature map using the sigmoid function. The module can improve the model's evaluation of the importance of spatial regions, thereby improving the model's localization accuracy in complex scenarios of the distribution network;
[0129] As Figure 4 shown, the described CSA can adaptively select and adjust the channel weights and spatial weights of the feature map, which can capture image features better while reducing the computational amount;
[0130] As Figure 5 shown, a CSA module is also added before CBS in the Neck; The final improved YOLOv8 structure diagram is as Figure 6 shown;
[0131] S3.5: In the described YOLOv8, the classification loss function is VFL (Varifocal Loss), and the regression loss adopts the form of CIoU Loss + DFL (Distribution Focal Loss);
[0132]
[0133] In the formula, S R and S p are the areas of the ground truth box and the predicted box respectively, IoU is the intersection over union of the predicted box and the ground truth box, m is the predicted class score, n is the predicted object score. If it is the true class, n = IoU; if it is other classes, n = 0;
[0134]
[0135]
[0136] In the formula, c and c t represent the center points of the predicted target box and the ground truth target box respectively, l is the diagonal length of the smallest box that can enclose both the predicted box and the ground truth box, β is the penalty coefficient, w t and w are the widths of the ground truth box and the predicted box respectively, h t and h are the lengths of the ground truth box and the predicted box;
[0137]
[0138] In the formula, y i is the predicted value, and y is the true value.
[0139] S3.6: To address the problem of insufficient sample quantity for distribution network faults or emergencies, transfer learning technology is used to optimize the model. The early layers of the pre-trained model are used to extract general features such as edges, textures, and shapes in the images, and these features are shared among different tasks. After the initial pre-training, the weights of the early layers are frozen to ensure the stability of these basic features during the subsequent fine-tuning process.
[0140] Furthermore, in step S4, the data augmentation process includes horizontal flipping, scaling, cropping, color adjustment, and adding slight blur, brightness change, and color distortion to the pictures taken on-site or by drones.
[0141] Furthermore, it includes:
[0142] In step S3, during the training of the image recognition model, the weights of the early layers of the image recognition model are frozen;
[0143] In step S5, during the training of the image recognition model, the weights of the later layers of the image recognition model are adjusted.
[0144] Among them, the large open-source dataset is a dataset publicly available on the Internet, and the specific task dataset is a dataset constructed according to specific tasks.
[0145] In this embodiment, the multi-modal distribution network emergency operation ticket generation model not only receives the text tokenization results, but also integrates the image recognition results to achieve deep semantic understanding of text and images. The model can accurately identify the equipment operation relationship and automatically generate the operation ticket content that complies with the safety regulations. This model solves the problem of information asymmetry caused by data fragmentation in traditional methods, enabling the system to accurately adapt to the actual situation when generating operation tickets. In particular, the timeliness and accuracy of generating operation tickets in emergency situations have been significantly improved.
[0146] In this embodiment, adaptive tokenization is performed through an improved Jieba tool to accurately extract the key terms in the operation ticket. The improved Jieba tokenization method combines dynamic dictionary adaptive update, named entity recognition (NER), and context-based tokenization optimization technology, which can identify and adapt to new equipment nouns at any time, ensuring accurate tokenization of industry terms by the system, greatly reducing the mis-tokenization phenomenon in traditional methods, and improving the precision of text processing. The tokenization results are deeply analyzed by the Transformer model to automatically generate operation tickets that comply with the distribution network operation standards, ensuring the semantic accuracy of the generated tickets. Through transfer learning technology, the image recognition model is optimized in the small-sample emergency fault scenario. Transfer learning enables the model to extract general features from large general datasets. In the case of insufficient specific fault samples, it can quickly adapt to different fault scenarios, significantly improving the generalization ability and accuracy of the model and reducing the dependence on a large amount of data. For emergency image data in emergency scenarios, an improved YOLOv8 model is adopted. By adding a channel spatial attention module (CSA), GhostBottleneck, and deformable convolution structure, the model can better adapt to the on-site environment with diverse equipment forms and complex situations, effectively solving the problem of poor adaptability of traditional image recognition technology to complex scenarios. The application of this multi-modal fusion technology ensures the accuracy and real-time performance in the operation ticket generation process.
[0147] Embodiment 3:
[0148] This embodiment provides a Figure 7 distribution network emergency operation ticket generation system based on multi-modal fusion as shown in
[0149] Cleaning module: Obtain the historical work ticket data as the source sequence, and clean the source sequence to obtain the cleaned source sequence;
[0150] Text training module: Build an invoicing model, extract keywords from the cleaned source sequence, and then train the invoicing model to obtain a text invoicing model;
[0151] Image pre-training module: Obtain power distribution network emergencies, build an image dataset and an image recognition model, and pre-train the image recognition model according to the image dataset to obtain a pre-trained image recognition model;
[0152] Specific module: Build a specific task dataset and perform data augmentation processing to obtain an enhanced specific task dataset;
[0153] Multi-modal module: Train the pre-trained image recognition model and the text invoicing model according to the enhanced specific task dataset to obtain a multi-modal power distribution network emergency operation ticket generation model.
[0154] In this embodiment, the multi-modal power distribution network emergency operation ticket generation model not only receives text tokenization results, but also integrates image recognition results to achieve deep semantic understanding of text and images. The model can accurately identify the operation relationships of devices and automatically generate operation ticket content that complies with safety regulations. This model solves the problem of information asymmetry caused by data fragmentation in traditional methods, enabling the system to accurately adapt to the actual situation when generating operation tickets. In particular, the timeliness and accuracy of generating operation tickets in emergency situations have been significantly improved.
[0155] Obviously, the above embodiments of the present invention are merely examples for clearly explaining the present invention, and are not intended to limit the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.
Claims
1. A method for generating an emergency operation ticket for a distribution network based on multimodal fusion, characterized in that Including: S1: Obtain historical work ticket data as the source sequence, and clean the source sequence to obtain the cleaned source sequence; S2: Construct an invoicing model, extract keywords from the cleaned source sequence, and then train the invoicing model to obtain a text invoicing model; S3: Select images similar to the on-site situation of sudden distribution network failures in a large open-source dataset to construct an image dataset; construct an image recognition model, and pre-train the image recognition model according to the image dataset to obtain a pre-trained image recognition model; S4: Construct a specific task dataset and perform data augmentation processing to obtain an augmented specific task dataset; S5: Train the pre-trained image recognition model and the text invoicing model according to the augmented specific task dataset to obtain a multi-modal distribution network emergency operation ticket generation model.
2. The method for generating an emergency operation ticket for a distribution network based on multimodal fusion according to claim 1, wherein In step S1, cleaning the source sequence includes: correcting typos, removing irrelevant characters, and standardizing the expression by unifying the term format.
3. The method for generating an emergency operation ticket for a distribution network based on multimodal fusion according to claim 1, wherein In step S2, extracting keywords from the cleaned source sequence includes: performing word segmentation through an improved version of Jieba.
4. The method for generating an emergency operation ticket for a distribution network based on multimodal fusion according to claim 3, wherein, Including: S2.1: In the improved version of Jieba, the domain adaptation dictionary and dynamic update module construct an adaptive dictionary and an automatic update mechanism for distribution network operation tickets according to the common words and specific scenario terms in historical work tickets. By analyzing a large number of historical work ticket texts, extract specific professional terms, equipment names, and operation action words, and automatically update them to the custom dictionary according to the frequency. And during the operation, monitor the newly input unrecognized words, calculate the word frequency of each new word ω, and weight according to the domain characteristics of the vocabulary to enhance the weights of professional terms and high-frequency equipment names. The calculation formula is as follows: where \(F(\omega)\) is the weighted word frequency, \(f(\omega)\) is the frequency of the word in all texts, \(N\) is the total number of words, \(f domain (\omega)\) is the frequency of the word \(\omega\) in the domain dictionary, \(N domain is the total number of words in the domain dictionary, and \(\alpha\) and \(\beta\) are adjustment parameters used to balance the influence of the domain dictionary and the ordinary word frequency. Usually, \(\alpha+\beta = 1\). Set a dynamic threshold so that new words with a frequency exceeding this threshold automatically enter the domain dictionary. The update formula for the threshold is: where θ t is the threshold value updated at the t-th time, γ is the adjustment coefficient of the threshold value, and θ t-1 is the previous threshold value, and η is the learning rate; M is the number of newly appeared words in the current round, and F(ωi) is the weighted word frequency of these new words; The threshold update ensures the adaptability of the system, enabling the system to automatically adapt to different scenarios and perform dynamic updates as new equipment is added; S2.2: The custom word segmentation rule and named entity recognition module in Jieba can accurately identify equipment names, action instructions, locations, and times. Using the custom word segmentation rule function of Jieba, set the key entities in the operation ticket as the words to be preferentially recognized. At the same time, combine with the deep learning-based NER model to improve the recognition of uncommon nouns, and define word segmentation rules such as "circuit breaker number + serial number" and action word recognition rules such as "cut off + equipment name". When targeting specific tasks, define the following weighted dictionary optimization formula: weight(ω i ) = β·TF-IDF(ω i ) + γ·X(ω i ) Among them, β and γ are hyperparameters used to balance the influence of word frequency and context factors; TF-IDF(ω i ) is the traditional word weight based on word frequency and document frequency; X(ω i ) is the context-based weight that reflects the importance of a word in a specific operation step; Definition Let the defined word set be W, where each word ωi ∈ W, and its corresponding named entity category C(ω i ) ∈ {device, action, location, time}. The system represents the category of each word through the NER model. According to the domain adaptive dictionary D and the dynamic update mechanism, the word frequency and entity category are weighted and adjusted: Among them, score(ω i ) represents the original score of the word in the source text, based on word frequency, context information, etc.; α(C(ω i )) is the coefficient weighted according to the entity category of the word to meet the requirements of a specific field; the weight adjustment of action instructions and device names is dynamically learned through frequency analysis of historical data; S2.3: The context-based word segmentation optimization module in the improved version of Jieba dynamically optimizes the word segmentation strategy according to different contexts of the operation ticket. By analyzing different contexts in historical work tickets, adjust the word segmentation strategy, pay more attention to identifying equipment status words in the safety inspection step, and pay more attention to identifying action words in the operation step, and automatically increase the weights of the corresponding words.
5. The method for generating an emergency operation ticket for a distribution network based on multimodal fusion according to claim 1, wherein In step S2, the invoicing model is a Transformer model.
6. The method for generating an emergency operation ticket for a distribution network based on multimodal fusion according to claim 2, wherein, Including: S2.4: Perform word embedding in the Transformer model, that is, convert the keyword information into a vector of a fixed size; S2.5: Add positional encoding in the Transformer model to provide the model with the position information of words, and add the positional encoding to the word embedding vector; S2.6: The self-attention mechanism in the Transformer model allows the model to consider all other words in the sequence while processing a certain word, thereby capturing long-range dependencies. This feature gives the Transformer great advantages in dealing with the sorting problem in operation tickets. In addition, each layer in the Transformer contains a self-attention mechanism and a feed-forward neural network, which can gradually extract the deep features of the text, and then accurately understand the polysemous words and specific operation intentions in the work ticket; S2.7: A series of vectors output by the Transformer model are fed into the decoder for decoding to generate the target text, that is, to generate a normalized and sequential electronic operation ticket.
7. The method for generating an emergency operation ticket for a distribution network based on multimodal fusion according to claim 1, wherein, In step S3, it includes: the image recognition model is an improved YOLOv8 model, specifically: S3.1: Replace the Bottleneck in C2f of YOLOv8 with GhostBottleneck; The described GhostBottleneck is composed of two stacked Ghost modules. Batch normalization and ReLu activation function are applied after the first Ghost module, and only batch normalization is performed after the second Ghost module; S3.2: Add deformable convolution (DCNv2) after C2f in the Backbone. The features output by the standard convolution mode for the center point a0 are: Where x is the input feature map, y is the output feature map, a0 is the true coordinate of the center point in the output feature map, and a n is the position of the convolution kernel, is the weight at the position of a n and N is the number of sampling points; After sampling the input feature map by deformable convolution, an offset Δa is introduced n ; S3.3: Introduce a modulation mechanism so that the deformable convolutional neural network module can not only adjust the offset of the input feature perception, but also adjust the amplitude of the input features from different spatial positions; where Δl n is the weight modulation parameter at the sampling point a n , and p(a n ) is the penalty term; S3.4: Add a channel-spatial attention integration module with adaptive ability after SPPF in the Backbone; The channel attention module first processes the input feature map with global average pooling and global max pooling to extract channel information, then generates the weight coefficients of the channels through a convolutional layer, and finally passes through a sigmoid layer to apply the weights to the original feature map. The described channel attention module can effectively enhance the key feature channels, improve the sensitivity of the model to important information, and thus improve the detection and recognition accuracy of on-site pictures in the emergency situation of the distribution network; The spatial attention module first performs global average pooling and global max pooling on the input feature map to generate two-channel feature maps, then inputs them into a convolutional layer in a splicing manner to generate a spatial attention map, and finally applies the spatial attention map to the original feature map with a sigmoid function. The module can improve the model's evaluation of the importance of spatial regions, thereby improving the localization accuracy of the model in complex scenarios of the distribution network; The described CSA can adaptively select and adjust the channel weights and spatial weights of the feature map, and can reduce the computational amount while better capturing image features; Add a CSA module before the CBS in the Neck; S3.5: The classification loss function is VFL, and the regression loss adopts the form of CIoU Loss + DFL; Where S R and S p are the areas of the ground truth box and the predicted box respectively, IoU is the intersection over union of the predicted box and the ground truth box, m is the predicted class score, and n is the predicted object score. If it is the true class, n = IoU; if it is other classes, n = 0. where c and c t represent the centers of the predicted target box and the ground truth box respectively, l is the diagonal length of the smallest box that can enclose both the predicted box and the ground truth box, β is the penalty coefficient, w t and w are the widths of the ground truth box and the predicted box respectively, h t and h are the lengths of the ground truth box and the predicted box; where y i is the predicted value and y is the true value; S3.6: Use transfer learning technology to optimize the model. The early layers of the pre-trained model are used to extract general features such as edges, textures, and shapes in the image. After preliminary pre-training, the weights of the early layers are frozen.
8. The method for generating an emergency operation ticket for a distribution network based on multi-modal fusion according to claim 1, wherein In step S4, data augmentation processing includes horizontal flipping, scaling, cropping, color adjustment, and adding slight blurring, brightness changes, and color distortion to the pictures taken on-site or by drones.
9. The method for generating an emergency operation ticket for a distribution network based on multimodal fusion according to claim 1, wherein Including: In step S3, during the training of the image recognition model, freeze the weights of the early layers of the image recognition model; In step S5, during the training of the image recognition model, adjust the weights of the later layers of the image recognition model.
10. A distribution network emergency operation ticket generation system based on multimodal fusion, characterized in that, Including: Cleaning module: Obtain historical work ticket data as the source sequence, and clean the source sequence to obtain the cleaned source sequence; Text training module: Build an invoicing model, extract keywords from the cleaned source sequence, and then train the invoicing model to obtain a text invoicing model; Image pre-training module: Obtain power distribution network emergencies, build an image dataset and an image recognition model, and pre-train the image recognition model according to the image dataset to obtain a pre-trained image recognition model; Specific module: Build a specific task dataset and perform data augmentation processing to obtain an enhanced specific task dataset; Multimodal module: Train the pre-trained image recognition model and the text invoicing model according to the enhanced specific task dataset to obtain a multimodal power distribution network emergency operation ticket generation model.
Citation Information
Cited By
Safety measure ticket intelligent semantic error correction method and system based on OCR technology
CN121170803A