Fire hazard identification method and device for power transmission line based on multi-modal large model

By constructing a multimodal large model and combining image and text feature fusion, the problem of the inability to directly determine the threat of fire to power grid equipment in existing technologies has been solved, realizing an upgrade from smoke and fire detection to equipment threat assessment and improving risk assessment capabilities.

CN120852793BActive Publication Date: 2025-11-21STATE GRID SICHUAN ELECTRIC POWER CORP ELECTRIC POWER RES INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511340011.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-19
Publication Date
2025-11-21
Estimated Expiration
2045-09-19

AI Technical Summary

Technical Problem

While existing deep learning-based target detection algorithms are highly accurate in detecting wildfires, they cannot directly determine the threat posed by fires to power grid equipment, leading to resource waste and untimely crisis response.

Method used

A method for identifying wildfire hazards based on a multimodal large model is constructed. By extracting image features, text encoding, feature fusion and large language model, combined with a wildfire hazard identification guidance dialogue set, a deep fusion of smoke and fire images and text semantics is achieved, generating wildfire hazard output dialogue to determine whether there is a risk to power transmission lines.

Benefits of technology

This technology upgrade, encompassing smoke and fire detection and equipment threat assessment, avoids resource waste and delayed crisis response, and enhances the ability to assess the risks posed by smoke and fire to power transmission lines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120852793B_ABST
    Figure CN120852793B_ABST
Patent Text Reader

Abstract

The application discloses a power transmission line mountain fire hidden danger identification method and device based on a multi-modal large model, and aims at the image containing smoke or flame output by a traditional target detection model to construct a mountain fire hidden danger identification model. The model combines a mountain fire hidden danger identification guided dialogue set to fuse dialogue context and image visual features, thereby realizing deep fusion of text semantics and smoke and fire image features. Finally, the features after the deep fusion are input into a large language model, the deep semantic understanding capability of the large language model is utilized to obtain a mountain fire hidden danger output dialogue of the power transmission line, and then a mountain fire hidden danger identification result of the power transmission line is obtained. Thus, the application realizes judgment on whether the smoke and fire will endanger the power transmission line, and completes technical upgrading from smoke and fire existence detection to equipment threat research and judgment. Based on this, the problems of resource waste and delayed crisis response caused by the traditional technology of reporting all smoke and fire information can be avoided.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of power grid monitoring, in particular to a forest fire hidden danger identification method and device for a power transmission line based on a multi-modal large model. BACKGROUND

[0002] With the continuous expansion of the power grid scale, power transmission lines are distributed more and more densely. In addition, the complex terrain and dense vegetation in the vast forest areas in China make a large number of power transmission lines have to pass through forest areas. When a forest fire occurs in a forest area, the forest fire will pose a direct threat to external insulation facilities, increasing the risk of tripping and outage accidents of power transmission lines. Therefore, it is crucial to realize real-time monitoring of forest fires and accurately assess whether forest fires pose a threat to power grid equipment to ensure the safe operation of power grid equipment.

[0003] At present, with the widespread application of power transmission line channel visualization systems, target detection algorithms based on deep learning have achieved good results in the field of hidden danger monitoring such as forest fires by using real-time video images collected. In practical applications, although target detection algorithms based on deep learning have high accuracy in forest fire detection, they cannot directly determine the threat of fire to power grid equipment. In addition, due to the large amount of data and hidden targets, the workload of online monitoring and operation and maintenance has increased dramatically. Therefore, if only all smoke and fire information is detected and reported, it is easy to cause resource waste, thereby affecting the timely response to real crises of the line. Therefore, based on the foregoing deficiencies, how to provide a forest fire hidden danger identification method for a power transmission line based on a multi-modal large model that can directly determine the threat of fire to power grid equipment has become a problem to be solved. SUMMARY

[0004] The technical problem to be solved by the present application is the identification of forest fire hidden dangers of power transmission lines. The purpose is to provide a forest fire hidden danger identification method and device for a power transmission line based on a multi-modal large model, which solves the problem of resource waste and reduced crisis response efficiency caused by the fact that traditional technology can only detect forest fires but cannot directly determine the threat of fire to power grid equipment, and all smoke and fire information is reported.

[0005] The present application is achieved by the following technical solutions:

[0006] In a first aspect, a forest fire hidden danger identification method for a power transmission line based on a multi-modal large model is provided, comprising:

[0007] obtaining a power transmission line image with a smoke and fire target and a forest fire hidden danger identification guide conversation set;

[0008] The wildfire hazard identification guide conversation set and the power transmission line image are input into a wildfire hazard identification model to obtain a wildfire hazard output conversation of the power transmission line, so as to obtain a wildfire hazard identification result of the power transmission line according to the wildfire hazard output conversation, and the wildfire hazard identification result includes that a fire target exists or does not exist a hazard to the power transmission line.

[0009] The wildfire hazard identification model includes an image feature extraction module, a text encoding module, a feature fusion module and a large language model module, and the image feature extraction module is used for performing feature extraction processing on the power transmission line image to obtain image features.

[0010] The text encoding module is used for mapping the wildfire hazard identification guide conversation set into a text semantic vector.

[0011] The feature fusion module is used for performing feature fusion processing on the image features and the text semantic vector by using a cross-attention mechanism to obtain enhanced image features with fused text semantics and enhanced text features with fused image features, and performing feature fusion processing on the enhanced image features and the text semantic vector to obtain enhanced conversation text.

[0012] The large language model module is used for generating the wildfire hazard output conversation based on the enhanced image features, the enhanced text features and the enhanced conversation text.

[0013] Based on the above disclosure, the present application constructs a multi-modal large model (i.e. a wildfire hazard identification model) for the output of a traditional target detection model, wherein the model integrates image feature extraction, text encoding, feature fusion and a large language model, and the present application also adds a wildfire hazard identification guide conversation set to the power transmission line image. In this way, the wildfire hazard identification guide conversation set and the power transmission line image can be combined and the model can be used to realize wildfire hazard identification. Specifically, the image feature extraction module is used to realize feature extraction of the power transmission line image, then the text encoding module is used to map the wildfire hazard identification guide conversation set into a text semantic vector. Then, the feature fusion module based on the cross-attention mechanism is used to realize deep fusion of text semantics and image features to obtain enhanced image features with fused text semantics and enhanced text features with fused image features. Meanwhile, the module also performs feature fusion on the enhanced image features and the text semantic vector to obtain enhanced conversation text. Finally, the enhanced image features, the enhanced text features and the enhanced conversation text are input into the large language model module, and the deep semantic understanding ability of the large language model is used to obtain the wildfire hazard output conversation of the power transmission line. Based on this, the wildfire hazard identification result of the power transmission line can be obtained according to the wildfire hazard output conversation output by the model, that is, whether the wildfire exists a hazard to the power transmission line.

[0014] Through the above design, the application constructs a forest fire hidden danger recognition model for the image containing smoke or flame output by the traditional target detection model, the model fuses the dialogue context and the image visual features by combining the forest fire hidden danger recognition guide dialogue set, so that the depth fusion of the text semantic and the smoke image features is realized, finally, the aforementioned depth fused features are input into the large language model, then the large language model can be used to obtain the forest fire hidden danger output dialogue of the power transmission line, and then the forest fire hidden danger recognition result of the power transmission line is obtained, thus, the application realizes the judgment of whether the smoke and flame will endanger the power transmission line, and completes the technical upgrade from the smoke and flame existence detection to the equipment threat research and judgment, based on this, the resource waste and the problem of not timely crisis response caused by the traditional technology of reporting all smoke and flame information can be avoided, therefore, the application is very suitable for large-scale application and promotion.

[0015] In a possible design, the image feature extraction module comprises an image feature extraction unit and an image text projection unit.

[0016] The image feature extraction unit is configured to extract image semantic features from the power transmission line image to obtain first initial image features, and the image feature extraction unit comprises a SigLIP visual model.

[0017] The image text projection unit is configured to perform secondary feature extraction processing on the first initial image features to obtain second initial image features, and perform feature dimension adjustment processing on the second initial image features to obtain the image features with the same feature dimension as the text semantic vector.

[0018] In a possible design, the image text projection unit comprises a first linear layer and a second linear layer connected in sequence, wherein the first linear layer is configured to perform secondary feature extraction processing on the first initial image features to obtain second initial image features, and the second linear layer is configured to perform feature dimension adjustment processing on the second initial image features to obtain the image features after the feature dimension adjustment processing.

[0019] In a possible design, the first prompt word image fusion unit, the second prompt word image fusion unit and the dialogue text image fusion unit;

[0020] The first prompt word image fusion unit is configured to perform one cross-modal feature fusion processing on the image features and the text semantic vector by using a cross-attention mechanism to obtain image weighted features fused with text semantics and text weighted features fused with image features, and perform feature enhancement processing on the text weighted features to obtain first initial enhanced text features fused with image features.

[0021] The second prompt word image fusion unit is configured to perform secondary cross-modal feature fusion processing on the first initial enhanced text feature and the image weighted feature by using a cross-attention mechanism, to obtain a second initial enhanced text feature of a fusion image feature and an enhanced image feature of a fusion text semantic;

[0022] The second prompt word image fusion unit is further configured to perform feature enhancement processing on the second initial enhanced text feature of the fusion image feature, to obtain an enhanced text feature of the fusion image feature after the feature enhancement processing.

[0023] The dialogue text image fusion unit is configured to perform feature fusion processing on the enhanced image feature and the text semantic vector, to obtain an enhanced dialogue text after the feature fusion processing.

[0024] In one possible design, the first prompt word image fusion unit includes an image-text cross-attention layer, a text-image cross-attention layer, a self-attention layer, and a feed-forward neural network layer.

[0025] The image-text cross-attention layer is configured to calculate a first attention weight of the image feature on the text semantic vector, and perform feature weighting processing on the text semantic vector by using the first attention weight, to obtain an image weighted feature of a fusion text semantic after the feature weighting processing.

[0026] The text-image cross-attention layer is configured to calculate a second attention weight of the text semantic vector on the image weighted feature, and perform feature weighting processing on the image weighted feature by using the second attention weight, to obtain a text weighted feature of a fusion image feature after the feature weighting processing.

[0027] The self-attention layer is configured to perform primary feature enhancement processing on the text weighted feature, to obtain a primary enhanced text feature.

[0028] The feed-forward neural network layer is configured to perform secondary feature enhancement processing on the primary enhanced text feature, to obtain a first initial enhanced text feature of the fusion image feature after the secondary feature enhancement processing.

[0029] In one possible design, the first attention weight is calculated in the following manner.

[0030] A first query matrix is generated based on the image feature, and a first key matrix is generated according to the text semantic vector.

[0031] The first query matrix and the first key matrix are used to calculate the first attention weight in the following formula (1).

[0032] (1)

[0033] In the above formula (1), y represents the first attention weight, SoftMax() represents an activation function of the image-text cross attention layer, represents the first query matrix, K l represents the first key matrix, M l represents a mask matrix of the text semantic vector, d h represents an attention scaling factor, wherein, , , and V represents the image feature, L represents the text semantic vector, all represent projection matrices of the image-text cross attention layer.

[0034] In one possible design, the dialogue text-image fusion unit comprises: a cross attention layer based on a gating mechanism, a first feature concatenation layer, a feedforward neural network layer based on a gating mechanism, and a second feature concatenation layer;

[0035] The cross attention layer based on the gating mechanism is configured to perform image-text feature fusion processing on the text semantic vector and the enhanced image feature to obtain a first initial enhanced dialogue text.

[0036] The first feature concatenation layer is configured to perform feature concatenation processing on the first initial enhanced dialogue text and the text semantic vector to obtain a concatenated dialogue text.

[0037] The feedforward neural network layer based on the gating mechanism is configured to perform feature enhancement processing on the concatenated dialogue text to obtain a second initial enhanced dialogue text.

[0038] The second feature concatenation layer is configured to perform feature concatenation processing on the second initial enhanced dialogue text and the concatenated dialogue text to obtain the enhanced dialogue text after the feature concatenation processing.

[0039] In one possible design, the first initial enhanced dialogue text output by the cross attention layer based on the gating mechanism is generated using the following formula (2).

[0040] (2)

[0041] In formula (2), y1 represents the first initial enhanced dialogue text, y' represents the text semantic vector, a represents a first learning parameter, tanh() represents a hyperbolic tangent function, and Attention(x, y) represents output data of a cross attention network in the cross attention layer based on the gating mechanism.

[0042] ​Correspondingly, the second initial enhanced dialogue text output by the feedforward neural network layer based on the gating mechanism is generated using the following formula (3);

[0043] (3)

[0044] In formula (3), y2 represents the second initial enhanced dialogue text, β represents a second learning parameter, and FFN (y1) represents output data of the feedforward neural network in the feedforward neural network layer based on the gating mechanism.

[0045] In one possible design, the image feature extraction module includes an image feature extraction unit and an image text projection unit, and the wildfire hazard identification model is obtained by training in the following manner.

[0046] A training data set is obtained, where the training data set includes a plurality of power transmission line sample images and a wildfire hazard identification guide sample dialogue set corresponding to each power transmission line sample image.

[0047] The initial wildfire hazard identification model is trained by taking each power transmission line sample image and the corresponding wildfire hazard identification guide sample dialogue set in the training data set as input and taking the corresponding wildfire hazard output dialogue as output, so that the wildfire hazard identification model is obtained after training.

[0048] In the training process, the learning rate of the initial wildfire hazard identification model is determined according to the training round number, and the update step size in each training round is determined using the learning rate in each training round, so that the model parameters of the initial wildfire hazard identification model are adjusted in reverse using the update step size in each training round. In the training process, the low-rank adaptive fine-tuning algorithm is used to adjust the model parameters in the large language model module, and the full-volume fine-tuning algorithm is used to adjust the model parameters in the feature fusion module and the image text projection unit.

[0049] In a second aspect, a wildfire hazard identification device for a power transmission line based on a multi-modal large model is provided, which includes:

[0050] An acquisition unit is configured to acquire a power transmission line image with a wildfire target and a wildfire hazard identification guide dialogue set.

[0051] A hazard identification unit is configured to input the wildfire hazard identification guide dialogue set and the power transmission line image into a wildfire hazard identification model to obtain a wildfire hazard output dialogue for the power transmission line, and to obtain a wildfire hazard identification result for the power transmission line according to the wildfire hazard output dialogue. The wildfire hazard identification result includes whether the wildfire target is harmful or harmless to the power transmission line.

[0052] The mountain fire hidden danger identification model comprises an image feature extraction module, a text encoding module, a feature fusion module, and a large language model module, and the image feature extraction module is configured to perform feature extraction processing on the power transmission line image to obtain image features.

[0053] The text encoding module is configured to map the mountain fire hidden danger identification guided dialogue set into a text semantic vector.

[0054] The feature fusion module is configured to perform feature fusion processing on the image features and the text semantic vector by using a cross-attention mechanism to obtain enhanced image features fused with text semantics and enhanced text features fused with image features, and to perform feature fusion processing on the enhanced image features and the text semantic vector to obtain enhanced dialogue text.

[0055] The large language model module is configured to generate the mountain fire hidden danger output dialogue based on the enhanced image features, the enhanced text features, and the enhanced dialogue text.

[0056] In a third aspect, another device for identifying a mountain fire hidden danger of a power transmission line based on a multi-modal large model is provided, which is taken as an electronic device and comprises a memory, a processor, and a transceiver connected in sequence and in communication, wherein the memory is configured to store a computer program, the transceiver is configured to transceive messages, and the processor is configured to read the computer program and execute the method for identifying a mountain fire hidden danger of a power transmission line based on a multi-modal large model as in the first aspect or any possible design in the first aspect.

[0057] In a fourth aspect, a storage medium is provided, and the storage medium stores instructions, which, when executed on a computer, execute the method for identifying a mountain fire hidden danger of a power transmission line based on a multi-modal large model as in the first aspect or any possible design in the first aspect.

[0058] In a fifth aspect, a computer program product containing instructions is provided, which, when executed on a computer, causes the computer to execute the method for identifying a mountain fire hidden danger of a power transmission line based on a multi-modal large model as in the first aspect or any possible design in the first aspect.

[0059] Compared with the prior art, the present application has the following advantages and beneficial effects:

[0060] (1) The present application aims at the image containing smoke or flame output by the traditional target detection model, and constructs a forest fire hidden danger recognition model. The model combines the forest fire hidden danger recognition guided dialogue set to fuse the dialogue context and image visual features, thereby realizing the deep fusion of text semantics and smoke and fire image features. Finally, the aforementioned deeply fused features are input into a large language model, and the deep semantic understanding ability of the large language model is used to obtain the forest fire hidden danger output dialogue of the power transmission line, and then the forest fire hidden danger recognition result of the power transmission line is obtained. Therefore, the present application realizes the judgment of whether the smoke and fire will endanger the power transmission line, completes the technical upgrading from smoke and fire existence detection to equipment threat judgment, and based on this, the problems of resource waste and crisis response not in time caused by the traditional technology of reporting all smoke and fire information can be avoided, so the present application is very suitable for large-scale application and popularization.

[0061] (2) The present application provides a feature fusion module with double cross-modal cross-attention mechanisms and gate attention mechanisms, which uses the text semantic vector generated by the multi-round dialogue data to deeply fuse the text semantics and image features, so that the potential correlation between the fire point visual features and the dialogue text semantics can be more accurately captured, and the risk judgment ability of the smoke and fire threat to the power transmission line can be effectively improved. BRIEF DESCRIPTION OF DRAWINGS

[0062] In order to more clearly illustrate the technical solutions of the example embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can also be obtained without creative labor. In the drawings:

[0063] Figure 1 The step flowchart of the forest fire hidden danger recognition method for power transmission lines based on the multi-modal large model provided by the embodiments of the present application;

[0064] Figure 2 The structure diagram of the forest fire hidden danger recognition model provided by the embodiments of the present application;

[0065] Figure 3 The structure diagram of the feature fusion module provided by the embodiments of the present application;

[0066] Figure 4 The structure diagram of the forest fire hidden danger recognition device for power transmission lines based on the multi-modal large model provided by the embodiments of the present application;

[0067] Figure 5 The structure diagram of the electronic device provided by the embodiments of the present application. DETAILED DESCRIPTION

[0068] In order to make the objects, technical solutions and advantages of the present application clearer, further detailed description will be made to the present application in combination with embodiments and drawings, the illustrative embodiments of the present application and the description thereof are only used to explain the present application and do not constitute limitation to the present application; it should be understood that although the terms first, second, etc. can be used herein to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another unit. For example, the first unit can be called the second unit, and similarly, the second unit can be called the first unit, without departing from the scope of the example embodiments of the present application.

[0069] Embodiment 1: refer to Figure 1 As shown in the figure, the mountain fire hidden danger identification method for power transmission lines based on the multi-modal large model provided in this embodiment provides a mountain fire hidden danger identification model for power transmission lines based on a multi-modal large model, wherein the SigLIP visual model is used to extract image features of the power transmission lines, then a text encoding module is used to map the mountain fire hidden danger identification guided dialogue set into a continuous text semantic vector, then an image text projection unit containing two linear layers connected in sequence is used to keep the image features and the text features consistent in the feature space and refine the image features, so as to facilitate the deep fusion of the image features with the text mode in the subsequent feature fusion module, wherein the feature fusion module is constructed based on a double cross-modal cross-attention mechanism and a gated attention mechanism, and the role is to focus the large language model on the flame, smoke and power facility related areas in the image through the deep fusion of the text and image information, and deeply understand the multi-round dialogue context, so as to improve the judgment accuracy; finally, the fused image features and text features are sent into the large language model Phi-2 together, which is used to generate the final mountain fire hidden danger output dialogue, so that the mountain fire hidden danger identification result can be obtained according to the mountain fire hidden danger output dialogue; based on this, the method realizes the judgment of whether the smoke and fire will endanger the power transmission lines, and completes the technical upgrade from smoke and fire detection to equipment threat research and judgment, thus avoiding the problems of resource waste and crisis response not in time caused by the traditional technology of reporting all smoke and fire information, and therefore, the method is very suitable for large-scale application and promotion; wherein the method can be but is not limited to running at the power transmission line monitoring end or the server side, it can be understood that the foregoing execution subject does not constitute limitation to the embodiments of the present application, and correspondingly, the running steps of the method can be but are not limited to the steps S1-S2 as shown below.

[0070] S1. Obtain the power transmission line image with the fire target and the set of fire hazard identification guided dialogues; in specific applications, for example, the power transmission line image can be obtained first, then the power transmission line image is input into the target detection model, so as to determine whether there is a fire target in the power transmission line image, wherein if there is a fire target, the power transmission line image is retained, otherwise it is discarded; in this way, the power transmission line image with the fire target can be obtained; for example, the target detection model can be trained by using a trained YOLOV8 model, that is, the target detection model is trained by using a plurality of power transmission line sample images as input and the fire detection results of each power transmission line sample image as output; of course, other target detection models can also be used, which are not limited to the foregoing examples.

[0071] After obtaining the power transmission line image with the fire target, the set of fire hazard identification guided dialogues can be obtained, wherein for example, the set of fire hazard identification guided dialogues can include multi-round fire hazard identification dialogue data, which contains fire hazard identification prompt words, that is, guided sentences for fire threat assessment of power transmission lines; at the same time, for example, the multi-round fire hazard identification dialogue data can be pre-set, and the detailed dialogue data can be shown in Table 1.

[0072] Table 1 is the multi-round fire hazard identification dialogue data.

[0073] Table 1

[0074] Question Answer Is there fire or smoke in the picture? 1. Yes, there is smoke in the picture; 2. Yes, there is both smoke and fire in the picture; 3. Yes, there is fire in the picture; Describe the size of the smoke or fire 1. The smoke is very small; 2. The smoke is very large; 3. The fire is very small; 4. The fire is very large 5. The smoke is very large and the fire is very large; 6. The smoke is very small and the fire is very large; 7. The smoke is very large and the fire is very small; 8. The smoke is very small and the fire is also very small; Will the fire or smoke affect the power infrastructure? 1. Yes, the smoke or fire is close to the power infrastructure and may cause significant damage; 2. No, the smoke or fire may not directly affect the power infrastructure, but if not properly controlled, it may cause air pollution or other hazards; 3. No, although there is no power infrastructure nearby, there is currently a large amount of smoke and fire that needs to be controlled in a timely manner; Please provide the bounding box coordinates of the smoke or fire Provide the bounding box coordinates according to the specific situation

[0075] In this way, the multi-round fire hazard identification dialogue data can be obtained through Table 1; then, the multi-round fire hazard identification dialogue data and the power transmission line image can be input into the fire hazard identification model, so as to obtain the fire hazard output dialogue of the power transmission line, and then determine whether the fire will cause harm to the power transmission line according to the fire hazard output dialogue; wherein the fire hazard identification process can be as shown in the following step S2.

[0076] S2. Input the set of fire hazard identification guided dialogues and the power transmission line image into the fire hazard identification model to obtain the fire hazard output dialogue of the power transmission line, so as to obtain the fire hazard identification result of the power transmission line according to the fire hazard output dialogue, and the fire hazard identification result includes whether the fire target causes harm to the power transmission line or not; in this embodiment, the model output is the dialogue answer to the third question in Table 1, that is, whether the fire or smoke will affect the power infrastructure; in this way, the fire hazard identification result can be obtained according to the output dialogue.

[0077] In a specific application, the embodiment is a forest fire hazard identification, a multi-modal large model is constructed, which integrates a large language model, a visual model (SigLIP), and a text and image modal fusion module based on a double cross-modal attention mechanism and a gate attention mechanism (i.e. a feature fusion module); wherein the multi-modal large model uses the logical reasoning, context understanding and cross-modal semantic fusion capability of the multi-modal large language model to accurately determine whether the smoke and fire will pose a threat to the power transmission line; of course, the aforementioned forest fire hazard identification model is a trained model, that is, a set of power transmission line sample images and a corresponding forest fire hazard identification guide conversation set for each power transmission line sample image are input, and the output of the forest fire hazard identification conversation corresponding to each power transmission line sample image is output to obtain the trained model; optionally, the training process is described in detail below.

[0078] In a specific application, one of the network structures of the disclosed forest fire hazard identification model is combined with the network structure to describe the identification process of the forest fire hazard.

[0079] Among them, referring to Figure 2 As shown, for example, the forest fire hazard identification model can include but is not limited to an image feature extraction module, a text encoding module, a feature fusion module, and a large language model module; wherein the detailed working process of each module in the model is:

[0080] First, the text encoding module is used to map the forest fire hazard identification guide conversation set to a text semantic vector; in a specific implementation, for example, the text encoding module can include but is not limited to the tokenizer and embedding layer in the Phi-2 large language model, wherein the tokenizer in the Phi-2 large language model can be loaded first to convert the conversation data in the forest fire hazard identification guide conversation set into discrete token IDs using the tokenizer, and then the embedding layer in the Phi-2 large language model is used to map the foregoing discrete token IDs to continuous text semantic vectors.

[0081] Further, when performing semantic mapping, one round of conversation data can be used as input to the text encoding module in combination with the text semantic vector output by the text encoding module in the previous round, so that the text semantic vector of the current round of conversation data can be generated; specifically, the process is:

[0082] From the multi-round fire hazard identification dialogue data, the i-th round dialogue data (i.e., the first question and its answer in Table 1) is filtered out, and the initial text semantic vector corresponding to the i-1-th round dialogue data is obtained; then, the initial text semantic vectors corresponding to the i-th round dialogue data and the i-1-th round dialogue data are input into the text encoding module to obtain the initial text semantic vector corresponding to the i-th round dialogue data; then, i is incremented by 1, and the i-th round dialogue data is re-filtered from the multi-round fire hazard identification dialogue data, until i is equal to n, the initial text semantic vector corresponding to the i-th round dialogue data is taken as the text semantic vector of the fire hazard identification guide dialogue set; wherein the initial value of i is 1, n is the number of dialogue rounds of the multi-round fire hazard identification dialogue data, and when i is 1, the initial text semantic vector corresponding to the i-1-th round dialogue data is empty, that is, only the dialogue data of the first round is input in the first round.

[0083] In this way, through the aforementioned text encoding module, the multi-round dialogue data can be converted into a text semantic vector with continuous semantics and associated context.

[0084] After obtaining the text semantic vector, image feature extraction can be performed, that is:

[0085] The image feature extraction module is configured to perform feature extraction processing on the power transmission line image to obtain image features. In this embodiment, the image feature extraction module is mainly responsible for extracting image semantic features from the power transmission line image, and performing feature refinement and feature dimension adjustment on the extracted features to obtain image features with the same dimension as the output of the text encoding module.

[0086] As shown in Figure 2 The image feature extraction module can include, but is not limited to, an image feature extraction unit and an image text projection unit. In specific applications, the image feature extraction unit is configured to extract image semantic features from the power transmission line image to obtain first initial image features, i.e., the first initial image features contain image information such as target categories, scene types, and texture shapes in the power transmission line image. Then, the image text projection unit is configured to perform secondary feature extraction processing on the first initial image features to obtain second initial image features, and perform feature dimension adjustment processing on the second initial image features to obtain image features with the same feature dimension as the text semantic vector.

[0087] In this embodiment, the image feature extraction unit may, but is not limited to, include a SigLIP vision model, that is, the image semantic features in the power transmission line image are extracted by using the SigLIP vision model. Similar to CLIP (a multi-modal pre-training model), the SigLIP vision model uses a contrastive learning mechanism to realize cross-modal semantic alignment through semantic matching of image-text pairs, and thus the image semantic features in the power transmission line image can be accurately extracted.

[0088] Meanwhile, one of the network structures of the disclosed image-text projection unit is as follows:

[0089] In this embodiment, the image-text projection unit may, but is not limited to, include a first linear layer and a second linear layer connected in sequence. The first linear layer is configured to perform secondary feature extraction on the first initial image features to obtain second initial image features, and the second linear layer is configured to perform feature dimension adjustment on the second initial image features to obtain the image features after the feature dimension adjustment. In this way, the first linear layer performs feature extraction, that is, removes redundant information in the first initial image features (for example, by adding L1 regularization to the weight matrix to make some weights zero, thereby filtering out some features), so that the image features are more compact and are ready for subsequent fusion with text features. The second linear layer converts the image modal (that is, the second initial image features) into the same feature dimension as the text modal (that is, the text semantic vector). Through the conversion of this layer, the image features are matched with the text features in terms of dimension, and can be effectively aligned in the feature space, thereby facilitating the subsequent feature fusion module to deeply fuse the image and text information.

[0090] In this way, after the image feature extraction module and the text encoding module obtain the corresponding image features and the text semantic vector of the fusion dialogue context, the feature fusion module is used to deeply fuse the image and text semantics, so that the large language model module can focus on the flame, smoke and power facility related regions in the image and deeply understand the multi-round dialogue context, thereby improving the judgment accuracy.

[0091] The working process of the feature fusion module is as follows:

[0092] The feature fusion module is configured to perform feature fusion processing on the image feature and the text semantic vector by using a cross-attention mechanism to obtain enhanced image features with fused text semantics and enhanced text features with fused image features, and perform feature fusion processing on the enhanced image features and the text semantic vector to obtain enhanced dialogue text. In this embodiment, the feature fusion module first performs bidirectional cross-modal feature interaction between the image and the text by using a cross-attention mechanism, and then performs feature enhancement on the features obtained after the cross-modal interaction by using an attention mechanism, thereby obtaining the enhanced text features. Then, the module further performs feature fusion on the enhanced image features with fused text semantics and the text semantic vector by using a cross-attention module based on a gating mechanism, thereby obtaining the enhanced dialogue text. Finally, the feature fusion module outputs the enhanced image features, enhanced text features and enhanced dialogue text to the large language model module, thereby generating the power line mountain fire hazard output dialogue.

[0093] Optionally, one of the following network structures of the feature fusion module is disclosed:

[0094] Referring to Figure 3 The feature fusion module may, but is not limited to, include a first prompt word image fusion unit, a second prompt word image fusion unit and a dialogue text image fusion unit. The first prompt word image fusion unit and the second prompt word image fusion unit have the same structure, Figure 3 In the figure, the number x2 represents two prompt word image fusion units, and the output of the first prompt word image fusion unit is input to the second prompt word image fusion unit. In this way, the image and text features will undergo two times of bidirectional cross-modal feature interaction.

[0095] The specific working process of each network unit is as follows:

[0096] The first prompt word image fusion unit is configured to perform cross-modal feature fusion processing on the image feature and the text semantic vector by using a cross-attention mechanism to obtain image weighted features with fused text semantics and text weighted features with fused image features, and perform feature enhancement processing on the text weighted features to obtain first initial enhanced text features with fused image features. In this embodiment, the cross-attention mechanism is first used to perform bidirectional feature weighting between the image feature and the text semantic vector, so that the two types of features can perceive and absorb the key information of each other during the weighting process, thereby realizing deep fusion of the text features and the image features. Finally, the attention mechanism is used for feature enhancement, and the first initial enhanced text features are obtained.

[0097] In specific implementation, referring to Figure 3The specific structure of the first prompt word image fusion unit is disclosed, and the bidirectional feature weighting (i.e., cross-modal feature interaction) process is described.

[0098] In a specific embodiment, the first prompt word image fusion unit can include, but is not limited to, an image-text cross-attention layer, a text-image cross-attention layer, a self-attention layer, and a feed-forward neural network layer. The specific working process of each network layer is as follows:

[0099] The image-text cross-attention layer is used to calculate the first attention weight of the image feature on the text semantic vector, and to perform feature weighting processing on the text semantic vector using the first attention weight, so as to obtain the image weighted feature of the fused text semantic after feature weighting processing. In a specific application, a query matrix (Query) is generated based on the image feature, a key (Key) and a value (Value) are generated based on the text feature, and the first attention weight of the image on the text is calculated based on the query matrix, the key matrix, and the value matrix. Then, the value information (Value) of the text feature is weighted and aggregated using the first attention weight, thereby outputting the image weighted feature of the fused text semantic.

[0100] Further, the process of generating the first attention weight is as follows: first, a first query matrix is generated based on the image feature, and a first key matrix is generated based on the text semantic vector; second, the first attention weight is calculated using the first query matrix and the first key matrix. For example, the first attention weight can be calculated using the following formula (1).

[0101] (1)

[0102] In the above formula (1), represents the first attention weight, SoftMax() represents the activation function of the image-text cross-attention layer, represents the first query matrix, K l represents the first key matrix, M l represents the mask matrix of the text semantic vector, d h represents the attention scaling factor, T represents the transpose operation, wherein, , , and V represents the image feature, L represents the text semantic vector, all represent the projection matrix of the image-text cross attention layer (which is the parameter of the network, and the initial value can be preset, and then updated during the training process until the training is completed, and then the optimal projection matrix is obtained); at the same time, the mask matrix is the filling mask of the text semantic vector, which is used to shield the invalid part in the text feature, and is generated in the data preprocessing process. The filling or invalid position in the text token can be marked in advance, that is, for each position of the text semantic vector, if any position is a filling field, it is marked as 1, otherwise it is marked as 0; in this way, the mask matrix can be obtained.

[0103] Therefore, by the foregoing formula (1), after calculating the first attention weight, the value information of the text semantic vector can be weighted processed with the first attention weight, so as to obtain the image weighted feature fused with the text semantic; wherein the text semantic vector can be taken as a value (Value) to directly perform a weighting operation with the first attention weight.

[0104] After completing the feature interaction of image-text, the feature interaction of text-image can be performed, that is, the text-image cross attention layer is used to calculate the second attention weight of the text semantic vector to the image weighted feature, and the image weighted feature is weighted processed by using the second attention weight, so as to obtain the text weighted feature fused with the image feature after the feature weighting processing; in the embodiment, the text-image cross attention layer generates a query matrix based on the text semantic vector, then generates a key and a value based on the image weighted feature, and calculates the second attention weight based on the key and the value. Then, the visual information of the image feature is weighted aggregated (that is, the image weighted feature is weighted aggregated) by using the second attention weight, so as to obtain the text weighted feature fused with the image feature.

[0105] Similarly, the calculation formula of the second attention weight is:

[0106]

[0107] In the formula, indicates the second attention weight, indicates the second query matrix, indicates the second key matrix, indicates the mask matrix of the image weighted feature, wherein, , , and indicates the image weighted feature, and L indicates the text semantic vector, all represent the projection matrix of the text-image cross attention layer.

[0108] Thus, the second attention weight can be calculated based on the foregoing formula, and then the image weighting feature is subjected to feature weighting processing, and the text weighting feature of the fusion image feature can be obtained. At this time, one-way feature weighting can be completed. Then, the self-attention layer can be used for feature enhancement processing, and the process is as follows.

[0109] The self-attention layer is used for one-time feature enhancement processing of the text weighting feature to obtain one-time enhanced text feature. In the embodiment, the working process of the self-attention layer is as follows: first, the input text weighting feature src is added to the position encoding pos, wherein the position encoding pos is used to add position information to each element in the text sequence, and the generation process is based on the periodicity of the sine and cosine functions, that is, the tensor pos_tensor representing the position (a tensor can be initialized for different positions in the text weighting feature) is mapped to a sine / cosine signal with different frequencies, the odd dimension uses cosine encoding, and the even dimension uses sine encoding, thereby generating a unique high-dimensional vector representation for each position. Then, the text weighting feature with added position information is converted into a query vector q, a key vector k and a value vector value. Since the query, the key and the value are all derived from the same input due to the characteristics of the self-attention mechanism, the value vector value directly uses the input text weighting feature src, and the query vector q and the key vector k are obtained by adding the input text weighting feature src to the position encoding pos. Then, the attention score is calculated through the multi-head attention layer, and src_mask (i.e., the mask matrix) is used to shield invalid positions (such as padding areas) in the process, and the attention score is multiplied by the foregoing value vector, thereby generating an enhanced feature src2 (i.e., the foregoing one-time enhanced text feature) containing context dependency. Of course, the calculation principle of the attention score of the self-attention mechanism is the same as the foregoing first attention weight, and the calculation process is not described again.

[0110] After one-time feature enhancement is completed, the feedforward neural network layer can be used for second-time feature enhancement, that is, the feedforward neural network layer is used for second-time feature enhancement processing of the one-time enhanced text feature, so as to obtain the first initial enhanced text feature of the fusion image feature after the second-time feature enhancement processing.

[0111] In a specific application, the secondary feature enhancement process is as follows: first, the primary enhanced text feature is connected in residual connection with the original input (i.e., the text weighted feature is spliced with the aforementioned text weighted feature), to obtain the spliced feature; then, the feature is normalized by the first normalization layer, to obtain the normalized feature (which serves to stabilize the feature distribution); then, the normalized feature is input into the feedforward neural network composed of linear1, ReLU activation function, Dropout and linear2 to further extract nonlinear features; finally, the feature is connected in residual connection with the spliced feature, and after the feature is normalized by the second normalization layer, the aforementioned first initial enhanced text feature can be obtained.

[0112] After the primary bidirectional feature weighting and feature enhancement are completed, secondary bidirectional feature weighting and feature enhancement can be performed, i.e., the second prompt word image fusion unit is configured to perform secondary cross-modal feature fusion processing on the first initial enhanced text feature and the image weighted feature by using a cross-attention mechanism, to obtain an enhanced image feature fused with text semantics and a second initial enhanced text feature fused with image features; then, the second prompt word image fusion unit is further configured to perform feature enhancement processing on the second initial enhanced text feature fused with image features, to obtain an enhanced text feature of the fused image features after the feature enhancement processing.

[0113] In this embodiment, the input of the image-text cross-attention layer in the second prompt word image fusion unit is the first initial enhanced text feature output by the first prompt word image fusion unit and the image weighted feature output by the image-text cross-attention layer in the first prompt word image fusion unit; then, the attention weight of the image weighted feature on the first initial enhanced text feature is calculated in the aforementioned same manner, and the first initial enhanced text feature is weighted, to obtain an enhanced image feature fused with text semantics; then, the enhanced image feature fused with text semantics is input into the text-image cross-attention layer in the second prompt word image fusion unit for feature cross between text and image, to obtain a second initial enhanced text feature fused with image features; then, the second initial enhanced text feature fused with image features is sequentially input into the self-attention layer and the feedforward neural network for feature enhancement, to obtain an enhanced text feature of the fused image features.

[0114] After the bidirectional feature weighting is completed, the aforementioned enhanced image feature and multi-round dialogue text (i.e., text semantic vector) can be input into the dialogue text-image fusion unit to output dialogue context-related cross-modal features, i.e., the dialogue text-image fusion unit is configured to perform feature fusion processing on the enhanced image feature and the text semantic vector, to obtain the enhanced dialogue text, i.e., the aforementioned dialogue context-related cross-modal features, after the feature fusion processing.

[0115] Optionally, one of the following network structures of the aforementioned dialogue text-image fusion unit is disclosed:

[0116] Referring to Figure 3 As shown, the dialogue text image fusion unit can include, but is not limited to, a cross-attention layer based on a gating mechanism, a first feature splicing layer (e.g. Figure 3 indicates feature splicing), a feedforward neural network layer based on a gating mechanism, and a second feature splicing layer; the specific working process of each network layer is as follows: The cross-attention layer based on a gating mechanism is used for image and text feature fusion processing of the text semantic vector and the enhanced image feature to obtain a first initial enhanced dialogue text. In this embodiment, the cross-attention layer based on a gating mechanism adds a hyperbolic tangent function on the basis of a traditional cross-attention layer, that is, the cross-attention layer based on a gating mechanism first performs feature cross between the text semantic vector and the enhanced image feature, that is, calculates the attention weight of the text semantic vector to the enhanced image feature, and then uses the attention weight to weight the enhanced image feature to obtain a dialogue text fused with the image feature. Then, the first initial enhanced dialogue text is generated by inputting to a hyperbolic tangent function layer. For example, the following formula (2) can be used to generate the first initial enhanced dialogue text.

[0117]

[0118] (2) In formula (2), y1 represents the first initial enhanced dialogue text, y' represents the text semantic vector, a represents a first learning parameter, tanh() represents a hyperbolic tangent function, and Attention(x, y) represents the output data of the cross-attention network in the cross-attention layer based on a gating mechanism. In this embodiment, the initial value of a is 0 during training, so that an optimal learning parameter can be learned through continuous training, and then the value of a is mapped to the interval (-1, 1) through the hyperbolic tangent function. In this way, through the training of the first learning parameter a, the model can dynamically adjust the contribution of the output after the cross-attention network to the result.

[0119] After the cross-attention layer based on a gating mechanism, feature splicing can be performed, that is, the first feature splicing layer is used for feature splicing processing of the first initial enhanced dialogue text and the text semantic vector to obtain a spliced dialogue text. Then, the feedforward neural network layer based on a gating mechanism can be used for feature enhancement, and the process is as follows:

[0120]

[0121] ​The feedforward neural network layer based on the gating mechanism is used for feature enhancement processing on the spliced dialogue text to obtain a second initial enhanced dialogue text. In this embodiment, the output of the feedforward neural network layer based on the gating mechanism is as follows:

[0122] (3)

[0123] In formula (3), y2 represents the second initial enhanced dialogue text, β represents a second learning parameter, and FFN(y1) represents output data of the feedforward neural network in the feedforward neural network layer based on the gating mechanism.

[0124] After the feature enhancement by the feedforward neural network layer based on the gating mechanism, residual connection can be performed on the input, that is, a second feature splicing layer is used for feature splicing processing on the second initial enhanced dialogue text and the spliced dialogue text, so as to obtain the enhanced dialogue text after the feature splicing processing.

[0125] Therefore, by using the feature fusion module based on the double cross-modal cross-attention mechanism and the gating attention mechanism, the deep fusion of the text semantic and the image feature can be performed, so that the potential correlation between the fire point visual feature and the dialogue text semantic can be more accurately captured, and the judgment ability of the subsequent large language model on the risk of the forest fire threat to the power transmission line can be effectively improved.

[0126] After obtaining the enhanced image feature, the enhanced text feature, and the enhanced dialogue text, the enhanced image feature, the enhanced text feature, and the enhanced dialogue text can be input into the large language model module, so as to output the forest fire hidden danger output dialogue of the power transmission line by using the deep semantic understanding ability of the large language model. Specifically, the large language model module can be but not limited to Phi-2 large language model, and the output is the answer to the third question in Table 1. In this embodiment, YES and NO in the answer are intercepted as the final output dialogue.

[0127] Therefore, by using the logical reasoning, context understanding, and cross-modal semantic fusion ability of the forest fire hidden danger identification model, the forest fire hidden danger output dialogue of the power transmission line can be output, so that whether the forest fire will threaten the power transmission line can be accurately judged based on the output dialogue.

[0128] In one possible design, one of the training methods of the forest fire hidden danger identification model is provided as follows:

[0129] The first step is to obtain a training data set, wherein the training data set includes a plurality of power line sample images and a corresponding forest fire hazard identification guide sample dialogue set for each power line sample image. In this embodiment, the forest fire hazard identification guide sample dialogue set for each power line sample image is the same as the aforementioned forest fire hazard identification guide dialogue set, and will not be described here.

[0130] Meanwhile, the following discloses the detailed construction process of the aforementioned data set:

[0131] First, collect the image data around the power line from the power grid terminal device cluster, then remove the blurred images (such as using a deep learning-based image quality evaluation algorithm, or outputting the aforementioned image data to the operation and maintenance terminal to enable the operation and maintenance terminal to respond to image removal human-computer interaction operations to complete the removal of blurred images); then, output the images after removing the blurred images to the operation and maintenance terminal for labeling processing, so as to extract effective scene images containing smoke / flame features in response to labeling operations; then, select a plurality of images from the effective scene images as an initial sample set and perform label labeling processing to construct a standardized initial labeling data set containing all 16 types of suspected smoke / fire target types and pixel-level positioning information; then, use the standardized initial labeling data set to train the YOLOV8 target detection model, and then use the trained YOLOV8 target detection model to implement automatic batch labeling on the remaining images, and finally perform labeling deviation correction, so as to form a complete labeling data set covering all scenes and high-precision, that is, the power line sample images in the aforementioned training data.

[0132] After obtaining the power line sample images, the multi-round dialogue set of the images can be constructed for training and testing of the multi-modal large model (i.e., the aforementioned forest fire hazard identification model); Specifically, but not limited to, a JSON labeling template can be first constructed to provide a unified format framework for subsequent labeling work; then, a plurality of image data are selected from the plurality of power line sample images and output to the operation and maintenance terminal for dialogue labeling processing, so as to obtain labeled dialogue data in response to dialogue labeling human-computer interaction operations, wherein these labeled examples will be used as the content of Few-Shot prompts; then, the labeled examples are integrated with the pre-defined template (i.e., the aforementioned JSON labeling template) to construct Few-Shot prompts; thereafter, the remaining image data are batch labeled using the large model to greatly improve the labeling efficiency, and finally, the labeling results are corrected to ensure the accuracy and consistency of the labeling data; wherein the multi-round dialogue data set format is shown in Table 1, and each image corresponds to one of the cases in Table 1 according to the actual situation.

[0133] Thus, the training data set can be constructed in the foregoing manner. Then, the training data set can be used to train the model, and the process is shown in the following second step.

[0134] Second step: training the initial forest fire hazard identification model by taking each power transmission line sample image in the training data set and the corresponding set of forest fire hazard identification guide sample dialogues of each power transmission line sample image as input and taking the output of the forest fire hazard of each power transmission line sample image as output, so as to obtain the forest fire hazard identification model after training, wherein the initial forest fire hazard identification model is an untrained forest fire hazard identification model. In this embodiment, the processing process of any power transmission line sample image and its corresponding set of forest fire hazard identification guide sample dialogues in the model can be referred to the first aspect of the embodiment, and the principle will not be described again.

[0135] Meanwhile, in the training process, if the learning rate is too large, gradient explosion will easily occur, leading to NaN (i.e., exceeding the range of floating-point representation) of loss, and if the learning rate is too small, the model will converge slowly. Therefore, in the training process, the learning rate of the initial forest fire hazard identification model in each training round is determined according to the training round number, and the update step in each training round is determined by using the learning rate in each training round, so as to adjust the model parameters of the initial forest fire hazard identification model in reverse by using the update step in each training round.

[0136] wherein the cosine annealing scheduling method is used to adjust the learning rate, and the learning rate in the tth training round is calculated by using the following formula (4).

[0137] (4)

[0138] In the above formula (4), denotes the learning rate in the tth training round, denotes the minimum learning rate, denotes the maximum learning rate, T cur denotes the number of training rounds that have been performed at present, T max denotes the total number of training rounds contained in a complete cosine annealing period in the training process (which is a preset value).

[0139] Thus, after the learning rate in the tth training round is calculated, the loss function in the tth training round is obtained by using the sample in the tth training round, and then the loss function gradient is obtained. Finally, the model parameters of the forest fire hazard identification model are updated by using the learning rate in the tth training round and the corresponding loss function gradient. For example, the total training round is set to 80, and the data batch size in each training is set to 8. Of course, the training parameters can be set according to actual use, and are not limited to the foregoing example.

[0140] In this way, the foregoing attenuation method can enable the model to adjust the parameters after convergence, avoid parameter oscillation caused by excessively high learning rate or slow training caused by excessively low learning rate in the later period, and thus improve the training stability and convergence quality.

[0141] Further, in the training process of the present embodiment, the visual model sigLIP is frozen, then a low-rank adaptive fine-tuning algorithm is used to adjust the model parameters in the large language model module, and a full-amount fine-tuning algorithm is used to adjust the model parameters in the feature fusion module and the image-text projection unit; wherein the core idea of LoRA fine-tuning is to change the strategy of directly updating model parameters in traditional full-amount fine-tuning to learning the incremental change of parameters, specifically, LoRA introduces two low-rank decomposed matrices in a specific layer (such as a linear layer) of the pre-trained model, realizes efficient fine-tuning through the construction of weight increment, in the structure design, LoRA retains the weight matrix of the original layer and keeps its parameters fixed, while adding two low-rank matrices, the input data passes through the original layer and the added low-rank matrix, and the final output is the superposition of the two; in this way, it can achieve similar performance to the full-amount fine-tuning model while significantly reducing the size of trainable parameters; of course, both full-amount fine-tuning and LoRA fine-tuning are common ways of model training, and their principles are not repeated here.

[0142] In the present embodiment, the cross-entropy loss function is used for the forest fire hazard identification model, and AdamW is selected as the optimizer during training, and the accuracy is used to evaluate the performance of the model.

[0143] Further, in the model training step, the cross-entropy loss function is used to optimize the model output, the specific process is as follows: first, the forest fire hidden danger identification model outputs the hidden state hidden_states containing the context semantics through the decoder, and maps it to the original prediction score logits of the vocabulary dimension through the language model head (LM Head) (the aforementioned decoder and language model head belong to the network structure in the phi-2 large language model), the shape is [batch_size, seq_length, vocab_size], wherein batch_size represents the batch size, seq_length represents the sequence length, and vocab_size represents the vocabulary size. In order to adapt to the training target of “predicting subsequent tokens after previous tokens” in the sequence generation task, logits and labels are shifted and aligned: taking the first seq_length-1 positions of logits as the predicted value (denoted as shift_logits), and taking the real tokens from the 2nd position in labels as the target value (denoted as shift_labels), forming the training pair of “token-i predicting token-i+1”; then, shift_logits and shift_labels are flattened into [batch_size×(seq_length-1), vocab_size] and [batch_size×(seq_length-1)] vectors, respectively, and input into the cross-entropy loss function. By minimizing the cross-entropy loss function, the model adjusts the parameters to make the predicted distribution approach the real label distribution, thereby improving the model performance.

[0144] In addition, model performance evaluation is a core link of deep learning projects, and its purpose is to ensure that the model functions meet the design expectations and meet the accuracy and efficiency requirements of actual applications. In order to verify the effectiveness of the model output, the accuracy rate is selected as the core evaluation index in this embodiment, and the performance evaluation is carried out on the test data set. The main test target is the binary classification accuracy rate of the model output “yes / no” for the question “Will fire or smoke affect power infrastructure?”.

[0145] In this way, through the foregoing training step, the training of the multi-modal large model can be completed.

[0146] Thus, in actual use, the power transmission line image and the forest fire hidden danger identification guided dialogue set can be directly input into the multi-modal large model, that is, the power transmission line image is input into the SigLIP visual model to obtain initial image features, and the forest fire hidden danger identification guided dialogue set is input into the text encoding module to obtain a text semantic vector; then, the modal projection can be performed, that is, the initial image features are input into the image-text projection unit for feature extraction and dimension adjustment to obtain image features with the same dimension as the text semantic vector; then, the feature fusion module is entered, that is, the text features and the image features are first input into the first prompt word image fusion module, the first prompt word image fusion module sends the information of the two modalities into a cross-attention unit from image to text and a cross-attention unit from text to image in sequence, the two attention units help to align the characteristics of different modalities, then the text features are enhanced by using a self-attention mechanism, and finally the initial text features after feature enhancement are obtained through a feedforward neural network.

[0147] Then, the enhanced text features are obtained by inputting into the second prompt word image fusion module and performing the same operation; subsequently, the enhanced image features and the text features of the dialogue are sent into a cross-attention layer with a gating mechanism to obtain an output, and then the output is input into a feedforward neural network (FFN) with a gating mechanism to obtain fused dialogue text features (i.e., enhanced dialogue text); finally, the enhanced text features, the image and the dialogue text are input into the Phi-2 large language model to perform risk judgment, so as to obtain a forest fire hidden danger output dialogue, and to finally determine whether the forest fire is harmful to the power transmission line.

[0148] Through the multi-modal large model-based power transmission line forest fire hidden danger identification method described in detail in the foregoing steps S1 and S2, the present application realizes the judgment of whether the forest fire is harmful to the power transmission line, completes the technical upgrade from the detection of the existence of the forest fire to the equipment threat judgment, and based on this, the problems of resource waste and delayed crisis response caused by the reporting of all forest fire information in the traditional technology can be avoided, so the present application is very suitable for large-scale application and promotion.

[0149] In one possible design, the third aspect of the present embodiment provides a simulation example, and the specific data are as follows:

[0150] The present embodiment screens 1120 power transmission line image data sets based on real power grid environment, and evaluates the performance of the model designed in the present application in judging the influence of forest fire or smoke on the power transmission line.

[0151] Among them, first, 10000 power transmission line surrounding images with fireworks are detected from the trained YOLO algorithm, and 1120 images at different positions are manually screened out, of which 900 images are used as a training set and 220 images are used as a test set, and their multi-round dialogue data sets are constructed; second, the 900 images and their corresponding multi-round dialogue data sets are used as input to train the model, and the training parameter settings are as follows: the data batch size is set to 8, the model is trained for a total of 80 rounds; the learning rate is set to 2e-4, and a cosine annealing scheduler is used to dynamically adjust to improve the training effect; AdamW is selected as the optimizer; at the same time, in the training process, the visual model SigLIP is frozen, and only the projection module and the modal fusion module are fully fine-tuned, and the large language model Phi-2 is fine-tuned using LoRA, and the cross-entropy loss function is used to optimize the model during fine-tuning.

[0152] Finally, the trained model is evaluated using the 220 screened images, and the test results are shown in Table 2. The trained model performs well on the test set, with an identification accuracy of 0.818, which is much higher than traditional image classification models and current popular multi-modal large models (without fine-tuning).

[0153] Table 2 is a comparison table of model identification accuracy.

[0154] Table 2

[0155] Model Recognition accuracy AlexNet 0.627 SVM 0.573 VGGNet 0.686 ResNet152 0.573 DenseNet 0.664 MobileNetV2 0.646 Llava(13B) 0.568 Qwen-VL(7B) 0.591 Janus-Pro(7B) 0.486 The model proposed in this embodiment 0.818

[0156] Therefore, from Table 2 above, it can be seen that the identification accuracy of the model provided in the present embodiment is much higher than traditional image classification models and current popular multi-modal large models (without fine-tuning).

[0157] As shown in Figure 4 the fourth aspect of the present embodiment provides a hardware device for implementing the power transmission line forest fire hazard identification method based on a multi-modal large model as described in the first aspect of the embodiment, comprising:

[0158] An acquisition unit is configured to acquire power transmission line images with fire targets and a forest fire hazard identification guide dialogue set.

[0159] A hazard identification unit is configured to input the forest fire hazard identification guide dialogue set and the power transmission line image into a forest fire hazard identification model to obtain a forest fire hazard output dialogue of the power transmission line, and to derive a forest fire hazard identification result of the power transmission line according to the forest fire hazard output dialogue, wherein the forest fire hazard identification result includes whether the fire target poses a hazard to the power transmission line or not.

[0160] The working process, working details and technical effects of the device provided in the present embodiment can be referred to the first aspect of the embodiment, which will not be repeated here.

[0161] As Figure 5 shown, the fifth aspect of the embodiment provides another power transmission line fire hazard identification device based on a multi-modal large model. Taking the device as an electronic device for example, it includes a memory, a processor and a transceiver connected in sequence. The memory is used to store a computer program. The transceiver is used to transceive messages. The processor is used to read the computer program and execute the power transmission line fire hazard identification method based on a multi-modal large model as described in the first aspect of the embodiment.

[0162] Specifically, the memory can include, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a flash memory, a first-in-first-out memory (FIFO) and / or a first-in-last-out memory (FILO), etc. Specifically, the processor can include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor can be implemented in at least one of the hardware forms of DSP (digital signal processing), FPGA (field programmable gate array) and PLA (programmable logic array). Meanwhile, the processor can include a main processor and a coprocessor. The main processor is a processor for processing data in the wake-up state, also known as CPU (central processing unit). The coprocessor is a low-power processor for processing data in the standby state.

[0163] In some embodiments, the processor can be integrated with a GPU (image processor) for rendering and drawing the content required to be displayed on the display screen. For example, the processor can be, but is not limited to, a microprocessor of STM32F105 series, a RISC (reduced instruction set computer) microprocessor, an X86 architecture processor or a processor integrated with an embedded neural network processing unit (NPU). The transceiver can be, but is not limited to, a wireless fidelity (WIFI) transceiver, a Bluetooth transceiver, a general packet radio service (GPRS) transceiver, a ZigBee transceiver, a 3G transceiver, a 4G transceiver and / or a 5G transceiver, etc. In addition, the device can further include, but is not limited to, a power module, a display screen and other necessary components.

[0164] The working process, working details and technical effects of the electronic device provided by the embodiment can be referred to the first aspect of the embodiment, which will not be repeated here.

[0165] The sixth aspect of the embodiment provides a storage medium storing instructions of the method for identifying a forest fire hidden danger of a power transmission line based on a multi-modal large model according to the first aspect of the embodiment, that is, the storage medium stores the instructions, and when the instructions are run on a computer, the method for identifying a forest fire hidden danger of a power transmission line based on a multi-modal large model according to the first aspect of the embodiment is executed.

[0166] The storage medium refers to a carrier for storing data, which can include, but is not limited to, a floppy disk, an optical disk, a hard disk, a flash memory, a USB flash disk, a memory stick and the like, and the computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable devices.

[0167] The working process, working details and technical effects of the storage medium provided by the embodiment can be referred to the first aspect of the embodiment, and will not be described here.

[0168] The seventh aspect of the embodiment provides a computer program product containing instructions, which, when run on a computer, causes the computer to execute the method for identifying a forest fire hidden danger of a power transmission line based on a multi-modal large model according to the first aspect of the embodiment, wherein the computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable devices.

[0169] The specific embodiments described above further illustrate the purposes, technical solutions and beneficial effects of the present application. It should be understood that the above description is only a specific embodiment of the present application and is not used to limit the protection scope of the present application. Any modification, equivalent replacement, improvement and the like within the spirit and principle of the present application should be included in the protection scope of the present application.

Claims

1. A method for identifying wildfire hazards of a power transmission line based on a multi-modal large model, characterized in that, The method comprises the following steps: obtaining a power transmission line image of a pyrotechnic target and a forest fire hazard identification guided dialogue set; inputting the forest fire hazard identification guided dialogue set and the power transmission line image into a forest fire hazard identification model to obtain a forest fire hazard output dialogue of the power transmission line, so as to obtain a forest fire hazard identification result of the power transmission line according to the forest fire hazard output dialogue, and the forest fire hazard identification result comprises that the pyrotechnic target has a hazard or no hazard to the power transmission line; wherein the forest fire hazard identification model comprises an image feature extraction module, a text encoding module, a feature fusion module and a large language model module, and the image feature extraction module is used for performing feature extraction processing on the power transmission line image to obtain image features; the text encoding module is used for mapping the forest fire hazard identification guided dialogue set into a text semantic vector; the feature fusion module is used for performing feature fusion processing on the image features and the text semantic vector by using a cross-attention mechanism to obtain enhanced image features with fused text semantics and enhanced text features with fused image features, and performing feature fusion processing on the enhanced image features and the text semantic vector to obtain enhanced dialogue text; the feature fusion module comprises a first prompt word image fusion unit, a second prompt word image fusion unit and a dialogue text image fusion unit; the first prompt word image fusion unit is used for performing one-time cross-modal feature fusion processing on the image features and the text semantic vector by using a cross-attention mechanism to obtain image weighted features with fused text semantics and text weighted features with fused image features, and performing feature enhancement processing on the text weighted features to obtain first initial enhanced text features with fused image features; the second prompt word image fusion unit is used for performing two-time cross-modal feature fusion processing on the first initial enhanced text features and the image weighted features by using a cross-attention mechanism to obtain enhanced image features with fused text semantics and second initial enhanced text features with fused image features; the second prompt word image fusion unit is also used for performing feature enhancement processing on the second initial enhanced text features with fused image features to obtain the enhanced text features with fused image features after the feature enhancement processing; the dialogue text image fusion unit is used for performing feature fusion processing on the enhanced image features and the text semantic vector to obtain the enhanced dialogue text after the feature fusion processing; the large language model module is used for generating the forest fire hazard output dialogue based on the enhanced image features, the enhanced text features and the enhanced dialogue text.

2. The method of claim 1, wherein, the image feature extraction module comprises an image feature extraction unit and an image text projection unit; wherein the image feature extraction unit is used for extracting image semantic features from the power transmission line image to obtain first initial image features, and the image feature extraction unit comprises a SigLIP visual model; The image text projection unit is configured to perform secondary feature extraction on the first initial image feature to obtain a second initial image feature, and perform feature dimension adjustment on the second initial image feature to obtain the image feature with the same feature dimension as the text semantic vector.

3. The method of claim 2, wherein, The image text projection unit comprises a first linear layer and a second linear layer connected in sequence, wherein the first linear layer is configured to perform secondary feature extraction on the first initial image feature to obtain a second initial image feature, and the second linear layer is configured to perform feature dimension adjustment on the second initial image feature to obtain the image feature after the feature dimension adjustment.

4. The method of claim 1, wherein, The first prompt word image fusion unit comprises an image text cross-attention layer, a text image cross-attention layer, a self-attention layer, and a feedforward neural network layer. The image text cross-attention layer is configured to calculate a first attention weight of the image feature on the text semantic vector, and perform feature weighting on the text semantic vector by using the first attention weight to obtain a weighted image feature of the fused text semantic after the feature weighting. The text image cross-attention layer is configured to calculate a second attention weight of the text semantic vector on the weighted image feature, and perform feature weighting on the weighted image feature by using the second attention weight to obtain a weighted text feature of the fused image feature after the feature weighting. The self-attention layer is configured to perform primary feature enhancement on the weighted text feature to obtain a primary enhanced text feature. The feedforward neural network layer is configured to perform secondary feature enhancement on the primary enhanced text feature to obtain a first initial enhanced text feature of the fused image feature after the secondary feature enhancement.

5. The method of claim 4, wherein, The first attention weight is calculated in the following manner: A first query matrix is generated based on the image feature, and a first key matrix is generated based on the text semantic vector. The first attention weight is calculated by using the first query matrix and the first key matrix and using the following formula (1). (1) In the above equation (1), denotes the first attention weight, SoftMax() denotes an activation function of the image-text cross-attention layer, denotes the first query matrix, K l denotes the first key matrix, M l denotes a mask matrix of the text semantic vector, d h denotes an attention scaling factor, T denotes a transpose operation, wherein, , , and V denotes the image feature, L denotes the text semantic vector, both denote a projection matrix of the image-text cross-attention layer.

6. The method of claim 1, wherein, The dialogue text image fusion unit comprises a gated cross-attention layer, a first feature splicing layer, a gated feedforward neural network layer, and a second feature splicing layer. The gated cross-attention layer is configured to perform image and text feature fusion on the text semantic vector and the enhanced image feature to obtain a first initial enhanced dialogue text. The first feature splicing layer is configured to perform feature splicing on the first initial enhanced dialogue text and the text semantic vector to obtain a spliced dialogue text. The gated feedforward neural network layer is configured to perform feature enhancement on the spliced dialogue text to obtain a second initial enhanced dialogue text. The second feature splicing layer is configured to perform feature splicing on the second initial enhanced dialogue text and the spliced dialogue text to obtain the enhanced dialogue text after the feature splicing.

7. The method of claim 6, wherein, The first initial enhanced dialogue text output by the gated cross-attention layer is generated in the following formula (2). (2) In formula (2), y1 represents the first initial enhanced dialogue text, y' represents the text semantic vector, a represents a first learning parameter, tanh() represents a hyperbolic tangent function, and Attention(x, y) represents output data of a cross-attention network in the cross-attention layer based on the gating mechanism. Correspondingly, the second initial enhanced dialogue text output by the feedforward neural network layer based on the gating mechanism is generated by using the following formula (3). (3) In formula (3), y2 represents the second initial enhanced dialogue text, b represents a second learning parameter, and FFN(y1) represents output data of a feedforward neural network in the feedforward neural network layer based on the gating mechanism.

8. The method of claim 1, wherein, The image feature extraction module includes an image feature extraction unit and an image text projection unit, and the forest fire hidden danger identification model is obtained by training in the following manner: A training data set is obtained, wherein the training data set includes a plurality of power transmission line sample images and a set of forest fire hidden danger identification guide sample dialogues corresponding to each power transmission line sample image. An initial forest fire hidden danger identification model is trained by taking each power transmission line sample image and the set of forest fire hidden danger identification guide sample dialogues corresponding to the power transmission line sample image as input and taking a forest fire hidden danger output dialogue corresponding to each power transmission line sample image as output, so as to obtain the forest fire hidden danger identification model after training. In the training process, the learning rate of the initial forest fire hidden danger identification model is determined according to the training round number, and the update step length in each training round is determined by using the learning rate in each training round, so as to adjust the model parameters of the initial forest fire hidden danger identification model by using the update step length in each training round. In the training process, the model parameters in the large language model module are adjusted by using a low-rank adaptive fine-tuning algorithm, and the model parameters in the feature fusion module and the image text projection unit are adjusted by using a full-amount fine-tuning algorithm. 9.A wildfire hazard identification device for a power transmission line based on a multi-modal large model, characterized by, The method comprises the following steps: An acquisition unit is configured to acquire a power transmission line image in which a forest fire target exists and a set of forest fire hidden danger identification guide dialogues. A hidden danger identification unit is configured to input the set of forest fire hidden danger identification guide dialogues and the power transmission line image into a forest fire hidden danger identification model to obtain a forest fire hidden danger output dialogue of the power transmission line, so as to obtain a forest fire hidden danger identification result of the power transmission line according to the forest fire hidden danger output dialogue, wherein the forest fire hidden danger identification result includes whether the forest fire target is harmful or harmless to the power transmission line. The forest fire hidden danger identification model includes an image feature extraction module, a text encoding module, a feature fusion module, and a large language model module, wherein the image feature extraction module is configured to perform feature extraction processing on the power transmission line image to obtain image features. The text encoding module is configured to map the set of forest fire hidden danger identification guide dialogues into a text semantic vector. The feature fusion module is configured to perform feature fusion processing on the image feature and the text semantic vector by using a cross-attention mechanism to obtain enhanced image features with fused text semantics and enhanced text features with fused image features, and perform feature fusion processing on the enhanced image features and the text semantic vector to obtain enhanced dialogue text. The feature fusion module includes a first prompt word image fusion unit, a second prompt word image fusion unit, and a dialogue text image fusion unit. The first prompt word image fusion unit is configured to perform cross-modal feature fusion processing on the image feature and the text semantic vector by using a cross-attention mechanism to obtain image weighted features with fused text semantics and text weighted features with fused image features, and perform feature enhancement processing on the text weighted features to obtain first initial enhanced text features with fused image features. The second prompt word image fusion unit is configured to perform secondary cross-modal feature fusion processing on the first initial enhanced text features and the image weighted features by using a cross-attention mechanism to obtain enhanced image features with fused text semantics and second initial enhanced text features with fused image features. The second prompt word image fusion unit is further configured to perform feature enhancement processing on the second initial enhanced text features with fused image features to obtain the enhanced text features with fused image features after the feature enhancement processing. The dialogue text image fusion unit is configured to perform feature fusion processing on the enhanced image features and the text semantic vector to obtain the enhanced dialogue text after the feature fusion processing. The large language model module is configured to generate the forest fire hidden danger output dialogue based on the enhanced image features, the enhanced text features, and the enhanced dialogue text.

Citation Information

Patent Citations

  • Semantic guidance pedestrian re-identification method and system based on text prompt

    CN120032307A

  • Text-guided multi-modal relationship extraction method and apparatus

    WO2025130069A1