Model training method, fire monitoring method based on prompt learning and related equipment

By adopting a model training method based on prompt learning in forest fire monitoring, and updating the prompt parameters of the prediction model in combination with visual and text detection results, the problems of low monitoring efficiency and non-universality in the prior art are solved, and higher pyrotechnic recognition accuracy and monitoring efficiency are achieved.

CN120047732APending Publication Date: 2025-05-27TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510109029.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

In forest fire monitoring, existing deep learning networks are prone to false detection or missed detection due to high dependence on training data, interference from environmental noise and camera jitter, long model training time, and large differences in the distribution of actual scenarios and training data, resulting in low monitoring efficiency and no universality.

Method used

Using a model training method based on prompt learning, the pyrotechnic monitoring images and pyrotechnic type tags in the sample data are obtained, and the visual model and text model of the prediction model are used to detect images and text respectively. The pyrotechnic prediction probability is determined based on visual and text detection results, and the prompt parameters of the prediction model are updated until the target prediction model with the preset conditions is met.

Benefits of technology

The accuracy of firework recognition is improved, the accuracy and robustness of the prediction model in various scenarios is enhanced, the target prediction model is universal, and the efficiency of forest fire monitoring is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047732A_ABST
    Figure CN120047732A_ABST
Patent Text Reader

Abstract

The invention provides a model training method, a fire monitoring method based on prompt learning and related equipment, and the method comprises the steps: obtaining sample data which comprises a smoke and fire monitoring image of a sample place and a smoke and fire type label corresponding to the smoke and fire monitoring image; determining a visual detection result by using a visual model of a prediction model based on the smoke and fire monitoring image; determining a text detection result by utilizing a text model of the prediction model based on the smoke and fire type label; determining a smoke and fire prediction probability based on the visual detection result and the text detection result, and determining a smoke and fire prediction label based on the smoke and fire prediction probability; on the basis of the smoke and fire prediction label and the smoke and fire type label, prompting parameters of the prediction model are updated; and under the condition that the updated prediction model reaches a preset condition, obtaining a target prediction model. The target prediction model obtained by the method can improve the monitoring efficiency of forest fire.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of natural disaster monitoring, and particularly to a model training method, a fire monitoring method based on prompt learning, and related devices. Background Art

[0002] In forest fire monitoring, deep learning networks are usually used to identify early fire features such as smoke and flames, so as to realize the monitoring and early warning of forest fires. However, the commonly used deep learning networks (such as Convolutional Neural Network (CNN)) have a high dependence on training data and are easily interfered by factors such as environmental noise and camera jitter. In addition, due to the large number of model parameters and the relatively deep network layer settings, the training time of the model is relatively long. In actual application scenarios, when the actual scenario has a large difference from the training data distribution, misdetection or missed detection may occur, and it cannot adapt to diverse fire field situations and does not have universality. Summary of the Invention

[0003] Embodiments of this application disclose a model training method, a fire monitoring method based on prompt learning, and related devices, which solve the technical problem that the current technical solutions for monitoring forest fires are not universal and result in low monitoring efficiency.

[0004] This application provides a model training method, and the method includes: obtaining sample data, where the sample data includes a smoke and fire monitoring image of a sample location and a smoke and fire type label corresponding to the smoke and fire monitoring image; based on the smoke and fire monitoring image, using a visual model of a prediction model to determine a visual detection result; based on the smoke and fire type label, using a text model of the prediction model to determine a text detection result; based on the visual detection result and the text detection result, determining a smoke and fire prediction probability, and determining a smoke and fire prediction label based on the smoke and fire prediction probability; based on the smoke and fire prediction label and the smoke and fire type label, updating a prompt parameter of the prediction model; and obtaining a target prediction model when the updated prediction model meets a preset condition.

[0005] In some embodiments of the present application, the vision model includes a linear layer, a projection layer, and N Vision Transformer layers. Determining a vision detection result using the vision model of the prediction model based on the fireworks monitoring image includes: concatenating the output result of the i-th Vision Transformer layer with the prompt vector of the (i + 1)-th Vision Transformer layer to determine the image sequence of the (i + 1)-th Vision Transformer layer, where 1 ≤ i ≤ N and N is a positive integer; wherein, inputting the fireworks monitoring image into the linear layer and the projection layer for linear transformation and size alignment to obtain projection features; concatenating the projection features with the prompt vector of the first Vision Transformer layer to determine the image sequence input into the first Vision Transformer layer. Based on the image sequence of the (i + 1)-th Vision Transformer layer, using the (i + 1)-th Vision Transformer layer to determine the output result of the (i + 1)-th Vision Transformer layer; taking the output result of the last Vision Transformer layer as the vision detection result.

[0006] In some embodiments of the present application, determining the output result of the (i + 1)-th Vision Transformer layer using the (i + 1)-th Vision Transformer layer based on the image sequence of the (i + 1)-th Vision Transformer layer includes: performing a linear mapping on the image sequence of the (i + 1)-th Vision Transformer layer based on a preset projection matrix to obtain first intermediate features; performing a residual connection and normalization on the first intermediate features and the image sequence to obtain second intermediate features; obtaining third intermediate features based on a first feed-forward network and the second intermediate features; performing a residual connection and normalization on the third intermediate features and the second intermediate features to obtain the output result of the (i + 1)-th Vision Transformer layer.

[0007] In some embodiments of the present application, the text model includes a word embedding layer and M Transformer layers. Determining a text detection result using the text model of the prediction model based on the firework type label includes: concatenating the output result of the j-th Transformer layer with the prompt vector of the (j + 1)-th Transformer layer to determine the text sequence of the (j + 1)-th Transformer layer, where 1 ≤ j ≤ M and M is a positive integer; wherein, inputting the firework type label into the word embedding layer for word segmentation and vectorization processing to obtain an initial text vector; mapping the initial text vector to the dimensional space corresponding to the prompt vector of the first Transformer layer, and concatenating the initial text vector mapped to the dimensional space with the prompt vector of the first Transformer layer to determine the text sequence input into the first Transformer layer; based on the text sequence of the (j + 1)-th Transformer layer, using the (j + 1)-th Transformer layer to determine the output result of the (j + 1)-th Transformer layer; taking the output result of the last Transformer layer as the text detection result.

[0008] In some embodiments of the present application, determining the output result of the (j + 1)-th Transformer layer using the (j + 1)-th Transformer layer based on the image sequence of the (j + 1)-th Transformer layer includes: performing a linear mapping on the text sequence of the (j + 1)-th Transformer layer based on a preset projection matrix to obtain a fourth intermediate feature; performing a residual connection and normalization on the fourth intermediate feature and the text sequence to obtain a fifth intermediate feature; obtaining a sixth intermediate feature based on a second feed-forward network and the fifth intermediate feature; performing a residual connection and normalization on the sixth intermediate feature and the fifth intermediate feature to obtain the output result of the (j + 1)-th Transformer layer.

[0009] In some embodiments of the present application, obtaining a target prediction model when the updated prediction model meets a preset condition includes: calculating the loss function of the updated prediction model; determining an evaluation result based on a preset evaluation dataset and the updated prediction model; when the loss function meets a first condition and / or the evaluation result meets a second condition, determining that the updated prediction model meets the preset condition, and taking the updated prediction model as the target prediction model.

[0010] In some embodiments of the present application, determining the fireworks prediction probability based on the visual detection result and the text detection result includes: calculating the similarity between the visual detection result and the text detection result; determining the fireworks prediction probability based on the similarity.

[0011] The present application also provides a fire monitoring method based on prompt learning, including: obtaining fireworks monitoring data at a specified location; based on the fireworks monitoring data, using a target prediction model to obtain the monitoring result of the specified location, and the target prediction model is obtained by the above-mentioned model training method.

[0012] The present application also provides an electronic device, which includes a processor and a memory. When the processor executes a computer program stored in the memory, it implements the above-mentioned model training method or the fire monitoring method based on prompt learning.

[0013] The present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the above-mentioned model training method or the fire monitoring method based on prompt learning.

[0014] In the model training method provided by the present application, the sample data is divided into fireworks monitoring images and fireworks type labels. The fireworks monitoring images and the fireworks type labels are processed respectively, including: based on the fireworks monitoring images, using the visual model of the prediction model to determine the visual detection result; based on the fireworks type labels, using the text model of the prediction model to determine the text detection result, providing a data basis for subsequent mining of the correlation between images and texts. Based on the visual detection result and the text detection result, the fireworks prediction probability is determined. The fireworks prediction probability determined by the visual detection result and the text detection result can effectively mine the correlation between images and texts, so that the prediction model can improve the accuracy and robustness in various scenarios. Based on the fireworks prediction label and the fireworks type label, the prompt parameters of the prediction model are updated. When the updated prediction model meets the preset conditions, the training of the prediction model is completed to obtain the target prediction model. Since the target prediction model can well learn the correlation between images and texts and adapt to various different scenarios, the target prediction model has universality. To a certain extent, when using the target prediction model to monitor forest fires, it has a high monitoring efficiency. Description of the Drawings

[0015] Figure 1 is a schematic diagram of the application scenario of the model training method and the fire monitoring method based on prompt learning provided by the embodiments of the present application.

[0016] Figure 2 is a flowchart of the model training method provided by the embodiments of the present application.

[0017] Figure 3 It is a schematic structural diagram of the visual model provided by the embodiments of the present application.

[0018] Figure 4 It is a schematic structural diagram of each Vision Transformer layer provided by the embodiments of the present application.

[0019] Figure 5 It is a schematic structural diagram of the text model provided by the embodiments of the present application.

[0020] Figure 6 It is a schematic structural diagram of each Transformer layer provided by the embodiments of the present application.

[0021] Figure 7 It is a schematic diagram of the calculation of the firework prediction probability provided by the embodiments of the present application.

[0022] Figure 8 It is a schematic diagram of the fire monitoring method based on prompt learning provided by the embodiments of the present application.

[0023] Figure 9 It is a schematic structural diagram of the electronic device provided by the embodiments of the present application. Detailed implementation manners

[0024] For ease of understanding, some explanations of concepts related to the embodiments of the present application are exemplarily given for reference.

[0025] It should be noted that "at least one" in the present application means one or more, and "a plurality" means two or more than two. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B may be singular or plural. The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims, and drawings of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence.

[0026] For ease of understanding, some explanations of concepts related to the embodiments of the present application are exemplarily given for reference.

[0027] 1) Prompt learning technology (Prompt Learning)

[0028] Prompt Learning is a method for efficient transfer of large-scale pre-trained models in downstream tasks. Its basic principle is to insert or learn "prompt" vectors or tokens in the input sequence, transforming the traditional fine-tuning process that requires a large number of parameter updates into a situation where adaptation to different tasks can be achieved with few or even zero parameter updates. Prompt Learning technology utilizes the original knowledge potential of the language model, "prompting" the task requirements to the model and enabling the model to make inferences based on the prompts, thereby achieving efficient adaptation to downstream tasks while maintaining the generality of the model.

[0029] 2) Patch Embedding Sublayer

[0030] The Patch Embedding Sublayer is commonly used in image processing tasks, especially in the Vision Transformer architecture. Its basic idea is to divide the input image into several fixed-size "patches" and map these local image patches into one-dimensional vector representations through linear transformation or convolution, etc., thus converting two-dimensional pixel information into a one-dimensional sequence representation for subsequent processing by the Transformer. Through this segmentation and embedding method, the model can capture the local structural features of the image and provide richer context information for subsequent attention calculation.

[0031] 3) Position Embedding Sublayer

[0032] The Position Embedding Sublayer is a key component in the Transformer structure used to represent the position information of sequence elements. Since the Self-Attention mechanism does not have explicit position information when capturing the dependencies within the input sequence, the position embedding sublayer injects position encoding into the sequence representation by adding sine and cosine functions or learnable parameters, enabling the model to distinguish elements at different positions in the sequence. This is crucial for natural language processing, image patch sequences, and the modeling of other sequence data.

[0033] 4) Word Embedding Layer

[0034] The word embedding layer is a representation learning technique that maps discrete lexical symbols into a continuous vector space. By converting lexical indices into dense vectors of a fixed dimension, it transforms language symbols into a numerical form. Its core goal is to capture the semantic and syntactic relationships between words in a low-dimensional vector space, such that semantically similar words have similar geometric distances in the vector space. The word embedding layer is usually the first layer of a deep learning model, providing high-dimensional, continuous input representations for subsequent neural network processing, thereby supporting the effective transmission and computation of semantic information.

[0035] 5) Multi-Head Self-Attention Layer

[0036] The Multi-Head Self-Attention Layer is a core module of the Transformer architecture. By executing multiple self-attention mechanisms in parallel, the model can learn the dependencies between elements in a sequence in different subspaces. Each attention head maps the input sequence to queries, keys, and values, then calculates attention scores, and finally combines the output results of multiple heads. In fields such as natural language processing and computer vision, the multi-head self-attention layer can help the model flexibly and efficiently capture global dependencies and context information.

[0037] 6) Layer Normalization

[0038] The Layer Normalization plays a role in stabilizing the training process and accelerating convergence in deep neural networks. By normalizing the activation values in a layer to have zero mean and unit variance, it can effectively prevent problems such as gradient explosion and gradient vanishing. In the self-attention mechanism and feed-forward network, placing the layer normalization after the input or residual connection further enhances the stability and generalization performance of the model.

[0039] 7) Feed-Forward Neural Network (FFN)

[0040] The basic structure of the feed-forward neural network is two layers of linear transformation, with an activation function in between, used for non-linear mapping and feature extraction of the representation at each position in the sequence. Through the process of dimensionality increase and decrease, the feed-forward neural network can deeply combine and abstract local information, providing richer semantic representations for the next attention calculation.

[0041] 8) Residual Connection

[0042] The residual connection was initially proposed by the Deep Residual Network (ResNet) to alleviate the problems of vanishing gradients and model degradation caused by the increase in network depth. In Transformer, the residual connection directly adds the input of each sub-layer (such as multi-head self-attention, MLP, etc.) to the sub-layer output, forming a "skip" path. This approach enables the identity transfer of information, avoids gradient decay caused by multiple layers of stacking, and promotes more effective feature learning and training stability.

[0043] In forest fire monitoring, deep learning networks are usually used to identify early fire features such as smoke and flames, so as to achieve the monitoring and early warning of forest fires. However, commonly used deep learning networks (such as Convolutional Neural Network (CNN)) have a high dependence on training data and are easily interfered by factors such as environmental noise and camera jitter. In addition, due to the large number of model parameters and the relatively deep network layer settings, the training time of the model is relatively long. In actual application scenarios, when the actual scenario has a large difference from the training data distribution, false detection or missed detection may occur, and it cannot adapt to diverse fire field situations and does not have universality.

[0044] To solve the technical problem of low monitoring efficiency caused by the lack of universality of the current technical solutions for monitoring forest fires, the embodiments of this application provide a model training method, a fire monitoring method based on prompt learning, and related devices, which can improve the accuracy of smoke and fire recognition. First, the application scenarios of the model training method and the fire monitoring method based on prompt learning of this application will be described below.

[0045] Figure 1It is a schematic diagram of the application scenario of the model training method and the fire monitoring method based on prompt learning provided by the embodiments of the present application. The model training method and the fire monitoring method based on prompt learning provided by the embodiments of the present application are applied to the electronic device 10, and the electronic device 10 is communicatively connected to the electronic device 20. The communication connection method can be a wired connection or a wireless connection. The wired connection method can include one or more of wired connection methods such as Universal Serial Bus (USB) and Controller Area Network (CAN). The wireless connection method can include one or more of wireless communication connection methods such as Wireless Fidelity (Wi-Fi), Bluetooth (BT), mobile communication network, Frequency Modulation (FM), Near Field Communication (NFC), and Infrared (IR).

[0046] Among them, the electronic device 10 can be a mobile phone, a tablet computer, a smart wearable device, an Augmented Reality (AR) / Virtual Reality (VR) device, a laptop computer, a netbook, a single server, a server cluster composed of multiple servers, a cloud server, etc. The electronic device 20 can be a mobile phone (such as an Android mobile phone, an IOS mobile phone, etc.), a laptop computer, a tablet computer, a personal digital assistant, a Mobile Internet Devices (MID), a PAD, a desktop computer, a smart TV, and other computer devices with a display screen.

[0047] In some embodiments of the present application, the electronic device 10 can include a model training module and a model application module. The model training method is executed through the model training module, and the fire monitoring method based on prompt learning is executed through the model application module. The electronic device 10 real-time monitors the smoke and fire monitoring data at a specified location, and calls the trained target prediction model in the model application module to predict the smoke and fire monitoring data, and sends the detection result of the specified location to the electronic device 20. The user can view the detection result through the display screen ( Figure 1 not shown) of the electronic device 20.

[0048] The schematic Figure 1 is only an example of the application scenario and does not constitute a limitation on the application scenario. It may include more or fewer components than shown in the figure, or combine some components, or different components. For example, the electronic device 10 may further include an input / output device, a network access device, a display device, etc.

[0049] Figure 2 is a flowchart of the model training method provided by an embodiment of the present application, which is applied to an electronic device (such as Figure 1 the electronic device 10). According to different requirements, the order of steps in this flowchart can be changed, and some steps can be omitted.

[0050] Step S201, obtain sample data.

[0051] In some embodiments of the present application, the sample location can be any location where fires and smokes are likely to occur, such as forests, grasslands, mountainous areas, wilderness border areas, etc. The fire and smoke monitoring data collected at the sample location can be obtained from the monitoring devices installed at the sample location. The present application does not limit the acquisition method of the fire and smoke monitoring data.

[0052] The fire and smoke monitoring data includes fire and smoke monitoring images, and fire and smoke type labels are generated based on the fire and smoke monitoring images. The fire and smoke monitoring images and the fire and smoke type labels are used as sample data. Among them, the fire and smoke type labels can be represented by different strings. For example, the fire and smoke type labels can be "burning waste and miscellaneous", "cooking smoke", "industrial emissions", etc.

[0053] Step S202, based on the fire and smoke monitoring data, use the visual model of the prediction model to determine the visual detection result.

[0054] In some embodiments of the present application, the prediction model is a Transformer network structure, which can include a visual model. As Figure 3 shown, the structure of this visual model can include a linear layer, a projection layer, and N sequentially stacked Vision Transformer layers. As Figure 4 shown in (1), each Vision Transformer layer includes a first module, a second module, a third module, and a fourth module. Among them, the first projection layer includes a patch embedding layer and a position encoding layer. As Figure 4 shown in (1) and (2), the first module includes a multi-head self-attention layer. The second module includes a normalization layer and a residual connection. The third module includes a feed-forward neural network. The fourth module includes a normalization layer and a residual connection. The Vision Transformer is used to update the node state. By stacking N Vision Transformer layers, the depth of the visual model can be increased, so that the visual model can better capture the features of the fire and smoke monitoring images.

[0055] Input the firework monitoring image into the vision model, which is processed successively through a linear layer, a projection layer, and N Vision Transformer layers. Concatenate the output result of the i-th Vision Transformer layer with the prompt vector of the (i + 1)-th Vision Transformer layer to determine the image sequence of the (i + 1)-th Vision Transformer layer, where 1 ≤ i ≤ N and N is a positive integer.

[0056] Among them, for the input of the first Vision Transformer layer, it can be determined by using the linear layer and the projection layer. Specifically, input the firework monitoring image into the linear layer and the projection layer for linear transformation and size alignment to obtain projection features. Concatenate the projection features with the prompt vector of the first Vision Transformer layer to determine the image sequence of the first Vision Transformer layer. Use the image sequence of the first Vision Transformer layer as the input of the first Vision Transformer layer. The first Vision Transformer layer processes the image sequence of the first Vision Transformer layer to obtain the output result of the first Vision Transformer layer.

[0057] In some embodiments of the present application, after determining the image sequence of the (i + 1)-th Vision Transformer layer, input the image sequence of the (i + 1)-th Vision Transformer layer into the (i + 1)-th Vision Transformer layer, and use the (i + 1)-th Vision Transformer layer to determine the output result of the (i + 1)-th Vision Transformer layer. Use the output result of the last Vision Transformer layer as the vision detection result.

[0058] The following combines Figure 4 Describe the processing process of the (i + 1)-th Vision Transformer layer. The image sequence of the (i + 1)-th Vision Transformer layer is represented by the following formula:

[0059]

[0060] In the formula, X (l) represents the image sequence of the (i + 1)-th Vision Transformer layer.

[0061] The image sequence X of the (i + 1)-th Vision Transformer layer (l)Input the (i + 1)-th Vision Transformer layer, X (l) First, it passes through the multi-head self-attention layer of the first module. In the multi-head self-attention layer, based on the multi-head self-attention mechanism, a linear mapping is performed on the image sequence of the (i + 1)-th Vision Transformer layer according to a preset projection matrix to obtain the first intermediate feature. It is expressed by the formula as follows:

[0062] Q = X (l) W Q , K = X (l) W K , V = X (l) W V ;

[0063] In the formula, W Q , W K , W V represent the preset projection matrix. Q represents the query. K represents the key. V represents the value.

[0064] According to Q, K, and V, calculate the first intermediate feature Z 1 . It is expressed by the formula as follows:

[0065]

[0066] In the formula, d represents the vector dimension (or scaling factor).

[0067] After obtaining the first intermediate feature, input the first intermediate feature into the second module. In the second module, perform a residual connection and layer normalization on the first intermediate feature and the image sequence of the (i + 1)-th Vision Transformer layer to obtain the second intermediate feature. It is expressed by the formula as follows:

[0068]

[0069] In the formula, represents the second intermediate feature.

[0070] After obtaining the second intermediate feature, input the second intermediate feature into the third module. In the third module, process the second intermediate feature using a feed-forward network (FFN) to obtain the third intermediate feature.

[0071] After obtaining the third intermediate feature, input the third intermediate feature into the fourth module. In the fourth module, perform a residual connection and normalization on the third intermediate feature and the second intermediate feature to obtain the output result of the (i + 1)-th Vision Transformer layer. It is expressed by the formula as follows:

[0072]

[0073] wherein, X (l+1) represents the output result of the (i + 1)-th Vision Transformer layer.

[0074] Concatenate the output result of the (i + 1)-th Vision Transformer layer with the prompt vector of the (i + 2)-th Vision Transformer layer to obtain the input of the (i + 2)-th Vision Transformer layer, that is, the image sequence of the (i + 2)-th Vision Transformer layer. Input the image sequence of the (i + 2)-th Vision Transformer layer into the (i + 2)-th Vision Transformer layer for processing to obtain the output result of the (i + 2)-th Vision Transformer layer.

[0075] Traverse all Vision Transformer layers according to the above calculation method, and use the output result of the last Vision Transformer layer as the visual detection result of the visual model.

[0076] Step S203, based on the firework type label, use the text model of the prediction model to determine the text detection result.

[0077] In some embodiments of the present application, the prediction model further includes a text model, as Figure 5 shown, the structure of the text model may include a word embedding layer and M stacked Transformer layers in sequence. As Figure 6 shown in (1), each Transformer layer specifically includes a fifth module, a sixth module, a seventh module, and an eighth module. As Figure 6 shown in (1) and (2), the fifth module includes a multi-head self-attention layer. The sixth module includes a normalization layer and a residual connection. The seventh module includes a feed-forward neural network. The eighth module includes a normalization layer and a residual connection. Stack the M Transformer layers so that the text model can better capture the features of the firework type label.

[0078] Input the firework type label into the text model, and process it through the word embedding layer and M Transformer layers in sequence. Concatenate the output result of the j-th Transformer layer with the prompt vector of the (j + 1)-th Transformer layer to determine the text sequence of the (j + 1)-th Transformer layer, 1 ≤ j ≤ M, and M is a positive integer.

[0079] Among them, for the input of the first Transformer layer, it can be determined by using a word embedding layer. Specifically, the firework type label is input into the word embedding layer for word segmentation and vectorization processing to obtain an initial text vector. The initial text vector is mapped to the dimensional space corresponding to the prompt vector of the first Transformer layer, and the initial text vector mapped to the dimensional space is concatenated with the prompt vector of the first Transformer layer to determine the text sequence of the first Transformer layer. The text sequence of the first Transformer layer is input into the first Transformer layer to obtain the output result of the first Transformer layer.

[0080] In some embodiments of the present application, after determining the text sequence of the (j + 1)-th Transformer layer, the text sequence of the (j + 1)-th Transformer layer is input into the (j + 1)-th Transformer layer, and the (j + 1)-th Transformer layer is used to determine the output result of the (j + 1)-th Transformer layer. The output result of the last Transformer layer is used as the text detection result.

[0081] The text sequence of the (j + 1)-th Transformer layer is expressed by the following formula:

[0082]

[0083] In the formula, X (u) represents the text sequence of the (j + 1)-th Transformer layer.

[0084] The text sequence X (u) of the (j + 1)-th Transformer layer is input into the (j + 1)-th Transformer layer, and X (u) first passes through the multi-head self-attention layer of the fifth module. In the multi-head self-attention layer, based on the multi-head self-attention mechanism, a linear mapping is performed on the text sequence of the (j + 1)-th Transformer layer according to a preset projection matrix to obtain a fourth intermediate feature. It is expressed by the following formula:

[0085] Q = X (u) W Q , K = X (u) W K , V = X (u) W V ;

[0086] In the formula, W Q , W K , W V represents the preset projection matrix. Q represents the query. K represents the key. V represents the value.

[0087] Calculate the fourth intermediate feature Z based on Q, K, and V 2 . It is expressed by the formula as follows:

[0088]

[0089] In the formula, d represents the vector dimension (or scaling factor).

[0090] After obtaining the fourth intermediate feature, input the fourth intermediate feature into the sixth module. In the sixth module, perform residual connection and normalization on the fourth intermediate feature and the text sequence of the (j + 1)-th Transformer layer to obtain the fifth intermediate feature. It is expressed by the formula as follows:

[0091]

[0092] In the formula, represents the fifth intermediate feature.

[0093] After obtaining the fifth intermediate feature, input the fifth intermediate feature into the seventh module. In the seventh module, process the fifth intermediate feature using a feed-forward network to obtain the sixth intermediate feature.

[0094] After obtaining the sixth intermediate feature, input the sixth intermediate feature into the eighth module. In the eighth module, perform residual connection and normalization on the fifth intermediate feature and the sixth intermediate feature to obtain the output result of the (i + 1)-th Vision Transformer layer. It is expressed by the formula as follows:

[0095]

[0096] In the formula, X (u+1) represents the output result of the (i + 1)-th Vision Transformer layer.

[0097] Concatenate the output result of the (j + 1)-th Transformer layer with the prompt vector of the (j + 2)-th Transformer layer to obtain the input of the (j + 2)-th Transformer layer, that is, the text sequence of the (j + 2)-th Transformer layer. Input the text sequence of the (j + 2)-th Transformer layer into the (j + 2)-th Transformer layer for processing to obtain the output result of the (j + 2)-th Transformer layer.

[0098] Traverse all Transformer layers according to the above calculation method, and use the output result of the last Transformer layer as the text detection result of the text model.

[0099] Step S204: Based on the visual detection result and the text detection result, determine the firework prediction probability, and determine the firework prediction label based on the firework prediction probability.

[0100] In some embodiments of the present application, calculate the similarity between the visual detection result and the text detection result, and through the Sigmoid function, obtain the firework prediction probability within the range of 0 to 1.

[0101] As Figure 7 shown, it is a schematic diagram of the calculation of the firework prediction probability. Input the firework monitoring image into the visual model. In the visual model, use the linear layer and the projection layer to perform linear transformation and size alignment on the firework monitoring image to obtain the projection feature. Concatenate the projection feature with the prompt vector of the first Vision Transformer layer to obtain the input of the first Vision Transformer layer, and then sequentially traverse all the Vision Transformer layers. Among them, the way to sequentially traverse all the Vision Transformer layers can refer to Figure 2 Step S202 of the embodiment. Take the output result of the last Vision Transformer layer as the visual detection result.

[0102] As Figure 7 shown, input the firework type label into the text model. In the text model, use the word embedding layer to perform word segmentation and vectorization processing on the firework type label to obtain the initial text vector. Map the initial text vector to the dimension space corresponding to the prompt vector of the first Transformer layer, and concatenate the initial text vector mapped to the dimension space with the prompt vector of the first Transformer layer to obtain the input of the first Transformer layer. Then sequentially traverse all the Transformer layers. Among them, the way to sequentially traverse all the Transformer layers can refer to Figure 2 Step S203 of the embodiment. Take the output result of the last Transformer layer as the text detection result.

[0103] As Figure 7 shown, calculate the similarity between the visual detection result output by the Nth Vision Transformer layer and the text detection result output by the Mth Transformer layer to obtain the firework prediction probability.

[0104] In one example, based on the similarity, the candidate labels and their corresponding probabilities are obtained as follows: "Burning of weeds" (0.7), "Cooking smoke" (0.2), "Industrial emissions" (0.1). Then, the candidate label corresponding to the maximum probability is used as the firework prediction label. In this example, "Burning of weeds" corresponding to the maximum probability of 0.7 is used as the firework prediction label.

[0105] Step S205: Update the prompt parameters of the prediction model based on the firework prediction label and the firework type label.

[0106] In some embodiments of the present application, the prompt parameters of the prediction model are updated according to the difference between the firework prediction label and the firework type label.

[0107] Step S206: Obtain the target prediction model when the updated prediction model meets the preset conditions.

[0108] In some embodiments of the present application, the preset conditions can be set according to the actual application scenario. For example, the preset conditions may include that the loss function of the prediction model reaches convergence, the evaluation accuracy rate of the preset model no longer increases, etc.

[0109] In some embodiments of the present application, during the process of training the prediction model, the total number of training iterations can be preset to 50 times, and the initial learning rate is 0.001. The weighted binary cross-entropy loss function is adopted to address the problem of sample imbalance. The loss function of the updated prediction model is calculated and expressed by the following formula:

[0110]

[0111] In the formula, N is the total number of samples, y i is the true label of the i-th sample, p i is the predicted probability of the i-th sample, w 1 and w 0 are the weight coefficients of the positive and negative class samples respectively.

[0112] If the loss function of the updated prediction model no longer decreases in 50 iterations, it is determined that the loss function meets the first condition.

[0113] In some embodiments of the present application, in order to better evaluate the effectiveness of the classification of the prediction model, evaluation metrics based on the confusion matrix are used to calculate the accuracy, recall rate, and F1 score of the prediction model. Specifically, an evaluation data set for evaluating the prediction model is obtained. Among them, the evaluation data set can be a part of the sample data that has not participated in the training. For example, 80% of the sample data is used as the training data (including fire monitoring images and fire type labels) in the above training process, and 20% is used as the evaluation data set. The evaluation data set can also be other data sets that do not belong to the sample data. The present application places no restrictions on the source of the evaluation data set. The evaluation data set is input into the updated prediction model for processing to obtain an evaluation result. Based on the evaluation result, the following formulas are used to calculate the accuracy, recall rate, and F1 score to verify whether the classification result (prediction result) output by the updated prediction model is consistent with the classification result (label) in the evaluation data set. Among them, consistency means that the classification results of the updated prediction model and the evaluation data set are the same, and non-consistency means that the classification results of the updated prediction model and the evaluation data set are different.

[0114]

[0115] Among them, TP is the number of samples in the evaluation data set whose label is positive and the prediction result of the updated prediction model is also positive. TN is the number of samples in the evaluation data set whose label is negative and the prediction result of the updated prediction model is negative. FP is the number of samples in the evaluation data set whose label is negative but the prediction result of the updated prediction model is positive. FN is the number of samples in the evaluation data set whose label is positive but the prediction result of the updated prediction model is negative.

[0116] Based on the evaluation of the evaluation data set, if the accuracy rate reaches a relatively stable and high value, the recall rate is greater than or equal to the first preset threshold, and the F1 score is greater than or equal to the second preset threshold, it is determined that the evaluation result meets the second condition.

[0117] If the loss function meets the first condition and / or the evaluation result meets the second condition, it is determined that the updated prediction model reaches the preset condition, and the updated prediction model is used as the target prediction model.

[0118] After verification, the accuracy rate of the target prediction model is 0.861, the recall rate is 0.957, and the F1 score is 0.867. Then the target prediction model can accurately extract the features of the fire image, thereby completing the monitoring of forest fires.

[0119] In the related art, it is difficult to fully understand the correlation between images and texts only by extracting image features from a visual perspective using traditional CNNs. Based on the above embodiments, the sample data is divided into fire and smoke monitoring images and fire and smoke type labels. The visual model of the prediction model is used to process the fire and smoke monitoring images to obtain visual detection results. The text model of the prediction model is used to process the fire and smoke type labels to obtain text detection results. By processing the fire and smoke monitoring images and fire and smoke type labels respectively, the correlation between cross-modal data (images and texts) can be effectively mined, thereby improving the accuracy and robustness of the prediction model in various scenarios such as target recognition and image-text matching.

[0120] In addition, the prediction model provides richer general features during training, enabling the target prediction model to significantly enhance the performance of downstream tasks and reduce the risk of overfitting under the condition of only a small amount of labeled data or unlabeled data when performing prompt learning. By flexibly adjusting text prompts in the prompt structure, multi-modal semantic associations can be better captured, improving the adaptability of the prediction model to complex scenarios. Since the structure of the prediction model provided by the above embodiments is simple, deploying the prediction model in an electronic device consumes less, has better performance, and higher operating efficiency.

[0121] Figure 8 It is a schematic diagram of a fire monitoring method based on prompt learning provided by an embodiment of the present application. As Figure 8 shown, it includes the following steps.

[0122] Step S801, obtain fire and smoke monitoring data at a specified location.

[0123] In some embodiments of the present application, the specified location and the sample location may be the same type of site. The fire and smoke monitoring data includes fire and smoke images.

[0124] Step S802, based on the fire and smoke monitoring data, use the target prediction model to obtain the monitoring result at the specified location.

[0125] In some embodiments of the present application, the fire and smoke image data is input into the target prediction model, and the target prediction model uses the visual model to predict the fire and smoke image data to obtain the monitoring result at the specified location. Among them, the training process of the target prediction model can refer to the embodiments as Figure 2 shown, and will not be repeated here.

[0126] Based on the above embodiments, using the target prediction model can accurately extract fire and smoke images, thereby completing the monitoring of forest fires.

[0127] Figure 9It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. The electronic device 10 may include a communication module 101, a memory 102, a processor 103, an input / output (I / O) interface 104, and a bus 105. The processor 103 is respectively coupled to the communication interface 101, the memory 102, and the I / O interface 104 through the bus 105.

[0128] The communication module 101 may include a wired communication module and / or a wireless communication module. The wired communication module may provide one or more of the solutions for wired communication such as Universal Serial Bus (USB), Controller Area Network (CAN), etc. The wireless communication module may provide one or more of the solutions for wireless communication such as Wireless Fidelity (Wi-Fi), Bluetooth (BT), mobile communication network, Frequency Modulation (FM), Near Field Communication (NFC), Infrared (IR), etc.

[0129] The memory 102 may include one or more Random Access Memories (RAMs) and one or more Non-Volatile Memories (NVMs). The random access memory can be directly read and written by the processor 103, and can be used to store the operating system or executable programs (such as machine instructions) of other running programs, and can also be used to store user and application data, etc.

[0130] The random access memory may include Static Random-Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDR SDRAM), etc.

[0131] The non-volatile memory can also store executable programs, user and application data, etc., which can be pre-loaded into the random access memory for direct reading and writing by the processor 103. The non-volatile memory can include a disk storage device, a flash memory.

[0132] The memory 102 is used to store one or more computer programs. The one or more computer programs are configured to be executed by the processor 103. The one or more computer programs include a plurality of instructions. When the plurality of instructions are executed by the processor 103, a model training method and a fire monitoring method based on prompt learning that can be executed on the electronic device 10 can be realized.

[0133] In other embodiments, the electronic device 10 further includes an external memory interface for connecting to an external memory to expand the storage capacity of the electronic device 10.

[0134] The processor 103 may include one or more processing units. For example, the processor 103 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.

[0135] The processor 103 provides computing and control capabilities. For example, the processor 103 is used to execute the computer program stored in the memory 102 to implement the above-mentioned model training method and the fire monitoring method based on prompt learning.

[0136] The I / O interface 104 is used to provide a channel for user input or output. For example, the I / O interface 104 can be used to connect various input and output devices, such as a mouse, a keyboard, a touch device, a display screen, etc., so that the user can enter information or visualize the information.

[0137] The bus 105 is at least used to provide a communication channel between the communication module 101, the memory 102, the processor 103, and the I / O interface 104 in the electronic device 10.

[0138] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the electronic device 10. In other embodiments of the present application, the electronic device 10 may include more or fewer components than those illustrated, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0139] The embodiments of the present application further provide a computer-readable storage medium, on which a computer program is stored. The computer program includes program instructions, and the method implemented when the program instructions are executed may refer to the methods in the above various embodiments of the present application.

[0140] Among them, the computer-readable storage medium may be the internal memory of the electronic device in the above embodiments, such as the hard disk or memory of the electronic device. The computer-readable storage medium may also be an external storage device of the electronic device, such as a plug-in hard disk equipped on the electronic device, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc.

[0141] In some embodiments, the computer-readable storage medium may include a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function, etc.; the data storage area may store data created according to the use of the electronic device, etc.

[0142] In the above embodiments, the descriptions of the various embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0143] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician may use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0144] In the embodiments provided in the present application, it should be understood that the disclosed device / terminal device and method can be implemented in other ways. For example, the device / terminal device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.

[0145] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0146] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A model training method, characterized in that: The method comprises: Acquire sample data, the sample data including a fireworks monitoring image of a sample location and a fireworks type label corresponding to the fireworks monitoring image; Based on the fireworks monitoring image, determining a visual detection result using a visual model of a prediction model; Determining a text detection result using a text model of the prediction model based on the fireworks type label; Determining a fireworks prediction probability based on the visual detection result and the text detection result, and determining a fireworks prediction label based on the fireworks prediction probability; Based on the fireworks prediction label and the fireworks type label, updating prompt parameters of the prediction model; When the updated prediction model meets the preset conditions, the target prediction model is obtained.

2. The model training method according to claim 1, characterized in that: The visual model includes a linear layer, a projection layer, and N Vision Transformer layers. The visual detection result is determined based on the fireworks monitoring image using the visual model of the prediction model, including: The output result of the i-th Vision Transformer layer is concatenated with the hint vector of the i+1-th Vision Transformer layer to determine the image sequence of the i+1-th Vision Transformer layer, 1≤i≤N, N is a positive integer; wherein the fireworks monitoring image is input into the linear layer and the projection layer for linear transformation and size alignment to obtain projection features; the projection features are concatenated with the hint vector of the first Vision Transformer layer to determine the image sequence input into the first Vision Transformer layer; Based on the image sequence of the i+1th Vision Transformer layer, using the i+1th Vision Transformer layer to determine an output result of the i+1th Vision Transformer layer; The output result of the last Vision Transformer layer is used as the visual detection result.

3. The model training method according to claim 2, characterized in that: The method of determining an output result of the i+1th Vision Transformer layer based on the image sequence of the i+1th Vision Transformer layer by using the i+1th Vision Transformer layer comprises: Performing linear mapping on the image sequence of the (i+1)th Vision Transformer layer based on a preset projection matrix to obtain a first intermediate feature; Performing residual connection and normalization on the first intermediate feature and the image sequence to obtain a second intermediate feature; Obtaining a third intermediate feature based on the first feedforward network and the second intermediate feature; The third intermediate feature and the second intermediate feature are residually connected and normalized to obtain an output result of the i+1th Vision Transformer layer.

4. The model training method according to claim 1, characterized in that: The text model includes a word embedding layer and M Transformer layers, and the determining of the text detection result based on the fireworks type label and using the text model of the prediction model includes: The output result of the jth Transformer layer is concatenated with the prompt vector of the j+1th Transformer layer to determine the text sequence of the j+1th Transformer layer, 1≤j≤M, M is a positive integer; wherein the fireworks type label is input into the word embedding layer for word segmentation and vectorization processing to obtain an initial text vector; the initial text vector is mapped to the dimensional space corresponding to the prompt vector of the first Transformer layer, and the initial text vector mapped to the dimensional space is concatenated with the prompt vector of the first Transformer layer to determine the text sequence input into the first Transformer layer; Based on the text sequence of the j+1th Transformer layer, using the j+1th Transformer layer to determine the output result of the j+1th Transformer layer; The output result of the last Transformer layer is used as the text detection result.

5. The model training method according to claim 4, characterized in that: The determining the output result of the j+1th Transformer layer by using the j+1th Transformer layer based on the image sequence of the j+1th Transformer layer includes: Linearly map the text sequence of the j+1th Transformer layer based on a preset projection matrix to obtain a fourth intermediate feature; Performing residual connection and normalization on the fourth intermediate feature and the text sequence to obtain a fifth intermediate feature; Based on the second feedforward network and the fifth intermediate feature, a sixth intermediate feature is obtained; The sixth intermediate feature and the fifth intermediate feature are residually connected and normalized to obtain the output result of the j+1th Transformer layer.

6. The model training method according to claim 1, characterized in that: The step of obtaining a target prediction model when the updated prediction model meets a preset condition comprises: Calculate the loss function of the updated prediction model; Determining an evaluation result based on a preset evaluation data set and the updated prediction model; When the loss function satisfies the first condition and / or the evaluation result satisfies the second condition, it is determined that the updated prediction model meets the preset condition, and the updated prediction model is used as the target prediction model.

7. The model training method according to claim 1, characterized in that: The determining of the fireworks prediction probability based on the visual detection result and the text detection result includes: Calculating the similarity between the visual detection result and the text detection result; The fireworks prediction probability is determined based on the similarity.

8. A fire monitoring method based on prompt learning, characterized in that: include: Obtain fireworks monitoring data at a specified location; Based on the fireworks monitoring data, a monitoring result of the designated location is obtained by using a target prediction model, wherein the target prediction model is obtained by the model training method according to any one of claims 1 to 7.

9. An electronic device, characterized in that: The electronic device includes a processor and a memory. When the processor executes a computer program stored in the memory, it implements the model training method as described in any one of claims 1 to 7, or implements the fire monitoring method based on prompt learning as described in claim 8.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one instruction, and when the at least one instruction is executed by the processor, it implements the model training method as described in any one of claims 1 to 7, or implements the fire monitoring method based on prompt learning as described in claim 8.