A smoke semantic segmentation method and system based on reliable query and prototype driving
By designing a smoke semantic segmentation network and utilizing multi-scale feature extraction and feature enhancement techniques, the accuracy problem of smoke image segmentation in complex scenarios was solved, achieving higher accuracy in fire monitoring.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-25
- Publication Date
- 2026-03-10
AI Technical Summary
In existing technologies, the accuracy and versatility of smoke image segmentation in fire monitoring are difficult to meet the requirements, especially in complex scenarios where the misjudgment rate is high due to smoke and other interference conditions.
Design a smoke semantic segmentation network, including a backbone sub-network, a classification branch network, a segmentation branch network, and a prediction module. By acquiring basic features at multiple different scales, the network utilizes the popular spatial deep classification of the classification branch network and the reliable target query extraction module of the segmentation branch network, combined with a local-global interaction module, to enhance the accuracy of feature extraction and prediction.
It significantly improves the accuracy of smoke image segmentation and the ability to locate fine-grained targets, thereby enhancing the accuracy of fire monitoring.
Smart Images

Figure CN119445580B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image segmentation, and in particular to a smoke semantic segmentation method and system based on reliable query and prototype-driven approach. Background Technology
[0002] With the rapid development of computer vision technology, image segmentation has been gradually applied to various security monitoring scenarios. Among them, fire, as a disaster that occurs frequently and seriously endangers social security, has naturally become a key security monitoring scenario for the application of image segmentation technology.
[0003] In current technologies, fire monitoring still relies heavily on manual image analysis and smoke detection. However, manual inspection and monitoring are extremely labor-intensive and costly. Therefore, image segmentation technology has gradually gained importance. However, fire monitoring applications are complex and subject to numerous interference conditions. In addition to common factors such as changes in lighting conditions and weather, smoke can also present many similar-looking targets that could lead to false positives, such as clouds and haze. Furthermore, the color, texture, shape, and transparency of smoke itself vary greatly, and the boundaries of smoke are particularly difficult to define accurately. Consequently, the accuracy and versatility of smoke image segmentation cannot meet the current requirements for fire safety monitoring.
[0004] Therefore, how to design an image segmentation method to improve the accuracy of smoke image segmentation is an urgent problem to be solved. Summary of the Invention
[0005] Based on this, the purpose of this invention is to provide a smoke semantic segmentation method and system based on reliable query and prototype-driven approaches. By designing a smoke semantic segmentation network, including a backbone sub-network, a classification branch network, a segmentation branch network, and a prediction module, the backbone sub-network acquires basic features at multiple scales to obtain semantic information at multiple scales. The classification branch network focuses on clustering the coordinates of components of the same category on the latent manifold while dispersing components of different categories to obtain feature prototypes that are more conducive to smoke component classification. The segmentation branch network, based on the reliable target query extraction module, emphasizes attention within the predicted foreground region rather than focusing on the entire feature map, providing richer and more accurate target location priors for each query feature, greatly promoting fine-grained target localization. Furthermore, the local-global interaction module progressively processes the features at multiple scales acquired by the backbone sub-network to uncover richer details, thereby effectively performing local and global relationship aggregation. This invention significantly improves the accuracy of smoke image segmentation.
[0006] This invention proposes a smoke semantic segmentation method based on reliable query and prototype-driven approach, comprising:
[0007] A single-frame smoke image is acquired in real time and preprocessed, and then the preprocessed smoke image is input into a smoke semantic segmentation network, which includes a backbone sub-network, a classification branch network, a segmentation branch network, and a prediction module.
[0008] The backbone subnetwork acquires multiple basic features at different scales, including low-order basic features, mid-order basic features, and high-order basic features. All of the basic features are input into the segmentation branch network, and the high-order basic features are input into the classification branch network.
[0009] The classification branch network performs deep classification of the higher-order basic features based on the popularity space to obtain multiple feature prototypes. Each feature prototype corresponds to a single higher-order basic feature. The feature prototypes are then input into the segmentation branch network.
[0010] The segmentation branch network performs feature enhancement on the high-order basic features according to the reliable target query extraction module to obtain reliable target query features, and performs feature enhancement on the low-order basic features, the intermediate basic features and the feature prototypes according to the local-global interaction module to obtain detail enhancement features. The reliable target query features and the detail enhancement features are then input into the prediction module.
[0011] The prediction module obtains the final prediction result based on the reliable target query features and the detail enhancement features, and performs correction and optimization based on the smoke focusing loss function.
[0012] In summary, based on the aforementioned smoke semantic segmentation method driven by reliable queries and prototypes, this invention designs a smoke semantic segmentation network comprising a backbone sub-network, a classification branch network, a segmentation branch network, and a prediction module. The backbone sub-network acquires basic features at multiple different scales to obtain semantic information at multiple scales. The classification branch network clusters the coordinates of components of the same category on the latent manifold as much as possible while dispersing components of different categories as much as possible to obtain feature prototypes that are more conducive to the classification of smoke components. The segmentation branch network, based on the reliable target query extraction module, emphasizes attention within the predicted foreground region rather than focusing on the entire feature map, providing richer and more accurate target location priors for each query feature, greatly promoting fine-grained target localization. Furthermore, the local-global interaction module progressively processes the features at multiple scales acquired by the backbone sub-network to mine richer details, thereby effectively performing local and global relation aggregation. This invention significantly improves the accuracy of smoke image segmentation. Specifically, a single-frame smoke image is acquired in real time and preprocessed. The preprocessed smoke image is then input into a smoke semantic segmentation network, which includes a backbone sub-network, a classification branch network, a segmentation branch network, and a prediction module. The backbone sub-network acquires multiple basic features at different scales, including low-order, mid-order, and high-order basic features. All of these basic features are input into the segmentation branch network, and the high-order basic features are input into the classification branch network to obtain semantic information at multiple scales. The classification branch network performs deep classification on the high-order basic features based on the popularity space to obtain multiple feature prototypes. Each feature prototype corresponds uniquely to a high-order basic feature. These feature prototypes are then input into the segmentation branch network. The network obtains feature prototypes that are more conducive to the classification of smoke components. The segmentation branch network enhances the high-order basic features according to the reliable target query extraction module to obtain reliable target query features. It enhances the low-order basic features, the mid-order basic features, and the feature prototypes according to the local-global interaction module to obtain detail enhancement features. The reliable target query features and the detail enhancement features are input into the prediction module. While obtaining richer and more accurate target location priors, it also discovers richer details. The prediction module obtains the final prediction result according to the reliable target query features and the detail enhancement features, and corrects and optimizes it according to the smoke focusing loss function. This invention greatly improves the accuracy of smoke image segmentation.
[0013] Furthermore, the step of the classification branch network performing deep classification of the higher-order basic features based on the popular space to obtain multiple feature prototypes specifically includes:
[0014] The classification branch network is based on the symmetric positive definite matrix of the Riemannian manifold. The classification branch network includes a Gaussian clustering layer, a nonlinear transformation layer, a BN block, a pooling block, a global pooling block, a fully connected layer, and a prediction layer.
[0015] The Gaussian clustering layer acquires Gaussian effective information of higher-order basic features to output a symmetric positive definite matrix corresponding to the Gaussian effective information to the nonlinear transformation layer. The nonlinear transformation layer then performs eigenvalue decomposition and nonlinear transformation. The specific algorithm of the nonlinear transformation layer is as follows:
[0016] ,
[0017] ,
[0018] in, This represents the output of the nonlinear transformation layer. This represents the input to the nonlinear transform layer. This represents the eigenvectors obtained from eigenvalue decomposition. This represents the eigenvalue matrix obtained from eigenvalue decomposition. Represents a nonlinear transformation function. , T Indicates transpose. Represents the x-coordinate of a pixel. Represents the ordinate of a pixel;
[0019] The BN block and pooling block map the symmetric positive definite matrix from the Riemannian manifold to the Euclidean space using a logarithmic function, and then remap the symmetric positive definite matrix back to the manifold space using an exponential function. The global pooling block performs global average pooling to obtain global context prior information. The specific algorithms for the BN block, pooling block, and global pooling block are as follows:
[0020] ,
[0021] ,
[0022] ,
[0023] in, Indicates the output. Indicates input, Indicates BN block, Indicates a pooling block. Represents a global pooling block. , and These represent batch normalization, average pooling, and global average pooling operations in Euclidean space, respectively. Represents an exponential function. Represents a logarithmic function;
[0024] The specific algorithms for the fully connected layer and the prediction layer are as follows:
[0025] ,
[0026] in, This represents the prediction result output by the classification network. This represents the input to the fully connected layer. Indicates a fully connected layer. Represents a non-linear activation function;
[0027] The feature prototype is obtained based on the output of the pooling block and the global pooling block, and the probability that the image contains a smoke target is obtained based on the prediction result output by the classification network.
[0028] Furthermore, the step of the segmentation branch network performing feature enhancement on the higher-order basic features according to the reliable target query extraction module to obtain reliable target query features specifically includes:
[0029] The reliable target query extraction module includes an attention enhancement block and a region enhancement mapping block;
[0030] The attention enhancement module performs feature weighting processing on higher-order basic features. The feature weighting processing includes channel relationship weighting processing and spatial position weighting processing. The spatial position weighting processing includes horizontal position weighting processing and vertical position weighting processing.
[0031] The specific algorithm for the attention enhancement block is as follows:
[0032] ,
[0033] ,
[0034] ,
[0035] ,
[0036] ,
[0037] in, Represents higher-order fundamental features. This represents the inverse bottleneck layer, which comprises three non-linearly activated fully connected layers. Represents learnable weight tensors from different perspectives. Indicates a flip operation. This indicates a channel-by-channel addition operation. Representing feature dimension, This indicates the height, width, and number of channels of the feature. This represents the activation function. and This represents the horizontal position weighted vector and the vertical position weighted vector. This represents the weighted vector of channel relationships. This represents the output features of the attention enhancement module;
[0038] The output features of the attention enhancement module are input into the region enhancement mapping block to obtain reliable target query features.
[0039] Furthermore, the region enhancement mapping block specifically includes:
[0040] The region enhancement mapping block obtains reliable target query features based on the output features of the attention enhancement module. The specific algorithm for obtaining reliable target query features is as follows:
[0041] ,
[0042] ,
[0043] ,
[0044] in, Indicates reliable target query characteristics. This represents the activation function. It is a dropout operation with a probability of 0.3. Indicates a fully connected layer. This represents the intermediate layer output after dimensional transformation. Indicates will b × c × h × w The size of a tensor is converted to b ×1×( c × h × w ), This indicates average pooling with a sampling rate of 2. Indicates batch normalization. This represents the output after a 1×1 convolution. Represents a 1×1 convolution. This represents the output features of the attention enhancement module, which is expanded into a two-dimensional tensor.
[0045] The obtained reliable target query features are then post-processed. The post-processing algorithm is as follows:
[0046] ,
[0047] ,
[0048] ,
[0049] ,
[0050] in, This represents the key matrix involved in attention calculation. This represents the matrix of values involved in the attention calculation. This indicates attention calculation. This indicates the reliable target query characteristics of the final output. T Indicates transpose. This represents matrix multiplication.
[0051] Furthermore, the step of performing feature enhancement on the low-order basic features, the mid-order basic features, and the feature prototype based on the local-global interaction module to obtain detailed enhanced features specifically includes:
[0052] The local-global interaction module includes an inter-domain attention learning block and an NL block;
[0053] The specific algorithm for the inter-domain attention learning block is as follows:
[0054] ,
[0055] ,
[0056] ,
[0057] in, Indicates will Copy 64 times Represents a class prototype tensor, the size of which is b × c × h / 8× w / 8, Representing the key matrix and value matrix, express Divided into 64 sub-regions of the same size, The input features represent the inter-domain attention learning block, and the size of the input features is... b × c × h × w , This indicates attention calculation. This represents the activation function. This represents the output of the inter-domain attention learning block. This means 64× b × c × h / 8× w The size of / 8 tensors is converted to b× c × h × w , It represents the Hadamardi (or Hadama) stack;
[0058] The specific algorithm for the NL block is as follows:
[0059] ,
[0060] ,
[0061] in, Represents the query matrix. Indicates the input of the NL block, Will Mapped to the corresponding query matrix Q Key matrix K Sum matrix V , Represents a 1×1 convolution. Indicates the output of the NL block. Key matrix K The length of the vector T Indicates transpose. The object label representing the service provided by the convolution function;
[0062] Detail-enhancing features are obtained from the outputs of the inter-domain attention learning block and the NL block.
[0063] Furthermore, the step of the prediction module obtaining the final prediction result based on the reliable target query features and the detail enhancement features specifically includes:
[0064] The prediction module includes an upsampling convolution module, a convolution module, and a sigmoid non-linear activation layer;
[0065] The upsampling convolution module includes a 2x bilinear interpolation upsampling layer, a 3×3 convolutional layer, a BN layer, and a ReLU layer;
[0066] The convolutional module includes a 3×3 convolutional layer, a BN layer, and a ReLU layer;
[0067] The prediction module maps reliable target query features and detail enhancement features from the feature space to the semantic space to obtain the final segmentation result.
[0068] Furthermore, the step of correcting and optimizing based on the smoke focusing loss function specifically includes:
[0069] The smoke focusing loss function is corrected and optimized, and the specific algorithm for the smoke focusing loss function is as follows:
[0070] ,
[0071] ,
[0072] ,
[0073] ,
[0074] ,
[0075] ,
[0076] ,
[0077] ,
[0078] in, This represents the smoke focusing loss function. and This indicates the weighting of foreground and background losses. Indicates a loss of prospects. Indicates background loss. express Number of mid-foreground pixels express Number of mid-foreground pixels express Number of background pixels express Number of background pixels The segmentation label represents the final segmentation result after flattening. This represents the final segmentation result after flattening and binarization. Represents a vector consisting entirely of 1s. Indicates transpose. Indicates the flattening operation. The segmentation label represents the final segmentation result. This represents the final segmentation result of binarization. This indicates the final segmentation result. Indicates the coordinates of the feature location. Indicates the feature height of the final segmentation result. Indicates the feature width of the final segmentation result. Represents a positive integer.
[0079] This invention proposes a smoke semantic segmentation system based on reliable querying and prototype-driven methods, comprising:
[0080] The preprocessing module is used to acquire single-frame smoke images in real time and perform preprocessing, so that the preprocessed smoke images are input into the smoke semantic segmentation network, which includes a backbone sub-network, a classification branch network, a segmentation branch network and a prediction module.
[0081] The basic feature extraction module is used to obtain multiple basic features at different scales in the backbone sub-network. The basic features include low-order basic features, mid-order basic features and high-order basic features. All the basic features are input into the segmentation branch network and the high-order basic features are input into the classification branch network.
[0082] A classification module is used for the classification branch network to perform deep classification of the higher-order basic features based on the popular space to obtain multiple feature prototypes, each feature prototype having a single correspondence with the higher-order basic features, and inputting the feature prototypes into the segmentation branch network.
[0083] The segmentation module is used to perform feature enhancement on the high-order basic features according to the reliable target query extraction module to obtain reliable target query features, and to perform feature enhancement on the low-order basic features, the mid-order basic features and the feature prototype according to the local-global interaction module to obtain detail enhancement features. The reliable target query features and the detail enhancement features are then input into the prediction module.
[0084] The prediction optimization module is used by the prediction module to obtain the final prediction result based on the reliable target query features and the detail enhancement features, and to perform correction and optimization based on the smoke focusing loss function.
[0085] The present invention also provides a storage medium that stores one or more programs that, when executed by a processor, implement the smoke semantic segmentation method based on reliable query and prototype-driven methods as described above.
[0086] The present invention also provides a computer device, the computer device including a memory and a processor, wherein:
[0087] The memory is used to store computer programs;
[0088] When the processor executes the computer program stored in the memory, it implements the smoke semantic segmentation method based on reliable query and prototype-driven methods as described above. Attached Figure Description
[0089] Figure 1 The flowchart shows the smoke semantic segmentation method based on reliable query and prototype-driven proposed in the first embodiment of the present invention.
[0090] Figure 2The flowchart shows the smoke semantic segmentation method based on reliable query and prototype-driven proposed in the second embodiment of the present invention.
[0091] Figure 3 This is a schematic diagram of the smoke semantic segmentation system based on reliable query and prototype-driven approach proposed in the third embodiment of the present invention.
[0092] Figure 4 This is a flowchart illustrating the structure of the smoke semantic segmentation network proposed in the first embodiment of the present invention.
[0093] The following detailed description, in conjunction with the accompanying drawings, will further illustrate the present invention. Detailed Implementation
[0094] To facilitate understanding of the present invention, a more complete description will be given below with reference to the accompanying drawings. Several embodiments of the invention are illustrated in the drawings. However, the invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete.
[0095] It should be noted that when a component is said to be "fixed to" another component, it can be directly on the other component or there may be an intervening component. When a component is said to be "connected to" another component, it can be directly connected to the other component or there may be an intervening component. The terms "vertical," "horizontal," "left," "right," and similar expressions used in this document are for illustrative purposes only.
[0096] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0097] Please see Figure 1 The diagram shows a flowchart of the smoke semantic segmentation method based on reliable query and prototype-driven approach proposed in the first embodiment of the present invention. This smoke semantic segmentation method based on reliable query and prototype-driven approach includes steps S01 to S05, wherein:
[0098] Step S01: Acquire single-frame smoke images in real time and preprocess them, so that the preprocessed smoke images can be input into the smoke semantic segmentation network;
[0099] It should be noted that the smoke image input smoke semantic segmentation network in this embodiment includes a backbone sub-network, a classification branch network, a segmentation branch network, and a prediction module. For the specific structure, please refer to [reference needed]. Figure 4 .
[0100] Step S02: The backbone subnetwork acquires multiple basic features at different scales, including low-order basic features, mid-order basic features and high-order basic features. All basic features are input into the segmentation branch network, and the high-order basic features are input into the classification branch network.
[0101] Step S03: The classification branch network performs deep classification of high-order basic features based on the popular space to obtain multiple feature prototypes, and inputs the feature prototypes into the segmentation branch network.
[0102] It should be noted that in this embodiment, the classification branch network is based on the symmetric positive definite matrix of the Riemannian manifold. The classification branch network includes a Gaussian clustering layer, a nonlinear transformation layer, a BN block, a pooling block, a global pooling block, a fully connected layer, and a prediction layer.
[0103] The Gaussian clustering layer acquires Gaussian effective information of higher-order basic features to output a symmetric positive definite matrix corresponding to the Gaussian effective information to the nonlinear transformation layer. The nonlinear transformation layer then performs eigenvalue decomposition and nonlinear transformation. The specific algorithm of the nonlinear transformation layer is as follows:
[0104] ,
[0105] ,
[0106] in, This represents the output of the nonlinear transformation layer. This represents the input to the nonlinear transform layer. This represents the eigenvectors obtained from eigenvalue decomposition. This represents the eigenvalue matrix obtained from eigenvalue decomposition. Represents a nonlinear transformation function. , T Indicates transpose. Represents the x-coordinate of a pixel. Represents the ordinate of a pixel;
[0107] The BN block and pooling block map the symmetric positive definite matrix from the Riemannian manifold to the Euclidean space using a logarithmic function, and then remap the symmetric positive definite matrix back to the manifold space using an exponential function. The global pooling block performs global average pooling to obtain global context prior information. The specific algorithms for the BN block, pooling block, and global pooling block are as follows:
[0108] ,
[0109] ,
[0110] ,
[0111] in, Indicates the output. Indicates input, Indicates BN block, Indicates a pooling block. Represents a global pooling block. , and These represent batch normalization, average pooling, and global average pooling operations in Euclidean space, respectively. Represents an exponential function. Represents a logarithmic function;
[0112] The specific algorithms for the fully connected layer and the prediction layer are as follows:
[0113] ,
[0114] in, This represents the prediction result output by the classification network. This represents the input to the fully connected layer. Indicates a fully connected layer. Represents a non-linear activation function;
[0115] The feature prototype is obtained based on the output of the pooling block and the global pooling block, and the probability that the image contains a smoke target is obtained based on the prediction result output by the classification network.
[0116] Step S04: The segmentation branch network performs feature enhancement on the high-order basic features according to the reliable target query extraction module to obtain reliable target query features, and performs feature enhancement on the low-order basic features, mid-order basic features and feature prototypes according to the local-global interaction module to obtain detailed enhancement features. The reliable target query features and detailed enhancement features are then input into the prediction module.
[0117] It should be noted that the reliable target query extraction module in this embodiment includes an attention enhancement block and a region enhancement mapping block;
[0118] The attention enhancement module performs feature weighting processing on higher-order basic features. The feature weighting processing includes channel relationship weighting processing and spatial position weighting processing. The spatial position weighting processing includes horizontal position weighting processing and vertical position weighting processing.
[0119] The specific algorithm for the attention enhancement block is as follows:
[0120] ,
[0121] ,
[0122] ,
[0123] ,
[0124] ,
[0125] in, Represents higher-order fundamental features. This represents the inverse bottleneck layer, which comprises three non-linearly activated fully connected layers. Represents learnable weight tensors from different perspectives. Indicates a flip operation. This indicates a channel-by-channel addition operation. Representing feature dimension, This indicates the height, width, and number of channels of the feature. This represents the activation function. and This represents the horizontal position weighted vector and the vertical position weighted vector. This represents the weighted vector of channel relationships. This represents the output features of the attention enhancement module;
[0126] The output features of the attention enhancement module are input into the region enhancement mapping block to obtain reliable target query features;
[0127] In this embodiment, the region enhancement mapping block obtains reliable target query features based on the output features of the attention enhancement module. The specific algorithm for obtaining reliable target query features is as follows:
[0128] ,
[0129] ,
[0130] ,
[0131] in, Indicates reliable target query characteristics. This represents the activation function. It is a dropout operation with a probability of 0.3. Indicates a fully connected layer. This represents the intermediate layer output after dimensional transformation. Indicates will b × c × h × w The size of a tensor is converted to b ×1×( c × h × w ), This indicates average pooling with a sampling rate of 2. Indicates batch normalization. This represents the output after a 1×1 convolution. Represents a 1×1 convolution. This represents the output features of the attention enhancement module, which is expanded into a two-dimensional tensor.
[0132] The obtained reliable target query features are then post-processed. The post-processing algorithm is as follows:
[0133] ,
[0134] ,
[0135] ,
[0136] ,
[0137] in, This represents the key matrix involved in attention calculation. This represents the matrix of values involved in the attention calculation. This indicates attention calculation. This indicates the reliable target query characteristics of the final output. T Indicates transpose. Represents matrix multiplication;
[0138] In this embodiment, the local-global interaction module includes an inter-domain attention learning block and an NL block;
[0139] The specific algorithm for the inter-domain attention learning block is as follows:
[0140] ,
[0141] ,
[0142] ,
[0143] in, Indicates will Copy 64 times Represents a class prototype tensor, the size of which is b × c × h / 8× w / 8, Representing the key matrix and value matrix, express Divided into 64 sub-regions of the same size, The input features represent the inter-domain attention learning block, and the size of the input features is... b × c ×h × w , This indicates attention calculation. This represents the activation function. This represents the output of the inter-domain attention learning block. This means 64× b × c × h / 8× w The size of / 8 tensors is converted to b × c × h × w , It represents the Hadamardi (or Hadama) stack;
[0144] The specific algorithm for the NL block is as follows:
[0145] ,
[0146] ,
[0147] in, Represents the query matrix. Indicates the input of the NL block, Will Mapped to the corresponding query matrix Q Key matrix K Sum matrix V , Represents a 1×1 convolution. Indicates the output of the NL block. Key matrix K The length of the vector T Indicates transpose. The object label representing the service provided by the convolution function;
[0148] Detail-enhancing features are obtained from the outputs of the inter-domain attention learning block and the NL block.
[0149] Step S05: The prediction module obtains the final prediction result based on the reliable target query features and detail enhancement features, and performs correction and optimization based on the smoke focusing loss function;
[0150] It should be noted that, in this embodiment, the prediction module includes an upsampling convolution module, a convolution module, and a sigmoid nonlinear activation layer;
[0151] The upsampling convolution module includes a 2x bilinear interpolation upsampling layer, a 3×3 convolutional layer, a BN layer, and a ReLU layer;
[0152] The convolutional module includes a 3×3 convolutional layer, a BN layer, and a ReLU layer;
[0153] The prediction module maps reliable target query features and detail enhancement features from the feature space to the semantic space to obtain the final segmentation result;
[0154] In this embodiment, the smoke focusing loss function is corrected and optimized. The specific algorithm for the smoke focusing loss function is as follows:
[0155] ,
[0156] ,
[0157] ,
[0158] ,
[0159] ,
[0160] ,
[0161] ,
[0162] ,
[0163] in, This represents the smoke focusing loss function. and This indicates the weighting of foreground and background losses. Indicates a loss of prospects. Indicates background loss. express Number of mid-foreground pixels express Number of mid-foreground pixels express Number of background pixels express Number of background pixels The segmentation label represents the final segmentation result after flattening. This represents the final segmentation result after flattening and binarization. Represents a vector consisting entirely of 1s. Indicates transpose. Indicates the flattening operation. The segmentation label represents the final segmentation result. This represents the final segmentation result of binarization. This indicates the final segmentation result. Indicates the coordinates of the feature location. Indicates the feature height of the final segmentation result. Indicates the feature width of the final segmentation result. Represents a positive integer.
[0164] In summary, based on the aforementioned smoke semantic segmentation method driven by reliable queries and prototypes, this invention designs a smoke semantic segmentation network comprising a backbone sub-network, a classification branch network, a segmentation branch network, and a prediction module. The backbone sub-network acquires basic features at multiple different scales to obtain semantic information at multiple scales. The classification branch network clusters the coordinates of components of the same category on the latent manifold as much as possible while dispersing components of different categories as much as possible to obtain feature prototypes that are more conducive to the classification of smoke components. The segmentation branch network, based on the reliable target query extraction module, emphasizes attention within the predicted foreground region rather than focusing on the entire feature map, providing richer and more accurate target location priors for each query feature, greatly promoting fine-grained target localization. Furthermore, the local-global interaction module progressively processes the features at multiple scales acquired by the backbone sub-network to mine richer details, thereby effectively performing local and global relation aggregation. This invention significantly improves the accuracy of smoke image segmentation. Specifically, a single-frame smoke image is acquired in real time and preprocessed. The preprocessed smoke image is then input into a smoke semantic segmentation network, which includes a backbone sub-network, a classification branch network, a segmentation branch network, and a prediction module. The backbone sub-network acquires multiple basic features at different scales, including low-order, mid-order, and high-order basic features. All of these basic features are input into the segmentation branch network, and the high-order basic features are input into the classification branch network to obtain semantic information at multiple scales. The classification branch network performs deep classification on the high-order basic features based on the popularity space to obtain multiple feature prototypes. Each feature prototype corresponds uniquely to a high-order basic feature. These feature prototypes are then input into the segmentation branch network. The network obtains feature prototypes that are more conducive to the classification of smoke components. The segmentation branch network enhances the high-order basic features according to the reliable target query extraction module to obtain reliable target query features. It enhances the low-order basic features, the mid-order basic features, and the feature prototypes according to the local-global interaction module to obtain detail enhancement features. The reliable target query features and the detail enhancement features are input into the prediction module. While obtaining richer and more accurate target location priors, it also discovers richer details. The prediction module obtains the final prediction result according to the reliable target query features and the detail enhancement features, and corrects and optimizes it according to the smoke focusing loss function. This invention greatly improves the accuracy of smoke image segmentation.
[0165] Please see Figure 2The diagram shows a flowchart of a smoke semantic segmentation method based on reliable query and prototype-driven approach proposed in the second embodiment of the present invention. This smoke semantic segmentation method based on reliable query and prototype-driven approach includes steps S11 to S16, wherein:
[0166] Step S11: Acquire single-frame smoke images in real time and preprocess them, so that the preprocessed smoke images can be input into the smoke semantic segmentation network;
[0167] Step S12: The backbone subnetwork acquires multiple basic features at different scales, including low-order basic features, mid-order basic features and high-order basic features. All basic features are input into the segmentation branch network and the high-order basic features are input into the classification branch network.
[0168] Step S13: The Gaussian clustering layer of the classification branch network obtains Gaussian effective information of high-order basic features to output a symmetric positive definite matrix corresponding to the Gaussian effective information to the nonlinear transformation layer. The nonlinear transformation layer then performs eigenvalue decomposition and nonlinear transformation. The BN block and pooling block map the symmetric positive definite matrix from the Riemannian manifold to the Euclidean space according to the logarithmic function, and then remap the symmetric positive definite matrix back to the manifold space through the exponential function. The global pooling block performs global average pooling to obtain global context prior information. The feature prototype is obtained according to the output of the pooling block and the global pooling block. The probability that the image contains a smoke target is obtained according to the prediction result output by the classification network.
[0169] Step S14: The reliable target query extraction module includes an attention enhancement block and a region enhancement mapping block. The attention enhancement module performs feature weighting processing on the high-order basic features, and the output features of the attention enhancement module are input into the region enhancement mapping block to obtain reliable target query features.
[0170] Step S15: The local-global interaction module includes an inter-domain attention learning block and an NL block, and obtains detailed enhancement features based on the outputs of the inter-domain attention learning block and the NL block;
[0171] Step S16: The prediction module maps reliable target query features and detail enhancement features from the feature space to the semantic space to obtain the final segmentation result, and performs correction and optimization based on the smoke focusing loss function.
[0172] In summary, based on the aforementioned smoke semantic segmentation method driven by reliable queries and prototypes, this invention designs a smoke semantic segmentation network comprising a backbone sub-network, a classification branch network, a segmentation branch network, and a prediction module. The backbone sub-network acquires basic features at multiple different scales to obtain semantic information at multiple scales. The classification branch network clusters the coordinates of components of the same category on the latent manifold as much as possible while dispersing components of different categories as much as possible to obtain feature prototypes that are more conducive to the classification of smoke components. The segmentation branch network, based on the reliable target query extraction module, emphasizes attention within the predicted foreground region rather than focusing on the entire feature map, providing richer and more accurate target location priors for each query feature, greatly promoting fine-grained target localization. Furthermore, the local-global interaction module progressively processes the features at multiple scales acquired by the backbone sub-network to mine richer details, thereby effectively performing local and global relation aggregation. This invention significantly improves the accuracy of smoke image segmentation. Specifically, a single-frame smoke image is acquired in real time and preprocessed. The preprocessed smoke image is then input into a smoke semantic segmentation network, which includes a backbone sub-network, a classification branch network, a segmentation branch network, and a prediction module. The backbone sub-network acquires multiple basic features at different scales, including low-order, mid-order, and high-order basic features. All of these basic features are input into the segmentation branch network, and the high-order basic features are input into the classification branch network to obtain semantic information at multiple scales. The classification branch network performs deep classification on the high-order basic features based on the popularity space to obtain multiple feature prototypes. Each feature prototype corresponds uniquely to a high-order basic feature. These feature prototypes are then input into the segmentation branch network. The network obtains feature prototypes that are more conducive to the classification of smoke components. The segmentation branch network enhances the high-order basic features according to the reliable target query extraction module to obtain reliable target query features. It enhances the low-order basic features, the mid-order basic features, and the feature prototypes according to the local-global interaction module to obtain detail enhancement features. The reliable target query features and the detail enhancement features are input into the prediction module. While obtaining richer and more accurate target location priors, it also discovers richer details. The prediction module obtains the final prediction result according to the reliable target query features and the detail enhancement features, and corrects and optimizes it according to the smoke focusing loss function. This invention greatly improves the accuracy of smoke image segmentation.
[0173] Please see Figure 3 The figure shows a schematic diagram of the smoke semantic segmentation system based on reliable query and prototype-driven approach proposed in the third embodiment of the present invention. The system includes:
[0174] The preprocessing module 10 is used to acquire a single frame of smoke image in real time and perform preprocessing, so as to input the preprocessed smoke image into the smoke semantic segmentation network. The smoke semantic segmentation network includes a backbone sub-network, a classification branch network, a segmentation branch network and a prediction module.
[0175] The basic feature extraction module 20 is used to obtain multiple basic features at different scales in the backbone sub-network. The basic features include low-order basic features, mid-order basic features and high-order basic features. All the basic features are input into the segmentation branch network and the high-order basic features are input into the classification branch network.
[0176] The classification module 30 is used for the classification branch network to perform deep classification on the higher-order basic features based on the popularity space to obtain multiple feature prototypes, each feature prototype having a single correspondence with the higher-order basic features, and inputting the feature prototypes into the segmentation branch network.
[0177] The segmentation module 40 is used to perform feature enhancement on the high-order basic features according to the reliable target query extraction module to obtain reliable target query features, and to perform feature enhancement on the low-order basic features, the intermediate basic features and the feature prototypes according to the local-global interaction module to obtain detail enhancement features, and to input the reliable target query features and the detail enhancement features into the prediction module.
[0178] The prediction optimization module 50 is used by the prediction module to obtain the final prediction result based on the reliable target query features and the detail enhancement features, and to perform correction and optimization based on the smoke focusing loss function.
[0179] The present invention also proposes a computer storage medium storing one or more programs that, when executed by a processor, implement the aforementioned smoke semantic segmentation method based on reliable querying and prototype-driven methods.
[0180] The present invention also proposes a computer device, including a memory and a processor, wherein the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory to implement the above-described smoke semantic segmentation method based on reliable query and prototype-driven approach.
[0181] Those skilled in the art will understand that the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can mean any means that can contain stored, communicated, propagated, or transmitted programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0182] More specific examples of computer-readable media (a non-exhaustive list) include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.
[0183] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0184] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0185] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be determined by the appended claims.
Claims
1. A method for smoke semantic segmentation based on reliable query and prototype driving, characterized in that, The method comprises the following steps: Real-time acquisition of a single frame of smoke image and preprocessing, so as to input the preprocessed smoke image into a smoke semantic segmentation network, wherein the smoke semantic segmentation network comprises a backbone sub-network, a classification branch network, a segmentation branch network and a prediction module; The backbone sub-network acquires a plurality of basic features of different scales, wherein the basic features comprise low-order basic features, middle-order basic features and high-order basic features, all the basic features are input into the segmentation branch network, and the high-order basic features are input into the classification branch network; The classification branch network performs deep classification on the high-order basic features based on a popular space to acquire a plurality of feature prototypes, wherein the feature prototypes are in one-to-one correspondence with the high-order basic features, and the feature prototypes are input into the segmentation branch network; The segmentation branch network performs feature enhancement on the high-order basic features according to a reliable target query extraction module to acquire reliable target query features, performs feature enhancement on the low-order basic features, the middle-order basic features and the feature prototypes according to a local-global interaction module to acquire detail enhancement features, and inputs the reliable target query features and the detail enhancement features into the prediction module; The prediction module acquires a final prediction result according to the reliable target query features and the detail enhancement features, and performs correction and optimization according to a smoke focus loss function.
2. The reliable query and prototype-driven based smoke semantic segmentation method according to claim 1, wherein, The step of performing deep classification on the high-order basic features based on a popular space to acquire a plurality of feature prototypes comprises the following steps: The classification branch network is based on a symmetric positive definite matrix of a Riemannian manifold, and comprises a Gaussian aggregation layer, a nonlinear transformation layer, a BN block, a pooling block, a global pooling block, a full connection layer and a prediction layer; The Gaussian aggregation layer acquires Gaussian effective information of the high-order basic features to output a symmetric positive definite matrix corresponding to the Gaussian effective information to the nonlinear transformation layer, and the nonlinear transformation layer further performs eigenvalue decomposition and nonlinear transformation, and the specific algorithm of the nonlinear transformation layer is as follows: , , wherein, represents an output of the nonlinear transformation layer, represents an input of the nonlinear transformation layer, represents an eigenvector obtained by eigenvalue decomposition, represents an eigenvalue matrix obtained by eigenvalue decomposition, represents a nonlinear transformation function, , T represents a transpose, represents a horizontal coordinate of a pixel point, represents a vertical coordinate of a pixel point; The BN block and the pooling block map the symmetric positive definite matrix from the Riemannian manifold to the Euclidean space according to a logarithmic function, and then remap the symmetric positive definite matrix back to the manifold space through an exponential function, and the global pooling block performs global average pooling to acquire global context prior information, and the specific algorithm of the BN block, the pooling block and the global pooling block is as follows: , , , wherein, represents an output, represents an input, represents a BN block, represents a pooling block, represents a global pooling block, , and represent batch normalization, average pooling and global average pooling operations in the Euclidean space, respectively, represents an exponential function, represents a logarithmic function; The specific algorithm of the full connection layer and the prediction layer is as follows: , wherein, denotes the prediction result output by the classification network, denotes the input to the fully connected layer, denotes the fully connected layer, denotes a non-linear activation function; The feature prototypes are acquired according to the outputs of the pooling block and the global pooling block, and the probability of containing a smoke target in the image is acquired according to the prediction result output by the classification network.
3. The reliable query and prototype-driven based smoke semantic segmentation method according to claim 1, wherein, The step of performing feature enhancement on the high-order basic features according to a reliable target query extraction module to acquire reliable target query features comprises the following steps: The reliable target query extraction module comprises an attention enhancement block and a region enhancement mapping block; The attention enhancement block performs feature weighting processing on the high-order basic features, wherein the feature weighting processing comprises channel relationship weighting processing and spatial position weighting processing, and the spatial position weighting processing comprises horizontal position weighting processing and vertical position weighting processing; The specific algorithm of the attention enhancement block is as follows: , , , , , wherein, represents a high-order base feature, represents an inverse bottleneck layer, the inverse bottleneck layer comprising three non-linear activation fully connected layers, represents a learnable weight tensor of different angles, represents a flip operation, represents a channel-wise addition operation, represents a feature dimension, represents a height, a width, a number of channels of a feature, represents an activation function, and represents a horizontal position weighting vector and a vertical position weighting vector, represents a channel relationship weighting vector, represents an output feature of the attention enhancement module; The output feature of the attention enhancement module is input into a region enhancement mapping block to obtain reliable target query features.
4. The method of claim 3, wherein, The region enhancement mapping block specifically comprises: The region enhancement mapping block obtains reliable target query features according to the output feature of the attention enhancement module, and the specific algorithm for obtaining the reliable target query features is as follows: , , , wherein, denotes a reliable target query feature, denotes an activation function, is a dropout operation with a probability of 0.3, denotes a fully connected layer, denotes an intermediate layer output after dimension transformation, denotes a concatenation of b × c × h × w the size of the tensor is converted to b ×1×( c × h × w ), denotes an average pooling with a sampling rate of 2, denotes batch normalization, denotes an output after 1×1 convolution, denotes 1×1 convolution, denotes an output feature of the attention enhancement module unfolded into a two-dimensional tensor; The obtained reliable target query features are further processed, and the algorithm for the processing is as follows: , , , , wherein, denotes a key matrix participating in attention computation, denotes a value matrix participating in attention computation, denotes attention computation, denotes a reliable target query feature for final output, T denotes transpose, denotes matrix multiplication.
5. The reliable query and prototype-driven based smoke semantic segmentation method according to claim 1, wherein, The step of obtaining the detail enhancement feature according to the output of the inter-domain attention learning block and the NL block specifically comprises: The inter-domain attention learning block comprises an inter-domain attention learning block and an NL block. The specific algorithm of the inter-domain attention learning block is as follows: , , , wherein, represents copying 64 times, represents a class prototype tensor, a size of the class prototype tensor being b × c × h / 8× w / 8, represents a key matrix and a value matrix, represents divided into 64 sub-regions of the same size, represents an input feature of an inter-domain attention learning block, a size of the input feature being b × c × h × w , represents attention calculation, represents an activation function, represents an output of the inter-domain attention learning block, represents converting a size of a 64× b × c × h / 8× w / 8 tensor into b × c × h × w , represents a Hadamard product; The specific algorithm of the NL block is as follows: , , wherein, denotes a query matrix, denotes an input of the NL block, maps to a corresponding query matrix Q , a key matrix K and a value matrix V , denotes a 1 x 1 convolution, denotes an output of the NL block, denotes a key matrix K vector length, T denotes a transpose, denotes an object label served by a convolution function The detail enhancement feature is obtained according to the output of the inter-domain attention learning block and the NL block.
6. The reliable query and prototype-driven based smoke semantic segmentation method according to claim 1, wherein, The step of obtaining the final prediction result according to the reliable target query feature and the detail enhancement feature by the prediction module specifically comprises: The prediction module comprises an up-sampling convolution module, a convolution module and a sigmoid nonlinear activation layer. The up-sampling convolution module comprises a 2 times bilinear interpolation up-sampling layer, a 3*3 convolution layer, a BN layer and a ReLu layer. The convolution module comprises a 3*3 convolution layer, a BN layer and a ReLu layer. The prediction module maps the reliable target query feature and the detail enhancement feature from a feature space to a semantic space to obtain a final segmentation result.
7. The reliable query and prototype-driven based smoke semantic segmentation method according to claim 1, wherein, The step of performing correction optimization according to the smoke focus loss function specifically comprises: The smoke focus loss function is specifically as follows: , , , , , , , , wherein, denotes a smoke focus loss function, and denotes foreground and background loss assignment weights, denotes foreground loss, denotes background loss, denotes foreground pixel count in, denotes foreground pixel count in, denotes background pixel count in, denotes background pixel count in, denotes segmentation label of flattened final segmentation result, denotes binarized final segmentation result, denotes all-ones vector, denotes transpose, denotes flattening operation, denotes segmentation label of final segmentation result, denotes binarized final segmentation result, denotes final segmentation result, denotes feature position coordinates, denotes feature height of final segmentation result, denotes feature width of final segmentation result, denotes positive integer.
8. A smoke semantic segmentation system based on reliable query and prototype driving, characterized in that, The pre-processing module is configured to acquire a single-frame smoke image in real time and perform pre-processing, so as to input the pre-processed smoke image into a smoke semantic segmentation network. The basic feature extraction module is configured to acquire a plurality of basic features of different scales by the backbone sub-network, and the basic features comprise low-order basic features, middle-order basic features and high-order basic features. The classification module is configured to perform deep classification on the high-order basic features based on a popular space by the classification branch network, so as to obtain a plurality of feature prototypes, and the feature prototypes are in one-to-one correspondence with the high-order basic features. The segmentation module is configured to perform feature enhancement on the high-order basic features by a reliable target query extraction module, so as to obtain reliable target query features, and perform feature enhancement on the low-order basic features, the middle-order basic features and the feature prototypes by a local-global interaction module, so as to obtain detail enhancement features. The prediction optimization module is configured to obtain a final prediction result according to the reliable target query features and the detail enhancement features by the prediction module, and perform correction optimization according to a smoke focus loss function. 9. A storage medium, characterized by The storage medium stores one or more programs, which are executed by the processor to implement the smoke semantic segmentation method based on reliable query and prototype driving according to any one of claims 1-7.
10. A computer device, comprising: The computer device comprises a memory and a processor, wherein: The memory is used to store a computer program; The processor is used to execute the computer program stored on the memory to implement the smoke semantic segmentation method based on reliable query and prototype driving according to any one of claims 1-7.
Citation Information
Patent Citations
Full-convolution real-time video instance segmentation method
CN115171020A
Image smoke fine detection method based on texture perception
CN115731401A