A method for detecting pulse defocused regions based on spatiotemporal information interaction
By constructing a pulse defocusing region detection network based on spatiotemporal information interaction, the problem that existing methods are not applicable to three-dimensional pulse sequences is solved, and efficient defocusing region detection and segmentation of pulse sequences is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-21
- Publication Date
- 2026-06-30
Smart Images

Figure CN122090063B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of pulse defocusing region detection technology based on spatiotemporal information interaction, and particularly relates to a pulse defocusing region detection method based on spatiotemporal information interaction. Background Technology
[0002] Neuromorphic sensors, which mimic the perceptual structures and mechanisms of biological vision systems, offer significant advantages in dynamic perception, high temporal resolution, and low power consumption. In recent years, they have shown broad application prospects in fields such as autonomous driving, UAV visual navigation, industrial inspection, and visual surveillance. While pulse cameras, as a novel biomimetic vision sensor, currently still widely use traditional optical lenses, the physical limitations of optical lenses cause objects outside the depth of field to appear out of focus. Multi-focus fusion technology can fuse image information from multiple locally focused areas into a single image, providing richer visual information. This is particularly advantageous for assisting in biological research and medical diagnosis, industrial flaw detection, and target detection.
[0003] In multi-focus fusion, accurate focus region information is needed to achieve more precise focusing of the foreground and background. After initial global focusing on the pulse focus stack, the corresponding pulse sequence positions and pulse data can be obtained. Given a digital image, the goal of defocus blur detection (DBD) is to segment out relatively complete out-of-focus regions. However, current DBD research mainly focuses on digital image parts. Furthermore, due to the special representation of pulse data and the noise mixed in the pulse data, existing DBD methods are not suitable for three-dimensional pulse sequences. Summary of the Invention
[0004] To address the shortcomings of current technologies, this invention proposes a pulse defocusing region detection method based on spatiotemporal information interaction, thus solving the problems mentioned in the background technology.
[0005] To achieve the above objectives, the present invention provides the following technical solution: a pulse defocusing region detection method based on spatiotemporal information interaction, comprising:
[0006] Step S1: Input the continuous pulse sequence acquired by the pulse camera into the spatiotemporal aggregation pulse feature extraction module to extract spatiotemporal features. The spatiotemporal features include global pulse sequence features, first pulse subsequence features, second pulse subsequence features, and third pulse subsequence features.
[0007] Step S2: Construct a defocused region detection network, which includes a feature pyramid, an enhanced multi-scale spatiotemporal feature extraction module, a spatiotemporal interaction aggregation module, and a spatiotemporal interaction decoder; construct a feature pyramid, set the acquired spatiotemporal features as the bottom layer features of the feature pyramid, and extract multi-scale features from the spatiotemporal features;
[0008] Step S3: Input the extracted multi-scale features into the enhanced multi-scale spatiotemporal feature extraction module for fusion to obtain the fused spatiotemporal features;
[0009] Step S4: Input the first pulse subsequence features, the second pulse subsequence features, and the third pulse subsequence features into the spatiotemporal interaction aggregation module for processing to obtain the output of the spatiotemporal interaction aggregation module. Input the output of the spatiotemporal interaction aggregation module and the fused spatiotemporal features into the spatiotemporal interaction decoder for processing to generate a preliminary segmentation feature map.
[0010] Step S5: Input the global pulse sequence features and preliminary segmentation feature map from step S1 into the multi-scale edge refinement module to refine the edges and obtain the target region segmentation result;
[0011] Step S6: Set a loss function to optimize the performance of the defocused region detection network, and reconstruct and evaluate how closely the target region segmentation results approximate the real image.
[0012] Furthermore, the specific process of extracting spatiotemporal features in step S1 is as follows:
[0013] Step S11: Continuous pulse sequence Along the time axis, the pulse camera's frame rate adaptively divides the data into three non-overlapping pulse subsequences. These three pulse subsequences are: the first pulse subsequence, the second pulse subsequence, and the third pulse subsequence. Second pulse subsequence The third pulse subsequence ;
[0014] Step S12: For the continuous pulse sequence Global feature extraction was performed using the thirty-second 1×1 convolution on the continuous pulse sequence. After dimensionality reduction, the 33rd 3×3 convolution is used for extraction to obtain global pulse sequence features. ;
[0015] Step S13: The first pulse subsequence The first local feature extraction is processed to obtain the first processed feature, which is then combined with the global pulse sequence feature. The pulse features are weighted and fused channel by channel (C) to obtain the features of the first pulse subsequence. ;
[0016] The second pulse subsequence The input is processed by the second local feature extraction to obtain the second processed feature. The second processed feature is then combined with the global pulse sequence feature. The pulse features are weighted and fused channel by channel (C) to obtain the second pulse subsequence features. ;
[0017] The third pulse subsequence The input is processed by third local feature extraction to obtain third processed features, which are then combined with global pulse sequence features. The pulse features are weighted and fused channel by channel (C) to obtain the features of the third pulse subsequence. The spatiotemporal characteristics are , , and .
[0018] Furthermore, the specific process for extracting multi-scale features in step S2 is as follows:
[0019] Extracted from step S1 , , The features are set as the bottom layer of the feature pyramid, and a four-layer feature pyramid is constructed. The four-layer feature pyramid consists of a first feature code, a second feature code, a third feature code, and a fourth feature code. The first feature code has 32 feature channels, the second feature code has 64 feature channels, the third feature code has 128 feature channels, and the fourth feature code has 256 feature channels.
[0020] Will , , and The features are fused to obtain the first fused feature. This first fused feature is then sequentially input into the first feature encoder, the second feature encoder, the third feature encoder, and the fourth feature encoder for convolution and PReLU activation function operations, respectively, to obtain the output of the first feature encoder. The output of the second feature encoding The output of the third feature encoding The output of the fourth feature encoding ;
[0021] Output of the first feature encoding The output of the second feature encoding The output of the third feature encoding The output of the fourth feature encoding This refers to multi-scale features.
[0022] Furthermore, the specific process for obtaining the fused spatiotemporal features in step S3 is as follows:
[0023] Specifically, a residual module is added after the first feature encoding, the second feature encoding, the third feature encoding, and the fourth feature encoding;
[0024] The output of the first feature encoding The output of the second feature encoding The output of the third feature encoding The output of the fourth feature encoding The data are processed in the first residual module, the second residual module, the third residual module, and the fourth residual module, respectively, to obtain the outputs of the first residual module, the second residual module, the third residual module, and the fourth residual module, respectively.
[0025] The output of the first residual module is input into the first enhanced multi-scale spatiotemporal feature extraction module for processing to obtain the output of the first enhanced multi-scale spatiotemporal feature extraction module. The output of the first enhanced multi-scale spatiotemporal feature extraction module and the output of the first residual module are subjected to pulse feature channel-wise weighted fusion C to obtain the second fused feature.
[0026] The output of the second residual module is input into the second enhanced multi-scale spatiotemporal feature extraction module for processing to obtain the output of the second enhanced multi-scale spatiotemporal feature extraction module. The output of the second enhanced multi-scale spatiotemporal feature extraction module and the output of the second residual module are then fused together with pulse features channel-wise weighted C to obtain the third fused feature.
[0027] The output of the third residual module is input into the third enhanced multi-scale spatiotemporal feature extraction module for processing to obtain the output of the third enhanced multi-scale spatiotemporal feature extraction module. The output of the third enhanced multi-scale spatiotemporal feature extraction module and the output of the third residual module are then fused together with pulse features channel-wise weighted C to obtain the fourth fused feature.
[0028] The output of the fourth residual module is input into the fourth enhanced multi-scale spatiotemporal feature extraction module for processing to obtain the output of the fourth enhanced multi-scale spatiotemporal feature extraction module. The output of the fourth enhanced multi-scale spatiotemporal feature extraction module and the output of the fourth residual module are then fused together with pulse features channel-wise weighted C to obtain the fifth fused feature.
[0029] The fused spatiotemporal features are the second fused feature, the third fused feature, the fourth fused feature, and the fifth fused feature.
[0030] Furthermore, the specific process of obtaining the target region segmentation result in step S5 is as follows:
[0031] Will , , The input is processed by the spatiotemporal interaction aggregation module to obtain the output of the spatiotemporal interaction aggregation module. The output of the spatiotemporal interaction aggregation module and the fifth fusion feature are then input together into the first spatiotemporal interaction decoder for processing to obtain the output of the first spatiotemporal interaction decoder.
[0032] The output of the first spatiotemporal interactive decoder and the fourth fusion feature are input into the second spatiotemporal interactive decoder for processing to obtain the output of the second spatiotemporal interactive decoder.
[0033] The output of the second spatiotemporal interactive decoder and the third fused feature are input together into the third spatiotemporal interactive decoder for processing to obtain the output of the third spatiotemporal interactive decoder.
[0034] The output of the third spatiotemporal interactive decoder and the second fused feature are input together into the fourth spatiotemporal interactive decoder for processing, resulting in the output of the fourth spatiotemporal interactive decoder, which is the preliminary segmentation feature map. ;
[0035] Will and The data are then processed together using a multi-scale edge refinement module to obtain the target region segmentation result.
[0036] Furthermore, the specific process of obtaining the output of the spatiotemporal interaction aggregation module is as follows:
[0037] Will , , In the spatiotemporal interaction aggregation module, for , , Decoupling and feature interaction in the spatiotemporal dimensions are performed, and continuous correlation features in the temporal dimension and contextual correlation features in the spatial dimension are extracted respectively;
[0038] The continuous correlation features and contextual correlation features after interaction are aggregated and fused to obtain the fused original pulse spatiotemporal context features, which is the output of the spatiotemporal interaction aggregation module.
[0039] Furthermore, the specific process for obtaining the output of the first enhanced multi-scale spatiotemporal feature extraction module is as follows:
[0040] The first enhanced multi-scale spatiotemporal feature extraction module consists of a spatiotemporal multi-scale expansion and fusion module, a feature fusion module, a first 1x1 convolution, a second 1x1 convolution, a third 1x1 convolution, a fourth 1x1 convolution, a Sigmoid activation function, a fifth 1x1 convolution, and a first ReLU activation function.
[0041] Among them, the first enhanced multi-scale spatiotemporal feature extraction module, the second enhanced multi-scale spatiotemporal feature extraction module, the third enhanced multi-scale spatiotemporal feature extraction module and the fourth enhanced multi-scale spatiotemporal feature extraction module have the same structure;
[0042] The output of the first residual module is input into the spatiotemporal multi-scale expansion fusion module for processing to obtain the output of the spatiotemporal multi-scale expansion fusion module. The output of the spatiotemporal multi-scale expansion fusion module is input into the feature fusion module for processing to obtain the output of the feature fusion module. The output of the feature fusion module is input into the second 1x1 convolution, the third 1x1 convolution, and the fourth 1x1 convolution for processing to obtain the spatial feature matrix Q, the temporal feature matrix K, and the value matrix V, respectively.
[0043] The spatial feature matrix Q and the temporal feature matrix K are multiplied element by element. This yields the output of the first multiplication.
[0044] The output of the first multiplication is input into the Sigmoid activation function for normalization, generating 0. 1. Spatiotemporal attention weighting graph of interval 1;
[0045] The spatiotemporal attention weight map is multiplied element-wise with the value matrix V. This yields the output of the second multiplication.
[0046] The output of the feature fusion module is input into the first 1x1 convolution for processing to obtain the output of the first 1x1 convolution.
[0047] The output of the second multiplication is added element-wise to the output of the first 1x1 convolution. This yields the output of the second sum;
[0048] The output of the second summation is sequentially input into the fifth 1x1 convolution and the first ReLU activation function for processing to obtain the output of the first enhanced multi-scale spatiotemporal feature extraction module.
[0049] Furthermore, the specific process of obtaining the output of the spatiotemporal multi-scale extended fusion module is as follows:
[0050] The spatiotemporal multi-scale expansion and fusion module consists of a feature concatenation layer, a sixth 1x1 convolution, a seventh 1x1 convolution, an eighth 1x1 convolution, a ninth 1x1 convolution, a tenth 1x3 convolution, an eleventh 1x5 convolution, a twelfth 1x7 convolution, a thirteenth 3x1 convolution, a fourteenth 5x1 convolution, a fifteenth 7x1 convolution, a sixteenth 3x1 convolution, a seventeenth 3x3 convolution, an eighteenth 3x3 convolution, a nineteenth 1x1 convolution, a feature concatenation + 3x3 convolution, and a second ReLU activation function.
[0051] The output of the first residual module is input into the feature concatenation layer for processing to obtain the output of the feature concatenation layer;
[0052] The output of the feature splicing layer is input into the sixth, seventh, eighth, and ninth 1x1 convolutions for processing, respectively, to obtain the outputs of the sixth, seventh, eighth, ninth, and nineteenth 1x1 convolutions.
[0053] The output of the seventh 1x1 convolution is sequentially input into the eleventh 1x3 convolution, the thirteenth 3x1 convolution, and the sixteenth 3x1 convolution for processing, resulting in the output of the sixteenth 3x1 convolution;
[0054] The output of the eighth 1x1 convolution is sequentially input into the eleventh 1x5 convolution, the fourteenth 5x1 convolution, and the seventeenth 3x3 convolution for processing, to obtain the output of the seventeenth 3x3 convolution;
[0055] The output of the ninth 1x1 convolution is sequentially input into the twelfth 1x7 convolution, the fifteenth 7x1 convolution, and the eighteenth 3x3 convolution for processing, to obtain the output of the eighteenth 3x3 convolution;
[0056] The outputs of the sixth 1x1 convolution, the sixteenth 3x1 convolution, the seventeenth 3x3 convolution, and the eighteenth 3x3 convolution are all input into a feature concatenation + 3x3 convolution for processing, resulting in the feature concatenation + 3x3 convolution. This result is then added element-wise to the output of the nineteenth 1x1 convolution. The output of the third sum is obtained, and the output of the third sum is input into the second ReLU activation function for processing to obtain the output of the spatiotemporal multiscale expansion fusion module.
[0057] Furthermore, the specific process for obtaining the target region segmentation result is as follows:
[0058] The output of the fourth spatiotemporal interactive decoder is input into the twentieth 3x3 convolution for processing, resulting in the output of the twentieth 3x3 convolution. ;
[0059] The output of the 20th 3x3 convolution The input is processed sequentially through a first average pooling step and a twenty-first 3x3 convolution, yielding the output of the twenty-first 3x3 convolution. ;
[0060] The output of the twenty-first 3x3 convolution The input is processed sequentially through a second average pooling step and a twenty-second 5x5 convolution, yielding the output of the twenty-second 5x5 convolution. ;
[0061] The output of the 22nd 5x5 convolution The input is processed sequentially using the third average pooling and the twenty-third 7x7 convolution to obtain the output of the twenty-third 7x7 convolution. ;
[0062] Will The input is processed in the first edge enhancement module to obtain the output of the first edge enhancement module;
[0063] Will The input is processed in the second edge enhancement module to obtain the output of the second edge enhancement module;
[0064] Will The input is processed in the third edge enhancement module to obtain the output of the third edge enhancement module;
[0065] The outputs of the first edge enhancement module, the second edge enhancement module, and the third edge enhancement module are compared with... The pulse features are weighted and fused channel by channel to obtain the sixth fused feature (C). This sixth fused feature is then input into a spatial attention mechanism for processing to obtain a spatial attention weight map. Finally, the spatial attention weight map is multiplied element-wise with the sixth fused feature. The output of the third multiplication is obtained. The output of the third multiplication is then fed into the twenty-fourth 1x1 convolution and the twenty-fifth 3x3 convolution for processing, resulting in the output of the twenty-fifth 3x3 convolution.
[0066] The output of the 25th 3x3 convolution is added element-wise to the output of the 4th spatiotemporal interactive decoder. This yields the output of the fourth summation;
[0067] Preliminary segmentation feature map The input is processed in the 26th 1x1 convolution to obtain the output of the 26th 1x1 convolution;
[0068] The output of the 26th 1x1 convolution is combined with the output of the fourth convolution by performing pulse feature channel-wise weighted fusion C to obtain the 7th fusion feature. The 7th fusion feature is then sequentially input into the 27th, 28th, and 29th convolutions and the Sigmoid activation function for processing, and the target region segmentation result is output.
[0069] Furthermore, the specific process of obtaining the output of the first edge enhancement module is as follows:
[0070] The first edge enhancement module consists of the Sobel operator, the fourth average pooling, the max pooling, the thirtieth convolution, the thirty-first convolution, and the channel attention mechanism;
[0071] The first edge enhancement module, the second edge enhancement module, and the third edge enhancement module have the same structure.
[0072] Will The inputs are processed by the Sobel operator, the fourth average pooling, and the max pooling respectively, to obtain the outputs of the Sobel operator, the fourth average pooling, and the max pooling.
[0073] The output of the fourth average pooling is compared with Subtract element by element This yields the first subtraction output;
[0074] The output of max pooling and Subtract element by element This yields the second subtraction output;
[0075] The pulse features of the first subtraction output, the second subtraction output, and the Sobel operator output are weighted and fused channel by channel to obtain the eighth fused feature;
[0076] The eighth fusion feature is input into the thirtieth convolution for processing, yielding the output of the thirtieth convolution. The output of the thirtieth convolution is then input into the thirty-first convolution for processing, yielding the output of the thirty-first convolution. The output of the thirty-first convolution is then input into the channel attention mechanism for processing, yielding the output of the channel attention mechanism. Finally, the output of the channel attention mechanism is multiplied element-wise with the output of the thirty-first convolution. The output of the fourth multiplication is then multiplied by... Add element by element The output of the first edge enhancement module is obtained.
[0077] Compared with existing technologies, this invention has the following advantages: This invention discloses a pulse defocusing region detection method based on spatiotemporal information interaction. The method mainly includes the following parts: a spatiotemporal aggregated pulse feature extraction module Spike Embedding, a multi-layer feature encoder HF-Encoder, a spatiotemporal interactive decoder STI-Decoder, an enhanced multi-scale spatiotemporal feature extraction module (EMTSE), and a multi-scale edge refinement module (MSER). Traditional defocusing region detection research mainly focuses on digital image parts. The existing DBD method is not applicable to three-dimensional pulse sequences. However, this invention is for continuous pulse sequences and uses deep learning methods to design a defocusing region detection network, which can directly use continuous pulse sequences to obtain the mask information of the focused and defocused regions. Attached Figure Description
[0078] Figure 1 This is a structural diagram of the spatiotemporal aggregation pulse feature extraction module provided by the present invention.
[0079] Figure 2 This is a structural diagram of the defocused area detection network provided by the present invention.
[0080] Figure 3 The structural diagram of the enhanced multi-scale spatiotemporal feature extraction module provided by the present invention.
[0081] Figure 4 This is a structural diagram of the spatiotemporal multi-scale extended fusion module provided by the present invention.
[0082] Figure 5 This is a structural diagram of the multi-scale edge refinement module provided by the present invention.
[0083] Figure 6 This is a structural diagram of the edge enhancement module provided by the present invention. Detailed Implementation
[0084] like Figures 1-2 As shown, the present invention provides a technical solution: a pulse defocusing region detection method based on spatiotemporal information interaction, comprising:
[0085] Step S1: Input the continuous pulse sequence acquired by the pulse camera into the spatiotemporal aggregation pulse feature extraction module to extract spatiotemporal features. The spatiotemporal features include global pulse sequence features, first pulse subsequence features, second pulse subsequence features, and third pulse subsequence features.
[0086] Step S2: Construct a defocused region detection network, which includes a feature pyramid, an enhanced multi-scale spatiotemporal feature extraction module, a spatiotemporal interaction aggregation module, and a spatiotemporal interaction decoder; construct a feature pyramid, set the acquired spatiotemporal features as the bottom layer features of the feature pyramid, and extract multi-scale features from the spatiotemporal features;
[0087] Step S3: Input the extracted multi-scale features into the enhanced multi-scale spatiotemporal feature extraction module for fusion to obtain the fused spatiotemporal features;
[0088] Step S4: Input the first pulse subsequence features, the second pulse subsequence features, and the third pulse subsequence features into the spatiotemporal interaction aggregation module for processing to obtain the output of the spatiotemporal interaction aggregation module. Input the output of the spatiotemporal interaction aggregation module and the fused spatiotemporal features into the spatiotemporal interaction decoder for processing to generate a preliminary segmentation feature map.
[0089] Step S5: Input the global pulse sequence features and preliminary segmentation feature map from step S1 into the multi-scale edge refinement module to refine the edges and obtain the target region segmentation result.
[0090] Step S6: Set a loss function to optimize the performance of the defocused region detection network and reconstruct the target region segmentation result to approximate the real image.
[0091] like Figure 1 As shown, the specific process of extracting spatiotemporal features in step S1 is as follows:
[0092] Step S11, the continuous pulse sequence Along the time axis, the pulse camera's frame rate adaptively divides the data into three non-overlapping pulse subsequences. These three pulse subsequences are: the first pulse subsequence, the second pulse subsequence, and the third pulse subsequence. Second pulse subsequence The third pulse subsequence ;
[0093] Step S12: In order to obtain the spatial information of the overall scene, the continuous pulse sequence is processed. Global feature extraction (Global FE) is performed (a lightweight feature extractor for pulse data adds impulse noise suppression layers to all four residual modules), using the thirty-second 1×1 convolution on continuous pulse sequences. After dimensionality reduction, the global pulse sequence features are extracted using the thirty-third 3×3 convolution. ;
[0094] Step S13, the first pulse subsequence The first local feature extraction (Local FE) is processed to obtain the first processed feature. This first processed feature is then combined with the global pulse sequence feature. The pulse features are then fused channel by channel with weights C (the fusion weights are adaptively assigned by the channel variance of the global features; the larger the variance, the higher the weight), to obtain the features of the first pulse subsequence. ;
[0095] The second pulse subsequence The input is processed by the second local feature extraction to obtain the second processed feature. The second processed feature is then combined with the global pulse sequence feature. The pulse features are weighted and fused channel by channel (C) to obtain the second pulse subsequence features. ;
[0096] The third pulse subsequence The input is processed by third local feature extraction to obtain third processed features, which are then combined with global pulse sequence features. The pulse features are weighted and fused channel by channel (C) to obtain the features of the third pulse subsequence. The spatiotemporal characteristics are , , and .
[0097] like Figure 2 As shown, the specific process of extracting multi-scale features in step S2 is as follows:
[0098] Extracted from step S1 , , The features are set as the bottom layer of the feature pyramid, and a four-layer feature pyramid is constructed. The four-layer feature pyramid consists of a first feature code, a second feature code, a third feature code, and a fourth feature code. The first feature code has 32 feature channels, the second feature code has 64 feature channels, the third feature code has 128 feature channels, and the fourth feature code has 256 feature channels. A residual module is added after the first feature code, the second feature code, the third feature code, and the fourth feature code.
[0099] Will , , and The features are fused to obtain the first fused feature. This first fused feature is then sequentially input into the first feature encoder, the second feature encoder, the third feature encoder, and the fourth feature encoder for convolution and PReLU activation function operations, respectively, to obtain the output of the first feature encoder. The output of the second feature encoding The output of the third feature encoding The output of the fourth feature encoding ;
[0100] Output of the first feature encoding The output of the second feature encoding The output of the third feature encoding The output of the fourth feature encoding This refers to multi-scale features.
[0101] In this process, after each convolution within the first, second, third, and fourth feature encodings, a continuous pulse sequence-adapted PReLU activation function is added. The slope of the negative half-axis of the PReLU activation function is adaptively adjusted by the sparsity of the continuous pulse sequence. When the sparsity of the first fused feature is >0.8, the slope is 0.1; otherwise, it is 0.25, in order to adapt to the sparsity characteristics of the continuous pulse sequence and retain effective pulse features.
[0102] like Figure 2 As shown, the output of the first feature encoding The output of the second feature encoding The output of the third feature encoding The output of the fourth feature encoding The data are processed in the first residual module, the second residual module, the third residual module, and the fourth residual module, respectively, to obtain the outputs of the first residual module, the second residual module, the third residual module, and the fourth residual module.
[0103] like Figure 2 As shown, the specific process for obtaining the fused spatiotemporal features in step S3 is as follows:
[0104] The output of the first residual module is input into the first enhanced multi-scale spatiotemporal feature extraction module for processing to obtain the output of the first enhanced multi-scale spatiotemporal feature extraction module. The output of the first enhanced multi-scale spatiotemporal feature extraction module and the output of the first residual module are subjected to pulse feature channel-wise weighted fusion C to obtain the second fused feature.
[0105] The output of the second residual module is input into the second enhanced multi-scale spatiotemporal feature extraction module for processing to obtain the output of the second enhanced multi-scale spatiotemporal feature extraction module. The output of the second enhanced multi-scale spatiotemporal feature extraction module and the output of the second residual module are then fused together with pulse features channel-wise weighted C to obtain the third fused feature.
[0106] The output of the third residual module is input into the third enhanced multi-scale spatiotemporal feature extraction module for processing to obtain the output of the third enhanced multi-scale spatiotemporal feature extraction module. The output of the third enhanced multi-scale spatiotemporal feature extraction module and the output of the third residual module are then fused together with pulse features channel-wise weighted C to obtain the fourth fused feature.
[0107] The output of the fourth residual module is input into the fourth enhanced multi-scale spatiotemporal feature extraction module for processing to obtain the output of the fourth enhanced multi-scale spatiotemporal feature extraction module. The output of the fourth enhanced multi-scale spatiotemporal feature extraction module and the output of the fourth residual module are then fused together with pulse features channel-wise weighted C to obtain the fifth fused feature.
[0108] The fused spatiotemporal features are the second fused feature, the third fused feature, the fourth fused feature, and the fifth fused feature.
[0109] like Figure 2 As shown, the specific process of generating the preliminary segmentation feature map in step S4 is as follows:
[0110] Will , , The input to the spatiotemporal interaction aggregation module SIA is processed to obtain the output of the spatiotemporal interaction aggregation module. The output of the spatiotemporal interaction aggregation module and the fifth fusion feature are then input together to the first spatiotemporal interaction decoder for processing to obtain the output of the first spatiotemporal interaction decoder.
[0111] The output of the first spatiotemporal interactive decoder and the fourth fusion feature are input into the second spatiotemporal interactive decoder for processing to obtain the output of the second spatiotemporal interactive decoder.
[0112] The output of the second spatiotemporal interactive decoder and the third fused feature are input together into the third spatiotemporal interactive decoder for processing to obtain the output of the third spatiotemporal interactive decoder.
[0113] The output of the third spatiotemporal interactive decoder and the second fused feature are input together into the fourth spatiotemporal interactive decoder for processing, resulting in the output of the fourth spatiotemporal interactive decoder, which is the preliminary segmentation feature map. .
[0114] The specific process of obtaining the output of the spatiotemporal interaction aggregation module is as follows:
[0115] Will , , In the spatiotemporal interaction aggregation module, for , , Decoupling and feature interaction in the spatiotemporal dimensions are performed, and continuous correlation features in the temporal dimension and contextual correlation features in the spatial dimension are extracted respectively;
[0116] The continuous correlation features and contextual correlation features after interaction are aggregated and fused to obtain the fused original pulse spatiotemporal context features, which is the output of the spatiotemporal interaction aggregation module.
[0117] like Figure 2 As shown, the specific process for obtaining the target region segmentation result in step S5 is as follows:
[0118] The output of the fourth spatiotemporal interactive decoder and The data are then processed together using a multi-scale edge refinement module to obtain the target region segmentation result.
[0119] like Figure 3 As shown, the specific process for obtaining the output of the first enhanced multi-scale spatiotemporal feature extraction module is as follows:
[0120] The first enhanced multi-scale spatiotemporal feature extraction module consists of a spatiotemporal multi-scale expansion and fusion module, a feature fusion module, a first 1x1 convolution, a second 1x1 convolution, a third 1x1 convolution, a fourth 1x1 convolution, a Sigmoid activation function, a fifth 1x1 convolution, and a first ReLU activation function.
[0121] The first enhanced multi-scale spatiotemporal feature extraction module, the second enhanced multi-scale spatiotemporal feature extraction module, the third enhanced multi-scale spatiotemporal feature extraction module, and the fourth enhanced multi-scale spatiotemporal feature extraction module have the same structure.
[0122] The output of the first residual module is input into the spatiotemporal multi-scale expansion fusion module for processing to obtain the output of the spatiotemporal multi-scale expansion fusion module. The output of the spatiotemporal multi-scale expansion fusion module is input into the feature fusion module for processing to obtain the output of the feature fusion module. The output of the feature fusion module is input into the second 1x1 convolution, the third 1x1 convolution, and the fourth 1x1 convolution for processing to obtain the spatial feature matrix Q, the temporal feature matrix K, and the value matrix V, respectively.
[0123] The spatial feature matrix Q and the temporal feature matrix K are multiplied element by element. This yields the output of the first multiplication.
[0124] The output of the first multiplication is input into the Sigmoid activation function for normalization, generating 0. 1. Spatiotemporal attention weighting graph of interval 1;
[0125] The spatiotemporal attention weight map is multiplied element-wise with the value matrix V. This yields the output of the second multiplication.
[0126] The output of the feature fusion module is input into the first 1x1 convolution for processing to obtain the output of the first 1x1 convolution.
[0127] The output of the second multiplication is added element-wise to the output of the first 1x1 convolution. This yields the output of the second sum;
[0128] The output of the second summation is sequentially input into the fifth 1x1 convolution and the first ReLU activation function for processing to obtain the output of the first enhanced multi-scale spatiotemporal feature extraction module.
[0129] The output calculation formula of the first enhanced multi-scale spatiotemporal feature extraction module is as follows:
[0130] ;
[0131] In the formula, This is the output of the first enhanced multi-scale spatiotemporal feature extraction module; This is the output of the spatiotemporal multi-scale extended fusion module; Use the Sigmoid activation function; It is the product of the feature dimension d and the sparsity s of the pulse sequence; This is a transpose.
[0132] like Figure 4 As shown, the specific process for obtaining the output of the spatiotemporal multi-scale extended fusion module is as follows:
[0133] The spatiotemporal multi-scale expansion and fusion module consists of a feature concatenation layer, a sixth 1x1 convolution, a seventh 1x1 convolution, an eighth 1x1 convolution, a ninth 1x1 convolution, a tenth 1x3 convolution, an eleventh 1x5 convolution, a twelfth 1x7 convolution, a thirteenth 3x1 convolution, a fourteenth 5x1 convolution, a fifteenth 7x1 convolution, a sixteenth 3x1 convolution, a seventeenth 3x3 convolution, an eighteenth 3x3 convolution, a nineteenth 1x1 convolution, a feature concatenation + 3x3 convolution, and a second ReLU activation function.
[0134] The output of the first residual module is input into the feature concatenation layer for processing to obtain the output of the feature concatenation layer;
[0135] The output of the feature splicing layer is input into the sixth, seventh, eighth, and ninth 1x1 convolutions for processing, respectively, to obtain the outputs of the sixth, seventh, eighth, ninth, and nineteenth 1x1 convolutions.
[0136] The output of the seventh 1x1 convolution is sequentially input into the eleventh 1x3 convolution, the thirteenth 3x1 convolution, and the sixteenth 3x1 convolution for processing, resulting in the output of the sixteenth 3x1 convolution;
[0137] The output of the eighth 1x1 convolution is sequentially input into the eleventh 1x5 convolution, the fourteenth 5x1 convolution, and the seventeenth 3x3 convolution for processing, to obtain the output of the seventeenth 3x3 convolution;
[0138] The output of the ninth 1x1 convolution is sequentially input into the twelfth 1x7 convolution, the fifteenth 7x1 convolution, and the eighteenth 3x3 convolution for processing, to obtain the output of the eighteenth 3x3 convolution;
[0139] The outputs of the sixth 1x1 convolution, the sixteenth 3x1 convolution, the seventeenth 3x3 convolution, and the eighteenth 3x3 convolution are all input into a feature concatenation + 3x3 convolution for processing, resulting in the feature concatenation + 3x3 convolution. This result is then added element-wise to the output of the nineteenth 1x1 convolution. The output of the third sum is obtained, and the output of the third sum is input into the second ReLU activation function for processing to obtain the output of the spatiotemporal multiscale expansion fusion module.
[0140] like Figure 5 As shown, the specific process for obtaining the target region segmentation result is as follows:
[0141] The output of the fourth spatiotemporal interactive decoder is input into the twentieth 3x3 convolution for processing, resulting in the output of the twentieth 3x3 convolution. ;
[0142] The output of the 20th 3x3 convolution The input is processed sequentially through a first average pooling step and a twenty-first 3x3 convolution, yielding the output of the twenty-first 3x3 convolution. ;
[0143] The output of the twenty-first 3x3 convolution The input is processed sequentially through a second average pooling step and a twenty-second 5x5 convolution, yielding the output of the twenty-second 5x5 convolution. ;
[0144] The output of the 22nd 5x5 convolution The input is processed sequentially using the third average pooling and the twenty-third 7x7 convolution to obtain the output of the twenty-third 7x7 convolution. ;
[0145] Will The input is processed in the first edge enhancement module to obtain the output of the first edge enhancement module;
[0146] Will The input is processed in the second edge enhancement module to obtain the output of the second edge enhancement module;
[0147] Will The input is processed in the third edge enhancement module to obtain the output of the third edge enhancement module;
[0148] The outputs of the first edge enhancement module, the second edge enhancement module, and the third edge enhancement module are compared with... The pulse features are weighted and fused channel by channel to obtain the sixth fused feature (C). This sixth fused feature is then input into a spatial attention mechanism for processing to obtain a spatial attention weight map. Finally, the spatial attention weight map is multiplied element-wise with the sixth fused feature. The output of the third multiplication is obtained. The output of the third multiplication is then fed into the twenty-fourth 1x1 convolution and the twenty-fifth 3x3 convolution for processing, resulting in the output of the twenty-fifth 3x3 convolution.
[0149] The output of the 25th 3x3 convolution is added element-wise to the output of the 4th spatiotemporal interactive decoder. This yields the output of the fourth summation;
[0150] Preliminary segmentation feature map The input is processed in the 26th 1x1 convolution to obtain the output of the 26th 1x1 convolution;
[0151] The output of the 26th 1x1 convolution is combined with the output of the fourth convolution by performing pulse feature channel-wise weighted fusion C to obtain the 7th fusion feature. The 7th fusion feature is then sequentially input into the 27th, 28th, and 29th convolutions and the Sigmoid activation function for processing, and the target region segmentation result is output.
[0152] like Figure 6 As shown, the specific process for obtaining the output of the first edge enhancement module is as follows:
[0153] The first edge enhancement module consists of the Sobel operator, the fourth average pooling, the max pooling, the thirtieth convolution, the thirty-first convolution, and the channel attention mechanism;
[0154] The first edge enhancement module, the second edge enhancement module, and the third edge enhancement module have the same structure.
[0155] Will The inputs are processed by the Sobel operator, the fourth average pooling, and the max pooling respectively, to obtain the outputs of the Sobel operator, the fourth average pooling, and the max pooling.
[0156] The output of the fourth average pooling is compared with Subtract element by element This yields the first subtraction output;
[0157] The output of max pooling and Subtract element by element This yields the second subtraction output;
[0158] The pulse features of the first subtraction output, the second subtraction output, and the Sobel operator output are weighted and fused channel by channel to obtain the eighth fused feature;
[0159] The eighth fusion feature is input into the thirtieth convolution for processing, yielding the output of the thirtieth convolution. The output of the thirtieth convolution is then input into the thirty-first convolution for processing, yielding the output of the thirty-first convolution. The output of the thirty-first convolution is then input into the channel attention mechanism for processing, yielding the output of the channel attention mechanism. Finally, the output of the channel attention mechanism is multiplied element-wise with the output of the thirty-first convolution. The output of the fourth multiplication is then multiplied by... Add element by element The output of the first edge enhancement module is obtained.
[0160] The specific process of step S6 is as follows:
[0161] The total loss function is a combination of the Binary Cross Entropy (BCE) loss function, the Structural Similarity (SSIM) loss function, and the Intersection-over-Union (IoU) loss function, expressed as:
[0162] ;
[0163] In the formula, Represents the total loss function; The binary cross-entropy loss function; The loss function is the pulse spatiotemporal structural similarity function. The intersection-union ratio loss function; α, β, λ They are respectively , , The weighting coefficients.
[0164] Among them, The binary cross-entropy loss function incorporates pulse mask pixel noise to measure the difference between the network's predicted probability and the true label in defocused region detection, defined as:
[0165] ;
[0166] In the formula, This represents the summation operation; These are the pixel values corresponding to the actual image; It is the natural logarithm; The mask pixel values are the target region segmentation results predicted by the defocused region detection network.
[0167] The pulse spatiotemporal structural similarity loss function extends the traditional spatial SSIM to a two-dimensional spatiotemporal SSIM, enabling simultaneous computation of structural similarity in both temporal and spatial dimensions. Assume there exists a real image and the target region segmentation result as follows: and The SSIM loss function is expressed as follows:
[0168] ;
[0169] In the formula, For real images Segmentation results with target region Structural similarity index; , Representing real images respectively Segmentation results with target region The local mean; and Representing real images respectively Segmentation results with target region Local variance; To stabilize the mean-related terms in the SSIM formula and avoid constants with a denominator of 0; To stabilize the variance correlation term in the SSIM formula and avoid constants with a denominator of 0.
[0170] and ,when The larger the value, the more realistic the image. Segmentation results with target region The higher the structural similarity between them, the better. (This refers to real images.) Target region segmentation results The corresponding target region segmentation result is output by the defocused region detection network. With real images The true label.
[0171] The SSIM metric loss is defined as follows:
[0172] ;
[0173] In the formula, The network infers the target region segmentation result for the defocused region detection network; This is a real image.
[0174] Defined as:
[0175] .
[0176] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A pulse defocusing region detection method based on spatiotemporal information interaction, characterized in that, include: Step S1: Input the continuous pulse sequence acquired by the pulse camera into the spatiotemporal aggregation pulse feature extraction module to extract spatiotemporal features. The spatiotemporal features include global pulse sequence features, first pulse subsequence features, second pulse subsequence features, and third pulse subsequence features. Step S2: Construct a defocused region detection network, which includes a feature pyramid, an enhanced multi-scale spatiotemporal feature extraction module, a spatiotemporal interaction aggregation module, and a spatiotemporal interaction decoder; construct a feature pyramid, set the acquired spatiotemporal features as the bottom layer features of the feature pyramid, and extract multi-scale features from the spatiotemporal features; Step S3: Input the extracted multi-scale features into the enhanced multi-scale spatiotemporal feature extraction module for fusion to obtain the fused spatiotemporal features; Step S4: Input the first pulse subsequence features, the second pulse subsequence features, and the third pulse subsequence features into the spatiotemporal interaction aggregation module for processing to obtain the output of the spatiotemporal interaction aggregation module. Input the output of the spatiotemporal interaction aggregation module and the fused spatiotemporal features into the spatiotemporal interaction decoder for processing to generate a preliminary segmentation feature map. Step S5: Input the global pulse sequence features and preliminary segmentation feature map from step S1 into the multi-scale edge refinement module to refine the edges and obtain the target region segmentation result; Step S6: Set a loss function to optimize the performance of the defocused region detection network, and reconstruct and evaluate how closely the target region segmentation results approximate the real image.
2. The pulse defocusing region detection method based on spatiotemporal information interaction according to claim 1, characterized in that: The specific process of extracting spatiotemporal features in step S1 is as follows: Step S11: Continuous pulse sequence Along the time axis, the pulse camera's frame rate adaptively divides the data into three non-overlapping pulse subsequences. These three pulse subsequences are: the first pulse subsequence, the second pulse subsequence, and the third pulse subsequence. Second pulse subsequence The third pulse subsequence ; Step S12: For the continuous pulse sequence Global feature extraction was performed using the thirty-second 1×1 convolution on the continuous pulse sequence. After dimensionality reduction, the 33rd 3×3 convolution is used for extraction to obtain global pulse sequence features. ; Step S13: The first pulse subsequence The first local feature extraction is processed to obtain the first processed feature, which is then combined with the global pulse sequence feature. The pulse features are weighted and fused channel by channel (C) to obtain the features of the first pulse subsequence. ; The second pulse subsequence The input is processed by the second local feature extraction to obtain the second processed feature. The second processed feature is then combined with the global pulse sequence feature. The pulse features are weighted and fused channel by channel (C) to obtain the second pulse subsequence features. ; The third pulse subsequence The input is processed by third local feature extraction to obtain third processed features, which are then combined with global pulse sequence features. The pulse features are weighted and fused channel by channel (C) to obtain the features of the third pulse subsequence. The spatiotemporal characteristics are , , and .
3. The pulse defocusing region detection method based on spatiotemporal information interaction according to claim 2, characterized in that: The specific process of extracting multi-scale features in step S2 is as follows: Extracted from step S1 , , The features are set as the bottom layer of the feature pyramid, and a four-layer feature pyramid is constructed. The four-layer feature pyramid consists of a first feature code, a second feature code, a third feature code, and a fourth feature code. The first feature code has 32 feature channels, the second feature code has 64 feature channels, the third feature code has 128 feature channels, and the fourth feature code has 256 feature channels. Will , , and The features are fused to obtain the first fused feature. This first fused feature is then sequentially input into the first feature encoder, the second feature encoder, the third feature encoder, and the fourth feature encoder for convolution and PReLU activation function operations, respectively, to obtain the output of the first feature encoder. The output of the second feature encoding The output of the third feature encoding The output of the fourth feature encoding ; Output of the first feature encoding The output of the second feature encoding The output of the third feature encoding The output of the fourth feature encoding This refers to multi-scale features.
4. The pulse defocusing region detection method based on spatiotemporal information interaction according to claim 3, characterized in that: The specific process for obtaining the fused spatiotemporal features in step S3 is as follows: Specifically, a residual module is added after the first feature encoding, the second feature encoding, the third feature encoding, and the fourth feature encoding; The output of the first feature encoding The output of the second feature encoding The output of the third feature encoding The output of the fourth feature encoding The data are processed in the first residual module, the second residual module, the third residual module, and the fourth residual module, respectively, to obtain the outputs of the first residual module, the second residual module, the third residual module, and the fourth residual module, respectively. The output of the first residual module is input into the first enhanced multi-scale spatiotemporal feature extraction module for processing to obtain the output of the first enhanced multi-scale spatiotemporal feature extraction module. The output of the first enhanced multi-scale spatiotemporal feature extraction module and the output of the first residual module are subjected to pulse feature channel-wise weighted fusion C to obtain the second fused feature. The output of the second residual module is input into the second enhanced multi-scale spatiotemporal feature extraction module for processing to obtain the output of the second enhanced multi-scale spatiotemporal feature extraction module. The output of the second enhanced multi-scale spatiotemporal feature extraction module and the output of the second residual module are then fused together with pulse features channel-wise weighted C to obtain the third fused feature. The output of the third residual module is input into the third enhanced multi-scale spatiotemporal feature extraction module for processing to obtain the output of the third enhanced multi-scale spatiotemporal feature extraction module. The output of the third enhanced multi-scale spatiotemporal feature extraction module and the output of the third residual module are then fused together with pulse features channel-wise weighted C to obtain the fourth fused feature. The output of the fourth residual module is input into the fourth enhanced multi-scale spatiotemporal feature extraction module for processing to obtain the output of the fourth enhanced multi-scale spatiotemporal feature extraction module. The output of the fourth enhanced multi-scale spatiotemporal feature extraction module and the output of the fourth residual module are then fused together with pulse features channel-wise weighted C to obtain the fifth fused feature. The fused spatiotemporal features are the second fused feature, the third fused feature, the fourth fused feature, and the fifth fused feature.
5. The pulse defocusing region detection method based on spatiotemporal information interaction according to claim 4, characterized in that: The specific process of obtaining the target region segmentation result in step S5 is as follows: Will , , The input is processed by the spatiotemporal interaction aggregation module to obtain the output of the spatiotemporal interaction aggregation module. The output of the spatiotemporal interaction aggregation module and the fifth fusion feature are then input together into the first spatiotemporal interaction decoder for processing to obtain the output of the first spatiotemporal interaction decoder. The output of the first spatiotemporal interactive decoder and the fourth fusion feature are input into the second spatiotemporal interactive decoder for processing to obtain the output of the second spatiotemporal interactive decoder. The output of the second spatiotemporal interactive decoder and the third fused feature are input together into the third spatiotemporal interactive decoder for processing to obtain the output of the third spatiotemporal interactive decoder. The output of the third spatiotemporal interactive decoder and the second fused feature are input together into the fourth spatiotemporal interactive decoder for processing, resulting in the output of the fourth spatiotemporal interactive decoder, which is the preliminary segmentation feature map. ; Will and The data are then processed together using a multi-scale edge refinement module to obtain the target region segmentation result.
6. The pulse defocusing region detection method based on spatiotemporal information interaction according to claim 5, characterized in that: The specific process of obtaining the output of the spatiotemporal interaction aggregation module is as follows: Will , , In the spatiotemporal interaction aggregation module, for , , Decoupling and feature interaction in the spatiotemporal dimensions are performed, and continuous correlation features in the temporal dimension and contextual correlation features in the spatial dimension are extracted respectively; The continuous correlation features and contextual correlation features after interaction are aggregated and fused to obtain the fused original pulse spatiotemporal context features, which is the output of the spatiotemporal interaction aggregation module.
7. The pulse defocusing region detection method based on spatiotemporal information interaction according to claim 6, characterized in that: The specific process for obtaining the output of the first enhanced multi-scale spatiotemporal feature extraction module is as follows: The first enhanced multi-scale spatiotemporal feature extraction module consists of a spatiotemporal multi-scale expansion and fusion module, a feature fusion module, a first 1x1 convolution, a second 1x1 convolution, a third 1x1 convolution, a fourth 1x1 convolution, a Sigmoid activation function, a fifth 1x1 convolution, and a first ReLU activation function. Among them, the first enhanced multi-scale spatiotemporal feature extraction module, the second enhanced multi-scale spatiotemporal feature extraction module, the third enhanced multi-scale spatiotemporal feature extraction module and the fourth enhanced multi-scale spatiotemporal feature extraction module have the same structure; The output of the first residual module is input into the spatiotemporal multi-scale expansion fusion module for processing to obtain the output of the spatiotemporal multi-scale expansion fusion module. The output of the spatiotemporal multi-scale expansion fusion module is input into the feature fusion module for processing to obtain the output of the feature fusion module. The output of the feature fusion module is input into the second 1x1 convolution, the third 1x1 convolution, and the fourth 1x1 convolution for processing to obtain the spatial feature matrix Q, the temporal feature matrix K, and the value matrix V, respectively. The spatial feature matrix Q and the temporal feature matrix K are multiplied element by element. This yields the output of the first multiplication. The output of the first multiplication is input into the Sigmoid activation function for normalization, generating 0.
1. Spatiotemporal attention weighting graph of interval 1; The spatiotemporal attention weight map is multiplied element-wise with the value matrix V. This yields the output of the second multiplication. The output of the feature fusion module is input into the first 1x1 convolution for processing to obtain the output of the first 1x1 convolution. The output of the second multiplication is added element-wise to the output of the first 1x1 convolution. This yields the output of the second sum; The output of the second summation is sequentially input into the fifth 1x1 convolution and the first ReLU activation function for processing to obtain the output of the first enhanced multi-scale spatiotemporal feature extraction module.
8. The pulse defocusing region detection method based on spatiotemporal information interaction according to claim 7, characterized in that: The specific process of obtaining the output of the spatiotemporal multi-scale extended fusion module is as follows: The spatiotemporal multi-scale expansion and fusion module consists of a feature concatenation layer, a sixth 1x1 convolution, a seventh 1x1 convolution, an eighth 1x1 convolution, a ninth 1x1 convolution, a tenth 1x3 convolution, an eleventh 1x5 convolution, a twelfth 1x7 convolution, a thirteenth 3x1 convolution, a fourteenth 5x1 convolution, a fifteenth 7x1 convolution, a sixteenth 3x1 convolution, a seventeenth 3x3 convolution, an eighteenth 3x3 convolution, a nineteenth 1x1 convolution, a feature concatenation + 3x3 convolution, and a second ReLU activation function. The output of the first residual module is input into the feature concatenation layer for processing to obtain the output of the feature concatenation layer; The output of the feature splicing layer is input into the sixth, seventh, eighth, and ninth 1x1 convolutions for processing, respectively, to obtain the outputs of the sixth, seventh, eighth, ninth, and nineteenth 1x1 convolutions. The output of the seventh 1x1 convolution is sequentially input into the eleventh 1x3 convolution, the thirteenth 3x1 convolution, and the sixteenth 3x1 convolution for processing, resulting in the output of the sixteenth 3x1 convolution; The output of the eighth 1x1 convolution is sequentially input into the eleventh 1x5 convolution, the fourteenth 5x1 convolution, and the seventeenth 3x3 convolution for processing, to obtain the output of the seventeenth 3x3 convolution; The output of the ninth 1x1 convolution is sequentially input into the twelfth 1x7 convolution, the fifteenth 7x1 convolution, and the eighteenth 3x3 convolution for processing, to obtain the output of the eighteenth 3x3 convolution; The outputs of the sixth 1x1 convolution, the sixteenth 3x1 convolution, the seventeenth 3x3 convolution, and the eighteenth 3x3 convolution are all input into a feature concatenation + 3x3 convolution for processing, resulting in the feature concatenation + 3x3 convolution. This result is then added element-wise to the output of the nineteenth 1x1 convolution. The output of the third sum is obtained, and the output of the third sum is input into the second ReLU activation function for processing to obtain the output of the spatiotemporal multiscale expansion fusion module.
9. The pulse defocusing region detection method based on spatiotemporal information interaction according to claim 8, characterized in that: The specific process for obtaining the target region segmentation result is as follows: The output of the fourth spatiotemporal interactive decoder is input into the twentieth 3x3 convolution for processing, resulting in the output of the twentieth 3x3 convolution. ; The output of the 20th 3x3 convolution The input is processed sequentially through a first average pooling step and a twenty-first 3x3 convolution, yielding the output of the twenty-first 3x3 convolution. ; The output of the twenty-first 3x3 convolution The input is processed sequentially through a second average pooling step and a twenty-second 5x5 convolution, yielding the output of the twenty-second 5x5 convolution. ; The output of the 22nd 5x5 convolution The input is processed sequentially using the third average pooling and the twenty-third 7x7 convolution to obtain the output of the twenty-third 7x7 convolution. ; Will The input is processed in the first edge enhancement module to obtain the output of the first edge enhancement module; Will The input is processed in the second edge enhancement module to obtain the output of the second edge enhancement module; Will The input is processed in the third edge enhancement module to obtain the output of the third edge enhancement module; The outputs of the first edge enhancement module, the second edge enhancement module, and the third edge enhancement module are compared with... The pulse features are weighted and fused channel by channel to obtain the sixth fused feature; The sixth fusion feature is input into the spatial attention mechanism for processing to obtain the spatial attention weight map; the spatial attention weight map is then multiplied element-wise with the sixth fusion feature. The output of the third multiplication is obtained. The output of the third multiplication is then fed into the twenty-fourth 1x1 convolution and the twenty-fifth 3x3 convolution for processing, resulting in the output of the twenty-fifth 3x3 convolution. The output of the 25th 3x3 convolution is added element-wise to the output of the 4th spatiotemporal interactive decoder. This yields the output of the fourth summation; Preliminary segmentation feature map The input is processed in the 26th 1x1 convolution to obtain the output of the 26th 1x1 convolution; The output of the 26th 1x1 convolution is combined with the output of the fourth convolution by performing pulse feature channel-wise weighted fusion C to obtain the 7th fusion feature. The 7th fusion feature is then sequentially input into the 27th, 28th, and 29th convolutions and the Sigmoid activation function for processing, and the target region segmentation result is output.
10. The pulse defocusing region detection method based on spatiotemporal information interaction according to claim 9, characterized in that: The specific process of obtaining the output of the first edge enhancement module is as follows: The first edge enhancement module consists of the Sobel operator, the fourth average pooling, the max pooling, the thirtieth convolution, the thirty-first convolution, and the channel attention mechanism; The first edge enhancement module, the second edge enhancement module, and the third edge enhancement module have the same structure. Will The inputs are processed by the Sobel operator, the fourth average pooling, and the max pooling respectively, to obtain the outputs of the Sobel operator, the fourth average pooling, and the max pooling. The output of the fourth average pooling is compared with Subtract element by element This yields the first subtraction output; The output of max pooling and Subtract element by element This yields the second subtraction output; The pulse features of the first subtraction output, the second subtraction output, and the Sobel operator output are weighted and fused channel by channel to obtain the eighth fused feature; The eighth fusion feature is input into the thirtieth convolution for processing, yielding the output of the thirtieth convolution. The output of the thirtieth convolution is then input into the thirty-first convolution for processing, yielding the output of the thirty-first convolution. The output of the thirty-first convolution is then input into the channel attention mechanism for processing, yielding the output of the channel attention mechanism. Finally, the output of the channel attention mechanism is multiplied element-wise with the output of the thirty-first convolution. The output of the fourth multiplication is then multiplied by... Add element by element The output of the first edge enhancement module is obtained.
Citation Information
Patent Citations
Novel real-time measuring method and device of light intensity distribution in focal depth area of light beam
CN101614585A
Unmanned aerial vehicle target detection method based on spiking neural network
CN119559527A