Newborn limb detection method based on variable kernel convolution and local and global feature fusion
By using variable nuclear convolution and local global feature fusion methods in neonatal limb detection, the sampling position and convolutional kernel shape are dynamically adjusted, and local and global features are fused, the problem of low limb detection accuracy in neonatal limb detection is solved, and a higher accuracy limb motion detection is achieved.
Patent Information
- Application Number
- CN202510154632.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-06-03
AI Technical Summary
The accuracy of existing neonatal limb detection is low, making it difficult to accurately locate and identify the neonatal limb movements, affecting the accuracy of pain assessment.
The neonatal limb detection method based on variable kernel convolution and local global feature fusion is adopted. Through the feature extraction module, feature fusion module and dynamic detection head module, the sampling position of the input feature map and the shape of the convolution kernel are dynamically adjusted, and local features and global features are fused.
It improves the accuracy of neonatal limb detection, can capture local details and global movement dynamics of neonatal limbs more accurately, improves the detection ability of subtle movements in complex scenarios, and ensures the integrity of feature expression.
Smart Images

Figure CN120088856A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a neonatal limb detection method based on variable kernel convolution and local-global feature fusion, belonging to the technical field of limb detection. Background Technique
[0002] Neonatal pain estimation is an important part of neonatal medical care. Accurately assessing whether a neonate is in pain is crucial for timely taking intervention measures and reducing suffering. In traditional clinical practice, pain assessment mainly relies on observing the facial expressions of neonates, such as frowning, crying, and eyes tightly closed. Although these facial expressions can reflect the degree of pain in some cases, single expression analysis is often insufficient to comprehensively and accurately evaluate the pain perception of neonates. Especially when the facial expressions are weak or interfered by other factors, the accuracy of the assessment will be affected.
[0003] To solve this problem, pain assessment auxiliary means based on neonatal limb movements have been introduced. Research shows that when infants perceive pain, they will not only show it through facial expressions but also be accompanied by obvious limb movement changes, such as irregular limb swings, leg kicks, and finger clenches. Therefore, as an auxiliary means for neonatal pain assessment, limb movement recognition can effectively complement facial expression analysis and improve the reliability and accuracy of the overall assessment. However, the current accuracy of neonatal limb detection is relatively low.
[0004] The core purpose of neonatal limb detection is to locate the limbs of neonates and provide basic data for further action recognition. By detecting the movement trajectories and posture changes of the limbs, medical staff can more comprehensively understand the pain response patterns of neonates. Summary of the Invention
[0005] The purpose of the present invention is to provide a neonatal limb detection method based on variable kernel convolution and local-global feature fusion to solve the problem that the accuracy of neonatal limb detection in the existing technology needs to be improved.
[0006] The technical solution of the present invention is as follows:
[0007] A neonatal limb detection method based on variable kernel convolution and local-global feature fusion includes the following steps:
[0008] S1. Collect neonatal videos in different states, intercept key frame images from the neonatal videos, and label the limbs and head of the neonates in the key frame images to construct a neonatal limb detection image set;
[0009] S2. Build a neonatal limb detection model based on deformable kernel convolution and local-global feature fusion. The neonatal limb detection model based on deformable kernel convolution and local-global feature fusion includes a feature extraction module, a feature fusion module, and a dynamic detection head module. Preprocess the input image to generate a tensor F 0 , and the feature extraction module extracts the feature tensors of neonatal limbs from the tensor F 0 and outputs them to the feature fusion module. The feature fusion module includes a Path Aggregation Network (PANet) and a Local-Global Feature Fusion Module (CAFM). The PANet generates new feature tensors F' i , F i+1 , F i+2 by aggregating the input feature tensors F i , F i+1 , F i+2 of the i-th to (i + 2)-th neonatal limbs; the CAFM concatenates the feature tensors F i , F i+1 , F i+2 of neonatal limbs and the new feature tensors F' i , F' i+1 , F' i+2 along the channel dimension respectively to generate fused feature tensors F fusion,i , F fusion,i+1 , F fusion,i+2 . The fused feature tensors F fusion,i , F fusion,i+1 , F fusion,i+2 are successively subjected to convolution operations and channel shuffling to extract local features, and the fused feature tensors F fusion,i , F fusion,i+1 , F fusion,i+2 are subjected to convolution operations and attention mechanisms to extract global features, and then the local features and global features are fused to generate output feature tensors F CAFM,i , F CAFM,i+1 , F CAFM,i+2 ; the dynamic detection head module receives the output feature tensors F CAFM,i , F CAFM,i+1 , F CAFM,i+2 , aligns the features with the feature tensor F CAFM,i+1 as a reference, and then concatenates the three aligned features along the channel dimension to generate a new feature vector F input ; then the new feature vector F input is processed by scale-aware attention, space-aware attention, and task-aware attention to generate the final feature tensor F out for target detection tasks such as object classification, center regression, and bounding box regression. After post-processing the final feature tensor F out , the detected image is obtained;
[0010] S3. After training the neonatal limb detection model based on variable kernel convolution and local-global feature fusion using the neonatal limb detection image set, the trained neonatal limb detection model based on variable kernel convolution and local-global feature fusion is obtained;
[0011] S4. Input the neonatal limb detection video to be tested into the trained neonatal limb detection model based on variable kernel convolution and local-global feature fusion for limb detection.
[0012] Furthermore, in step S2, the feature extraction module includes M groups of variable kernel convolution layers AKConv, a cross-stage partial bottleneck layer with two convolutions, namely the C2f layer, and a fast-spatial pyramid pooling layer, namely the SPPF layer. The tensor F 0 passes through M groups of variable kernel convolution layers AKConv and the C2f layer in sequence and then is input into the SPPF layer. In the i-th group of variable kernel convolution layers AKConv and the C2f layer, the variable kernel convolution layer AKConv generates a feature tensor fi from the input feature tensor F i-1 The C2f layer is used to transform the feature tensor fi into a higher-quality feature tensor F i and output it; the SPPF layer outputs the M-th layer feature tensor F M .
[0013] Furthermore, the variable kernel convolution layer AKConv generates a feature tensor fi from the input feature tensor F i-1 Specifically,
[0014] First, the feature tensor F i-1 generates an offset ΔP through a two-dimensional convolution operation Conv2d n ; according to the generated offset ΔP n , the initial regular sampling coordinates P 0 are adjusted to obtain the modified sampling coordinates P' n : P' n =P 0 +ΔP n ;
[0015] Using the modified sampling coordinates P' n , the feature tensor F i is resampled, and the feature value corresponding to the sampling position is calculated through bilinear interpolation to obtain the resampled feature tensor F sample,i ;
[0016] The dimension of the resampled feature tensor F sample,i is transformed, and finally the output feature tensor fi is generated.
[0017] Furthermore, the dimension of the resampled feature tensor F sample,iPerform dimensional transformation and finally generate the output feature tensor fi, specifically: taking the width as the column direction and the height as the row direction, for the feature tensor F sample,i Stack along the column direction, splice the multi-channel feature maps in the column dimension, then extract features through a row convolution with a convolution kernel size of N×1 and a stride of N×1, and finally generate the output feature tensor F i+1 .
[0018] Furthermore, the local-global feature fusion module CAFM will fuse the feature tensor F fusion,i Successively perform convolution operations and channel shuffle to extract the local feature F local,i :
[0019] F local,i = W 3×3×3 (CS(W 1×1 (F fusion,i )))
[0020] where W 1×1 and W 3×3×3 represent 1×1 convolution and 3×3×3 convolution respectively, and CS represents the channel shuffle operation.
[0021] Furthermore, in the local-global feature fusion module CAFM, the fused feature tensor F fusion,i Undergoes convolution operations and attention mechanisms to extract global features. Specifically,
[0022] The fused feature tensor F fusion,i First, perform 1×1 convolution to adjust the channel dimension into three groups of feature tensors. These three groups of feature tensors respectively generate the query vector Q, the key vector K, and the value vector V through 3×3 depth convolution operations; the query vector Q and the key vector K are respectively rearranged into matrix forms: Q is converted to K is converted to
[0023] Subsequently, calculate the attention map A through matrix multiplication:
[0024]
[0025] where α is a learnable scaling parameter used to control the amplitude of the matrix multiplication result;
[0026] Multiply the attention map A by the value vector V to generate the attention feature tensor; then, after performing channel fusion on the attention feature tensor through a 1×1 convolution operation, add it to the input tensor F fusion,i to generate the global feature F global,i :
[0027] F global,i = W 1×1 (V·A)+Ffusion,i
[0028] Among them, W 1×1 represents a 1×1 convolution operation for adjusting the attention features.
[0029] Furthermore, in the dynamic detection head module, the new feature vector F input passes through scale-aware attention, spatial-aware attention, and task-aware attention to generate the final feature tensor F for the object detection tasks of object classification, center regression, and bounding box regression out , specifically,
[0030] The new feature tensor F input successively performs scale-aware attention, spatial-aware attention, and task-aware attention to obtain the final feature tensor F out :
[0031] F out = π C (π S (π L (F input )F input )·F input )·F input
[0032] Among them, π L (·), π S (·), and π C (·) respectively represent the scale-aware attention function, the spatial-aware attention function, and the task-aware attention function.
[0033] The beneficial effects of the present invention are:
[0034] First, this neonatal limb detection method based on variable kernel convolution and local-global feature fusion can dynamically adjust the sampling position of the input feature map and the shape of the convolution kernel, effectively fuse local features and global features, and can effectively improve the accuracy of neonatal limb detection.
[0035] Second, in the present invention, the variable kernel convolution layer AKConv in the feature extraction module can dynamically adjust the position of each convolution kernel to adapt to the shape of the target. At the same time, unlike traditional convolution kernels, the number of parameters of AKConv does not increase with the square of the convolution kernel size. This improvement enables the convolution kernel to no longer be limited to regular sampling points, but can flexibly adapt to the changes in irregular limb postures, thereby enhancing the ability to capture neonatal limb movements, especially the detection performance when limb postures change randomly.
[0036] III. The neonatal limb detection method based on variable kernel convolution and local-global feature fusion effectively fuses the features of PANet in the feature extraction module and the feature fusion module through the local-global feature fusion module CAFM, which can more precisely capture the local details of neonatal limbs and the global action dynamics, improve the detection ability for subtle actions in complex scenes, and ensure the integrity of feature representation.
[0037] IV. The neonatal limb detection method based on variable kernel convolution and local-global feature fusion adopts a dynamic detection head module. Its scale attention can handle the multi-scale problems of neonatal limbs and heads, helping the model to be more accurate when detecting parts of different sizes; the spatial attention can handle the spatial problems caused by changes in shooting angles and distances, ensuring that the model can adapt to different shooting angle and distance changes; the channel attention can adaptively adjust the weights of each channel, ensuring that the model focuses on the feature channels related to limbs and ignores irrelevant or noisy channels, improving the overall detection accuracy and robustness. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 is a schematic flowchart of the neonatal limb detection method based on variable kernel convolution and local-global feature fusion according to an embodiment of the present invention;
[0039] Figure 2 is a schematic illustration of the neonatal limb detection model based on variable kernel convolution and local-global feature fusion in the embodiment;
[0040] Figure 3 is a schematic illustration of the feature extraction module in the embodiment;
[0041] Figure 4 is a schematic illustration of the variable kernel convolution layer AKConv in the embodiment;
[0042] Figure 5 is a schematic illustration of the feature fusion module in the embodiment;
[0043] Figure 6 is a schematic illustration of the local-global feature fusion module CAFM in the feature fusion module in the embodiment;
[0044] Figure 7 is a schematic illustration of the dynamic detection head module in the embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0045] The preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0046] The embodiment provides a neonatal limb detection method based on variable kernel convolution and local-global feature fusion, asFigure 1 , including the following steps,
[0047] S1. Collect neonatal videos in different states, intercept key-frame images from the neonatal videos, and label the limbs and head of the neonate in the key-frame images to construct a neonatal limb detection image set.
[0048] In step S1, use the labeling tool labelImg to label the limbs and head of the neonate in the key-frame images in different action states.
[0049] S2. Construct a neonatal limb detection model based on deformable kernel convolution and local-global feature fusion. The neonatal limb detection model based on deformable kernel convolution and local-global feature fusion includes a feature extraction module, a feature fusion module, and a dynamic detection head module, and preprocesses the input image to generate a tensor where C 0 is the number of input channels, H 0 and W 0 are the height and width of the input image respectively. The feature extraction module extracts the feature tensor of the neonatal limbs from the tensor F 0 and outputs it to the feature fusion module. The feature fusion module includes a Path Aggregation Network (PANet) and a local-global feature fusion module (CAFM). The PANet generates a new feature tensor from the input feature tensors F i 、F i+1 、F i+2 of the neonatal limbs at the i-th to (i + 2)-th where, C' i 、C' i+1 、C' i+2 are the number of channels, H i 、H i+1 、H i+2 represent the height, and W i 、W i+1 、W i+2 represent the width; the CAFM concatenates the feature tensors F i 、F i+1 、F i+2 of the neonatal limbs and the new feature tensors F' i 、F' i+1 、F' i+2 along the channel dimension in groups with the same subscript to generate a fused feature tensor where, C i + C' i 、C i+1 + C' i+1 、Ci+2 +C' i+2 is the number of channels, H i , H i+1 , H i+2 represents the height, W i , W i+1 , W i+2 represents the width; the fused feature tensor F fusion,i , F fusion,i+1 , F fusion,i+2 successively undergoes convolution operations and channel shuffling to extract local features, and the fused feature tensor F fusion,i , F fusion,i+1 , F fusion,i+2 undergoes convolution operations and an attention mechanism to extract global features, and then the local features and global features are fused to generate an output feature tensor Among them, C CAFM,i , C CAFM,i+1 , C CAFM,i+2 is the number of channels, H i , H i+1 , H i+2 represents the height, W i , W i+1 , W i+2 represents the width; the dynamic detection head module receives the output feature tensors F CAFM,i , F CAFM,i+1 , F CAFM,i+2 , and aligns the features with the feature tensor F CAFM,i+1 as the benchmark, and then concatenates the three aligned features in the channel dimension to generate a new feature vector F input ; then the new feature vector F input is processed by scale-aware attention, space-aware attention, and task-aware attention to generate the final feature tensor F out for the object detection tasks of object classification, center regression, and bounding box regression. After post-processing the final feature tensor F out , the detected image is obtained.
[0050] In step S2, the feature extraction module includes M groups of deformable kernel convolutional layers AKConv, a cross-stage partial bottleneck layer with two convolutions, i.e., the C2f layer, and a fast-spatial pyramid pooling layer, i.e., the SPPF layer. The value of M is an integer from 3 to 7. The tensor F 0 successively passes through M groups of deformable kernel convolutional layers AKConv and the C2f layer and then is input into the SPPF layer. In the i-th group of deformable kernel convolutional layers AKConv and the C2f layer, the deformable kernel convolutional layer AKConv generates the feature tensor fi from the input feature tensor F i-1 , and the C2f layer is used to convert the output feature tensor fi of the deformable kernel convolutional layer AKConv into a higher-quality feature tensor Among them, C i is the number of channels, and the value of i is an integer from 1 to M - 2; the SPPF layer outputs the feature tensor of the Mth layer Among them, C M is the number of output feature channels, H M and W M are the height and width of the output feature map respectively. The Fast-Spatial Pyramid Pooling layer SPPF can improve the robustness of the network in processing objects of different sizes.
[0051] For example Figure 3 , the embodiment is described as follows with the feature extraction module including 5 Atrous Kernel Convolution layers AKConv, 4 C2f layers, and one SPPF layer: Input the neonatal image, and after convolution operations, features of different scales are output. Among them, the C2f layer and the SPPF layer after the 3rd - 4th Atrous Kernel Convolution layers AKConv output the feature tensors F 3 、F 4 、F 5 of the 3rd - 5th neonatal limbs respectively, which are 1 / 8, 1 / 16, and 1 / 32 of the input image size.
[0052] For example Figure 4 , the Atrous Kernel Convolution layer AKConv generates the feature tensor fi from the input feature tensor F i-1 . Specifically
[0053] First generate the offset ΔP through the two-dimensional convolution operation Conv2d where N is the size of the convolution kernel, H i-1 and W i-1 are the height and width of the input feature map respectively; according to the generated offset ΔP n , adjust the initial regular sampling coordinates P 0 to obtain the modified sampling coordinates P' n : P' n = P 0 + ΔP n ; this process reflects the variable kernel characteristic of the Atrous Kernel Convolution layer AKConv, enabling the convolution kernel to dynamically adjust the sampling shape according to different regions of the input features; subsequently n , use the modified sampling coordinates P' i to resample the feature tensor F After completing the resampling sample,i perform dimensional transformation on the resampled feature tensor F where Ci is the number of channels of the output feature map, H i and W i are the height and width of the output feature map respectively.
[0054] Perform dimensional transformation on the resampled feature tensor F sample,i to finally generate the output feature tensor fi. Specifically: with the width as the column direction and the height as the row direction, stack the feature tensor F sample,i along the column direction, concatenate the multi-channel feature maps in the column dimension, and then extract features through a row convolution with a convolution kernel size of N×1 and a stride of N×1 to finally generate the output feature fi.
[0055] After the input features enter the deformable kernel convolution layer AKConv, it calculates the offset by learning the dynamic sampling positions, and then generates the positions of the new sampling points according to the calculation results of the offset convolution, rather than sampling on a fixed grid; since the final sampling point positions may fall between pixels, bilinear interpolation is used in the next step to combine the offset coordinates to ensure smooth sampling of the features, ensuring that the features at the new positions can better represent local information; in the next step, these resampled features are integrated into new features to complete the feature reconstruction; subsequently, a dynamic convolution kernel is adaptively generated based on the input features, and by adjusting the weights or shapes of the convolution kernels, it is made more suitable for the current feature distribution and task requirements; finally, the new features after offset are convolved to obtain the output features.
[0056] For example Figure 5 , in the feature fusion module, the feature tensors F 3 , F 4 , F 5 of the neonatal limb and the new feature tensors F 3 ’, F 4 ’, F 5 ’ obtained through the path aggregation network PANet. Then the new feature tensors F 3 ’, F 4 ’, F 5 ’ are respectively input into the local-global feature fusion module CAFM together with the feature tensors F 3 , F 4 , F 5 of the neonatal limb for same-level feature fusion to generate the output feature tensor.
[0057] For example Figure 6 , taking the feature tensor F 3 of the neonatal limb as an example, it is described as follows: the feature tensor F 3 of the neonatal limb and the new feature tensor F' 3The initial fusion is performed by simple addition to generate a fused feature tensor; then, local features and global features are obtained through depthwise separable convolution and multi-head attention mechanism in the local branch and global branch respectively; finally, the local features and global features are added to obtain the output feature tensor.
[0058] In the local-global feature fusion module CAFM, in the local feature processing part, the input fused feature tensor F fusion,i First, a 1×1 convolution is performed to adjust the channel dimension, and the channel information is further mixed through a channel shuffle operation. In the shuffle operation, the tensor is divided into several groups, and each group undergoes a depth convolution operation for channel fusion. Subsequently, the output tensors of each group are concatenated along the channel dimension to generate a new feature tensor F mid,i ; Subsequently, features are extracted through a 3×3×3 convolution operation to obtain the local feature tensor F local,i :
[0059] F local,i = W 3×3×3 (CS(W 1×1 (F fusion,i )))
[0060] where W 1×1 and W 3×3×3 represent 1×1 convolution and 3×3×3 convolution respectively, and CS represents the channel shuffle operation.
[0061] In the local-global feature fusion module CAFM, in the global feature processing part, the fused feature tensor F fusion,i is passed through a convolution operation and an attention mechanism to extract global features. Specifically,
[0062] The fused feature tensor F fusion,i First, a 1×1 convolution is performed to adjust the channel dimension into three groups of feature tensors. These three groups of feature tensors are respectively passed through a 3×3 depth convolution operation to generate a query vector Q, a key vector K, and a value vector V; the query vector Q and the key vector K are respectively rearranged into matrix forms: Q is converted to K is converted to
[0063] Subsequently, the attention map A is calculated through matrix multiplication:
[0064]
[0065] where Softmax represents the normalized exponential function, and α is a learnable scaling parameter used to control the magnitude of the matrix multiplication result;
[0066] Multiply the attention map A by the value vector V to generate an attention feature tensor; then, after performing a 1×1 convolution operation on the attention feature tensor for channel fusion, add it to the fused feature tensor to generate the global feature F global,i :
[0067] F global,i = W 1×1 (V·A)+F fusion,i
[0068] where W 1×1 represents a 1×1 convolution operation for adjusting the attention features.
[0069] Finally, add the local feature F local,i and the global feature F global,i in the channel dimension to obtain the output feature tensor F CAFM,i , and its formula is:
[0070] F CAFM,i = F global,i + F local,i
[0071] Output feature tensor where (C local,i + C global,i ) is the number of output channels, and H i and W i are the height and width of the output feature map, respectively.
[0072] Similarly, the local-global feature fusion module CAFM processes the fused feature tensors F fusion,i+1 , F fusion,i+2 , respectively, to obtain local and global features, and then fuses the local and global features to generate the output feature tensors F CAFM,i+1 , F CAFM,i+2 .
[0073] The dynamic detection head module will receive the output feature tensors F CAFM,i , F CAFM,i+1 , F CAFM,i+2 , and align the features with F CAFM,i+1 as the reference. Specifically, the spatial dimensions of F CAFM,i and F CAFM,i+2 are adjusted to the height and width H CAFM,i+1 × W i+1 × W i+1 of F by upsampling or downsampling. By aligning the spatial feature maps at different levels, the dimensions of the feature maps are ensured to be consistent, preparing for subsequent fusion. The aligned feature tensors are concatenated in the channel dimension to generate a new feature vector i + C i+1+C i+2 is the number of channels after stacking.
[0074] In the dynamic detection head module, the new feature vector F input passes through scale-aware attention, spatial-aware attention, and task-aware attention to generate the final feature tensor F for the object detection tasks of object classification, center regression, and bounding box regression out , specifically,
[0075] The new feature tensor F input successively performs scale-aware attention, spatial-aware attention, and task-aware attention to obtain the final feature tensor F out :
[0076] F out = π C (π S (π L (F input )·F input )·F input )·F input
[0077] where π L (·), π S (·), and π C (·) represent the scale-aware attention function, spatial-aware attention function, and task-aware attention function respectively.
[0078] The new feature tensor F input first dynamically fuses features from different scales through the scale-aware attention mechanism. The calculation formula of the scale-aware attention function π L (·) is:
[0079]
[0080] where f(·) is a linear function implemented by a 1×1 convolution; σ(x) is the Sigmoid activation function, which is used to generate scale-aware weights to weight and fuse features of different scales;
[0081] The feature tensor W L (F input ) after scale-aware attention processing is:
[0082] W L (F input ) = π L (F input )·F input
[0083] Subsequently, the feature tensor W L (Finput ) Through the spatial perception attention function, it further focuses on the more discriminative spatial regions, and the spatial perception attention function π S (·) is calculated as follows:
[0084]
[0085] where K represents the number of sparse sampling positions, and Δp k represents the offset calculated through self-learning, and Δm k represents the importance weight of the sampling position;
[0086] The feature tensor W S (F input ) is as follows:
[0087] W S (F input ) = π S (W L (F input )) · W L (F input )
[0088] Finally, the feature tensor W S (F input ) completes the feature enhancement under different task requirements through the task perception attention function by dynamically activating or suppressing the channels related to specific tasks. The task perception attention obtains the feature vector W C (F input ) after the task perception attention processing as the final feature tensor F out
[0089] F out = W C (F input ) = π C (W S (F input )) · W S (F input )
[0090] where the final feature tensor has the same shape as the input reference feature tensor and is used for object detection tasks such as object classification, center point regression, and bounding box regression; the calculation formula of the task perception attention function π C (·) is as follows:
[0091] π C (W S (F input )) · W S (Finput )
[0092] = max(α 1 (W S (F input ))·F c + β 1 (W S (F input )), α 2 (W S (F input ))·F c + β 2 (W S (F input )))
[0093] where F c represents the feature of the c-th channel, and α 1 , α 2 , β 1 , β 2 are weights learned through global average pooling and fully connected layers.
[0094] For example, Figure 7 , in the dynamic detection head module, after feature map alignment, the three aligned features are concatenated in the channel dimension to generate a new feature vector F input . Next, it enters the scale perception stage. The scale perception attention module generates weights for each layer based on the context information of the input features (such as the result of adaptive average pooling). These weights reflect the importance of different hierarchical features in the current scene. According to the generated weights, the features of each layer are weighted and fused to ensure that each layer of features can contribute reasonably to the final fusion. Then, the spatial perception attention module generates a position weight matrix based on the global feature information and weights each position of the feature map to highlight the key regions. The task perception attention module is used to dynamically adjust the channel weights of the features according to the requirements of the classification and regression tasks. The finally fused feature map is used to predict the bounding box of the object, providing accurate results for the detection task.
[0095] S3. After training the neonatal limb detection model based on variable kernel convolution and local-global feature fusion using the neonatal limb detection image set, the trained neonatal limb detection model based on variable kernel convolution and local-global feature fusion is obtained.
[0096] S4. Input the neonatal limb detection video to be tested into the trained neonatal limb detection model based on variable kernel convolution and local-global feature fusion for limb detection.
[0097] This neonatal limb detection method based on variable kernel convolution and local-global feature fusion involves establishing a neonatal limb detection image set; constructing a neonatal limb detection model based on variable kernel convolution and local-global feature fusion; training the neonatal limb detection model with the neonatal limb detection image set; and using the trained neonatal limb detection model to perform limb detection on each image frame of a newly input test video. This method can dynamically adjust the sampling positions of the input feature maps and the shapes of the convolution kernels, effectively fuse local and global features, and improve the accuracy of neonatal limb detection.
[0098] This neonatal limb detection method based on variable kernel convolution and local-global feature fusion constructs a neonatal limb detection model based on variable kernel convolution and local-global feature fusion, which includes a feature extraction module, a feature fusion module, and a dynamic detection head module. The feature extraction module extracts features from the input image, then the feature fusion module fuses the feature tensors to obtain features at different levels, and finally the dynamic detection head performs object classification and bounding box regression on the fused features to complete the detection of neonatal limbs.
[0099] In the present invention, the variable kernel convolution layer AKConv in the feature extraction module can dynamically adjust the position of each convolution kernel to adapt to the shape of the target. At the same time, unlike traditional convolution kernels, the number of parameters of AKConv does not increase with the square of the convolution kernel size. This improvement enables the convolution kernel to be no longer limited to regular sampling points, but can flexibly adapt to the changes in irregular limb postures, thereby enhancing the ability to capture neonatal limb movements, especially the detection performance when the limb postures change randomly.
[0100] This neonatal limb detection method based on variable kernel convolution and local-global feature fusion effectively fuses the features of PANet in the feature extraction module and the feature fusion module through the local-global feature fusion module CAFM, which combines local convolutional features and global attention features. It can more accurately capture the local details of neonatal limbs and the global action dynamics, improve the detection ability for subtle movements in complex scenes, and ensure the integrity of feature representation.
[0101] This neonatal limb detection method based on variable kernel convolution and local-global feature fusion adopts a dynamic detection head module. Its scale attention can handle the multi-scale problems of neonatal limbs and heads, helping the model to be more accurate when detecting parts of different sizes; spatial attention can handle the spatial problems caused by changes in shooting angles and distances, ensuring that the model can adapt to different shooting angle and distance changes; channel attention can adaptively adjust the weights of each channel, ensuring that the model focuses on the feature channels related to limbs and ignores irrelevant or noisy channels, improving the overall detection accuracy and robustness.
[0102] This neonatal limb detection method based on variable kernel convolution and local-global feature fusion is a neonatal limb detection method based on variable kernel convolution and dual-branch multi-level feature fusion. It can dynamically adjust the sampling positions of the input feature maps and the shapes of the convolution kernels, effectively fuse local features and global features, and can effectively improve the accuracy of neonatal limb detection.
[0103] As described above, it is only the specific implementation manner in the present invention, but the protection scope of the present invention is not limited thereto. Any transformation or replacement that can be understood and conceived by those familiar with the technology within the technical scope disclosed by the present invention should be covered within the scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. A method for detecting neonatal limbs based on variable kernel convolution and local-global feature fusion, characterized in that: The following steps are included: S1. Collect videos of newborns in different states, capture key frame images from the newborn videos, and mark the limbs and head of the newborns in the key frame images to construct a newborn limb detection image set; S2. Construct a neonatal limb detection model based on variable kernel convolution and local global feature fusion. The neonatal limb detection model based on variable kernel convolution and local global feature fusion includes a feature extraction module, a feature fusion module and a dynamic detection head module. The input image is preprocessed to generate a tensor F0. The feature extraction module extracts the feature tensor of the neonatal limb from the tensor F0 based on the variable kernel convolution and outputs it to the feature fusion module. The feature fusion module includes a path aggregation network PANet and a local global feature fusion module CAFM. The path aggregation network PANet converts the input feature tensors F0 of the i-th to i+2-th neonatal limbs into a local global feature fusion module CAFM. i 、F i+1 、F i+2 Generate a new feature tensor F' i 、F' i+1 、F' i+2 ; The local-global feature fusion module CAFM transforms the feature tensor F of the newborn’s limbs i 、F i+1 、F i+2 With the new feature tensor F' i 、F' i+1 、F' i+2 Corresponding to the splicing along the channel dimension to generate the fusion feature tensor F fusion,i 、F fusion,i+1 、F fusion,i+2 , the fused feature tensor F fusion,i 、F fusion,i+1 、F fusion,i+2 The local features are extracted through convolution operations and channel shuffling, and the fused feature tensor F fusion,i 、F fusion,i+1 、F fusion,i+2 After the convolution operation and attention mechanism, the global features are extracted, and then the local features are combined with the global features. Fusion, generate output feature tensor F CAFM,i 、F CAFM,i+1 、F CAFM,i+2 ; The dynamic detection head module will receive the output feature tensor F CAFM,i 、F CAFM,i+1 、F CAFM,i+2 , with the feature tensor F CAFM,i+1 The three aligned features are then concatenated in the channel dimension to generate a new feature vector F input ; Then the new eigenvector F input After scale-aware attention, space-aware attention, and task-aware attention processing, the final feature tensor F is generated for the object detection task of object classification, center regression, and bounding box regression. out , for the final feature tensor F out After post-processing, the detection image is obtained; S3, after using the newborn limb detection image set to train the newborn limb detection model based on variable kernel convolution and local global feature fusion, a trained newborn limb detection model based on variable kernel convolution and local global feature fusion is obtained; S4. Input the newborn limb detection video to be tested into the trained newborn limb detection model based on variable kernel convolution and local-global feature fusion for limb detection.
2. The method for detecting neonatal limbs based on variable kernel convolution and local-global feature fusion as claimed in claim 1, characterized in that: In step S2, the feature extraction module includes M groups of variable kernel convolution layers AKConv and a cross-stage partial bottleneck layer with two convolutions, namely the C2f layer, and a fast-spatial pyramid pooling layer, namely the SPPF layer. The tensor F0 passes through M groups of variable kernel convolution layers AKConv and C2f layers in sequence and is input into the SPPF layer. In the i-th group of variable kernel convolution layers AKConv and C2f layers, the variable kernel convolution layer AKConv converts the input feature tensor F i-1 Generate feature tensor fi, the C2f layer is used to transform the feature tensor fi into a higher quality feature tensor F i And output; SPPF layer outputs the Mth layer feature tensor F M .
3. The method for detecting neonatal limbs based on variable kernel convolution and local-global feature fusion as claimed in claim 1, characterized in that: The variable kernel convolution layer AKConv converts the input feature tensor F i-1 Generate feature tensor fi, specifically, First, the feature tensor F i-1 The offset ΔP is generated by the two-dimensional convolution operation Conv2d n ; Based on the generated offset ΔP n , adjust the initial regular sampling coordinate P0, and obtain the modified sampling coordinate P' n :P' n =P0+ΔP n ; Using the modified sampling coordinates P' n , for the feature tensor F i Resample and calculate the eigenvalues corresponding to the sampling positions through bilinear interpolation to obtain the resampled eigentensor F sample,i ; The resampled feature tensor F sample,i Perform dimension transformation and finally generate the output feature tensor fi.
4. The method for detecting neonatal limbs based on variable kernel convolution and local-global feature fusion as claimed in claim 3, characterized in that: The resampled feature tensor F sample,i Perform dimensional transformation and finally generate the output feature tensor fi. Specifically, the feature tensor F is transformed with width as column direction and height as row direction. sample,i Stack along the column direction, concatenate the multi-channel feature maps in the column dimension, and then extract features through row convolution with a convolution kernel size of N×1 and a step size of N×1, and finally generate the output feature tensor F i+1 .
5. The method for detecting neonatal limbs based on variable kernel convolution and local-global feature fusion according to any one of claims 1 to 4, characterized in that: The local-global feature fusion module CAFM fuses the feature tensor F fusion,i The local features F are extracted by convolution operation and channel shuffling in sequence. local,i : F local,i =W 3×3×3 (CS(W 1×1 (F fusion,i ))) Among them, W 1×1 and W 3×3×3 They represent 1×1 convolution and 3×3×3 convolution respectively, and CS represents the channel shuffle operation.
6. The method for detecting neonatal limbs based on variable kernel convolution and local-global feature fusion according to any one of claims 1 to 4, characterized in that: In the local-global feature fusion module CAFM, the fused feature tensor F fusion,i After convolution operation and attention mechanism, global features are extracted, specifically, Input tensor F fusion,i First, a 1×1 convolution is performed to adjust the channel dimension into three groups of feature tensors. These three groups of feature tensors are respectively generated through 3×3 deep convolution operations to generate query vector Q, key vector K and value vector V; query vector Q and key vector K are rearranged into matrix form: Q is converted to K is converted to Then, the attention map A is calculated by matrix multiplication: Among them, α is a learnable scaling parameter used to control the amplitude of the matrix multiplication result; The attention map A is multiplied by the value vector V to generate the attention feature tensor; the attention feature tensor is then fused with the input tensor F after a 1×1 convolution operation. fusion,i Add to generate the global feature F global,i : F global,i =W 1×1 (V·A)+F fusion,i Among them, W 1×1 Represents a 1×1 convolution operation, which is used to adjust the attention features.
7. The method for detecting neonatal limbs based on variable kernel convolution and local-global feature fusion according to any one of claims 1 to 4, characterized in that: In the dynamic detection head module, the new feature vector F input After scale-aware attention, space-aware attention, and task-aware attention, the final feature tensor F for the object detection task of object classification, center regression, and bounding box regression is generated. out , specifically, The new feature tensor F input Execute scale-aware attention, space-aware attention, and task-aware attention in sequence to obtain the final feature tensor F out : F out =π C (p S (p L (F input )·F input )·F input )·F input Among them, π L (·),π S (·) and π C (·) represent scale-aware attention function, space-aware attention function and task-aware attention function respectively.
Citation Information
Cited By
Method and device for detecting multi-scale target of unmanned aerial vehicle
CN122416325A