Newborn limb movement recognition method based on double-branch multi-scale feature fusion
Through the fusion of double-branch multi-scale feature and multi-scale convolutional attention mechanism, the problem of insufficient subtle motion capture ability in the recognition of limb movements in neonatals is solved, and higher recognition accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510182218.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-02-18
AI Technical Summary
The prior art is difficult to accurately capture the subtle changes in limb movements of newborns and the low accuracy of motion recognition, especially in complex backgrounds, which are difficult to distinguish between movements and backgrounds.
A neonatal limb movement recognition method based on the fusion of multi-scale features of dual branches is adopted. The action features are extracted through the first branch and the second branch respectively, and the multi-scale convolutional attention mechanism is used to fusion and adjust the feature to improve the ability to capture subtle movement changes.
It effectively improves the accuracy and robustness of body movement recognition for newborns, can capture movement differences more accurately and adapt to complex scenarios, with a small amount of calculation.
Smart Images

Figure CN120014710A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a method for recognizing neonatal limb movements based on dual-branch multi-scale feature fusion, and belongs to the technical field of movement recognition. Background Art
[0002] Action is a special way of communication analysis. When they cannot express themselves in words, we can reveal their inner state and needs by observing their subtle body movements and posture changes. This reaction usually occurs naturally without consciousness and is difficult to hide or control. It is often directly related to a person's true emotions and can reveal their true feelings and psychological conditions. As a branch of action recognition research, the body movements of newborns focus on analyzing and interpreting the body language of infants when they cannot express themselves in words. It can provide a deeper understanding of the non-verbal communication methods of newborns, thereby providing them with more accurate and timely care. Therefore, newborn-related action recognition has very important applications in human-computer interaction, early health detection, medical diagnosis assistance, and home care.
[0003] As an emerging research field, newborn body movement recognition has attracted the attention of many researchers. Although researchers are also trying to introduce attention mechanisms and temporal pyramid networks for action recognition, existing methods are often subtle and short-lived for newborn body movements, and existing technologies may find it difficult to accurately capture the beginning and end of these tiny movements.
[0004] In actual application scenarios, the movements of newborns will be affected by the complex background and the large amount of model calculation, which may result in the inability of existing methods to effectively distinguish between movements and background, and high difficulty in obtaining information. Therefore, it is difficult to train a high-accuracy and robust motion recognition model, resulting in low motion recognition accuracy. Summary of the invention
[0005] The purpose of the present invention is to provide a method for recognizing neonatal limb movements based on dual-branch multi-scale feature fusion to solve the problems in the prior art of low ability to capture subtle movement changes and the need to improve the accuracy of movement recognition.
[0006] The technical solution of the present invention is:
[0007] A method for identifying neonatal limb movements based on dual-branch multi-scale feature fusion, comprising the following steps:
[0008] S1. Collect videos of newborns in different limb movement states, segment the videos according to the time interval T to obtain video segments, and use the first frame of each video segment as the key frame, detect the key limb parts of the newborn in the key frame, including the head, upper limbs and lower limbs, generate corresponding bounding boxes, mark the newborn limb movement labels, integrate the video segments and labels, and construct a newborn limb movement video sample set;
[0009] S2. Construct a newborn limb movement recognition model based on dual-branch multi-scale feature fusion. The newborn limb movement recognition model based on dual-branch multi-scale feature fusion includes a first branch feature extraction network, a second branch feature extraction network, a feature fusion module and a classifier, wherein:
[0010] The first branch feature extraction network: The deep convolutional neural network Net1 is used to extract action features F at N levels from the input video segment. 1,p , where p = 1, 2, ..., N, and the action features F extracted at the Nth level 1,N Output to the feature fusion module; where N is an integer ranging from 3 to 5;
[0011] The second branch feature extraction network: The input video segment is frame-sampled at intervals to obtain a low frame rate video. The low frame rate video is extracted by the deep convolutional neural network Net2 at N levels to extract action features F 2,p , and the action features F extracted from the first three levels 2,p With action feature F 1,p The features obtained by introducing the multi-scale convolutional attention mechanism are then combined with the action features F 2,p After the residual connection, the residual connection feature F is obtained r,p , as the input feature of the next 3D residual block in the deep convolutional neural network Net2, the action feature F extracted at the Nth level 2,N Output to the feature fusion module;
[0012] Feature fusion module: used to fuse the newborn limb movement features output by the first branch feature extraction network and the second branch feature extraction network at the Nth level to obtain the fused feature F;
[0013] Classifier: classifies and identifies the fused feature F and outputs the body movement category;
[0014] S3, using the newborn limb movement video sample set obtained in step S1 to train the newborn limb movement recognition model based on dual-branch multi-scale feature fusion constructed in step S2, to obtain a newborn limb movement recognition model;
[0015] S4. Use the newborn limb movement recognition model obtained in step S3 to perform movement recognition on the newly input test video.
[0016] Furthermore, in step S1, the newborn's limb movement labels include head shaking, fingers opening, fist clenching, arm waving, leg kicking and standing still.
[0017] Furthermore, in step S1, the time interval T ranges from 8 frames to 64 frames.
[0018] Furthermore, in step S2, the second branch feature extraction network includes a frame extraction module, a deep convolutional neural network Net2, a deep convolutional neural network Net2, three multi-scale convolutional attention modules and three residual connection modules.
[0019] Frame extraction module: extracts frames from the input video segment at intervals to obtain low frame rate video output to the deep convolutional neural network Net2;
[0020] Deep convolutional neural network Net2: Extract action features F at N levels for low frame rate videos 2,p ;
[0021] Multi-scale convolutional attention module: respectively input action features F 1,p and action feature F 2,p Correspondingly, a multi-scale convolutional attention mechanism is introduced and a multi-scale feature tensor Fpˋ is output;
[0022] Residual connection module: The input multi-scale feature tensor Fpˋ is combined with the action feature F output by the previous layer of the deep convolutional neural network Net2. 2,p After the residual connection, the output is sent to the 3D residual block of the next layer of the deep convolutional neural network Net2.
[0023] Furthermore, in the multi-scale convolutional attention module, the input action features F 1,p and action feature F 2,p Correspondingly, a multi-scale convolutional attention mechanism is introduced and a multi-scale feature tensor Fpˋ is output, specifically,
[0024] 1) The action feature F 1,p and action feature F 2,p Concatenate in the channel dimension to obtain the feature tensor Fp;
[0025] 2) Use the multi-scale convolutional attention mechanism to obtain the feature tensor F p The attention weight ω p for:
[0026]
[0027] Among them, σ represents the Sigmoid activation function, Conv 1×1×1represents a point-by-point convolution with a convolution kernel of 1×1×1, L is the number of depth-wise separable convolutions, L∈{2,3,4,5}; DW j denotes the jth banded convolution kernel as k×1×1, 1×k×1 and 1×1×k depthwise separable convolution, k∈{3,5,7,9,11};
[0028] 3) The attention weight ω p Channel by channel and feature tensor F p Multiply them together and perform point-by-point convolution with a convolution kernel of 1×1×1 to output a multi-scale feature tensor F p 'for:
[0029] F p ′=Conv 1×1×1 (ω p ·F p )
[0030] Among them, Conv 1×1×1 It indicates a point-by-point convolution with a convolution kernel of 1×1×1.
[0031] Furthermore, the feature fusion module includes the SE attention module and the efficient feature fusion block.
[0032] SE attention module: a squeeze-excitation operation is applied to the outputs of the first feature extraction network and the second branch feature extraction at the Nth level to obtain the enhanced first action feature tensor and the enhanced second action feature tensor;
[0033] Efficient feature fusion block: After adaptive pooling and convolution processing on the enhanced first action feature tensor and the enhanced second action feature tensor, the convolution first action feature and the convolution second action feature are obtained. The activation function Sigmoid is used to obtain the weights, and the weights are multiplied channel by channel with the enhanced action feature tensor and the enhanced second action tensor to obtain the first intermediate action feature and the second intermediate action feature, which are then concatenated to obtain the fused feature F.
[0034] Furthermore, the classifier includes a pooling block and a fully connected layer.
[0035] Pooling block: The fused feature F is subjected to maximum pooling and average pooling operations to obtain the pooled feature. The pooled feature is stretched and flattened to output the feature vector V.
[0036] Fully connected layer: maps the feature vector V to the probability distribution of the corresponding prediction category and outputs the body movement category.
[0037] Furthermore, in step S3, when training the newborn limb movement recognition model based on dual-branch multi-scale feature fusion constructed in step S2, the focal loss function is used to optimize the training:
[0038]
[0039] Among them, loss i,c represents the binary cross entropy loss, FL represents the focal loss, K represents the total number of categories, c∈[1,K], y i,c Indicates the true value that the i-th sample belongs to the c-th category, i∈[1,n], n represents the total number of samples, It indicates the model prediction value that the i-th sample belongs to the c-th category, α is the balancing factor, which is used to adjust the weights of positive and negative samples, and β is the adjustment factor, which is used to make the weight of difficult-to-classify labels larger.
[0040] The beneficial effects of the present invention are:
[0041] 1. Compared with the existing methods, this method for recognizing newborn limb movements based on dual-branch multi-scale feature fusion uses dual branches to extract movement features from videos with different frame rates, improves the ability to capture subtle movement changes, and adopts a multi-scale convolutional attention mechanism to focus on the key features of limb movements, effectively improving the accuracy of newborn limb movement recognition and its robustness to complex scenes, with less computational effort.
[0042] 2. In order to address the problem that the sampling frequency of a single branch cannot flexibly adapt to the diversity of actions, which leads to information redundancy or insufficiency and affects the training speed, the present invention uses dual branches to extract action features from videos with different frame rates, capture rapidly changing short-term dynamic details, alleviate the huge imbalance between effective action features and invalid background features, and enhance the synergy of features at different time scales.
[0043] 3. This method of neonatal limb movement recognition based on dual-branch multi-scale feature fusion introduces a multi-scale convolutional attention mechanism, integrates the feature extraction capabilities of different receptive fields, and combines the weighted strategy of global and local information to achieve accurate capture of multi-scale features, effectively improving the model's ability to recognize detailed features in complex scenes, and using deep separable convolution instead of traditional convolution to effectively reduce the computational complexity of the model. In the task of neonatal limb movement recognition, the multi-scale convolutional attention mechanism can further improve the accuracy and effectiveness of the model.
[0044] 4. This method of newborn limb movement recognition based on dual-branch multi-scale feature fusion, by fusing the dual-branch movement features, dynamically adjusts the weights of feature channels, effectively highlights the expressiveness of key feature areas, enables the model to focus on the key features of limb movements, dynamically adapts to movements of different time scales, accurately captures subtle movement differences, and effectively improves the accuracy and robustness of movement recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 It is a structural schematic diagram of a method for recognizing newborn limb movements based on dual-branch multi-scale feature fusion according to an embodiment of the present invention.
[0046] Figure 2 3 is a schematic diagram illustrating a newborn baby limb movement recognition model based on dual-branch multi-scale feature fusion in an embodiment.
[0047] Figure 3 Schematic diagram of the multi-scale convolutional attention module in the embodiment.
[0048] Figure 4 Schematic diagram for explaining the feature fusion module in the embodiment. DETAILED DESCRIPTION
[0049] The preferred embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0050] Embodiment A method for identifying newborn limb movements based on dual-branch multi-scale feature fusion, such as Figure 1 , including the following steps:
[0051] S1. Collect videos of newborns in different limb movement states, segment the videos according to time interval T to obtain video segments, and use the first frame of each video segment as the key frame. Detect the key limb parts of the newborn in the key frame, including the head, upper limbs and lower limbs, generate corresponding bounding boxes, annotate the newborn's limb movement labels, integrate the video segments and labels, and construct a newborn's limb movement video sample set.
[0052] In step S1, the YOLO model is used to detect the key limb parts of the newborn such as the head, upper limbs, lower limbs, etc. in the key frame, and the corresponding bounding box is generated. The annotation tool VGG Image Annotator is used to fine-tune the position and size of the bounding box, and the newborn limb movement labels are marked on the bounding box. The newborn limb movement labels include head shaking, fingers open, fists, arm waving, kicking, and standing still. The value range of the time interval T is 8 frames to 64 frames. In this embodiment, T is equal to 32 frames. The newborn limb movement video sample set of this embodiment includes newborn limb movement video samples and labels with a time length of 10 seconds and a frame rate of 30-fps.
[0053] S2. Construct a newborn limb movement recognition model based on dual-branch multi-scale feature fusion. The newborn limb movement recognition model based on dual-branch multi-scale feature fusion includes a first branch feature extraction network, a second branch feature extraction network, a feature fusion module and a classifier, wherein:
[0054] The first branch feature extraction network: The deep convolutional neural network Net1 is used to extract action features F at N levels from the input video segment. 1,p , where p = 1, 2, ..., N, and the action features F extracted at the Nth level 1,N Output to the feature fusion module; where N is an integer ranging from 3 to 5.
[0055] like Figure 2 The first branch feature extraction network takes N=3 and the deep convolutional neural network Net1 uses 3DResNet-50 as an example. The first branch passes through the deep convolutional neural network Net1. The first branch obtains more timing information by capturing the fast motion information in the video. ResNet-50 contains 1 3D convolution layer, 1 maximum pooling layer, and 16 3D residual blocks. The 16 3D residual blocks are represented as 3D residual blocks B1 to 3D residual blocks B2. 16 The output features of N 3D residual blocks are taken as the action features F extracted at N levels. 1,p , the features output by 3D ResNet-50 at the first three different levels are represented as Among them, p represents the sequence number of the feature level, p∈{1,2,3}, H p Represents the height of the feature map, W p Represents the width of the feature map, C 1,p Represents the number of channels of the deep convolutional neural network Net1, Q represents the channel scaling ratio, which is used to achieve lightweight model. Let Q be an integer ranging from 2 to 16. 2,pRepresents the number of channels of the deep convolutional neural network Net2. The time interval T is specified as 32 frames, Q=8, and the number of output channels of the 3D convolution layer is set to 8, which is used to achieve lightweight model, and provide less redundant spatial information to the first branch through fewer channels and weaker spatial information processing capabilities. The 3D convolution layer uses a 5×7×7 convolution kernel of size to convolve the input feature tensor of the branch, with a stride of 1×2×2; the pooling layer uses a 1×3×3 maximum pooling layer to pool the input feature tensor; the 3D residual block uses 3 different 3D convolution layers, and the 3D residual block also uses m2 3×1×1, 1×3×3 and 1×1×1 convolution kernels to convolve the input feature tensor, where m2 can be 16, 32, 64. Select 3D residual block B3, 3D residual block B7, and 3D residual block B from the 16 3D residual blocks. 13 Output feature tensor F 1,p Input into the multi-scale convolution attention module, the feature tensor scales of the three residual blocks output are respectively
[0056] The second branch feature extraction network: The input video segment is frame-sampled at intervals to obtain a low frame rate video. The low frame rate video is extracted by the deep convolutional neural network Net2 at N levels to extract action features. Among them, p represents the sequence number of the feature level, H p Represents the height of the feature map, W p Represents the width of the feature map, C 2,p represents the number of channels of the deep convolutional neural network Net2, and the action features F extracted from the first three levels 2,p With action feature F 1,p The features obtained by introducing the multi-scale convolutional attention mechanism are then combined with the action features F 2,p After concatenation, we get the residual connection feature F r,p , as the input feature of the next 3D residual block in the next deep convolutional neural network Net2, the action feature F extracted at the Nth level 2,N Output to the feature fusion module.
[0057] In step S2, the second branch feature extraction network includes a frame extraction module, a deep convolutional neural network Net2, three multi-scale convolutional attention modules and three residual connection modules:
[0058] Frame extraction module: The input video segment is intermittently extracted at a frame extraction frequency M to obtain a low frame rate video output to the deep convolutional neural network Net2; the value range of the frame extraction frequency M is an integer from 2 to 16.
[0059] Deep convolutional neural network Net2: Extract action features F at N levels for low frame rate videos 2,p ;
[0060] Multi-scale convolutional attention module: respectively input action features F 1,p and action feature F 2,p Correspondingly, a multi-scale convolutional attention mechanism is introduced and a multi-scale feature tensor Fpˋ is output; Figure 3 , specifically:
[0061] 1) The action feature F 1,p and action feature F 2,p Splice and get the feature tensor F p ;
[0062] 2) The feature tensor F p The input is fed into three parallel depth-wise separable convolution blocks, where the three depth-wise separable convolution blocks use m 3 3 groups of 3×3×3 convolution kernels and m3 groups of 3 band convolution kernels are used for multi-scale information extraction, where the optional values of the band convolution kernel are 7, 9, 11, and the optional values of m3 are 128, 256, 512; the output 3 feature tensors are consistent with the feature tensor F p Add and use the activation function to get the attention weight ω p for:
[0063]
[0064] Among them, σ represents the Sigmoid activation function, Conv 1×1×1 represents a point-by-point convolution with a convolution kernel of 1×1×1, L is the number of depth-wise separable convolution blocks, L∈{2,3,4,5}; DW j represents a depthwise separable convolution with the jth banded convolution kernels being k×1×1, 1×k×1, and 1×1×k, k∈{3,5,7,9,11}; in this embodiment, L=3, the banded convolution kernel sizes in the first depthwise separable convolution are 7×1×1, 1×7×1, and 1×1×7, the banded convolution kernel sizes in the second depthwise separable convolution are 9×1×1, 1×9×1, and 1×1×9, and the banded convolution kernel sizes in the third depthwise separable convolution are 11×1×1, 1×11×1, and 1×1×11.
[0065] Depthwise Separable Convolution DW j By transforming the feature tensor F in the channel dimension p Divide into features and Features Among them, feature G p,0 Stay still, feature G p,1 Split equally into features in the channel dimension Among them, s represents the sequence number of the segmentation feature, s∈{1,2,3,4}, and a 3×3×3 convolution kernel is used to check the feature. Convolution is performed, and the remaining features are respectively convolved using k×1×1, 1×k×1, and 1×1×k banded convolution kernels. Perform convolution operation, concatenate the outputs of each branch in the channel dimension to obtain the feature tensor, the depth-separable convolution table DW j The expression is:
[0066]
[0067] Among them, Concat represents feature concatenation, Dw 3×3×3 Denotes a depthwise convolution with a convolution kernel of 3×3×3, Dw k×1×1 Denotes a depthwise convolution with a banded convolution kernel of k×1×1, Dw 1×k×1 Denotes a depthwise convolution with a banded convolution kernel of 1×k×1, Dw 1×1×k Indicates that the banded convolution kernel is a depthwise convolution of 1×1×k;
[0068] 3) Weight ω p Channel by channel and feature tensor F p Multiply them and perform point-by-point convolution with a convolution kernel of 1×1×1 to output a multi-scale feature tensor for:
[0069] F p ′=Conv 1×1×1 (ω p ·F p )
[0070] Among them, Conv 1×1×1 It indicates a point-by-point convolution with a convolution kernel of 1×1×1.
[0071] Taking N=3 as an example, the following is explained: Among the three multi-scale convolutional attention modules, the first multi-scale convolutional attention module is used to extract the shallow feature tensor F output by the first branch and the second branch feature extraction network. 1,1 and F 2,1 A multi-scale convolutional attention mechanism is introduced to focus on the key features of subtle body movements; the second multi-scale convolutional attention module is used to extract the intermediate layer feature tensor F output by the first branch and the second branch feature extraction network. 1,2 and F 2,2 A multi-scale convolutional attention mechanism is introduced to enhance the information interaction between low frame rate features and high frame rate features; the third multi-scale convolutional attention module is used to extract the deep feature tensor F output by the first branch and the second branch feature extraction network. 1,3 and F 2,3 A multi-scale convolutional attention mechanism is introduced to focus on the key features of large-scale limb movements.
[0072] Residual connection module: The input multi-scale feature tensor Fpˋ is combined with the action feature F output by the previous layer of the deep convolutional neural network Net2. 2,p Output residual connection features after residual connection In the 3D residual block of the latter layer of the deep convolutional neural network Net2.
[0073] like Figure 2 The second branch feature extraction network uses N=3 and the deep convolutional neural network Net2 using 3DResNet-50 as an example to illustrate as follows: Input the newborn video segment, the second branch captures the spatial semantic features in static or slowly changing content, uses sparse sampling to reduce the computational cost, and retains the key frame information. The spatiotemporal size of the unprocessed video segment is 32×224×224, where 32 represents the time interval T and 224×224 represents the height and width after cropping. The second branch passes through the frame extraction module and the deep convolutional neural network Net2, and the features output by 3D ResNet-50 at the first three different levels are represented as F 2,p , where p represents the sequence number of the feature level, p∈{1,2,3}. 3D ResNet-50 contains 1 3D convolution layer, 1 maximum pooling layer, and 16 3D residual blocks. The output features of N 3D residual blocks in the deep convolutional neural network Net2 are taken as the action features F extracted at N levels. 2,p , Figure 2 Only three 3D residual blocks are shown in FIG. 3 , and the three 3D residual blocks are represented as residual block A1 to residual block A2. 16 . In this embodiment, the frame extraction frequency M=8, and the frame extraction module extracts one frame every 8 frames to reduce the frame rate of the video, and the time interval T is reduced from the original 32 frames to 4 frames. The number of output channels of the 3D convolution layer is set to 64, and the stride is 1×2×2; the 3D convolution layer uses a convolution kernel of size 1×7×7 to convolve the input feature tensor; the pooling layer uses a maximum pooling layer of 1×3×3 to pool the input feature tensor, and the stride is 1×2×2; the residual block uses 3 different 3D convolution layers, and each residual block uses m1 1×1×1, 1×3×3 and 1×1×1 convolution kernels to convolve the input feature tensor, where the optional values of m1 are 128, 256, and 512, and residual blocks A3, residual blocks A7, and residual blocks A are selected from the 16 residual blocks. 13 Output feature tensor F 2,p They are input into the p-th multi-scale convolutional attention module respectively, and the output F p ′ and F 2,p The residual connection of is taken as output, and the residual connection feature is output The feature tensor scales of the three residual blocks are respectively
[0074] Residual connection module: The input feature Fpˋ is combined with the action feature F output by the previous deep convolutional neural network Net2. 2,p After residual connection, the output is sent to the next deep convolutional neural network Net2.
[0075] In this embodiment, since N=3 is consistent with the number of multi-scale convolutional attention, the second branch feature extraction network connects the residual feature F r,3 As the action feature F extracted at the Nth level 2,N The features input to the feature fusion module; when N>3, the features input to the feature fusion module by the second branch feature extraction network are F 2,N .
[0076] Feature fusion module: used to fuse the newborn limb movement features output by the first branch feature extraction network and the second branch feature extraction network at the Nth level to obtain the fused feature F.
[0077] like Figure 4 , the feature fusion module includes the SE attention module and the efficient feature fusion block:
[0078] SE attention module: The outputs of the first feature extraction network and the second branch feature extraction at the Nth level are squeezed and excited to obtain the enhanced first action feature tensor and the enhanced second action feature tensor. The convolution layer inside the SE attention module uses a 3×3×3 convolution kernel for convolution, and the number of channels in the first convolution layer becomes the original number of channels. After the second convolutional layer, the number of channels is restored to its original value.
[0079] Efficient feature fusion block: After adaptive pooling and convolution processing on the enhanced first action feature tensor and the enhanced second action feature tensor, the convolution first action feature and the convolution second action feature are obtained. The activation function Sigmoid is used to obtain the weights, and the weights are multiplied channel by channel with the enhanced action feature tensor and the enhanced second action tensor to obtain the first intermediate action feature and the second intermediate action feature, which are then concatenated to obtain the fused feature F.
[0080] like Figure 4 The efficient feature fusion block includes an adaptive average pooling layer, a convolution layer, and an activation function Sigmoid, which are used for multi-scale feature extraction, smoothing, and enhancement tasks, respectively. The efficient feature fusion block performs an adaptive pooling operation to compress the spatial dimensions of the output features, using a convolution kernel size of C. k The convolution layer performs convolution processing on the obtained features, where the size of the variable convolution kernel is C kThe convolution kernel size can be dynamically calculated according to the number of channels t of the feature tensor, and its expression is:
[0081]
[0082] The Sigmoid activation function is used to obtain the weights, and the weights are multiplied by the original feature tensor channel by channel, the obtained features are concatenated, and the fused features are output. Among them, H N Represents the height of the feature map, W N Represents the width of the feature map, C 1,N and C 2,N Indicates the number of channels, In this embodiment, the number of channels t includes the number of channels t1 of the first branch and the number of channels t2 of the second branch, wherein t1=2048 and t2=256.
[0083] Classifier: classifies and identifies the action feature vector V and outputs the limb action category. The classifier includes a pooling block and a fully connected layer, where the pooling block: performs maximum pooling and average pooling operations on the fused feature F to obtain the pooled feature, and the pooled feature is stretched and flattened to output the feature vector V; the fully connected layer: maps the feature vector V to the probability distribution of the corresponding predicted category and outputs the limb action category.
[0084] S3. Use the newborn limb movement video sample set obtained in step S1 to train the newborn limb movement recognition model based on dual-branch multi-scale feature fusion constructed in step S2 to obtain a newborn limb movement recognition model.
[0085] In step S3, when training the newborn limb movement recognition model based on dual-branch multi-scale feature fusion constructed in step S2, the focal loss function is used to optimize the training:
[0086]
[0087] Among them, loss i,c represents the binary cross entropy loss, FL represents the focal loss, K represents the total number of categories, c∈[1,K], y i,c Indicates the true value that the i-th sample belongs to the c-th category, i∈[1,n], n represents the total number of samples, Indicates the model prediction value that the i-th sample belongs to the c-th category, α indicates the balance factor, which is used to adjust the weights of positive and negative samples, and β indicates the adjustment factor, which is used to make the weights of difficult-to-classify labels larger. The balance factor α can be 0 to 2, and the adjustment factor β can be 0 to 5. In this embodiment, α = 0.75 and β = 2.
[0088] In step S3, the classifier uses the focal loss function to optimize the training of the model. The output of the model is converted into the predicted probability of each category through the Sigmoid activation function. The predicted probability is converted into a binary label according to the preset threshold 0.5. The category with probability greater than the threshold is marked as 1, otherwise it is marked as 0, so that it can be directly compatible with the calculation formula of the focal loss function, effectively deal with the problem of category imbalance, and increase the attention to difficult-to-classify samples.
[0089] S4. Use the newborn limb movement recognition model obtained in step S3 to perform movement recognition on the newly input test video.
[0090] This method of neonatal limb movement recognition based on dual-branch multi-scale feature fusion uses a dual-branch feature extraction module to extract movement features from videos with different frame rates, improving the ability to capture subtle movement changes, and adopts a multi-scale convolutional attention module to focus on the key features of limb movements, effectively improving the accuracy of neonatal limb movement recognition and its robustness to complex scenes, with less computational effort.
[0091] This method of neonatal limb movement recognition based on dual-branch multi-scale feature fusion introduces a multi-scale convolutional attention mechanism, integrates the feature extraction capabilities of different receptive fields, and combines the weighted strategy of global and local information to achieve accurate capture of multi-scale features, effectively improving the model's ability to recognize detailed features in complex scenes, and using deep separable convolution instead of traditional convolution to effectively reduce the computational complexity of the model. In the task of neonatal limb movement recognition, the multi-scale convolutional attention mechanism can further improve the accuracy and effectiveness of the model.
[0092] This method for recognizing newborn limb movements based on dual-branch multi-scale feature fusion can dynamically adjust the weights of feature channels by fusing dual-branch movement features, effectively highlight the expressiveness of key feature areas, and enable the model to focus on the key features of limb movements. It can dynamically adapt to movements at different time scales, accurately capture subtle movement differences, and effectively improve the accuracy and robustness of movement recognition.
[0093] The above description is only a specific implementation of the present invention, but the protection scope of the present invention is not limited thereto. Any person familiar with the technology can understand and think of any changes or substitutions within the technical scope disclosed by the present invention, which should be included in the scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. A method for identifying newborn limb movements based on dual-branch multi-scale feature fusion, characterized in that: The following steps are included: S1. Collect videos of newborns in different limb movement states, segment the videos according to the time interval T to obtain video segments, and use the first frame of each video segment as the key frame, detect the key limb parts of the newborn in the key frame, including the head, upper limbs and lower limbs, generate corresponding bounding boxes, mark the newborn limb movement labels, integrate the video segments and labels, and construct a newborn limb movement video sample set; S2. Construct a newborn limb movement recognition model based on dual-branch multi-scale feature fusion. The newborn limb movement recognition model based on dual-branch multi-scale feature fusion includes a first branch feature extraction network, a second branch feature extraction network, a feature fusion module and a classifier, wherein: The first branch feature extraction network: The deep convolutional neural network Net1 is used to extract action features F at N levels from the input video segment. 1,p , where p = 1, 2, ..., N, and the action features F extracted at the Nth level 1,N Output to the feature fusion module; where N is an integer ranging from 3 to 5; The second branch feature extraction network: The input video segment is frame-sampled at intervals to obtain a low frame rate video. The low frame rate video is extracted by the deep convolutional neural network Net2 at N levels to extract action features F 2,p , and the action features F extracted from the first three levels 2,p With action feature F 1,p The features obtained by introducing the multi-scale convolutional attention mechanism are then combined with the action features F 2,p After residual connection, we get the residual connection feature F r,p , as the input feature of the next 3D residual block in the deep convolutional neural network Net2, the action feature F extracted at the Nth level 2,N Output to the feature fusion module; Feature fusion module: used to fuse the newborn limb movement features output by the first branch feature extraction network and the second branch feature extraction network at the Nth level to obtain the fused feature F; Classifier: classifies and identifies the fused feature F and outputs the body movement category; S3, using the newborn limb movement video sample set obtained in step S1 to train the newborn limb movement recognition model based on dual-branch multi-scale feature fusion constructed in step S2, to obtain a newborn limb movement recognition model; S4. Use the newborn limb movement recognition model obtained in step S3 to perform movement recognition on the newly input test video.
2. The method for identifying newborn limb movements based on dual-branch multi-scale feature fusion as claimed in claim 1, characterized in that: In step S1, the newborn's limb movement labels include head shaking, fingers opening, fist clenching, arm waving, leg kicking and standing still.
3. The method for identifying newborn limb movements based on dual-branch multi-scale feature fusion as claimed in claim 1, characterized in that: In step S1, the time interval T ranges from 8 frames to 64 frames.
4. The method for identifying neonatal limb movements based on dual-branch multi-scale feature fusion according to any one of claims 1 to 3, characterized in that: In step S2, the second branch feature extraction network includes a frame extraction module, a deep convolutional neural network Net2, three multi-scale convolutional attention modules and three residual connection modules. Frame extraction module: extracts frames from the input video segment at intervals to obtain low frame rate video output to the deep convolutional neural network Net2; Deep convolutional neural network Net2: Extract action features F at N levels for low frame rate videos 2,p ; Multi-scale convolutional attention module: respectively input action features F 1,p and action feature F 2,p Correspondingly, a multi-scale convolutional attention mechanism is introduced and a multi-scale feature tensor Fp` is output; Residual connection module: The input multi-scale feature tensor Fp` is combined with the action feature F output by the previous layer of the deep convolutional neural network Net2. 2,p After the residual connection, the output is sent to the 3D residual block of the next layer of the deep convolutional neural network Net2.
5. The method for identifying newborn limb movements based on dual-branch multi-scale feature fusion as claimed in claim 4, characterized in that: In the multi-scale convolutional attention module, the input action features F 1,p and action feature F 2,p Correspondingly, a multi-scale convolutional attention mechanism is introduced and a multi-scale feature tensor Fp` is output, specifically, 1) The action feature F 1,p and action feature F 2,p Concatenate in the channel dimension to obtain the feature tensor Fp; 2) Use the multi-scale convolutional attention mechanism to obtain the attention weight ω of the feature tensor Fp p for: Among them, σ represents the Sigmoid activation function, Conv 1×1×1 represents a point-by-point convolution with a convolution kernel of 1×1×1, L is the number of depth-wise separable convolutions, L∈{2,3,4,5}; DW j denotes the jth banded convolution kernel as k×1×1, 1×k×1 and 1×1×k depthwise separable convolution, k∈{3,5,7,9,11}; 3) The attention weight ω p Channel by channel and feature tensor F p Multiply them together and perform point-by-point convolution with a convolution kernel of 1×1×1 to output a multi-scale feature tensor F p 'for: F p ′=Conv 1×1×1 (ω p ·F p ) Among them, Conv 1×1×1 It indicates a point-by-point convolution with a convolution kernel of 1×1×1.
6. The method for identifying neonatal limb movements based on dual-branch multi-scale feature fusion according to any one of claims 1 to 3, characterized in that: The feature fusion module includes the SE attention module and the efficient feature fusion block. SE attention module: a squeeze-excitation operation is applied to the outputs of the first feature extraction network and the second branch feature extraction at the Nth level to obtain the enhanced first action feature tensor and the enhanced second action feature tensor; Efficient feature fusion block: After adaptive pooling and convolution processing on the enhanced first action feature tensor and the enhanced second action feature tensor, the convolution first action feature and the convolution second action feature are obtained. The activation function Sigmoid is used to obtain the weights, and the weights are multiplied channel by channel with the enhanced action feature tensor and the enhanced second action tensor to obtain the first intermediate action feature and the second intermediate action feature, which are then concatenated to obtain the fused feature F.
7. The method for identifying neonatal limb movements based on dual-branch multi-scale feature fusion according to any one of claims 1 to 3, characterized in that: The classifier consists of a pooling block and a fully connected layer. Pooling block: The fused feature F is subjected to maximum pooling and average pooling operations to obtain the pooled feature. The pooled feature is stretched and flattened to output the feature vector V. Fully connected layer: maps the feature vector V to the probability distribution of the corresponding prediction category and outputs the body movement category.
8. The method for identifying neonatal limb movements based on dual-branch multi-scale feature fusion according to any one of claims 1 to 3, characterized in that: In step S3, when training the newborn limb movement recognition model based on dual-branch multi-scale feature fusion constructed in step S2, the focal loss function is used to optimize the training: Among them, loss i,c represents the binary cross entropy loss, FL represents the focal loss, K represents the total number of categories, c∈[1,K], y i,c Indicates the true value that the i-th sample belongs to the c-th category, i∈[1,n], n represents the total number of samples, It indicates the model prediction value that the i-th sample belongs to the c-th category, α is the balancing factor, which is used to adjust the weights of positive and negative samples, and β is the adjustment factor, which is used to make the weight of difficult-to-classify labels larger.
Citation Information
Patent Citations
Video action recognition method based on high and low frequency double branches
CN116434343A
Newborn limb movement monitoring method based on multi-task classification network
CN116486320A
Low-frame-rate video stream multi-target tracking method and model training method
CN118261939A
Live face detection system applying two-branch three-dimensional convolutional model, terminal and storage medium
WO2021248733A1
Cited By
Artificial limb control method based on motion intention recognition
CN122251164A