Racing track multi-target real-time segmentation method and system based on deep learning
Through the deep learning method of dual-branch spatiotemporal feature extraction and adaptive attention mechanism, the problem of goal segmentation in marathon event scenarios is solved, real-time and accurate segmentation of multiple goals is achieved, and feature expression and segmentation accuracy is improved.
Patent Information
- Application Number
- CN202510425411.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-07
AI Technical Summary
The prior art goal segmentation method in marathon event scenarios is difficult to effectively utilize timing information, lacks feature expression ability, cannot accurately distinguish targets with similar visual features, lacks adaptive attention mechanisms, and does not design a special goal classification strategy for event scenarios, resulting in unsatisfactory segmentation accuracy.
Using a dual-branch spatiotemporal feature extraction and adaptive attention mechanism based on deep learning, spatial and temporal features are extracted through a dual-branch architecture, combined with a multi-level attention mechanism and a multi-scale convolutional layer, a dynamic target classification branch network is designed to achieve real-time segmentation of multiple goals.
Real-time and accurate segmentation of multiple goals in marathon event scenarios is achieved, feature expression ability is improved, and it can adaptively pay attention to important features, suppress background interference, accurately capture multi-scale features, and improve classification and segmentation accuracy.
Smart Images

Figure CN120279466A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multi-target real-time segmentation of race tracks, and in particular, to a method and system for multi-target real-time segmentation of race tracks based on deep learning. Background Art
[0002] With the rapid development of sports event live broadcast and analysis technologies, higher and higher requirements are put forward for target segmentation in large-scale event scenarios such as marathons. Accurately identifying and segmenting the participating athletes and staff in the race track scene is of great significance for applications such as event live broadcast, security monitoring, and trajectory analysis. At present, computer vision and deep learning technologies have made remarkable progress in the field of target segmentation, but still face many challenges when dealing with special scenarios such as marathon events.
[0003] The existing technologies mainly adopt semantic segmentation methods based on deep learning. For example, Patent CN118941796A discloses a deep asymmetric bottleneck real-time semantic segmentation method that fuses multi-scale information. This method realizes the segmentation of targets at different scales through a self-correction DAB module and a CCSP-ASPP module. Although this method has achieved certain results in static image segmentation, it mainly has the following deficiencies: First, this method only considers the spatial features of a single-frame image and ignores the rich temporal information in the video sequence, resulting in limited performance when dealing with moving targets; Second, the adopted feature extraction method is relatively simple, and it fails to fully consider the target motion characteristics and scene complexity, making it difficult to handle complex situations such as dense crowds and occlusion overlaps in marathon events.
[0004] In addition, there are also some common problems in other existing segmentation methods: One is the insufficient feature expression ability, which is difficult to accurately distinguish different category targets with similar visual features, such as participating athletes and staff; The second is the lack of an effective attention mechanism, which cannot adaptively highlight important features and suppress background interference; The third is that no special target classification strategy is designed for the characteristics of the event scene, resulting in unsatisfactory classification accuracy. These problems seriously restrict the practical application of segmentation technologies in marathon event scenarios.
[0005] Therefore, there is an urgent need to develop a target segmentation method that can effectively utilize temporal information, has a strong feature expression ability, and adapts to the characteristics of the event scene. The present invention precisely aims at the above technical problems and proposes a target segmentation method based on dual-branch spatio-temporal feature extraction and an adaptive attention mechanism to achieve real-time and accurate segmentation of multiple targets in marathon event scenarios. Summary of the Invention
[0006] In view of this, the present invention provides a method for multi-target real-time segmentation of race tracks based on deep learning, aiming to achieve real-time and accurate segmentation of multiple targets in marathon event scenarios and provide effective technical support for event video analysis and processing.
[0007] To achieve the above object, a multi-object real-time segmentation method for a race track based on deep learning provided by the present invention includes the following steps:
[0008] S1: Obtain the race track video data stream and perform dual-branch spatio-temporal feature extraction and fusion to obtain the fused image;
[0009] S2: Enhance the fused image based on the adaptive attention mechanism to obtain the enhanced image;
[0010] S3: Use a multi-scale convolutional layer to perform feature parsing and recombination on the enhanced feature map, construct a multi-scale feature representation, and obtain a multi-scale feature map;
[0011] S4: Based on the multi-scale feature map, design a dynamic target classification branch network to achieve the recognition of different category targets and obtain a feature map with category information;
[0012] S5: Use the target feature map with category information to complete the real-time segmentation of multiple targets in the race track scene.
[0013] As a further improvement method of the present invention:
[0014] Optionally, in the step S1 of obtaining the race track video data stream and performing dual-branch spatio-temporal feature extraction and fusion to obtain the fused image, it includes:
[0015] S11: Obtain a single-frame image and an optical flow image sequence in the race track video data stream, specifically:
[0016] Obtain a single-frame image I t in the race track video data stream at time t, and at the same time obtain an optical flow image sequence O = {O t-r , …, O t , …, O t+r} including a total of T frames before and after this frame, where:
[0017]
[0018] O t = Φ(I t , I t+1 ), t ∈ [t - r, t + r];
[0019] Among them, t is the time index of the current frame; is the floor operation; T is the size of the time sequence window, and the value is an odd number; r is the time sequence radius; I t and I t+1 are the images at times t and t + 1 respectively; O t is the optical flow image at time t; Φ is the optical flow estimation function;
[0020] S12: Perform spatial feature extraction on the input single-frame image I t Specifically as follows:
[0021]
[0022] Among them, F spt (x, y) is the feature value of the spatial feature map F spt at the pixel coordinates (x, y); W spt (P, q) is the value of the spatial convolution kernel weight matrix W spt at the row and column numbers (p, q) of the weight matrix; I t (x + p, y + q) is the pixel value of the single-frame image I at time t t at the pixel coordinates (x + p, y + q);
[0023] S13: Perform temporal feature extraction on the sequence of consecutive T-frame optical flow images, specifically as follows:
[0024]
[0025] Among them, F tmp (x, y) is the feature value of the temporal feature map F tmp at the pixel coordinates (x, y); W tmp (m, n) is the value of the temporal convolution kernel weight matrix W tmp at the row and column numbers (m, n) of the weight matrix; O k’ (x + m, y + n) is the value of the k'-th frame optical flow image at the pixel coordinates (x + m, y + n); γ is the normalization function; D k’ (x, y) is the motion displacement vector of the k'-th frame at the pixel coordinates (x, y);
[0026] S14: Fuse the spatial feature and the temporal feature to obtain the fused image H, specifically as follows:
[0027] H(x, y) = F spt (x, y) ⊙ σ(F tmp (x, y) + F tmp (x, y) ⊙ σ(F spt (x, y)));
[0028] Among them, H(x, y) is the feature value of the fused image H at the pixel coordinates (x, y); σ is the sigmoid function; ⊙ is element-wise multiplication.
[0029] Optionally, in the step S2, enhancing the fused image based on the adaptive attention mechanism to obtain the enhanced image includes:
[0030] S21: Construct a channel attention module to calculate the importance weights of each channel of the fused image, specifically:
[0031]
[0032] where H c (x, y) is the pixel value of the c-th channel of the fused image H at the pixel coordinates (x, y); Z c is the global average pooling value of the c-th channel; and are the width and length of the fused image H respectively; ω c is the attention weight of the c-th channel; N is the number of channels; e is the natural constant; Z d is the global average pooling value of the d-th channel, where d is the dimension index and takes positive integer values from 1 to N;
[0033] S22: Construct a spatial attention module to calculate the attention map of each spatial position of the fused image H, specifically:
[0034]
[0035] where Q(x, y) is the gradient response intensity at the pixel coordinates (x, y); is the spatial gradient of the c-th channel of the fused image H at the pixel coordinates (x, y); R(x, y) is the spatial attention value at the pixel coordinates (x, y);
[0036] S23: Fuse the channel attention and spatial attention to obtain the enhanced image E, specifically:
[0037]
[0038] where E c (x, y) is the pixel value of the c-th channel of the enhanced image E at the pixel coordinates (x, y); σ′ is the ReLU function; is the square root of N, and is normalized by dividing by to avoid numerical instability problems.
[0039] Optionally, in the S3 step, a multi-scale convolutional layer is used to perform feature parsing and recombination on the enhanced image, construct a multi-scale feature representation, and obtain a multi-scale feature map, including:
[0040] S31: Generate decomposition maps of different scales, specifically:
[0041] P k = Conv k (E), k ∈ {1, 2, 3};
[0042]
[0043] Among them, P k is the decomposition graph of scale k, and P k (x, y) is the pixel value of the decomposition graph P of scale k at pixel coordinates (x, y); Conv k is the convolution operation of scale k; k and and are the width and length of the decomposition graph of scale k respectively; α k is the global representation of the decomposition graph of scale k;
[0044] S32: Generate multi-scale feature maps, specifically:
[0045]
[0046] Among them, is the feature map of scale k; Conv re is the reconstruction convolution operation; Res k (E) is the feature map obtained by performing a residual connection operation on the enhanced image E at scale k.
[0047] Optionally, in step S4, based on the multi-scale feature maps, design a dynamic target classification branch network to realize the recognition of different category targets and obtain a feature map with category information, including:
[0048] S41: Construct a category-aware attention module to extract category-related features, specifically:
[0049]
[0050]
[0051]
[0052] Among them, μ k is the feature of the feature map of scale k after being convolved by the category-aware convolution Conv cls ; is the category feature representation of the -th category target, which is extracted by the category-specific convolution ; Pool is the global pooling operation; is the attention weight of the -th category target; is the number of target categories; and are respectively the category features and The feature vector obtained by performing a fully connected layer transformation can map inputs in any range to the interval from 0 to 1 through a combined operation of an exponential function and normalization, while maintaining the relative importance between different classes;
[0053] S42: Output the target feature map with class information, specifically:
[0054]
[0055] where, is the final feature representation of the -th class target at pixel coordinates (x, y); Conv map is the feature map convolution; Ω(x, y) is the feature set containing all class information at pixel coordinates (x, y).
[0056] Optionally, in the S5 step, the target feature map with class information is used to complete the real-time segmentation of multiple targets in the track scene, including:
[0057] S51: Construct a feature enhancement module to fuse class information and location information, specifically:
[0058] Λ(x, y) = Conv enhance (Ω(x, y)) + PE(x, y);
[0059] where, Λ(x, y) is the fused feature at pixel coordinates (x, y); Conv enhance is the enhancement convolution; PE(x, y) is the position encoding of pixel coordinates (x, y);
[0060] S52: Generate a mask prediction to obtain the segmentation result, specifically:
[0061]
[0062] where, ξ(x, y) is the instance mask prediction at pixel coordinates (x, y), is the probability that the pixel at predicted pixel coordinates (x, y) belongs to the -th class target; Conv mask is the mask prediction convolution;
[0063] The segmentation result is:
[0064]
[0065] where, τ(x, y) is the class label at pixel coordinates (x, y).
[0066] The present invention also discloses a track multi-target real-time segmentation system based on deep learning, including:
[0067] Fusion module: Obtain the track video data stream and perform dual-branch spatio-temporal feature extraction and fusion;
[0068] Enhancement module: Enhance the fused image based on the adaptive attention mechanism;
[0069] Multi-scale module: Use multi-scale convolutional layers to perform feature parsing and recombination on the enhanced feature map, and construct a multi-scale feature representation;
[0070] Target feature extraction module: Based on the multi-scale feature map, design a dynamic target classification branch network to achieve the recognition of different categories of targets;
[0071] Segmentation module: Use the target feature map with category information to complete the real-time segmentation of multiple targets in the track scene.
[0072] Compared with the prior art, the present invention has at least the following beneficial effects:
[0073] The present invention proposes a target segmentation method based on dual-branch spatio-temporal feature extraction and adaptive attention mechanism. By designing a dual-branch architecture to extract spatial features and temporal features respectively, and adopting an innovative feature fusion strategy, accurate modeling of dynamic targets is achieved. This method can make full use of temporal information when processing video data.
[0074] The present invention innovatively designs a multi-level attention mechanism, including a channel attention module and a spatial attention module, and proposes an adaptive feature enhancement strategy. This multi-level attention mechanism can adaptively focus on important features, effectively suppress background interference, and improve the discriminability of feature expression.
[0075] The present invention adopts an innovative architecture of multi-scale feature representation and dynamic target classification branch network. By using multi-scale convolutional layers to parse and recombine features, the detection problem of targets with different scales is effectively solved. This method can accurately capture the multi-scale features of targets, and extract category-related features through the category-aware attention module, and finally achieve accurate target classification and segmentation. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] Figure 1 It is a schematic flowchart of a method for real-time multi-target segmentation of a track based on deep learning according to an embodiment of the present invention;
[0077] Figure 2 It is a multi-target segmentation result diagram according to an embodiment of the present invention, (a) the image to be segmented; (b) the schematic diagram of the segmentation result;
[0078] Figure 3 It is a deep learning network training diagram according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0079] The present invention will be further described below in conjunction with the accompanying drawings, but the present invention is not limited in any way. Any transformation or replacement based on the teachings of the present invention falls within the protection scope of the present invention.
[0080] Embodiment 1: A method for real-time multi-object segmentation of a track based on deep learning, as Figure 1 shown, includes the following steps:
[0081] S1: Obtain the track video data stream and perform dual-branch spatio-temporal feature extraction and fusion to obtain the fused image:
[0082] S11: Obtain a single-frame image and an optical flow image sequence in the track video data stream, specifically:
[0083] Obtain a single-frame image I t at time t in the track video data stream, and at the same time obtain an optical flow image sequence O = {O t-r , …, O t , …, O t+r} containing a total of T frames before and after this frame, where:
[0084]
[0085] O t = Φ(I t , I t+1 ), t ∈ [t - r, t + r];
[0086] Among them, t is the time index of the current frame; is the floor operation; T is the size of the time sequence window, which is an odd number and is 7 in this embodiment; r is the time sequence radius; I t and I t+1 are the images at times t and t + 1 respectively; O t is the optical flow image at time t; Φ is the optical flow estimation function;
[0087] S12: Perform spatial feature extraction on the input single-frame image I t , specifically:
[0088]
[0089] Among them, F spt (x, y) is the feature value of the spatial feature map F spt at the pixel coordinates (x, y); W spt (p, q) is the value of the spatial convolution kernel weight matrix W spt at the row and column numbers (p, q) of the weight matrix; I t (x + p, y + q) is the single-frame image I tThe pixel value at pixel coordinates (x + p, y + q); the spatial convolution kernel weight matrix W spt Specifically:
[0090]
[0091] Where e is the natural constant; λ spt Is the spatial scale parameter, which is 0.9 in this embodiment; η spt Is the direction enhancement coefficient, which is 0.5 in this embodiment; θ spt (p, q) is the local direction angle at the position (p, q), specifically:
[0092]
[0093] Where G x (x + p, y + q) and G y (x + p, y + q) are the horizontal and vertical direction gradients of the single-frame image I at time t t At pixel coordinates (x + p, y + q);
[0094] S13: Perform temporal feature extraction on the sequence of optical flow images for continuous T frames, specifically:
[0095]
[0096] Where F tmp (x, y) is the eigenvalue of the temporal feature map F tmp At pixel coordinates (x, y); W tmp (m, n) is the value of the temporal convolution kernel weight matrix W tmp At the row and column numbers (m, n) of the weight matrix; O k’ (x + m, y + n) is the value of the k'-th frame optical flow image at pixel coordinates (x + m, y + n); γ is the normalization function; D k’ (x, y) is the motion displacement vector of the k'-th frame at pixel coordinates (x, y); the temporal convolution kernel weight matrix W tmp Specifically:
[0097]
[0098] Where λ tmp Is the temporal scale parameter, which is 1.5 in this embodiment; η tmp Is the motion enhancement coefficient, which is 0.3 in this embodiment; M k′ (m, n) is the motion amplitude of the k'-th frame at the position (m, n), specifically:
[0099]
[0100] Where and are the gradients of the optical flow image of the k'-th frame in the m and n directions, respectively;
[0101] S14: Fuse the spatial features and temporal features to obtain the fused image H, specifically:
[0102] H(x,y) = F spt (x,y) ⊙ σ(F tmp (x,y) + F tmp (x,y) ⊙ σ(F spt (x,y)));
[0103] where H(x,y) is the feature value of the fused image H at the pixel coordinates (x,y); σ is the sigmoid function; ⊙ is element-wise multiplication.
[0104] In this step, by designing a dual-branch spatio-temporal feature extraction and fusion architecture, the key problems of object detection and segmentation in the track scene are effectively solved. In the spatial feature extraction branch, a direction-aware adaptive convolution kernel is used for feature extraction, which can highlight the key edge and contour information in the track scene; in the temporal feature extraction branch, by modeling the motion information of multiple-frame optical flow sequences, the temporal dynamic features of moving objects in the scene can be effectively captured.
[0105] S2: Enhance the fused image based on the adaptive attention mechanism to obtain the enhanced image:
[0106] S21: Construct a channel attention module to calculate the importance weights of each channel of the fused image, specifically:
[0107]
[0108]
[0109] where H c (x,y) is the pixel value of the c-th channel of the fused image H at the pixel coordinates (x,y); Z c is the global average pooling value of the c-th channel; and are the width and length of the fused image H, respectively; ω c is the attention weight of the c-th channel; N is the number of channels, which is 3 in this embodiment; Z d is the global average pooling value of the d-th channel, where d is the dimension index, and the value range is a positive integer from 1 to N;
[0110] S22: Construct a spatial attention module to calculate the attention map of each spatial position of the fused image H, specifically:
[0111]
[0112] Among them, Q(x, y) is the gradient response intensity at the pixel coordinates x(, y); is the spatial gradient of the c-th channel of the fused image H at the pixel coordinates (x, y), which is calculated using the Sobel operator in this embodiment; R(x, y) is the spatial attention value at the pixel coordinates (x, y);
[0113] S23: Fuse the channel attention and spatial attention to obtain the enhanced image E, specifically:
[0114]
[0115] where E c (x, y) is the pixel value of the c-th channel of the enhanced image E at the pixel coordinates (x, y); σ′ is the ReLU function; is the square root of N, and by dividing by for normalization processing, it can avoid the problem of numerical instability.
[0116] In this step, the features are adaptively enhanced by designing a dual attention mechanism, effectively improving the expression ability and discriminability of the features. In the channel dimension, the importance weights of each channel are calculated through global average pooling, enabling the system to automatically identify and strengthen the feature channels that are more important for the target detection and segmentation tasks, while suppressing the interference of irrelevant channels. In the spatial dimension, the spatial attention map is constructed by calculating the gradient response intensity, enabling the system to highlight the feature expressions in the key regions.
[0117] S3: Use a multi-scale convolutional layer to perform feature parsing and recombination on the enhanced feature map, construct a multi-scale feature representation, and obtain a multi-scale feature map:
[0118] S31: Generate decomposition maps of different scales, specifically:
[0119] P k = Conv k (E), k ∈ {1, 2, 3};
[0120]
[0121] where P k is the decomposition map of scale k, and P k (x, y) is the pixel value of the decomposition map P of scale k at the pixel coordinates (x, y); Conv k at the pixel coordinates (x, y); Conv kis a convolution operation with scale k. When k = 1, a 3×3 convolution kernel is used, and 128 channels are output. When k = 2, a 5×5 convolution kernel is used, and 128 channels are output. When k = 3, a 7×7 convolution kernel is used, and 128 channels are output. All Conv k operations use a stride of 1; and are the width and length of the decomposition graph with scale k, respectively; α k is the global representation of the decomposition graph with scale k;
[0122] S32: Generate multi-scale feature maps, specifically:
[0123]
[0124] Among them, is the feature map with scale k; Conv re is a reconstruction convolution operation. In this embodiment, Conv re is a 1×1 convolution operation, with the number of output channels being 3 and the stride being 1; Res k is the residual connection operation at scale k. In this embodiment, Res k consists of two cascaded 1×1 convolutional layers. The number of output channels of the first 1×1 convolutional layer is 64, and the number of output channels of the second 1×1 convolutional layer is 3. A ReLU activation function is added between the two convolutional layers, and the strides are both 1; Res k (E) is the feature map obtained by performing a residual connection operation on the enhanced image E at scale k.
[0125] In this step, by designing a multi-scale convolutional layer network structure, multi-level feature extraction and effective fusion of feature maps are achieved. By using convolution kernels of different sizes for feature decomposition, the system can simultaneously capture local detail features and large-scale context information in the marathon track scene. Small-sized convolution kernels focus on extracting fine textures and edge features of the target, medium-sized convolution kernels can capture local structural information of the target, and large-sized convolution kernels help to understand the overall shape of the target and the context relationship of the scene.
[0126] S4: Based on the multi-scale feature maps, design a dynamic target classification branch network to achieve the recognition of different category targets and obtain feature maps with category information:
[0127] S41: Construct a category-aware attention module to extract category-related features, specifically:
[0128]
[0129] Among them, μ k is the feature map with scale k after being processed by the category-aware convolution Convcls After the feature, in this embodiment, Conv cls is a 3×3 convolution, the number of output channels is 128, and the stride is 1; is the class feature representation of the th class of target, extracted by class-specific convolution In this embodiment, is a 1×1 convolution, the number of output channels is 64, and the stride is 1; Pool is a global pooling operation; is the th class of target's attention weight; is the number of target classes. In this embodiment, it is 3, corresponding to background, participating athletes, and staff respectively; FC is a fully connected layer. In this embodiment, the input dimension is 64 and the output dimension is 1; and are the feature vectors obtained by performing fully connected layer transformation on the class features and respectively. Through the combined operation of the exponential function and normalization, the input in any range can be mapped to the interval from 0 to 1 while maintaining the relative importance between different classes;
[0130] S42: Output the target feature map with class information, specifically:
[0131]
[0132] Among them, is the final feature representation of the th class of target at the pixel coordinates (x, y); Conv map is the feature mapping convolution. In this embodiment, Conv map is the feature mapping convolution, adopting a 1×1 convolution structure, the number of output channels is 32, and the stride is 1; Ω(x, y) is the feature set containing all class information at the pixel coordinates (x, y).
[0133] In this step, by designing a dynamic target classification branch network, accurate recognition of multiple types of dynamic targets in the marathon event scenario is achieved. The introduced class-aware attention module can adaptively extract and enhance the feature representations related to specific target classes, effectively distinguishing different types of targets such as participating athletes, staff, and spectators. Through the fusion of multi-scale features and class-specific convolution operations, this step can simultaneously process the feature expressions of targets at different scales, capturing both the overall features of large-scale targets and the detailed information of small-scale targets.
[0134] S5: Use the target feature map with class information to complete the real-time segmentation of multiple targets in the track scene:
[0135] S51: Construct a feature enhancement module to fuse category information and location information, specifically:
[0136] Λ(x,y) = Conv enhance (Ω(x,y)) + PE(x,y);
[0137] where Λ(x,y) is the fused feature at pixel coordinates (x,y); Conv enhance is the enhanced convolution. In this embodiment, Conv enhance adopts a two-layer 3×3 convolution structure. The number of output channels of the first layer is 256, and the number of output channels is 128. Each layer is followed by BatchNorm and ReLU activation functions, and the stride is 1 for both layers; PE(x,y) is the position encoding of pixel coordinates (x,y). In this embodiment, the sine-cosine position encoding method is adopted, and the encoding dimension is 128;
[0138] S52: Generate a mask prediction to obtain the segmentation result, specifically:
[0139] ξ(x,y)) = Softmax(Conv mask (Λ(x,y)));
[0140] where ξ(x,y) is the instance mask prediction at pixel coordinates (x,y), is the probability that the pixel at the predicted pixel coordinates (x,y) belongs to the th class target; Conv mask is the mask prediction convolution. In this embodiment, a 1×1 convolution structure is adopted, and the number of output channels is with a stride of 1;
[0141] The segmentation result is:
[0142]
[0143] where τ(x,y) is the class label at pixel coordinates (x,y);
[0144] In this embodiment, the network training is carried out in an end-to-end manner. The training parameters include all convolution parameters in steps S3 to S5. During the training process, the cross-entropy loss function is used to calculate the difference between the predicted mask and the true label, and the L2 regularization is used to constrain the model parameters. The Adam optimizer is used for parameter optimization. The initial learning rate is 0.001, and the learning rate is reduced to 0.1 of the original every 50 training rounds. The batch size is set to 16, and the total number of training rounds is 200. The training process is as Figure 3 shown.
[0145] To verify the effectiveness of the present invention, a complete experimental plan was designed, and a progressive method was adopted to verify the effectiveness of each module. As shown in Table 1, starting from the baseline model, the spatio-temporal feature fusion module, the attention mechanism, and multi-scale features were gradually added, and finally a complete solution was formed. The experimental results show that the spatio-temporal feature fusion module increased the mAP from 81.2% to 85.6%, which proves that the dual-branch spatio-temporal feature extraction architecture can effectively capture the temporal information of dynamic targets. The introduction of the attention mechanism further improved the performance to 88.9%, indicating that the adaptive attention module can effectively enhance the expression of key features. The integration of multi-scale features brought the performance to 91.5%, demonstrating that multi-scale feature representation can better handle the detection and segmentation of targets at different scales. The mAP of the final complete solution reached 93.2%, and the real-time performance of the system reached the level of 22.3 FPS, fully meeting the actual application requirements.
[0146] Table 1 Comparison of Performance of Different Modules
[0147]
[0148] Regarding the recognition effect of different category targets, as shown in Table 2, the background category obtained the highest F1 score of 96.5%, which benefits from the relative stability of the background region features and the feature enhancement mechanism proposed in the present invention. The F1 score of the participating athlete category reached 94.2%, mainly due to the effective modeling of motion information by the spatio-temporal feature fusion module. The F1 score of the staff reached 90.8%, indicating that the multi-scale feature representation and the dynamic target classification branch network proposed in the present invention can effectively identify and segment this type of target.
[0149] Table 2 Segmentation Accuracy of Each Category
[0150] Target category Accuracy Recall rate F1 score Background 97.1 95.9 96.5 Competitive athletes 94.8 93.6 94.2 Staff members 91.5 90.1 90.8
[0151] Example 2: The present invention also discloses a multi-target real-time segmentation system for a race track based on deep learning, including the following five modules:
[0152] The present invention also discloses a multi-target real-time segmentation system for a race track based on deep learning, including:
[0153] Fusion module: Obtain the race track video data stream and perform dual-branch spatio-temporal feature extraction and fusion;
[0154] Enhancement module: Enhance the fused image based on the adaptive attention mechanism;
[0155] Multi-scale module: Use multi-scale convolutional layers to perform feature parsing and recombination on the enhanced feature map to construct a multi-scale feature representation;
[0156] Target feature extraction module: Based on multi-scale feature maps, a dynamic target classification branch network is designed to achieve the recognition of different types of targets.
[0157] Segmentation module: Use the target feature map with class information to complete the real-time segmentation of multiple targets in the track scene.
[0158] It should be noted that the serial numbers of the above embodiments of the present invention are only for description and do not represent the superiority or inferiority of the embodiments. And the term "including", "comprising" or any other variant thereof in this article is intended to cover non-exclusive inclusion, so that a process, device, article or method including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such a process, device, article or method. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, device, article or method including the element.
[0159] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above and includes several instructions to enable a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present invention.
[0160] The above are only the preferred embodiments of the present invention and do not limit the patent scope of the present invention accordingly. Any equivalent structural or equivalent process transformation made by using the description of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. A real-time multi-object segmentation method for a track based on deep learning, characterized in that, It includes the following steps: S1: Obtain the track video data stream and perform dual-branch spatio-temporal feature extraction and fusion to obtain the fused image; S2: Enhance the fused image based on the adaptive attention mechanism to obtain the enhanced image; S3: Use multi-scale convolutional layers to perform feature parsing and recombination on the enhanced feature map, construct multi-scale feature representations, and obtain multi-scale feature maps; S4: Based on the multi-scale feature maps, design a dynamic target classification branch network to realize the recognition of different category targets, and obtain the feature map with category information: First, perform category-aware convolution operations on the feature maps of each scale to obtain initial features, then perform specific convolution and global pooling operations on each target category respectively to obtain category feature representations, and finally obtain the attention weights of each category through the fully connected layer and softmax operations; Based on the category feature representations and attention weights, output the target feature map with category information: Weight the category attention weights with the category features after feature map convolution to obtain the final feature representation of each category, and combine the feature representations of all categories to form a complete feature set; S5: Use the target feature map with category information to complete the real-time segmentation of multiple targets in the track scene.
2. The method for real-time multi-object segmentation of a track based on deep learning according to claim 1, wherein, The step S1 includes: S11: Obtain a single-frame image and an optical flow image sequence in the track video data stream, specifically: Obtain a single-frame image I from the track video data stream at time t t , and at the same time obtain an optical flow image sequence O = {O t-T , …, O t , …, O t+r} that includes a total of T frames before and after this frame, where: O t = Φ(I t , I t+1 ), t ∈ [t - r, t + r]; where t is the time index of the current frame; is the floor operation; T is the size of the time series window, which takes an odd value; r is the time series radius; I t and I t+1 are the images at times t and t + 1 respectively; O t is the optical flow image at time t; Φ is the optical flow estimation function; S12: Extract spatial features from the input single-frame image I t Specifically, it is as follows: Among them, F spt (x, y) is the feature value of the spatial feature map F spt at the pixel coordinates (x, y); W spt (p, q) is the value of the spatial convolution kernel weight matrix W spt at the row and column numbers (p, q) of the weight matrix; I t (x + p, y + q) is the pixel value of the single-frame image I t at the pixel coordinates (x + p, y + q); S13: Perform temporal feature extraction on the optical flow image sequence of continuous T frames, specifically: Among them, F tmp (x, y) is the eigenvalue of the temporal feature map F tmp at the pixel coordinates (x, y); W tmp (m, n) is the weight matrix of the temporal convolution kernel W tmp at the row and column numbers (m, n) of the weight matrix; O k’ (x + m, y + n) is the value of the k'-th frame optical flow image at the pixel coordinates (x + m, y + n); γ is the normalization function; D k’ (x, y) is the motion displacement vector of the k'-th frame at the pixel coordinates (x, y); S14: Fuse the spatial feature and the temporal feature to obtain the fused image G, specifically: H(x,y) = F spt (x,y) ⊙ σ(F tmp (x,y) + F tmp (x,y) ⊙ σ(F spt (x,y))); where, H(x, y) is the feature value of the fused image H at the pixel coordinates (x, y); σ is the sigmoid function; ⊙ is element-wise multiplication.
3. The method for real-time multi-object segmentation of a track based on deep learning according to claim 2, wherein The step S2 includes: S21: Construct a channel attention module to calculate the importance weights of each channel of the fused image, specifically: Among them, H c (x, y) is the pixel value of the c-th channel of the fused image H at the pixel coordinates (x, y); Z c is the global average pooling value of the c-th channel; and are the width and length of the fused image H respectively; ω c is the attention weight of the c-th channel; N is the number of channels; e is the natural constant; Z d is the global average pooling value of the d-th channel, where d is the dimension index and takes positive integer values ranging from 1 to N; S22: Construct a spatial attention module to calculate the attention map of each spatial position of the fused image H, specifically: where Q(x, y) is the gradient response intensity at the pixel coordinates (x, y); is the spatial gradient of the c-th channel of the fused image H at the pixel coordinates (x, y); R(x, y) is the spatial attention value at the pixel coordinates (x, y); S23: Fuse the channel attention and the spatial attention to obtain the enhanced image E, specifically: where E c (x, y) is the pixel value of the c-th channel of the enhanced image E at the pixel coordinates (x, y); σ′ is the ReLU function; is the square root of N.
4. The multi-object real-time segmentation method for a track based on deep learning according to claim 3, wherein The S3 includes the following steps: S31: Generate decomposition maps of different scales, specifically: P k = Conv k (E), k ∈ {1, 2, 3}; Among them, P k is the decomposition graph of scale k, and P k (x, y) is the pixel value of the decomposition graph P k at the pixel coordinates (x, y); Conv k is the convolution operation of scale k; and are the width and length of the decomposition graph of scale k respectively; α k is the global representation of the decomposition graph of scale k; S32: Generate multi-scale feature maps, specifically: Among them, is the feature map with scale k; Conv re is the reconstruction convolution operation; Res k (E) is the feature map obtained by performing a residual connection operation on the enhanced image E at scale k.
5. The method for real-time multi-object segmentation of a race track based on deep learning according to claim 4, wherein The S4 includes the following steps: S41: Construct a category-aware attention module to extract category-related features, specifically: Among them, μ k is the feature map of scale k after being processed by the class-aware convolution Conv cls ; is the class feature representation of the th class target, which is extracted by the class-specific convolution ; Pool is the global pooling operation; is the attention weight of the th class target; is the number of target classes; and are the feature vectors obtained by performing a fully connected layer transformation on the class features and respectively. Through the combined operation of the exponential function and normalization, the input in any range can be mapped to the interval from 0 to 1 while maintaining the relative importance between different classes; S42: Output the target feature map with class information, specifically: Among them, is the final feature representation of the class target at the pixel coordinates (x, y); Conv map is the feature map convolution; Ω(x, y) is the feature set containing all class information at the pixel coordinates (x, y).
6. The method for real-time multi-object segmentation of a track based on deep learning according to claim 5, wherein The S5 includes the following steps: S51: Construct a feature enhancement module to fuse category information and location information, specifically: Λ(x,y) = Conv enhance (Ω(x,y)) + PE(x,y); Among them, Λ(x, y) is the fused feature at the pixel coordinates (x, y); Conv enhance is the enhanced convolution; PE(x, y) is the position encoding of the pixel coordinates (x, y). S52: Generate mask predictions to obtain the segmentation result, specifically: ξ(x,y) = Softmax(Conv mask (Λ(x,y))); where, ξ(x,y) is the instance mask prediction at the pixel coordinates (x,y), is the probability that the pixel at the predicted pixel coordinates (x,y) belongs to the th class of objects; Conv mask is the mask prediction convolution; The segmentation result is: where, τ(x, y) is the category label at the pixel coordinates (x, y).
7. A multi-object real-time segmentation system for a race track based on deep learning, characterized in that, It includes: Fusion module: Obtain the track video data stream and perform dual-branch spatio-temporal feature extraction and fusion; Enhancement module: Enhance the fused image based on the adaptive attention mechanism; Multi-scale module: Use multi-scale convolutional layers to perform feature parsing and recombination on the enhanced feature map and construct multi-scale feature representations; Target feature extraction module: Based on the multi-scale feature maps, design a dynamic target classification branch network to realize the recognition of different category targets; Segmentation module: Use the target feature map with class information to complete the real-time segmentation of multiple targets in the track scene; To implement a method for real-time segmentation of multiple targets on a track based on deep learning as described in any one of claims 1-6.
Citation Information
Patent Citations
Behavior recognition method based on space-time attention enhancement feature fusion network
CN111709304A
End-to-end video action detection and positioning system
WO2022134655A1