Image segmentation and dynamic target identification method based on artificial intelligence
Through technical means such as multimodal spatiotemporal alignment, dual-stream feature fusion and deformable segmentation network, the data consistency and feature extraction problems in image segmentation and dynamic target recognition are solved, the segmentation accuracy and target recognition stability are improved, and efficient tracking of dynamic targets is achieved.
Patent Information
- Application Number
- CN202510795937.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-09-23
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies in image segmentation and dynamic target recognition have problems such as inconsistency in different modal data, difficulty in capturing changes in dynamic target edge information, and incomplete feature extraction, which lead to inaccurate segmentation results, especially when the target and background are similar or the target moves quickly.
By adopting multimodal spatiotemporal alignment processing, dual-stream feature fusion, Markov random field model and spatiotemporal confidence evaluation, combined with deformable segmentation network and graph attention network, we gradually decouple multi-level features, optimize the segmentation results and complete the target trajectory.
It effectively eliminates the temporal misalignment of modal data, strengthens dynamic edge information, improves segmentation accuracy and target recognition performance, solves occlusion and loss problems, and achieves stable target tracking.
Smart Images

Figure CN120689620A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an image segmentation and dynamic target recognition method based on artificial intelligence. Background Art
[0002] Currently, image segmentation and dynamic target recognition still face some challenges, including the following aspects: data of different modalities may be inconsistent in time and space, which will affect the subsequent feature extraction and target recognition effects; the edge information of dynamic targets will continue to change during movement, and traditional image segmentation methods may not be able to fully capture this dynamic edge information, resulting in inaccurate segmentation results, especially when the target is similar to the background or the target moves quickly; in image segmentation and target recognition, features at different levels have different functions, and traditional feature extraction methods may only focus on features at a certain level, resulting in incomplete feature representation and inability to adapt well to complex scenes and targets. To this end, the present invention proposes an image segmentation and dynamic target recognition method based on artificial intelligence. Summary of the Invention
[0003] The purpose of the present invention is to solve the problems in the background technology and to propose an image segmentation and dynamic target recognition method based on artificial intelligence.
[0004] In order to achieve the above object, the present invention adopts the following technical solutions:
[0005] An artificial intelligence-based image segmentation and dynamic target recognition method, comprising:
[0006] S1, collects temporal RGB-D image sequences, enhances dynamic edge information through multimodal spatiotemporal alignment processing and dual-stream feature fusion, and combines the Markov random field model with spatiotemporal confidence evaluation to output a dynamic candidate region set R1 with spatiotemporal labels;
[0007] S2, input the dynamic candidate region set R1 into the multi-level feature decoupler to achieve the gradual decoupling and enhancement of the bottom-level, middle-level and high-level features, thereby generating the spatiotemporal enhanced candidate region set R2;
[0008] S3. Initialize the deformable segmentation network based on the spatiotemporal features of the spatiotemporal enhancement candidate region set R2, encode the boundary constraints of the spatiotemporal enhancement candidate region set as the offset parameters of the deformable convolution kernel, and introduce the edge motion matching loss function for comparison analysis to generate a dynamic segmentation mask M1 that integrates geometric boundaries and motion continuity;
[0009] S4, based on the generated dynamic segmentation mask M1, calculate the optical flow field and back-propagate the segmentation boundary in the reverse direction to further optimize the segmentation result to determine the target area; after segmentation confidence evaluation and propagation, correct the segmentation boundary and output a dynamic segmentation mask M2 that is consistent in time and space;
[0010] S5. Build a target trajectory graph based on the dynamic segmentation mask M2, re-identify the target and complete the trajectory through the graph attention network, and predict the subsequent frame position based on the completed trajectory to achieve stable target tracking.
[0011] Furthermore, by multimodal spatiotemporal alignment processing and dual-stream feature fusion to enhance dynamic edge information, and combining the Markov random field model with spatiotemporal confidence evaluation, the process of outputting the dynamic candidate region set R1 with spatiotemporal labels includes:
[0012] Based on the collected temporal RGB-D image sequence, the following steps are performed to achieve multimodal spatiotemporal alignment: the temporal RGB-D image sequence is used as the input of a 3D convolutional network, and the 3D convolution kernel performs convolution operations simultaneously in the spatial and temporal dimensions; the 3D convolutional network automatically learns the spatiotemporal features of local regions in the temporal RGB-D image sequence through multiple layers of 3D convolution, pooling, and activation function mapping operations; the local spatiotemporal features are extracted through the 3D convolutional network, and finally a synchronized spatiotemporal alignment dataset is generated;
[0013] Based on a spatiotemporally aligned dataset, a two-stream feature fusion network is constructed to enhance dynamic edges: a first-stream network is set up to capture edge gradients in rigid motion regions; a second-stream network is set up to detect non-rigid deformation boundaries using a Canny edge detector combined with normal vector changes. Two branches are set up in the two-stream feature fusion network to process the edge maps of the first-stream network and the second-stream network respectively. A competition module is set up to compare and select the outputs of the two branches, thereby fusing the two edge maps, enhancing the motion saliency of dynamic targets, and generating a dynamic edge probability map. In the final dynamic edge probability map, high-probability regions correspond to potential target boundaries.
[0014] Based on the dynamic edge probability map, a Markov random field model is constructed to generate candidate regions: the Markov random field model is constructed with superpixels as nodes and edge strength as edge weights. By using superpixels as nodes of the Markov random field model and edge strength as edge weights, similar superpixels are gradually merged to form larger regions based on the similarity and edge strength between adjacent superpixels, thus achieving candidate region growth.
[0015] A spatiotemporal confidence evaluation function is introduced to evaluate the reliability of all candidate regions, and a built-in non-maximum suppression strategy is used for screening, ultimately forming a dynamic candidate region set with spatiotemporal labels.
[0016] Multi-dimensional features are extracted for each candidate region and fused to form a feature tensor: ResNet-50 is used to extract CNN features to represent the spatial structure information of the candidate region, thereby obtaining spatial dimension features; the optical flow deformation network is used to align the features of adjacent frames to generate a motion vector field, and then obtain motion dimension features; among them, the motion vector field encodes the target motion pattern; LSTM is used to record and analyze the motion changes of the target between adjacent frames, and track the dynamic changes of the target in the time series, thereby obtaining temporal dimension features; the features of the spatial dimension, motion dimension and temporal dimension are fused to form a multi-scale feature tensor.
[0017] Furthermore, the dynamic candidate region set R1 is input into the multi-level feature decoupler to achieve the gradual decoupling and enhancement of the bottom-level, middle-level, and high-level features, thereby generating the spatiotemporal enhanced candidate region set R2. The process includes:
[0018] The candidate region feature tensor is input into the pyramidal decoupling network: the bottom layer uses deformable convolution to extract local edge motion details, preliminarily captures the motion information of the candidate region edge, and forms the initial framework of the bottom-level motion features; the middle layer uses the trajectory prediction network to preliminarily aggregate the long-range motion trajectories within the candidate region and construct the initial framework of the middle-level trajectory features; the high-level uses the association analysis network to model the interaction relationship between candidate regions, form the preliminary structure of high-level association features, and preliminarily characterize the spatial association between candidate regions; after completing the initialization of the multi-level decoupling architecture, that is, preliminarily forming a three-dimensional feature framework of the bottom-level motion features, middle-level trajectory features, and high-level association features, further analysis and processing will be performed on the features of each level;
[0019] Based on the motion features initially extracted from the underlying layer, edge optical flow difference calculation is performed on the edge pixels of the candidate region to generate a fine motion vector field at the edge of the candidate region. Background motion interference is smoothed using Gaussian filtering, and the target motion vector is separated by combining the motion saliency mask, outputting a denoised pixel-level motion vector set.
[0020] The denoised pixel-level motion vectors are input into the trajectory prediction network. The optical flow estimation algorithm is used to process each pair of adjacent frames in the input candidate region image sequence, and the motion vector of each pixel in the candidate region image is re-estimated to obtain the inter-frame optical flow estimation results. Based on these estimated motion vectors, the motion trajectory of the candidate region is predicted frame by frame to further enhance the mid-level trajectory features. A motion trajectory coincidence loss function is set to constrain the displacement deviation between the predicted trajectory and the inter-frame optical flow estimation results, while optimizing the motion trajectory smoothness. Finally, the enhanced semantic trajectory features are output.
[0021] Based on the correlation features of the high-level preliminary modeling, a candidate region association graph is constructed. The nodes of the candidate region association graph are the center points of the candidate regions, and the edge weights are calculated jointly by motion similarity and spatial proximity. The features of adjacent nodes are aggregated through a graph convolutional network to generate a global correlation vector to represent the spatiotemporal dependency relationship between candidate regions.
[0022] A cross-level gating unit is set up to dynamically allocate the fusion ratio of low-level motion features, mid-level trajectory features, and high-level correlation features through the Sigmoid function. The features fused according to this ratio are input into the feature decoding network, which gradually restores the spatial resolution of the features through the spatial resolution reconstruction operation and adjusts the boundaries of the initially generated candidate regions through the boundary regression algorithm in combination with the position information of the original candidate regions. Finally, a spatiotemporal enhanced candidate region set R2 containing the enhanced boundaries and correlation relationships of the motion features is generated.
[0023] Furthermore, the deformable segmentation network is initialized based on the spatiotemporal features of the spatiotemporal enhancement candidate region set R2. The boundary constraints of the spatiotemporal enhancement candidate region set are encoded as the offset parameters of the deformable convolution kernel. The edge motion matching loss function is introduced for comparison analysis. The process of generating a dynamic segmentation mask M1 that integrates geometric boundaries and motion continuity includes:
[0024] The core component of the deformable segmentation network is the deformable convolution layer, which adaptively adjusts the shape and position of the convolution kernel according to the input spatiotemporal features to capture the geometric boundaries and motion changes of the target. During the initialization process, the initial parameters of the deformable convolution layer are set according to the spatiotemporal features of the spatiotemporal enhancement candidate region set R2. At the same time, the initial offset parameters of the deformable convolution kernel are initialized based on the motion features of the spatiotemporal enhancement candidate region set R2 to enhance the boundary information.
[0025] Extract the boundary information of the candidate region from the spatiotemporal enhancement candidate region set R2; encode the extracted boundary constraints as the offset parameters of the deformable convolution kernel, that is, for each sampling point on the deformable convolution kernel, calculate its spatial offset according to the boundary information; fuse the calculated offset parameters with the initial parameters of the deformable convolution kernel to update the parameters of the deformable convolution kernel;
[0026] For the candidate regions of adjacent frames in the spatiotemporal enhancement candidate region set R2, the displacement vector field of their edges is calculated; during the forward propagation of the deformable segmentation network, the segmentation edges of the current frame are extracted in real time; at the same time, based on the prediction results of the deformable segmentation network and the segmentation information of the historical frames, the segmentation edge change trend of the next frame is predicted; an edge motion matching loss function is introduced to measure the difference between the edge displacement vector field of the candidate regions of adjacent frames and the segmentation edge change trend;
[0027] The deformable segmentation network is combined with the edge motion matching loss function for training; during the training process, labeled data is used for iterative optimization, and the deformable segmentation network parameters are updated through the backpropagation algorithm; the stochastic gradient descent method is used to optimize the deformable segmentation network parameters during training; after training, the test image is input into the trained deformable segmentation network, and the deformable segmentation network generates a dynamic segmentation mask M1 based on the spatiotemporal characteristics and boundary constraints of the input image.
[0028] Furthermore, based on the generated dynamic segmentation mask M1, the optical flow field is calculated and the segmentation boundary is back-propagated in the reverse direction to further optimize the segmentation result to determine the target area. After segmentation confidence evaluation and propagation correction of the segmentation boundary, the process of outputting a spatiotemporally consistent dynamic segmentation mask M2 includes the following steps:
[0029] The Lucas-Kanade optical flow method is used to calculate the optical flow field between adjacent frames. The segmentation boundary in the dynamic segmentation mask M1 is used as the starting point and backpropagation is performed in the opposite direction of the optical flow calculated based on the adjacent frames. The optical flow backpropagation is used to simulate the motion trajectory of the boundary in the time dimension and determine whether the motion of the segmentation boundary between different frames is continuous. During the optical flow backpropagation process, a motion continuity threshold is set to determine the motion continuity of the segmentation boundary. If the deviation between the boundary position after backpropagation and the boundary position of the actual target in the adjacent frame is less than the motion continuity threshold, the boundary is considered to be continuous in motion; otherwise, it is considered to be discontinuous in motion.
[0030] A continuous frame image sequence containing the target area is input into a 3D convolutional neural network; 3D convolution is used to capture the features of the image in the spatial and temporal dimensions and extract the apparent temporal features of the target area; through the 3D convolution operation, the apparent change pattern of the target between different frames is obtained; based on the extracted 3D convolution features, the stability of the target appearance between adjacent frames is analyzed: the feature similarity of the target between different frames is calculated. If the similarity is lower than the preset appearance stability threshold, the appearance of the target area may be unstable between adjacent frames, and the area of apparent instability is determined;
[0031] The segmentation confidence of each candidate region in the dynamic segmentation mask M1 is evaluated. The segmentation confidence is calculated based on the dynamic edge strength of the segmentation boundary and the apparent feature similarity of the target appearance. High segmentation confidence is assigned to regions with continuous motion and stable appearance, and low segmentation confidence is assigned to regions with continuous motion and stable appearance. A dynamic segmentation confidence propagation mechanism is established to diffuse the prediction results of low segmentation confidence regions to high segmentation confidence neighborhoods. After segmentation confidence propagation, the dynamic segmentation mask M1 is updated to obtain a spatiotemporally consistent dynamic segmentation mask M2.
[0032] Furthermore, the target trajectory graph is constructed based on the dynamic segmentation mask M2, and the target is re-identified and the trajectory is completed through the graph attention network. The subsequent frame position is predicted based on the completed trajectory to achieve stable target tracking. The process includes:
[0033] The dynamic segmentation mask M2 is used as input and the pixels within the dynamic segmentation mask M2 are clustered using the saliency clustering algorithm. After clustering, the core area of each target segmentation area is extracted based on the clustering results. The centroid of each cluster area is calculated and the area within a fixed radius is selected as the core area with the centroid as the center.
[0034] After extracting the core area of the segmentation mask, the centroid of each core area is used as a node of the target trajectory map; each node corresponds to the position information of a target in a single frame of the video sequence. By recording the coordinates of the node and the frame number of the frame, the spatial distribution of the target in the video sequence frame is described;
[0035] Determine the weight of the edge in the target trajectory graph by calculating the cross-frame motion similarity;
[0036] Construct the target trajectory graph based on the determined nodes and calculated edge weights;
[0037] The constructed target trajectory graph is input into the graph attention network. The graph attention network automatically learns the association relationship between nodes and assigns corresponding attention weights to each node. The target is re-identified through the graph attention network. For missing parts of the trajectory, the motion extrapolation of the associated nodes is used to complete the trajectory. That is, based on the motion relationship and edge weights between the nodes in the target trajectory graph, the position and state of the target in the missing frame are predicted, thereby completing the trajectory.
[0038] Based on the completed target trajectory graph, the position and state of the target in subsequent frames of the video sequence are predicted using the motion extrapolation method of associated nodes. Motion extrapolation estimates the target's position in the next frame based on the target's historical motion trajectory and the position information of associated nodes. By continuously iterating this process, continuous tracking results are generated, achieving stable tracking of the target throughout the entire video sequence.
[0039] Compared with the prior art, the present invention has the following advantages: through multimodal spatiotemporal alignment processing, the temporal misalignment between different modal data is effectively eliminated, thereby improving data accuracy; dual-stream feature fusion enhances dynamic edge information and highlights target boundary features; combining the Markov random field model with spatiotemporal confidence assessment, a dynamic candidate region set with spatiotemporal labels is output, providing a reliable basis for subsequent processing; the dynamic candidate region set is input into a multi-level feature decoupler, which gradually decouples and enhances the multi-layer features, enabling more comprehensive mining of the feature information of the candidate regions, thereby generating a spatiotemporally enhanced candidate region set and improving the performance of subsequent segmentation and recognition; by initializing a deformable segmentation network, the boundary constraints are encoded as deformable convolution kernel offset parameters, which can better adapt to the target geometric boundaries; segmentation boundaries are corrected through segmentation confidence assessment and propagation, effectively correcting inaccurate segmentation areas, and outputting a spatiotemporally consistent dynamic segmentation mask, making the target region determination more accurate and improving segmentation quality; by utilizing a graph attention network to re-identify and complete the trajectory of the target, problems such as target occlusion and loss are effectively resolved; and the position of subsequent frames is predicted based on the completed trajectory, which can achieve stable target tracking and enhance the practicality of the system in real-world scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 This is a flowchart of an artificial intelligence-based image segmentation and dynamic target recognition method proposed by the present invention. DETAILED DESCRIPTION
[0041] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the implementation regulations described are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0042] Reference Figure 1 , an image segmentation and dynamic target recognition method based on artificial intelligence, comprising:
[0043] S1, collects a temporal RGB-D image sequence, enhances dynamic edge information through multimodal spatiotemporal alignment processing and dual-stream feature fusion, and combines the Markov random field model with spatiotemporal confidence evaluation to output a dynamic candidate region set R1 with spatiotemporal labels. It can be understood that a temporal RGB-D image is a series of image sequences collected in chronological order that contain color information RGB and depth information D;
[0044] S2, input the dynamic candidate region set R1 into the multi-level feature decoupler to achieve the gradual decoupling and enhancement of the bottom-level, middle-level and high-level features, thereby generating the spatiotemporal enhanced candidate region set R2;
[0045] S3. Initialize the deformable segmentation network based on the spatiotemporal features of the spatiotemporal enhancement candidate region set R2, encode the boundary constraints of the spatiotemporal enhancement candidate region set as the offset parameters of the deformable convolution kernel, and introduce the edge motion matching loss function for comparison analysis to generate a dynamic segmentation mask M1 that integrates geometric boundaries and motion continuity;
[0046] S4. Based on the generated dynamic segmentation mask M1, the optical flow field is calculated and the segmentation boundary is back-propagated in the reverse direction to further optimize the segmentation result to determine the target area. After segmentation confidence evaluation and propagation correction of the segmentation boundary, a dynamic segmentation mask M2 with consistent time and space is output. This mask is the final target segmentation result.
[0047] S5. Build a target trajectory graph based on the dynamic segmentation mask M2, re-identify the target and complete the trajectory through the graph attention network, and predict the subsequent frame position based on the completed trajectory to achieve stable target tracking.
[0048] It should be further explained that in the specific implementation process, through multimodal spatiotemporal alignment processing, dual-stream feature fusion to enhance dynamic edge information, and combined with the Markov random field model and spatiotemporal confidence evaluation, the process of outputting the dynamic candidate region set R1 with spatiotemporal labels is as follows:
[0049] Based on the collected temporal RGB-D image sequence, the following steps are performed to achieve multimodal spatiotemporal alignment: the temporal RGB-D image sequence is used as the input of the 3D convolutional network, and the 3D convolution kernel performs convolution operations in both spatial (width, height) and temporal dimensions; the 3D convolutional network automatically learns the spatiotemporal features of local regions in the temporal RGB-D image sequence with the help of multiple layers of 3D convolution, pooling, and activation function mapping operations. These features contain information about spatial structure and temporal dynamic changes; the local spatiotemporal features are extracted through the 3D convolutional network, and a synchronized spatiotemporal alignment dataset is finally generated; in the spatiotemporal alignment dataset, the motion gradient amplitude of the dynamic region is significantly higher than that of the static background region, where the motion gradient amplitude is obtained by calculating the degree of change of the motion vector of the pixel point in the image; in the dynamic region, since the object is moving, the motion vector of its pixel point changes greatly, so the motion gradient amplitude is higher; while in the static background region, the motion vector of the pixel point remains almost unchanged, and the motion gradient amplitude is lower; based on this difference, the dynamic region and the static background region are distinguished;
[0050] Based on a spatiotemporally aligned dataset, a two-stream feature fusion network is constructed to enhance dynamic edges. The first-stream network is set up to capture edge gradients in rigid motion regions. Edge gradients refer to the rate of change of the grayscale values of pixels on either side of an edge in an image, reflecting the strength and direction of the edge. Edge gradient information is obtained by processing the image using edge detection operators (such as the Sobel operator). The first-stream network captures edge gradient features in rigid motion regions through learning.
[0051] The second-stream network is configured to detect non-rigidly deformable boundaries using a Canny edge detector combined with normal vector changes. Normal vector changes are caused by surface deformation, resulting in changes in the direction and magnitude of normal vectors at different locations. The Canny edge detector is an edge detection algorithm used to detect edges in an image. The second-stream network combines edge information obtained from the Canny edge detector with normal vector change information to detect non-rigidly deformable boundaries.
[0052] In the dual-stream feature fusion network, two branches are set up to process the edge maps of the first-stream network and the second-stream network respectively. A competition module is set up to compare and select the outputs of the two branches (the first-stream network and the second-stream network). For example, the similarity of the outputs of the two branches is calculated. For areas with low similarity, they are considered to contain important edge information and are given higher weights. In this way, the two edge maps are fused, the motion saliency of dynamic targets is enhanced, and a dynamic edge probability map is generated. In the final dynamic edge probability map, high-probability areas correspond to potential target boundaries, that is, areas with higher probability values are more likely to be the boundaries of dynamic targets.
[0053] Based on the dynamic edge probability map, a Markov random field model is constructed to generate candidate regions: the Markov random field model is constructed with superpixels as nodes and edge strength as edge weights. Superpixels are image blocks formed by aggregating adjacent pixels with similar color, texture, and other features in an image, effectively reducing the computational complexity of image processing while retaining important structural information of the image. By using superpixels as nodes in the Markov random field model and edge strength as edge weights, similar superpixels are gradually merged to form larger regions based on the similarity and edge strength between adjacent superpixels, thus achieving candidate region growth.
[0054] A spatiotemporal confidence evaluation function is introduced to evaluate the reliability of all candidate regions, and a built-in non-maximum suppression strategy is used for screening, ultimately forming a dynamic candidate region set with spatiotemporal labels:
[0055] Specifically, a spatiotemporal confidence evaluation function is constructed based on motion saliency × region compactness × temporal continuity. The spatiotemporal confidence value of each candidate region is obtained by calculation. This value is used to quantify the reliability of the candidate region as a dynamic target region. The spatiotemporal confidence evaluation function expression is: FV = MS × RC × TC, where FV is the spatiotemporal confidence value and MS is the motion saliency ( Where PN represents the total number of pixels in the candidate region, i is the pixel index in the dynamic edge probability map, PVi represents the probability value of the i-th pixel in the dynamic edge probability map (0-1, generated by the two-stream network), and RC is the region compactness ( Where PA is the area of the superpixel region, PS is the perimeter of the superpixel region, and TC is the temporal continuity Where Δp pred is the displacement vector predicted by optical flow (coordinate change between adjacent frames), Δp gt is the true displacement vector, σ is the attenuation factor for controlling the matching error);
[0056] All candidate regions are sorted in descending order according to their spatiotemporal confidence values, ensuring that regions with high spatiotemporal confidence are ranked first; starting from the candidate region with the highest spatiotemporal confidence, the degree of overlap between this region and other candidate regions is calculated in turn, using the intersection over union (IoU) as the measurement indicator; for any two candidate regions, if their IoU is greater than the preset overlap threshold, proceed to the next step of comparison; compare the spatiotemporal confidence values of the two candidate regions whose IoU is greater than the overlap threshold, retain the candidate region with higher spatiotemporal confidence, and suppress (remove) the candidate region with lower spatiotemporal confidence, and continue the above screening process until all candidate regions are traversed; after the screening process, the final set of candidate regions retained is the dynamic candidate region set with spatiotemporal labels, and the candidate region boundaries in this set cover more than 90% of the dynamic target area;
[0057] Multi-dimensional features are extracted for each candidate region and fused to form a feature tensor: ResNet-50 is used to extract CNN features to represent the spatial structure information of the candidate region, thereby obtaining spatial dimension features; the optical flow deformation network is used to align the features of adjacent frames to generate a motion vector field, thereby obtaining motion dimension features; LSTM is used to record and analyze the motion changes of the target between adjacent frames, and track the dynamic changes of the target in the time series, thereby obtaining temporal dimension features; the features of the spatial dimension, motion dimension and temporal dimension are fused to form a multi-scale feature tensor.
[0058] It should be further explained that, in the specific implementation process, the dynamic candidate region set R1 is input into the multi-level feature decoupler to achieve the gradual decoupling and enhancement of the bottom-level, middle-level, and high-level features, thereby generating the spatiotemporal enhancement candidate region set R2 as follows:
[0059] The candidate region feature tensor is input into the pyramidal decoupling network: the bottom layer uses deformable convolution to extract local edge motion details, preliminarily captures the motion information of the candidate region edge, and forms the initial framework of the bottom-level motion features; the middle layer uses the trajectory prediction network to preliminarily aggregate the long-range motion trajectories within the candidate region and construct the initial framework of the middle-level trajectory features; the high-level uses the association analysis network to model the interaction relationship between candidate regions, form the preliminary structure of high-level association features, and preliminarily characterize the spatial association between candidate regions; after completing the initialization of the multi-level decoupling architecture, that is, preliminarily forming a three-dimensional feature framework of the bottom-level motion features, middle-level trajectory features, and high-level association features, further analysis and processing will be performed on the features of each level;
[0060] Based on the motion features initially extracted at the bottom layer, edge optical flow difference calculation is performed on the edge pixels of the candidate region to generate a fine motion vector field of the edge of the candidate region; it is understandable that the edge optical flow difference calculation extracts the pixel-level motion vector of the edge of the candidate region by performing a differential operation on the optical flow information of the edge pixels of the dynamic candidate region between adjacent frames to capture the local motion details of the edge; background motion interference is smoothed by Gaussian filtering, and the target motion vector is separated by combining the motion saliency mask to output the denoised pixel-level motion vector set; it is understandable that after the multi-level decoupling architecture is initialized, the motion features initially extracted at the bottom layer contain information about target motion and background motion; background motion refers to the fact that in a dynamic scene, the background of the candidate region is not static, but there are various complex motions, such as the motion of dynamic background areas such as wind blowing leaves and clouds floating. These motions will interfere with the feature extraction of the target candidate region, so denoising is required; the motion saliency mask is generated according to the saliency of the motion information, and is used to highlight the target motion area and suppress the background motion area;
[0061] The denoised pixel-level motion vectors are input into the trajectory prediction network. An optical flow estimation algorithm (such as FlowNet and RAFT) is used to process each pair of adjacent frames in the input candidate region image sequence, and the motion vector of each pixel in the candidate region image is re-estimated to obtain the inter-frame optical flow estimation results. Based on these estimated motion vectors, the motion trajectory of the candidate region is predicted frame by frame to further enhance the mid-level trajectory features. A motion trajectory coincidence loss function is set to constrain the displacement deviation between the predicted trajectory and the inter-frame optical flow estimation result, ensuring that the predicted trajectory is consistent with the actual motion situation and optimizing the smoothness of the motion trajectory.
[0062] Among them, the expression of the motion trajectory coincidence loss function is: L traj =α·L motion +β·L smooth , where L traj is the motion trajectory overlap loss value, L motion is the displacement deviation loss ( is the displacement vector of the predicted trajectory at the tth frame (output by the trajectory prediction network), expressed as (x t ,y t ); t is the index of the frame, T is the number of frames; x and y are the coordinate positions of the pixel, respectively, with the x axis extending to the right and the y axis extending downward; is the displacement vector of the optical flow estimation result in the tth frame (generated by the optical flow estimation algorithm), expressed as (u t ,v t ); u and v are the components of the optical flow vector, that is, the motion displacement of the pixel between adjacent frames, u is the horizontal displacement component (lateral motion), and v is the vertical displacement component (longitudinal motion);
[0063] L smooth is the trajectory smoothness loss α is the weight of displacement deviation loss; β is the weight of trajectory smoothness loss;
[0064] The final output is the enhanced semantic trajectory feature, and the displacement prediction error between the starting point and the end point of the trajectory is less than 2 pixels;
[0065] Based on the correlation features of the high-level preliminary model, a candidate region association graph is constructed. The nodes of the candidate region association graph are the center points of the candidate regions, and the edge weights are calculated by jointly calculating the motion similarity (cosine similarity) and spatial proximity (inverse Euclidean distance). The features of adjacent nodes are aggregated through a graph convolutional network to generate a global correlation vector, which represents the spatiotemporal dependency relationship between candidate regions.
[0066] A cross-level gating unit is set up to dynamically allocate the fusion ratio of the bottom-level motion features (weight γ1), the middle-level trajectory features (weight γ2), and the high-level association features (weight γ3) through the Sigmoid function; the features fused according to this ratio are input into the feature decoding network, which uses spatial resolution reconstruction operations (including deconvolution operations, bilinear interpolation, transposed convolution, etc.) to gradually restore the spatial resolution of the features, and adjusts the boundaries of the initially generated candidate regions through the boundary regression algorithm in combination with the position information of the original candidate regions; it is understandable that the position information of the original candidate regions is obtained by the target detection or candidate region generation algorithm; the boundary regression algorithm is a technology used to optimize the position of the bounding box in target detection or candidate region generation. In the target detection task, the initially generated candidate region bounding box may not be able to accurately frame the target object. The boundary regression algorithm is used to adjust the position and size of these bounding boxes to make them fit the target more accurately; finally, a spatiotemporal enhanced candidate region set R2 containing the enhanced boundaries and association relationships of the motion features is generated.
[0067] It should be further explained that in the specific implementation process, the deformable segmentation network is initialized based on the spatiotemporal features of the spatiotemporal enhancement candidate region set R2, the boundary constraints of the spatiotemporal enhancement candidate region set are encoded as the offset parameters of the deformable convolution kernel, and the edge motion matching loss function is introduced for comparison analysis. The process of generating a dynamic segmentation mask M1 that integrates geometric boundaries and motion continuity is as follows:
[0068] The core component of the deformable segmentation network is the deformable convolution layer, which adaptively adjusts the shape and position of the convolution kernel according to the input spatiotemporal features to capture the geometric boundaries and motion changes of the target. During the initialization process, the initial parameters of the deformable convolution layer, such as the size of the convolution kernel, are set according to the spatiotemporal features in the spatiotemporal enhancement candidate region set R2. At the same time, based on the motion features in the spatiotemporal enhancement candidate region set R2, the boundary information is enhanced and the initial offset parameters of the deformable convolution kernel are initialized (which can be set to zero or initialized according to empirical values at the beginning). This enables the deformable convolution kernel to have a certain target boundary perception ability at the initialization stage.
[0069] Extract the boundary information of the candidate region from the spatiotemporal enhancement candidate region set R2, where the extracted boundary information of the candidate region is the embodiment of the motion feature enhanced boundary in the spatiotemporal enhancement candidate region set R2, and includes the approximate outline and position of the target in the image; encode the extracted boundary constraint as the offset parameter of the deformable convolution kernel, that is, for each sampling point on the deformable convolution kernel, calculate its spatial offset according to the boundary information. Specifically, by calculating the distance and direction between the sampling point and the boundary, the size and direction of its offset are determined, and the offset parameter is calculated using a gradient-based method to ensure that the offset can make the convolution kernel better adapt to the geometric boundary of the target; fuse the calculated offset parameter with the initial parameter of the deformable convolution kernel, and update the parameters of the deformable convolution kernel, so that the deformable convolution kernel can be adaptively adjusted according to the geometric boundary of the target in subsequent convolution operations, thereby improving the accuracy of segmentation;
[0070] For the adjacent frame candidate regions in the spatiotemporal enhancement candidate region set R2, the displacement vector field of their edges is calculated, wherein the displacement vectors of the edge pixels of the candidate regions between adjacent frames are calculated based on a feature matching method to obtain an edge displacement vector field, which reflects the motion change of the target between adjacent frames; during the forward propagation process of the deformable segmentation network, the segmentation edge of the current frame is extracted in real time; at the same time, based on the prediction results of the deformable segmentation network and the segmentation information of the historical frames, the segmentation edge change trend of the next frame is predicted, for example, by analyzing the geometric features (such as curvature) of the segmentation edge to obtain the segmentation edge change trend; it can be understood that the deformable segmentation network performs forward propagation calculation on the current frame input image, and predicts the image segmentation through its network structure and parameters, thereby obtaining the segmentation prediction result of the current frame; the above-mentioned historical frame comes from the frame image located before the current frame in the temporal RGB-D image sequence; an edge motion matching loss function is introduced to measure the difference between the edge displacement vector field of the candidate regions of adjacent frames and the segmentation edge change trend: the Euclidean distance between the edge displacement vector field and the predicted segmentation edge change trend is calculated and used as the value of the edge motion matching loss function;
[0071] The deformable segmentation network is combined with the edge motion matching loss function for training; during the training process, labeled data is used for iterative optimization, and the deformable segmentation network parameters are updated through the back-propagation algorithm; it is understandable that the labeled data is obtained by manually accurately segmenting and annotating the target objects in the temporal RGB-D image sequence, and is used to guide the training and optimization of the deformable segmentation network; the stochastic gradient descent method is used to optimize the deformable segmentation network parameters during training; after training, the test image is input into the trained deformable segmentation network, and the deformable segmentation network generates a dynamic segmentation mask M1 based on the spatiotemporal characteristics and boundary constraints of the input image; among them, the dynamic segmentation mask M1 integrates geometric boundaries and motion continuity information, and can accurately segment the target object.
[0072] It should be further explained that in the specific implementation process, based on the generated dynamic segmentation mask M1, the optical flow field is calculated and the segmentation boundary is back-propagated in the reverse direction to further optimize the segmentation result to determine the target area; after the segmentation confidence is evaluated and the segmentation boundary is corrected through propagation, the process of outputting the spatiotemporally consistent dynamic segmentation mask M2 is as follows:
[0073] The Lucas-Kanade optical flow method is used to calculate the optical flow field between adjacent frames; the segmentation boundary in the dynamic segmentation mask M1 is used as the starting point, and backpropagation is performed along the opposite direction of the optical flow calculated based on the adjacent frames; wherein, the optical flow backpropagation is used to simulate the motion trajectory of the boundary in the time dimension and determine whether the motion of the segmentation boundary between different frames is continuous; in the process of optical flow backpropagation, a motion continuity threshold is set to determine the motion continuity of the segmentation boundary; if the deviation between the boundary position after backpropagation and the boundary position of the actual target in the adjacent frame is less than the motion continuity threshold, then the boundary is considered to be continuous in motion; otherwise, it is considered that there is motion discontinuity; at the same time, the area of motion discontinuity and its related information, such as position, degree, etc., are recorded to provide a basis for subsequent confidence propagation and mask optimization to correct the candidate area boundary with inaccurate segmentation; it can be understood that the Lucas-Kanade optical flow method is an optical flow estimation algorithm that can calculate the optical flow field to quantify the motion information of pixels between adjacent frames, providing a key data basis for subsequent optical flow backpropagation of the segmentation boundary to simulate the motion trajectory and determine the motion continuity;
[0074] A sequence of continuous frame images containing the target area is input into the 3D convolutional neural network; it can be understood that the above-mentioned target area is a region that may contain the target based on the generated dynamic segmentation mask M1; 3D convolution is used to capture the characteristics of the image in the spatial and temporal dimensions and extract the apparent temporal characteristics of the target area; through the 3D convolution operation, the apparent change pattern of the target between different frames is obtained; based on the extracted 3D convolution features, the stability of the target appearance between adjacent frames is analyzed: the feature similarity of the target between different frames is calculated. If the similarity is lower than the preset appearance stability threshold, the appearance of the possible target area is unstable between adjacent frames, and the area of apparent instability is determined to provide a basis for subsequent confidence propagation so as to adjust the inaccurately segmented area;
[0075] The segmentation confidence of each candidate region in the dynamic segmentation mask M1 is evaluated; wherein, the segmentation confidence is calculated according to the dynamic edge strength of the segmentation boundary and the apparent feature similarity of the target appearance (the dynamic edge strength is generated by the dynamic edge probability map, which indicates the significance of the segmentation boundary pixels; the apparent feature similarity is calculated by the apparent temporal features extracted by the 3D convolutional network to calculate the similarity between adjacent frames), and high segmentation confidence is given to the region with continuous motion and stable appearance, and low segmentation confidence is given otherwise; it is understandable that the segmentation confidence evaluation is a comprehensive evaluation based on the verification results of motion flow and appearance flow, which is used to determine which candidate regions are more likely to be the real target regions; a dynamic segmentation confidence propagation mechanism is established to propagate low segmentation confidence regions to the target regions. The prediction results are diffused to the high segmentation confidence neighborhood. Specifically, for each low segmentation confidence area, its high segmentation confidence neighborhood at the adjacent position in the frame sequence is found, and the prediction results of the neighborhood are propagated to the low segmentation confidence area according to factors such as the segmentation confidence and distance of the neighborhood. The weighted average method is adopted in the propagation process, and the weight is determined according to factors such as the neighborhood distance and segmentation confidence. After the segmentation confidence is propagated, the dynamic segmentation mask M1 is updated to obtain a spatiotemporally consistent dynamic segmentation mask M2. The dynamic segmentation mask M2 integrates the information of motion continuity and apparent temporal stability, and can more accurately segment the target object. The dynamic segmentation mask M2 obtained at this time is the final target area segmentation result.
[0076] It should be further explained that in the specific implementation process, the target trajectory graph is constructed based on the dynamic segmentation mask M2, the target is re-identified and the trajectory is completed through the graph attention network, and the subsequent frame position is predicted based on the completed trajectory to achieve stable target tracking. The process is as follows:
[0077] The dynamic segmentation mask M2 is used as input, and the pixels within the dynamic segmentation mask M2 are clustered using the saliency clustering algorithm. The saliency clustering algorithm clusters pixels belonging to the same target based on pixel color and other features, thereby identifying the segmentation regions of different targets. After clustering, the core region of each target segmentation region is extracted based on the clustering results. The core region represents the main part of the target and contains the salient features of the target, providing key information for the subsequent target trajectory construction. The centroid of each cluster region is calculated, and the region within a fixed radius is selected as the core region with the centroid as the center. The determination of the fixed radius range integrates the regional feature distribution density and the prior knowledge of the target size. The regional feature distribution density is used to reflect the concentration of features in the target region, and the prior knowledge of the target size reasonably determines the range of the core region based on the actual size of the target, ensuring that the core region can accurately represent the main part of the target.
[0078] After extracting the core area of the segmentation mask, the centroid of each core area is used as a node in the target trajectory graph; each node corresponds to the position information of a target in a single frame of the video sequence (a video sequence consists of a series of continuous image frames). By recording the coordinates of the node and the frame number of the frame, the spatial distribution of the target in the video sequence frame is described, providing an accurate positioning basis for the subsequent construction of the target trajectory;
[0079] The weights of the edges in the target trajectory graph are determined by calculating the cross-frame motion similarity: the cross-frame motion similarity is calculated by jointly calculating the trajectory offset and the apparent feature distance. The trajectory offset is obtained by calculating the displacement of the target center of mass in adjacent frames, and the apparent feature distance is obtained by comparing the features of the target core area in adjacent frames (such as color histograms). The trajectory offset and the apparent feature distance are weighted and fused to obtain the cross-frame motion similarity, which is used as the edge weight.
[0080] Based on the determined nodes and the calculated edge weights, a target trajectory graph is constructed. It can be understood that the target trajectory graph intuitively shows the motion relationship and correlation degree of the target between different frames.
[0081] The constructed target trajectory graph is input into the graph attention network; the graph attention network automatically learns the association relationship between nodes and assigns corresponding attention weights to each node; through the graph attention network, the target is re-identified, and the same target can be accurately identified even when the target is occluded, deformed, or the scene illumination changes; for the missing part of the trajectory caused by occlusion, target deformation, etc., the motion extrapolation of the associated nodes is used to complete it, that is, according to the motion relationship and edge weights between the nodes in the target trajectory graph, the position and state of the target in the missing frame are predicted, thereby completing the trajectory; it is understandable that in the target trajectory In the figure, associated nodes refer to nodes that are connected to the current node by edges. These nodes are adjacent to the current node or have motion association with the current node in the video sequence frames. Their motion information provides a basis for completing the trajectory of the current node. For example, during the target movement, the nodes in the adjacent frames are associated nodes. The motion extrapolation method is to predict the position and state of the target in the frame where the trajectory is missing based on the historical motion trajectory of the target (such as speed) and the position information of the associated nodes. For example, if the target moves at a certain speed and direction in the previous frame, and the position of the associated node also conforms to this motion trend, then this information is used to predict the position of the target in the missing frame.
[0082] Based on the completed target trajectory graph, the position and state of the target in subsequent frames of the video sequence are predicted using the motion extrapolation method of associated nodes. Motion extrapolation estimates the target's position in the next frame based on the target's historical motion trajectory and the position information of associated nodes. By continuously iterating this process (i.e., starting from trajectory completion using motion extrapolation of associated nodes to predicting the target's position and state in subsequent frames of the video sequence based on the completed trajectory graph), continuous tracking results are generated, achieving stable tracking of the target throughout the entire video sequence (consisting of a series of consecutive image frames).
[0083] It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0084] It should be understood that determining B based on A does not mean determining B based solely on A. B can also be determined based on A and / or other information.
[0085] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
[0086] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. An artificial intelligence-based image segmentation and dynamic target recognition method, characterized by: S1, collects temporal RGB-D image sequences, enhances dynamic edge information through multimodal spatiotemporal alignment processing and dual-stream feature fusion, and combines the Markov random field model with spatiotemporal confidence evaluation to output a dynamic candidate region set R1 with spatiotemporal labels; S2, input the dynamic candidate region set R1 into the multi-level feature decoupler to achieve the gradual decoupling and enhancement of the bottom-level, middle-level and high-level features, thereby generating the spatiotemporal enhanced candidate region set R2; S3. Initialize the deformable segmentation network based on the spatiotemporal features of the spatiotemporal enhancement candidate region set R2, encode the boundary constraints of the spatiotemporal enhancement candidate region set as the offset parameters of the deformable convolution kernel, and introduce the edge motion matching loss function for comparison analysis to generate a dynamic segmentation mask M1 that integrates geometric boundaries and motion continuity; S4, based on the generated dynamic segmentation mask M1, calculate the optical flow field and back-propagate the segmentation boundary in the reverse direction to further optimize the segmentation result to determine the target area; after segmentation confidence evaluation and propagation, correct the segmentation boundary and output a dynamic segmentation mask M2 that is consistent in time and space; S5. Build a target trajectory graph based on the dynamic segmentation mask M2, re-identify the target and complete the trajectory through the graph attention network, and predict the subsequent frame position based on the completed trajectory to achieve stable target tracking.
2. The method for image segmentation and dynamic target recognition based on artificial intelligence according to claim 1, characterized in that: The process of outputting the dynamic candidate region set R1 with spatiotemporal labels by multimodal spatiotemporal alignment processing, dual-stream feature fusion, and dynamic edge information enhancement, combined with the Markov random field model and spatiotemporal confidence evaluation, includes: Based on the collected temporal RGB-D image sequence, the following steps are performed to achieve multimodal spatiotemporal alignment: the temporal RGB-D image sequence is used as the input of a 3D convolutional network, and the 3D convolution kernel performs convolution operations simultaneously in the spatial and temporal dimensions; the 3D convolutional network automatically learns the spatiotemporal features of local regions in the temporal RGB-D image sequence through multiple layers of 3D convolution, pooling, and activation function mapping operations; the local spatiotemporal features are extracted through the 3D convolutional network, and finally a synchronized spatiotemporal alignment dataset is generated; Based on a spatiotemporally aligned dataset, a two-stream feature fusion network is constructed to enhance dynamic edges: a first-stream network is set up to capture edge gradients in rigid motion regions; a second-stream network is set up to detect non-rigid deformation boundaries using a Canny edge detector combined with normal vector changes. Two branches are set up in the two-stream feature fusion network to process the edge maps of the first-stream network and the second-stream network respectively. A competition module is set up to compare and select the outputs of the two branches, thereby fusing the two edge maps, enhancing the motion saliency of dynamic targets, and generating a dynamic edge probability map. In the final dynamic edge probability map, high-probability regions correspond to potential target boundaries. Based on the dynamic edge probability map, a Markov random field model is constructed to generate candidate regions: the Markov random field model is constructed with superpixels as nodes and edge strength as edge weights. By using superpixels as nodes of the Markov random field model and edge strength as edge weights, similar superpixels are gradually merged to form larger regions based on the similarity and edge strength between adjacent superpixels, thus achieving candidate region growth. A spatiotemporal confidence evaluation function is introduced to evaluate the reliability of all candidate regions, and a built-in non-maximum suppression strategy is used for screening, ultimately forming a dynamic candidate region set with spatiotemporal labels. Multi-dimensional features are extracted for each candidate region and fused to form a feature tensor: ResNet-50 is used to extract CNN features to represent the spatial structure information of the candidate region, thereby obtaining spatial dimension features; the optical flow deformation network is used to align the features of adjacent frames to generate a motion vector field, and then obtain motion dimension features; among them, the motion vector field encodes the target motion pattern; LSTM is used to record and analyze the motion changes of the target between adjacent frames, and track the dynamic changes of the target in the time series, thereby obtaining temporal dimension features; the features of the spatial dimension, motion dimension and temporal dimension are fused to form a multi-scale feature tensor.
3. The method for image segmentation and dynamic target recognition based on artificial intelligence according to claim 2, characterized in that: The dynamic candidate region set R1 is input into the multi-level feature decoupler to achieve the gradual decoupling and enhancement of the bottom-level, middle-level, and high-level features, thereby generating the spatiotemporal enhancement candidate region set R2. The process includes: The candidate region feature tensor is input into the pyramidal decoupling network: the bottom layer uses deformable convolution to extract local edge motion details, preliminarily captures the motion information of the candidate region edge, and forms the initial framework of the bottom-level motion features; the middle layer uses the trajectory prediction network to preliminarily aggregate the long-range motion trajectories within the candidate region and construct the initial framework of the middle-level trajectory features; the high-level uses the association analysis network to model the interaction relationship between candidate regions, form the preliminary structure of high-level association features, and preliminarily characterize the spatial association between candidate regions; after completing the initialization of the multi-level decoupling architecture, that is, preliminarily forming a three-dimensional feature framework of the bottom-level motion features, middle-level trajectory features, and high-level association features, further analysis and processing will be performed on the features of each level; Based on the motion features initially extracted from the underlying layer, edge optical flow difference calculation is performed on the edge pixels of the candidate region to generate a fine motion vector field at the edge of the candidate region. Background motion interference is smoothed using Gaussian filtering, and the target motion vector is separated by combining the motion saliency mask, outputting a denoised pixel-level motion vector set. The denoised pixel-level motion vectors are input into the trajectory prediction network. The optical flow estimation algorithm is used to process each pair of adjacent frames in the input candidate region image sequence, and the motion vector of each pixel in the candidate region image is re-estimated to obtain the inter-frame optical flow estimation results. Based on these estimated motion vectors, the motion trajectory of the candidate region is predicted frame by frame to further enhance the mid-level trajectory features. A motion trajectory coincidence loss function is set to constrain the displacement deviation between the predicted trajectory and the inter-frame optical flow estimation results, while optimizing the motion trajectory smoothness. Finally, the enhanced semantic trajectory features are output. Based on the correlation features of the high-level preliminary modeling, a candidate region association graph is constructed. The nodes of the candidate region association graph are the center points of the candidate regions, and the edge weights are calculated jointly by motion similarity and spatial proximity. The features of adjacent nodes are aggregated through a graph convolutional network to generate a global correlation vector to represent the spatiotemporal dependency relationship between candidate regions. A cross-level gating unit is set up to dynamically allocate the fusion ratio of low-level motion features, mid-level trajectory features, and high-level correlation features through the Sigmoid function. The features fused according to this ratio are input into the feature decoding network, which gradually restores the spatial resolution of the features through the spatial resolution reconstruction operation and adjusts the boundaries of the initially generated candidate regions through the boundary regression algorithm in combination with the position information of the original candidate regions. Finally, a spatiotemporal enhanced candidate region set R2 containing the enhanced boundaries and correlation relationships of the motion features is generated.
4. The method for image segmentation and dynamic target recognition based on artificial intelligence according to claim 1, characterized in that: The deformable segmentation network is initialized based on the spatiotemporal features of the spatiotemporal enhancement candidate region set R2. The boundary constraints of the spatiotemporal enhancement candidate region set are encoded as the offset parameters of the deformable convolution kernel. The edge motion matching loss function is introduced for comparison analysis. The process of generating a dynamic segmentation mask M1 that integrates geometric boundaries and motion continuity includes the following: The core component of the deformable segmentation network is the deformable convolution layer, which adaptively adjusts the shape and position of the convolution kernel according to the input spatiotemporal features to capture the geometric boundaries and motion changes of the target. During the initialization process, the initial parameters of the deformable convolution layer are set according to the spatiotemporal features of the spatiotemporal enhancement candidate region set R2. At the same time, the initial offset parameters of the deformable convolution kernel are initialized based on the motion features of the spatiotemporal enhancement candidate region set R2 to enhance the boundary information. Extract the boundary information of the candidate region from the spatiotemporal enhancement candidate region set R2; encode the extracted boundary constraints as the offset parameters of the deformable convolution kernel, that is, for each sampling point on the deformable convolution kernel, calculate its spatial offset according to the boundary information; fuse the calculated offset parameters with the initial parameters of the deformable convolution kernel to update the parameters of the deformable convolution kernel; For the candidate regions of adjacent frames in the spatiotemporal enhancement candidate region set R2, the displacement vector field of their edges is calculated; during the forward propagation of the deformable segmentation network, the segmentation edges of the current frame are extracted in real time; at the same time, based on the prediction results of the deformable segmentation network and the segmentation information of the historical frames, the segmentation edge change trend of the next frame is predicted; an edge motion matching loss function is introduced to measure the difference between the edge displacement vector field of the candidate regions of adjacent frames and the segmentation edge change trend; The deformable segmentation network is combined with the edge motion matching loss function for training; during the training process, labeled data is used for iterative optimization, and the deformable segmentation network parameters are updated through the backpropagation algorithm; the stochastic gradient descent method is used to optimize the deformable segmentation network parameters during training; after training, the test image is input into the trained deformable segmentation network, and the deformable segmentation network generates a dynamic segmentation mask M1 based on the spatiotemporal characteristics and boundary constraints of the input image.
5. The method for image segmentation and dynamic target recognition based on artificial intelligence according to claim 1, characterized in that: Based on the generated dynamic segmentation mask M1, the optical flow field is calculated and the segmentation boundary is back-propagated in the reverse direction to further optimize the segmentation result to determine the target area; After segmentation confidence evaluation and propagation correction of segmentation boundaries, the process of outputting a spatiotemporally consistent dynamic segmentation mask M2 includes: The Lucas-Kanade optical flow method is used to calculate the optical flow field between adjacent frames. The segmentation boundary in the dynamic segmentation mask M1 is used as the starting point and backpropagation is performed in the opposite direction of the optical flow calculated based on the adjacent frames. The optical flow backpropagation is used to simulate the motion trajectory of the boundary in the time dimension and determine whether the motion of the segmentation boundary between different frames is continuous. During the optical flow backpropagation process, a motion continuity threshold is set to determine the motion continuity of the segmentation boundary. If the deviation between the boundary position after backpropagation and the boundary position of the actual target in the adjacent frame is less than the motion continuity threshold, the boundary is considered to be continuous in motion; otherwise, it is considered to be discontinuous in motion. A continuous frame image sequence containing the target area is input into a 3D convolutional neural network; 3D convolution is used to capture the features of the image in the spatial and temporal dimensions and extract the apparent temporal features of the target area; through the 3D convolution operation, the apparent change pattern of the target between different frames is obtained; based on the extracted 3D convolution features, the stability of the target appearance between adjacent frames is analyzed: the feature similarity of the target between different frames is calculated. If the similarity is lower than the preset appearance stability threshold, the appearance of the target area may be unstable between adjacent frames, and the area of apparent instability is determined; The segmentation confidence of each candidate region in the dynamic segmentation mask M1 is evaluated. The segmentation confidence is calculated based on the dynamic edge strength of the segmentation boundary and the apparent feature similarity of the target appearance. High segmentation confidence is assigned to regions with continuous motion and stable appearance, and low segmentation confidence is assigned to regions with continuous motion and stable appearance. A dynamic segmentation confidence propagation mechanism is established to diffuse the prediction results of low segmentation confidence regions to high segmentation confidence neighborhoods. After segmentation confidence propagation, the dynamic segmentation mask M1 is updated to obtain a spatiotemporally consistent dynamic segmentation mask M2.
6. The method for image segmentation and dynamic target recognition based on artificial intelligence according to claim 1, characterized in that: The target trajectory graph is constructed based on the dynamic segmentation mask M2. The target is re-identified and the trajectory is completed through the graph attention network. The subsequent frame position is predicted based on the completed trajectory to achieve stable target tracking. The process includes: The dynamic segmentation mask M2 is used as input and the pixels within the dynamic segmentation mask M2 are clustered using the saliency clustering algorithm. After clustering, the core area of each target segmentation area is extracted based on the clustering results. The centroid of each cluster area is calculated and the area within a fixed radius is selected as the core area with the centroid as the center. After extracting the core area of the segmentation mask, the centroid of each core area is used as a node of the target trajectory map; each node corresponds to the position information of a target in a single frame of the video sequence. By recording the coordinates of the node and the frame number of the frame, the spatial distribution of the target in the video sequence frame is described; Determine the weight of the edge in the target trajectory graph by calculating the cross-frame motion similarity; Construct the target trajectory graph based on the determined nodes and calculated edge weights; The constructed target trajectory graph is input into the graph attention network. The graph attention network automatically learns the association relationship between nodes and assigns corresponding attention weights to each node. The target is re-identified through the graph attention network. For missing parts of the trajectory, the motion extrapolation of the associated nodes is used to complete the trajectory. That is, based on the motion relationship and edge weights between the nodes in the target trajectory graph, the position and state of the target in the missing frame are predicted, thereby completing the trajectory. Based on the completed target trajectory graph, the position and state of the target in subsequent frames of the video sequence are predicted using the motion extrapolation method of associated nodes. Motion extrapolation estimates the target's position in the next frame based on the target's historical motion trajectory and the position information of associated nodes. By continuously iterating this process, continuous tracking results are generated, achieving stable tracking of the target throughout the entire video sequence.
Citation Information
Cited By
Real-time monitoring method for copper-plated steel strip electroplating production line
CN120894364A
SPR response region identification method based on image semantic segmentation and time sequence alignment
CN120894544A
Backlight effect image edge enhancement method based on intelligent identification
CN120953153A
Intelligent recognition-based backlight effect image edge enhancement method
CN120953153B
Traffic density detection method and system based on video semantic segmentation
CN120954242A