A video segmentation method based on temporal context correlation
Through the video segmentation method based on timing context correlation, the optical flow network and the time multi-scale memory network are used to solve the problems of inter-frame feature differences and insufficient labeling in ultrasonic videos, and high-precision target segmentation and diagnostic assistance are achieved.
Patent Information
- Application Number
- CN202510003482.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-01-02
AI Technical Summary
The existing medical image/video segmentation methods are difficult to effectively capture the time information between frames in ultrasonic video segmentation tasks, and lack sufficient labeling data, resulting in difficult to accurately segment the target boundaries and differences in inter-frame characteristics lead to segmentation errors.
Using a video segmentation method based on timing context association, the effectiveness pixels are extracted through the optical flow network and encoded into context encoding. Combined with a time multi-scale memory network, it includes features of the past N frames and short-term features, and adjusts the similarity matching to read out more unified features.
It realizes automatic segmentation of target cavity/organ in ultrasonic video sequence, reduces manual recognition, improves diagnostic efficiency, and reduces segmentation errors caused by inter-frame feature differences, and is suitable for segmentation scenarios of ultrasonic video and natural images.
Smart Images

Figure CN119941798B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of video segmentation, and relates to a video segmentation method based on temporal context association. Background Art
[0002] This video segmentation method is not limited to natural images and is mainly intended to be adapted to more complex medical ultrasound videos. Ultrasound images are commonly used tools in clinical diagnosis. Segmenting ultrasound images can be used to calculate some important physiological parameters. Taking cardiac ultrasound as an example, after segmenting different chambers of the heart (left atrium, left ventricle, etc.), the volume of the cardiac chamber can be accurately measured. By measuring at different cardiac cycles, important indicators such as the ejection fraction of the heart can be obtained, and these data are of crucial value for evaluating cardiac function. Ultrasound videos can record the dynamic change process of organs or tissues. For example, when observing an ultrasound video of a fetal heart, through video segmentation, the morphology, movement trajectory, etc. of the heart at different time points can be tracked. This is of great significance for detecting whether the fetal heart has congenital heart disease.
[0003] Ultrasound video segmentation is limited by the following technical difficulties. First, affected by factors such as the inhomogeneity inside tissues, the scattering and reflection of ultrasonic waves, speckle noise will be generated during the ultrasonic imaging process. This noise makes the pixel gray values of the image show a granular distribution, seriously interfering with the boundary recognition of the target tissues and organs. Second, the shape and position of human tissues and organs in ultrasound images will change due to factors such as individual differences, body postures, and respiratory movements. Third, ultrasound videos are a type of time-series data. In addition to containing the spatial information of the images, they also contain the temporal information between frames. This requires the segmentation algorithm to be able to not only process the features of single-frame images but also capture the dynamic change information in the video. Fourth, medical images usually lack sufficient annotation information, and the annotation information for video format data is even more scarce. This makes it difficult for the algorithm to continuously and accurately track the changes of the segmentation target and complete high-precision frame-by-frame segmentation.
[0004] Currently, there are many medical image / video segmentation methods. Most of the mainstream methods are based on UNet and Transformer. The method based on UNet has a contracting path and an expanding path. The contracting path extracts high-level semantic features through convolutional and pooling operations, and the expanding path gradually restores the resolution through upsampling and uses skip connections to fuse the features of the contracting path with the features of the corresponding layers of the expanding path. This structure can effectively combine low-level detail information and high-level semantic information. The method based on Transformer uses self-attention mechanisms to capture the long-term dependencies and global information in the image.
[0005] However, the above methods still have limitations in the ultrasonic video segmentation task. The UNet architecture mainly focuses on extracting and fusing the spatial features of single-frame images. In ultrasonic video segmentation, it is difficult to effectively capture the temporal information between frames, such as the movement trajectories of tissues and organs, the dynamic processes of morphological changes, etc. Secondly, the convolution operation of Unet is difficult to distinguish whether the blurred boundary in processing the motion blur and speckle noise in ultrasonic videos is due to real tissue structure changes or motion, and may mis-segment the noise as target features. In addition, when there are large-scale movements, morphological changes of the target organ or tissue in the ultrasonic video, or new targets appear, the generalization ability of the Unet model will be tested. Because its training process is mainly based on fixed spatial feature patterns, for dynamic change situations beyond the range of training samples, it may not be able to adjust the segmentation strategy in time.
[0006] Transformer requires a large amount of memory and computing time to calculate the self-attention matrix, and these models require a large amount of labeled data for training to fully exert their advantages. However, in the field of medical ultrasonic videos, it is very difficult to obtain large-scale and high-quality labeled data. Secondly, although Transformer can well capture global information and long-range dependencies, in ultrasonic video segmentation, it may ignore local detail information such as tissue and organ boundaries, and minor lesions.
[0007] It can be seen from this that the problems faced by existing medical image / video segmentation methods in ultrasonic video segmentation are as follows:
[0008] 1. Due to the presence of speckle noise or unclear self-boundaries in ultrasonic images, it is difficult to accurately segment the target boundaries.
[0009] 2. The morphological changes of moving objects result in differences in inter-frame features. In addition, there is often a lack of sufficient annotations, making it difficult to comprehensively use the existing and insufficient inter-frame features to complete the frame-by-frame segmentation of the target.
[0010] Therefore, a video segmentation method that can obtain more sufficient and higher-quality inter-frame features and then achieve higher-precision segmentation results is needed to solve the above technical problems. Summary of the Invention
[0011] The technical solution adopted by the present invention to solve the technical problems is: a video segmentation method based on temporal context association, including the following steps:
[0012] Step 1: Obtain the original video;
[0013] Step 2: For adjacent frames in the original video, use the optical flow network FlowNet to extract the valid pixels generated in the optical flow calculation;
[0014] Step 3: Encode the valid pixels into context encoding to emphasize the valid positions of the existing features after temporal changes;
[0015] Step 4: Construct a temporal multi-scale memory network by combining the context encoding in Step 3. The temporal multi-scale memory network contains the features of the past N frames (long-term memory) and short-term features (short-term memory);
[0016] Step 5: Read out the similarity matching between the features of the past N frames and the short-term features and the current frame. Before adjustment, the long-term and short-term memories are obtained based on the past frames; after being adjusted by the context encoding, their features are changed and the similarity matching result is affected, and more unified features with the current frame are read out through the temporal multi-scale memory network;
[0017] The method of this application can be adapted to ultrasonic images that are more complex than natural images. Transplant the trained video segmentation model based on temporal context correlation to the B-ultrasound image mobile device, input the ultrasonic image or video and specify the segmented organ, and output the segmentation result of the corresponding organ.
[0018] Preferably, taking ultrasonic video segmentation as an example, Step 1 specifically includes the following sub-steps:
[0019] Step 1-1: Obtain the original ultrasonic video from the ultrasonic imaging device and perform annotation to obtain paired masks;
[0020] Step 1-2: Delete the content in the original ultrasonic video that is irrelevant to the ultrasonic scanning fan-shaped area;
[0021] Step 1-3: Organize the obtained samples into a training set, a validation set, and a test set.
[0022] More preferably, Step 2 specifically includes:
[0023] Input two adjacent frames I T-1 and I T in the video into the optical flow network F f to obtain the optical flow at time T
[0024]
[0025] Generate the two-dimensional pixel grid coordinates of the mask at time T-1 through the optical flow g T-1 g and use to obtain the two-dimensional pixel grid coordinates of the mask at time T T T :
[0026]
[0027] In formula (1), I T-1 and I T respectively represent the frame at time T-1 and the frame at time T in the video.
[0028] More preferably, step 3 specifically includes:
[0029] Sampling based on the two-dimensional grid coordinates of the pixels of the mask at time T g T to generate a valid pixel mask and further generating an encoded tensor representing the valid pixels when changing from T-1 to T
[0030]
[0031] Taking the four vertices of to delimit an attention box to indicate the valid region when changing over time:
[0032]
[0033] In formula (4), σ is a preset threshold to define that values greater than the threshold are encoded as context embeddings.
[0034] More preferably, step 5 specifically includes the following sub-steps:
[0035] Step 5-1: Encoding all frames I in the video t into a tensor to generate a query q t and a key
[0036] Step 5-2: Encoding the image at time T-1 image I T-1 and the paired mask M T-1 into values and concatenating them with all the values before time T-1 to form the values of the long-term memory Concatenating the memory read out at T-1 and the short-term memory at T-2 to form the short-term memory
[0037] Step 5-3: Using the MSA in formula (5) to adjust the query, key, and value of the long-term memory, and the value of the short-term memory:
[0038]
[0039] In formula (6), q T , k T-1 respectively represent the query of the long-term memory, the key of the long-term memory, the value of the long-term memory, and the value of the short-term memory;
[0040] Step 5-4: Perform similarity matching W based on the adjusted memory to read out the long-term memory and fuse the short-term memory as the finally read multi-scale memory
[0041]
[0042] In Formula (7) and Formula (8), respectively represent the long-term feature and the short-term feature.
[0043] The beneficial effects of the present invention are as follows:
[0044] 1. The video segmentation method based on temporal context association of the present invention can automatically segment the target cavity / organ in the ultrasound video sequence. For example, when inputting a transesophageal echocardiogram (video) to observe a thrombus, this method can segment the left atrial appendage frame by frame (the thrombus is located in the left atrial appendage). Therefore, the present invention can assist physicians in automatically locating the target area, reducing manual recognition, and improving the diagnosis efficiency.
[0045] 2. The memory network of the present invention can reduce the segmentation error caused by the feature difference between frames in video segmentation, and slow down the situation where the features in the existing technology are not temporally aligned with the current frame due to the lack of sufficient annotation of video data and the changes in the position and shape of the target between frames; Therefore, the present invention effectively solves the above problems commonly existing in the memory network by modeling the inter-frame association in the time dimension.
[0046] 3. The present invention has great scalability. Considering that ultrasound images contain more high-order noise than natural images, the present invention can not only be applied to ultrasound videos, but also be migrated to other natural image segmentation scenarios. Description of the Drawings
[0047] Figure 1 is a schematic diagram of a video segmentation method based on temporal context association of the present invention;
[0048] Figure 2 is a process diagram of the temporal context association of the present invention;
[0049] Figure 3 is a comparison diagram of the segmentation results between the present invention and the prior art;
[0050] Figure 4 is a schematic diagram of the target effect before and after segmentation in the implementation steps of the present invention. Detailed Embodiments
[0051] Next, the related technologies in the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0052] Reference Figures 1 to 4 , in this embodiment, the video segmentation method based on temporal context association includes the following steps:
[0053] 1.1. Obtain the original video;
[0054] 1.2. For adjacent frames in the video, use an optical flow network FlowNet to extract the valid pixels generated in the optical flow calculation;
[0055] 1.3. Encode the valid pixels into context encoding to emphasize the valid positions of existing features after temporal changes;
[0056] 1.4. Construct a temporal multi-scale memory network in combination with the context encoding constructed in 1.3 (the prototype of the memory network here is the long short-term memory network, which is not innovatively proposed by the present invention. The context encoding constructed by the present invention is obtained through the proposed temporal context network <i.e., the process described in 1.1 - 1.3>, and then the long short-term memory network is improved using the context encoding to form a temporal multi-scale memory network);
[0057] 1.5. In the temporal multi-scale memory network, it contains the features of the past N frames (long-term memory) and short-term features (short-term memory), which are read out through similarity matching with the current frame, and the long and short-term memories are obtained based on the past frames before adjustment. After being adjusted by the context encoding, their features are changed and the similarity matching results are affected, thereby reading out more unified features with the current frame.
[0058] More specifically, the video segmentation method based on temporal context association is as Figure 2 shown. Taking ultrasonic video segmentation as an example, it is described as follows:
[0059] 1. Data acquisition
[0060] 1.1 Obtain the original ultrasonic video from the ultrasonic imaging device and perform annotation to obtain paired masks;
[0061] 1.2 Delete the content in the video that is irrelevant to the ultrasonic scanning fan-shaped area;
[0062] 1.3 Organize the obtained samples into a training set, a validation set, and a test set.
[0063] 2. Context Information Attention Based on Temporal Context Network
[0064] 2.1 For two adjacent frames I in the video T-1 and I T , input them into the optical flow network F f to obtain the optical flow at time T Generate the two-dimensional pixel grid coordinates g of the mask at time T-1 T-1 , and use to obtain the two-dimensional pixel grid coordinates g of the mask at time T T :
[0065]
[0066]
[0067] 2.2 Based on the two-dimensional pixel grid coordinates of the mask at time T g T Perform sampling to generate a valid pixel mask Furthermore, generate an encoded tensor representing the valid pixels when changing from time T-1 to T
[0068]
[0069] where σ is a preset threshold to limit Values greater than the threshold are encoded as context embeddings
[0070] 2.3 Take of the four vertices to delimit the attention box to indicate the valid region when the time changes:
[0071]
[0072] 3. Implement a Temporal Multi-Scale Memory Network Based on the Encoded Context Information:
[0073] 3.1 Encode all frames I in the video t (t = 1, 2, …, T-1, T, … N) into a tensor to generate a query q t and a key
[0074] 3.2 Encode the image at time T-1 image I T-1 and the paired mask M T-1 into values, and splice them with all the values before time T-1 to form the values of the long-term memory
[0075] Concatenate the memory read out at T-1 and the short-term memory at T-2 to form short-term memory
[0076] 3.3 Use MSA to adjust the query, key, and value of long-term memory, and the value of short-term memory (short-term memory only involves values, not queries and keys):
[0077]
[0078] 3.4 Perform similarity matching W based on the adjusted memory to read out long-term memory, and fuse short-term memory as the finally read multi-scale memory
[0079]
[0080] 4. Deployment and diagnostic application of the video segmentation model based on temporal context association.
[0081] Transfer the trained video segmentation model based on temporal context association to the mobile device for obtaining B-ultrasound, input the ultrasonic image or video and specify the segmented organ, and output the segmentation result of the corresponding organ. If it is video data, the segmentation result is frame-by-frame, which helps to reduce complex and laborious manual recognition and automatically locate the target organ.
[0082] Embodiment
[0083] Take the segmentation of the left atrial appendage in cardiac ultrasound as an example, as Figure 4 shown, during the ultrasonic video segmentation process,[[]] Figure 4 the contour line in (a) represents the contour of the left atrial appendage calibrated before segmentation. At the beginning of segmentation, the model frames a larger area of the left atrial appendage based on limited context information, as Figure 4 shown in (b); as the segmentation progresses, the model begins to gradually learn the features of the left atrial appendage. At this time, there is still interference from similar cavities, as Figure 4 shown in (c), and the area below the left of the left atrial appendage is a similar cavity. So far, many segmentation methods have been difficult to effectively distinguish between similar cavities and the target area of the left atrial appendage. In this embodiment, with the accumulation of temporal context, the correct area of the left atrial appendage is gradually determined under the guidance of context encoding, and the target area is gradually and accurately recognized, and finally a higher segmentation accuracy is obtained, as Figure 4 shown in (d).
[0084] The key point of this embodiment is:
[0085] 1. Temporal context network: Whether it is a medical video (or ultrasound video in a medical video) or a natural image video, one of the biggest problems faced during segmentation is the lack of sufficient labels (even short videos can easily have hundreds or thousands of frames, and the cost of manual frame-by-frame labeling is very high), so it is difficult to obtain sufficient inter-frame information when performing target segmentation. The temporal context network proposed in the present invention is a method based on valid pixels in optical flow propagation. Given two frames, the optical flow can be calculated without additional labeling, which is more efficient than manual labeling. At the same time, the present invention also takes into account that the optical flow will be affected by noise, because we do not directly apply the optical flow, but use the valid pixels in the optical flow calculation to encode the temporal context information, which is not seen in the existing optical flow-related methods.
[0086] 2. Temporal multi-scale memory network: Memory network is a commonly used method in video segmentation, and there are many variants. The present invention improves the long short-term memory network. It cannot be regarded as proposing a new memory network, but the ideas inside are very important. Because in the long short-term memory network, the query and key vectors of the long-term memory are generated once according to the image of the corresponding frame, but their values change frame by frame, which strips away the temporal correlation between the query, key, and value. The present invention improves this through the temporal context network, so that the query and key are adjusted frame by frame to match the long-term memory value that changes frame by frame. Short-term memory is similar. It splices the memory read out at the previous moment frame by frame. The memory at each moment is adjusted by the temporal context network, so that the short-term memory is more aligned with the features required by the current frame.
[0087] 3. The temporal context module and the temporal multi-scale memory module constitute a temporal context-related video segmentation method. It is not a simple segmentation model. Because it can complete the segmentation independently (add a decoder to decode the read memory on the original basis), and can also be used with other segmentation network models (such as inputting the read memory into the decoder of other segmentation models, as prompt information for other models, etc.) to help these segmentation models obtain more sufficient and higher-quality inter-frame features, thereby achieving higher-precision segmentation results.
[0088] In summary, the method of the present invention can automatically segment the target cavity / organ in the ultrasound video sequence, thereby assisting the physician to automatically locate the target area, reducing manual identification and improving the diagnostic efficiency. Therefore, the present invention has a wide range of application prospects in the field of ultrasound video segmentation.
[0089] It should be emphasized that the above are only preferred embodiments of the present invention and do not impose any form of limitation on the present invention. Any simple modification of the above embodiments based on the technical essence of the present invention also falls within the protection scope of the present invention. Other equivalent changes and modifications still fall within the scope of the technical solution of the present invention.
Claims
1. A video segmentation method based on temporal context association, characterized in that, It includes the following steps: Step 1: Obtain the original video; Step 2: For adjacent frames in the original video, use the optical flow network FlowNet to extract the valid pixels generated in the optical flow calculation; Step 3: Encode the valid pixels into context encoding to emphasize the valid positions of existing features after temporal changes; Step 4: Construct a temporal multi-scale memory network by combining the context encoding in Step 3; Step 5: Read out features that are more unified with the current frame through the temporal multi-scale memory network; The specific content of Step 3 includes: The two-dimensional grid coordinates g of pixels based on the mask at time T T are sampled to generate a valid pixel mask and then a coded tensor representing the valid pixels when changing from time T-1 to time T is produced Take The four vertices of are used to delimit an attention box to indicate the valid region when the time changes: In formula (4), σ is a preset threshold to define values greater than the threshold are encoded as context embeddings; The specific content of Step 5 includes the following sub-steps: Step 5-1: Encode all frames I in the video t into tensors to generate queries q t and keys Step 5-2: Encode the image at time T-1 Image I T-1 and the paired mask M T-1 Encode them as values and concatenate with all the values before time T-1 to form the values of long-term memory Concatenate the memory read at T-1 and the short-term memory at time T-2 to form the short-term memory Step 5-3: Use MSA to adjust the queries, keys, and values of the long-term memory, and the values of the short-term memory; In formula (6), respectively represent the query of long-term memory, the key of long-term memory, the value of long-term memory, and the value of short-term memory; Step 5-4: Perform similarity matching W based on the adjusted memory to read out the long-term memory and fuse the short-term memory as the finally read out multi-scale memory In Formula (7) and Formula (8), respectively represent long-term features and short-term features.
2. The video segmentation method based on temporal context association according to claim 1, characterized in that, The specific content of Step 1 includes the following sub-steps: Step 1-1: Obtain the original video from the video imaging device and perform annotation to obtain the paired mask; Step 1-2: Delete the redundant information in the original video that is irrelevant to the video content, and the redundant information includes device information; Step 1-3: Organize the obtained samples into a training set, a validation set, and a test set.
3. A video segmentation method based on temporal context association according to claim 1, characterized in that, The specific content of Step 2 includes: Input two adjacent I-frames in the video T-1 and I T into the optical flow network F f to obtain the optical flow at time T Through optical flow Generate the two-dimensional grid coordinates g of the pixels of the mask at time T-1 T-1 , and use to obtain the two-dimensional grid coordinates g of the pixels of the mask at time T T : In formula (1), I T-1 and I T respectively represent the frame at time T-1 and the frame at time T in the video.
Citation Information
Patent Citations
Lightweight video object segmentation method based on big data memory storage
CN114882076A
Semi-supervised video target segmentation method and device
CN117994702A