Video segmentation method based on time sequence context association

By adopting a video segmentation method based on timing context association in ultrasonic video segmentation, using optical flow networks and time multi-scale memory networks, the problem of difficulty in capturing time information and lack of labeled data in the prior art is solved, and high-precision ultrasonic video segmentation and diagnostic efficiency are improved.

CN119941798AActive Publication Date: 2025-05-06XI AN JIAOTONG UNIV

Patent Information

Application Number
CN202510003482.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-05-06
Estimated Expiration
2045-01-02

AI Technical Summary

Technical Problem

The existing medical image/video segmentation methods are difficult to effectively capture the time information between frames in ultrasonic video segmentation tasks, and due to the lack of sufficient labeling data, it is difficult to achieve high-precision frame-by-frame segmentation.

Method used

Using a video segmentation method based on timing context association, the effectiveness pixels in optical flow calculation are extracted through the optical flow network, encoded into context encoding, and a time multi-scale memory network is constructed. Combining the characteristics of the past N frames and short-term features, it matches the similarity of the current frame, and reads out features that are more unified with the current frame.

Benefits of technology

It realizes automatic segmentation of the target cavity/organ in the ultrasound video sequence, assists the physician to automatically locate the target area, reduces manual recognition, improves diagnostic efficiency, and reduces segmentation errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941798A_ABST
    Figure CN119941798A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of video segmentation, and relates to a video segmentation method based on time sequence context association, which comprises the following steps of: 1, acquiring an original video; 2, for adjacent frames in the original video, extracting effective pixels generated in optical flow calculation by using an optical flow network; 3, coding the effective pixels into context coding so as to emphasize the effective positions of the existing features after time change; 4, constructing a time multi-scale memory network in combination with context coding; 5, reading features more unified with the current frame through the time multi-scale memory network; the method not only can be used for natural images in a video format, but also can be used for medical video data with more complex noise conditions, for example, an ultrasonic video is input, a segmentation organ is specified, and a segmentation result of the corresponding organ is output; according to the method, the target cavity / organ can be automatically segmented in the ultrasonic video sequence, so that a physician is assisted in automatically positioning the target area, manual recognition is reduced, and the diagnosis efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of video segmentation and relates to a video segmentation method based on temporal context association. Background Art

[0002] This video segmentation method is not limited to natural images, and is mainly intended to be adapted to more complex medical ultrasound videos. Ultrasound images are commonly used tools in clinical diagnosis, and segmented ultrasound images can be used to calculate some important physiological parameters. Taking cardiac ultrasound as an example, after segmenting different chambers of the heart (left atrium, left ventricle, etc.), the volume of the heart cavity can be accurately measured. By measuring in different cardiac cycles, important indicators such as the ejection fraction of the heart can be obtained, and these data are of key value for evaluating cardiac function. Ultrasound videos can record the dynamic changes of organs or tissues. For example, when observing the ultrasound video of the fetal heart, the morphology and movement trajectory of the heart at different time points can be tracked through video segmentation. This is of great significance for detecting whether the fetal heart has congenital heart disease.

[0003] Ultrasound video segmentation is limited by the following technical difficulties. First, due to factors such as internal tissue heterogeneity, scattering and reflection of ultrasound waves, speckle noise is generated during ultrasound imaging. This noise causes the pixel grayscale values ​​of the image to present a granular distribution, which seriously interferes with the boundary identification of target tissues and organs. Second, the shape and position of human tissues and organs in ultrasound images will change due to individual differences, body posture, respiratory movement and other factors. Third, ultrasound video is a time series data that contains not only the spatial information of the image, but also the time information between frames. This requires the segmentation algorithm to be able to process not only the features of a single frame image, but also the dynamic change information in the video. Fourth, medical images usually lack sufficient annotation information, and the annotation information is even more scarce for data in video format. This makes it difficult for the algorithm to continuously and accurately track the changes in the segmentation target and complete high-precision frame-by-frame segmentation.

[0004] At present, there are many medical image / video segmentation methods, and most of the mainstream methods are based on UNet and Transformer. The UNet-based method has a contraction path and an expansion path. The contraction path extracts high-level semantic features through convolution and pooling operations, and the expansion path gradually restores the resolution through upsampling, and uses jump connections to fuse the features of the contraction path with the features of the corresponding layer of the expansion path. This structure can effectively combine low-level detail information and high-level semantic information. The Transformer-based method uses a self-attention machine to capture long-term dependencies and global information in the image.

[0005] However, the above methods still have limitations in ultrasound video segmentation tasks. The UNet architecture mainly focuses on extracting and fusing the spatial features of a single frame. In ultrasound video segmentation, it is difficult to effectively capture the temporal information between frames, such as the motion trajectory of tissues and organs, the dynamic process of morphological changes, etc. Secondly, when dealing with motion blur and speckle noise in ultrasound videos, the convolution operation of Unet has difficulty distinguishing whether the blurred boundaries are real tissue structure changes or caused by motion, and may mistakenly segment the noise as a target feature. In addition, when the target organs or tissues in the ultrasound video show large movements, morphological changes, or new targets appear, the generalization ability of the Unet model will be tested. Because its training process is mainly based on fixed spatial feature patterns, it may not be able to adjust the segmentation strategy in time for dynamic changes that exceed the range of training samples.

[0006] Transformer requires a lot of memory and computing time to calculate the self-attention matrix, and these models require a lot of labeled data to train to fully exert their advantages. However, in the field of medical ultrasound videos, it is very difficult to obtain large-scale, high-quality labeled data. Secondly, although Transformer can capture global information and long-range dependencies well, in ultrasound video segmentation, it may ignore local details such as tissue and organ boundaries and small lesions.

[0007] It can be seen that the existing medical image / video segmentation methods face the following problems when performing ultrasound video segmentation:

[0008] 1. Due to the presence of speckle noise or unclear boundaries in ultrasound images, it is difficult to accurately segment the target boundaries;

[0009] 2. The morphological changes of moving objects lead to differences in inter-frame features. In addition, there is often a lack of sufficient annotations, which makes it difficult to integrate the existing and insufficient inter-frame features to complete the frame-by-frame segmentation of the target.

[0010] Therefore, a video segmentation method is needed to obtain more sufficient and higher-quality inter-frame features and thus achieve higher-precision segmentation results to solve the above technical problems. Summary of the invention

[0011] The technical solution adopted by the present invention to solve the technical problem is: a video segmentation method based on temporal context association, comprising the following steps:

[0012] Step 1: Get the original video;

[0013] Step 2: For adjacent frames in the original video, use the optical flow network FlowNet to extract valid pixels generated in the optical flow calculation;

[0014] Step 3: Encode the validity pixels into contextual codes to emphasize the valid positions of existing features after time changes;

[0015] Step 4: Combine the context encoding in step 3 to construct a temporal multi-scale memory network, which contains the features of the past N frames (long-term memory) and short-term features (short-term memory);

[0016] Step 5: The features and short-term features of the past N frames are read out by similarity matching with the current frame, and the long short-term memory is obtained based on the past frames before adjustment; after the context encoding adjustment, its features are changed and the similarity matching results are affected, and the features that are more unified with the current frame are read out through the time multi-scale memory network;

[0017] The method of the present application can be adapted to ultrasound images that are more complex than natural images. The trained video segmentation model based on temporal context association is transplanted to the B ultrasound image mobile terminal device, the ultrasound image or video is input and the segmented organs are specified, and the segmentation results of the corresponding organs are output.

[0018] Preferably, taking ultrasound video segmentation as an example, step 1 specifically includes the following sub-steps:

[0019] Step 1-1: Obtain the original ultrasound video from the ultrasound imaging device and annotate it to obtain a pairing mask;

[0020] Step 1-2: Delete the contents irrelevant to the ultrasound scanning sector area in the original ultrasound video;

[0021] Step 1-3: Organize the obtained samples into training set, validation set and test set.

[0022] Preferably, the step 2 specifically includes:

[0023] The two adjacent frames I in the video T-1 and I T Input to the optical flow network F f Get the optical flow at time T

[0024]

[0025] Through optical flow Generate the 2D grid coordinates of the pixels of the mask at time T-1 g T-1 and use Get the pixel two-dimensional grid coordinates of the mask at time T g T :

[0026]

[0027] In formula (1), I T-1 and I T They represent the frame at time T-1 and the frame at time T in the video respectively.

[0028] Preferably, the step 3 specifically includes:

[0029] Pixel 2D grid coordinates based on mask at time T g T Sampling to generate validity pixel mask Then generate the encoding tensor representing the validity of pixels from time T-1 to time T

[0030]

[0031] Pick The four vertices of delimit the attention box to indicate the validity region over time:

[0032]

[0033] In formula (4), σ is the preset threshold value to limit Values ​​greater than the threshold are encoded as context embeddings.

[0034] Preferably, the step 5 specifically includes the following sub-steps:

[0035] Step 5-1: All frames in the video I t Encoding as a tensor To generate a query q t and key

[0036] Step 5-2: Encode the image at time T-1 Image I T-1 and the pairing mask M T-1 Encoded as a value and concatenated with all values ​​before time T-1 to form a long-term memory value The memory read out at T-1 and the short-term memory at T-2 are concatenated to form short-term memory.

[0037] Step 5-3: Use the MSA in formula (5) to adjust the query, key and value of long-term memory and the value of short-term memory:

[0038]

[0039] In formula (6), q T , k T-1 Respectively represent the query of long-term memory, the key of long-term memory, the value of long-term memory, and the value of short-term memory;

[0040] Step 5-4: Perform similarity matching based on the adjusted memory to read out the long-term memory, and integrate the short-term memory as the final read-out multi-scale memory

[0041]

[0042] In formula (7) and formula (8), They represent long-term characteristics and short-term characteristics respectively.

[0043] The beneficial effects of the present invention are:

[0044] 1. The video segmentation method based on temporal context association of the present invention can automatically segment the target cavity / organ in the ultrasound video sequence. For example, by inputting a transesophageal echocardiogram (video) to observe thrombus, the method can segment the left atrial appendage frame by frame (the thrombus is located in the left atrial appendage). Therefore, the present invention can assist doctors to automatically locate the target area, reduce manual identification, and improve diagnostic efficiency.

[0045] 2. The memory network of the present invention can reduce the segmentation error caused by the difference in inter-frame features in video segmentation, and alleviate the situation in the prior art where the video data lacks sufficient annotations and the changes in the target position and shape between frames cause the features to be misaligned with the current frame in time sequence; therefore, the present invention effectively solves the above-mentioned problems that are prevalent in memory networks by modeling the relationship between frames in the time dimension.

[0046] 3. The present invention has great scalability. Considering that ultrasound images contain more high-order noise than natural images, the present invention can not only be applied to ultrasound videos, but also can be migrated to other natural image segmentation scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 It is a schematic diagram of a video segmentation method based on temporal context association of the present invention;

[0048] Figure 2 is a temporal context association process diagram of the present invention;

[0049] Figure 3 It is a comparison diagram of the segmentation results of the present invention and the prior art;

[0050] Figure 4 It is a schematic diagram of the target effect before and after segmentation in the implementation steps of the present invention. DETAILED DESCRIPTION

[0051] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the relevant technologies in the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0052] refer to Figures 1 to 4 In this implementation, the video segmentation method based on temporal context association includes the following steps:

[0053] 1.1. Get the original video;

[0054] 1.2. For adjacent frames in the video, an optical flow network FlowNet is used to extract valid pixels generated in the optical flow calculation;

[0055] 1.3. Encode the validity pixels as context codes to emphasize the valid positions of existing features after time changes;

[0056] 1.4. Combine the context coding constructed in 1.3 to construct a time-multi-scale memory network (the prototype of the memory network here is the long short-term memory network, which is not innovatively proposed by the present invention. The context coding constructed by the present invention is derived through the proposed temporal context network <i.e., the process described in 1.1-1.3>, and then the long short-term memory network is improved by using the context coding to form a time-multi-scale memory network);

[0057] 1.5. In the temporal multi-scale memory network, the features of the past N frames (long-term memory) and short-term features (short-term memory) are read out through similarity matching with the current frame, and the long short-term memory is derived from the past frames before adjustment. After the context encoding adjustment, its features are changed and the similarity matching results are affected, thereby reading out more unified features with the current frame.

[0058] More specific video segmentation methods based on temporal context association are as follows Figure 2 As shown, taking ultrasound video segmentation as an example, the description is as follows:

[0059] 1. Data acquisition

[0060] 1.1 Obtain the original ultrasound video from the ultrasound imaging device and annotate it to obtain the pairing mask;

[0061] 1.2 Delete the content in the video that is not related to the ultrasound scanning sector area;

[0062] 1.3 Organize the obtained samples into training set, validation set and test set.

[0063] 2. Realizing contextual information attention based on temporal context network

[0064] 2.1 For two adjacent frames in the video T-1 and I T , which is input into the optical flow network F f Get the optical flow at time T Generate the pixel 2D grid coordinates g of the mask at time T-1 T-1 and use Get the pixel two-dimensional grid coordinates g of the mask at time T T :

[0065]

[0066]

[0067] 2.2 Pixel 2D grid coordinates based on mask at time T g T Sampling is performed to generate a validity pixel mask Then generate the encoding tensor representing the validity of pixels from time T-1 to time T

[0068]

[0069] Where σ is the preset threshold to limit Values ​​greater than the threshold are encoded as context embeddings.

[0070] 2.3 Take The four vertices of delimit the attention box to indicate the validity region over time:

[0071]

[0072] 3. Realize temporal multi-scale memory network based on encoded context information:

[0073] 3.1 All frames in the video I t (t=1,2,…,T-1,T,…N) is encoded as a tensor To generate a query q t and key

[0074] 3.2 Encoding the image at time T-1 Image I T-1 and the pairing mask M T-1 Encoded as a value and concatenated with all values ​​before time T-1 to form a long-term memory value

[0075] The memory read out at T-1 and the short-term memory at T-2 are concatenated to form short-term memory.

[0076] 3.3 Use MSA to adjust the query, key and value of long-term memory, and the value of short-term memory (short-term memory only involves values, not queries and keys):

[0077]

[0078] 3.4 Perform similarity matching based on the adjusted memory to read out the long-term memory, and integrate the short-term memory as the final read-out multi-scale memory

[0079]

[0080] 4. Deployment and diagnostic application of video segmentation models based on temporal context association.

[0081] The trained temporal context-based video segmentation model is transplanted to the mobile device that obtains B-ultrasound. The ultrasound image or video is input and the segmented organs are specified, and the segmentation results of the corresponding organs are output. If it is video data, the segmentation results are frame by frame, which helps to reduce complex and laborious manual recognition and automatically locate the target organ.

[0082] Example

[0083] Take the segmentation of the left atrial appendage in cardiac ultrasound as an example. Figure 4 As shown in Figure 2, during the ultrasound video segmentation process, Figure 4 The contour line in (a) represents the LAA contour marked before segmentation. At the beginning of segmentation, the model frames the LAA as a larger area based on limited context information, such as Figure 4 As shown in (b), as the segmentation progresses, the model gradually begins to learn the characteristics of the left atrial appendage. At this time, there is still interference from similar cavities, such as Figure 4 As shown in (c), the area below the left side of the left atrial appendage is a similar cavity. So far, many segmentation methods have been difficult to effectively distinguish the similar cavity and the left atrial appendage target area. In this embodiment, with the accumulation of temporal context, the correct area of ​​the left atrial appendage is gradually determined under the guidance of context coding, and the target area is gradually accurately identified, and finally a higher segmentation accuracy is obtained, as shown in FIG. Figure 4 (d) as shown.

[0084] The key points of this embodiment are:

[0085] 1. Temporal context network: Whether it is a medical video (or ultrasound video in a medical video) or a natural image video, one of the biggest problems faced during segmentation is the lack of sufficient labels (even short videos can easily have hundreds or thousands of frames, and the cost of manual frame-by-frame labeling is very high), so it is difficult to obtain sufficient inter-frame information when performing target segmentation. The temporal context network proposed in the present invention is a method based on valid pixels in optical flow propagation. Given two frames, the optical flow can be calculated without additional labeling, which is more efficient than manual labeling. At the same time, the present invention also takes into account that the optical flow will be affected by noise, because we do not directly apply the optical flow, but use the valid pixels in the optical flow calculation to encode the temporal context information, which is not seen in the existing optical flow-related methods.

[0086] 2. Temporal multi-scale memory network: Memory network is a commonly used method in video segmentation, and there are many variants. The present invention improves the long short-term memory network. It cannot be regarded as proposing a new memory network, but the ideas inside are very important. Because in the long short-term memory network, the query and key vectors of the long-term memory are generated once according to the image of the corresponding frame, but their values ​​change frame by frame, which strips away the temporal correlation between the query, key, and value. The present invention improves this through the temporal context network, so that the query and key are adjusted frame by frame to match the long-term memory value that changes frame by frame. Short-term memory is similar. It splices the memory read out at the previous moment frame by frame. The memory at each moment is adjusted by the temporal context network, so that the short-term memory is more aligned with the features required by the current frame.

[0087] 3. The temporal context module and the temporal multi-scale memory module constitute a temporal context-related video segmentation method. It is not a simple segmentation model. Because it can complete the segmentation independently (add a decoder to decode the read memory on the original basis), and can also be used with other segmentation network models (such as inputting the read memory into the decoder of other segmentation models, as prompt information for other models, etc.) to help these segmentation models obtain more sufficient and higher-quality inter-frame features, thereby achieving higher-precision segmentation results.

[0088] In summary, the method of the present invention can automatically segment the target cavity / organ in the ultrasound video sequence, thereby assisting the physician to automatically locate the target area, reducing manual identification and improving the diagnostic efficiency. Therefore, the present invention has a wide range of application prospects in the field of ultrasound video segmentation.

[0089] It should be emphasized that the above are only preferred embodiments of the present invention and do not impose any form of limitation on the present invention. Any simple modification of the above embodiments based on the technical essence of the present invention also falls within the protection scope of the present invention. Other equivalent changes and modifications still fall within the scope of the technical solution of the present invention.

Claims

1. A video segmentation method based on temporal context association, characterized in that: The following steps are involved: Step 1: Get the original video; Step 2: For adjacent frames in the original video, use the optical flow network FlowNet to extract valid pixels generated in the optical flow calculation; Step 3: Encode the validity pixels into contextual codes to emphasize the valid positions of existing features after time changes; Step 4: Combine the context encoding in step 3 to construct a temporal multi-scale memory network; Step 5: Read out features that are more consistent with the current frame through the temporal multi-scale memory network.

2. The video segmentation method based on temporal context association according to claim 1, characterized in that: The step 1 specifically includes the following sub-steps: Step 1-1: Obtain the original video from the video imaging device and annotate it to obtain the pairing mask; Step 1-2: Deleting redundant information in the original video that is irrelevant to the video content, wherein the redundant information includes device information; Step 1-3: Organize the obtained samples into training set, validation set and test set.

3. The video segmentation method based on temporal context association according to claim 1, characterized in that: The step 2 specifically includes: The two adjacent frames I in the video T-1 and I T Input to the optical flow network F f Get the optical flow at time T Through optical flow Generate the 2D grid coordinates of the pixels of the mask at time T-1 g T-1 and use Get the pixel two-dimensional grid coordinates g of the mask at time T T : In formula (1), I T-1 and I T They represent the frame at time T-1 and the frame at time T in the video respectively.

4. The video segmentation method based on temporal context association according to claim 3, characterized in that: The step 3 specifically includes: The 2D grid coordinates g of the pixels based on the mask at time T T Sampling is performed to generate a validity pixel mask Then generate the encoding tensor representing the validity of pixels from time T-1 to time T Pick The four vertices of delimit the attention box to indicate the validity region over time: In formula (4), σ is the preset threshold value to limit Values ​​greater than the threshold are encoded as context embeddings.

5. The video segmentation method based on temporal context association according to claim 4, characterized in that: The step 5 specifically includes the following sub-steps: Step 5-1: All frames in the video I t Encoding as a tensor To generate the query q t and key Step 5-2: Encode the image at time T-1 Image I T-1 and the pairing mask M T-1 Encoded as a value and concatenated with all values ​​before time T-1 to form a long-term memory value The memory read out at T-1 and the short-term memory at T-2 are concatenated to form short-term memory. Step 5-3: Use MSA to adjust the query, key and value of long-term memory and the value of short-term memory: In formula (6), q T , k T-1 , Respectively represent the query of long-term memory, the key of long-term memory, the value of long-term memory, and the value of short-term memory; Step 5-4: Perform similarity matching based on the adjusted memory to read out the long-term memory, and integrate the short-term memory as the final read-out multi-scale memory In formula (7) and formula (8), They represent long-term characteristics and short-term characteristics respectively.

Citation Information

Patent Citations

  • A video semantic segmentation method and device based on prediction for feature propagation

    CN109919044A

  • Lightweight video object segmentation method based on big data memory storage

    CN114882076A

  • Fast video portrait segmentation method and device based on FastPortrait model and medium

    CN115908794A

  • Semi-supervised video target segmentation method and device

    CN117994702A

  • Unified referring video object segmentation network

    US20210383171A1

Cited By

  • Object segmentation method and training method of object segmentation model

    CN122336302A