Method, apparatus for processing moving image, electronic device and medium

JP2024150410A5Pending Publication Date: 2026-06-04NEC CORP

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
NEC CORP
Filing Date
2024-03-26
Publication Date
2026-06-04

AI Technical Summary

Technical Problem

Existing endoscopic systems require significant medical expertise to accurately detect and highlight blood vessels and lesions, as current methods struggle with spatiotemporal consistency in video data, leading to unreliable edge enhancement.

Method used

A video processing method that utilizes edge detection across multiple frames to establish spatial offsets and update edge attributes based on registration, incorporating spatiotemporal cues to enhance edge detection and provide local hints for lesions.

Benefits of technology

Improves the accuracy and reliability of edge detection in endoscopic videos, reducing the reliance on medical expertise by providing robust edge enhancement and visual aids for lesion identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide a method for detecting and enhancing a blood vessel or a lesioned portion in an endoscopic moving image, offering localized hints for a lesion and supporting an inspection by a doctor.SOLUTION: A method includes obtaining a first edge of a first frame of a moving image and a second edge of at least one adjacent region frame of the first frame. The method further includes determining a spatial offset from the pixel of the first edge to the corresponding pixel of the second edge on the basis of registration between the first frame and at least one adjacent region frame. The method further includes updating a first edge attribute of the pixel of the first edge on the basis of at least a second edge attribute of the corresponding pixel of the second edge and the spatial offset. By utilizing a time space clue of the moving image, the method improves accuracy and reliability of edge detection and enhancement, reduces adverse effects from low-quality images in the moving image, and enhances robustness of moving image processing.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] TECHNICAL FIELD Embodiments of the present disclosure relate to the field of image processing technology, and more particularly, to video processing methods, apparatus, electronic devices, computer-readable storage media, and computer program products. [Background technology]

[0002] Endoscopy is a type of optical equipment examination that is sent from outside the body through the natural lumen of the human body to examine internal diseases. When using it, doctors introduce the endoscope into the organ to be examined, operate the endoscope to move it, and directly look into the relevant area, and record images and videos.

[0003] Some endoscopy systems can enhance the contrast between blood vessels and the surroundings, and can also magnify tens to hundreds of times to help doctors observe vascular structures and lesions more clearly. However, direct observation of vascular structures and lesions with an endoscopy system requires a doctor to have strong and rich experience. Therefore, it is necessary to provide an improvement scheme to detect and highlight blood vessels or lesions in endoscopy videos, and give local hints to the lesions to help doctors complete the examination. Summary of the Invention [Problem to be solved by the invention]

[0004] In view of this, an embodiment of the present disclosure provides a technical proposal for video processing. [Means for solving the problem]

[0005] According to a first aspect of the present disclosure, there is provided a video processing method. The method includes obtaining a first edge of a first frame of a video and a second edge of at least one adjacent region frame of the first frame. The method further includes determining a spatial offset from a pixel of the first edge to a corresponding pixel of the second edge based on a registration between the first frame and the at least one adjacent region frame. The method further includes updating a first edge attribute of the pixel of the first edge based on at least a second edge attribute of the corresponding pixel of the second edge and the spatial offset.

[0006] According to a second aspect of the present disclosure, there is provided a video processing apparatus. The apparatus includes an edge acquisition unit, a spatial offset determination unit, and an update unit. The edge acquisition unit is configured to acquire a first edge of a first frame of the video and a second edge of at least one adjacent region frame of the first frame. The spatial offset determination unit is configured to determine a spatial offset from a pixel of the first edge to a corresponding pixel of the second edge based on a registration between the first frame and the at least one adjacent region frame. The update unit is configured to update a first edge attribute of the pixel of the first edge based at least on the second edge attribute of the corresponding pixel and the spatial offset.

[0007] According to a third aspect of the present disclosure, there is provided an electronic device comprising at least one processing unit and a memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, performing the first aspect of the present disclosure.

[0008] According to a fourth aspect of the present disclosure, there is provided a computer-readable storage medium comprising machine-executable instructions which, when executed by a device, cause the device to perform the method according to the first aspect of the present disclosure.

[0009] According to a fifth aspect of the present disclosure, there is provided a computer program product comprising machine executable instructions which, when executed by a device, cause said device to perform a method according to the first aspect of the present disclosure.

[0010] The following sections are provided to introduce in a simplified form a selection of concepts that are further described in the specific embodiments below, and are not intended to indicate critical or required features of the disclosure, nor are they intended to limit the scope of the disclosure. [Brief description of the drawings]

[0011] The above and other objects, features, and advantages of the present disclosure will become more apparent by describing in more detail the exemplary embodiments of the present disclosure with reference to the accompanying drawings, in which like reference numerals generally represent like parts in the exemplary embodiments of the present disclosure. [Figure 1A] 1 illustrates a schematic diagram of an exemplary environment in which several embodiments of the present disclosure may be implemented. [Figure 1B] 1 illustrates example video screenshots to which embodiments of the present disclosure may be applied. [Diagram 2] 1 shows a schematic flowchart of a video processing method according to an embodiment of the present disclosure. [Diagram 3] 1 shows a schematic diagram of a video processing system according to an embodiment of the present disclosure. [Figure 4A] 1 shows a schematic diagram of applying edge detection to a moving image according to an embodiment of the present disclosure; [Figure 4B] 1 shows a schematic diagram of applying image registration to a moving image according to an embodiment of the present disclosure. [Diagram 5] 1 shows a schematic flow chart of a process for edge enhancement according to an embodiment of the present disclosure. [Figure 6] 1 shows a schematic flow chart of a process for contour tracking according to an embodiment of the present disclosure. [Figure 7]1 shows a schematic flow chart of a process for generating contour segments according to an embodiment of the present disclosure. [Figure 8] 1 shows a schematic flow chart of a process for contour matching according to an embodiment of the present disclosure. [Figure 9] 1 shows a schematic block diagram of a video processing device according to an embodiment of the present disclosure. [Figure 10] FIG. 1 shows a schematic block diagram of an example device that can be used to implement embodiments of the present disclosed subject matter. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0012] Hereinafter, preferred embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Preferred embodiments of the present disclosure are shown in the accompanying drawings, but it should be understood that the present disclosure should not be limited to the embodiments described herein and can be implemented in various forms. Rather, these embodiments are provided to make the present disclosure clearer and more complete, and to fully convey the scope of the present disclosure to those skilled in the art.

[0013] As used herein, the term "including" and variations thereof refer to an open inclusion, i.e., "including, but not limited to." Unless otherwise noted, the term "or" refers to "and / or." The term "based on" refers to "based at least in part on." The terms "one exemplary embodiment" and "one embodiment" refer to "at least one exemplary embodiment." The term "another embodiment" refers to "at least one other embodiment." Terms such as "first," "second," and the like can refer to different objects or the same object. Other explicit and implicit definitions may be included below.

[0014] As mentioned above, some endoscope systems can enhance the contrast between blood vessels and the surroundings, which is helpful for observing blood vessels, and can also magnify tens to hundreds of times to make vascular structures and lesions more clearly visible. For example, there is a gastroscope or intestinal endoscope with narrow band imaging (NBI) + magnifying endoscopy (ME), in which NBI can highlight capillaries and ME can magnify capillaries. Compared with ordinary white light endoscopy, NBI + ME is a lesion inspection method that has been increasingly used in recent years, and can observe the lesion situation of the mucosal surface from the microstructure. However, to directly observe the lesion of the edge of the capillary with NBI + ME, a doctor's strong and rich experience is required, and it is useful to perform edge enhancement on the capillary and provide local hints for the lesion.

[0015] Some methods use deep learning techniques to extract microvascular and microstructural features from NBI+ME images and present the characterized images to endoscopists. However, these methods focus on single frame images to extract vascular features, and are not suitable for video-format image data obtained during endoscopy. For example, edge enhancement of single frame images will not match the actual scene due to image quality effects such as blur and lighting. And these methods only perform image enhancement on a frame-by-frame basis, which lacks spatiotemporal cues and the same edges are not preserved between video frames, making it difficult to perform spatiotemporal consistency analysis in sequence.

[0016] In view of this, an image enhancement method that better matches the actual scene of endoscopy is needed to effectively enhance edge features of the video to help doctors view the video data obtained by the examination. According to an embodiment of the present disclosure, a video processing method is provided. The method obtains edges of a given frame of a video and edges of an adjacent region frame of the given frame through edge detection. Through edge detection, each pixel at the edge has a respective edge attribute. In some embodiments, the edge attribute may be, for example, the width of color change of the image shown in the form of a probability, or the degree of change of the lesion shown in the form of a concept. Then, a registration is performed between the given frame and the adjacent region frame, and the corresponding pixel at the edge of the adjacent region frame for each pixel at the edge of the given frame is determined, i.e., a mapping between pixels between the two frames is established, and thus a spatial offset is obtained from a pixel of the given frame to a corresponding pixel of the adjacent region frame. The spatial offset may be, for example, a difference between the coordinates of the pixels. Then, adjust the edge attributes of the pixels at the edge of the given frame according to the edge attributes of the neighboring region frames and the determined spatial offset. Such a scheme can utilize the spatiotemporal cues of the video to improve the accuracy and reliability of edge detection and enhancement, reduce the adverse effects of low quality images in the video, and improve the robustness of video processing. Hereinafter, the implementation details of the embodiment of the present disclosure will be described in detail with reference to Figures 1 to 8.

[0017] FIG. 1A illustrates a schematic diagram of an exemplary environment 100 in which embodiments of the present disclosure may be implemented. As illustrated, the environment 100 relates to an endoscopic system including an optional control device 101 operable by an operator to control the movement of an endoscope lens 102 within the human body to capture images and videos. The captured images and videos may be transmitted and stored in a computing device 105. A display 103 is coupled to the computing device 105 and may display the images and videos captured through the lens 102 in real time. In some embodiments, the endoscopic system may be a narrow band imaging and magnifying endoscope (NBI+ME), and the resulting video may be an NBI+ME video. Although FIG. 1 illustrates an exemplary gastroendoscopic system, it should be understood that embodiments of the present disclosure may be applicable to other endoscopic systems, including, but not limited to, ENT endoscopes, oral endoscopes, dental endoscopes, neuroscopes, urethrocystoscopes, electrotomies, laparoscopes, arthroscopes, sinusoscopes, laryngoscopes, enteroscopes, and the like.

[0018] The computing device 105 may be a general-purpose computing device or a device dedicated to endoscopy, and is configured for image and video processing, particularly medical image processing. The computing device 105 may be a terminal or a server device. If the computing device 105 is a server, it may be an independent server, or a server network or server cluster of servers, including but not limited to a computer, a network host, a single network server, a set of multiple network servers, or a cloud server built with multiple servers. The computing device 105 may comprise a desktop terminal or a mobile terminal, such as but not limited to a desktop, a mobile phone, a tablet, a laptop, a medical support instrument, etc.

[0019] As mentioned above, the conventional solutions cannot stably and accurately inspect and present the edges of microstructures (e.g., the edges of capillaries and lesions) in the video, and highly require the specialized technical knowledge and experience of doctors. According to the embodiment of the present disclosure, the computing device 105 can enhance edges of endoscopic videos and provide local hints for lesions, helping doctors perform more effective medical examinations.

[0020] An exemplary environment in which embodiments of the present disclosure may be implemented has been described above with reference to Fig. 1A. It should be understood that Fig. 1A is merely schematic, and the environment may include more modules or systems, or may omit some modules or systems, or may recombine the illustrated modules or systems. The embodiments of the present disclosure may be implemented in environments different from the environment illustrated in Fig. 1A, and the present disclosure is not limited thereto.

[0021] FIG. 1B shows a screenshot of an exemplary video to which an embodiment of the present disclosure may be applied. The screenshot in FIG. 1B is a gastrointestinal endoscopy image acquired by an NBI+ME system. As shown, the image contains abundant capillary structures and possible areas of lesions, and there may be low-quality video frames, making it difficult for a physician to accurately and quickly locate the area of ​​interest therefrom. An embodiment of the present disclosure is suitable for edge detection and enhancement on such videos, and provides corresponding local hints (e.g., tracking the edges of interest) to reduce the difficulty of browsing the videos.

[0022] Note that while the video processing methods and systems according to the embodiments of the present disclosure are described with reference to NBI+ME video, the embodiments of the present disclosure may be equally applicable when performing similar processing on any other type of video or image, and the present disclosure is not limited thereto.

[0023] 2 shows a schematic flow chart of a video processing method 200 according to an embodiment of the present disclosure. The method 200 may be implemented, for example, by the computing device 105 shown in FIG. 1. It should be understood that the method 200 may also include additional operations not shown and / or omit operations shown, and the scope of the present disclosure is not limited in this respect. The method 200 is described in detail below with reference to FIG. 1A.

[0024] The method 200 relates to video processing implemented by the computing device 105. The video may be endoscopic video obtained by an endoscopic sampling procedure, such as narrow band imaging (NBI) and magnifying endoscopy (ME) video, or other types of video. In some embodiments, the computing device 105 may receive the video in real time or near real time during an endoscopy procedure and process the video. Alternatively, the computing device 105 may receive and store the video, temporarily without processing, and process it after the complete video is received.

[0025] In block 210, a first edge of a first frame of a video and a second edge of at least one adjacent region frame of the first frame are obtained. Here, the adjacent region frame may be defined as some frames located before or after the first frame. For example, 1, 2, 5, 10, 20, or any other number of frames before and after the first frame (i.e., the current frame) may be defined as the adjacent region frame, and the number of frames before and after may be the same or different, and is not limited in this disclosure. In some implementations, the adjacent region frame may include frames before, after, or in two directions before and after the first frame.

[0026] The computing device 105 performs edge detection on the video to generate edges for each frame of the video. Edge detection may be for different purposes and may use different algorithms. For example, the edge detection algorithm may include an algorithm for microstructure edges (e.g., capillaries) or lesion edges. Thus, the obtained first and second edges may be microstructure edges or lesion edges. In some implementations, edge detection may be based on deep network prediction. For example, annotation samples may be collected to train a deep network, and then the trained deep network may be used to generate predictions for edge attributes for each frame. Alternatively, a feature extraction classifier (e.g., a support vector machine) may be employed to train and predict, or image features may be extracted and a weighting scheme may be employed to obtain category probabilities and compared with a threshold, etc.

[0027] The edge detection may generate an edge attribute for each pixel in the first frame of the video and the adjacent region frames of the first frame, and the edge attribute may indicate the probability that the corresponding pixel belongs to an edge. For example, for microstructure edge detection, the edge attribute may have a range of 0 to 255, from small to large representing a weak microstructure edge to a strong microstructure edge. For lesion detection, the edge attribute may also have a range of 0 to 255. This disclosure does not limit the range of the edge attribute. The computing device 105 may then determine the first edge and the second edge by determining pixels in the first frame and the adjacent region frames whose edge attribute is greater than a threshold (e.g., 0 or other value) as pixels in the first edge and the second edge, respectively. Note that the edge attribute or edge probability obtained here is calculated by the computing device 105 based on a single frame and may be adjusted or optimized using spatiotemporal information of the video.

[0028] In block 220, a spatial offset from a pixel of the first edge to a corresponding pixel of the second edge is determined based on the registration between the first frame and at least one adjacent region frame. Here, the computing device 105 employs image registration to obtain a mapping relationship of edge points between frames, and obtains a spatial offset of edge points. The spatial offset can represent a difference in coordinates between two frames, for example, a two-dimensional vector, for each pixel.

[0029] In some embodiments, the computing device 105 may register only adjacent frames (e.g., frames immediately following the first frame) to improve calculation efficiency and ensure real-time performance. For non-adjacent frames, the spatial offsets of adjacent frames may be accumulated. In some embodiments, to further improve calculation efficiency, matching may be performed only on edges. The edge texture has clear characteristics, so the matching degree is high. This allows each pixel on the first edge to be mapped to a pixel on the second edge of the adjacent region frame.

[0030] Methods used for frame-to-frame registration may include sparse optical flow fields, such as the KLT method, which focuses on edge points and then uses local gradient changes to calculate pixel mapping points. Alternatively, registration methods may include deep learning methods, such as FlowNet, which uses deep learning networks to automatically obtain spatial offsets of pixels.

[0031] In block 230, the first edge attribute of the pixel of the first edge is updated based on at least the second edge attribute of the corresponding pixel of the second edge and the spatial offset. The computing device 105 uses the neighboring region frames of the first frame to enhance the edge attribute of the first edge of the first frame, considering that the single frame edge detection is sensitive to image quality and deviation between the edge attribute obtained and the true attribute in the case of exposure, low light, dirt, and blur.

[0032] In some embodiments, the computing device 105 may search for some frames that meet pre-set conditions, such as frames with high quality and satisfying offset degree, from the neighboring region frames preceding and following the first frame, and use these frames to enhance the edge of the first frame. Specifically, for a given pixel having a first edge attribute in the first edge of the first frame, the computing device 105 may accumulate a second edge attribute of a corresponding pixel in the neighboring region frame of the pixel to the first edge attribute to adjust the edge attribute of the given pixel. For example, a weighted summation scheme may be used to adjust the edge attribute of the first edge. The weight for the edge attribute of the neighboring region frame may depend on the spatial offset between the pixels. The weight may also depend on the time between the neighboring region frame and the given frame.

[0033] This method introduces spatiotemporal consistency information about edges in a video, and the edge attributes of the edge pixels in the first frame include the edge attributes of its adjacent region frames, and the edge information of the adjacent region frames is obtained from high-quality frames, thereby reducing the adverse effects of low-quality images in the video and improving the accuracy and reliability of edge detection and enhancement.

[0034] Fig. 3 shows a schematic diagram of a video processing system 300 according to an embodiment of the present disclosure. The video processing system 300 may be implemented in the computing device 105 in the form of software or hardware and is suitable for implementing the method 200 described with reference to Fig. 2. It should be understood that the video processing system 300 shown in Fig. 3 is only exemplary and may include more or fewer modules or units, and some modules shown in Fig. 3 may be divided into multiple modules, and two or more modules may be implemented in the same module.

[0035] The video processing system 300 can receive endoscopic video frames 301, 302, 303, etc. as input in real time, where chronologically, frame 301 is the frame before frame 302, which is the frame before frame 303, and so on. Alternatively, the video processing system 300 can read video stored in a storage device of the computing device 105.

[0036] The video processing system 300 includes an edge detection module 310, an inter-frame registration module 320, and a spatio-temporal edge update and vectorization module 330. The edge detection module 310 is used to extract edges in a single frame. By applying an edge detection algorithm, the edge detection module 310 may generate, for each pixel of the frame, an edge attribute indicating, for example, a probability that the pixel may belong to an edge. It should be understood that such a probability is relative and not absolute. To improve the real-time performance of the video processing, the edge detection module 310 may use an edge detection algorithm that requires less computational resources. In some embodiments, the edge detection module 310 may apply multiple edge detection algorithms to generate multiple types of edges, for example, microstructure edges, lesion edges, etc. Thus, the edge detection module 310 generates edge attributes for frame 301, frame 302, frame 303, etc.

[0037] 4A shows a schematic diagram of applying edge detection to a video according to an embodiment of the present disclosure. The edge detection module 310 performs edge detection on a given frame of a video as shown on the left side of FIG. 4A, obtains edge attributes for each pixel, and then generates an edge detection result based on the edge attributes as shown on the right side of FIG. 4A.

[0038] The inter-frame registration module 320 is used to obtain inter-frame offsets for pixels using image registration in a video frame sequence, and establish a mapping relationship between edges between frames. Considering that edge textures have distinct features and a high degree of matching, matching may be performed only on the edges of frames. That is, pixels at the edge of a given frame are mapped to corresponding pixels at the edge of an adjacent region frame. In some embodiments, to ensure real-time performance, the inter-frame registration module 320 may perform registration only on two adjacent frames. For example, the inter-frame registration module 320 may perform registration between frames 301 and 302, and map the edge pixels of frame 301 to the edge pixels of frame 302, and may perform registration between frames 302 and 302, and map the edge pixels of frame 302 to the edge pixels of frame 303, and so on.

[0039] Using the frame-to-frame registration, an edge mapping relationship can be obtained. The edge mapping relationship may be expressed as a spatial offset between two pixels, such as the difference between coordinates. Based on the superposition (e.g., vector addition) of the spatial offsets between pixels of adjacent frames, the spatial offset of the pixels of the edge of a given frame in any adjacent region frame can be obtained.

[0040] 4B shows a screenshot of applying image registration to a video according to an embodiment of the present disclosure. The inter-frame registration module 320 performs registration sequentially on a series of adjacent frames on the left side of FIG. 4B to obtain a corresponding pixel of a given pixel in the next frame and determine a spatial offset of the two pixels. As shown in FIG. 4B, a given pixel may be presented as a motion between frames (as indicated by the arrow) using the obtained spatial offset.

[0041] The spatio-temporal edge update and vectorization module 330 is used to utilize information of neighboring region frames (e.g., spatial offset of pixels with time, edge attributes, etc.) according to the spatio-temporal consistency information of the video, and adjust the edge attributes of the frames to highlight edges in the video. Additionally, the spatio-temporal edge update and vectorization module 330 is also used to establish and track the contour segment of the object of interest from the edge.

[0042] 5 shows a schematic flow chart of a process 500 for edge enhancement according to an embodiment of the present disclosure. The process 500 may be an example embodiment of the block 230 shown in FIG. 2 and may be implemented by the spatio-temporal edge update and vectorization module 330 shown in FIG.

[0043] The process 500 is used to enhance edges of a given frame (i.e., a "first frame") based on neighboring region frames. The neighboring region frames may be at least one frame in the vicinity of the given frame whose edge attributes are updated. The range of the neighboring region frames may consist of or refer to a number of frames preceding the first frame, a number of frames following the first frame. In some implementations, the neighboring region frames may include frames before the first frame, or may include frames after the first frame, or may include both. The present disclosure does not limit the range of the neighboring region frames. Overall, the process 500 searches for frames with high quality and offset degrees as matching frames from the vicinity of the first frame, and utilizes these matching frames to update the edge attributes of the first frame to achieve edge enhancement.

[0044] In block 501, one adjacent region frame of a first frame is obtained. In some embodiments, the adjacent region of the first frame may be searched in a certain order, for example, from the earliest adjacent region frame to the latest adjacent region frame according to the frame time. Alternatively, the search may be performed based on the temporal distance between the adjacent region frame and the first frame, starting with the nearest adjacent region frame from the first frame, and successively searching the farther adjacent region frames to obtain one adjacent region frame.

[0045] In block 502, it is determined whether the quality of the obtained adjacent region frame exceeds a threshold. The quality may be obtained by deep learning or by feature extraction classifier. If the quality exceeds the threshold, the process 500 proceeds to block 503; if not, the adjacent region frame is considered to be a low-quality frame, and the use of this frame to enhance the first frame may be abandoned, and the process 500 may proceed to block 505.

[0046] In block 503, it is determined whether the average spatial offset between the edge of the adjacent region frame and the edge of the first frame is within a preset range. The average spatial offset refers to the average value of the spatial offset between the pixel at the edge of the first frame and the corresponding pixel in the adjacent region frame. For example, the magnitude of the spatial offset of each pixel may be calculated, and the average value of the magnitude of the spatial offset may be calculated. The preset range may include a lower threshold and an upper threshold, and if the average spatial offset is greater than the lower threshold and less than the upper threshold, the process 500 proceeds to block 504. If not, the frame may be considered not suitable to be used for edge enhancement, and the frame may be abandoned and the process 500 may proceed to block 505.

[0047] In block 504, an adjacent region frame is selected. The selected adjacent region frame is used for edge enhancement of the first frame. The frame of the selected adjacent region frame may be recorded, and the edge attributes of the edges of the first frame may be updated according to all the selected adjacent region frames after waiting for all the adjacent region frames to have been processed.

[0048] At block 505, it is determined whether all adjacent region frames have been traversed. In some embodiments, a threshold corresponding to the total number of frames in the adjacent region range may be set, and when the number of processed adjacent region frames reaches this threshold, all frames in the adjacent region range of the first frame are considered to have been processed. If at block 505, it is determined that there are no unprocessed adjacent region frames, the process 500 proceeds to block 506 and updates the edge attributes of the first edge of the first frame using the selected matching frame. If it is determined that there are unprocessed adjacent region frames, the process 500 returns to block 501.

[0049] In response to all the adjacent region frames being processed and all the selected adjacent region frames being determined, update edge attributes of edge pixels of the first frame in block 506. For each pixel in the first edge of the first frame, its corresponding pixel and corresponding spatial offset are determined from the selected adjacent region frames, and the updated edge attribute of the pixel is determined based on the following equation: TIFF2024150410000002.tif12152, where Edge()' is the edge attribute after updating the pixel, Edge() is the edge attribute obtained by edge detection, Pi is the pixel at the edge of the first frame, P j are the corresponding pixels at the edge of the neighboring region frame, Dj represents the spatial offset distance, and δ and σ are empirical parameters.

[0050] 6 shows a schematic flow chart of a process 600 for contour tracking according to an embodiment of the present disclosure. The process 600 may be an additional operation of the method 200 shown in FIG. 2 and may be implemented by the spatio-temporal edge update and vectorization module 330 shown in FIG.

[0051] The pixel-to-pixel mapping relationship and updated edge attributes obtained by edge matching are discrete and irregular features. However, doctors are usually interested in the contour segments of lesions, which are connected in series to track pixel points with the same attributes and are conducive to unified analysis.

[0052] In block 610, a contour segment of the object of interest in the first frame is determined based on the updated edge attributes. The contour segment may be, for example, a combination of the detected edge segments and may reflect a structure or lesion having clinical significance. To facilitate subsequent tracking, the pixel points of the edge may be vectorized into a contour segment according to the edge attributes. For every frame in the video, a contour segment may be determined in the frame and tracked throughout the video. Alternatively, a contour segment may be determined for a portion of frames in the adjacent area range of the first frame and tracked in a portion of the video. For each frame, one or more contour segments may be determined. In some implementations, there may be multiple contour segments in a single frame, and a contour segment to be tracked is selected and tracked according to user input.

[0053] In block 620, a global spatial offset from the contour segment to the subsequent frame is determined based on the registration between the first frame and the subsequent frame of the first frame. Depending on the implementation, the subsequent frame may be an adjacent region frame that is immediately adjacent to the first frame. In such a case, the adjacent frames may be registered to obtain the spatial offset from the pixels in the contour segment to the corresponding pixels, and the average value of the spatial offsets of all the pixels of the contour segment may be determined as the global spatial offset. Alternatively, the subsequent frame may be an adjacent region frame that is a certain distance away from the first frame. The spatial offsets obtained by the registration of the adjacent frames may be accumulated to obtain the spatial offset from the pixels of the contour segment of the first frame to the corresponding pixels of the subsequent frame, and thus the global spatial offset may be determined.

[0054] In block 630, a search is performed for another contour segment in the subsequent frame that matches the contour segment in the first frame based on the global spatial offset. In some embodiments, the other contour segment to be matched is the contour segment in the subsequent frame that has the smallest distance from the contour segment in the first frame.

[0055] The process 600 may be performed chronologically for frames in a video to track an object of interest. In some embodiments, as the video plays, a contour segment may be highlighted in a first frame, another contour segment in a subsequent frame, and so on, providing a visual aid to assist the physician in tracking an object of interest in the video.

[0056] 7 shows a schematic diagram of a process 700 for generating contour segments according to an embodiment of the present disclosure. The process 700 may be an example implementation of the operation shown in block 610 of FIG.

[0057] In block 710, super-pixel extraction is performed on the first frame. In the first frame, the super-pixel extraction is performed with the updated edge attributes as a guide to obtain super-pixel contours. The super-pixel extraction method includes, but is not limited to, Simple Linear Iterative Clustering (SLIC), etc. Here, the original gradient calculation of the super-pixel extraction is replaced with the updated edge attributes to ensure that the super-pixel contours are close to high probability edges. Then, a contour segment of the object of interest may be generated based on the super-pixel contours.

[0058] Specifically, segments are obtained by edge vectorization in block 720. In some embodiments, the super-pixel contour may be divided into a set of segments based on intersection points in the extracted super-pixel contour. Each segment may include a sequence of pixels.

[0059] In block 730, a contour segment of the object of interest is obtained by merging the segments. For the obtained set of segments, it is analyzed which adjacent segments are similar, and the similar adjacent segments are combined to form a contour segment of the object of interest. Specifically, for two adjacent segments, the similarity between the two segments may be calculated, and if the similarity exceeds a threshold, they may be merged to generate a contour segment.

[0060] In some embodiments, the cues for determining the similarity of adjacent segments may include intra-frame cues and inter-frame cues. The intra-frame cues may include the absolute value of the difference between the gradients of two adjacent segments. The segment gradient may be defined as the average gradient {gx(f),gy(f)} of all pixels in a segment f, in which the gradient direction of a pixel is the offset amount of the upper and lower adjacent region pixels of this pixel. In some embodiments, the difference between the Euclidean distances of the segment gradients may be the corresponding inter-frame cues, in which the Euclidean distance is the Euclidean norm of the segment gradients {gx(f),gy(f)}. The inter-frame cues may further include the absolute value of the difference between the edge probabilities Edge(f) of two adjacent segments, in which Edge(f) is the average value of the edge probabilities of all pixels on the segment. Here, the edge probabilities may include the edge probabilities of the microstructure edges and the edge probabilities of the lesion edges.

[0061] The inter-frame cues may include the difference between the spatial offsets {dx(f), dy(f)} of two adjacent segments. The spatial offset of a segment may be defined as the average value of the spatial offsets from all pixels in segment f to the corresponding pixels in the adjacent region frames.

[0062] In some embodiments, the similarity of adjacent segments may be determined according to the following equation: TIFF2024150410000003.tif11164 Among them, Err grad is the difference in segment gradients, and Err edge is the difference between the edge probabilities of the segments, and Err disease is the difference in the segment foci probability, Err offset is the difference in spatial offset of the segments, r1, r2, r3, and r4 are the weights of each cue, and δ1, δ2, δ3, and δ4 are regularization parameters. rr If is greater than some threshold, then merge two adjacent segments.

[0063] 8 shows a schematic diagram of a process 800 for contour matching according to an embodiment of the present disclosure. The process 800 may be an example implementation of the operation shown in block 630 of FIG. 6. Since pixels in a contour segment have corresponding spatial offsets in neighboring regions, the average of the spatial offsets of all pixels may be taken as the overall offset, and then the closest contour in the neighboring region may be searched for based on the overall offset.

[0064] In block 810, a projection of the contour segment onto a subsequent frame is determined based on the global offset. Let the current contour segment be Frag_t1, and its global offset in the neighboring region frame be {dx, dy}. Then, we may obtain the projection of the current contour segment onto this subsequent frame according to the global offset, i.e., Frag_t1'=Frag_t1+{dx, dy}.

[0065] In block 820, a number of candidate contour segments in the subsequent frame are determined. Considering the locality of the contour segments, only contour segments in the vicinity of the projection in the subsequent frame may be considered. In some embodiments, a region in the subsequent frame may be determined based on the position of the projection in the subsequent frame. The region may be the largest bounding rectangle of the projection. To facilitate the calculation and improve the operation efficiency, the sides of the rectangle may be aligned along the horizontal and vertical directions of the frame. The contour segments covered by the rectangle are defined as a candidate contour segment set Set_t2.

[0066] In block 830, a contour segment having the smallest distance from the contour segment of interest is determined. For each candidate contour segment Frag_t2 in the candidate contour segment set Set_t2, the distance between the projection Frag_t1' and this candidate contour segment Frag_t2 is determined by determining the minimum distance between each pixel of the projection and the pixel of the candidate contour segment. The candidate contour segment having the smallest determined distance may then be determined as another contour segment matching the contour segment of the first frame. The distance between the projection Frag_t1' and the candidate contour segment Frag_t2 may be calculated by the following equation: TIFF2024150410000004.tif13162 Among them, each pixel p of Frag_t1' i For that, the closest matching pixel p j If the distance exceeds the threshold T, the distance between the two pixels is defined as T; otherwise, the true distance |p i -p j The threshold T is used to ensure the stability of the computation of contour matching between frames.

[0067] Above, a video processing method according to an embodiment of the present disclosure has been described with reference to Figs. 1 to 8. Compared with conventional solutions, the embodiment of the present disclosure can improve the accuracy and reliability of edge detection and enhancement by utilizing spatiotemporal cues in the video, reduce the adverse effects of low-quality images in the video, and improve the robustness of video processing. Some embodiments of the present disclosure further realize edge vectorization and contour tracking based on spatiotemporal consistency, which can realize useful visual assistance and assist doctors in tracking objects of interest in the video.

[0068] 9 shows a schematic block diagram of a video processing apparatus 900 according to an embodiment of the present disclosure. The apparatus 900 may be implemented by the computing device 105, and may be implemented as software, hardware, or a combination of software and hardware.

[0069] As shown, the apparatus 900 includes an edge acquisition unit 910, a spatial offset determination unit 920, and an update unit 930. The edge acquisition unit 910 is configured to acquire a first edge of a first frame of a moving image and a second edge of at least one adjacent region frame of the first frame. The spatial offset determination unit 920 is configured to determine a spatial offset from a pixel of the first edge to a corresponding pixel of the second edge based on a registration between the first frame and the at least one adjacent region frame. The update unit 930 is configured to update a first edge attribute of a pixel of the first edge based on at least a second edge attribute of a corresponding pixel of the second edge and the spatial offset.

[0070] In some embodiments, the edge obtaining unit 910 may be further configured to determine an edge attribute of each pixel in the first frame and the at least one adjacent region frame, the edge attribute indicating a probability that the pixel belongs to an edge, and to determine the first edge and the second edge by determining pixels in the first frame and the at least one adjacent region frame whose edge attribute is greater than a threshold.

[0071] In some embodiments, the update unit 930 may be further configured to determine whether the at least one adjacent region frame satisfies a preset condition; and in response to determining that the at least one adjacent region frame satisfies the preset condition, update a first edge attribute of a pixel of the first edge based on the second edge attribute of a corresponding pixel of the second edge and the spatial offset.

[0072] In some embodiments, the update unit 930 may be further configured to determine whether the at least one adjacent region frame satisfies a preset condition by determining whether an image quality level of the at least one adjacent region frame exceeds a quality threshold and determining whether an average value of spatial offsets from pixels of the first edge to corresponding pixels of the second edge is within a preset range.

[0073] In some embodiments, the apparatus 900 may further include a tracking unit configured to: determine a contour segment of the object of interest in the first frame based on the first edge attribute after updating the pixels of the first edge, determine a global spatial offset from the contour segment to the subsequent frame based on a registration between the first frame and the subsequent frame of the first frame, and search for another contour segment matching the contour segment in the subsequent frame based on the global spatial offset.

[0074] In some embodiments, the tracking unit may be configured to determine a contour segment of the object of interest in the first frame by generating a super-pixel contour of the first frame based on the updated first edge attribute, and generating a contour segment of the object of interest based on the super-pixel contour.

[0075] In some embodiments, the tracking unit may be further configured to generate a contour segment of the object of interest by dividing the super-pixel contour into a set of segments based on intersection points in the super-pixel contour, and generating a contour segment by merging the set of segments.

[0076] In some embodiments, the tracking unit may be further configured to generate the contour segment by determining a similarity between two adjacent segments of the set of segments and, in response to the similarity exceeding a threshold, merging the two adjacent segments to generate the contour segment.

[0077] In some embodiments, the tracking unit may be further configured to determine the similarity between two adjacent segments of the set of segments by any one of the following: a difference between average gradients of pixels of the two adjacent segments, a difference between average edge probabilities of pixels of the two adjacent segments, and a difference between average spatial offsets from pixels of the two adjacent segments to adjacent region frames.

[0078] In some embodiments, the tracking unit may be further configured to search for another contour segment in the subsequent frame that matches the contour segment by obtaining a projection of the contour segment onto the subsequent frame based on the overall spatial offset, determining a number of candidate contour segments in the subsequent frame, and searching for another contour segment in the number of candidate contour segments that is closest in distance to the projection.

[0079] In some embodiments, the tracking unit may be further configured to determine a plurality of candidate contour segments in the subsequent frame by determining a region of the subsequent frame based on a position of the projection in the subsequent frame, and determining a plurality of candidate contour segments of the subsequent frame that are at least partially within the region.

[0080] In some embodiments, the tracking unit may be further configured to search for another contour segment that is closest in distance to the projection by determining, for each of a plurality of candidate contour segments, a distance between the projection and the candidate contour segment by determining a minimum distance between each pixel of the projection and a pixel of the candidate contour segment, and determining the candidate contour segment with the smallest determined distance as the other contour segment.

[0081] In some embodiments, the tracking unit may be further configured to highlight the contour segment and another contour segment for tracking during playback of the video.

[0082] In some embodiments, the first edge and the second edge may each include at least one of an edge of a microstructure and an edge of a lesion.

[0083] In some embodiments, the video may include narrow band imaging (NBI) and magnifying endoscopy (ME) video.

[0084] FIG. 10 shows a schematic block diagram of an exemplary device 1000 according to an embodiment that may be used to implement the teachings of the present disclosure. For example, a computing device 105 according to an embodiment of the present disclosure may be implemented by the device 1000. As shown, the device 1000 includes a central processing unit (CPU) 1001 that can perform various suitable operations and processes according to computer program instructions stored in a read-only memory (ROM) 1002 or loaded from a storage unit 1008 into a random access memory (RAM) 1003. The RAM 1003 may also store various programs and data necessary for the operation of the storage device 1000. The CPU 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.

[0085] Several components of the device 1000 are connected to the I / O interface 1005, including an input unit 1006 such as a keyboard and a mouse, an output unit 1007 such as various displays and speakers, a storage unit 1008 such as a magnetic disk and an optical disk, and a communication unit 1009 such as a network interface card, a modem, a wireless communication transceiver, etc. The communication unit 1009 allows the device 1000 to exchange information and data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0086] Each process and operation described above, such as methods 200, 500, 600, 700 and / or 800, may be executed by the processing unit 1001. For example, in some embodiments, the methods 200, 500, 600, 700 and / or 800 may be implemented as a computer software program tangibly contained in a machine-readable medium, such as the storage unit 1008. In some embodiments, some or all of the computer program may be loaded and / or installed in the device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into the RAM 1003 and executed by the CPU 1001, one or more operations of the methods 200, 500, 600, 700 and / or 800 described above may be executed.

[0087] The present disclosure may be a method, an apparatus, a system, and / or a computer program product, which may include a computer-readable storage medium having computer-readable program instructions for carrying out aspects of the present disclosure.

[0088] A computer readable storage medium may be a tangible device that can hold and store instructions used by an instruction execution device. A computer readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the above. More specific examples (not an exhaustive list) of computer readable storage media include portable computer disks, hard disks, random access memory (RAM), read only memory (ROM), erasable programmable read only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, machine-encoded devices such as punch cards or slot-in-projection structures on which instructions are stored, and any suitable combination of the above. A computer readable storage medium as used herein is not to be construed as a momentary signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagated through a wave guide or other transmission medium (e.g., light pulses through a fiber optic cable), or electrical signals transmitted through electrical wiring.

[0089] The computer readable program instructions described herein may be downloaded from the computer readable storage medium to each computing / processing device or to an external computer or storage device via a network, e.g., the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, fiber optic transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer readable program instructions from the network and transfers the computer readable program instructions for storage in the computer readable storage medium in each computing / processing device.

[0090] The computer program instructions for carrying out the operations of the present disclosure may be assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, or source or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, and traditional process-based programming languages ​​such as "C" or similar programming languages. The computer-readable program instructions may be executed entirely on the user computer, partially on the user computer, as a separate software package, partially on the user computer and partially on a remote computer, or entirely on a remote computer or server. When a remote computer is involved, the remote computer may be connected to the user computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., connected via the Internet using an Internet service provider). In some embodiments, the state information of the computer-readable program instructions is used to implement aspects of the present disclosure by customizing electronic circuitry, e.g., programmable logic circuits, field programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), that can execute the computer-readable program instructions.

[0091] Aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It will be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, may be implemented by computer-readable program instructions.

[0092] These computer-readable program instructions may be provided to a processing unit of a general purpose computer, special purpose computer, or other programmable data processing apparatus such that, when these instructions are executed by a processing unit of the computer or other programmable data processing apparatus, a machine is produced that implements the functions / operations specified in one or more blocks in the flowcharts and / or block diagrams. These computer-readable program instructions may be stored on a computer-readable storage medium, and these instructions cause a computer, programmable data processing apparatus, and / or other device to operate in a particular manner, and thus a computer-readable medium on which instructions are stored includes an article of manufacture containing instructions that implement each aspect of the functions / operations specified in one or more blocks in the flowcharts and / or block diagrams.

[0093] The computer readable program instructions may be loaded into a computer, other programmable data processing apparatus, or other device and cause the computer, other programmable data processing apparatus, or other device to perform a series of operational steps to generate a computer-implemented process, such that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / operations specified in one or more blocks in the flowcharts and / or block diagrams.

[0094] The flowcharts and block diagrams in the drawings illustrate the architecture, functionality, and operation of the systems, methods, and computer program products according to the embodiments of the present disclosure. In this regard, each block in the flowcharts or block diagrams may represent a module, program segment, or part of instructions, including one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, or may be executed in a reverse order depending on the function. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts may be implemented in a hardware-based dedicated system that performs a certain function or operation, or may be implemented in a combination of dedicated hardware and computer instructions.

[0095] Although each embodiment of the present disclosure has been described above, the above description is illustrative, not exhaustive, and not limited to each disclosed embodiment. Many modifications and changes will be obvious to those skilled in the art without departing from the scope and spirit of each described embodiment. The selection of terms used in the present specification is intended to best interpret the principles, practical applications, or improvements to technology in the marketplace of each embodiment, or to enable those skilled in the art to understand each embodiment disclosed in the present specification.

Claims

1. A method of video processing, The process involves obtaining the first edge of the first frame of the aforementioned video and the second edge of at least one adjacent region frame of the first frame, Based on the registration between the first frame and the at least one adjacent region frame, a spatial offset is determined from the pixel of the first edge to the corresponding pixel of the second edge, This includes updating the first edge attribute of the pixel of the first edge based at least on the second edge attribute of the corresponding pixel of the second edge and the spatial offset, method.

2. Acquiring the first edge of the first frame of the aforementioned video and the second edge of the at least one adjacent region frame is: Determining the edge attribute of each pixel in the first frame and the at least one adjacent region frame, wherein the edge attribute indicates the probability that the pixel belongs to an edge. The method includes determining the first edge and the second edge by determining a pixel whose edge attribute is greater than a threshold in the first frame and the at least one adjacent region frame, The method according to claim 1.

3. Updating the first edge attribute of the pixel of the first edge is: Determining whether the at least one adjacent region frame satisfies a predetermined condition, In response to determining that at least one adjacent region frame satisfies a predetermined condition, the first edge attribute of the pixel of the first edge is updated based on the second edge attribute of the corresponding pixel of the second edge and the spatial offset. The method according to claim 1.

4. Determining whether the at least one adjacent region frame satisfies the predetermined conditions is: Determining whether the image quality level of the frame of at least one adjacent region exceeds a quality threshold, The method includes at least one of the following: determining whether the average value of the spatial offset from the pixel of the first edge to the corresponding pixel of the second edge is within a predetermined range. The method according to claim 3.

5. Based on the updated first edge attributes of the pixels of the first edge, the contour segment of the object of interest in the first frame is determined, Based on the registration between the first frame and the subsequent frame of the first frame, the overall spatial offset from the contour segment to the subsequent frame is determined, The further step includes searching for another contour segment in the subsequent frame that matches the contour segment based on the overall spatial offset, The method according to claim 1.

6. Determining the contour segment of the object of interest in the first frame is: Based on the updated first edge attributes, the superpixel contour of the first frame is generated, This includes generating the contour segment of the object of interest based on the superpixel contour, The method according to claim 5.

7. Generating the contour segment of the object of interest is, Based on the intersections in the superpixel contour, the superpixel contour is divided into a set of segments, The process includes generating the contour segment by merging the aforementioned set of segments, The method according to claim 6.

8. Generating the contour segment by merging the aforementioned set of segments is, The process involves determining the similarity between two adjacent segments from the aforementioned set of segments, The process includes, in response to the similarity exceeding a threshold, merging the two adjacent segments to generate the contour segment, The method according to claim 7.

9. Determining the similarity between two adjacent segments of the aforementioned set of segments is: The difference between the average gradients of the pixels of the two adjacent segments, The difference between the mean edge probabilities of pixels in the two adjacent segments, This includes determining one of the following terms: the difference between the average spatial offset from the pixels of the two adjacent segments to the adjacent region frame, The method according to claim 8.

10. Based on the overall spatial offset, searching for another contour segment that matches the contour segment in the subsequent frame is: Based on the overall spatial offset, the projection of the contour segment onto the subsequent frame is obtained, Determining multiple candidate contour segments in the subsequent frame, This includes searching for another contour segment among the plurality of candidate contour segments that is closest in distance to the projection, The method according to claim 7.

11. Determining multiple candidate contour segments in the subsequent frame is: The region of the subsequent frame is determined based on the position of the projection in the subsequent frame, This includes determining the plurality of candidate contour segments that are at least partially located within the region of the subsequent frame, The method according to claim 10.

12. A video processing device, An edge acquisition unit configured to acquire a first edge of a first frame of the aforementioned video and a second edge of at least one adjacent region frame of the first frame, A spatial offset determination unit configured to determine a spatial offset from a pixel of the first edge to a corresponding pixel of the second edge based on the registration between the first frame and the at least one adjacent region frame, The system comprises an update unit configured to update the first edge attribute of the pixel of the first edge based on at least the second edge attribute of the corresponding pixel and the spatial offset, Video processing device.

13. At least one processing unit, The device comprises a memory connected to the at least one processing unit and storing instructions to be executed by the at least one processing unit, wherein when an instruction is executed by the at least one processing unit, the method according to any one of claims 1 to 11 is performed. Electronic devices.

14. Includes a machine-executable instruction, when executed by a device, that causes the device to perform the method according to any one of claims 1 to 11, A computer-readable storage medium.

15. Includes a machine-executable instruction, when executed by a device, that causes the device to perform the method according to any one of claims 1 to 11, Computer program.