A memory bank-based intestinal video target detection method and system

By constructing a memory bank and combining it with optical flow analysis, the problems of high time cost and low accuracy in target detection in colonoscopy videos were solved, achieving more efficient target recognition in colonoscopy videos and improving detection accuracy and robustness.

CN120032107BActive Publication Date: 2026-01-13WUHAN PONY PENTIUM MEDICAL TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510105722.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2026-01-13
Estimated Expiration
2045-01-23

AI Technical Summary

Technical Problem

Existing target detection methods in colonoscopy videos are time-consuming and costly, and their accuracy depends on the experience of the screening personnel. Traditional image or video detection methods have poor accuracy in colonoscopy videos and are difficult to effectively handle image jitter caused by camera movement and interference caused by rapid target movement.

Method used

An intestinal video target detection system based on a memory bank is adopted, including a memory storage module, a memory update module, and an optical flow memory assistance module. By storing features of neighboring image frames in the memory bank, attention mechanisms and optical flow methods are used to enhance and update features, thereby improving detection accuracy.

Benefits of technology

It improves the efficiency and accuracy of target detection in colonoscopy, effectively utilizes the temporal information of video sequences, reduces the influence of irrelevant interference frames, and achieves more efficient target recognition in colonoscopy videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032107B_ABST
    Figure CN120032107B_ABST
Patent Text Reader

Abstract

The application provides a memory bank-based intestinal tract video target detection method and system, and belongs to the technical field of computer vision.The system comprises a memory storage module, which is used for performing feature enhancement on a current image frame by using a memory frame in a memory bank in a target detection stage of the current image frame; a memory updating module, which is used for performing feature similarity and importance analysis on image features of the current image frame and image features of the memory frame in the memory bank, and deciding whether to update the memory bank according to an analysis result; and an optical flow memory auxiliary module, which is used for generating an attention feature map by using an optical flow method to track changes of optical flow key points in a video sequence, and performing image enhancement on the current image frame.The application fully explores time sequence information of a colonoscopy video, enhances main target features of the colonoscopy by using continuity of optical flow, eliminates possible interference in the colonoscopy video, realizes feature enhancement, and improves detection precision.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer vision, and in particular to an intestinal video target detection system based on a memory bank. BACKGROUND

[0002] The unique time sequence information of a colonoscopy video can well reflect the motion trajectory of a foreground target and the change in a colonoscopy screening stage, and is helpful for better developing a colonoscopy target detection.

[0003] A traditional method for target detection based on a colonoscopy video is an artificial screening method, which consumes a large amount of time cost and the accuracy of detection is seriously dependent on the experience of a screener; further, since a colonoscopy video is different from a common fixed camera video, the movement of a camera will cause rapid picture shaking, rapid motion of a target, result in a great difference between adjacent foreground features, and introduce a large number of irrelevant interference frames, increasing the difficulty of applying a traditional image or video detection method to the field of colonoscopy videos. Existing intestinal target methods based on deep learning design are only designed at the image level, or simply interact with adjacent frames at the video level for feature enhancement, resulting in poor target recognition accuracy.

[0004] Therefore, there is an urgent need for a new technical solution for intestinal target detection based on a colonoscopy video. SUMMARY

[0005] The present application provides an intestinal video target detection method and system based on a memory bank to solve at least one defect in the prior art.

[0006] In a first aspect, the present application provides an intestinal video target detection system based on a memory bank, comprising: a memory storage module, a memory updating module, and an optical flow memory auxiliary module.

[0007] The memory storage module is configured to utilize a memory frame in a memory bank to enhance the features of a current image frame at a target detection stage of the current image frame; the image features of a plurality of memory frames stored in the memory bank are image features of a plurality of adjacent image frames before the current image frame.

[0008] The memory updating module is configured to analyze the feature similarity and importance of the image features of the current image frame and the image features of the memory frames in the memory bank, and determine whether to update the memory bank according to the analysis result.

[0009] The optical flow memory auxiliary module is configured to generate an attention feature map by tracking the change of an optical flow key point in a video sequence using an optical flow method, and to enhance the image of the current image frame.

[0010] According to the memory bank-based intestinal tract video target detection system provided by the application, the memory bank storage module is specifically used for storing the image features of the previous fixed number of memory frames adjacent to the current image frame to be detected into the memory bank, and during the detection process of the current image frame, the image features of the memory frames in the memory bank are used to enhance the features of the current image frame by using the attention mechanism.

[0011] According to the memory bank-based intestinal tract video target detection system provided by the application, the memory bank storage module is specifically used for storing the image features of the previous fixed number of memory frames adjacent to the current image frame to be detected into the memory bank, and during the detection process of the current image frame, the image features of the memory frames in the memory bank are used to enhance the features of the current image frame by using the attention mechanism.

[0012] According to the memory bank-based intestinal tract video target detection system provided by the application, the memory bank storage module is specifically used for storing the image features of the previous fixed number of memory frames adjacent to the current image frame to be detected into the memory bank, and during the detection process of the current image frame, the image features of the memory frames in the memory bank are used to enhance the features of the current image frame by using the attention mechanism.

[0013] According to the memory bank-based intestinal tract video target detection system provided by the application, the memory bank storage module is specifically used for storing the image features of the previous fixed number of memory frames adjacent to the current image frame to be detected into the memory bank, and during the detection process of the current image frame, the image features of the memory frames in the memory bank are used to enhance the features of the current image frame by using the attention mechanism.

[0014] According to the memory bank-based intestinal tract video target detection system provided by the application, the memory bank storage module is specifically used for storing the image features of the previous fixed number of memory frames adjacent to the current image frame to be detected into the memory bank, and during the detection process of the current image frame, the image features of the memory frames in the memory bank are used to enhance the features of the current image frame by using the attention mechanism.

[0015] According to the memory bank-based intestinal tract video target detection system provided by the application, the optical flow memory auxiliary module is specifically used for: determining whether there is a target at the end of target detection of a previous frame, extracting optical flow key points if there is a target, performing optical flow tracking on the optical flow key points to obtain optical flow vectors, analyzing the speed and angle of the optical flow vectors to generate an attention feature map, and performing image enhancement on a current image frame by using the attention feature map.

[0016] According to the memory bank-based intestinal tract video target detection system provided by the application, the video sequence is a colonoscopy video sequence.

[0017] In a second aspect, the application further provides a memory bank-based intestinal tract video target detection method, which applies the intestinal tract video target detection system as described above, and comprises the following steps:

[0018] The optical flow memory auxiliary module is used for performing image enhancement on a current image frame;

[0019] The memory bank is used for performing feature enhancement on the current image frame by using the memory frames in the memory bank in a target detection stage of the current image frame; and the memory bank stores image features of the memory frames adjacent to the current image frame.

[0020] The memory bank is used for performing similarity and importance analysis on the image features of the current image frame and the memory frames in the memory bank, and determining whether to update the memory bank according to the analysis result.

[0021] In a third aspect, the application provides an electronic device, which comprises a memory, a processor, and a computer program stored in the memory and capable of running on the processor, and the processor implements the steps of the memory bank-based intestinal tract video target detection method as described above when executing the program.

[0022] The memory bank-based intestinal tract video target detection method and system provided by the application have the following beneficial effects compared with the prior art:

[0023] (1) The application designs a complete colonoscopy video detection process, provides a complete system of memory feature (image feature) encoding, memory bank updating, memory bank interaction with a current frame, and optical flow motion statistical analysis, and improves the detection efficiency and accuracy of the colonoscopy target.

[0024] (2) The application constructs a memory bank that can well store the time sequence and spatial information of image frames at past time points, and uses the detection results of previous frames to perform cross-attention calculation on the features of the current frame and the features of the image frames in the memory bank to learn the time sequence information of the video sequence when detecting the next frame.

[0025] (3) The application fully considers that the memory bank needs to have representative problems, so when updating the memory bank, the traditional first-in first-out principle is not adopted, but a set of evaluation criteria for memory bank updating is specially designed to evaluate the importance and current real specificity of the existing image frames in the memory bank, to determine whether the current frame needs to enter the memory bank and whether the existing image frames in the memory bank need to be removed.

[0026] (4) The application considers that the video is dynamic, the target has a certain motion trend with the progression of the video image frame, introduces light flow memory analysis auxiliary for motion modeling of the target with the progression of the video, learns the motion condition of the target through statistical analysis of the speed and direction of the light flow vector, and generates an attention map through the light flow analysis result to enhance the image features. BRIEF DESCRIPTION OF DRAWINGS

[0027] In order to more clearly illustrate the technical solutions in the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0028] Figure 1 is a schematic diagram of the framework for target detection under colonoscopy video by constructing a memory bank in the embodiments of the present application;

[0029] Figure 2 is a schematic diagram of the framework for target detection under colonoscopy video by constructing a memory bank in the embodiments of the present application;

[0030] Figure 3 is a schematic diagram of the framework for target detection under colonoscopy video by constructing a memory bank in the embodiments of the present application;

[0031] Figure 4 is a schematic diagram of the framework for target detection under colonoscopy video by constructing a memory bank in the embodiments of the present application;

[0032] Figure 5 is a schematic diagram of the framework for target detection under colonoscopy video by constructing a memory bank in the embodiments of the present application;

[0033] Figure 6 is a schematic diagram of the framework for target detection under colonoscopy video by constructing a memory bank in the embodiments of the present application; DETAILED DESCRIPTION

[0034] In order to make the objects, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below in conjunction with the drawings in the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0035] It should be noted that, in the description of the embodiments of the present application, the terms "comprise", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitation, the element defined by the sentence "comprises a" does not exclude the presence of another identical element in the process, method, article or device comprising the element. The specific meaning of the above terms in the present application can be understood according to the specific circumstances by those of ordinary skill in the art.

[0036] The present application provides a memory bank-based colonoscopy video target detection method and system, which can well utilize the time sequence characteristics of colonoscopy video and make full use of the information of representative image frames before the current frame to assist the detection of the current frame, thereby fully exploiting motion and time sequence information, avoiding interference and improving detection performance.

[0037] Figure 1 is a schematic diagram of the framework for target detection in colonoscopy video by constructing a memory bank in the embodiments of the present application, Figure 2 is a schematic diagram of the flow of target detection in colonoscopy video by constructing a memory bank in the embodiments of the present application, referring to Figure 1 and Figure 2 The present application provides a complete technical solution for memory bank construction, memory bank interaction, memory bank updating and video detection processing guided by memory based on optical flow memory assistance:

[0038] (1) Memory storage module: used for storing the image features of a fixed number of memory frames adjacent to the current image frame to be detected in the memory bank, and based on the image features of the memory frames in the memory bank, the current image frame is enhanced in feature by using the attention mechanism during the detection process of the current image frame; the image features of the plurality of memory frames stored in the memory bank are the image features of the plurality of image frames adjacent to the current image frame.

[0039] To achieve memory bank-based video sequence modeling and enhance the image features of the current frame through cross-attention mechanism, a memory bank needs to be first constructed to store the image features of past images in the video. The design aims to store the image features of the memory frames into the memory bank, and to enhance the detection effect of the current image by retrieving and updating the image features in the memory bank. In this way, not only the information of the current frame is relied on, but also the features of past image frames are utilized to improve the accuracy and robustness of image understanding.

[0040] In this example, each image frame of the video is input into a pre-trained feature encoder network to extract high-dimensional feature vectors. These high-dimensional spatial features are enhanced through memory attention and optical flow auxiliary analysis with the existing information features in the memory bank, and then passed through a decoder network. The output of the decoder is processed to generate detection results and is also passed through a memory encoder network to re-encode the features of the current frame containing the current frame detection information and store them in the memory bank for subsequent frame detection as a memory attention prompt for subsequent frame detection.

[0041] Regarding the process of using the existing information in the memory bank to prompt the current frame, specifically, cross-attention analysis is performed between the image features of the current frame after encoding and the spatial features of the past image frames stored in the memory bank. Each time a new image is processed, the features of the current image are extracted by the previous encoder, the features of the memory image frames stored in the memory bank are generated by the memory encoder during the construction of the memory bank, and the cross-attention analysis operation is performed between the current image frame and each image frame stored in the memory bank. The features of the current frame are used as queries, while the features in the memory bank are used as keys and values. Through the cross-attention mechanism, the attention score is calculated, which introduces the effective spatio-temporal information in the memory frames with significant reference significance into the current frame detection to update the features of the current image using the memory frame information.

[0042] In the detection process of the current image frame, the image features of the memory frames in the memory bank are used to enhance the features of the current image frame based on the attention mechanism, including:

[0043] A query is generated based on the image features of the current image frame, and a corresponding key-value pair is generated based on the image features of the memory frames in the memory bank. Cross-attention is calculated for the query of the current image frame and the key-value of each memory frame in the memory bank to obtain the attention score corresponding to each memory frame. Based on the attention score, the image features of the memory frames in the memory bank are used to enhance the features of the current image frame to generate enhanced features.

[0044] The specific implementation is as follows:

[0045] First, the image features of the current frame are extracted from the image features of the current frame Generate query For each frame feature map in the memory bank, a corresponding key-value pair is obtained using a feature extractor And Then, the query of the current frame obtained is matched with the key of each frame in the history bank (i.e., the memory bank) to obtain an attention score, and the specific calculation method is as follows:

[0046]

[0047] Wherein, t represents the image frame at time t that needs to be detected at present, n represents the nth memory frame in the memory bank, The dimension of . The attention score of the cross-attention interaction between the image at time t and the nth memory frame.

[0048] The attention score is normalized by softmax to obtain the attention weight of a frame. The current frame feature is weighted using the weight, so that the feature of the nth memory frame is optimized for the current frame feature. Specifically, the normalized attention weighting score is applied to the value in the history bank (memory bank) :

[0049]

[0050] The weighted sum obtains the final output, which represents the feature of the current frame and has combined the information of the memory frame. In actual calculation according to the correlation, the current frame and each frame in the memory bank need to be interacted, so the key-value pairs in the memory bank can be spliced together to form a matrix for unified attention calculation:

[0051]

[0052]

[0053] Replace k and v in the above formulas 1 and 2 with K and V, then the memory attention calculation for all image frames in the memory bank can be realized, and the past frame features are fused by weighting the current frame using the cross-attention result to obtain the final enhanced features , N The number of memory frames in the memory bank.

[0054] (2) Memory updating module: used for analyzing the feature similarity and importance of the image features of the current image frame and the image features of the memory frames in the memory bank, and deciding whether to update the memory bank according to the analysis result.

[0055] The memory update mainly occurs before a frame detection enters the next frame detection, because the detection result of the current frame is more adjacent to the next frame detection, which can theoretically provide more references. The existing memory frame in the memory bank has existed for a long time, which may have lost its representativeness due to the distance being too far in the subsequent frame detection. At the same time, due to the existence of irrelevant redundant frames, the image frames already stored in the memory bank may not be representative and need to be removed immediately. Therefore, the importance of the memory frame needs to be investigated and the unimportant memory frame needs to be removed.

[0056] Because the main reason for updating the memory bank is actually due to the length of the memory bank being limited by the memory, the interaction speed and efficiency of the feature interaction using the memory image frame features in the memory bank are limited by the computing medium. Therefore, the number of image frames that can be stored in the memory bank is limited. On the other hand, the colonoscopy video is a long video sequence, and a complete colonoscopy video sequence is often divided into tens of thousands of images. The colonoscopy video goes through multiple stages such as entering the anus, anal canal, intestinal tract, reaching the terminal ileum, and exiting. The intestinal images observed by the colonoscopy in different stages have obvious differences. Therefore, it is not necessary to use images with a long distance to interact with the current frame, and only the temporal characteristics of the neighborhood video segment need to be considered. Therefore, it is necessary to limit the size of the memory bank and update it.

[0057] The update of the memory bank includes two aspects: one is to evaluate the importance of the memory image frame already existing in the memory bank. The importance can be directly reflected by the contribution value of the memory frame in the current frame detection. The contribution value can be directly reflected by the similarity score matrix of the attention interaction between the memory bank and the current frame. In addition, generally speaking, the longer the distance between the current frame and the image frame, the more limited the reference significance of the current image frame. At the same time, if the target disappears and appears between two frames, it means that the two adjacent frames reflect an important stage, and the corresponding adjacent frames need to be more cautious to delete.

[0058] Specifically, the memory update module is configured to: calculate the cosine similarity according to the image features of the current image frame and the memory frames in the memory bank; obtain an affinity matrix based on all the cosine similarities; and analyze the affinity matrix to update the memory bank.

[0059] The memory updating module analyzes the affinity matrix to update the memory bank, including: determining the maximum value in the affinity matrix; in the case that the maximum value is greater than or equal to a preset threshold, performing feature fusion on the image features of the current image frame and the memory frame corresponding to the maximum value to generate enhanced image features of the memory frame, so as to update the memory bank; in the case that the maximum value is less than the preset threshold, comprehensively evaluating the importance of each memory frame in the memory bank to the current image frame to obtain a comprehensive evaluation score; replacing the image features of the memory frame with the lowest comprehensive evaluation score with the image features of the current image frame to update the memory bank.

[0060] Figure 3 is a flowchart of updating the memory bank provided by the application, referring to Figure 3 , the memory updating module of the application will be further described below Figure 3 .

[0061] First, the detection result of the current frame is processed by the memory encoder to obtain the memory features of the current frame (image features of the current frame), the memory features of the current frame and the memory features of the memory frame (image features of the memory frame) are respectively calculated for cosine similarity .

[0062]

[0063] The cosine similarities calculated between the current frame and all frames in the memory bank are spliced to obtain an affinity matrix . reflects the similarity degree between the memory features of the current frame and the image features already existing in the memory bank. The extreme value of the affinity matrix is analyzed, if the maximum value a exceeds 0.7 (a preset threshold), it indicates that the current frame and the memory frame already existing in the memory bank are highly similar, and the features of the current frame are directly fused with the corresponding frame to obtain enhanced memory features .

[0064]

[0065] If it is not satisfied, it indicates that the memory features of the current frame do not exist in the memory features in the existing memory bank, and therefore it is necessary to evaluate the importance of the memory frames in the memory bank to determine the comprehensive evaluation score, so as to eliminate the least important (i.e. the memory frame with the lowest comprehensive evaluation score), and update the memory features of the current frame to the memory bank.

[0066] ​Wherein, for the memory frame already existing in the memory bank, the method for determining the comprehensive evaluation score comprises: determining the comprehensive evaluation score according to the contribution value of the memory frame to the current frame detection, the existing duration of the memory frame and whether the memory frame has target mutation.

[0067] Specifically, the contribution value of the memory frame can be directly based on the attention score between the memory frame and the current image frame It is determined that the existing duration of the memory frame is the difference between the frame number of the memory frame in the entire video and the frame number of the current image frame It is reflected that whether the memory frame has target mutation is analyzed and constructed through the detection result of each frame in the history , The value is determined by the detection result of the current frame and the previous frame, if the previous frame has target and the next frame does not have target or the previous frame does not have target and the current frame has target, then 1 is stored in the corresponding position of the existing situation identification matrix of the memory feature corresponding to the current frame, otherwise 0.

[0068] Finally, the comprehensive evaluation score of the existing memory frame in the memory bank is obtained by comprehensively considering the three indexes:

[0069] , (6)

[0070] Wherein is the weight value of each index, which can be 0.4, 0.4 and 0.2 respectively. The memory bank can be updated by removing the memory frame with the lowest score and replacing it with the memory feature of the current frame.

[0071] (3) Optical flow memory auxiliary module: the optical flow method is used to track the change of the optical flow key point in the video sequence to generate an attention feature map, and the current image frame is image enhanced.

[0072] At the end of the previous frame target detection, it is determined whether there is a target, if there is a target, the optical flow key point is extracted, the optical flow tracking is performed on the optical flow key point, the optical flow vector is obtained, the speed and angle of the optical flow vector are analyzed to generate an attention feature map, and the current image frame is image enhanced by using the attention feature map.

[0073] There are a large number of irrelevant image frames in the detection process of the colonoscopy video, and since the camera is movable, the images between adjacent frames are generally similar, but there is a small displacement due to the movement of the camera. If the deviation of the detection foreground caused by this small movement can be well captured, it will be of great help to the subsequent detection. The continuity trend of the movement can be learned to adjust the attention of the detection network on the whole picture, so that the previous detection result can give a guide to the more attention area of the subsequent detection.

[0074] Specifically, when the output detects the target, the LK algorithm is used to extract the LK optical flow inside the detection frame from the current frame to track the necessary optical flow key points, and the results of the optical flow key points extracted by the optical flow operation of this frame are stored. From the next frame, the output of the key points extracted in the last frame detection result is used to track the motion of the optical flow key points of the last frame in this frame using the LK optical flow method. The optical flow key points of the last frame and the optical flow key points of the current frame are associated to obtain a series of optical flow vectors, and the changes in speed and direction of the optical flow vectors are counted to analyze the motion trend of the target of the last frame in the current frame.

[0075] There are mainly two situations that need to be distinguished and focused on for the motion of the target between image frames: one is that the target moves and the position changes, but still exists; the other is that the target disappears in this frame, and the video enters a pure background picture or a shaking, water flow, etc. irrelevant interference frame picture, which shows differences in optical flow statistics and differences in subsequent use of optical flow information. In order to better use the information obtained by optical flow for analysis, it is necessary to analyze these two situations first.

[0076] Through experimental observation and analysis of the principle of the optical flow method, optical flow describes the motion change of pixels in two consecutive frames of an image, reflecting the displacement of each pixel or key point in time. The direction of optical flow tells us the moving direction of the target object in the image. If the optical flow directions of multiple key points are consistent, it means that these points may belong to the same object or the same motion direction; the size of optical flow reflects the motion speed of the target between two frames. Fast-moving targets will have larger optical flow, while stationary or slow-moving targets will have smaller optical flow. If the direction and speed of optical flow are consistent in the image (for example, all objects in the scene move in a certain direction), it can be considered that these movements are caused by a unified target, and if the optical flow changes are chaotic, it may be due to background changes or scenes without targets. If the change of optical flow in the image appears to be sudden (for example, the direction of optical flow changes very irregularly and the speed increases sharply), it may be due to rapid movement in the scene (such as camera movement, objects changing direction quickly, etc.). This sudden change can help detect abnormal situations in the image and handle them. Therefore, the speed and direction of all node optical flows are analyzed.

[0077] Since there is a big difference between the speed and direction of the optical flow if the target image frame enters the frame without a target, the mean value of the angle change of the direction of all the optical flows and the mean value of the speed change are calculated, and the standard deviation is calculated to reflect the floating, and a threshold is set as the standard for judgment; if there is a same target between the two image frames, the speed and direction change of the optical flow should show consistency with the movement of the target, and the mean value and standard deviation of the speed and direction of the optical flow are also calculated for judgment.

[0078] The Gaussian attention map is a weighted map generated by a Gaussian distribution function. The feature of Gaussian distribution is that its weight is highest at the center position and gradually decreases with the increase of the distance from the center. The basic idea of using Gaussian attention map to weight image features is to calculate a weighting coefficient according to the importance of the features at each position in the image, and then apply the weighting coefficient to the corresponding image features. In this way, the model can focus on the most valuable areas in the image and weaken the unimportant areas, thereby improving the performance. Using the results of the speed and direction of the optical flow trajectory statistics to generate a Gaussian center and a radius to obtain a suitable Gaussian distribution to generate an attention map to weight the image features can make the motion trend learned by the optical flow reflected on the attention degree of the foreground image features in different areas of the picture.

[0079] In specific implementation, it is also divided into two cases: the current frame has a target and the target disappears in the current frame:

[0080] For the case of target disappearance, it is explained that the features of the current frame usually do not need to pay attention to fine details, and a module similar to low-pass filtering is designed through the Gaussian attention map to reduce the complexity of local texture and make the feature map smoother. Through the uniform weighting or center attenuation of the Gaussian map, the feature map of the background area can show continuous smooth characteristics, avoiding excessive detail processing; at the same time, the feature map after the Gaussian map weighting will become simpler and more regular, which is helpful for the subsequent network to reduce the computational complexity.

[0081] For the case where the target still exists, first, the geometric center of the key nodes in the current frame is calculated to obtain the center of gravity of the key points, i.e. the center point of the Gaussian distribution. In fact, if the tracking results of the key points are relatively concentrated, the center point will be close to the actual position of the target; if the key points are scattered, the center point tends to the geometric center of the overall position. The standard deviation of the Gaussian distribution determines the smoothing radius of the Gaussian distribution. Through the change of the speed and direction of the key points, the standard deviation of the Gaussian distribution can be dynamically adjusted to adapt to the situation of target concentration or dispersion. Specifically, the smaller the change of the speed and direction of the optical flow vector, the more concentrated the motion of the key points, and the Gaussian distribution should be more focused; when it is larger, the Gaussian distribution should be smoother. The generated Gaussian attention map is multiplied with the original feature map pixel by pixel to enhance the region of interest and weaken the background.

[0082] Figure 4 is a flowchart of the optical flow memory auxiliary process provided by the present application, which is described below with reference to Figure 4 The specific implementation process is described as follows:

[0083] First, the extraction of the optical flow information of the previous frame occurs at the end of the detection of the previous frame. The decoding result of the decoder is analyzed. If there is a target, the optical flow key points are extracted. If the detection result indicates that there is no target, the optical flow key points are not extracted.

[0084] Before the current frame enters the encoder to extract features, the storage part of the optical flow memory auxiliary module is accessed to confirm whether there are key points extracted in the previous frame. If there are, the key points of the previous frame are tracked by the LK optical flow method in the current frame.

[0085] The optical flow key points of the previous frame and the tracking results of the current frame are matched and a set of optical flow vectors is obtained. The speed of the optical flow vectors, i.e., the Euclidean distance between two points, and the direction, i.e., the angle, are calculated. It is specified that the upper right is 0°, and the direction is from the point of the previous frame to the point of the current frame. The standard deviation and mean value of all speeds and directions are calculated.

[0086]

[0087] is each value in the data, is the mean value of the data ,n is the number of data. The data here can be speed and angle values. For the mean value and variance of the speed and direction, a threshold is first set to distinguish the tracking situation of the current frame, and the current frame is analyzed according to the tracking situation to determine whether there is a target. Through experimental analysis, when the mean value of the speed is greater than half of the smaller one of the length and width of the image and the mean value of the angle is greater than 180°, it is considered that there is no target, and the current frame is likely to be an irrelevant frame, which needs to be further analyzed to estimate the target situation of the current frame. Specifically, if the standard deviation value is less than 0.3 of the mean value, it is more inclined to consider that the current frame has a target, and vice versa.

[0088] According to the above analysis of the existence of the target in the current frame, the Gaussian attention map is generated in two cases: the current frame has no target and the current frame has a target.

[0089] If the current frame has no target, the center position of the Gaussian function is considered to be the geometric center of the image of the current frame, and the standard deviation value of the Gaussian distribution is selected to be a large value to ensure that the weight distribution is smoother. The two-dimensional weight value of the Gaussian of the background image without a target is calculated as follows:

[0090]

[0091] where is the position of the target in the image, is the center of the Gaussian distribution in the current frame where there is no target is replaced by the geometric center, Here we take 200, which intends to give a larger standard deviation to smooth the background image details without target.

[0092] If the current frame is considered to have a target, the calculated mean and standard deviation are used to generate the parameters of the Gaussian attention map:

[0093] First, the center point of the Gaussian distribution should be the cluster center of all current positions tracked by optical flow, that is, the geometric center point of all points:

[0094]

[0095] where, is the weight value of each coordinate position, which is calculated by the speed and direction change amount of the optical flow vector and the change mean, the weight value of the first i optical flow vector is calculated as follows:

[0096]

[0097] where and are the mean values of the speed and angle change amount, respectively.

[0098] The standard deviation value of the Gaussian distribution is obtained by the standard deviation values of the speed and direction, in order to comprehensively consider the changes of speed and direction, it is calculated by the following formula:

[0099]

[0100] where is the reference value of the smoothing radius, which is taken as 50 in the experiment, and are the standard deviations of the speed and angle change, which are normalized by max( , ) and 180°, respectively, and w and h in max( , ) are the width and height of the image size.

[0101] The obtained Gaussian image feature map of optical flow attention G(x, y) and the original image I(x, y) are multiplied element by element, so as to weight the original image by using the optical flow information:

[0102]

[0103] Figure 5is a schematic diagram of image enhancement using optical flow vector according to the present application, as shown Figure 5 After Gaussian attention weighting, the image can pay more attention to the image area near the target in the presence of the target, and can smooth the image features in the absence of the target, thereby reducing noise or irrelevant information interference and improving detection efficiency.

[0104] The method based on optical flow memory is only used when a target is detected in the previous frame, in order to further provide better features of the current frame for subsequent memory interaction.

[0105] The following table is a comparison table of experimental results before and after adding each module:

[0106] Comparison table of experimental results

[0107]

[0108] Among them:

[0109] Flow-A represents the use of optical flow memory auxiliary enhancement method.

[0110] Mem-A represents constructing a memory bank and using a memory attention mechanism for enhancement.

[0111] Mem-UP represents updating the memory bank using the updating method designed in the present application, otherwise using the FIFO first-in first-out method for updating.

[0112] P, R, and Error represent the accuracy, precision, and false detection rate of detection, respectively.

[0113] The present application also provides an intestinal video target detection method based on a memory bank, which applies any of the intestinal video target detection systems described above, and the method comprises:

[0114] Using an optical flow memory auxiliary module to enhance the current image frame;

[0115] Using a memory storage module to enhance the features of the current image frame using the memory frames in the memory bank during the target detection stage of the current image frame; the memory bank stores the image features of the adjacent memory frames before the current image frame;

[0116] Using a memory updating module to analyze the similarity and importance of the image features of the current image frame and the memory frames in the memory bank, and deciding whether to update the memory bank according to the analysis results.

[0117] The intestinal video target detection method based on the memory bank provided by the present application can also apply the intestinal video target detection system based on the memory bank described in any of the above embodiments, which will not be repeated here.

[0118] The memory bank-based intestinal video target detection method and system provided by the application have the following beneficial effects compared with the prior art:

[0119] (1) The application designs a complete colonoscopy video detection process, provides a complete system for memory feature coding construction, memory bank updating, memory bank and current frame interaction, and optical flow motion statistical analysis, and improves the detection efficiency and accuracy of the colonoscopy target.

[0120] (2) The application constructs a memory bank that can well store the time sequence and spatial information of the image frames at the past time points, and uses the detection results of the previous frames to calculate the cross attention of the features of the current frame and the features of the existing image frames in the memory bank, so as to learn the time sequence information of the video sequence.

[0121] (3) The application fully considers that the memory bank needs to have representative problems, so when updating the memory bank, the traditional first-in-first-out principle is not used, but a set of evaluation criteria for memory bank updating is specially designed to evaluate the importance of the existing image frames in the memory bank and the current specificity, to determine whether the current frame needs to enter the memory bank and whether the existing image frames in the memory bank need to be removed.

[0122] (4) The application considers that the video is dynamic, the target has a certain motion trend with the progress of the video image frames, introduces optical flow memory analysis to assist the motion modeling of the target with the progress of the video, learns the motion of the target through statistical analysis of the speed and direction of the optical flow vector, and generates an attention map through the optical flow analysis result to enhance the image features.

[0123] Figure 6 is a structural schematic diagram of an electronic device provided by the application, as Figure 6 shown, the electronic device can include a processor 610, a communications interface 620, a memory 630 and a communications bus 640, wherein the processor 610, the communications interface 620 and the memory 630 complete mutual communication through the communications bus 640. The processor 610 can call the logical instructions in the memory 630 to execute the memory bank-based intestinal video target detection method.

[0124] On the other hand, the application also provides a computer program product, which includes a computer program stored on a non-transitory computer readable storage medium, and the computer program includes program instructions, when the program instructions are executed by a computer, the computer can execute the memory bank-based intestinal video target detection method provided by each of the embodiments.

[0125] In yet another aspect, the present application also provides a non-transitory computer readable storage medium having stored thereon a computer program, which, when executed by a processor, implements the memory bank based intestinal video target detection method provided by each of the above embodiments.

[0126] Those skilled in the art can clearly understand the implementation of the embodiments by means of software and the necessary general hardware platform, and of course, the embodiments can also be implemented by hardware. Based on such understanding, the above technical solutions, essentially or in the sense of contribution to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, or an optical disc, and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each of the embodiments or some parts of the embodiments.

[0127] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for some technical features thereof; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A memory bank based intestinal video target detection system, characterized in that, The method comprises the following steps: The memory storage module, the memory update module, and the optical flow memory auxiliary module; The memory storage module is configured to perform feature enhancement on the current image frame by using the memory frames in the memory bank during a target detection stage of the current image frame; The image features of the plurality of memory frames stored in the memory bank are image features of a plurality of adjacent image frames before the current image frame; The memory update module is configured to perform feature similarity and importance analysis on the image features of the current image frame and the memory frames in the memory bank, and determine whether to update the memory bank according to the analysis result; The optical flow memory auxiliary module is configured to generate an attention feature map by using an optical flow method to track changes in optical flow of key points in the video sequence, and perform image enhancement on the current image frame; The memory update module is specifically configured to: Calculate a cosine similarity according to the image features of the current image frame and the memory frames in the memory bank; Obtain an affinity matrix based on all the cosine similarities; Analyze the affinity matrix to update the memory bank; The memory update module analyzes the affinity matrix to update the memory bank, which comprises: Determine the maximum value in the affinity matrix; In the case where the maximum value is greater than or equal to a preset threshold, perform feature fusion on the image features of the current image frame and the memory frame corresponding to the maximum value to generate enhanced image features of the memory frame, so as to update the memory bank; In the case where the maximum value is less than the preset threshold, comprehensively evaluate the importance of each memory frame in the memory bank to the current image frame to obtain a comprehensive evaluation score; Replace the image features of the memory frame with the lowest comprehensive evaluation score with the image features of the current image frame to update the memory bank.

2. The memory bank based enteroscopy video target detection system of claim 1, wherein, The memory storage module is specifically configured to: Store the image features of a fixed number of memory frames adjacent to the current image frame to be detected in the memory bank, and perform feature enhancement on the current image frame by using an attention mechanism based on the image features of the memory frames in the memory bank during the detection process of the current image frame.

3. The memory bank-based enteroscopy video target detection system of claim 2, wherein, The memory storage module performs feature enhancement on the current image frame by using an attention mechanism based on the image features of the memory frames in the memory bank, which comprises: Generate a query according to the image features of the current image frame, and generate a corresponding key-value pair according to the image features of the memory frames in the memory bank; Perform cross-attention calculation on the query of the current image frame and the key-value of each memory frame in the memory bank to obtain an attention score corresponding to each memory frame; Perform feature enhancement on the current image frame by using the image features of the memory frames in the memory bank based on the attention score to generate enhanced features.

4. The memory bank-based enteroscopy video target detection system of claim 1, wherein, The memory update module comprehensively evaluates the importance of each memory frame in the memory bank to the current image frame to obtain a comprehensive evaluation score, which comprises: Determine the comprehensive evaluation score according to the contribution value of the memory frame to the detection of the current frame, the existence duration of the memory frame, and whether the memory frame has a target mutation; The contribution value of the memory frame to the detection of the current image frame is determined according to the attention score between the memory frame and the current image frame.

5. The memory bank-based enteroscopy video target detection system of claim 1, wherein, The optical flow memory auxiliary module is specifically configured to: At the end of the target detection of the last frame, it is determined whether there is a target, and if there is a target, the optical flow key points are extracted and tracked to obtain optical flow vectors; The speed and angle of the optical flow vectors are analyzed to generate an attention feature map; The current image frame is enhanced using the attention feature map.

6. The memory bank-based enteroscopy video target detection system of claim 1, wherein, The video sequence is a colonoscopy video sequence.

7. A memory bank based intestinal video target detection method, applying the intestinal video target detection system according to any one of claims 1 to 6, characterized in that, The method comprises: The current image frame is enhanced using the optical flow memory auxiliary module; In the target detection stage of the current image frame, the memory frames in the memory bank are used to enhance the features of the current image frame; the image features of the adjacent memory frames before the current image frame are stored in the memory bank; The memory bank is updated according to the similarity and importance analysis of the image features of the current image frame and the memory frames in the memory bank.

8. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the computer program to realize the steps of the memory bank-based colonoscopy video target detection method according to claim 7.

Citation Information

Patent Citations

  • Lightweight video object segmentation method based on big data memory storage

    CN114882076A

  • Real-time motion detection method based on multi-scale feature fusion attention

    CN115131710A

  • Target tracking method based on space-time interaction attention mechanism

    CN116563355A